Take-Home Assignment Grading Rubric
Per-submission rubric scorecard: scores per dimension, evidence per score, second-grader calibration, final decision, candidate feedback. ATS updated with full grading record.
Before you start
- Take-home assignment with documented evaluation criteria
- Calibrated graders (2+ per submission to reduce bias)
- Approved time-cap (typically 3-4 hours of candidate work)
- Candidate submission
- Evaluation rubric (skills assessed, scoring scale)
- Reference solutions or scoring exemplars
The steps
- Verify submission completeness and integrity — Check: submission within time cap, all required components present, no obvious LLM hallucinations or copy-pastes from public sources. Flag anything anomalous for human review before grading.
- Apply rubric scoring per dimension — Score each rubric dimension independently (e.g., correctness, code quality, testing, communication, technical depth). Each dimension on a 1-5 scale with documented anchor descriptions. Avoid composite scores — track each dimension to identify strengths and weaknesses.
- Capture specific evidence per score — For each score, note 1-2 specific code/answer references. 'Correctness: 4/5 — handled edge cases for empty input but missed Unicode case' beats 'Correctness: 4/5'. Evidence makes scoring auditable.
- Run a second-grader review for calibration — Send to a second grader (different from the first) for independent scoring. Compare scores per dimension. If scores differ by >2 points on any dimension, schedule a 15-minute calibration discussion before final score.
- Synthesize feedback for the candidate — Whether the candidate advances or not, write 2-3 sentences of constructive feedback they could use. Focus on patterns, not nitpicks. Candidates have invested 3-4 hours — feedback is the courtesy.
- Log final outcome and route — Log: per-dimension scores, evidence, calibration result, final advance/reject decision, feedback for candidate. Route advance decisions to next interview stage; route reject decisions to rejection email workflow with feedback included.
If it goes wrong
Graders apply different mental models to the rubric
Quarterly calibration sessions: graders independently score the same exemplar submission, compare and discuss. Re-anchor the rubric definitions if drift detected.
Submissions show patterns of LLM use
Don't ban LLM use outright — use of tools is a real-world skill. But require an in-person follow-up conversation to verify the candidate understands their own submission. Mismatch = automatic disqualification.
Strong candidates declining the take-home as too much
Re-evaluate the time cap. 4 hours is the upper limit for senior roles. If declines spike, the assignment is too long — shorten it without losing signal.
All OpenLabor playbooks