Code reasoning evaluation program
Rubric design and expert-graded evaluation of model outputs across software engineering tasks, with reproducible agreement metrics.
Industry
AI lab, remote
Year
2026
Stack
- Argilla
- Python
- DuckDB
- Weights & Biases
01Problem
A model team needed dependable human judgment on multi-step coding tasks, but ad-hoc grading produced noisy, unrepeatable signals.
02Approach
- 01
Decomposed tasks into scoreable dimensions with worked calibration examples.
- 02
Ran paired grading with adjudication and tracked inter-rater agreement per dimension.
- 03
Built adversarial probes targeting known failure modes and regression suites.
- 04
Delivered weekly reports with drift analysis and rubric revisions.
03Outcome
- Agreement metrics stable enough to use as a release gate.
- Regression suite reused across successive model versions.
- Rubric adopted as the team's internal standard.
Similar problem?
Let's talk about yours.
Next chapter
Lending operations platform