Skip to content
Caldorin000/ Forging
← All work/

AI evaluation & data

Representative engagement

Code reasoning evaluation program

Rubric design and expert-graded evaluation of model outputs across software engineering tasks, with reproducible agreement metrics.

Industry

AI lab, remote

Year

2026

Stack

  • Argilla
  • Python
  • DuckDB
  • Weights & Biases

01Problem

A model team needed dependable human judgment on multi-step coding tasks, but ad-hoc grading produced noisy, unrepeatable signals.

02Approach

  1. 01

    Decomposed tasks into scoreable dimensions with worked calibration examples.

  2. 02

    Ran paired grading with adjudication and tracked inter-rater agreement per dimension.

  3. 03

    Built adversarial probes targeting known failure modes and regression suites.

  4. 04

    Delivered weekly reports with drift analysis and rubric revisions.

03Outcome

  • Agreement metrics stable enough to use as a release gate.
  • Regression suite reused across successive model versions.
  • Rubric adopted as the team's internal standard.

Similar problem?

Let's talk about yours.

Next chapter

Lending operations platform