Pin a prompt, a weighted rubric, and test cases. An independent judge scores every output; the
pass/fail gate is decided in code from your thresholds — the model never grades its own pass. Edit the prompt, re-run, and
diff versions to catch regressions before they ship.
Backend: Claude — target claude-haiku-4-5 / judge claude-haiku-4-5 · same model (judge≈target) · judge ≈ target — set a different judge model for a real diamond
Latest run
No runs yet — hit ▶ Run version to score the suite.