GritAI Prompt Regression Harness
LLM-as-Judge · Eval Gates
Pin a prompt, a weighted rubric, and test cases. An independent judge scores every output; the pass/fail gate is decided in code from your thresholds — the model never grades its own pass. Edit the prompt, re-run, and diff versions to catch regressions before they ship.
Backend: Claude — target claude-haiku-4-5 / judge claude-haiku-4-5 · same model (judge≈target) · judge ≈ target — set a different judge model for a real diamond

Suite · the contract

min case score
min pass-rate
regression Δ
Each run executes every case through the target prompt, then the judge scores it. 4 cases × 2 calls.

Latest run

No runs yet — hit ▶ Run version to score the suite.