AI evaluation · Concept
Catch the regression the average hides.
A concept for an LLM evaluation tool. A fine-tuned model scores higher on average, so it ships. Two weeks later, support is giving medical advice it used to refuse. I designed Rubric around the number that actually decides a release: the case that got worse.
- Role
- Product design, concept
- Type
- Self-directed
- Year
- 2026

The trap
Aggregate scores reward the happy path. A model that answers more questions correctly looks better, even when the new correctness comes from answering things it should have refused. The average moves up. The dangerous case moves down. Ship it, and you learn about the regression from a user.
The bet
So the screen does not open with the plus two percent on accuracy. It leads with the single case that regressed: the input, the old refusal, the new answer, and the assertion that failed, laid out side by side. The thing you would have missed is the first thing you see.
The gate
Promote to prod stays disabled while a regression stands. It is the same principle as the rest of my work: the model proposes, a person clears it. Trust is not a score. It is the moment someone can say no before it reaches a user.