AI evaluation · Concept

Catch the regression the average hides.

A concept for an LLM evaluation tool. A fine-tuned model scores higher on average, so it ships. Two weeks later, support is giving medical advice it used to refuse. I designed Rubric around the number that actually decides a release: the case that got worse.

Role
Product design, concept
Type
Self-directed
Year
2026
An eval run comparing two model versions, with one regressed metric flagged and the failing case shown side by side
One run, two models. Accuracy went up, so the average says ship. Rubric surfaces the one metric that regressed, the exact case behind it, and blocks the promote until a person clears it.

The trap

Aggregate scores reward the happy path. A model that answers more questions correctly looks better, even when the new correctness comes from answering things it should have refused. The average moves up. The dangerous case moves down. Ship it, and you learn about the regression from a user.

The bet

So the screen does not open with the plus two percent on accuracy. It leads with the single case that regressed: the input, the old refusal, the new answer, and the assertion that failed, laid out side by side. The thing you would have missed is the first thing you see.

The gate

Promote to prod stays disabled while a regression stands. It is the same principle as the rest of my work: the model proposes, a person clears it. Trust is not a score. It is the moment someone can say no before it reaches a user.