Skip to content

evaluation

banking77 canary: 87.2% pass, two dead runs, one false alarm

the pipeline is two open-source repos, working together. thomas is the training harness. gonogo is the evaluation and decision layer. the task was intent routing on banking77 — 13,083 real retail-bank customer messages across 77 intents. the model: ModernBERT-small-v2, a ~38M-parameter modern encoder, fine-tuned with a fresh 77-way classification head.

Your Agent Eval Is Lying to You: 47/50 Is Not 94%

A common situation: fifty real cases went through our agent, forty-seven came back right. So we call this "94% accuracy," and it is accepted without much protest by stakeholders.

But the 95% confidence interval on 47/50 runs from 83.8% to 97.9%.

So imagine in our situation the number you'd agreed to hit was 90%. You are in the awkward position of having neither hit nor missed your target. The truth is you still don't know, and what you should be saying out loud is that you need more cases. But it would be much better if your eval tool could tell you that up front. Sometimes it feels impossible to illustrate this point without a clever report.