Evaluation and evidence
"It seems to work" is not evidence. Before you let a workflow act on real cases, and continuously afterwards, you need a way to measure whether it is good enough.
Build a labelled evaluation set
Collect real (or realistic) examples with known correct answers: messages with the right extracted fields, invoices with the right totals, requests with the right routing. This set is the ruler you measure against, and it is worth more than any model choice.
Measure the things that matter operationally
Generic accuracy is rarely the right metric. Ask what a mistake actually costs:
- Precision vs recall. Is a false positive or a false negative worse here?
- Confidence calibration. When the system says it is unsure, is it actually unsure? That is what makes the human-approval routing trustworthy.
- Time-to-correct. How long does it take a person to fix a wrong output?
Watch the exceptions, not just the average
A model that is 95% right can still be unusable if the 5% fails silently on the cases that matter most. Segment your evaluation by scenario and look hard at the tail.
Keep evaluating in production
Inputs drift, models change, and customers do surprising things. Sample real cases, keep a human-review queue, and track whether quality holds over time. Regressions should be visible before a customer finds them.
Report honestly
When you show results to a customer or on a product page, state what was measured, on what data, and where it falls short. Overclaiming is the fastest way to lose a technical buyer's trust — and it violates the honesty rules this catalogue holds itself to.
Related
See measuring economic value to connect quality metrics to money.