Why it matters
Evaluation converts discomfort into trends leadership can fund. Without trends, teams relitigate anecdotes quarterly while customers experience the same failures.
Evaluation covers correctness, citations, policy fit, tone, refusals, and retrieval quality on sampled and canonical questions.
Eval without action thresholds is performance art. Define what scores block releases before debates get personal.
Adversarial paraphrase and typo-laden queries belong in eval sets. Insiders write canonical questions; customers do not.
How it works
Define rubrics per domain with automated and human components. Automate regressions; humans catch tone and edge cases.
Maintain golden sets with expected sources and claims. Update sets when policy changes.
Segment scores by intent and app. Aggregate scores hide billing failures behind easy FAQs.
Publish eval summaries to stakeholders outside engineering monthly.
Version eval suites with knowledge releases to keep scores comparable.
Rotate human evaluators quarterly to reduce normalization of bad tone or scope.
Version eval suites with knowledge releases so trend lines stay honest after policy changes.
Example
Nintendo evaluates chatbot refund answers monthly against expected Refund Policy language, citation rules, and prohibited-claim checks. Scores dropped after a Help Center rewrite triggered a doc fix before customers noticed.
Common mistakes
- 1Eval sets too small to cover tail questions
- 2Human eval without inter-rater guidelines
- 3Scores tracked but not tied to release gates
- 4Evaluating paraphrase too loosely on numeric facts
Test the answer before your customer does.
Run answer tests and evaluations against your governed knowledge before you ship.
See answer evaluations