Why it matters
Answer quality spans correctness, completeness, scope, tone, citations, and policy fit. A correct but rude answer fails. A polite wrong answer fails worse.
How it works
Define multidimensional rubrics. Sample production traffic for review. Correlate quality drops with source changes and model updates.
Example
Nintendo scores support AI on refund correctness, citation presence, empathy guidelines, and escalation appropriateness. Quality reviews feed doc updates and new tests.
Common mistakes
- 1Quality equals semantic similarity to one gold string
- 2Ignoring tone on customer-facing channels
- 3No sampling of live traffic only pre-release tests
Test the answer before your customer does.
Run answer tests and evaluations against your governed knowledge before you ship.
See answer evaluations