Why it matters
Retrieval quality measures how often the right document appears in top results. Low quality guarantees wrong or hallucination-filled answers no matter the model size.
How it works
Track recall at k, mean reciprocal rank, and human judgments on sample queries. Segment metrics by topic: billing, security, onboarding.
Example
Nintendo tracks retrieval quality on canonical refund prompts. When recall drops after a Help Center IA change, engineers adjust chunk boundaries and redirects before customer complaints spike.
Common mistakes
- 1Measuring only on easy head queries
- 2Conflating keyword match with semantic relevance
- 3No baseline before major CMS migrations
Definitions are useful. Governed knowledge is better.
Wiki helps you turn these concepts into a real system for the knowledge behind your AI.
See Wiki in action