Separate retrieval from generation
A model cannot ground an answer in evidence that retrieval did not find. Evaluate whether the relevant passages appear in the retrieved context, then separately assess whether the generated answer faithfully uses them. This makes failures easier to diagnose.
Build a representative evaluation set
Include frequent questions, ambiguous requests, outdated documents and questions that have no supported answer. Use real domain terminology and expected sources. Keep sensitive information out of evaluation exports unless the handling is explicitly approved.
Test permissions as a quality boundary
Authorization must apply before content reaches the model. Evaluate with users who have different access rights, including users with no access to a source. Cache keys, retrieval filters and source links must preserve the same boundary.
Define an abstention behavior
A production system needs a useful response when evidence is insufficient. Provide a clear uncertainty statement, source references where available and a route to a human or authoritative system. Do not substitute model confidence for verified evidence.
Track cost and drift
Record model and retrieval versions, token usage, latency and evaluation outcomes. Re-run the baseline when documents, embedding models, prompts or retrieval settings change. Treat these as application releases with explicit acceptance criteria.