Evaluate RAG Before Changing the Model
Key takeaway: When a RAG answer is wrong, first determine whether the evidence was missing, poorly retrieved or incorrectly used. Each failure points to a different fix.
Turn a demo into a set of testable questions
Retrieval-augmented generation gives a language model retrieved material to use when answering a question. Imagine an internal assistant that answers questions about product documentation. A convincing answer to a familiar question says little about how it handles an outdated manual, an ambiguous product name or a question the documents cannot answer.
Build a small, reviewed collection of representative questions. For each one, record the relevant document version, the evidence needed and an acceptable answer or refusal. Include questions that need more than one passage and questions with insufficient evidence. Keep a separate set for checking changes after tuning.
Inspect retrieval separately from generation
The RAGAS research distinguishes the relevance of retrieved context, the faithfulness of an answer to that context and the quality of the answer itself. That separation is useful even when the evaluation is a spreadsheet reviewed by a person rather than an automated framework. RAGAS: automated RAG evaluation
For a failed question, read the retrieved passages first. If the required information was never retrieved, investigate document coverage, parsing, chunk boundaries and ranking. If the right passage was present but the answer contradicted it, investigate how the model used the context. Changing both layers at once makes the improvement harder to explain.
Make citations earn their place
A citation is useful only if the cited passage supports the specific claim. Check the document version and whether a qualification was lost. “The feature is available on the enterprise plan” is different from “the feature is available”; retrieving the correct page does not excuse omitting the condition.
Test the boundary explicitly: when the documentation lacks an answer, the assistant should explain that limitation or ask for clarification. Also test access controls. A relevant document is not necessarily a document that the current user is allowed to see; permissions need to be enforced by the retrieval system.
Compare changes with an evaluation record
Keep the questions fixed while comparing one change at a time: a different chunking rule, ranking method or model configuration. Record retrieved passages, answers, latency and failures. An automatic judge can help prioritise review, but its scores need spot checks against human judgement and a clear rubric.
Before release, write down what improved, what became worse and which question types remain unreliable. The decision is then concrete: ship within a narrower scope, improve the document collection, or continue testing. A better-sounding answer alone is not enough evidence to choose.
