AI Engineering Field Note · Pathan Afnan Khan
Production RAG: retrieval quality matters more than model size
If retrieval is weak, even a strong model will produce a polished answer from the wrong evidence. Production RAG quality starts with search, ranking and evaluation.
The model cannot rescue bad retrieval
In production RAG systems, the language model is only one part of the quality chain. If the retriever returns weak or irrelevant evidence, a larger model may make the final answer sound better without making it more correct.
That is why I treat retrieval quality as a first-class engineering problem. The goal is not simply to return semantically similar chunks; it is to return the smallest set of evidence that is relevant, authorized, traceable and sufficient for the user's question.
Chunking should preserve meaning and structure
Naive fixed-size chunking can separate a heading from the paragraph it explains, split table rows from their labels, or destroy document hierarchy. Section-aware and structure-aware chunking usually gives retrieval a better representation of the source material.
Metadata should travel with every chunk. Domain, owner, document type, access scope, seniority, date or business unit can all become useful retrieval filters when they reflect real user intent.
Hybrid retrieval and reranking improve precision
Semantic similarity is powerful, but lexical matching still matters for product codes, acronyms, exact skills, identifiers and domain-specific terminology. Combining semantic and lexical signals often produces a stronger candidate set than either approach alone.
I prefer to retrieve broadly and then rerank a smaller candidate set with a relevance model or cross-encoder. This gives the generator a cleaner context window and reduces the amount of unrelated evidence that can distract the final answer.
Measure the retrieval layer separately
A RAG evaluation should distinguish retrieval failures from generation failures. Metrics such as hit rate, precision at k, MRR and citation relevance help show whether the right evidence was found before the LLM answered.
Answer groundedness, latency, token cost and user feedback then complete the picture. When these signals are tracked separately, teams can improve the actual failing component instead of changing prompts or models blindly.