Skip to content
October 3, 2026
Search
humaineeti AI engineered for your business
Technology • 5 min read

RAG engineering: your problem is almost always retrieval quality

When a RAG system gives bad answers, teams blame the model. Usually the fault is upstream, in retrieval. Here is how to think about RAG retrieval quality and fix the part that actually matters.

Here is the pattern I see again and again. A team builds a retrieval-augmented generation system, the answers are disappointing, and the first instinct is to swap the language model for a bigger one. They spend a fortnight on it, the answers get marginally better, and they are still disappointed. The model was never the problem. The retrieval was.

A RAG system has a simple truth at its core: the generator can only be as good as the context it is handed. If retrieval pulls the wrong passages, or the right passages chopped into unusable fragments, no model on earth will save the answer. It will just produce a fluent, confident response built on the wrong material. So if you are serious about RAG, RAG retrieval quality is where most of your attention belongs.

Chunking is a real decision, not a default

The least glamorous part of RAG, splitting documents into chunks, quietly determines a lot of your ceiling. Chunk too large and each piece contains several ideas, so the embedding becomes a blurry average and retrieval gets imprecise. Chunk too small and you sever the context a passage needs to make sense, so you retrieve a fragment that is technically relevant and practically useless.

There is no universal chunk size, which is exactly why treating it as a default you never revisit is a mistake. It depends on your documents. Dense technical material wants different handling from long narrative prose. The reliable move is to look at what your retrieval actually returns for real queries and adjust, rather than accepting whatever the tutorial set.

Embeddings decide what similar even means

Your embedding model defines the geometry of your search. Two passages are close if the model thinks they are close. And if the model never saw the kind of language in your domain during training, its judgement of closeness will be off in ways that are hard to see.

This is where domain matters. General-purpose embeddings are fine for general text and can quietly underperform on specialised vocabulary, product codes, legal phrasing, clinical terms. Before assuming your retrieval logic is broken, it is worth asking whether the embedding model even understands your domain’s language in the first place.

Pure vector search is not the finish line

Semantic search is powerful and it has a known blind spot: exact terms. Vector search can sail right past the precise product code, name, or rare keyword a user typed, because semantically it looks unremarkable. Keyword search catches exactly those cases and misses the conceptual ones.

This is why hybrid search, combining semantic and keyword retrieval, so often beats either alone. A reranking step on top of first-pass retrieval is frequently the highest-return improvement you can make. A reranker takes your candidate passages and reorders them with a more careful relevance judgement than the initial fast search could afford. If you have never added reranking, it is usually the first thing I would try.

RAG retrieval quality: you cannot improve what you do not measure

The most important shift in RAG retrieval quality is to evaluate retrieval on its own, separately from the final answer. If you only judge the end result, you can never tell whether a bad answer came from bad retrieval or bad generation. So you keep tuning the wrong half.

Measure retrieval directly. For a set of real queries, did the passages that actually contain the answer get retrieved at all? And did they land near the top, not buried at rank twenty? Ask two questions: is the right material being found, and does it rank highly? The answers tell you more about your system’s health than any amount of staring at final outputs. Once retrieval is measured, improving it becomes an engineering loop instead of guesswork.

The order of operations that saves time

When a RAG system underperforms, resist the urge to reach for a bigger model first. Work the retrieval path in roughly this order. Confirm the right passages are even being retrieved for failing queries. If not, look at chunking and at whether your embeddings understand your domain. Add hybrid search if you are relying on vectors alone. Add reranking. Measure retrieval quality directly at each step so you know whether you are actually improving it. Only once retrieval is genuinely good should you turn your attention to the generation side.

Nine times out of ten, the disappointing RAG system gets fixed somewhere in that list, long before you touch the model.

This is unglamorous, high-leverage work. We spend a lot of time on it at humaineeti when we help teams get RAG systems into production. RAG retrieval quality is rarely the part anyone is excited to work on. Yet it is almost always the part that decides whether the whole system is trustworthy or just confident.

Comments 1

Leave a Reply

Your email address will not be published. Required fields are marked *