RAG

RAG, explained as a retrieval pipeline

A grounded introduction to chunking, embeddings, retrieval, and the evaluation work that makes RAG useful.

RAG pipeline from source documents through retrieval into a language model response.
On this page

Retrieval-augmented generation (RAG) gives a language model relevant material at answer time. Instead of expecting the model to remember every private or changing fact, an application searches a collection and places selected passages in the prompt.

RAG is a pipeline, not a vector database feature. Its output is only as useful as the content, retrieval, and evaluation around it.

Prepare content for retrieval

Start with source documents that are current and allowed to be used. Parse them into meaningful sections, then split those sections into chunks that preserve enough context to make each passage understandable.

Very small chunks lose context. Very large chunks can bury the relevant detail and consume the prompt budget. Test a few chunking strategies against real questions rather than choosing a size by habit.

Retrieve candidate passages

An embedding model maps text to vectors so semantically similar passages can be found. Many systems also combine semantic search with keyword search and metadata filters. The right approach depends on the vocabulary and shape of the questions people actually ask.

text
Question
  -> query rewrite (optional)
  -> retrieve candidates
  -> rank and filter
  -> assemble context
  -> generate answer with citations

Generate with evidence

Give the model a concise instruction to answer from the supplied context and to say when the evidence is insufficient. Return citations that point back to actual source passages, not just a document title.

This does not make the model infallible. It makes the evidence available and gives the application a way to check whether a response is grounded.

Evaluate each stage

Build a small set of representative questions with expected source passages and acceptable answers. Measure retrieval quality separately from answer quality. If the right passage never reaches the prompt, changing the generation prompt will not fix retrieval.

Track no-result questions, stale sources, permission-filter behavior, and latency. Re-indexing should remove deleted or newly restricted content, not only add new files.

When RAG is not the answer

If the information is short, stable, and already known to the application, a direct lookup or ordinary prompt may be simpler. Use retrieval when a changing or sizable knowledge collection needs to be searched and its sources matter.

The practical starting point is one data source, a handful of real questions, inspectable retrieved passages, and a repeatable evaluation. Expand from evidence rather than from the diagram.

Keep learning with Sri

More practical tutorials and experiments on the channel.

Watch on YouTube
Back to articles