RAG Architecture Deep Dive: Vector Embeddings, Hybrid Search & Reranking
A production engineering guide to building accurate Retrieval-Augmented Generation pipelines using dense embeddings, BM25 keyword matching, and cross-encoder rerankers.
On this page
Retrieval-Augmented Generation (RAG) bridges private knowledge bases with the generative capabilities of large language models. However, standard naive RAG pipelines suffer from low precision, irrelevant context retrieval, and hallucinations.
In this guide, we explore the multi-stage architecture required for enterprise-grade retrieval.
The Limitations of Naive RAG
Naive RAG relies solely on cosine similarity over fixed-chunk vector embeddings:
- Text is split into naive 500-token chunks.
- An embedding model converts chunks into dense vectors.
- User questions are embedded and top-$k$ nearest neighbors are retrieved.
This fails whenever the user's query depends on exact product codes, acronyms, or multi-hop logic that dense semantic embeddings tend to blur.
Stage 1: Chunking with Context Preservation
Rather than slicing text arbitrarily at character counts, effective chunking respects markdown structure and document semantics:
from typing import List
def chunk_markdown_by_section(markdown_text: str) -> List[dict]:
"""Splits markdown by headings to maintain contextual coherence."""
sections = markdown_text.split("\n## ")
chunks = []
for idx, section in enumerate(sections):
if not section.strip():
continue
lines = section.split("\n")
title = lines[0].replace("#", "").strip()
body = "\n".join(lines[1:]).strip()
chunks.append({
"chunk_id": f"chunk-{idx}",
"header": title,
"text": f"Section: {title}\n{body}"
})
return chunksStage 2: Hybrid Search with Reciprocal Rank Fusion
Hybrid search combines the semantic generalization of dense embeddings with the exact keyword precision of sparse BM25:
def reciprocal_rank_fusion(dense_ranks: list, sparse_ranks: list, k: int = 60) -> list:
"""Combines ranking lists using the RRF algorithm."""
scores = {}
for rank, doc_id in enumerate(dense_ranks):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
for rank, doc_id in enumerate(sparse_ranks):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
# Sort documents by combined fusion score
sorted_docs = sorted(scores.items(), key=lambda item: item[1], reverse=True)
return [doc_id for doc_id, score in sorted_docs]Stage 3: Cross-Encoder Reranking
Bi-encoder embedding models compute query and document representations independently. Cross-encoders, on the other hand, evaluate both simultaneously through all self-attention layers:
from sentence_transformers import CrossEncoder
reranker = CrossEncoder('cross-encoder/ms-marco-MiniLM-L-6-v2')
def rerank_top_k(query: str, retrieved_docs: list, top_n: int = 3) -> list:
pairs = [[query, doc["text"]] for doc in retrieved_docs]
scores = reranker.predict(pairs)
for doc, score in zip(retrieved_docs, scores):
doc["score"] = float(score)
return sorted(retrieved_docs, key=lambda d: d["score"], reverse=True)[:top_n]Conclusion
A modern RAG architecture is not a single vector lookup. It is an optimized multi-tier information retrieval system that balances recall at the candidate generation layer with high precision at the reranking layer.
Keep learning with Sri
More practical tutorials and experiments on the channel.