Stop letting your AI guess. Give it a library.#
LangChain agents can act, but they still do not know your data. RAG fixes that. You will build a system that reads your PDFs, finds the right passage for a question, and answers from real text. RAG is the pattern behind most "chat with your docs" apps.
Scaler Academy β every example below is shown with data, tables and worked steps.
What You Can Already Do#
Where You Are, and the Two Problems Left#
1) It doesn't know your stuff β your company's docs, your PDFs, last week's release notes. 2) When it doesn't know, it makes things up (hallucinates) confidently. What follows is the single biggest fix in AI engineering. Same agent loop β but now it can look things up in a library you built.
ποΈ Speaker note Β· landing the bridge. If the shop-assistant agent felt unclear, this is the cleanup. Say it plainly: "With agents, the model chose a tool. Now the tool is its own library β and the main idea moves from which tool to how do we find the right page?." That's RAG, in one sentence.
Why Does RAG Exist?#
The 4 Problems RAG Solves#
An LLM is a smart intern who finished training a year ago and was never allowed inside your office. RAG fixes both β fresh data, and your private data β by handing the intern a relevant page at question-time.
- ποΈ Knowledge cutoffThe model's training ended months ago. It doesn't know yesterday's
RBI rate cut or last week's Budget numbers.
β RAG: fetch today's article - π Private / internal dataIt has never seen your HR policy, your contracts, your
product wiki, or your customer tickets β and you can't paste them all into one prompt.
β RAG: search your docs - π HallucinationsWhen unsure, the model invents confident-sounding answers. Bad in
customer support, dangerous in legal, fatal in healthcare.
β RAG: ground in real text - π No citationsAsk "where did you get that?" and a raw LLM has nothing. Real products
need "see source: page 14".
β RAG: return the source chunk
π‘ The 1-line definition. RAG = Retrieval-Augmented Generation. Before the model answers, you retrieve the most relevant snippets from your own data and stuff them into the prompt. The model then answers from those snippets. No retraining. No fine-tuning. Just open-book exam instead of closed-book.
Industry spotlight Β· where you've seen this already β almost every "AI feature" you've used is RAG underneath. Notion AI answering from your workspace, Glean searching your company's apps, Perplexity citing sources, Intercom's Fin support bot, ChatGPT's "search the web", Cursor finding code in your repo, even the new "ask about this PDF" button in your browser β same pattern, every time. Master this, and you can clone any of them.
π‘ Quick reality-check (mention this). People sometimes ask: "Why not just fine-tune the model on our data?" Three reasons it's usually wrong: (1) expensive and slow to redo every time data changes, (2) still can't cite sources, (3) still hallucinates. RAG is faster, cheaper, traceable, and almost always the right first move.
RAG in One Simple Picture#
The Smart Librarian#
Forget the buzzwords. The whole pipeline is just a smart librarian sitting between your question and the model.
Imagine you ask your friend a question about a 800-page book they've never read. Bad idea β they'll guess. Smart move? You find the 3 most relevant pages, hand them over, and then ask. They read the pages and answer confidently with the real text. That's RAG. The "librarian" who finds those 3 pages is the only new thing we're building today.
The 5-Box Pipeline (Memorise This)#
flowchart LR
Q[β Question<br>user types it] --> S[π Search<br>find relevant chunks]:::hl
S --> F[π Stuff<br>add chunks to prompt]:::hl
F --> L[π§ LLM<br>answer using them]
L --> A[π¬ Answer<br>+ source citation]:::good
Two phases Β· don't confuse them.
Phase 1 β Indexing (offline, once): read your docs β break into chunks β turn each chunk into numbers (embedding) β store in a vector database. Slow, but you only do it when documents change.
Phase 2 β Querying (online, every question): turn the question into numbers β find nearest chunks in the database β paste them into the prompt β LLM answers. Milliseconds per query.
π‘ The important idea. You are not retraining the
model. The model stays exactly the same gpt-4o-mini you called through the raw
OpenAI API. We're only changing what goes into the prompt. RAG is, at its
core, very fancy prompt engineering β automated.
Embeddings: Words Become Coordinates#
What Is an Embedding?#
Before we can "search by meaning", we need a way to turn meaning into numbers. That's an embedding. And once you see what it does, every confusing thing about RAG becomes easier to understand.
Definition Β· keep this in your head. An embedding is a list of numbers (a vector) that captures the meaning of a piece of text. Similar meanings βΆ similar numbers βΆ nearby points in space. Different meanings βΆ far apart.
| Word | x | y | Colour group |
|---|---|---|---|
| banana | 80 | 90 | #4ec9b0 |
| mango | 60 | 115 | #4ec9b0 |
| grape | 100 | 75 | #4ec9b0 |
| cat | 280 | 80 | #e0af68 |
| dog | 300 | 105 | #e0af68 |
| tiger | 265 | 60 | #e0af68 |
| king | 180 | 175 | #f48771 |
| queen | 215 | 155 | #f48771 |
| man | 120 | 215 | #8cc8ff |
| woman | 160 | 200 | #8cc8ff |
| laptop | 60 | 240 | #c586c0 |
| computer | 90 | 255 | #c586c0 |
Imagine a 2D World First β Height & Weight#
Forget AI for a second. If I plot people by height vs weight, people of similar build end up close together on the chart. Same idea, scaled up. Real embeddings have 384, 768 or even 3072 dimensions β way more than we can draw β but the principle is identical: similar βΆ close, different βΆ far.
Because meaning is rich. Two words can be similar in many ways at once β topic, tone, formality, language, sentiment. Each dimension captures a different axis of similarity. 768 isn't arbitrary; it's just enough to separate millions of distinct ideas.
Embedding Math Works#
This example shows why engineers trust embeddings. After training on enough text, vectors carry relationships that you can use in arithmetic. Here are four examples:
| Vector equation | Nearest answer |
|---|---|
| king β man + woman | queen |
| paris β france + india | delhi |
| tokyo β japan + germany | berlin |
| walking β walk + run | running |
π‘ Why this is the foundation of everything here. If woman and queen can be found by simple arithmetic, then "find the chunk most similar to my question" is also just arithmetic β fast, scalable, and runs on commodity hardware. Every RAG system uses exactly this idea. Search by meaning = nearest point in vector space.
See It in 2D β Same Idea in a Drawing#
This is a real word-vector plot reduced from 768 dimensions to 2 so we can draw it. The examples show that the direction between related words stays similar:
| Relationship | Pair A | Pair B | What to notice |
|---|---|---|---|
| Royalty | king β queen | man β woman | Both arrows point in a similar direction. |
| Capital-of | france β paris | india β delhi | The country-to-capital relationship is a direction. |
| Food vs aeroplane | banana and grape | aeroplane | banana and grape are close. aeroplane is far away, with cos β 0.05. |
The Score We Use: Cosine Similarity#
Two vectors close together = the angle between them is small. The cosine of that angle is our score. It ranges from -1 (opposite) to +1 (identical). Practically: above 0.7 means "very similar", below 0.3 means "barely related". No math needed to use it β just remember: bigger number = more similar.
Two clock hands pointing at the same time βΆ angle = 0 βΆ cosine = 1 βΆ perfectly similar. Pointing at 12 and 6 βΆ angle = 180Β° βΆ cosine = β1 βΆ opposites. We don't care how long the hands are, only the angle. That's why text length doesn't break the score.
Example Scores from the Similarity Meter#
This example compares pairs of sentences with cosine similarity. It shows the scores for preset pairs and the scoring rules. The scores come from a small built-in table of topics, so common topics look realistic.
Default example. Sentence A: A cat is sleeping on the couch. Sentence B: A kitten is napping on the sofa. Score: 0.89.
| Preset label | Sentence A | Sentence B | Score |
|---|---|---|---|
| synonyms | A cat is sleeping on the couch. | A kitten is napping on the sofa. | 0.89 |
| paraphrase | I love programming. | I enjoy coding. | 0.86 |
| same city, different names | Flight to Mumbai | Plane to Bombay | 0.92 |
| related concepts | I am hungry | I want food | 0.83 |
| unrelated | I love programming | I hate vegetables | 0.08 |
| totally unrelated | The stock market crashed | A cat is sleeping | 0.04 |
| Baked token pair A | Baked token pair B | Score |
|---|---|---|
| cat sleeping couch | kitten napping sofa | 0.89 |
| cat sleeping couch | dog running park | 0.32 |
| cat sleeping couch | stock market crash | 0.04 |
| love programming | enjoy coding | 0.86 |
| love programming | hate vegetables | 0.08 |
| love programming | i write software | 0.71 |
| flight to mumbai | plane to bombay | 0.92 |
| i am hungry | i want food | 0.83 |
| how to fix bug | debugging tips | 0.79 |
| Score band | Verdict text |
|---|---|
| 0.85 to 1.00 | π’ Near-synonyms β same meaning, different words. |
| 0.65 to 0.84 | π’ Strongly related β same topic. |
| 0.40 to 0.64 | π‘ Loosely related β shares some concepts. |
| 0.20 to 0.39 | π Distantly related β barely overlapping. |
| Below 0.20 | π΄ Unrelated β totally different vector neighbourhoods. |
Fallback formula for other text. For text outside the presets, the example scorer lowercases the text, removes punctuation, keeps words longer than 2 characters, counts exact token overlap, checks shared topic buckets, then returns min(0.97, max(0.02, jaccard * 0.55 + bucketScore + 0.05)). bucketScore is min(0.75, matchingBuckets * 0.4).
| Bucket | Words in this bucket |
|---|---|
| food | eat, food, hungry, breakfast, lunch, dinner, meal, restaurant, recipe, cook, tasty, sweet, spicy, rice, dal, curry, biryani |
| animal | cat, dog, kitten, puppy, tiger, lion, animal, pet, bird, fish |
| tech | code, coding, program, programming, software, bug, debug, python, java, laptop, computer, API, LLM, AI |
| travel | flight, plane, travel, journey, trip, vacation, airport, train, bombay, mumbai, delhi, bangalore |
| work | office, meeting, job, work, career, salary, manager, team |
| sleep | sleep, sleeping, nap, napping, rest, tired, bed, couch, sofa |
| money | money, price, cost, stock, market, rupees, dollar, income, wealth |
How Do We Actually Get an Embedding? One Function Call.#
You don't have to train anything β somebody else already did. Open-source models from Hugging
Face (sentence-transformers) or APIs from OpenAI / Cohere give you embeddings in one
line:
# pip install sentence-transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2") # β free, fast, 384 dims
vec = model.encode("A cat is sleeping on the couch") # β‘ that's an embedding!
print(vec.shape) # β (384,) β 384 numbers
print(vec[:5]) # β [-0.05, 0.12, 0.41, -0.08, 0.22] (something like that)
# to compare two pieces of text β encode both β take cosine similarity
from numpy import dot
from numpy.linalg import norm
v1 = model.encode("A cat is sleeping on the couch")
v2 = model.encode("A kitten is napping on the sofa")
similarity = dot(v1, v2) / (norm(v1) * norm(v2)) # cosine
print(similarity) # β 0.87 (very similar π)
π Decoder
- β all-MiniLM-L6-v2 is a tiny but excellent open-source
embedding model β runs on your laptop, no API key needed. For production, look at
BAAI/bge-large-en-v1.5(best English) or OpenAI'stext-embedding-3-small(paid, very good multilingual). - β‘
.encode("text")returns the vector. That's the whole API surface β one call, 384 numbers back. Now your text is searchable by meaning.
ποΈ Speaker note Β· how to pitch this section. This is the "physics" of RAG. Don't rush. Spend a full 5 minutes letting them play with the meter and the king-queen equation. Once they internalize "similar meaning = nearby vector", everything else here is bookkeeping.
- π Vectors = coordinatesEvery chunk of text gets a point in space.
- π Cosine similarityScore from β1 to +1. Bigger = closer.
- πͺ Math you can doking β man + woman β queen. Really.
- β‘ One function call
.encode(text)β that's it.
Chunking: Cut the Book into Snippets#
Why Chunking Decides Everything#
Before we embed anything, we have to cut it up. You can't embed a 200-page PDF as one vector β meaning gets averaged away to mush. So we split it into chunks. How you chunk decides how good your RAG actually is. Engineers underestimate this; the best ones obsess over it.
One whole pizza is hard to share β too big. Cut it into 100 confetti-sized bits and nobody can taste anything. Slice sizes matter. Chunks are pizza slices: too big and you lose precision (the relevant bit is buried), too small and you lose context (each crumb is meaningless).
The 4 Chunking Strategies Worth Knowing#
- π Fixed sizeEvery N characters or tokens β done. Very simple, fast, your
default. Downside: chops mid-sentence, mid-table, mid-thought.
Best for: prototypes, blog posts - π RecursiveTry to split on paragraphs first; if too big, split on sentences; if still
too big, on words. A clear fallback ladder. The default in LangChain.
Best for: most real apps - π§ SemanticEmbed every sentence; group consecutive sentences that "talk about the same
thing". Chunks follow meaning, not size.
Best for: long flowing prose - π Structure-awareUse the document's own structure β markdown headings, code blocks,
HTML sections. Each section becomes a chunk.
Best for: docs, code, wikis
Compare All 4 Strategies β Same Text, Four Different Cuts#
The same source passage is cut in four ways. The table shows where the cuts land and how chunk shape changes. The differences are the lesson.
| Strategy | Setting | Output chunks | Stats | What to notice |
|---|---|---|---|---|
| Fixed size | 80 characters for the worked example. The old range was 40β240 characters, default 120. | Chunk 1 (80 chars; mid-word cut) # Bengaluru Overview ## Tech Industry Bengaluru is known as India's Silicon Val Chunk 2 (80 chars; mid-word cut) ley. Tech parks like Electronic City and Whitefield host thousands of tech compa Chunk 3 (80 chars; mid-word cut) nies. Major firms include Infosys, Wipro, and TCS. ## Climate The city sits at Chunk 4 (80 chars; mid-word cut) 920 meters altitude. This gives it pleasantly cool weather year-round. Average t Chunk 5 (80 chars; mid-word cut) emperatures rarely exceed 30 degrees. ## Food Bengaluru's food scene is legenda Chunk 6 (80 chars; mid-word cut) ry. South Indian classics like masala dosa and idli thrive here. Filter coffee s Chunk 7 (29 chars; mid-word cut) hops dot every street corner. | 7 chunks; 73 average chars; 7 mid-word cuts | Cuts land at character N. They do not respect words, sentences, or paragraphs. This is cheap to write, but poor for retrieval. |
| Recursive | 80 characters, with paragraph β sentence β word fallback. | Chunk 1 (20 chars) # Bengaluru Overview Chunk 2 (62 chars) ## Tech Industry Bengaluru is known as India's Silicon Valley. Chunk 3 (80 chars) Tech parks like Electronic City and Whitefield host thousands of tech companies. Chunk 4 (44 chars) Major firms include Infosys, Wipro, and TCS. Chunk 5 (48 chars) ## Climate The city sits at 920 meters altitude. Chunk 6 (49 chars) This gives it pleasantly cool weather year-round. Chunk 7 (46 chars) Average temperatures rarely exceed 30 degrees. Chunk 8 (44 chars) ## Food Bengaluru's food scene is legendary. Chunk 9 (60 chars) South Indian classics like masala dosa and idli thrive here. Chunk 10 (44 chars) Filter coffee shops dot every street corner. | 10 chunks; 50 average chars; 0 mid-word cuts | Same character budget, cleaner cuts. LangChain uses RecursiveCharacterTextSplitter as the safe default for about 90% of RAG apps. |
| Semantic | No size knob. Sentences are tagged by topic and adjacent same-topic sentences are merged. | tech (171 chars) Bengaluru is known as India's Silicon Valley. Tech parks like Electronic City and Whitefield host thousands of tech companies. Major firms include Infosys, Wipro, and TCS. climate (134 chars) The city sits at 920 meters altitude. This gives it pleasantly cool weather year-round. Average temperatures rarely exceed 30 degrees. food (142 chars) Bengaluru's food scene is legendary. South Indian classics like masala dosa and idli thrive here. Filter coffee shops dot every street corner. | 3 chunks; 3 topics detected; size range 134-171 | Chunks are shaped by meaning, not size. Tech sentences join together. Climate sentences join together. Food sentences join together. |
| Structure-aware | No size knob. Markdown headers set the boundaries. | Section 1 (20 chars) # Bengaluru Overview Section 2 (188 chars) ## Tech Industry Bengaluru is known as India's Silicon Valley. Tech parks like Electronic City and Whitefield host thousands of tech companies. Major firms include Infosys, Wipro, and TCS. Section 3 (145 chars) ## Climate The city sits at 920 meters altitude. This gives it pleasantly cool weather year-round. Average temperatures rarely exceed 30 degrees. Section 4 (150 chars) ## Food Bengaluru's food scene is legendary. South Indian classics like masala dosa and idli thrive here. Filter coffee shops dot every street corner. | 4 chunks; 4 sections; 126 average chars | Each ## section becomes one chunk with its header attached. This is best for docs, wikis, source code, Markdown, and HTML. |
| Default setting | Computed stats |
|---|---|
| Fixed size 120 | 5 chunks; 102 average chars; 2 mid-word cuts |
| Recursive size 120 | 8 chunks; 62 average chars; 0 mid-word cuts |
Read this sequence: Fixed at size 80 creates mid-word cuts. Recursive at the same size removes those cuts. Semantic creates three chunks that match the three topics. Structure uses the markdown ## headers as ready-made boundaries.
The recall vs precision tradeoff Β· keep this in mind. Big chunks β high recall (the answer is probably in there), but low precision (lots of fluff around it). Small chunks β high precision (the chunk is exactly the answer), but low recall (the relevant bit might be split across two chunks and you only pulled one). Most teams sweep chunk size as their first RAG tuning knob.
One Line of Real Code#
from langchain_text_splitters import RecursiveCharacterTextSplitter
# β choose chunk size and overlap so nearby context stays connected
splitter = RecursiveCharacterTextSplitter(
chunk_size=800, # aim for ~800 chars per chunk
chunk_overlap=100, # adjacent chunks share 100 chars (context glue)
)
# β‘ split the long document into chunks ready for embedding
chunks = splitter.split_text(your_long_document)
# β’ print how many chunks will go into the vector search index
print(len(chunks)) # β e.g. 47 chunks ready to embed
π‘ Why chunk_overlap? If your question's answer sits exactly at
a chunk boundary, you'd miss it. Overlap (50β100 chars) makes sure every sentence appears in at
least one chunk fully. Cheap insurance.
Vector Databases: a Search Engine for Meaning#
What a Vector Database Is#
You have thousands of embeddings. For every question, you need the top-k nearest ones β fast. That's what a vector database does, and that's all it does.
A vector database is a search engine where instead of "find documents containing this word", the query is "find vectors closest to this vector". Same idea as Google, swapped engine.
The Three Names You'll Hear All the Time#
- π’ Chroma β
todayOpen-source, runs in-process (no server!), one
pip install. Perfect for prototypes & up to ~10M vectors. What we'll use today. - π΅ PineconeFully managed SaaS. Zero ops, scales to billions, but paid. The boring-and-reliable choice for production.
- π£ QdrantOpen-source and production-grade. Self-host or use their cloud. Great middle ground when you outgrow Chroma.
π― Picking one (don't overthink). Prototyping or under 1M chunks? Chroma. Want zero ops & have a budget? Pinecone. Need to self-host at scale? Qdrant or Weaviate. The good news: LangChain wraps all of them with the same interface, so swapping later is a one-line change.
What It Actually Does β Three Operations#
A vector DB is just these three calls:
db.add(documents, embeddings)β store chunks and their vectors. (Indexing.)db.query(query_vector, k=3)β return the 3 chunks whose vectors are closest. (Retrieval.)db.delete(ids)β remove chunks when a doc is deleted. That's it. Three calls.
Industry spotlight Β· how it stays fast β HNSW, the trick behind every vector DB. Searching billions of vectors naively means computing billions of distances per query. HNSW (Hierarchical Navigable Small World β a 2018 algorithm) builds a "ladder" graph so each query takes only ~log(N) hops. Result: sub-10ms lookups on 100M vectors. You'll never write this yourself β every vector DB ships it built-in.
Build a Real RAG in 30 Lines#
Step 1 β Index Your Documents (Once)#
We've talked about every piece. Now we wire them up. This is the whole pattern β every RAG system you'll see in production is just a fancier version of this.
index.py
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_huggingface import HuggingFaceEmbeddings
from langchain_chroma import Chroma
# β load your documents (any text β for now, hardcoded)
docs = [
"Our return policy allows refunds within 30 days of purchase.",
"Shipping is free for orders above βΉ999 across India.",
"For corporate orders above 50 units, contact sales@example.com.",
"Our office is in Indiranagar, Bangalore. Open Mon-Fri 10am-7pm.",
]
# β‘ split into chunks (small docs here, but production = thousands of pages)
splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50)
chunks = splitter.create_documents(docs)
# β’ pick an embedding model (free, runs locally, no API key)
embedder = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
# β£ build the vector store from chunks + embeddings (saves to disk)
db = Chroma.from_documents(chunks, embedder, persist_directory="./chroma_db")
print(f"Indexed {len(chunks)} chunks π")
Step 2 β Ask a Question (Every Time)#
from langchain_huggingface import HuggingFaceEmbeddings
from langchain_chroma import Chroma
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from dotenv import load_dotenv
load_dotenv()
# β€ open the same vector store we built in step 1
embedder = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
db = Chroma(persist_directory="./chroma_db", embedding_function=embedder)
model = ChatOpenAI(model="gpt-4o-mini", temperature=0)
# β₯ the prompt: instructions + chunks + question
prompt = ChatPromptTemplate.from_template("""
Answer the question using ONLY the context below. If the context doesn't contain
the answer, say "I don't know." Be concise and quote facts directly.
Context:
{context}
Question: {question}
""")
def rag_answer(question):
chunks = db.similarity_search(question, k=3) # β¦ retrieve top 3
context = "\n\n".join(c.page_content for c in chunks)
chain = prompt | model # β§ same LangChain chain trick
return chain.invoke({"context": context, "question": question}).content
print(rag_answer("How long do I have to return something?"))
# β "30 days from the purchase date." β
from your data, not a guess
π Decoder β the 8 numbered steps
- β Your data. In real life this is PDFs, web pages, Notion exports β anything you can turn into text.
- β‘ Chunk it. Block 17 in action.
chunk_size=500is a sane default. - β’ Pick an embedder. We're using a free local one β swap to
OpenAIEmbeddings()for a paid, slightly better version. - β£
Chroma.from_documentsembeds every chunk and stores both the text and the vector. Done once per dataset. - β€ Reopen the same database. The embeddings are already on disk β no re-encoding.
- β₯ The important prompt template. Notice
"ONLY the context below"β this single line is the most important hallucination reducer in RAG. - β¦
similarity_search= embed the question, find the 3 nearest chunks, return them. The librarian. - β§ Same LangChain pipe (
prompt | model) you used for LangChain chains. RAG didn't replace your previous knowledge β it slotted right in.
Follow the Pipeline Step by Step#
Here is what happens inside rag_answer("How long do I have to return something?"). Each item is one step in the RAG flow:
- 1 Β· User question.
"How long do I have to return something?" - 2 Β· Embed the question. Turn the question into a 384-number vector with
model.encode(question). Think of it as the question's coordinates. - 3 Β· Search the vector DB.
db.similarity_search(query, k=3)finds the 3 nearest chunks: refunds within 30 days...shipping is free above βΉ999corporate orders contact sales@ - 4 Β· Put the chunks into the prompt. The retrieved chunks become
{context}. The question becomes{question}. One string goes to the LLM. - 5 Β· LLM answers from context. GPT-4o-mini replies: "You can return items within 30 days of purchase." β The answer came from the chunk, not from memory.
π― Take a moment β this is the canonical pattern. That's it. Every RAG system in the world β from a hobby project to Perplexity β is a variant of these 8 lines. From here we're just making it better: smarter retrieval, smarter prompts, multiple passes. The core never changes.
Make It Better: Rerank + Hybrid Search#
Problem 1: the Top-5 from Vector Search Isn't the Best 5#
A vanilla RAG works. A good RAG works well. The two tricks that close that gap β and that every senior engineer asks about β are reranking and hybrid search. Both are easy to add.
Embedding similarity is fast, but it sometimes ranks shallow word-matches above deep semantic matches. So we use a two-stage retrieval β a fast first pass to narrow down, then a slow accurate pass to pick the real winners.
You can't interview 1,000 people (it would take a year). You also can't pick someone by gut feel from a stack of resumes (you'd hire badly). So you do two passes: a fast resume scan to shortlist 25, then a real 30-minute interview with each of those 25. Bi-encoder and cross-encoder are exactly these two passes.
- π Bi-encoder Β· like a resume scanReads the query and the document
separately, turns each into a vector, then compares the two vectors with cosine
similarity. The encoder never sees them together β it judges each "card" alone.
β‘ Fast β sub-10ms over millions of docs
πΎ Pre-computable β encode all your docs once, store forever, only encode the query at runtime
π Shallow β misses subtle relevance because the model never compares them side-by-sideβ use it to grab top 50β100 candidates - ποΈ Cross-encoder Β· like a 30-min interviewFeeds the query and the
document into one transformer together, lets the model attend to both at once, then
outputs a single relevance score. The model can compare them token by token, weigh trade-offs,
spot nuance.
π― Way more accurate β sees queryβdoc word interactions directly
π’ Slow β full transformer run per (query, doc) pair
β Not pre-computable β each score is pair-specific, you can't cache anythingβ use it to rerank those 50 β top 3
flowchart LR
A[1,000,000<br>all your chunks] -->|π bi-encoder Β· ~10 ms| B[50<br>candidates]:::hl
B -->|ποΈ cross-encoder Β· ~200 ms| C[3<br>to the LLM]:::good
Remember this β the two-stage pattern is universal.
Stage 1 (bi-encoder): grab top 50-100 candidates from the vector DB. Fast.
Stage 2 (cross-encoder): re-score those candidates with a smarter model. Slow per item, but you're only re-scoring 50, not 50 million. Keep the top 3-5 for the LLM.
Reranking Flips the Order#
Same query, same candidates. Left: what cosine similarity returned. Right: what a cross-encoder reranker returned. The relevant answer rises to the top:
The code to add reranking is short:
rerank.py
from sentence_transformers import CrossEncoder
# β load a cross-encoder that can score a question with one chunk
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2") # free, fast
def retrieve_with_rerank(question, top_k=3):
# β retrieve more cheap candidates than you finally need
candidates = db.similarity_search(question, k=25) # grab 25 cheap candidates
# β‘ pair the same question with each candidate chunk for scoring
pairs = [(question, c.page_content) for c in candidates]
# β’ score every question chunk pair with the cross-encoder
scores = reranker.predict(pairs) # cross-encoder scores each pair
# β£ sort by score and return the best chunks
return [c for _, c in sorted(zip(scores, candidates), reverse=True)[:top_k]]
Problem 2: Sometimes You Need the Exact Word#
Pure semantic search ignores keywords. Ask "what's the price of SKU-4429?" and the embedding doesn't care about that specific code β it just sees "price" and "product". For codes, names, IDs, dates β keyword search beats semantic hands down. The fix isn't to pick one or the other. It's to use both, with very different strengths.
Remember our librarian from Block 15? She has two assistants who search the stacks completely differently. One is a literalist who lives for exact words. The other is a philosopher who lives for meaning. Send your query to the right one and you get great results. Send your query to both β and let them merge their rankings β and you get strong results. That's hybrid search.
- π€ Librarian A Β· BM25 Β· the literalistObsessed with exact words. Type
"SKU-4429" or "Article 21" and she'll find every single doc containing those
exact characters. But ask for "running shoes" and she'll skip the doc titled
"marathon footwear" β different words, even though same meaning.
β Brilliant at product codes, SKUs, model numbers, version strings
β Brilliant at proper nouns, dates, legal references, technical jargon
β Blind to synonyms, paraphrase, intent β words are just characters to herβ catches what vector misses - π§ Librarian B Β· Vector Β· the philosopherObsessed with meaning. Ask for
"running shoes" and she surfaces "marathon footwear", "jogging
trainers", even "sneakers for athletes". But ask for "SKU-4429" and
she shrugs β random letters and numbers are just noise to her.
β Brilliant at natural-language questions, "what's this about?"
β Brilliant at unclear queries, paraphrase, multilingual matching
β Blind to exact codes, IDs, technical strings β drowns them in averagesβ catches what BM25 misses
Hybrid = Both Librarians on the Case#
Send the same query to both librarians at the same time. Each ranks the docs by their own logic. You then merge their two ranked lists into one using a tiny formula called Reciprocal Rank Fusion (RRF):
β‘ Reciprocal Rank Fusion in one line. For each doc, final score =
1 / (60 + rank in BM25) + 1 / (60 + rank in Vector). Docs that both
librarians ranked highly bubble to the top. Docs only one of them liked still get a fair shot. No
tuning, no thresholds, no main idea numbers (well, 60 β but it almost never matters). That's
it. Used by almost every production RAG system in the wild.
Real user queries are messy β mostly natural language, but sprinkled with technical terms, product codes, names, or jargon the embedding model has never seen. Hybrid covers both halves of every messy query automatically: the natural-language part goes to Vector, the technical-term part goes to BM25, and RRF stitches the answers together. You're not picking sides β you're using each tool for what it's actually good at.
Compare Search Methods on "I Want Running Shoes"#
Same query, three rankings. Librarian A and Librarian B return different top picks. Hybrid merges the rankings:
score: 8.40
score: 3.10
score: 2.40
score: 1.20
Counts word overlap. It likes documents that include "shoes".
score: 0.83
score: 0.78
score: 0.71
score: 0.55
Understands that "running" relates to jogging and trainers.
fused: 0.0328
fused: 0.0325
fused: 0.0323
fused: 0.0318
Reciprocal rank fusion combines both ranked lists. It is a common production default.
Industry spotlight Β· this is what the leaderboard chases β every "+10% retrieval quality" paper is one of these tricks. Reranking + hybrid search are the two single biggest quality wins in RAG. Add them and you go from "demo works" to "production works". Almost every benchmark in the MTEB leaderboard uses some combination. You now know the playbook.
Mini-Project: Chat with Your PDF#
The Full App β 4 Files#
Time to ship. Point it at any PDF β your resume, a research paper, an annual report, your college notes β and chat with it. Every answer cites the page it came from. This is the project that gets people asking "how did you build this?"
flowchart LR
P[π PDF in<br>user gives a path] --> C[βοΈ Chunk<br>~800 chars]
C --> E[π Embed + store<br>Chroma]:::hl
E --> R[π Retrieve<br>top-3 chunks]:::hl
R --> T[π¬ Chat<br>terminal loop]:::good
Step 1 β Load the PDF & Index It Once#
# pip install langchain langchain-openai langchain-chroma langchain-huggingface \
# langchain-community pypdf sentence-transformers python-dotenv
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_huggingface import HuggingFaceEmbeddings
from langchain_chroma import Chroma
def build_index(pdf_path):
pages = PyPDFLoader(pdf_path).load() # β read all pages
splitter = RecursiveCharacterTextSplitter(chunk_size=800, chunk_overlap=100)
chunks = splitter.split_documents(pages) # β‘ chunk them
embedder = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
db = Chroma.from_documents(chunks, embedder) # β’ in-memory store
return db
Step 2 β Answer with Retrieval + Citations#
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from dotenv import load_dotenv
# β load the API key and create the PDF assistant model
load_dotenv()
model = ChatOpenAI(model="gpt-4o-mini", temperature=0)
# β‘ define the grounded answer prompt and source-citation format
prompt = ChatPromptTemplate.from_template("""
You are a helpful PDF assistant. Answer the question using ONLY the context below.
If the context doesn't contain the answer, say "I couldn't find that in the document."
After your answer, list the page numbers you used as: Sources: page X, page Y.
Context:
{context}
Question: {question}
""")
def ask(db, question):
# β retrieve the most relevant PDF chunks for the question
chunks = db.similarity_search(question, k=4)
# β‘ include page numbers beside each chunk before prompting the model
context = "\n\n".join(
f"[page {c.metadata['page']+1}] {c.page_content}" for c in chunks)
# β’ connect the prompt to the model for this answer
chain = prompt | model
# β£ ask the model using only the retrieved PDF context
return chain.invoke({"context": context, "question": question}).content
Step 3 β The Entry Point#
A tiny main.py indexes the PDF once, then answers questions in a loop:
from pdf_chat import build_index, ask
pdf_path = input("Path to your PDF: ")
db = build_index(pdf_path) # β index once, before the chat loop
print("β
PDF indexed! Ask me anything about it. (blank line to quit)")
while True:
question = input("\nYou: ").strip() # β‘ every question reuses the same index
if not question:
break
print("Bot:", ask(db, question))
π Decoder Β· the 3 small choices that matter
- β
build_indexruns once, before the loop, and the result is kept indb. Without that, every question would re-index the PDF from scratch (slow!). - β‘ Indexing only once means each question is instant. This is the same indexing/querying split from Block 15.
- β’ We pass
pagefrom the chunk metadata into the prompt (inask) β that's how the model knows which page to cite. Metadata is RAG's strength.
Run It#
Save pdf_chat.py (parts 1 + 2), main.py and your .env in one
folder, then run:
$ python main.py
Path to your PDF: annual_report.pdf
β
PDF indexed! Ask me anything about it. (blank line to quit)
You: What was the revenue this year?
- Point it at any PDF. Your resume, the Indian Constitution, your company's HR policy, last quarter's earnings call transcript. Anything text-based.
- Ask 3 questions. One factual ("when was X founded?"), one summary ("what's the main argument?"), one tricky ("compare X and Y"). Notice the citations.
- Try a question that's NOT in the doc. The answer should be "I couldn't find that in the document." β that's RAG refusing to hallucinate. Watch for this moment. It's the main point.
Walkthrough of the PDF Chat#
This walkthrough uses three sample documents that are already indexed. The table shows the documents, example questions, retrieved answers, and the refusal message for questions not in the document:
Index message from the old demo. Indexed "π HR Policy" β 1 page, ~8 chunks. Ask me anything! The same pattern applied to each sample document.
| Sample document | Indexed text | Answer keys and answers | Example questions |
|---|---|---|---|
| π HR Policy | Annual leave: 22 days/year for full-time employees, accrued monthly. Sick leave: 12 days/year, no carry-forward. Work from home: up to 2 days per week with manager approval. Maternity leave: 26 weeks paid as per Indian law. Paternity leave: 10 working days. Reimbursements: internet βΉ1500/month, mobile βΉ800/month. Notice period: 60 days for senior roles, 30 days for others. Probation: 6 months. Probation extension requires HR approval. | leave: Full-time employees get 22 days of annual leave (accrued monthly) plus 12 days of sick leave per year. Sources: page 1. wfh: Yes β up to 2 days per week with your manager's approval. Sources: page 1. work from home: Yes β up to 2 days per week with your manager's approval. Sources: page 1. maternity: 26 weeks of paid maternity leave, in line with Indian law. Sources: page 1. paternity: 10 working days of paternity leave. Sources: page 1. notice: Notice period is 60 days for senior roles, 30 days for others. Sources: page 1. reimbursement: Internet: βΉ1500/month. Mobile: βΉ800/month. Sources: page 1. internet: Internet reimbursement is βΉ1500/month. Sources: page 1. probation: Probation period is 6 months. Extensions require HR approval. Sources: page 1. | How many leaves do I get?, Can I work from home?, What is the notice period?, What is the maternity policy? |
| π° TCS Q3 Earnings | Revenue: βΉ62,613 cr, up 4.0% YoY in constant currency. Operating margin: 24.6%, up 50 bps QoQ. Net profit: βΉ12,380 cr. TCV (Total Contract Value): $13.2 bn, highest in 7 quarters. Headcount: 612,724 employees, net addition of 5,370 this quarter. Attrition: 13.0% (LTM), down from 13.3%. Cash and equivalents: βΉ58,200 cr. Dividend: βΉ76/share interim declared. BFSI segment grew 3.8%, retail 2.1%, manufacturing 5.9%. | revenue: Revenue was βΉ62,613 cr, up 4.0% YoY in constant currency. Sources: page 1. margin: Operating margin was 24.6%, up 50 basis points quarter-on-quarter. Sources: page 1. profit: Net profit was βΉ12,380 cr. Sources: page 1. tcv: TCV (Total Contract Value) hit $13.2 bn β the highest in 7 quarters. Sources: page 1. headcount: 612,724 employees, with a net addition of 5,370 this quarter. Sources: page 1. attrition: Attrition (LTM) was 13.0%, down from 13.3%. Sources: page 1. dividend: Interim dividend of βΉ76 per share was declared. Sources: page 1. bfsi: BFSI segment grew 3.8% this quarter. Sources: page 1. | What was the revenue?, How was the operating margin?, What is the attrition rate?, How big is the latest TCV? |
| π Indian Constitution (Part III) | Article 14: Equality before law β the State shall not deny equality to any person. Article 15: Prohibition of discrimination on grounds of religion, race, caste, sex or place of birth. Article 19: Six fundamental freedoms β speech, assembly, association, movement, residence, profession. Article 21: Right to life and personal liberty β no person shall be deprived except by procedure established by law. Article 21A: Right to education for children aged 6-14. Article 25: Freedom of conscience and free profession of religion. Article 32: Right to constitutional remedies β Supreme Court can be approached for enforcement. | article 14: Article 14 guarantees equality before law β the State shall not deny equality to any person. Sources: page 1. article 15: Article 15 prohibits discrimination on grounds of religion, race, caste, sex or place of birth. Sources: page 1. article 19: Article 19 grants six fundamental freedoms: speech, assembly, association, movement, residence, and profession. Sources: page 1. article 21: Article 21 protects the right to life and personal liberty β no person shall be deprived except by procedure established by law. Sources: page 1. article 32: Article 32 is the right to constitutional remedies β citizens can approach the Supreme Court for enforcement of Fundamental Rights. Sources: page 1. right to education: Article 21A guarantees the right to education for children aged 6 to 14. Sources: page 1. religion: Article 25 guarantees freedom of conscience and free profession of religion. Article 15 prohibits discrimination on grounds of religion. Sources: page 1. freedom of speech: Article 19 grants freedom of speech as one of six fundamental freedoms. Sources: page 1. | What does Article 21 say?, Tell me about Article 19, What is the right to education?, Which article covers religion? |
Unknown question response. I couldn't find that in the document. π€· This is RAG refusing to hallucinate.
pdf_chat.py uses actual embeddings + GPT with the same flow.Industry spotlight Β· this is the canonical AI product β you just built the
most-shipped AI app of 2024β26. "Chat with [your docs / your PDF / your codebase / your
Notion]" is the single most common AI feature on the market today β and almost all of them are
this exact pattern with a more polished UI. ChatGPT's "Browse my files", Claude Projects, Cursor's
@codebase, Glean, Notion AI β every one of them. You now own the recipe.
Want a Variation? Pick a Flavour, Swap the PDF#
The recipe is the same; the data makes it interesting. Try one of these on your own:
- π Chat with the ConstitutionIndex the Indian Constitution PDF. Ask "what are the
Fundamental Rights?" with citations.
data: indiacode.nic.in - π Chat with an earnings callIndex TCS or Infosys' last quarterly transcript. Ask about
margins, guidance, hiring.
data: investor relations sites - π Chat with your textbookIndex a chapter. Ask exam-style questions. Suddenly: a
personalised tutor.
data: your bookshelf - π§Ύ Chat with company HR policyGenuinely useful at your workplace. "How many leaves do I
have left?" "What's the WFH policy?"
data: your HR portal
Where This Is Going#
π€ Agentic RAG β When Retrieval Becomes a Decision#
RAG works. But the frontier is making it smarter β and that's where this course leads next.
The RAG you just built always retrieves. But sometimes you don't need to (small talk), and sometimes one retrieval isn't enough (multi-hop questions). Agentic RAG = combine the LangChain agent loop with this RAG library. The agent decides: "do I need to look this up? Is what I found enough? Should I rewrite my query and try again?"
Query β retrieve β grade chunks β if bad, rewrite query, search again β if still bad, ask user for clarification β answer. Same agent loop you built with LangChain agents, with the RAG librarian as one of its tools.
π Advanced RAG β the Tricks the Senior Engineers Use#
- βοΈ HyDEHypothetical Document Embeddings: have the LLM imagine what the answer would look like, then search for chunks that match the imagined answer. Counter-intuitive, works surprisingly well.
- πͺ Step-back promptingBefore searching, ask the LLM to generalize the question. "What's the formula for compound interest in this case?" β "What is compound interest?" β broader, better retrieval.
- πΈοΈ Graph RAGBuild a knowledge graph from your docs (entities + relationships). Now you can answer "who reports to X" by walking the graph, not just searching.
- π RAG evaluationHow do you measure if your RAG is good? Three metrics: relevance, faithfulness, correctness. We'll wire up a real eval pipeline.
π Where you are right now. By now you've built a chatbot, an agent, and a RAG system. You understand the four pieces of every AI product β model Β· prompt Β· tool Β· retrieval. Almost everything ahead is recombining these four in cleverer ways. You're past the steep part of the curve.
- π EmbeddingsText β vectors. Similar meaning = nearby.
- βοΈ ChunkingThe unglamorous knob that decides quality.
- ποΈ Vector DBThree calls: add, query, delete.
- π― You shippedA real "Chat with your PDF". π