Two months ago I watched a "knowledgeable" support bot confidently invent a refund policy that didn't exist. The model was fine. The problem was that nobody had ever shown it the refund policy — so it improvised one. That's the entire point of RAG in one sentence: stop asking the model to know your documents. Give the model your documents at the moment it needs them.
RAG (Retrieval-Augmented Generation) is a pipeline of five stages: load your documents → chunk them into small pieces → embed each piece into a vector → store the vectors → and at question time, retrieve the most similar pieces and hand them to the model as context.
In n8n, all five stages are visual nodes — no code. The three details that decide whether your bot is useful or embarrassing: the chunk size (section 5), the embedding model consistency between ingest and query (section 6), and the tool description that tells the agent when to search (section 8). The full importable template is in section 9.
On this page
- What RAG Solves (And When You Don't Need It)
- The 5-Stage Pipeline
- The Two-Phase Architecture: Ingest vs Chat
- Stage 1: Load Documents
- Stage 2: Chunking (Size Guide)
- Stage 3: Embeddings
- Stage 4: Vector Stores Compared
- Stage 5: Retrieve & Generate
- Free: Importable RAG Chat Template
- 8 Common RAG Errors + Fixes
- FAQ
- Sources & References
1. What RAG Solves (And When You Don't Need It)
A language model knows what it learned during training — and nothing else. Your company's refund policy, your product manual from last month, your internal runbook: none of it exists inside the model. Two ways to change that:
| Approach | How it works | Cost & maintenance | Verdict |
|---|---|---|---|
| Fine-tuning | Retrains the model's weights on your data | Expensive, slow, and outdated the day your documents change | Rarely worth it for documents |
| RAG | Keeps your data external; retrieves relevant pieces per question | Cheap, instant, and updated the moment you re-index a file | The default for document Q&A |
When NOT to build RAG: if your documents fit inside the model's context window, just paste them — RAG is overhead. And if users ask only a handful of fixed questions, a few well-crafted prompts beat a retrieval pipeline. RAG earns its complexity when the corpus is large, changing, or too big for any prompt.
2. The 5-Stage Pipeline
Every RAG system on earth is these five stages. n8n has a node for each:
| Stage | n8n node | What it does |
|---|---|---|
| 1. Load | Default Data Loader | Reads PDFs, text, or binary files into documents |
| 2. Chunk | Token Splitter | Breaks documents into small, embeddable pieces |
| 3. Embed | Embeddings OpenAI (or any provider) | Converts each chunk into a numerical vector |
| 4. Store | Vector Store (Qdrant, Pinecone, Supabase, In-Memory…) | Indexes vectors for similarity search |
| 5. Retrieve | Vector Store (as tool) + AI Agent | Finds the top-k relevant chunks and answers from them |
Keep this mental image: stages 1–4 run once per document. Stage 5 runs once per question. That split shapes everything — the architecture, the costs, and most of the errors in section 10.
3. The Two-Phase Architecture: Ingest vs Chat
The #1 conceptual mistake in RAG guides is blurring two completely different flows into one diagram. They are separate — usually separate workflows:
- Ingest phase (stages 1–4): runs when a document is added or changed. Trigger: a file upload, a Google Drive watch, a manual run. Output: fresh vectors in the store.
- Chat phase (stage 5): runs per question. Trigger: a chat message. Output: an answer grounded in retrieved context.
Why the separation matters in practice: you can re-index a document without touching your live bot, and your bot answers instantly without re-reading every file. When both phases share one vector store (Qdrant, Pinecone, Supabase), they never need to share a workflow.
4. Stage 1: Load Documents
The Default Data Loader takes a file from a previous node (or directly) and turns it into document objects. In practice, its input comes from one of:
- Form Trigger — a user uploads a PDF through n8n's built-in form (the official starter template uses exactly this)
- Google Drive / S3 / FTP — watch a folder and index every new file automatically
- Manual trigger + binary data — for testing with a file you pick yourself
5. Stage 2: Chunking (The Size Guide Nobody Gives You)
The Token Splitter slices documents into chunks small enough to embed and retrieve meaningfully. chunkSize is the single most-tuned parameter in RAG — and the ranges below come from running this on real support docs:
| Chunk size (tokens) | What happens | Best for |
|---|---|---|
| 100–200 | Precise hits, but sentences get cut in half and context is lost | FAQ-style short answers |
| 300–500 ✅ | The sweet spot — one idea per chunk, enough context to answer | Most document Q&A. Start here |
| 800–1500 | Whole sections retrieved at once — but search signal dilutes and token cost per question climbs | Long-form answers, legal clauses |
| 3000+ | Effectively "search returns half the document" — retrieval stops meaning anything | Almost never |
The tuning loop: pick 300–500, then ask your bot five real questions from real users. If answers miss details, your chunks are too large (the detail is buried in a big block). If answers feel cut-off mid-thought, your chunks are too small. Adjust, re-index, repeat — three rounds usually settles it.
6. Stage 3: Embeddings (The Consistency Rule)
An embedding model converts each chunk into a vector — a long list of numbers encoding its meaning. Embeddings OpenAI with text-embedding-3-large is the common default, but every provider has one (Google, Cohere, Ollama for fully local).
7. Stage 4: Vector Stores Compared
| Store | Setup | Persistence | Best for |
|---|---|---|---|
| In-Memory Vector Store | Zero — it's built in | ❌ Resets on restart | First tests, prototypes (the honest truth: it will empty itself at 2 AM and confuse you) |
| Qdrant | One Docker container on your VPS | ✅ Persistent | Self-hosters — pairs perfectly with our Docker guide stack |
| Pinecone | Cloud account | ✅ Managed | Teams that don't want to run infrastructure |
| Supabase (pgvector) | Cloud or self-hosted Postgres | ✅ Persistent | When your docs already live in Postgres |
My honest path: prototype on In-Memory (zero friction), ship on Qdrant (one container, full control), and know that moving between them is just swapping the vector store node — the rest of the pipeline doesn't change.
8. Stage 5: Retrieve & Generate (The Two Dials)
At question time, a second Vector Store node runs in retrieve mode and is attached to the AI Agent as a tool. Two settings decide whether the agent uses it — and uses it well:
8.1 Top-k: how many chunks per question
- 2–3: focused answers, lowest token cost — risks missing context
- 4–6: the balanced default most production bots use ✅
- 10+: more context, noticeably more tokens per question — and more noise for the model to get lost in
8.2 The tool description — the setting nobody tells you about
The agent only calls a tool when it understands when to call it. A tool named vector_store_1 with an empty description gets ignored — the agent answers from its own training, which is exactly the hallucination you built RAG to prevent. The fix is one sentence:
"Use this tool to search our knowledge base for information about refunds,
shipping, and account policies. Always search it before answering product
or policy questions."
9. Free Template: Importable RAG Chat Workflow
The chat phase as a complete, verified workflow: Chat Trigger → AI Agent, with the Qdrant vector store attached as a retrieval tool, Window Buffer Memory for follow-ups, and the OpenAI chat model as the brain. All node types and connection directions below match real n8n exports — import, fill the three credentials, run the ingest phase once (section 9 notes), and chat.
{
"name": "RAG Chat with Your Documents",
"nodes": [
{
"parameters": {
"options": {}
},
"id": "chat-trigger",
"name": "Chat Trigger",
"type": "@n8n/n8n-nodes-langchain.chatTrigger",
"typeVersion": 1.1,
"position": [250, 300]
},
{
"parameters": {
"model": "gpt-4o-mini",
"options": {}
},
"id": "model",
"name": "OpenAI Chat Model",
"type": "@n8n/n8n-nodes-langchain.lmChatOpenAi",
"typeVersion": 1.1,
"position": [250, 520],
"credentials": {
"openAiApi": {
"id": "YOUR_OPENAI_CREDENTIAL_ID"
}
}
},
{
"parameters": {
"mode": "retrieve-as-tool",
"toolName": "knowledge_base_search",
"toolDescription": "Use this tool to search the knowledge base for information about products, policies, and documentation. Always search it before answering questions about our content.",
"qdrantCollection": {
"__rl": true,
"mode": "id",
"value": "=YOUR_QDRANT_COLLECTION_NAME"
},
"topK": 4,
"options": {}
},
"id": "retriever",
"name": "Knowledge Base (Qdrant)",
"type": "@n8n/n8n-nodes-langchain.vectorStoreQdrant",
"typeVersion": 1,
"position": [650, 520],
"credentials": {
"qdrantApi": {
"id": "YOUR_QDRANT_CREDENTIAL_ID"
}
}
},
{
"parameters": {
"sessionIdType": "customKey",
"sessionKey": "={{ $('Chat Trigger').item.json.sessionId }}",
"contextWindowLength": 20
},
"id": "memory",
"name": "Window Buffer Memory",
"type": "@n8n/n8n-nodes-langchain.memoryBufferWindow",
"typeVersion": 1.3,
"position": [850, 520]
},
{
"parameters": {
"options": {
"systemMessage": "You are a helpful assistant that answers questions using the knowledge base tool. Ground every answer in retrieved information and say when the knowledge base has no answer."
}
},
"id": "agent",
"name": "AI Agent",
"type": "@n8n/n8n-nodes-langchain.agent",
"typeVersion": 1.7,
"position": [450, 300]
}
],
"connections": {
"Chat Trigger": {
"main": [
[
{
"node": "AI Agent",
"type": "main",
"index": 0
}
]
]
},
"OpenAI Chat Model": {
"ai_languageModel": [
[
{
"node": "AI Agent",
"type": "ai_languageModel",
"index": 0
}
]
]
},
"Window Buffer Memory": {
"ai_memory": [
[
{
"node": "AI Agent",
"type": "ai_memory",
"index": 0
}
]
]
},
"Knowledge Base (Qdrant)": {
"ai_tool": [
[
{
"node": "AI Agent",
"type": "ai_tool",
"index": 0
}
]
]
}
},
"active": false,
"settings": {
"executionOrder": "v1",
"saveManualExecutions": true
}
}
- Three credentials: OpenAI, Qdrant, and a Qdrant collection name — replace the three
YOUR_*placeholders. - The ingest side: build the four-node chain from sections 4–7 once —
Default Data Loader(fed by a Form Trigger or Drive) →Token Splitter(chunkSize300–500) →Embeddings OpenAI→ a second Qdrant node in insert mode with the same collection name. Run it per new document; the chat side picks changes up automatically. - Zero-infra first test: swap both Qdrant nodes for In-Memory Vector Store (
vectorStoreInMemory) — identical wiring, nothing to install, and it empties on restart (see error #7). - Honest version note: AI-node type versions drift between n8n releases. If the import complains, re-add the five nodes from the panel and wire them the same way — the settings above are the part that matters.
- Want this bot in Telegram? Our Telegram bot guide covers the messenger side — swap the Chat Trigger and you have a grounded bot answering in any chat.
10. The 8 Most Common RAG Errors & Fixes
| # | Error / Symptom | Cause | Fix |
|---|---|---|---|
| 1 | Retriever returns nothing / "no documents found" | Store is empty or the collection name is wrong | Run the ingest phase once and confirm the collection name matches on both nodes |
| 2 | Garbage answers that look relevant | Embedding model mismatch between ingest and query | Same provider + same model name on both sides (section 6) |
| 3 | Agent answers from its own knowledge, ignores the docs | Tool description missing or too vague | Name the tool clearly and describe when to use it (section 8.2) |
| 4 | Token bill climbs fast | Top-k too high or chunks too large | Cut top-k to 4–6 and chunk size to 300–500 (sections 5 & 8) |
| 5 | Answers miss details users need | Chunks too large — the detail is buried inside a huge block | Shrink chunk size and re-index; test with real questions |
| 6 | Duplicate or stale answers after re-indexing | Old vectors never removed from the store | Clear the store before re-ingesting (the clearStore option) or delete old points by file ID |
| 7 | Bot forgot everything after a restart | You're on the In-Memory Vector Store — it resets by design | Move to Qdrant, Pinecone, or Supabase (section 7) |
| 8 | Switched embedding providers, everything broke | Vectors from different providers are incompatible | Re-index all documents with the new model — there is no migration shortcut |
11. Frequently Asked Questions
Retrieval-Augmented Generation. You index your own documents into a vector store, and every user question first retrieves the most relevant chunks, which are then passed to the language model as context. The model answers from your data instead of guessing (section 2).
No. The whole pipeline — loading, chunking, embedding, storing, and retrieving — is visual in n8n with dedicated nodes. Code becomes optional, only for advanced logic.
Start with the In-Memory Vector Store — zero setup, perfect for testing. It resets on restart, so for production switch to Qdrant (self-hosted via Docker), Pinecone, or Supabase (section 7).
Start with 300–500 tokens per chunk. Larger chunks dilute the search signal and cost more tokens per retrieval; smaller chunks lose context and cut sentences in half. Tune by testing real questions from your users (section 5).
Usually the tool description is too vague, so the agent doesn't know when to use it. Name the tool clearly and describe when to call it — or the agent answers from its own training instead of your data (section 8.2).
The classic cause is a mismatch between the embedding model used at ingest and the one used at query time. Vectors from different models are not comparable. Use the same model on both sides (section 6).
Yes, but predictably: embeddings cost once per chunk at ingest time (tiny), and each question costs one embedding call plus the retrieved context in tokens. Top-k directly controls the per-question cost (section 8.1).
Fine-tuning retrains the model weights on your data — expensive and static. RAG keeps your data external and retrieves it per question — cheap, instantly updatable, and answers cite retrievable sources.
Yes — and you should. Window Buffer Memory keeps the recent conversation flowing, while the vector store supplies document knowledge. The agent uses both: chat history plus retrieved context. The template in section 9 wires both.
Sources & References
- n8n Blog — Build RAG pipelines in n8n
- n8n Templates — PDF RAG with OpenAI + Pinecone + Cohere reranking
- n8n Templates — Local RAG with Ollama + Qdrant
- n8n Documentation — RAG examples
- n8n Documentation — Vector stores
