📁 last Posts

n8n RAG Tutorial: Chat With Your Documents (2026 Complete Guide)

n8n RAG Tutorial: Chat With Your Documents (2026 Complete Guide)

Two months ago I watched a "knowledgeable" support bot confidently invent a refund policy that didn't exist. The model was fine. The problem was that nobody had ever shown it the refund policy — so it improvised one. That's the entire point of RAG in one sentence: stop asking the model to know your documents. Give the model your documents at the moment it needs them.

The Direct Answer

RAG (Retrieval-Augmented Generation) is a pipeline of five stages: load your documents → chunk them into small pieces → embed each piece into a vector → store the vectors → and at question time, retrieve the most similar pieces and hand them to the model as context.

In n8n, all five stages are visual nodes — no code. The three details that decide whether your bot is useful or embarrassing: the chunk size (section 5), the embedding model consistency between ingest and query (section 6), and the tool description that tells the agent when to search (section 8). The full importable template is in section 9.


1. What RAG Solves (And When You Don't Need It)

A language model knows what it learned during training — and nothing else. Your company's refund policy, your product manual from last month, your internal runbook: none of it exists inside the model. Two ways to change that:

ApproachHow it worksCost & maintenanceVerdict
Fine-tuning Retrains the model's weights on your data Expensive, slow, and outdated the day your documents change Rarely worth it for documents
RAG Keeps your data external; retrieves relevant pieces per question Cheap, instant, and updated the moment you re-index a file The default for document Q&A

When NOT to build RAG: if your documents fit inside the model's context window, just paste them — RAG is overhead. And if users ask only a handful of fixed questions, a few well-crafted prompts beat a retrieval pipeline. RAG earns its complexity when the corpus is large, changing, or too big for any prompt.


2. The 5-Stage Pipeline

Every RAG system on earth is these five stages. n8n has a node for each:

Stagen8n nodeWhat it does
1. LoadDefault Data LoaderReads PDFs, text, or binary files into documents
2. ChunkToken SplitterBreaks documents into small, embeddable pieces
3. EmbedEmbeddings OpenAI (or any provider)Converts each chunk into a numerical vector
4. StoreVector Store (Qdrant, Pinecone, Supabase, In-Memory…)Indexes vectors for similarity search
5. RetrieveVector Store (as tool) + AI AgentFinds the top-k relevant chunks and answers from them

Keep this mental image: stages 1–4 run once per document. Stage 5 runs once per question. That split shapes everything — the architecture, the costs, and most of the errors in section 10.

ℹ️
New to AI agents in n8n? Our How to Build an AI Agent with n8n guide covers the agent basics this article builds on.

3. The Two-Phase Architecture: Ingest vs Chat

The #1 conceptual mistake in RAG guides is blurring two completely different flows into one diagram. They are separate — usually separate workflows:

  • Ingest phase (stages 1–4): runs when a document is added or changed. Trigger: a file upload, a Google Drive watch, a manual run. Output: fresh vectors in the store.
  • Chat phase (stage 5): runs per question. Trigger: a chat message. Output: an answer grounded in retrieved context.

Why the separation matters in practice: you can re-index a document without touching your live bot, and your bot answers instantly without re-reading every file. When both phases share one vector store (Qdrant, Pinecone, Supabase), they never need to share a workflow.


4. Stage 1: Load Documents

The Default Data Loader takes a file from a previous node (or directly) and turns it into document objects. In practice, its input comes from one of:

  • Form Trigger — a user uploads a PDF through n8n's built-in form (the official starter template uses exactly this)
  • Google Drive / S3 / FTP — watch a folder and index every new file automatically
  • Manual trigger + binary data — for testing with a file you pick yourself
💡
Start with one file. Ten files × wrong chunk size = ten times the confusion. Get one PDF answering questions well end-to-end, then automate the folder watch. Every serious RAG I've shipped started as a one-file test.

5. Stage 2: Chunking (The Size Guide Nobody Gives You)

The Token Splitter slices documents into chunks small enough to embed and retrieve meaningfully. chunkSize is the single most-tuned parameter in RAG — and the ranges below come from running this on real support docs:

Chunk size (tokens)What happensBest for
100–200 Precise hits, but sentences get cut in half and context is lost FAQ-style short answers
300–500 ✅ The sweet spot — one idea per chunk, enough context to answer Most document Q&A. Start here
800–1500 Whole sections retrieved at once — but search signal dilutes and token cost per question climbs Long-form answers, legal clauses
3000+ Effectively "search returns half the document" — retrieval stops meaning anything Almost never

The tuning loop: pick 300–500, then ask your bot five real questions from real users. If answers miss details, your chunks are too large (the detail is buried in a big block). If answers feel cut-off mid-thought, your chunks are too small. Adjust, re-index, repeat — three rounds usually settles it.


6. Stage 3: Embeddings (The Consistency Rule)

An embedding model converts each chunk into a vector — a long list of numbers encoding its meaning. Embeddings OpenAI with text-embedding-3-large is the common default, but every provider has one (Google, Cohere, Ollama for fully local).

⚠️
The rule that prevents garbage answers: the embedding model used at query time must be the exact same model used at ingest time. Vectors from different models live in different mathematical spaces — comparing them returns nonsense that looks deceptively plausible. Same provider, same model name, both phases. Always.

7. Stage 4: Vector Stores Compared

StoreSetupPersistenceBest for
In-Memory Vector Store Zero — it's built in ❌ Resets on restart First tests, prototypes (the honest truth: it will empty itself at 2 AM and confuse you)
Qdrant One Docker container on your VPS ✅ Persistent Self-hosters — pairs perfectly with our Docker guide stack
Pinecone Cloud account ✅ Managed Teams that don't want to run infrastructure
Supabase (pgvector) Cloud or self-hosted Postgres ✅ Persistent When your docs already live in Postgres

My honest path: prototype on In-Memory (zero friction), ship on Qdrant (one container, full control), and know that moving between them is just swapping the vector store node — the rest of the pipeline doesn't change.


8. Stage 5: Retrieve & Generate (The Two Dials)

At question time, a second Vector Store node runs in retrieve mode and is attached to the AI Agent as a tool. Two settings decide whether the agent uses it — and uses it well:

8.1 Top-k: how many chunks per question

  • 2–3: focused answers, lowest token cost — risks missing context
  • 4–6: the balanced default most production bots use ✅
  • 10+: more context, noticeably more tokens per question — and more noise for the model to get lost in

8.2 The tool description — the setting nobody tells you about

The agent only calls a tool when it understands when to call it. A tool named vector_store_1 with an empty description gets ignored — the agent answers from its own training, which is exactly the hallucination you built RAG to prevent. The fix is one sentence:

"Use this tool to search our knowledge base for information about refunds,
shipping, and account policies. Always search it before answering product
or policy questions."
✅
Add memory and the bot becomes conversational: a Window Buffer Memory node next to the retriever keeps follow-ups natural ("and what about refunds for that?"). The full memory setup is in our n8n AI Agent Memory guide — the two articles together build a complete grounded, remembering bot.

9. Free Template: Importable RAG Chat Workflow

The chat phase as a complete, verified workflow: Chat Trigger → AI Agent, with the Qdrant vector store attached as a retrieval tool, Window Buffer Memory for follow-ups, and the OpenAI chat model as the brain. All node types and connection directions below match real n8n exports — import, fill the three credentials, run the ingest phase once (section 9 notes), and chat.

{
  "name": "RAG Chat with Your Documents",
  "nodes": [
    {
      "parameters": {
        "options": {}
      },
      "id": "chat-trigger",
      "name": "Chat Trigger",
      "type": "@n8n/n8n-nodes-langchain.chatTrigger",
      "typeVersion": 1.1,
      "position": [250, 300]
    },
    {
      "parameters": {
        "model": "gpt-4o-mini",
        "options": {}
      },
      "id": "model",
      "name": "OpenAI Chat Model",
      "type": "@n8n/n8n-nodes-langchain.lmChatOpenAi",
      "typeVersion": 1.1,
      "position": [250, 520],
      "credentials": {
        "openAiApi": {
          "id": "YOUR_OPENAI_CREDENTIAL_ID"
        }
      }
    },
    {
      "parameters": {
        "mode": "retrieve-as-tool",
        "toolName": "knowledge_base_search",
        "toolDescription": "Use this tool to search the knowledge base for information about products, policies, and documentation. Always search it before answering questions about our content.",
        "qdrantCollection": {
          "__rl": true,
          "mode": "id",
          "value": "=YOUR_QDRANT_COLLECTION_NAME"
        },
        "topK": 4,
        "options": {}
      },
      "id": "retriever",
      "name": "Knowledge Base (Qdrant)",
      "type": "@n8n/n8n-nodes-langchain.vectorStoreQdrant",
      "typeVersion": 1,
      "position": [650, 520],
      "credentials": {
        "qdrantApi": {
          "id": "YOUR_QDRANT_CREDENTIAL_ID"
        }
      }
    },
    {
      "parameters": {
        "sessionIdType": "customKey",
        "sessionKey": "={{ $('Chat Trigger').item.json.sessionId }}",
        "contextWindowLength": 20
      },
      "id": "memory",
      "name": "Window Buffer Memory",
      "type": "@n8n/n8n-nodes-langchain.memoryBufferWindow",
      "typeVersion": 1.3,
      "position": [850, 520]
    },
    {
      "parameters": {
        "options": {
          "systemMessage": "You are a helpful assistant that answers questions using the knowledge base tool. Ground every answer in retrieved information and say when the knowledge base has no answer."
        }
      },
      "id": "agent",
      "name": "AI Agent",
      "type": "@n8n/n8n-nodes-langchain.agent",
      "typeVersion": 1.7,
      "position": [450, 300]
    }
  ],
  "connections": {
    "Chat Trigger": {
      "main": [
        [
          {
            "node": "AI Agent",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "OpenAI Chat Model": {
      "ai_languageModel": [
        [
          {
            "node": "AI Agent",
            "type": "ai_languageModel",
            "index": 0
          }
        ]
      ]
    },
    "Window Buffer Memory": {
      "ai_memory": [
        [
          {
            "node": "AI Agent",
            "type": "ai_memory",
            "index": 0
          }
        ]
      ]
    },
    "Knowledge Base (Qdrant)": {
      "ai_tool": [
        [
          {
            "node": "AI Agent",
            "type": "ai_tool",
            "index": 0
          }
        ]
      ]
    }
  },
  "active": false,
  "settings": {
    "executionOrder": "v1",
    "saveManualExecutions": true
  }
}
  • Three credentials: OpenAI, Qdrant, and a Qdrant collection name — replace the three YOUR_* placeholders.
  • The ingest side: build the four-node chain from sections 4–7 once — Default Data Loader (fed by a Form Trigger or Drive) → Token Splitter (chunkSize 300–500) → Embeddings OpenAI → a second Qdrant node in insert mode with the same collection name. Run it per new document; the chat side picks changes up automatically.
  • Zero-infra first test: swap both Qdrant nodes for In-Memory Vector Store (vectorStoreInMemory) — identical wiring, nothing to install, and it empties on restart (see error #7).
  • Honest version note: AI-node type versions drift between n8n releases. If the import complains, re-add the five nodes from the panel and wire them the same way — the settings above are the part that matters.
  • Want this bot in Telegram? Our Telegram bot guide covers the messenger side — swap the Chat Trigger and you have a grounded bot answering in any chat.

10. The 8 Most Common RAG Errors & Fixes

#Error / SymptomCauseFix
1 Retriever returns nothing / "no documents found" Store is empty or the collection name is wrong Run the ingest phase once and confirm the collection name matches on both nodes
2 Garbage answers that look relevant Embedding model mismatch between ingest and query Same provider + same model name on both sides (section 6)
3 Agent answers from its own knowledge, ignores the docs Tool description missing or too vague Name the tool clearly and describe when to use it (section 8.2)
4 Token bill climbs fast Top-k too high or chunks too large Cut top-k to 4–6 and chunk size to 300–500 (sections 5 & 8)
5 Answers miss details users need Chunks too large — the detail is buried inside a huge block Shrink chunk size and re-index; test with real questions
6 Duplicate or stale answers after re-indexing Old vectors never removed from the store Clear the store before re-ingesting (the clearStore option) or delete old points by file ID
7 Bot forgot everything after a restart You're on the In-Memory Vector Store — it resets by design Move to Qdrant, Pinecone, or Supabase (section 7)
8 Switched embedding providers, everything broke Vectors from different providers are incompatible Re-index all documents with the new model — there is no migration shortcut
ℹ️
A RAG bot failing silently is worse than no bot at all. Route its errors through an Error Trigger workflow — the n8n Error Handling guide builds the complete alerting setup.

11. Frequently Asked Questions

What is RAG in n8n?

Retrieval-Augmented Generation. You index your own documents into a vector store, and every user question first retrieves the most relevant chunks, which are then passed to the language model as context. The model answers from your data instead of guessing (section 2).

Do I need to write code to build RAG in n8n?

No. The whole pipeline — loading, chunking, embedding, storing, and retrieving — is visual in n8n with dedicated nodes. Code becomes optional, only for advanced logic.

What is the best vector store for a beginner in n8n?

Start with the In-Memory Vector Store — zero setup, perfect for testing. It resets on restart, so for production switch to Qdrant (self-hosted via Docker), Pinecone, or Supabase (section 7).

What chunk size should I use for RAG?

Start with 300–500 tokens per chunk. Larger chunks dilute the search signal and cost more tokens per retrieval; smaller chunks lose context and cut sentences in half. Tune by testing real questions from your users (section 5).

Why does my AI agent ignore the documents?

Usually the tool description is too vague, so the agent doesn't know when to use it. Name the tool clearly and describe when to call it — or the agent answers from its own training instead of your data (section 8.2).

Why do I get garbage answers from my RAG bot?

The classic cause is a mismatch between the embedding model used at ingest and the one used at query time. Vectors from different models are not comparable. Use the same model on both sides (section 6).

Does RAG cost money to run in n8n?

Yes, but predictably: embeddings cost once per chunk at ingest time (tiny), and each question costs one embedding call plus the retrieved context in tokens. Top-k directly controls the per-question cost (section 8.1).

What is the difference between RAG and fine-tuning?

Fine-tuning retrains the model weights on your data — expensive and static. RAG keeps your data external and retrieves it per question — cheap, instantly updatable, and answers cite retrievable sources.

Can I combine memory and RAG in the same agent?

Yes — and you should. Window Buffer Memory keeps the recent conversation flowing, while the vector store supplies document knowledge. The agent uses both: chat history plus retrieved context. The template in section 9 wires both.


Sources & References

Explore More Guides

Ahmed Ayari — TriggerWorkflow The chunk-size table in section 5 is not theory — it's the result of re-indexing support docs three times until real questions stopped getting half-answers. And error #7 is personal: the night the in-memory store emptied itself at 2 AM taught me more about vector persistence than any guide ever did. The template in section 9 is the exact pattern our own docs bot runs on. Read more guides by Mr.Ayari
Comments