
Retrieval-augmented generation (RAG) is a way to make a large language model answer questions using your own business information. When a user asks something, the system first retrieves the most relevant passages from your documents or databases, then passes them to the model, which writes an answer grounded in those passages, usually with citations. RAG is the standard approach for company knowledge assistants, support bots and document Q&A because it keeps answers current, traceable and permission-aware without retraining a model.
This guide explains how RAG works, when to use it instead of fine-tuning, what a production architecture looks like, and how to evaluate and secure it.
Why Do Businesses Need RAG?
A language model on its own only knows what was in its training data. It doesn't know your pricing, your policies, last week's product update or the contract you signed yesterday. Ask it anyway and it may produce a confident, plausible and wrong answer.
RAG solves three problems at once:
- Freshness. Update a document and the next answer reflects it. No retraining.
- Traceability. Answers cite the passages they came from, so users and auditors can check them.
- Access control. Retrieval can respect who is allowed to see what, which a model's internal knowledge cannot.
How Does RAG Work?
RAG has two pipelines: an ingestion pipeline that prepares your content, and a query pipeline that answers questions.
The ingestion pipeline
- Connect sources. Pull content from shared drives, wikis, helpdesks, CRMs, databases and PDFs.
- Parse and clean. Extract text, tables and structure. Scanned documents need OCR or a multimodal model.
- Chunk. Split documents into passages small enough to retrieve precisely but large enough to keep context. Splitting along headings and sections works better than fixed character counts.
- Enrich. Attach metadata: source, title, date, department, and the access permissions of the original document.
- Embed and index. Convert each chunk into a vector embedding and store it in a search index, alongside a keyword index.
The query pipeline
- Understand the question. Optionally rewrite it, expand abbreviations or split a complex question into parts.
- Retrieve. Run hybrid search (vector similarity plus keyword search such as BM25), filtered by the user's permissions and metadata like date or product.
- Rerank. A reranking model reorders the top results by true relevance, so the best passages come first.
- Generate. The LLM receives the question plus the top passages and instructions to answer only from them, cite sources and say when it doesn't know.
- Check and return. Optional checks verify the answer is supported by the sources before it reaches the user.
RAG vs Fine-Tuning vs Long Context
RAG, fine-tuning and long-context prompting solve different problems, and they can be combined.
| RAG | Fine-tuning | Long context (paste everything) | |
|---|---|---|---|
| What it changes | What the model sees at question time | The model's weights and behavior | What the model sees, all at once |
| Best for | Company facts that change, Q&A over many documents | Consistent style, format or a narrow repeated task | A few documents, one-off analysis |
| Keeping knowledge current | Update the index | Retrain | Re-send documents each time |
| Citations | Natural | Hard | Possible, less precise |
| Per-user permissions | Yes, at retrieval | No | Only if you filter manually |
| Cost profile | Indexing plus modest per-query tokens | Training runs, then cheaper inference for that task | High tokens per query at scale |
| Main risk | Poor retrieval leads to poor answers | Outdated knowledge, harder to audit | Cost, latency, details lost in long inputs |
The practical rule: use RAG for knowledge, fine-tuning for behavior. Long context windows in 2026 models are large and useful for analyzing a handful of documents, but at business scale RAG is usually cheaper, faster and easier to govern.
What Does a Production RAG Architecture Look Like?
A production RAG system has more parts than a prototype. The main components are:
- Connectors that sync from each source on a schedule or on change, and remove deleted content.
- Document processing for parsing, OCR, table extraction and chunking.
- Embedding model (hosted or self-hosted) and an index that supports vector and keyword search. Common choices include PostgreSQL with pgvector, OpenSearch, Elasticsearch, Qdrant, Weaviate and Pinecone.
- Reranker to improve the ordering of results.
- LLM with a prompt that enforces grounding and citations.
- Permission layer that maps each user to the documents they may see.
- Application layer: chat UI, API, Slack or Teams bot, or a voice agent.
- Observability: logs of queries, retrieved passages, answers, feedback, latency and cost.
- Evaluation harness that runs test sets on every change.
Beyond basic RAG
Two patterns are increasingly common. Agentic RAG lets an AI agent decide which sources to search, run several searches and combine results, which helps with multi-part questions. Structured retrieval queries databases or APIs directly (for example, "orders over a value last month") instead of searching text, because some questions are better answered by SQL than by similarity search. Knowledge-graph approaches such as GraphRAG help when answers depend on relationships across many documents.
How Do You Evaluate a RAG System?
Evaluate retrieval and generation separately, using a test set built from real user questions. If you only score final answers, you won't know whether a failure came from search or from the model.
Retrieval metrics:
- Recall at k. Of the passages needed to answer, how many appear in the top k results?
- Precision and ranking. Are the relevant passages near the top?
Answer metrics:
- Faithfulness (groundedness). Is every claim supported by the retrieved passages?
- Correctness and completeness. Does it match the reference answer?
- Citation accuracy. Do the citations actually support the statements?
- Appropriate refusal. Does it say "I don't know" when the answer isn't in the sources?
Teams often use an LLM-as-judge to score faithfulness and relevance at scale, with frameworks such as Ragas, then spot-check with human reviewers. Re-run the full set whenever you change chunking, the embedding model, prompts or the LLM.
How Do You Keep RAG Secure?
RAG connects a model to your internal data, so security has to be designed in, not added later.
- Permission-aware retrieval. Filter results by the user's access rights before anything reaches the model. Never rely on the prompt to hide documents.
- Data minimization. Exclude or redact sensitive fields (personal data, credentials, salaries) at ingestion when they aren't needed.
- Prompt-injection defense. Retrieved documents, emails and web pages can contain text that tries to instruct the model. Treat retrieved content as data, constrain what the model can do, and don't give a RAG assistant write tools it doesn't need.
- Provider terms and residency. Confirm that your model and embedding providers don't train on your data, and that data is processed in regions you're allowed to use.
- Audit logging. Record who asked what, what was retrieved and what was answered.
Common RAG Failure Modes
- Bad parsing. Tables and scanned PDFs turned into garbage text.
- Poor chunking. Answers split across chunks, or chunks without the heading that gives them meaning.
- Stale or duplicate content. Old policy versions outranking new ones. Use dates and source-of-truth rules.
- Vector-only search. Missing exact matches on product codes, names and IDs. Add keyword search.
- No "I don't know". The model fills gaps with guesses. Instruct and test for refusals.
Getting Started With RAG
Start with one well-defined knowledge domain, such as support articles or internal policies, a test set of real questions and clear access rules. Prove accuracy there, then add sources. RAG is the foundation of most of the generative AI use cases businesses deploy today, and it is often the first step before an AI agent that can also act on what it finds.
Aaga's generative AI development team builds RAG systems with permission-aware retrieval, evaluation from day one and integrations into the tools your team already uses, including AI chatbots on your website or messaging channels. Talk to us about your documents and use case.

