Skip to content
RAG Explained Better

What Is Retrieval-Augmented Generation? RAG Explained

RAG answers from documents fetched at query time instead of weights alone. What each stage does, what it costs, and what it does not fix.

Retrieval-augmented generation is a technique that fetches documents at query time and hands them to a language model as context, so the answer is grounded in an external knowledge source rather than in the model’s frozen training weights alone. Patrick Lewis and colleagues named the method in “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” accepted at NeurIPS 2020, combining a pre-trained seq2seq model with a dense Wikipedia index. Production systems still retrieve, then generate: documents are chunked, embedded, and indexed once, then each question retrieves the top-k chunks and the model writes from the augmented prompt. A language model cannot see private or current documents and invents answers when it does not know; fetching the source text reduces those failures without retraining. The lookup store is a vector database such as Weaviate, Pinecone, Qdrant, or Elasticsearch, and every pipeline stage can be measured, and every stage can break.

How does RAG work?

RAG works in two phases. Index-time, done once: documents are ingested, split into chunks, embedded into vectors, and stored in a search index. Query-time, on every question: the question is embedded, used to retrieve the top-k most similar chunks from that index, and sent together with the question to a language model that generates an answer from that augmented prompt. AWS describes the same four moves as create external data, retrieve relevant information, augment the LLM prompt, and update the external data when it changes. IBM names the four components that carry those moves: a knowledge base, a retriever, an integration layer that stuffs the retrieved text into the prompt, and a generator. Mary Newhauser’s “Introduction to LLM RAG – Retrieval Augmented Generation Explained” on the Weaviate blog (15 October 2024) describes the same two stages, ingestion and inference, then extends them to agentic RAG, graph RAG, and evaluation.

The RAG pipeline: a query is embedded and used to retrieve chunks from an index built by ingesting and chunking documents; the retrieved context and query are sent to a language model to generate an answer. Evaluation measures the retrieve and generate stages; failures can occur at retrieval and generation.
Index-time builds the library. Query-time looks a passage up and writes from it. Evaluation scores those two query-time stages separately; the two most common breaks, wrong retrieval and ungrounded generation, are marked on the retrieve and generate boxes. A paste-and-run build of the same pipeline is at RAG pipeline build tutorial.

Weaviate’s retrieval-augmented generation documentation (as of September 2026) describes a RAG query as two parts in one request: a search query, and a prompt for a generative model. Weaviate first performs the search, then passes both the search results and the prompt to the model before returning the generated response. The search can be similarity, keyword, or hybrid, with filters, which means the retrieval half is not locked to dense vectors. Generation itself is offered in two shapes: a single prompt that interpolates object properties and returns one generated string per retrieved object, and a grouped task that returns one response over the whole result set. Runtime choice of generative provider was added in Weaviate v1.30, so a collection can keep a default model and override it per query.

That combined query is a convenience, not a change in the underlying job. The generator still only sees what retrieval returned, so a miss at search time becomes a confident wrong answer at generation time. Understanding the pipeline as retrieve-then-generate raises a historical question that most product pages skip: the 2020 paper that named the method was not this production pipeline, and the difference decides what “RAG” is allowed to mean.

Who named retrieval-augmented generation, and what did the 2020 paper build?

Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela named retrieval-augmented generation in a 2020 paper accepted at NeurIPS, “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks” (arXiv 2005.11401v4). The abstract calls RAG a general-purpose fine-tuning recipe for models that combine pre-trained parametric memory with non-parametric memory for language generation. The parametric memory is a pre-trained seq2seq model. The non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. NVIDIA’s 2023 explainer, updated later, quotes Lewis apologising for the acronym and locating the work at University College London and the then Facebook AI Research London lab, with Perez and Kiela then of New York University and Facebook AI Research.

IBM Technology’s YouTube video “What is Retrieval-Augmented Generation (RAG)?” walks through retrieve-then-generate in the same terms used here.

The paper compares two formulations. One conditions on the same retrieved passages across the whole generated sequence. The other can use different passages per token. The authors report state-of-the-art on three open-domain question-answering tasks, outperforming parametric seq2seq models and task-specific retrieve-and-extract architectures, and they find that on language generation tasks the RAG models produce more specific, diverse, and factual language than a parametric-only seq2seq baseline. Those claims are the paper’s own results, dated to the NeurIPS 2020 version of the work, not a transferrable score for a customer-support chatbot in 2026.

Wikipedia’s RAG entry (retrieved 17 September 2026) records the same origin: the term was introduced in a 2020 paper that combined a parametric language model with a non-parametric external memory accessed through retrieval at inference time. What changed after 2020 is the surrounding stack. Production RAG usually does not fine-tune the generator at all. It chunks a private corpus, embeds the chunks, indexes them in a vector store, retrieves at query time, and prompts a frozen LLM. Weaviate’s starter guide on RAG (as of September 2026) frames that same pairing, search then prompt, as a single query against a collection, with the language model supplied by a generative integration rather than trained jointly with the retriever. The 2020 recipe still explains why the name stuck. The production pipeline explains why retrieval in RAG is the stage the rest of the architecture is named after, and why a miss there cannot be repaired by a better prompt alone.

Why use RAG, and what problem does it solve?

RAG solves an access problem that a frozen language model cannot solve from its weights. AWS lists the failure modes that follow from that freeze: the model presents false information when it does not have the answer, presents out-of-date or generic information when the user expects a current one, creates a response from non-authoritative sources, and confuses terminology when different training sources used the same words for different things. IBM’s RAG explainer, updated 24 July 2026, puts the same limit in training terms: foundation-model training sets are finite and public, so internal organisational data, scholarly journals, and specialised datasets sit outside what the model can see unless they are fetched.

Weaviate’s RAG starter guide (as of September 2026) names the same two limits: language models can confidently produce incorrect or outdated information, and they might simply not be trained on the information you need. RAG’s remedy, in that guide, is the two-step process already described: retrieve relevant data, then prompt the model with the retrieved data plus the user query, so the model uses up-to-date context rather than recall from training. Pinecone’s RAG guide frames the same gap as four foundation-model limits: knowledge cutoffs, missing domain depth, no private data, and probabilistic output that loses trust when it cannot cite a source.

The practical consequence is that adding a document to the index makes it answerable without a training run. NVIDIA reports that Lewis and co-authors described an implementation on Hugging Face in as few as five lines of code, and treats that cheapness as the reason RAG spread faster than retraining. Cheap is not free: every query now pays for retrieval, and the answer is only as good as the chunk that came back. That is why the benefits and the limits belong in the same account, not in a list of wins with the failures postponed.

What are RAG’s benefits and limitations?

RAG’s benefits are real, and they are the reason the architecture shipped. IBM lists cost-efficient implementation without full retraining, access to current and domain-specific data, lower risk of hallucination, increased user trust, expanded use cases, more developer control over sources, and greater data security because sensitive corpora can stay in an index the model queries rather than in a training set. AWS groups the same wins as cost-effective implementation, current information, enhanced user trust through citations, and more developer control over which sources the model may see. Azure’s RAG glossary adds the citation trail as the mechanism behind that trust: retrieved chunks are the sources a reader can check.

The limitations sit on the same pipeline. Retrieval can fetch the wrong chunk, and the generator will answer from it with the same confidence it would have used on the right one; that failure is why RAG returns the wrong chunk. Wikipedia, citing Ars Technica, states the limit without softening it: RAG is not a direct solution, because the language model can still hallucinate around the source material in its response. IBM notes that RAG reduces the need for frequent retraining but does not remove it entirely, and that models may still generate answers when they should indicate uncertainty. Databricks’ RAG explainer lists implementation challenges on the other side of the ledger: chunking, embedding quality, index freshness, and evaluation. Latency and infrastructure are structural, not bugs: a vector store, an embedding call, and a retrieval hop now sit in front of every answer.

The honest summary is that RAG is an access fix, not an intelligence fix. It supplies knowledge the weights do not hold. It does not teach a weak reasoner to reason, and it does not make a wrong retrieved passage into a right answer. Whether a given system is any good is a measurement question, which is why how to evaluate a RAG system is the companion cluster to the definition. The sharpest form of that measurement question is the one users actually type: does RAG stop hallucinations, or only move them?

Does RAG stop hallucinations?

RAG reduces ungrounded hallucinations. It does not stop them. Giving the model retrieved text to quote cuts down guesses invented from the weights, which is the failure AWS, IBM, NVIDIA, and Pinecone all name as the reason to fetch sources. The residual failure is different: a wrong, stale, or out-of-context chunk produces a wrong answer that looks cited. Wikipedia’s RAG entry uses MIT Technology Review’s example of an AI-generated claim that the United States has had one Muslim president, Barack Hussein Obama, retrieved from the rhetorical chapter title “Barack Hussein Obama: America’s First Muslim President?” in Faith in the New Millennium. The retriever found a real source. The generator did not read the question mark.

That pattern is why “grounded” and “correct” are not the same metric. A faithfulness score asks whether the answer is supported by the retrieved context. It does not ask whether that context was the right document, still current, or used with its original meaning intact. Wikipedia also records that models with RAG are often programmed to prioritise the new information (“prompt stuffing”), which helps when the retrieved passage is right and hurts when it is not. IBM’s limitation note is the operational version of the same fact: without specific training, a model may generate an answer even when it should say it does not know.

The detection signature is a wrong answer whose retrieved chunks do not contain the claim, or contain it only by misreading. When the chunks are wrong, the fix is retrieval, chunking, or index freshness, not a sterner generate prompt. When the chunks are right and the model still invents, the fix is grounding, refusal, or a judge on the answer. Both failures live in why RAG systems fail, sorted by the stage that caused them. Knowing that RAG is not an anti-hallucination switch is also what keeps the next comparison honest: fine-tuning is not that switch either, and the two methods are not substitutes.

Is RAG the same as fine-tuning?

RAG is not the same as fine-tuning. IBM’s RAG explainer states the difference in one pair of jobs: RAG lets a language model query an external data source, while fine-tuning trains the model on domain-specific data. Both aim to make the model perform better in a specified domain. They are often contrasted and can be used together. Fine-tuning increases familiarity with the intended domain and output requirements. RAG assists the model in generating relevant, high-quality outputs from sources that were not baked into the weights.

Azure’s RAG glossary asks the same question as “What is the difference between RAG and LLM?” and answers it as a division of labour: the language model generates, RAG supplies the documents the generation should stick to. The comparison practitioners type is RAG versus fine-tuning, and when to use each. If the answer depends on data that changes or is private, fetch it at query time. If the model must act differently, in tone, format, or a skill that is not a lookup, change the weights. The one-line test is knowledge versus behaviour. The scored either/or/both call, with cost on the same table, is at RAG vs fine-tuning.

Prompt engineering is a third lever, not a fourth architecture. A prompt can tell the model to stay inside the retrieved context. It cannot fetch a document the retriever never returned. LoRA and other parameter-efficient fine-tunes sit on the fine-tuning side of IBM’s split: they change the model, they do not query an index. Teams that need both a house style and a live knowledge base run fine-tuning and RAG in tandem, which is the combination IBM describes rather than a winner-take-all choice. Once the method is not fine-tuning, the next question is what it is used for in practice, because a definition that cannot name a workload is still a brochure.

What is RAG used for?

RAG is used wherever an answer must come from documents the model was not trained on, with a trail back to those documents. IBM lists specialised chatbots and virtual assistants, research, content generation, market analysis and product development, knowledge engines, and recommendation services. NVIDIA’s examples are a medical index as an assistant for a clinician, market data for a financial analyst, and technical or policy manuals turned into knowledge bases for customer support, employee training, and developer productivity. Azure groups the same pattern as the common applications of grounding generation in enterprise content rather than in the public pre-training mix.

Wikipedia records healthcare as a studied application and, citing Amugongo et al. in PLOS Digital Health (2025), notes continuing challenges around evaluation, ethics, and clinical reliability. That caveat is the point, not a footnote: a domain with a high cost of being wrong still has to measure retrieval and generation separately, because a cited answer can be the wrong citation. Consumer products that browse the web or search attached files are doing retrieval-augmented generation in effect. The base language model still answers from weights when no retrieval step runs. ChatGPT’s browsing mode is that pattern, not a property of the underlying model; the model alone is not RAG, the product feature that reads outside sources is.

Workloads that look like RAG and are not lookups still need a different architecture. Multi-step questions that need two documents, agent loops that decide whether to retrieve, and graph-backed retrieval are named patterns, not synonyms of the 2020 recipe. Those shapes, and when they earn their extra machinery, live under from naive to agentic. The workload list also implies a storage question: how RAG works with a vector database, and whether one is required at all.

How does RAG work with a vector database?

RAG works with a vector database by storing each chunk as a vector and retrieving the nearest neighbours of the query vector at answer time. AWS describes that step as an embedding language model converting external data into numerical representations stored in a vector database, which becomes the knowledge library the generative model can search. Wikipedia’s process section says the same: referenced data is converted into embeddings and stored in a vector database to allow document retrieval, after which a retriever selects the most relevant documents and the model is prompted with them. That mechanism is how RAG uses a vector store, not a product roundup.

A vector database is the usual production store because similarity search over embeddings is how dense retrieval finds paraphrases that share no tokens with the query. Small prototypes can keep vectors in memory. Production systems typically run a real store so the index can filter, isolate tenants, and update without a full rebuild. Weaviate is the default example: a collection holds the objects, the vectors, and the inverted index for keyword search, and a generative query can retrieve and prompt in one request, as its RAG documentation describes. Weaviate, Pinecone, Qdrant, Elasticsearch, and the other stores in the same category do the retrieval half of that job with different APIs and different defaults. None of them makes generation correct if the wrong object was retrieved.

Weaviate’s own limits belong in the same paragraph as everyone else’s. Named-vector collections must include a target vector name in the query, or the database cannot choose which vector to compare (Weaviate RAG documentation, as of September 2026). Generation is a call to an external model provider, so latency, cost, and output variability are properties of that provider as well as of the search. The starter guide states the variability directly: generated text differs across models and across runs of the same model, which is expected. Keyword search exists because dense retrieval misses exact tokens, and how fusion works is the retrieval question that follows from keeping both rather than from the definition of RAG. Once the store is in the picture, each moving part of the pipeline has its own full explanation, because flattening ingest through generation into a single definition loses the failure that actually happened.

Where does each part of RAG go deep?

Nine pipeline clusters cover the build, and two truth-face clusters cover measurement and breakage. RAG architecture, stage by stage, is the spine that shows how ingest, chunk, embed, retrieve, and generate connect. Document ingestion is how files enter without losing the structure retrieval depends on. How to split documents that retrieve well is the split that decides what a retriever can find. The embedding step sets the ceiling on retrieval quality by choosing what a vector can represent. Indexing for RAG is the structure, trade-off, and limit layer most tutorials skip. What is scored, what is returned is the lookup the architecture is named after, including lexical ranking and hybrid fusion. A second-stage model reorders what retrieval returned when first-stage recall is not enough. What reaches the model is assembly: window budget, compression, order, and duplicates. Where grounding either holds or fails is the last stage, where citations and refusals live.

Measurement and failure have their own hubs because a definition cannot double as a test plan. How to evaluate a RAG system separates retrieval scores from generation scores so a fix can land on the broken stage. The scored, tool-neutral comparison of those instruments is RAG evaluation tools. The complete failure taxonomy sorts wrong answers by the stage that caused them. What it does to retrieval quality is the author layer: vector databases, embedding models, and frameworks, written through retrieval rather than as a vendor catalogue. RAG glossary holds one definition per term. The papers behind those terms are indexed at RAG research index.

The remaining top-level clusters are workloads, operations, and choices. RAG use cases are worked examples by document type and question type. RAG in production is serving, scaling, monitoring, and cost once the traffic is real. RAG security is prompt injection, leakage, access control, and compliance. Choosing between the options is the scored calls that come up before a build: RAG versus fine-tuning, which vector database, build versus buy.

Retrieval-augmented generation fetches documents at query time and writes from them. Patrick Lewis and colleagues named that pattern in 2020, pairing a parametric seq2seq model with a dense Wikipedia index. Production systems implement it as a measurable pipeline over a vector store such as Weaviate, Pinecone, Qdrant, or Elasticsearch, and the pipeline is only as right as the chunk it retrieved.