Retrieval engineering · architecture

RAG Architecture: From Documents to Grounded Answers

Retrieval-augmented generation gives a model selected evidence at answer time. It can improve grounding, but only when ingestion, retrieval, generation, and evaluation are designed as separate, testable stages.

The complete pipeline

A production flow begins before the user asks a question: collect approved sources, normalize them, split them into meaningful units, attach metadata, create embeddings, and store both text and vectors. At query time, retrieve candidates, optionally filter or rerank them, build a bounded context, generate an answer, and return citations.

Chunking is an information-design decision

A fixed character count is a starting point, not a strategy. Preserve headings, lists, tables, and semantic boundaries. A chunk should be specific enough to retrieve precisely and complete enough to answer without inventing the missing half.

Retrieval and generation fail differently

If the right evidence is absent from the retrieved set, rewriting the final prompt cannot repair retrieval. Measure whether relevant evidence appears in the top results before judging the generated answer. Then measure faithfulness: does the answer stay within that evidence?

Abstention is a feature

The system needs permission to say that the available sources do not support an answer. Define minimum relevance, require citations for factual claims, and provide an escalation path when evidence conflicts or is missing.

Production checklist

Version documents and embeddings, preserve source identifiers, enforce access control before retrieval, log which chunks supported each response, remove sensitive data, test prompt injection inside documents, and monitor cost and latency as the collection grows.