What a RAG pipeline needs before it goes to production

MK
Maryam Khan
AI engineer, Sep 4, 2026 · 9 min read

A retrieval-augmented assistant can look impressive in a demo within a day. Getting it ready for a whole company to rely on takes longer, and almost all of that time goes into five areas.

1. Ingestion you can rerun

Documents change. The pipeline that parses, chunks and embeds them has to be repeatable, track versions, and remove content that has been deleted at the source.

2. Access control at retrieval time

If a person cannot open a document, the assistant should not quote it to them. We filter by permissions inside the search itself, not after the answer is written.

3. Hybrid search and re-ranking

Pure vector search misses exact terms like policy numbers and product codes. Combining keyword and vector search, then re-ranking the top results, fixes most “it couldn’t find it” complaints.

4. An evaluation set

Collect a few hundred real questions with known good answers. Run them on every change to prompts, chunking or models, and track accuracy and citation correctness over time.

5. A clear “I don’t know”

The most trusted answer an assistant can give is an honest one about what it could not find.

When retrieval confidence is low, say so and route to a person. Users forgive a handoff far more readily than a confident wrong answer.

Back to all posts