Generative AI
Retrieval-augmented generation is a data project first
· 6 min read · LT Lab
The quality of an AI knowledge assistant depends far more on document preparation and access control than on the choice of model.
Retrieval-augmented generation (RAG) connects a language model to your own documents so it can answer questions with organizational knowledge. It is one of the most practical applications of generative AI. It is also frequently underestimated.
Most failures are retrieval failures
When an assistant gives a poor answer, the model is rarely the root cause. More often, the right passage was never retrieved: the document was outdated, poorly parsed, split in the wrong place or missing entirely.
What good preparation involves
- Identifying authoritative sources and retiring duplicates and old versions
- Parsing tables, headings and scanned documents correctly
- Chunking content along meaningful boundaries
- Attaching metadata such as owner, date, department and access level
- Keeping the index in sync as documents change
Permissions are not optional
An assistant must never reveal content to someone who could not open the original document. Access control needs to be enforced at retrieval time, using the same rules as the source systems.
Measure before and after launch
Build a test set of real questions with known good answers. Use it to compare retrieval strategies, prompts and models, and keep running it after launch. Evaluation turns an impressive demo into a dependable tool.