How it works
- Ingest. Documents are split into passages and converted into embeddings, numeric representations of meaning, stored in a vector index alongside the text and its permissions.
- Retrieve. A question is converted the same way, and the most relevant passages the user is allowed to see are retrieved.
- Generate. The model answers using only those passages, and cites them.
Why it works better than a chatbot alone
- Answers come from your current content, not the model's training data.
- Citations let people verify answers in seconds.
- Updating knowledge means updating documents, not retraining a model.
What separates a demo from production
- Permissions. Retrieval must respect existing access rights, or the assistant becomes a data leak.
- Content quality. Outdated and duplicated documents produce confident wrong answers. Curate the sources.
- Evaluation. Build a test set of real questions with known answers and measure accuracy before and after every change.
- Refusal. The assistant should say when the sources do not support an answer.
- Security. Retrieved content can contain hidden instructions, so treat it as untrusted and test for prompt injection.
- Hosting. Choose where embeddings, indexes and models run based on the sensitivity of the content.
Where it pays off
Policies and procedures, contracts, technical manuals, service desks, research libraries and board papers: anywhere people spend time searching for the right document and interpreting it.
How we helpTalk to us about private ai assistants, or start with the free AI readiness check.
