AI & Generative AI
Take a Retrieval-Based AI Assistant From Pilot to Production
By David Campodonico ·
Design an enterprise AI assistant around authorized retrieval, grounded answers, evaluation, and a clear operational handoff.
- Enterprise AI
- AI architecture
- Retrieval augmented generation
- AI governance
A retrieval-based assistant can answer questions using an organization's documents, but adding search to a model does not make the resulting service trustworthy by default. Production readiness depends on which information is retrieved, whether the user may access it, how answers are checked, and who responds when the system fails.
Begin with a narrow service boundary. An assistant that explains approved internal procedures is easier to evaluate than one expected to answer any question about the business. State what it covers, which source is authoritative, and when it should decline or route the question to a person.
Treat retrieval as a governed data path
Map the flow from source systems through ingestion, indexing, retrieval, and generation to the user. At each stage, identify the data owner, access rules, retention expectations, and failure behavior. Derived indexes and logs need attention as well as the original documents.
Apply authorization before information is made available to the model. Do not rely on a prompt telling the model to hide information that it has already received. Test with users who have different permissions, including a user whose access was recently removed. Confirm how changes in the source propagate to the index and caches.
Treat retrieved text as data, not trusted instructions. A document can contain irrelevant commands or malicious instructions. Keep the assistant's tool permissions bounded, separate content from control decisions, and test attempts to redirect the workflow. Retrieval does not remove the need for application security.
Require evidence for useful answers
For a hypothetical employee policy assistant, ask it to provide the applicable source and indicate when information is missing or conflicting. The presence of a citation is only a first check. Review whether the source actually supports the answer and whether it is current for the question being asked.
Create tests for answerable questions, questions outside scope, stale documents, conflicting versions, and inaccessible material. Measure supported answer quality, appropriate abstention, retrieval failures, and the time a user needs to verify a result. Keep a set of cases out of prompt development for independent checking.
AWS's Generative AI Lens provides a broader architecture review framework. Use it alongside workload-specific tests; a general review cannot establish whether an answer about your organization's policy is correct.
Plan for partial failure
Decide what the user sees if retrieval is unavailable, the model times out, or a document update has not reached the index. A clear unavailable message with an alternative route can be better than an unsupported answer generated from general model knowledge.
Bound retries, record enough diagnostic context to investigate, and avoid indiscriminately logging sensitive content. Track latency and cost across retrieval, generation, and human review. The service owner needs visibility into failed and abandoned requests, not just successful demonstrations.
Keep a rollback or pause mechanism that does not depend on changing a prompt under pressure. Rehearse it with the operations team and document how users reach the original source or support channel while the assistant is unavailable.
Hand over a service, not a prototype
Before expansion, confirm a named business owner, technical owner, content maintenance process, incident route, and evaluation cadence. Define which changes require re-evaluation: a new model, different retrieval settings, new document collections, or expanded permissions can change behavior.
Compare this architecture with the simpler alternatives in the enterprise AI pilot scorecard and use the Claude workflow evaluation guide when assessing that model. The production decision should rest on demonstrated workflow quality and operational readiness, with model choice as one part of the system.
