Back to the notebook
Oct 09, 20263 min read

Reliable AI Starts Before the Prompt

AIEngineering

A convincing answer is easy to mistake for a correct one. It has a tidy structure, an assured tone, and perhaps a few links at the bottom. None of that tells you whether the system found the right evidence.

For a source-grounded assistant, I would start the design conversation one step earlier. Before asking how the model should respond, ask what information it will receive, where that information came from, and what it is allowed to conclude.

Treat the document collection as a product

Imagine a question about eligibility for a public programme. The collection contains an introductory brochure, an older policy, and a revised application guide. All three mention eligibility. They do not all describe the same rules.

A useful document record needs more than extracted text. I would keep its title, publisher, publication date, version where available, and a stable link to the original. Where the material supports it, I would also record what supersedes what. Dates alone do not prove that one document replaces another.

That work is less visible than the chat interface, but it gives the application a way to explain which evidence it used.

Inspect retrieval separately from the answer

Retrieval-augmented generation supplies external content to a model. Microsoft's RAG overview identifies content preparation and retrieval relevance as important parts of the system, alongside access controls and response constraints.

My practical preference is to make retrieval inspectable. For a small set of representative questions, look at the passages returned before reading the generated answer. Are the relevant documents present? Does a passage retain the exception that follows the rule? Is an unrelated document appearing because it shares a familiar phrase?

Consider a section that says a service is available nationally, followed by a paragraph limiting the current pilot to two districts. Returning only the first sentence creates an incomplete basis for an answer. A more eloquent prompt does not restore the missing paragraph.

A citation should make checking easier

I would design citations around the reader's next action. Open the source, find the relevant section, and compare the claim with the evidence.

A document title and a useful location are more helpful than an unexplained identifier. When an answer makes several distinct claims, each should have an identifiable basis. A link to an entire report should not be treated as proof that every sentence above it is supported.

Citation presence and citation correctness deserve separate checks. The interface can have plenty of links and still direct the reader to the wrong evidence.

Make uncertainty an ordinary outcome

Some questions should produce a request for clarification. Others should produce a limited answer, or a clear statement that the collection does not cover the question.

I would include those cases in evaluation from the start: a missing document, conflicting versions, an ambiguous place name, and a question whose premise is wrong. A product that handles only answerable questions is being tested on its easiest day.

The goal is an assistant whose confidence has a basis the reader can inspect. That starts with the evidence pipeline, long before the final sentence appears on screen.