This article is part of our Journal archive. Any prior offers reflect its publication date. Read our current services and approach.
AI systems can sound certain when the evidence is weak. A useful workflow makes uncertainty visible and gives people a clear way to handle it.
A model will not always know the answer. That is not unusual, and it is not automatically a reason to avoid AI.
The real question is what the system does next.
A weak workflow lets the model produce a confident answer, passes that answer downstream, and leaves a person to discover the mistake later. A better workflow treats uncertainty as part of the design. It makes the limits visible, routes doubtful cases to the right person, and keeps enough evidence for that person to make a sound judgment.
This matters whether the model is sorting requests, extracting information from documents, drafting a response, or searching internal material. The model can assist with the work. It should not hide how little support it has for an answer.
Confidence is not the same as correctness
Models generate likely responses from context. A polished sentence may be well supported, partly supported, or invented. Tone alone does not tell us which one we have.
A confidence score can help in some systems, but it is not a universal truth meter. The score may describe the model's own estimate, a retrieval score, a rule set, or a combination of signals. Each means something different.
We therefore avoid treating one number as proof. We look at the conditions around the answer instead. Did the system find a relevant source? Does that source contain the claimed fact? Are several sources in conflict? Is required information missing? Is the requested action within the approved scope?
Those questions produce useful signals because they relate to the actual task.
Start by defining an acceptable answer
Uncertainty is difficult to handle when nobody has defined what a good answer requires.
For a document workflow, an acceptable answer might require a named field to appear in the source and a link back to the page where it was found. For an internal assistant, it might require an answer grounded only in approved company material. For request routing, it might require both a recognized topic and enough customer information to choose the correct queue.
The standard should match the consequence of being wrong. A rough internal summary may tolerate ambiguity that would be unacceptable in a payment decision, personnel matter, safety instruction, or public statement.
We decide that standard before we decide how much autonomy to give the system. Otherwise the workflow tends to automate whatever the model can produce rather than what the organization can safely use.
Give the model a way to decline
Many poor outputs begin with an instruction that quietly demands an answer every time.
A useful system needs an approved way to say that it does not have enough information. That response should be specific. It might identify the missing field, cite conflicting material, ask a focused follow-up question, or route the item for review.
This is more useful than a vague warning attached to a completed answer. If the evidence is insufficient, the workflow should stop at the point where more information is needed.
The decline path also needs to be acceptable to the people using the tool. If staff are punished for sending a case to review, they will work around the control. If customers receive a dead end with no next step, they will repeat the request or leave. The handoff is part of the product, not an exception to it.
Show the evidence with the output
Review is faster when the reviewer can see why the system reached its answer.
For work grounded in documents, we prefer outputs that point to the source passage rather than merely naming a file. For extracted data, we want the original value beside the normalized value. For a drafted message, we want the approved facts and instructions that shaped it.
This does not make the model infallible. Citations can be irrelevant, and source material can be wrong. It does make inspection possible.
Evidence also improves correction. A reviewer can tell whether the problem came from a missing document, poor retrieval, an unclear instruction, or the model's interpretation. Each failure calls for a different fix.
Route review according to consequence
Not every uncertain item needs the same reviewer or the same response time.
A low-consequence case may go into a normal work queue. A sensitive case may need a named owner before anything is sent, changed, or recorded. Some actions should remain human decisions even when the model appears certain.
We use practical boundaries such as subject, data type, requested action, and downstream effect. The workflow might allow the model to draft a reply but not send it, summarize a contract but not approve a term, or suggest a record update but not write to the system of record.
The boundary should be visible to the user. People should know whether they are reading a draft, reviewing a recommendation, or approving an action.
Keep the uncertain cases
The difficult cases are useful design material.
We keep a reviewable record of what the model received, what it produced, what evidence it used, why the case was escalated, and what the reviewer decided. The record should respect the organization's retention and access rules, but it should contain enough detail to diagnose the workflow.
Patterns usually matter more than isolated surprises. Repeated uncertainty may reveal missing source material, inconsistent terminology, an overly broad task, or a connector that does not expose the necessary information. It may also show that the work depends on judgment that should not be automated.
The goal is not to drive every uncertainty signal to zero. A system that never admits doubt may simply be hiding it. The goal is to make the uncertain cases legible and manageable.
Test the failure path in the prototype
A demonstration that covers only the easy examples tells us very little.
When we build a working prototype, we include incomplete requests, conflicting documents, unfamiliar language, and cases outside the intended scope. We look at what the system refuses, what it escalates, what evidence it shows, and what reaches a person.
That is where the operating model becomes concrete. The buyer can inspect not only what the system does when it succeeds, but also how it behaves when the answer is unclear.
If the work is a fit, we build a working prototype before you pay. There is no obligation and no money down. You see the normal path and the uncertain path before you decide whether to proceed.
AI does not need to be certain about everything to be useful. It does need a workflow that knows what uncertainty looks like, stops at the right boundary, and gives a person enough context to take over.