When AI should say “I don’t know”: test uncertainty in your brand answers
What the benchmark actually tested
AbstentionBench, published in June 2025, studied abstention across 20 datasets and 20 frontier models. Its authors reported that reasoning fine-tuning reduced abstention performance by an average of 24% in their comparisons.
That result does not mean 24% of all AI answers are false, or that no model can express uncertainty. It concerns a defined benchmark and scoring method. Different question types and model configurations can behave differently.
Build questions with a known verification boundary
Start with claims that matter to your customers: current prices, product availability, compatibility, contract terms or refund rules. Save the authoritative source and the date. Include a question the published evidence can answer and a closely related question it cannot settle.
For example, a page may publish the standard cancellation policy without establishing whether a specific booking qualifies for an exception. A good answer should state the standard rule and direct the customer to confirmation rather than inventing permission.
| Outcome | How to record it |
|---|---|
| Correct, supported answer | The response matches the dated source and keeps its restrictions. |
| Appropriate uncertainty | The response identifies a missing fact and does not invent it. |
| Unsupported assertion | The response states a fact the evidence does not establish. |
| Incorrect refusal | The response declines a question the supplied evidence can answer. |
Score uncertainty separately from visibility
Repeat the test and preserve the conditions
Keep the prompt, available tools, model version, market and date. Save the answer and sources. Repeat both answerable and unanswerable questions so a refusal rate is not inflated by a test set that contains only impossible requests.
Separate pricing errors from policy errors or unsupported product benefits. The fix depends on the source and failure type. A single average score can conceal a serious error in the one policy customers rely on.
Correct the evidence before adding persuasive copy
When a cited source is wrong, fix that page or request a correction from its publisher. When the evidence is missing, publish only the facts the business can substantiate. Keep restrictions next to the headline claim so the answer does not have to infer them.
Google says no special schema is required for its AI features. Structured data should match visible facts, not fill gaps with invented certainty. Schema is not a guarantee that an assistant will stop hallucinating.
Do not confuse a bot visit with an accurate answer
OpenAI's agent documentation describes the purposes of its crawlers and user-action agents. A request to your pricing page is not proof that the final answer quoted the price correctly. Inspect the captured answer to assess that outcome.
Report unsupported-claim rates with their raw counts, test-set composition and model settings. Track human support escalations separately. Where health, safety or contractual consequences matter, route unresolved questions to an appropriately qualified person.
The AI search metrics guide explains how to keep answer accuracy, brand mentions and human outcomes separate.
Explore the related measurement tools
See Aiso’s prompt, fan-out and source-analysis workflow, with its sampling and coverage limits.