Working paper 1 · August 2026
The belief failure taxonomy
Working paper 1 in a series on the integrity of AI-served knowledge
A grounded AI answer can still be wrong.
The source can be real. The database value can exist. The API can return successfully. The citation can be accurate.
And the resulting assertion can still be something your organization should not act on.
This working paper defines four distinct failure mechanisms: stale, contested, unauthorized, and orphaned.
Free to read and download. No signup required.
Grounded is not the same as governed by current reality
Most AI quality controls ask some version of the same question:
Is the answer supported by the evidence the system retrieved?
That is an important test. It is not the whole test.
The retrieved information may no longer be current. Two credible sources may disagree. A source may be accessible without ever being approved to govern that answer. A stored value may survive after the evidence or rules that produced it have disappeared.
In all four cases, the AI can faithfully use real information and still produce an assertion that should not govern a decision.
The problem is broader than document management. The underlying source might be a document, database, API, metric, model output, configuration value, business rule, cache, or other derived state.
Four failure mechanisms
These failures look similar at the answer boundary. Their causes are not the same.
Stale
The answer outlived its truth.
The information was correct. Reality changed. An older representation remained available and the AI continued to serve it.
A previous policy can remain searchable after a replacement is published. A reporting database can retain the old customer tier after the master system changes. A downstream rules table can continue using yesterday’s pricing threshold.
Missing decision: Which value governs now?
Contested
Two answers, no ruling.
Two credible sources make incompatible claims within the same scope, and nothing establishes which one should govern.
A global policy can conflict with regional guidance. A customer master can classify an account as enterprise while billing calls it mid-market. Finance and sales can hold different current revenue values.
The AI does not create the disagreement. It inherits it.
Missing decision: Which source has authority for this question?
Unauthorized
Nobody approved this.
The information is accessible, but no one decided it should be allowed to govern the answer.
A draft proposal can sit beside the published price book. An experimental scoring table can be reachable from production. A provisional forecast or sandbox rule can become evidence simply because the AI has permission to retrieve it.
Technical access silently becomes operational authority.
Missing decision: Was this source approved to govern this answer?
Orphaned
The justification is gone, but the assertion remains.
A stored or derived assertion survives after the evidence, system, or rule that originally supported it can no longer be followed.
A cached summary can outlive a deleted source. A risk classification can remain after the calculation pipeline is retired. A metric can survive after its definition disappears.
The assertion may even still be correct. The failure is that no one can demonstrate why.
Missing decision: Can this assertion still be traced to living evidence and rules?
Why calling all of this hallucination matters
“Hallucination” has become shorthand for “the AI was wrong.”
That collapses several very different problems into one category.
If the model invents an unsupported answer, improving generation, retrieval, or evaluation may help.
But better prompting cannot determine whether one database supersedes another.
A hallucination detector cannot tell you whether an experimental scoring table was authorized for production decisions.
Citation enforcement cannot establish whether a cached result still reflects current reality.
The taxonomy matters because different failure mechanisms require different controls.
The paper is an attempt to make those mechanisms visible enough to measure separately.
You can test one part of this today
You do not need to accept the taxonomy on faith.
Choose a fact in your organization that changed recently.
It might be:
- a policy threshold
- a customer classification
- an entitlement
- an employee assignment
- a price
- a business rule
- a metric definition
- an operational status
Identify both the previous value and the value that governs now.
Then ask your AI system the question whose answer changed.
Record whether it returns the current value, the superseded value, both values, a conflict, or uncertainty.
Do it more than once.
The interesting result is not whether the first answer is right.
The interesting result is the rate.
Read the full working paper
The working paper develops the taxonomy in detail, including:
- why belief failures are different from generation failures
- how the four mechanisms differ
- why structured data is just as susceptible as documents
- what existing groundedness checks do and do not establish
- how each failure can be detected
- a simple observable test you can run against your own environment
No registration. No email gate.
A series on the integrity of AI-served knowledge
This is the first working paper in an ongoing series examining what happens when AI systems become an operational interface to organizational knowledge.
The papers are intended to make the problem increasingly testable, moving from failure taxonomy to measurement and then to the controls required when organizational knowledge changes.
Coming next
Working paper 2: Why recency is not authority
When two pieces of information conflict, treating the newest one as correct sounds reasonable.
Inside an organization, it often is not.
The next paper examines why recency is an unreliable substitute for authority.
Follow the research
If you want the next working paper and the measurement methodology when they are published, leave your email.