ContinuousLogic Research

A grounded AI answer can still be wrong.

The source can be real. The database value can exist. The API can return successfully. The citation can be accurate.

And the resulting assertion can still be something your organization should not act on.

This working paper defines four distinct failure mechanisms: stale, contested, unauthorized, and orphaned.

Free to read and download. No signup required.

01

Grounded is not the same as governed by current reality

Most AI quality controls ask some version of the same question:

Is the answer supported by the evidence the system retrieved?

That is an important test. It is not the whole test.

The retrieved information may no longer be current. Two credible sources may disagree. A source may be accessible without ever being approved to govern that answer. A stored value may survive after the evidence or rules that produced it have disappeared.

In all four cases, the AI can faithfully use real information and still produce an assertion that should not govern a decision.

The problem is broader than document management. The underlying source might be a document, database, API, metric, model output, configuration value, business rule, cache, or other derived state.

Diagram showing documents, databases, APIs, metrics, configuration and derived state feeding evidence available to an AI. Groundedness tests whether the evidence supports the assertion, while the belief failure taxonomy tests whether the evidence is current, resolved, authorized and traceable.
Figure 1 · The relationship grounding does not test. Groundedness asks whether an assertion is supported by available evidence. Belief integrity also requires asking whether that evidence is current, resolved, authorized, and traceable. Open at full size
02

Four failure mechanisms

These failures look similar at the answer boundary. Their causes are not the same.

Stale

The answer outlived its truth.

The information was correct. Reality changed. An older representation remained available and the AI continued to serve it.

A previous policy can remain searchable after a replacement is published. A reporting database can retain the old customer tier after the master system changes. A downstream rules table can continue using yesterday’s pricing threshold.

Missing decision: Which value governs now?

Contested

Two answers, no ruling.

Two credible sources make incompatible claims within the same scope, and nothing establishes which one should govern.

A global policy can conflict with regional guidance. A customer master can classify an account as enterprise while billing calls it mid-market. Finance and sales can hold different current revenue values.

The AI does not create the disagreement. It inherits it.

Missing decision: Which source has authority for this question?

Unauthorized

Nobody approved this.

The information is accessible, but no one decided it should be allowed to govern the answer.

A draft proposal can sit beside the published price book. An experimental scoring table can be reachable from production. A provisional forecast or sandbox rule can become evidence simply because the AI has permission to retrieve it.

Technical access silently becomes operational authority.

Missing decision: Was this source approved to govern this answer?

Orphaned

The justification is gone, but the assertion remains.

A stored or derived assertion survives after the evidence, system, or rule that originally supported it can no longer be followed.

A cached summary can outlive a deleted source. A risk classification can remain after the calculation pipeline is retired. A metric can survive after its definition disappears.

The assertion may even still be correct. The failure is that no one can demonstrate why.

Missing decision: Can this assertion still be traced to living evidence and rules?

Matrix comparing stale, contested, unauthorized and orphaned belief failures with examples from documents and structured data, what the AI sees, and the missing organizational decision.
Figure 2 · Four mechanisms, not four versions of the same problem. Each failure can occur in documents or structured data, and each requires a different detection procedure and control. Open at full size
03

Why calling all of this hallucination matters

“Hallucination” has become shorthand for “the AI was wrong.”

That collapses several very different problems into one category.

If the model invents an unsupported answer, improving generation, retrieval, or evaluation may help.

But better prompting cannot determine whether one database supersedes another.

A hallucination detector cannot tell you whether an experimental scoring table was authorized for production decisions.

Citation enforcement cannot establish whether a cached result still reflects current reality.

The taxonomy matters because different failure mechanisms require different controls.

The paper is an attempt to make those mechanisms visible enough to measure separately.

04

You can test one part of this today

You do not need to accept the taxonomy on faith.

Choose a fact in your organization that changed recently.

It might be:

  • a policy threshold
  • a customer classification
  • an entitlement
  • an employee assignment
  • a price
  • a business rule
  • a metric definition
  • an operational status

Identify both the previous value and the value that governs now.

Then ask your AI system the question whose answer changed.

Record whether it returns the current value, the superseded value, both values, a conflict, or uncertainty.

Do it more than once.

The interesting result is not whether the first answer is right.

The interesting result is the rate.

Process showing an old value and a current value from documents, databases, APIs, configuration or derived metrics being tested through an AI assistant. Answers are classified as current, superseded, both, conflict or uncertainty and used to calculate a superseded-answer rate.
Figure 3 · A simple test for superseded answers. Repeat the test across changed facts from multiple source types and calculate how often the AI continues to serve the previous value as current. Open at full size

Read the full working paper

The working paper develops the taxonomy in detail, including:

  • why belief failures are different from generation failures
  • how the four mechanisms differ
  • why structured data is just as susceptible as documents
  • what existing groundedness checks do and do not establish
  • how each failure can be detected
  • a simple observable test you can run against your own environment

No registration. No email gate.

05

A series on the integrity of AI-served knowledge

This is the first working paper in an ongoing series examining what happens when AI systems become an operational interface to organizational knowledge.

The papers are intended to make the problem increasingly testable, moving from failure taxonomy to measurement and then to the controls required when organizational knowledge changes.

Coming next

Working paper 2: Why recency is not authority

When two pieces of information conflict, treating the newest one as correct sounds reasonable.

Inside an organization, it often is not.

The next paper examines why recency is an unreliable substitute for authority.

Follow the research

If you want the next working paper and the measurement methodology when they are published, leave your email.

Research updates only. Unsubscribe anytime.

About the author

Treb Gatte

Treb works at the intersection of data, AI, and operational decision-making, with a particular focus on what happens when AI systems become trusted interfaces to changing organizational knowledge.

This working-paper series documents an effort to define those failures precisely enough that they can be observed, measured, challenged, and eventually controlled.