When your AI agent contradicts itself, stabilising it can hide two documents that contradict each other
Two answers to one question: the cause may be the model, retrieval or two contradictory documents. The double-answer log tells them apart.
When an AI agent gives two different answers to the same question, the natural place to look is the model: sampling settings, system instructions, document chunking. The cause can sit one level lower, in two company documents that say two different things and that the agent quotes in turn. In that case, inconsistent AI agent answers are inherited from the corpus rather than produced by the engine. And the setting that makes them go away has an effect nobody asked for: an engineer picks, without knowing it, which of the two documents the company applies.
This piece is written for the CDO, CTO or head of knowledge management of a large group whose agent or internal assistant already answers employees from the document estate. It offers a three-level diagnosis, model, retrieval, corpus, of which only the last falls outside the AI team’s tooling. If your agent does not cite its sources, the argument will apply on the day that traceability is switched on.
The person this text makes uncomfortable is the agent’s product lead. Say the product lead of the HR assistant at an industrial group. At the monthly review, an HR manager puts two screenshots on screen. Same question, asked a few days apart by two line managers: can a manager on a fixed-days contract work from home on Fridays? The first answer says the line manager’s approval is enough. The second adds an HR sign-off. The product lead has already spent weeks on this case: instruction rewritten, setting lowered, chunking reworked. That day, someone finally opens the two cited sources. One is the original remote-work agreement, the other the amendment that changed it. Both are online.
The step this piece proposes costs nothing. For every question that received two answers this month, open the sources cited by each answer before touching the model. The output is an artefact described below, the double-answer log.
Two reflexes come first. One is the AI team’s routine: tune the model and retrieval until the answer settles. The other is the tool already bought: the agent evaluation platform, which measures consistency by replaying questions, or the document management system, which handles versions. All three work on answers or on versions; opening the two cited sources to see whether they contradict each other is part of none of them. On 23 September we described how far contradiction-detection tools reach; on 2 October, how the AI team compensates with instructions for what it does not route back to the author. This piece sits before both moments: at diagnosis, when an inconsistency appears and nobody yet knows where it comes from.
What agent consistency means, and its three causes
Analysts have made reliability the precondition for agent autonomy. In a webinar recorded on 23 September 2026, Gartner states that agent autonomy is earned through demonstrated reliability, notes that the underlying technologies perform unevenly from one task to the next, and recommends improving reliability “beyond the model itself”, through agent design, operating environment and human oversight.
Consistency is the simplest of these criteria to measure: ask the same question several times, or in two close phrasings, and compare the answers. The result describes the agent’s outputs. It says nothing about the cause, and three levels can produce it.
The first is the model. A large language model is not strictly reproducible, even when configured to be: work published in September 2025 by Thinking Machines Lab showed that the way servers batch requests is enough to make an answer vary. The defect is real and is fixed on the engine side.
The second is retrieval. Two close phrasings of the same question can surface two different passages, both accurate and compatible. The answer changes shape without changing substance, or becomes partial. That defect is fixed through retrieval and chunking.
The third is the corpus. Two live documents state two incompatible things about the same question. The agent quotes one or the other depending on what surfaces, and its inconsistency faithfully reproduces that of the document estate. Google Research work published in 2025 on conflicting sources shows that models handle these situations poorly, and answer better when explicitly told about the conflict. The consequence for a buyer fits in one sentence: the model cannot settle a disagreement the company has not settled.
The position we hold does not depend on these publications. The document layer comes before the engine, whatever reliability framework an analyst publishes this quarter.
What tuning fixes, and what it hides
The existing tooling deserves to be described at its best. A serious evaluation platform replays a reference question set, compares answers, flags gaps and checks the effect of each setting change. A competent AI team knows how to reduce model variability, stabilise retrieval, pin a source. A well-run document management system knows which version of a document is the latest.
Their boundary is precise. The evaluation platform sees the gap in the outputs without opening the cited documents. The document management system manages versions of one document and does not know that an amendment contradicts an agreement filed elsewhere. Tuning, for its part, produces a constancy that can mislead: when the team pins the amendment or excludes the original agreement, the agent always gives the same answer, and an engineer has chosen the rule being applied without knowing it.
This is the boundary where the product lead comes back in. The weeks spent on the remote-work case served the agent: it now always gives the same answer. She simply does not know whether that answer is the right one, and the original agreement is still online, opened as is by managers who consult it directly.
Another discipline settled this question long ago. The International Vocabulary of Metrology separates the dispersion caused by the instrument from so-called definitional uncertainty, which comes from the limited detail in the definition of what is being measured. That uncertainty sets a floor: no instrument, however precise, will go below it. An agent asked about a rule defined by two contradictory documents is in exactly that position, and changing models amounts to changing instruments. The transfer concerns the method, separating instrument dispersion from definitional dispersion, and not the object: nobody is certifying a corpus, and the metrology function is not the audience for this piece.
Inconsistent AI agent answers: the double-answer log
The step in the introduction produces an artefact, the double-answer log. It is the only new term in this piece.
What it contains. One line per question that received two incompatible answers: the question, rephrased without user identifier or customer data; the first answer and its cited sources; the second and its cited sources; the triage case; the follow-up and its date.
Where it lives. The reference question set the AI team already maintains for its tests, in its evaluation platform or, failing that, a shared file. The log adds two columns, “sources cited by each answer” and “triage case”. It is kept by the team that runs the agent.
When its first lines get written. At the next scheduled run of the reference question set, before any setting change. Never while the team is already tuning the model to make the gap disappear: that is precisely when the cause stops being visible.
A filled-in line. Question: can a manager on a fixed-days contract work from home on Fridays? First answer: yes, with line-manager approval; cited source: remote-work agreement, original version, article on eligible days. Second answer: yes, after HR sign-off; cited source: amendment to the remote-work agreement. Triage case: different and contradictory sources. Follow-up: sent to the authoring HR department, no setting changed, pending.
How to read it. Every line falls into one of three cases. Same sources, different answers: the variability comes from the model, and tuning is legitimate. Different but compatible sources: retrieval is at fault, and chunking or ranking is the place to work. Different and contradictory sources: the corpus is at fault, and no setting should choose between the two documents before the authoring department has said which one applies. Meanwhile, an instruction makes the agent say that two documents diverge on this point, cite both and route the user to the relevant department. After a month, if most lines fall into the first two cases, the AI team is working at the right level. If a visible share falls into the third, part of its tuning work is compensating for documents.
How it fails. If the exercise works, the AI team finds that some of its tuning tickets were corpus issues, and has to hand some of them to departments that were not expecting them. That is its cost: the time it takes to open two sources per case. If the agent does not cite its sources, the exercise fails at once, and that finding is a result in itself: this agent’s consistency cannot be diagnosed.
The record it creates. A line classified “contradictory” shows that on a given date the company knew two of its documents contradicted each other on a rule applied to its employees. It is our reading, with no published decision behind it, that a record followed by a dated handover is preferable to a silent setting that makes the gap disappear without resolving it. The weight of such a record depends on your context: take your legal department’s view before making it a group rule, and your DPO’s before any user question appears in it.
The log creates no new role and draws on the agent’s operating budget, which already funds the reference question set. The trade-off is a dependency: if the evaluation platform changes, the two columns have to be carried over to the next one.
What real fieldwork showed, and what it does not establish
On the technical corpus of a European energy and industrial group, a K-AI diagnostic detected 398 document conflicts: competing versions of the same procedure, inconsistencies between departments. Their targeted remediation came with an improvement in the reliability of AI answers, measured on that scope.
The boundary of this evidence must be stated. It establishes that a real technical corpus can carry several hundred contradictions, and that remediating them weighed on answer reliability. It does not measure consistency in the sense of this piece, the same question replayed twice, and it does not say what share of an agent’s inconsistencies comes from the corpus rather than the model. The double-answer log exists to answer that question on your own agent.
What a DKP adds, and where it stops
A Document Knowledge Platform (DKP) is the document quality and governance layer that runs upstream of AI systems: it governs the document estate (Govern), detects and handles anomalies, duplicates, obsolescence and contradictions (Clean), then activates the corpus for agents only once those two steps hold (Activate). It replaces neither the agent, nor the evaluation platform, nor the document management system. It is not sold as a model observability or evaluation tool and does not belong in a tender for those categories. Consistency measured on outputs is only the channel through which the document problem becomes visible.
Its contribution to the log lies in the third case. A “contradictory” line points to two documents. The DKP counts the other documents in the estate that carry the same statement or its opposite, so that fixing the amendment does not miss the how-to sheet that repeated the original agreement. That is its own unit of work, the contradiction counted across documents, which neither output evaluation nor version management produces.
K-AI is accountable for the count, its reproducibility and the routing of each case to the relevant department, with a prepared diagnosis. Choosing which document applies stays with the business. Technically, the analysis covers only the designated document content, within a contractual ingestion scope. User questions, agent answers and logs stay in your tools, because a DKP does not need them to count contradictions. No data is reused to train models.
And if you do nothing
The status quo has a predictable outcome. The AI team will keep stabilising the agent case by case, each setting silently choosing one document over another. The agent will look more constant at every review, and the corpus will stay as contradictory as before, read as is by everyone who opens it without going through the agent.
Conclusion: audit, clean, monitor
Audit: for every question that received two answers this month, open the sources cited by each and classify the case. Clean: send contradictory cases to the authoring department, and handle with them the other documents that carry the same contradiction. Monitor: make the log a written step of every run of the reference question set, so that inconsistency is diagnosed before it is tuned away.
At the next review, the HR assistant’s product lead will not be putting two screenshots on screen. She will present a line classified “contradictory”, HR’s answer on which document applies, and a setting change she did not need to make.
Frequently Asked Questions
Why does an AI agent give two different answers to the same question?
There are three possible causes: variability in the model itself, retrieval surfacing different passages depending on phrasing, or two documents in the corpus contradicting each other. Only reading the sources cited by each answer tells these cases apart.
Does setting the model’s temperature to zero make an agent consistent?
It reduces model variability without guaranteeing perfect reproducibility. When two documents contradict each other, it makes the agent more constant in choosing one of them, without anyone having decided which one applies.
Our evaluation platform already measures consistency. What is missing?
It measures the gap between answers. The double-answer log adds a reading of the cited sources and a classification of the case, which tells you which level to act on: model, retrieval or corpus.
What should the agent answer while the contradiction is unresolved?
The safest option is for it to flag the divergence, cite both documents and route the user to the relevant department. Silently choosing one of them means applying a rule nobody has validated.
What if our agent does not cite its sources?
Then the diagnosis is impossible, and that is the first thing to fix. Without citations, an inconsistency inherited from the corpus and one produced by the model cannot be told apart.
What confidentiality framework applies to a review of the document estate?
The ingestion scope is contractual and limited to designated document content, excluding user questions, usage logs and telemetry. No data is reused to train models. The scope is validated jointly by the business Document Owner and the CISO or DPO, never by IT alone.
Sources
- Gartner — From Demo to Production: Closing the AI Agent Reliability Gap, webinar recorded 23 September 2026 (Leinar Ramos, Birgi Tamersoy) — autonomy earned through demonstrated reliability; underlying technologies perform inconsistently across tasks; reliability to be improved “beyond the model itself”.
- Thinking Machines Lab — Defeating Nondeterminism in LLM Inference, September 2025 — server-side batch-size variation is enough to make inference non-reproducible, even at temperature zero.
- Google Research — (D)RAGged Into a Conflict: Detecting and Addressing Conflicting Sources in Retrieval-Augmented LLMs, 2025 — taxonomy of source conflicts; explicitly informing the model of the conflict improves answers, with large room for improvement.
- JCGM 200:2012, International Vocabulary of Metrology, 2.27 — definitional uncertainty — uncertainty component resulting from the finite detail in the definition of a measurand; practical minimum achievable uncertainty.
- K-AI — clients page — anonymised energy & industry case, 398 conflicts detected.
Where to Go From Here
K-AI Corpus Diagnostic — 10 business days on your document estate, full report of the 20 most critical anomalies, money-back guarantee if no meaningful anomaly is found. A one-hour conversation can start from a few lines of your double-answer log: together we look, for each contradictory case, at how many other documents carry the same statement or its opposite, and which department each case should go to. Reach the K-AI team: contact@k-ai.ai. The scope of every diagnostic is validated jointly by the business Document Owner and the CISO/DPO, never by IT alone.
K-AI already works with CMA CGM, Veolia, PwC, BNP Paribas, TotalEnergies and CEVA Logistics. Partners: AWS, Snowflake, Microsoft, Wavestone, Devoteam.
