A withdrawn document can keep answering through your AI
A withdrawal procedure handles the file, not the copies AI made of it. The withdrawal record checks that a withdrawn document is really gone.
Companies withdraw a document when it is superseded, when it is wrong, or when someone has asked for their personal data to be erased. A withdrawal procedure written before AI arrived deals with the source document. It knows nothing of the copies AI systems have made of it: a search index, a saved summary, an agent’s memory, a test question set. When AI is still citing a withdrawn document, the assistant is quietly undoing a documented decision, and nobody knows.
This article is for the CDO, CTO or head of knowledge management of a large group whose internal assistant or agent already reads the document estate. If your assistant queries documents live, with no index, no memory and no saved summaries, the argument only covers derived documents and exports; it will apply in full the day you add a memory or a second connector.
The person this text makes uncomfortable is the quality manager of an industrial site. After a non-conformity, he withdrew the goods-inward inspection procedure for a component and published a new version with a changed acceptance criterion. At the next lessons-learned review, an inspector says he asked the plant assistant for the criterion and the answer gave the old one. He only noticed because the sheet posted at his workstation said something else. The quality manager opens the document management system: the old version is archived as it should be. He has no idea where the assistant found it.
The step this article proposes costs nothing. In the archive log of your document management system, or the revision register of your quality system, find the last document formally withdrawn, and ask your assistant the question it used to answer. The result feeds an artefact described below, the withdrawal record.
Two answers come to mind. The first is the in-house process: the withdrawal workflow run by records management or the quality system, which archives the superseded version and notifies recipients. The second is the tool already bought: connector synchronisation, which removes a deleted file from the index, and retention and archiving policies, some of which now exclude archived content from the assistant’s indexing. These tools act on the file; the document’s content, copied or rephrased elsewhere, escapes them. In July we described a corpus that drifts silently as documents change. This article covers the opposite case: an explicit decision to make content disappear, which the AI’s copies do not follow.
Where a document’s copies live once AI has read it
Depending on its architecture, an enterprise assistant keeps traces of a document at several levels.
The search index splits the document into passages and stores representations of them. With a well-configured connector, deleting the file propagates to the index after a synchronisation delay. If the document is merely moved to an archive space the connector still reads, nothing propagates.
Content produced from the document outlives it: a summary saved in a team space, a job aid drafted with the assistant’s help, an answer pasted into an email and filed. Each carries the withdrawn content in another form, under another file name.
Agents’ persistent memories keep what they retained from an exchange. In an exploratory note published on 20 July 2026, France’s data protection authority, the CNIL, and the French AI and Digital Council observe that keeping interaction histories and using persistent memories increase the data these systems hold “on a variety of media”. The note is about personal data and creates no new obligation. It describes a mechanism that applies to documentary content too.
Finally, the AI team keeps its own copies: the reference question set whose expected answers cite the document, and instructions that mention it by name. On the day the document is withdrawn, the test keeps validating the old answer.
AI still citing a withdrawn document: what the tools cover, and what they leave
Existing tooling deserves to be described at its best. A well-configured connector propagates deletions. Access controls apply at query time, so a document made inaccessible no longer surfaces. In September 2026 Microsoft announced, for rollout between October and November, an archive option in Purview retention policies: inactive SharePoint and OneDrive files move to an archive, remain available for eDiscovery and leave Copilot indexing. The option is off by default.
The boundary of these tools is clear. They handle a file, in one environment. They do not know that a job aid repeats three paragraphs of the withdrawn procedure, that an agent memorised its key criterion, or that the reference question set still expects the old answer. Archiving by inactivity does not coincide with withdrawal either: a procedure withdrawn after a non-conformity was, the day before, one of the site’s most consulted documents.
This is where the quality manager comes back in. His withdrawal workflow worked: the superseded version is archived, recipients are notified. Nothing in that workflow asked him to check what the plant assistant had kept, and nobody else did.
Another discipline has already written this rule. On 10 February 2026, the European Data Protection Board adopted its report on the 2025 coordinated enforcement action on the right to erasure, carried out by 32 supervisory authorities with 764 controllers. Half of the responding authorities raised concerns about erasure in back-ups, and many organisations had no specific procedure for it. The report recommends keeping track of erasure requests so they can be applied to restored systems, verifying that erasure has been carried out and being able to demonstrate it. Data protection has therefore already accepted that erasure must reach the copies of the original. Copies created by AI call for the same catch-up for any withdrawn document, personal or not. What carries over is the method, following a withdrawal to its copies, and not the object: this article is not about GDPR compliance, and the DPO is only concerned for withdrawals that follow an erasure request.
The withdrawal record
The step from the introduction produces an artefact, the withdrawal record. It is the only new term in this article.
What it contains. For each withdrawn document: its identifier, the reason (superseded, wrong, erasure requested), the control question it used to answer, the assistant’s answer after withdrawal and its cited sources, the list of copies checked with their status (absent, found then removed, cannot be checked), and the date and team to which each uncheckable copy was referred.
Where it lives. The existing withdrawal procedure, in the document management system or the quality management system, run by records management or the document controller. The record adds one section, “AI copies checked”, completed with the team that operates the assistant.
When its first entries are written. At the next planned withdrawal, a version replaced by another in the normal course of revisions. Never during an emergency withdrawal after a non-conformity: that is exactly when nobody has time to look for copies.
A completed record. Document: goods-inward inspection procedure for the component, previous version. Reason: superseded after a non-conformity. Control question: what is the acceptance criterion for the component at goods-inward? Assistant’s answer after withdrawal: new criterion, citing the new version. Copies checked: plant assistant index, absent; night-shift job aid drafted with the assistant, found then replaced; reference question set, expected answer updated; planning agent memory, cannot be checked, referred to the AI team.
How to read it. If the assistant answers with the replacement document and every copy is absent or removed, the withdrawal is complete. If it cites or paraphrases the withdrawn document, a copy is escaping it: find it before closing the withdrawal. If a copy cannot be checked, the withdrawal stays open and the assistant’s operator says how to purge it. Meanwhile, the withdrawn document goes on the assistant’s exclusion list, and the control question joins the reference question set to be replayed at every run.
How it can fail. If the exercise works, the document controller discovers copies no procedure asked him to know about, and has to flag them to teams that did not expect it. That is its cost: one question asked and a few copies looked for per withdrawal. If the assistant does not cite its sources, the reading rule cannot be applied, and that finding is a result in itself: you cannot verify that a withdrawal reached this assistant.
The record it creates. A record showing an “uncheckable” copy establishes that, on a given date, the company knew withdrawn content might still be served. Our reading, with no published decision to support it, is that a dated trace followed by a referral is preferable to a withdrawal declared complete without verification. What such a record means for you depends on your context: take your legal team’s view before making it a group rule. For withdrawals that follow an erasure request, the record holds no personal data, only the request reference, and the DPO’s procedure applies.
The record creates no new role. It relies on the withdrawal procedure the quality system already funds and, if your organisation configures archiving of inactive content this autumn, on the project rolling it out. The trade-off is a dependency: if the assistant’s architecture changes, the list of copies to check changes with it.
What a real deployment showed, and what it does not establish
At TotalEnergies Retail Power & Gas, the customer chatbot relies on some 500 pages of official web documentation. During the diagnostic carried out with K-AI, 19% of those pages turned out to need correcting, including cases hard to spot by eye at that scale. Within three weeks, 53% of cases were resolved by handling the most critical first, with one or two experts giving half a day to a day a week. Business experts made the call on each case; K-AI prepared the diagnostic and indicated which expert each one should go to.
The boundary of this evidence must be stated. It shows that a production corpus can carry a significant share of content needing correction, and that prioritised remediation remains sustainable with limited resources. It does not measure how withdrawals propagate to an assistant’s copies. The withdrawal record exists precisely to examine that question in your organisation.
What a DKP adds, and where it stops
A Document Knowledge Platform (DKP) is the document quality and governance layer that runs upstream of AI systems: it governs the document estate (Govern), detects and treats anomalies, duplicates, obsolete content and contradictions (Clean), then activates the corpus for agents only once those two steps hold (Activate). It does not replace the document management system, the retention and archiving tool, or the assistant. It is not sold as a data protection compliance or data lifecycle management tool, and does not belong in a tender for those categories. Withdrawal is simply the moment the documentary problem becomes visible. This position does not depend on regulators’ or vendors’ calendars: the document layer comes before the engine.
Records management holds part of the answer: it knows which document was withdrawn, when and why. What a DKP contributes is what that discipline does not produce: the documents in the estate that repeat the withdrawn content under another name. For each withdrawn document, a DKP counts the job aids, notes and pages still carrying its statements, so that withdrawing a procedure does not miss the night-shift job aid that copied it. That is its own unit of work, a statement tracked across documents, which version control and retention do not produce.
K-AI is accountable for that count, its reproducibility and the routing of each documentary copy to the right team. The decision to withdraw stays with the business. Purging the assistant’s indexes, memories and instructions stays with the team that operates it and its vendor. Technically, the analysis covers only the designated document content, within a contractual ingestion scope; user questions, agent memories and logs stay in your tools. No data is reused to train models.
If you do nothing
The status quo has a predictable outcome. Every withdrawal will be declared complete when the file leaves the active area of the document management system. Depending on the case, the assistant will keep citing a derived document, relying on its memory or validating the old answer in its tests. Nobody will know until a user reports an answer based on a document the company thought was gone.
Conclusion: audit, clean, monitor
Audit: ask the assistant the question the last withdrawn document used to answer, and read the cited sources. Clean: look for copies, documentary and technical, and refer those beyond the document controller’s reach. Monitor: make the withdrawal record a written step of every withdrawal, and replay the control questions at every run of the reference question set.
At the next lessons-learned review, the quality manager will no longer ask where the assistant found the old criterion. He will present the withdrawal record for the goods-inward inspection procedure, with the list of copies checked and the assistant’s answer after withdrawal.
Frequently Asked Questions
Why does our AI still cite a document we deleted or archived?
The document may have left copies the withdrawal procedure does not cover: an index not yet synchronised or still reading the archive space, a derived document, a saved summary, an agent’s memory or a test question set. Reading the sources cited in the answer tells you which.
Does deleting a file remove it from an AI assistant’s index?
With a well-configured connector, deletion propagates to the index after a synchronisation delay. It does not touch content that repeats the document under another name, nor the assistant’s memories and tests.
How do we find outdated documents our AI is still using?
Start from recent withdrawals logged in your document management or quality system, ask the assistant the question each document used to answer, and read the cited sources. An answer that cites or paraphrases the withdrawn document points to a copy to find.
Does automatic archiving of inactive content solve the problem?
It removes archived files from indexing, which helps. It relies on inactivity, which does not coincide with a decision to withdraw a document, and it does not reach copies of the content stored elsewhere.
What if the withdrawal follows a personal data erasure request?
The DPO’s procedure applies. The withdrawal record then holds only the request reference, with no personal data, and serves to verify that erasure has reached the assistant’s copies, as the European Data Protection Board recommends for back-ups.
What confidentiality framework applies to a review of the document estate?
The ingestion scope is contractual and limited to designated document content, excluding user questions, agent memories, usage logs and telemetry. No data is reused to train models. The scope is validated jointly by the business Document Owner and the CISO or DPO, never by IT alone.
Sources
- CNIL — Agentic AI and personal data: exploratory note with the French AI and Digital Council, 20 July 2026 (French) — persistent memories and interaction histories, data held “on a variety of media”; existing framework applies, implementation to be adapted.
- EDPB — Coordinated Enforcement Action: implementation of the right to erasure by controllers, adopted 10 February 2026 (PDF) — 32 authorities, 764 controllers; issue 6 on back-ups; recommendation to verify erasure and be able to demonstrate it.
- Reed Smith — EDPB report on the right to erasure: key takeaways, 9 March 2026 — law firm commentary; only points verified in the primary report are used in the body.
- MC1472601 — Microsoft Purview: archive OneDrive and SharePoint files under retention policies, published 16 September 2026 (Message Center relay) — off by default; archived content excluded from Copilot indexing; available for eDiscovery.
- K-AI — customers page and TotalEnergies Retail Power & Gas case study — ~500 pages, 19% needing correction, 53% of cases resolved in three weeks.
Where to Go From Here
K-AI Corpus Diagnostic — 10 business days on your document estate, full report of the 20 most critical anomalies, money-back guarantee if no meaningful anomaly is found. A one-hour conversation can start from a few withdrawal records: together we look, for each withdrawn document, at how many other documents still carry its content, and which team each should go to. Reach the K-AI team: contact@k-ai.ai. The scope of every diagnostic is validated jointly by the business Document Owner and the CISO/DPO, never by IT alone.
K-AI already works with CMA CGM, Veolia, PwC, BNP Paribas, TotalEnergies and CEVA Logistics. Partners: AWS, Snowflake, Microsoft, Wavestone, Devoteam.
