Press
47 articles

The description of your documents already exists somewhere in your company. Your AI programme is rebuilding it
Gartner: through 2028, building your own unstructured metadata will cost 300% more than reusing what exists. The trade-off is settled at framing.

Your tools catch contradictions inside a package. Nobody counts the ones that cross packages
Two buyer-side corporate funds just financed automated document review. At package scale. The cross-functional document falls outside every perimeter.

An incorrect answer to a regulator is a separate infringement from the one it is asking about
Article 99(5) fines an incorrect reply: €7.5M or 1% of turnover. And that reply is assembled from documents that contradict each other.

Authoritative sources for enterprise AI: what you certify when you designate one
SharePoint now lets you mark a site authoritative for Copilot. The signal sits on the container; the contradiction lives in the documents.

Enterprise-Wide AI Agent Deployment: When One Document Contradiction Becomes a Serial Defect
Cisco opened MyAgent to 90,000 employees. At pilot scale the user is the anomaly detector. At 100% of the workforce, the detector is gone.

Cost per task in enterprise AI: the metric that turns document quality into a budget line
Gartner: only 22% of enterprises have scaled AI. Around 60% of an agentic task's cost goes to rework — a corpus problem, not a model problem.

AI Agent Decision Auditability: The Log Shows What the Agent Did, Not What It Chose Between
Glean gave agents their own identity and audit log. The log records which source the agent cited, never which valid source it discarded.

Enterprise AI Benchmarks: You Will Pick Your Knowledge Engine on a Score Someone Cleaned the Corpus to Produce
Pinecone Nexus: 47.4% on τ-Knowledge, the top published score — on 698 documents made consistent before measuring. Nobody cleaned yours.

Measuring Document AI Readiness: Why a Well-Scored Corpus Can Still Be a Broken One
7% of enterprises say their data is AI-ready — they consider it, no byte audited. The document corpus is the one layer measurable on evidence.

AI Already Writes a Third of Your Company's Documents. No 2026 Trend Asks If They're Reliable.
AvePoint 2026: 35.5% of document volume is already AI-generated, 42% within a year. Gartner's six trends govern the agent, never the corpus.

43% of Data Breaches Now Involve Shadow AI. Document Reliability Doesn't Show Up in Either Report.
IBM 2026: shadow AI jumped from 20% to 43% of data breaches. Locking down access says nothing about the reliability of the documents AI reads.

88% of AI Agent Pilots Never Reach Production. No Official Diagnostic Checked Your Documents.
Forrester/Anaconda, June 2026: 88% of AI agent pilots fail. None of the three causes cited ever asks about the corpus the agent actually read.

Gartner Just Ranked AI Agents Into Four Autonomy Levels. Document Reliability Isn't One of Them.
Gartner's May 2026 framework calibrates governance on an agent's ability to act. Not on the reliability of the document it just read to decide what to do.

AI Act Article 50: Labeling AI Content Leaves Your Source Documents Out of Scope
Article 50 applies from 2 August 2026. Marking an AI output proves an AI produced it — not which documents fed it, or whether they still carried authority.

Your Knowledge Platform Heals Itself. Nobody Is Watching Your Other Document Silos.
Bloomfire just won a 2026 award for self-healing knowledge. That mechanism heals one silo — the rest of the document estate keeps drifting unwatched.

Continuous AI Act Compliance vs. a Moving Calendar: What the Second Delay in Eighteen Months Means
Regulation 2026/1744 delays Annex III to December 2027 — but Article 50 still applies from 2 August 2026. Second delay in eighteen months: what to build.

Your AI Agents Are Governed Now. The Documents They Read Still Aren't.
Gartner just published its first AI Governance Platforms Magic Quadrant. But none of them governs the documents agents read — 2026's blind spot.

The AI Act's Traceability Deadline Just Moved to 2027. That Doesn't Buy You Time.
The Digital Omnibus pushes AI Act Articles 12-13 to December 2027. But document traceability can't be improvised — the deferral doesn't buy you time.

Extracting a Document Isn't the Same as Making It Trustworthy for AI
Apryse ships a new AI OCR engine. But a perfectly extracted document isn't trustworthy for it — extraction judges neither freshness nor authority.

Document Drift: Why Your RAG That Worked at Launch Is Lying Six Months Later
A corpus that's clean at launch doesn't stay that way. Document drift quietly degrades your RAG — with no alert, until someone finally checks.

RAG Isn't Broken. Your Documents Are.
67% of RAG failures trace to data quality, not retrieval. RAG isn't broken — your documents are. What the market is finally admitting in 2026.

Why Document Lifecycle Management Outweighs Model Choice in 2026 AI ROI
The model battle is settled. What quietly breaks enterprise AI in 2026 is the document corpus lifecycle — not the LLM you pick.

Copilot Ships by Default in Microsoft 365: What It Reveals About Your Document Governance
Since July 1, Copilot ships by default in M365 Business. Not security, not permissions: what it exposes is your document governance debt.

AI-Ready Documents: The Concrete Markers That Separate Usable Content From Merely Stored Files
Gartner: 60% of enterprise AI projects abandoned for lack of AI-ready data. Five verifiable markers to tell an AI-ready document from a merely stored one.

EU AI Act Article 10: What Data Governance Requires of Your Document Corpus
The EU Council formally adopted the Digital Omnibus on 29 June 2026. Article 10 on data governance is untouched: what it requires of an enterprise RAG corpus.

Context Engineering Fails at the Source: Why Knowledge Compilation Needs a Clean Corpus First
Pinecone, Glean, Vectara, Sinequa move RAG toward context compilation. None address what happens when the source corpus itself is contradictory.

Cross-Source Conflation: When RAG Gets the Fact Right and the Source Wrong
An enterprise RAG can answer correctly and still cite the wrong source. A silent failure mode, missed by standard benchmarks, that breaks provenance.

AI Document Quality Scorecard: Rate Your Corpus RAG-Readiness in 30 Minutes — and Know Where to Focus
72% of enterprise RAG deployments fail in the first year. Is your corpus the reason? A 5-dimension scoring framework gives you a diagnosis in 30 minutes.

Model Context Protocol and Enterprise Documents: What MCP Standardizes — and Leaves Ungoverned
MCP standardizes access to enterprise document sources for LLM agents — not their quality. What RAG architects must understand before connecting SharePoint.

EU AI Act August 2: The Article 50 Your Legal Team Underestimated — What Internal Chatbots Must Display
The Digital Omnibus postponed Annex III — not Article 50. By August 2, every internal chatbot must show an AI disclosure. The deployer carries it.

Unstructured Data Catalog vs Document Knowledge Platform: Why the Confusion Is Costing CIOs Millions in 2026
An unstructured data catalog is not a Document Knowledge Platform. Four major vendors blurred the line in 2026. What your CDO needs to know before choosing.

Agentic AI Governance 2026: 53% Operate Without Policies — Document Quality Is the Missing Foundation
A Sinequa survey of 740 executives found 53% have no agent-specific governance policy. The blind spot: the quality of the corpus those agents read.

RAGAS, DeepEval, LettuceDetect: Why Standard RAG Evaluation Frameworks Are Blind to Corpus-Level Failures
RAGAS, DeepEval, LettuceDetect measure fidelity to what was retrieved — not the reliability of your corpus. The blind spot in your RAG evaluation.

Atlassian Will Train Its AI on Your Confluence Data. What Enterprise Leaders Must Decide Before August 17.
Atlassian turns on Confluence and Jira AI training by default on August 17. The CIO reflex isn't where to click — it's auditing the corpus state first.

What Is a Document Knowledge Platform (DKP) — How It Differs from ECM, GED, and RAG: A 2026 Guide
ECM, SharePoint, GED, RAG — and DKP: not synonyms. Definition of a Document Knowledge Platform and how to know whether you need one.

RAGOps: The Data Half Nobody Operates — SLIs, SLOs, and a Control Plane for Corpus Health
RAGOps includes continuous corpus management. A year after the founding paper, nobody operationalizes it. Here are the SLIs of a production corpus.

Agentic AI Without Document Foundations: Why 64% of Enterprises Are Building on Sand
Hyland GA's its Context Engine, Semarchy quantifies the MDM gap. No one names the upstream layer: the document corpus.

EU AI Act August 2026: what the Digital Omnibus did not postpone — and the 60-day corpus plan
The Digital Omnibus pushed Annex III to December 2027 — not Article 50, not AI literacy, not Annex IV for systems already on the market. 60-day corpus plan.

Context engineering done right: why this post-RAG paradigm needs a clean corpus
Anthropic, Glean, LangChain, Pinecone, LlamaIndex set the grammar of context engineering. All work downstream of the corpus. No one says it.

RAG Doesn't Solve Hallucination, It Postpones It. The Failure Mode No One Talks About: Cross-Source Contradictions
Enterprise RAG is the 2026 default. Yet production deployments fail in series — and the root cause is neither the embedding nor the LLM.

AI Readiness Assessment 2026 — the 'Corpus' pillar every framework leaves out
Five 2026 AI Readiness frameworks — Cisco, Microsoft, Cloudera, Iris.ai, Atlan. None makes the document corpus a stand-alone pillar. Here is the gap.

Knowledge Graph vs. vector database for enterprise RAG — start with the corpus
Microsoft, Pinecone, Neo4j, Glean, Squirro, Writer all shipped a 'Knowledge Graph vs vector database' guide. Here's what neither architecture fixes.

If Copilot can't find your SharePoint documents, the bug isn't in Copilot — it's in your corpus
Microsoft hit 20M paid Copilot seats. On Microsoft Q&A, the same complaint keeps surfacing: Copilot can't retrieve our SharePoint documents.

The AI Act on the business side: the bank-insurance Compliance Director's checklist
The EU AI Act and the bank-insurance Compliance Director: 6 concrete obligations to put in place before August 2, 2026, and the trade-offs being made now.

Knowledge AI vs. Knowledge Management vs. DKP: untangling 3 enterprise AI categories
Three terms that sound alike, three categories that do entirely different work. The buyer's decoder for KM, Knowledge AI and DKP — and why conflating them…

Auditing an enterprise document corpus for AI — the K-AI 6-axis method
Anomalies, conflicts, divergent duplicates, unmarked obsolescence, traceability, freshness: six measurable axes we instrument before any serious AI deployment.

You think your RAG hallucinates because of the embedding model? Look at your corpus.
Pinecone has just admitted the model is no longer the bottleneck in enterprise RAG. Three numbers point to a different culprit: corpus rot.
