Est.

Retrieval-Augmented Generation and Cross-Tenant Data Leakage

RAG systems leak data across tenants by design, not accident.

Staff Writer, Incident Analysis and Risk · · 9 min read
Cover illustration for “Retrieval-Augmented Generation and Cross-Tenant Data Leakage”
Data Exposure Scenarios · October 6, 2026 · 9 min read · 1,996 words

Retrieval-augmented generation, or RAG, is now the default way companies put large language models to work on their own data, and RAG solves three problems that a plain language model cannot solve on its own: it keeps answers current by pulling from a live knowledge base instead of a frozen training set, it cuts down on hallucination by grounding what the model says in actual retrieved documents, and it lets a company use its own proprietary records without the expense of retraining or fine-tuning a model from scratch. The pipeline behind this is simple to describe: a user asks a question, that question triggers a search against a vector database, the chunks of text that come back get inserted into the model's context window next to the original question, and the model writes an answer grounded in whatever it was just handed. The knowledge base feeding this pipeline usually holds a company's internal wikis, CRM records, financial statements, legal contracts, and source code, so the same real-time retrieval that makes RAG useful is the mechanism by which that material can end up in front of the wrong person.

The three pipeline stages where cross-tenant leakage originates

Cross-tenant leakage in RAG systems is not something that happens because an engineer configured a system carelessly. The same three properties that make RAG work in the first place also cause cross-tenant leakage: tenant-agnostic vector search, permission loss at ingestion, and access-control staleness. If isolation isn't enforced at the index level, the search simply walks through everyone's data, because the algorithm sees only one data set to search. The second is permission loss at ingestion. Most vector databases end up storing embeddings as a flat pool with no inherited structure. The access rules that governed a file in its source system don't carry over when that file becomes a chunk in a vector index. By the time a retrieved chunk reaches the model, there's no trace left of which department, project, or named list of users the original document was restricted to. The third is access-control staleness. Permissions change constantly, someone leaves a team, a project gets reclassified, a contractor's access gets revoked, but enforcement in most pipelines only happens once, at ingestion time. A user who loses access on paper can keep pulling the same sensitive documents out of the pipeline for hours afterward, until someone reindexes. None of these three causes is a bug waiting to be patched.

A shared index and cross-tenant retrieval as a normal outcome

The most common way companies build multi-tenant RAG systems is to put every tenant's vectors into one shared index and separate them with a metadata filter, something along the lines of where tenant_id = 123. It's cheap, it's fast to build, and it scales without needing a new infrastructure footprint for every new customer. The trouble is that this pattern is only as strong as that one filter. If the filter fails, gets misapplied, or is simply absent from a given query path, no isolation remains, because the underlying index never distinguished between tenants. A team might try to patch this by doing post-retrieval filtering: pull the top-k most similar chunks from the shared index, then drop whichever ones the requesting user isn't allowed to see before the results reach the model. On its face that looks like a reasonable fix. It creates two problems of its own. First, the ranking itself is already contaminated, because the similarity scores were computed against the entire corpus, including documents the user will never be shown. The final filtered list isn't a clean top-k result for that user; it's a damaged version of one, with gaps where the best-matching but restricted documents used to be. Second, the act of filtering can itself leak information: timing differences and subtle shifts in scoring patterns can tell an attacker that a restricted document exists and roughly how closely it matched their query, even when its content never appears in the response. None of this means engineering teams are choosing the shared-index pattern out of ignorance. Building a fully isolated index per tenant multiplies infrastructure cost and operational complexity with every new customer onboarded, and for many companies that tradeoff is real enough to justify the risk, at least until the risk becomes a headline.

Why the "embeddings are safe" assumption accelerates the problem

A common belief among teams building these systems is that storing vector embeddings instead of raw text provides a kind of privacy by obscurity, since embeddings are just long lists of numbers and not readable documents. That belief does not hold up under testing. Research on sentence embeddings has shown they can leak a substantial amount of information about the text that produced them, and in some cases the original sentences can be reconstructed from the embedding itself: the embedding layer behaves less like a one-way function and more like a reversible encoding under the right conditions. The consequence of treating embeddings as inherently safe is that it removes the urgency that would otherwise push a team toward real isolation controls. A documented access-control bypass involving Pinecone exposed over 200,000 healthcare records, a case that shows what happens once retrieval-layer access control fails: the scale of exposure tracks the scale of the index itself, so the bigger and more centralized the knowledge base, the bigger the single point of failure becomes.

What EchoLeak showed about leakage in production systems

Disclosed by Aim Security researchers Pavan Reddy and Aditya Sanjay Gujral in June 2025 and rated critical by Microsoft at a CVSS score of 9.3, the flaw affected Microsoft 365 Copilot integrations across Word, Excel, PowerPoint, Outlook, and Teams. The mechanism required no click, no download, and no mistake by the victim. An attacker crafted an email containing hidden instructions, and once Copilot retrieved that email as part of its normal RAG context for an unrelated query, it carried out the attacker's commands and exfiltrated data: chat logs, OneDrive files, SharePoint content, and Teams messages, sent to a server the attacker controlled. The exploit worked by getting past several layers of defense that Microsoft had already built in, including the XPIA classifier meant to catch cross-prompt injection, link redaction, restrictions on automatic image fetching, and Content Security Policy, with that last barrier bypassed through an allowlisted Microsoft Teams asynchronous preview proxy. A related technique researchers called "RAG spraying" industrialized the attack further: by writing messages designed to score well against the range of topics a user might plausibly ask Copilot about, an attacker could raise the odds that a malicious payload gets pulled into context on some future query, with no targeting of a specific person required. Microsoft shipped a server-side patch in May 2025, disclosed the issue publicly the following month, and said no customer action was needed and no in-the-wild exploitation had been found. What the patch fixed was this specific exploit chain, not the underlying architectural fact that let it work in the first place: a system that retrieves content and acts on it without a reliable way to tell trusted material from attacker-crafted material.

Agentic RAG, multimodal retrieval, and the expanded leakage surface

Two developments pushing RAG forward right now, agentic RAG and multimodal RAG, don't bring new categories of risk with them. They take the same tenant-agnostic retrieval problem described earlier and spread it across a wider surface, often with action channels that carry higher consequences than a plain text answer would. A 2026 preprint by Al-Lawati and Wang, "Do Multimodal RAG Systems Leak Data?" (arXiv:2601.17644), runs a broad evaluation of membership inference and image caption retrieval attacks and finds that the leakage surface extends into images, a retrieval type where the tooling for access control lags well behind what exists for text. Several ordinary design choices inside standard RAG setups make this worse without anyone intending them to: retrieving a larger top-k set of chunks, using early-fusion concatenation to combine sources, or applying query rewriting inside query-based fusion pipelines all increase how much external content gets poured into the model's context window at once. More content in that window means a wider window for cross-tenant material to slip through. A broader RAG survey (arXiv:2407.13193) names security as one of the central open challenges facing industrial RAG deployment, and groups it alongside efficiency and graph-based retrieval as directions the field still has to work out. That a research survey is naming security as unresolved in 2026, well after RAG became standard enterprise practice, says something about how far implementation has outrun the safeguards meant to contain it.

Knowledge base poisoning and the retrieval trust model

A second attack class, separate from cross-tenant leakage but resting on the same structural foundation, is knowledge base poisoning. The language model cannot tell whether content was retrieved from a legitimate, trusted document or deliberately planted to manipulate it. That's a direct consequence of how context injection works, not a flaw specific to any one model that a future update might fix: everything that lands in the context window gets treated as trustworthy input, regardless of where it came from. Cornell research, cited in AI Security and Safety's RAG security guide, found that a single poisoned document placed in a knowledge base reliably caused multiple large language models, including GPT-4, Claude, and Gemini, to follow instructions embedded in that document, in some cases generating phishing content or leaking other documents. A related tactic, retrieval manipulation, crafts a document to score artificially high on similarity against a target query so the system preferentially retrieves it, targeting vector similarity scoring the way search engine manipulation targets keyword ranking. The two attack classes share a root cause: access control at the retrieval layer is necessary but solves only half the problem, because a user with every legitimate permission in place can still be served poisoned content that the pipeline has no mechanism to flag as untrustworthy.

What genuine mitigation requires

Closing the gap described across this piece takes security work at every stage of the pipeline, ingestion, retrieval, and generation, because the root causes span all three and a fix at only one stage leaves the other two open. At ingestion, OWASP's RAG Security Cheat Sheet calls for storing classification, owner, permitted roles, and permitted tenants at the level of each individual chunk rather than at the document level, since document-level permissions don't automatically carry over once a document gets split apart during chunking. That gap, chunking without carrying permissions forward, is the direct, traceable cause of the permission loss described earlier in this piece. Ingestion-stage defenses should also run content integrity checks, such as checksums and digital signatures, to catch unauthorized changes before a poisoned document makes it into the index, and should scan documents for injection patterns, using tools like LLM Guard or Vigil, to catch embedded adversarial instructions before ingestion completes. At the retrieval layer, the clearest fix for the shared-index problem is physical or namespace-based isolation per tenant, which removes the cross-contamination risk entirely but brings back the cost and complexity tradeoff that pushed teams toward shared indices to begin with. If you can't avoid a shared index for cost or scale reasons, you need to enforce access control at query time rather than only at ingestion, and you can run policy caching on conservative time-to-live windows and reevaluate it asynchronously so permission checks stay current without destroying query latency. None of these controls is a single silver bullet, and that's the honest limitation of this whole conversation. Ingestion-stage scanning does nothing for a stale permission that was valid when the document was indexed and invalid by the time someone queries for it. Query-time enforcement does nothing against a poisoned document that passed every integrity check because the person who planted it had legitimate write access. Real mitigation means layering these controls across every stage at once and treating the tradeoffs, cost, latency, and complexity, as the price of running a RAG system at the scale enterprises are now running them.

Sources

  1. Retrieval-Augmented Generation for Natural Language Processing: A Survey
  2. RAG Security - OWASP Cheat Sheet Series
  3. Embedding Inference Attack
  4. EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System
  5. A Systemic Evaluation of Multimodal RAG Privacy

More in Data Exposure Scenarios