Est.
ProcurementLong read

AI Data Processing Addendum Key Clauses

Model training on customer data requires explicit authorization, not silence in your contract.

Contributing Editor · · 12 min read
Cover illustration for “AI Data Processing Addendum Key Clauses”
Procurement · October 1, 2026 · 12 min read · 2,781 words

An AI data processing addendum has to do something a standard vendor contract was never built to do: control what happens to personal data after it disappears into a model. That is the central problem this piece works through, clause by clause, because the gap between a legacy DPA and an AI-specific one has become an active source of legal exposure rather than a drafting nicety. The difference starts at the architectural level. A traditional SaaS processor stores data, processes it, and deletes it on request, and this deterministic, auditable relationship makes risk allocation in the contract fairly straightforward. Once an AI system ingests personal data into its model weights, the data is no longer a discrete record that can be located and erased. It becomes statistical influence distributed across an architecture, produced in outputs that shift as the model keeps training.

That difference is why a single loosely worded sentence in a services agreement can do more damage with an AI vendor than it ever could with a conventional SaaS provider. A clause stating that the "vendor may use AI to improve services" sounds like standard boilerplate. In practice, it can be read to authorize model training, fine-tuning, indefinite retention, the creation of derivative models, and even product expansion built on a customer's own operational data and expertise. In a contract with an AI vendor, that same sentence hands over rights that are extraordinarily difficult to claw back once training has occurred, because the data cannot simply be deleted out of a model the way a row can be deleted out of a database.

The legal environment has caught up to this distinction from several directions at once, and it did so within a short span of time. GDPR Article 28(3) already requires a written DPA before any processor touches personal data, with fines under Article 83(4) awaiting both parties if the agreement is missing or inadequate. The Italian Garante's €15 million fine against OpenAI in December 2024 showed what happens when that requirement collides with a real training pipeline: the finding centered on the absence of an adequate legal basis for processing personal data to train ChatGPT, compounded by transparency failures, inadequate age verification, and a failed breach notification. On the regulatory-guidance side, the EDPB went further than the fine itself, reasoning that personal data unlawfully used to train a model can taint the deployment of that model too, unless it has been properly anonymized, which converts a sloppy training run into a liability that follows the model into every later use. CPPA regulations effective January 1, 2026 add cybersecurity-audit, risk-assessment, and automated-decision-making cooperation duties to CCPA service-provider contracts that 2020-era templates never addressed. And the EU AI Act's transparency obligations went live with enforcement beginning August 2, 2026, after the European Commission confirmed that date at the end of July.

A generic DPA cannot be patched into compliance with a side letter or a footnote. Procurement and legal teams now need to recognize the AI-specific addendum as its own document type, with its own required clauses, the same way they learned years ago to recognize a standard DPA as distinct from the master services agreement it accompanies.

The GDPR Article 28 requirements every AI DPA must satisfy as a mandatory baseline

Before any AI-specific clause gets negotiated, the addendum has to clear a floor that has nothing to do with artificial intelligence at all. Article 28(3) sets eight minimum terms for any processor agreement, and an AI DPA missing even one of them is not a compliant processor contract, no matter how carefully its training restrictions are written. The eight terms require the processor to act only on the controller's documented instructions, bind everyone with data access to confidentiality, implement the security measures described in Article 32, refrain from engaging a sub-processor without authorization, assist the controller with data-subject requests, help with breach notification and impact assessments, delete or return data at the end of the engagement, and make compliance records available for audit. These obligations function as a foundation; the AI-specific terms, such as the training prohibition and the definitional work around data categories, sit on top of that foundation rather than replacing any part of it.

AssemblyAI's published DPA offers a useful picture of what this baseline looks like once a vendor actually writes it into a contract rather than leaving it as a regulatory abstraction. Its documented-instructions clause requires processing only for purposes set out in the underlying agreement and consistent with the customer's documented instructions, and it obligates the processor to inform the customer before processing on any other legal basis, where the law permits that disclosure. The same document defines "Privacy Laws" broadly enough to span EU GDPR, UK GDPR, the Swiss Federal Act on Data Protection, the UK Data Protection Act 2018, the EU's ePrivacy regulations, and U.S. state comprehensive privacy statutes including the CCPA. That breadth is not incidental. It explicitly maps "Controller" and "Processor" to "Business" and "Service Provider" under the CCPA, so the same document carries both EU and California obligations.

One might argue that this baseline is table stakes, unrelated to what makes an AI vendor different from any other processor. The eight Article 28(3) terms are the price of admission, and the training language is the differentiator. A vendor that gets these eight terms right but leaves out how training data will be used has built a contract that looks compliant on a checklist while leaving open the exposure that matters most for an AI relationship. The sections that follow work through where that exposure tends to live.

The documented-instructions clause is the hinge on which every AI-specific restriction turns

The first of the eight Article 28(3) terms, the requirement that a processor act only on the controller's documented instructions, looks purely procedural until an AI vendor is on the other side of the table. For a conventional processor, documented instructions describe what to do with data that is stored and returned. For an AI vendor, the same clause has to answer a much harder question: does training a model on customer data fall inside those instructions, or outside them? Under GDPR Article 6, every distinct act of processing needs its own lawful basis, and training a model is a distinct act, not an incidental byproduct of delivering the underlying service. If the documented-instructions clause does not explicitly rule training out, a vendor has room to argue that training falls within the ordinary scope of "providing the service," and that argument can be made in good faith, because the contract never foreclosed it.

This is exactly the gap that produced the Garante's finding against OpenAI. The Garante/OpenAI case turned on exactly this gap: processing personal data for training without an adequate lawful basis was one of several findings, alongside transparency failures, a data breach notification failure, and inadequate age verification, that produced the fine, which has since been annulled by the Court of Rome on jurisdictional grounds. The jurisdictional reversal does not undo the underlying lesson: a documented-instructions clause that fails to name training as either authorized or prohibited leaves the question to be litigated after the fact, when the data has already gone into the model.

The clause also has to handle the opposite scenario, where a law requires the processor to do something the controller never instructed. AssemblyAI's DPA handles this directly, requiring the processor to inform the customer of any legal requirement to process data before doing so, to the extent the law allows that disclosure. OpenAI's DPA applies the same logic to law-enforcement requests specifically, committing to notify the customer, where legally permitted, if OpenAI receives a legally binding demand for disclosure of customer data. That commitment matters because AI vendors increasingly sit in the path of exactly this kind of request, and a documented-instructions clause that only addresses ordinary business processing, without addressing compelled disclosure, leaves a customer blind to exactly the scenario where visibility matters most.

A documented-instructions clause protects a customer only when it affirmatively lists what the vendor may do with the data, rather than prohibiting whatever the customer happens to think of first. A prohibition-only clause leaves every unlisted activity in a gray zone that a vendor's lawyers can argue either way. An affirmative enumeration closes that gray zone before it opens, and it sets up the training prohibition discussed next as the specific mechanism that enforces what the documented-instructions clause establishes in principle.

The AI training prohibition's required scope

If the documented-instructions clause establishes the principle that training needs its own authorization, the training prohibition is where that principle either holds or collapses under specific contract language. For most enterprise customers, this is the single most consequential clause in the entire AI addendum, and broad drafting is not a stylistic preference: narrow drafting creates exploitable gaps that vendors use, sometimes without any bad intent, simply because the contract left the gap open. A properly scoped prohibition has to cover customer data in every form it takes inside the vendor's systems: prompts, inputs, outputs, derivative outputs, metadata, and behavioral interaction data. It also has to reach every activity that touches a model, not just "training" in the narrowest technical sense, extending to fine-tuning, optimization, and any other modification of a model made available to third parties.

That breadth matters because vendors can and do draw a technical distinction between "training" and adjacent activities like "optimization" or "safety calibration," even when those activities have the same practical effect on customer data. A prohibition that names only "training" leaves those adjacent categories untouched. Prohibiting the use of any input or output for any model modification, rather than just for "training," closes that particular door.

Some platforms address this by design: architecture where customer data never enters a training pipeline by default, with no opt-out required because there is no pathway for the data to take. That kind of design changes the risk profile of the relationship, because the contractual prohibition is reinforcing what the system already does rather than constraining what it would otherwise do. Even so, the contract still has to say so explicitly. An architecture that excludes customer data from training is only as reliable, from a customer's standpoint, as the written commitment that describes it, because architecture can change with a product update in a way a signed addendum cannot.

A common and expensive failure occurs when procurement reviews the master services agreement, finds a clean training opt-out, and assumes the matter is settled, without checking whether the DPA itself contains more permissive language for specific processing activities. The conflict-resolution provision in the DPA frequently overrides the MSA's intent when the two documents disagree about training.

The most contested point in current negotiations is the anonymization carve-out, the clause that lets a vendor use "aggregated" or "anonymized" versions of customer data even where training on identifiable data is banned. The CNIL's penalty against Clearview established that data being "publicly available" does not exempt it from Article 6's lawful-basis requirement, and a parallel logic applies to "anonymized" data whenever re-identification remains technically plausible. Practitioners disagree about whether anonymization carve-outs can ever be acceptable, but the stronger position treats them as acceptable only with explicit technical specification, naming the anonymization standard used and who verifies it, rather than accepting a vendor's self-certification at face value. A best-practice checklist now asks for a governance warranty confirming those controls, including ISO 42001 certification where applicable, a standard that simply did not exist in earlier generations of DPA templates.

Government contracting offers a useful benchmark for how far this language can go. The U.S. GSA's proposed 2026 contracting clause explicitly prohibits using government data to train, fine-tune, or otherwise improve any AI model for any other customer or commercial purpose, which is about as unambiguous as this kind of restriction can be written. Enterprise customers negotiating commercial contracts have reason to ask for language at least that direct.

Precise definitions in the addendum determine what the training prohibition protects

A training prohibition is only as strong as the definitions that tell a reader what it covers. An addendum that bars training on "customer data" without defining the term is protective only to the extent of the vendor's narrowest possible reading of that phrase. This is where a great deal of negotiating leverage quietly disappears, because definitions read as dry boilerplate and get far less scrutiny than the operative prohibitions they support.

A properly built addendum defines customer data, output data, training data, telemetry data, derived data, model artifacts, embeddings, and vector databases as separate categories, each with its own treatment. That separation is not academic. Customer content, meaning both the inputs a customer submits and the outputs the system returns, warrants a strict no-training restriction paired with a strong deletion right. Usage telemetry, such as latency data or system performance metrics, carries a different risk profile and can reasonably be subject to broader vendor rights, since it reveals little or nothing about the substance of what a customer processed. Practitioners report that this bifurcated approach, tight restrictions on content and looser ones on operational telemetry, produces more workable negotiations than an absolutist demand that every category of data receive identical treatment.

The category that resists easy definition is the one covering embeddings and vector databases. Once customer data has been encoded as a vector embedding, it no longer exists as a legible record a person could read, but it has not necessarily been anonymized either, since embeddings can sometimes be reversed or matched back to source content under the right conditions. An addendum has to answer several concrete questions about this category: what happens to embeddings at the end of the engagement, whether they persist after termination, whether prompt histories are retained separately from the embeddings derived from them, whether outputs get logged independently, and whether any of the retained artifacts could allow information to be reconstructed.

A parallel illustration of why precise categorization matters comes from outside the DPA context entirely, in how the EU AI Act treats intended use. The lesson transfers directly to data categorization in a DPA: when a system's intended function and the data categories it touches are defined narrowly and precisely, both the vendor's obligations and the customer's exposure are fixed by those definitions. Loose definitions leave both sides negotiating against an addendum that can be read to mean almost anything once a dispute actually arises.

Sub-processor governance: maintaining the training prohibition through the full vendor chain

The sub-processor agreements underneath the primary DPA limit how far its protections extend. Under GDPR Article 28(4), the primary processor stays fully liable for every sub-processor it engages, so a customer's contractual protections are only as strong as the weakest link in that chain. That fact turns sub-processor governance from a housekeeping clause into a direct extension of the training prohibition itself.

A properly governed sub-processor arrangement needs several elements working together: specific or general written authorization, a current and published list of sub-processors, advance notice before any new sub-processor is added (Article 28 prescribes no specific period; 30 days is a common contractual standard), a right for the customer to object, and a termination right if that objection goes unresolved, the last of which is how OpenAI's DPA structures the mechanism. Each element closes a different way a sub-processor relationship could otherwise operate outside the customer's visibility.

None of that matters, though, unless the substantive protections in the primary DPA, especially the training prohibition, actually flow down to every sub-processor by contract. A primary DPA that bans training but never requires its sub-processors to honor the identical prohibition has recreated the exact gap it was written to close, just one layer removed from where the customer can see it. This is where sub-processor governance and the training prohibition become inseparable: one defines the rule, the other makes sure the rule survives contact with every vendor downstream of the primary relationship.

The structure gets genuinely complicated for AI platforms built on top of foundation model providers treated as sub-processors. The foundation model provider carries its own conformity obligations independent of any customer relationship, obligations that belong to it regardless of what any downstream contract says. But the primary vendor's DPA still has to ensure that the customer's data-processing restrictions bind how that foundation model actually interacts with the customer's data, because the customer has no direct contractual relationship with the foundation model provider at all. Every commitment the customer relies on, including the training prohibition negotiated at the top of the chain, has to be re-created contractually at every layer beneath it, or it stops applying at the exact point where the data leaves the primary vendor's direct control.

Sources

  1. Data Processing Addendum
  2. AI Addendums: Contract Clauses to Negotiate with AI Vendors
  3. AI Data Processing Addendum (DPA): Template, GDPR Article 28 Terms & 2026 Rules
  4. OpenAI Data Processing Addendum — OpenAI
  5. Law & Compliance in AI Security & Data Protection
Filed underProcurement

More in Procurement