← Back to the blog

10 August 2026 · Gilles Maury

Can you send your customers' data to an AI without risking a GDPR fine?

GDPRAIcomplianceriskdata protection

A confusion keeps coming up in the AI projects I run into: people talk about "anonymizing" data before sending it to an AI, thinking that's enough to fall outside GDPR. In the vast majority of cases, that's wrong — and the confusion can get expensive.

The difference that changes everything

Anonymization = data made permanently impossible to link back to a person. Once truly anonymized, it falls outside the scope of GDPR.

Pseudonymization = data where identity is masked but remains recoverable via a separate key. It stays under GDPR, with every obligation that implies.

The problem: what's sold to you as "anonymization" is almost always, in reality, pseudonymization.

Why true anonymization is a myth in practice

Researchers from UCLouvain and Imperial College London demonstrated that it's possible to precisely re-identify individuals within datasets presented as "anonymized," using machine learning. This isn't a theoretical hypothesis — it's a published research result.

The practical consequence: if a vendor guarantees total, irreversible anonymization of your data, ask them to explain exactly how — in most cases, the honest answer is that it's actually pseudonymization.

What to demand, concretely

Rather than chasing an anonymization that almost never exists, the right practice is rigorous pseudonymization:

  1. Replace identifiers (names, emails) with random identifiers before anything is sent to an AI
  2. Store the lookup table separately, encrypted, with strictly limited access
  3. Never transmit that key to the AI vendor
  4. Have a signed Data Processing Agreement (DPA) with that vendor

Concrete example: instead of sending "Client #5293, Alice Smith, [sensitive data]", you send "ID uuid-a7f9-4e2b, [sensitive data]" — the key that lets you trace back to Alice stays with you, encrypted, kept separate from the rest.

The point that trips up the most companies: the contract with your AI vendor

A DPA (Data Processing Agreement) is mandatory as soon as personal data is processed. Yet most consumer AI offerings don't provide one.

Vendor Is a DPA available? From which plan
OpenAI / ChatGPT Yes Business plan (minimum 2 seats)
Anthropic / Claude Yes Team plan (minimum 2 seats)
Mistral AI Yes Enterprise plan (EU hosting, Paris)
Grok / xAI No No professional plan with a DPA to date

Free or individual plans (ChatGPT Plus, Claude Pro, etc.) include no DPA at all — using them with your customers' or employees' data is a compliance risk, no matter how careful the pseudonymization upstream.

And a DPA alone isn't enough either: you also need to document the processing (Article 30 register), carry out an impact assessment when the processing warrants it, and train the teams who use these tools day to day.

Why a fallback plan matters as much as the initial choice

Depending on a single AI vendor for a sensitive business process (HR, healthcare, finance) is a risk in itself: a change in terms of service or an account suspension can block your operations overnight, with an emergency recovery cost far higher than the cost of having planned an alternative from the start.

This is a principle I apply to every architecture I build: your system should stay reversible — able to switch vendors without rebuilding everything — rather than locking you into a single provider you'd depend on entirely.

Update: an expert publicly corrected me, and they were right

This article prompted a detailed reply from a professional who works on AI compliance day to day. I'm choosing to share it and correct myself publicly rather than let an inaccuracy slide, because that's exactly the standard I apply before proposing anything to a client.

Three important corrections:

  1. Health data requires specific hosting. My original example understated this: processing health data requires hosting specifically certified for it (HDS in France), not standard cloud hosting even if GDPR-compliant otherwise. Without that certification, no health data — whatever the price.
  2. A DPA doesn't always guarantee 100% European processing. Some vendors presented as "European" don't contractually guarantee that 100% of processing stays in Europe. For sensitive data, this needs to be checked explicitly rather than trusting a vendor's advertised nationality.
  3. The impact assessment (DPIA) is too often skipped. Having a processing register isn't enough: as soon as a process carries significant risk (typically in HR), a formal impact assessment is required on top of it.

Something this expert also rightly reminded me of: too many AI tools in sensitive sectors (healthcare, HR) operate today in a gray zone, exposing the people ultimately responsible to risks they don't always measure themselves. That's exactly what I try to avoid: a client discovering after the fact that they were carrying a risk nobody told them about.

What this means for your project

If you're considering integrating AI into business processes that touch customer, prospect or employee data, three questions are enough to gauge your exposure: does your vendor have a signed DPA, is your data genuinely pseudonymized with a separate key, and do you have a plan if that vendor changes its terms tomorrow? If even one answer is no, that's the starting point for getting compliant — contact me if you'd like to look at it together.