AI and GDPR: what the CNIL really says, in plain terms
AI is not exempt from the GDPR. Legal basis, web scraping, your rights over the data used for training: what France's data authority says, without legal jargon.

Many companies think they must choose between using AI and complying with the GDPR. France’s data protection authority (the CNIL) settled the question back in July 2025, with its first recommendations on developing AI systems: the two do not conflict, but AI enjoys no exemption. Here is what this changes concretely, without legal jargon.
The principle: AI is not exempt from the GDPR
An AI system that processes personal data remains subject to the GDPR, full stop. The CNIL states this explicitly in its recommendations: the regulation’s principles apply fully, whatever the technical sophistication of the system. Developing or using an AI model does not create a lawless zone over the data it processes.
Which legal basis to train an AI?
The legal basis most used by private companies to train a model is legitimate interest, and the CNIL specifies its conditions rather than banning it. To be valid, the interest pursued must meet three cumulative conditions: be manifestly lawful under the law, be defined with enough clarity and precision, and be real and present — not merely hypothetical. In return, the CNIL requires strong safeguards: excluding certain categories of data from collection, reinforced transparency towards the people concerned, and making the exercise of their rights genuinely easy.
Web scraping is not forbidden, but it is regulated
Collecting public web data to build a training set — “scraping” — remains possible, but under precise conditions. The CNIL has published specific criteria for this case within the legitimate-interest framework: the collection must stay proportionate to the purpose pursued, and the same safeguards of transparency and rights apply. Just because data is public on the internet does not mean it escapes the GDPR once reused to train a model.
Provider or mere user: very different obligations
As with the AI Act, the GDPR does not impose the same level of requirement on whoever designs an AI system and whoever uses it day to day.
| Your situation | What it implies |
|---|---|
| You develop a model or AI tool with personal data | Legal basis to document (often legitimate interest), triple test, strong safeguards, traceability from the design stage |
| You use an existing AI tool (ChatGPT, Copilot, a business assistant) | Check the plan’s processing terms, inform your teams and clients if personal data is entered, do not paste sensitive data without a guarantee |
| You are a person concerned (your data may have served for training) | Right of access, rectification, objection — the CNIL ensures these rights stay exercisable in practice |
What it changes for your company, concretely
If you develop an internal tool that relies on personal data — clients, employees, applicants — GDPR compliance must be designed in from the start, not added afterwards: data traceability, documentation of choices, a procedure to answer an access request. It is the same reflex already needed for the AI Act, fully in force since 2 August 2026: the two texts stack up and point in the same direction, supervision and documentation. If you merely use off-the-shelf AI tools, the essential lies in a few simple habits when entering data — see the confidentiality habits before pasting your data.
Key takeaway
The GDPR does not hold AI back, it frames it: legitimate interest under conditions, web scraping tolerated but documented, people’s rights preserved. The heaviest obligations target those who design models; companies that merely use off-the-shelf tools mainly have a duty of vigilance and transparency. In case of doubt on a specific project, the CNIL’s practical fact sheets remain the up-to-date reference.
Sources
Frequently asked questions
Does the GDPR apply to generative AI?
Yes, without exception. The CNIL is clear: the principles of the GDPR remain fully applicable whatever the complexity of the AI system, as soon as personal data is involved, in training or in use.
Can a company train an AI without people's consent?
In some cases, yes, via the legal basis of legitimate interest — the most used by private organisations according to the CNIL. It is subject to a triple test: a lawful, precisely defined, and real interest, with strong safeguards (transparency, exclusion of certain data, easier exercise of rights).
Is web scraping to train an AI forbidden?
No, but it is regulated. The CNIL sets out the conditions under which collecting public web data (scraping) can rely on legitimate interest, with safeguards on how it is used.
Was my data used to train ChatGPT, Claude or Gemini?
It is possible if you used a consumer plan whose terms allow that use, or if your public data was collected on the web. You keep a right of access, rectification and objection: the CNIL insists on making these rights easy to exercise as a central safeguard.
What obligations for a company that uses (without developing) an AI tool?
Document the uses, inform the people concerned when personal data is processed, and check the guarantees of the plan used (professional vs consumer). These are first-level obligations, far lighter than those of model developers.