HomeNewsThe CNIL's new approach to AI models Memorisation, extractability and anonymisation under the GDPR
8 October 2026Dr. Tobias Höllwarth

The CNIL's new approach to AI models Memorisation, extractability and anonymisation under the GDPR

As generative AI becomes increasingly powerful, the CNIL’s 2026 guidance shifts the GDPR debate from how personal data is used in training to whether AI models themselves can retain and reveal information about individuals. By focusing on memorisation, extractability, and anonymisation, the framework offers a practical roadmap for assessing when AI models may still pose privacy risks and therefore remain subject to European data protection law.

The CNIL's new approach to AI models  Memorisation, extractability and anonymisation under the GDPR

The rapid deployment of generative artificial intelligence has forced regulators to confront a question that did not exist when the General Data Protection Regulation (GDPR) was adopted: can a trained AI model itself fall within the scope of data protection law?

Until recently, most discussions surrounding GDPR compliance in artificial intelligence focused on training datasets. Regulatory scrutiny centred on whether personal data was collected lawfully, whether appropriate transparency notices had been provided, and whether a valid legal basis existed for processing. However, the emergence of large language models and foundation models has shifted attention to a different issue: what happens after training has taken place?

In January 2026, the French data protection authority, the Commission Nationale de l'Informatique et des Libertés (CNIL), published guidance addressing the status of AI models under the GDPR. Rather than adopting a categorical position, the CNIL proposes a technical and risk-based assessment centred on three key concepts: memorisation, extractability and anonymisation.

The significance of this guidance extends beyond France. While the CNIL's analysis is rooted in established GDPR principles, it provides one of the clearest indications yet of how European regulators may assess the legal status of AI models themselves. As organisations across Europe continue to develop and deploy generative AI systems, the CNIL's framework offers important insight into the future direction of privacy regulation in the age of artificial intelligence.

Moving beyond training data
The starting point for any GDPR analysis is the definition of personal data contained in Article 4(1), which refers to "any information relating to an identified or identifiable natural person". Traditionally, the question has been whether personal data exists within a dataset. AI systems challenge this assumption because the information used during training is transformed into model parameters rather than stored in a conventional database.

For many years, this distinction encouraged the view that once data had been processed through machine learning systems, the resulting model could be treated as fundamentally different from the training data itself. Under this approach, privacy concerns primarily related to the collection and use of training data, rather than the model produced through training.

The CNIL's guidance suggests that this distinction may be overly simplistic.

Rather than asking whether a model was trained using personal data, the regulator focuses on whether the model itself may continue to contain or reveal information relating to identifiable individuals. This shift is important because it directs attention towards the behaviour and capabilities of the model rather than merely the provenance of the data used to create it.

The result is a more nuanced assessment that reflects the technical realities of modern AI systems.

Memorisation as a regulatory concern
The first component of the CNIL's analysis concerns memorisation.
Machine learning models are designed to identify patterns and relationships within training data. Ideally, a model learns generalisable features rather than reproducing specific examples. But important to note, researchers have demonstrated that large language models can sometimes reproduce portions of their training data, including names, contact information, code fragments, and other potentially sensitive content.
This phenomenon, commonly referred to as memorisation, has become a growing concern within both technical and regulatory communities.
The CNIL does not suggest that memorisation automatically renders a model subject to the GDPR. Such an approach would likely be impractical given the complexity of modern machine learning systems and the varying degrees to which information may be retained during training.
Instead, memorisation serves as an indicator that further analysis may be required. If a model is capable of retaining information relating to identifiable individuals, questions arise regarding whether that information remains accessible and whether privacy risks persist after training has been completed.
In this respect, memorisation represents the starting point rather than the conclusion of the legal analysis.

Extractability and the GDPR's identifiability test
The central element of the CNIL's framework is extractability.

The GDPR does not require regulators to consider every theoretical possibility of identification. Recital 26 provides that account should be taken of "all the means reasonably likely to be used" to identify an individual. This principle has long been interpreted as requiring a contextual and risk-based assessment rather than a purely theoretical inquiry.

The CNIL applies this reasoning directly to AI models.

The relevant question is not whether information could conceivably be extracted under laboratory conditions, but whether personal data can be recovered through means that are reasonably likely to be employed in practice. This assessment may include consideration of techniques such as model inversion attacks, membership inference attacks, adversarial prompting and other methods designed to reveal information embedded within model parameters.

By focusing on extractability, the CNIL effectively translates existing GDPR concepts into the context of artificial intelligence.

This approach is particularly notable because it avoids two extremes. On one hand, it rejects the argument that AI models should automatically be treated as personal data merely because they were trained on personal information. On the other hand, it also rejects the assumption that training necessarily removes all privacy risks.
Instead, the legal analysis depends on whether information relating to identifiable individuals can realistically be obtained from the model itself.

This reflects a broader trend within European privacy law towards assessments based on actual risk rather than formal categorisation.

The role of anonymisation
The third pillar of the CNIL's framework concerns anonymisation.
Under European data protection law, information falls outside the scope of the GDPR only where individuals are no longer identifiable by any means reasonably likely to be used. This threshold has traditionally been interpreted strictly by both regulators and courts.

The CNIL applies the same logic to AI systems.

The fact that information has been transformed into model parameters does not automatically mean that it has been anonymised. The decisive question remains whether personal information can still be linked to identifiable individuals through extraction, reconstruction or inference.

This is perhaps one of the most significant aspects of the guidance. It demonstrates that anonymisation in the context of AI is not determined by the technical process used during training, but by the practical consequences of that process.

A model will only fall outside the scope of the GDPR where the risk of identification has been reduced to a level consistent with established anonymisation standards.
This places a considerable burden on organisations seeking to argue that trained models no longer contain personal data. It also reinforces the importance of robust technical safeguards designed to reduce memorisation and extraction risks.

Why the guidance matters beyond France

Although the guidance originates from the French supervisory authority, its significance extends well beyond the French legal framework.

The CNIL has long played an influential role in European data protection regulation and frequently contributes to broader discussions within the European Data Protection Board (EDPB). As a result, its positions often influence wider regulatory thinking across the European Union.

Important to remember is that the issues addressed by the guidance are not unique to France. Questions concerning memorisation, model inversion attacks, extraction risks and anonymisation affect virtually every organisation developing or deploying generative AI systems. Whether an organisation is based in Paris, Dublin, Berlin or Madrid, the underlying technical challenges remain largely the same.

The guidance may therefore provide an early indication of how other European regulators will approach similar issues in the coming years.

This is particularly important given the growing interaction between the GDPR and the EU AI Act. While the AI Act introduces a dedicated framework for AI governance, it does not replace existing data protection obligations. Organisations will increasingly be expected to navigate both regimes simultaneously, ensuring compliance not only with AI-specific requirements but also with established privacy principles.

Practical implications for organisations
The CNIL's framework has several practical implications for organisations developing or deploying AI systems.

First, organisations should avoid assuming that a trained model automatically falls outside the scope of the GDPR simply because personal data is no longer stored in its original form. The regulatory focus is increasingly shifting towards the capabilities of the model itself.
Secondly, technical assessments of memorisation and extraction risks are likely to become an increasingly important component of privacy compliance programmes. Legal analyses may no longer be sufficient in isolation. Instead, organisations may need to draw upon expertise from machine learning engineers, security specialists and privacy professionals to assess whether identifiable information remains accessible within deployed models.
Thirdly, the guidance reinforces the importance of accountability. Organisations should be prepared to document the basis upon which they conclude that a model does or does not fall within the scope of the GDPR. As with many areas of European data protection law, the ability to demonstrate compliance may ultimately prove just as important as compliance itself.

Finally, organisations should recognise that this area of law remains in development. The technical capabilities of AI systems continue to evolve rapidly, as do the methods available for extracting information from them. Assessments conducted today may require reconsideration as technology advances.

Lessons learned

Conclusion
The CNIL's guidance on the status of AI models under the GDPR represents one of the most significant recent developments in European privacy law.

Rather than creating a new legal category for artificial intelligence, the regulator applies established GDPR principles to a novel technological context. By focusing on memorisation, extractability and anonymisation, the CNIL provides a framework that is both technically informed and legally grounded.

Perhaps most importantly, the guidance shifts the regulatory conversation away from a narrow focus on training datasets and towards the capabilities of AI models themselves. In doing so, it reflects an emerging recognition that privacy risks increasingly arise not only from the collection of personal data, but also from the behaviour of the systems built upon it.
As regulators, organisations and courts continue to grapple with the implications of generative AI, the CNIL's framework may prove to be an important reference point in determining when AI models fall within the scope of European data protection law.

Article provided by INPLP member: Charlotte Gerrish (Gerrish Legal SARL, France)

By Dr. Tobias Höllwarth← All news