Article · 18 MIN

Pseudonymization vs Anonymization Under the GDPR: When Personal Data Really Falls Outside the GDPR

Pseudonymization vs Anonymization Under the GDPR: When Personal Data Really Falls Outside the GDPR

Under the GDPR, pseudonymized data remains personal data when an individual can still be identified, directly or indirectly.

In its March 4, 2026 decision in the Criteo case, the French Conseil d’État clarified that data can be considered anonymized through pseudonymization only if the risk of identifying the individual is insignificant and identification would be practically impossible because it would require a disproportionate effort in terms of time, cost, and manpower.

This is a critical point for companies using data for advertising, analytics, artificial intelligence, profiling, data-sharing, product development, or research.

The practical rule is simple:

A company cannot declare data anonymous merely because direct identifiers have been removed. It must be able to demonstrate that re-identification is not reasonably possible.

Executive Summary

Many companies treat pseudonymization as if it were anonymization.

This is a legal mistake.

Pseudonymization reduces the direct link between data and an individual. It may replace a name, email address, customer number, or account identifier with a pseudonymous identifier.

However, under the GDPR, pseudonymization is still a processing operation applied to personal data.

It does not automatically remove the data from the scope of the GDPR.

The distinction matters because anonymized data falls outside the GDPR, while pseudonymized personal data remains subject to GDPR obligations, including legal basis, transparency, data subject rights, consent management, security, accountability, and joint controllership obligations where applicable.

In its March 4, 2026 decision involving Criteo, the French Conseil d’État confirmed a strict but practical test:

Data pseudonymized by a controller can be considered anonymous only if the risk of re-identification is insignificant and if re-identification would be practically impossible, especially because it would require disproportionate time, cost, and manpower.

The court rejected the idea that the absence of a direct re-identification key was enough.

In the Criteo case, the pseudonymous identifier remained associated with multiple categories of information, including IP address, geographic location, device identifier, partner identifiers, visited websites, purchases, viewed advertisements, and advertisements that led to purchases.

Taken individually, some of these elements may appear technical or indirect.

Taken together, they can reconstruct a profile.

That is the central lesson of the decision.

Direct Answer

Does pseudonymization make personal data anonymous under the GDPR?

No.

Pseudonymization does not, by itself, make personal data anonymous.

Pseudonymized data may be considered anonymous only if the risk of identifying the data subject is insignificant and if re-identification would be practically impossible because it would require a disproportionate effort in terms of time, cost, and manpower.

If the individual remains reasonably identifiable, directly or indirectly, the data remains personal data under Article 4(1) of the GDPR.

Why This Topic Matters for Companies

The distinction between pseudonymization and anonymization is not theoretical.

It affects the legal architecture of many modern business models.

Companies increasingly rely on large datasets for:

  • targeted advertising;
  • website analytics;
  • mobile app tracking;
  • artificial intelligence training;
  • customer segmentation;
  • fraud detection;
  • product personalization;
  • behavioral analytics;
  • research and development;
  • data monetization;
  • partner data-sharing;
  • and internal business intelligence.

In each of these contexts, organizations often try to reduce legal risk by removing obvious identifiers.

The problem is that removing direct identifiers does not necessarily eliminate the possibility of identifying a person.

If the remaining data still allows a person to be singled out, linked to a profile, or identified through cross-referencing, the GDPR may continue to apply.

This is why the legal classification of data is now a strategic governance issue.

The Legal Framework: Personal Data, Pseudonymization, and Anonymization

What Is Personal Data Under Article 4(1) GDPR?

Article 4(1) of the GDPR defines personal data as any information relating to an identified or identifiable natural person.

A person is identifiable if they can be identified directly or indirectly, including by reference to:

  • a name;
  • an identification number;
  • location data;
  • an online identifier;
  • or one or more factors specific to their physical, physiological, genetic, mental, economic, cultural, or social identity.

This definition is intentionally broad.

It does not require that the company already knows the person’s civil identity.

It is enough that the individual can be identified directly or indirectly.

What Is Pseudonymization Under Article 4(5) GDPR?

Article 4(5) of the GDPR defines pseudonymization as the processing of personal data in such a way that the data can no longer be attributed to a specific data subject without the use of additional information.

That additional information must be kept separately and protected by technical and organizational measures.

This definition is important because it confirms that pseudonymization is not the same as anonymization.

Pseudonymization is a security and data-protection technique.

It reduces risk.

It does not necessarily eliminate personal data status.

What Is Anonymization?

The GDPR does not apply to anonymous information.

Data is anonymous when the individual is no longer identifiable by means reasonably likely to be used.

This requires a practical assessment.

The question is not whether identification is theoretically imaginable in an abstract sense.

The question is whether re-identification is reasonably possible in practice.

This is where the Conseil d’État’s March 4, 2026 decision is important.

It confirms that anonymization requires an insignificant risk of identification and practical irreversibility in light of the effort required.

The Criteo Case: Why the Conseil d’État Refused to Treat the Data as Anonymous

The Criteo case arose from a CNIL sanction decision involving online advertising, cookies, user tracking, consent, transparency, and joint controllership issues.

One of Criteo’s arguments was that the data it processed was not personal data because each user was assigned a pseudonymous identifier.

The Conseil d’État rejected this argument.

The court examined the data ecosystem around the pseudonymous identifier.

It found that the identifier was not isolated.

It was associated with several categories of information, including:

  • the IP address of the user’s terminal;
  • geographic location linked to the IP address;
  • the device identifier;
  • partner-specific identifiers;
  • websites visited;
  • purchases made;
  • advertisements viewed;
  • and advertisements that resulted in purchases.

The court emphasized that Criteo’s processing aimed to deliver advertising adapted to users’ consumption habits and interests.

That purpose required the collection and cross-referencing of many elements linked to a given identifier.

In that context, the absence of a direct re-identification key was not decisive.

The court concluded that at least some individuals were identifiable by means that did not require disproportionate effort in terms of time, cost, and manpower.

Therefore, the data remained personal data under Article 4 GDPR.

The Core Legal Rule from the Conseil d’État

The rule can be summarized as follows:

Pseudonymized data can be treated as anonymous only where the risk of identifying the individual is insignificant and where identification would be practically impossible because it would require disproportionate effort in time, cost, and manpower.

This rule matters because it rejects two common assumptions.

First, the absence of names or direct identifiers is not enough.

Second, the absence of business interest in re-identifying individuals is not enough.

The relevant question is not whether the company wants to re-identify the person.

The relevant question is whether the company, or another reasonably likely actor depending on the context, can re-identify the person without disproportionate effort.

The Most Important Practical Lesson

The decisive issue is not what the dataset looks like at first sight.

The decisive issue is what the dataset can reveal when combined with other information.

A pseudonymous identifier may seem harmless when viewed alone.

However, when it is linked with behavioral, technical, location, device, browsing, and transaction data, it may allow the reconstruction of a detailed individual profile.

This is why companies must assess identifiability through a risk-based and context-specific analysis.

Pseudonymization vs Anonymization: Key Differences

Pseudonymization

Pseudonymization replaces or masks direct identifiers.

It usually allows the data to remain linkable to the same person or device across time.

It may still allow profiling, segmentation, measurement, personalization, or behavioral analysis.

It remains within the GDPR if the person is still reasonably identifiable.

Anonymization

Anonymization prevents identification in practice.

The data can no longer be attributed to an individual by reasonably likely means.

The risk of identification must be insignificant.

Re-identification must be practically impossible or require a disproportionate effort.

Properly anonymized data falls outside the GDPR.

Why the “No Re-Identification Key” Argument Is Not Enough

Many companies assume that data is anonymous because they do not hold a direct re-identification key.

The Criteo decision shows why this reasoning is incomplete.

A person can be identifiable without a formal re-identification key.

Identification may result from:

  • linkage with IP addresses;
  • linkage with device identifiers;
  • repeated online behavior;
  • location patterns;
  • browsing history;
  • transaction history;
  • partner identifiers;
  • advertising interaction data;
  • account activity;
  • or other datasets available within the same ecosystem.

The GDPR definition of personal data covers indirect identification.

That means the analysis must include both direct and indirect means of identification.

Why Intent Does Not Decide Whether Data Is Personal

Another important lesson is that the controller’s intention is not decisive.

A company cannot argue that data is anonymous merely because it has no interest in re-identifying individuals.

The Conseil d’État rejected that type of reasoning.

The question is not whether re-identification is commercially useful.

The question is whether re-identification is reasonably possible.

This distinction is essential for governance.

Companies must assess capability, not intention.

The “Disproportionate Effort” Test

The Conseil d’État refers to three practical criteria:

  • time;
  • cost;
  • manpower.

These criteria help determine whether re-identification is realistically possible.

Time

Would re-identification require only a short technical process, or would it require an extremely long and complex investigation?

Cost

Would re-identification be financially realistic, or would it require an economically irrational level of investment?

Manpower

Would re-identification require ordinary technical expertise, or a large and unrealistic deployment of human resources?

The more feasible re-identification is, the more likely the data remains personal data.

The more disproportionate the required effort is, the stronger the anonymization argument becomes.

How Companies Should Assess Re-Identification Risk

A robust anonymization assessment should examine several questions.

1. What Direct Identifiers Were Removed?

Examples include:

  • name;
  • email address;
  • phone number;
  • customer ID;
  • account number;
  • national identifier;
  • payment identifier;
  • employee number.

Removing these elements is important.

It is not sufficient.

2. What Indirect Identifiers Remain?

Examples include:

  • IP address;
  • device ID;
  • cookie ID;
  • advertising ID;
  • location data;
  • timestamps;
  • behavioral events;
  • transaction history;
  • browsing history;
  • product usage logs;
  • partner identifiers.

Indirect identifiers are often the main source of re-identification risk.

3. Can the Same Person Be Tracked Over Time?

If a dataset allows the same person, browser, account, or device to be followed over time, the risk increases.

Persistent identifiers are especially sensitive.

4. Can the Data Be Cross-Referenced?

Companies should assess whether the data can be combined with:

  • internal databases;
  • partner databases;
  • public datasets;
  • data broker information;
  • CRM systems;
  • analytics tools;
  • advertising platforms;
  • payment records;
  • or other operational systems.

Cross-referencing is often more important than direct identification.

5. What Is the Purpose of the Processing?

If the purpose of processing is profiling, personalization, targeting, measurement, or behavioral prediction, the dataset may be structurally designed to distinguish individuals or devices.

This increases the risk that the data remains personal.

6. Who Could Re-Identify the Person?

The analysis should not be limited to one employee or one system.

Companies should identify which actors may have access to additional data:

  • the controller;
  • joint controllers;
  • processors;
  • partners;
  • advertisers;
  • publishers;
  • analytics providers;
  • platform operators;
  • data brokers;
  • malicious actors;
  • or recipients of shared datasets.

7. Would Re-Identification Require Disproportionate Effort?

The company should document the time, cost, technical expertise, and manpower required to identify a person.

A vague statement is not enough.

The conclusion must be supported by technical and legal analysis.

Business Impact: Why Misclassification Creates Risk

Misclassifying pseudonymized data as anonymous can create major legal and operational consequences.

Legal Basis

If the data remains personal, the company must identify a valid legal basis under the GDPR.

Depending on the context, this may involve consent, legitimate interests, contract performance, legal obligation, vital interests, public interest, or another basis.

For advertising and tracking activities, consent may be required under applicable rules.

Transparency

If the data is personal, individuals must receive clear information about the processing.

This includes information about purposes, legal basis, recipients, retention, rights, and where applicable, joint controllership arrangements.

Data Subject Rights

If the data remains personal, individuals may have rights of access, erasure, restriction, objection, portability, and rectification.

A company cannot ignore those rights by labeling the dataset anonymous.

Consent Evidence

Where processing is based on consent, the controller must be able to demonstrate valid consent.

In complex advertising or partner ecosystems, this may require strong contractual, technical, and evidentiary mechanisms.

Joint Controllership

Where several actors determine purposes and means, joint controllership may arise.

This requires transparent allocation of responsibilities under Article 26 GDPR.

Data Protection Impact Assessment

High-risk processing may require a DPIA.

This is especially relevant for large-scale profiling, behavioral targeting, or systematic monitoring.

Security Measures

Pseudonymization may remain an important security measure.

However, it does not remove the need for appropriate technical and organizational safeguards.

Regulatory Sanctions

Incorrectly treating personal data as anonymous may expose companies to regulatory enforcement, financial penalties, corrective orders, and reputational harm.

Why This Matters for AI, Advertising, and Analytics Projects

The Conseil d’État’s reasoning is particularly important for AI, adtech, and analytics.

These projects often rely on datasets that appear non-identifying because names and emails have been removed.

However, AI and analytics environments often contain rich behavioral and technical signals.

Examples include:

  • user journeys;
  • browsing behavior;
  • device fingerprints;
  • purchase patterns;
  • timestamps;
  • location patterns;
  • app usage;
  • clickstream data;
  • search behavior;
  • recommendation histories;
  • conversion events.

These signals may allow singling out or re-identification even without a name.

For AI governance, the key lesson is clear:

A dataset used for training, optimization, targeting, or profiling should not be treated as anonymous merely because it has been pseudonymized.

Practical Governance Framework for Companies

Companies should create a formal anonymization assessment process.

Step 1: Map the Dataset

Identify all categories of data included in the dataset.

Include direct identifiers, indirect identifiers, technical data, behavioral data, transactional data, and metadata.

Step 2: Identify Remaining Linkability

Determine whether the dataset allows the same individual, device, account, or household to be followed over time.

Step 3: Assess Available Additional Information

Identify what additional information exists internally or externally and whether it could be used to re-identify individuals.

Step 4: Evaluate Re-Identification Scenarios

Build realistic re-identification scenarios.

Do not limit the analysis to the company’s intended use.

Step 5: Measure the Required Effort

Assess time, cost, manpower, technical expertise, and access requirements.

Step 6: Document the Conclusion

Record why the risk of identification is insignificant or why the data remains personal.

Step 7: Reassess Over Time

Anonymization is not always permanent.

Technological progress, new datasets, new partnerships, or new tools can change the re-identification risk.

Practical Checklist Before Calling Data Anonymous

Before treating pseudonymized data as anonymous, companies should be able to answer the following questions:

  • Have all direct identifiers been removed?
  • What indirect identifiers remain?
  • Can the same person or device be followed over time?
  • Can the data be linked with other internal databases?
  • Can the data be linked with partner data?
  • Can the data be linked with public or commercially available datasets?
  • Does the dataset include IP addresses, device IDs, cookie IDs, advertising IDs, location data, or behavioral history?
  • Is the purpose of the processing based on profiling, targeting, personalization, or analytics?
  • Could a technically skilled actor re-identify some individuals?
  • Would re-identification require disproportionate time, cost, and manpower?
  • Has the analysis been documented?
  • Has the legal team reviewed the conclusion?
  • Has the technical team validated the assessment?
  • Has the DPO approved the qualification?
  • Is the conclusion periodically reassessed?

If these questions have not been answered, the company should be cautious before treating the data as anonymous.

Common Mistakes Companies Should Avoid

Mistake 1: Confusing Masking With Anonymization

Replacing a name with a random identifier may reduce risk.

It does not automatically anonymize the dataset.

Mistake 2: Looking Only at Direct Identifiers

Indirect identifiers can be enough to make a person identifiable.

IP addresses, device identifiers, location data, and behavioral patterns can matter.

Mistake 3: Ignoring Cross-Referencing

A dataset may appear harmless in isolation but become identifying when combined with other data.

Mistake 4: Relying on Lack of Intent

A company’s lack of interest in re-identifying people is not decisive.

The legal issue is whether identification is reasonably possible.

Mistake 5: Treating Anonymization as a One-Time Label

Anonymization should be reassessed over time.

New tools, new datasets, or new data-sharing arrangements can change the risk.

Mistake 6: Failing to Document the Assessment

If challenged by a regulator, a company must be able to explain and support its conclusion.

A statement that “the data is anonymized” is not enough.

What CEOs and CFOs Should Understand

For executives, the distinction between pseudonymization and anonymization has direct business consequences.

If a company wrongly treats personal data as anonymous, it may build entire processes on a false legal assumption.

That can affect:

  • product design;
  • data monetization;
  • advertising revenue;
  • AI training datasets;
  • customer analytics;
  • vendor contracts;
  • investor due diligence;
  • M&A transactions;
  • regulatory exposure;
  • and valuation of data assets.

The question is not only whether the company has data.

The question is whether the company has legally usable data.

Poor data qualification can turn a valuable asset into a regulatory liability.

What Legal, DPO, Data, and IT Teams Should Do

The Criteo decision shows that anonymization must be assessed jointly.

Legal teams cannot do it alone.

Technical teams cannot do it alone.

DPOs cannot validate it without understanding the actual dataset and the processing context.

A strong governance process should involve:

  • legal;
  • DPO;
  • cybersecurity;
  • data science;
  • product;
  • marketing;
  • IT;
  • compliance;
  • and business owners.

The goal is not to slow innovation.

The goal is to ensure that data-driven projects are built on a legally accurate classification of the data.

Frequently Asked Questions

Is pseudonymized data personal data under the GDPR?

Yes, in most cases.

Pseudonymized data remains personal data if a person can still be identified directly or indirectly.

Can pseudonymized data ever become anonymous?

Yes, but only if the risk of identification is insignificant and re-identification would be practically impossible because it would require disproportionate time, cost, and manpower.

What did the French Conseil d’État decide in the Criteo case?

The Conseil d’État held that Criteo’s pseudonymized data remained personal data because users could still be identified indirectly through associated data such as IP address, geolocation, device identifiers, partner identifiers, browsing activity, purchases, and advertising interactions.

Does the absence of a re-identification key prove anonymization?

No.

The absence of a direct re-identification key is not sufficient if the data can still be linked, combined, or cross-referenced to identify individuals.

Does a company’s lack of interest in re-identification matter?

No.

The legal test focuses on whether identification is reasonably possible, not whether the company intends to re-identify individuals.

What does “disproportionate effort” mean?

It refers to the level of time, cost, and manpower required to identify a person.

If re-identification requires disproportionate effort, the anonymization argument is stronger.

If re-identification is reasonably feasible, the data likely remains personal.

Why is this important for AI projects?

AI projects often use large datasets containing behavioral, technical, or transactional signals.

Even when direct identifiers are removed, those signals may allow individuals to be singled out or re-identified.

Why is this important for advertising and analytics?

Advertising and analytics datasets often rely on identifiers, browsing behavior, location, device data, and interaction history.

Those elements can create re-identification risk even without names or email addresses.

What should a company document before treating data as anonymous?

A company should document the dataset, remaining identifiers, possible linkages, re-identification scenarios, required effort, technical safeguards, legal assessment, DPO review, and reassessment schedule.

Does anonymized data fall outside the GDPR?

Yes, properly anonymized data falls outside the GDPR.

However, the company must be able to support the conclusion that the data is truly anonymous in practice.

Practical Rule for Companies

The safest operational rule is this:

If a person can still be singled out, linked to a profile, or identified through reasonably available data, the dataset should not be treated as anonymous.

For business teams, this means that pseudonymization should be treated as a protective measure, not as a legal exit from the GDPR.

Strategic Takeaways

The Conseil d’État’s March 4, 2026 decision confirms a strict but pragmatic approach to anonymization.

The court does not require absolute impossibility in a purely theoretical sense.

It requires practical irreversibility.

The risk of identification must be insignificant.

Re-identification must require disproportionate time, cost, and manpower.

This approach is highly relevant for modern data ecosystems because identifiability often comes from the combination of signals rather than from a single direct identifier.

For companies, the message is clear:

Data qualification is not a label. It is a demonstrable legal and technical conclusion.

Conclusion

Pseudonymization remains a valuable GDPR tool.

It reduces risk.

It supports data minimization.

It can strengthen security.

It may help demonstrate accountability.

However, pseudonymization is not anonymization.

The French Conseil d’État’s March 4, 2026 Criteo decision makes this distinction unavoidable.

A dataset does not leave the GDPR simply because names, emails, or direct identifiers have been removed.

It leaves the GDPR only when individuals are no longer reasonably identifiable in practice.

For CEOs, CFOs, DPOs, legal teams, data teams, and AI leaders, the implication is decisive:

Before treating data as anonymous, prove that re-identification is not reasonably possible.

That proof is now part of responsible data governance.