EDPB Issues GDPR Guidelines on AI Web Scraping and Anonymisation
The EDPB has adopted new guidelines on data anonymisation, web scraping for generative AI, and blockchain. Learn what this GDPR guidance means for your privacy compliance.
The European Data Protection Board (EDPB) has adopted new guidance on three of the hardest problems in modern data protection: when data is truly anonymous, whether scraping the public web to train generative AI is lawful, and how blockchain projects can coexist with the GDPR. If your business claims to anonymise data, uses or builds AI, or touches distributed-ledger technology, this is guidance you will be measured against.
TL;DR — Key Takeaways| Key Facts | Details |
|---|---|
| What happened | EDPB adopted guidance on anonymisation and web scraping for generative AI; finalized blockchain guidelines |
| Who issued it | European Data Protection Board |
| Regulation | GDPR (EU) |
| Announced | July 8, 2026 |
| Impact level | High |
| Most affected | AI developers and deployers, analytics and data companies, blockchain projects, anyone claiming "anonymised" data |
What the EDPB Announced
According to the official EDPB announcement, the Board has issued guidance covering three connected areas. Each one closes a gray zone that businesses have been operating in for years.
Anonymisation: A Higher Bar for "Not Personal Data"
Under Recital 26 of the GDPR, truly anonymous data falls entirely outside data protection law — which is why "we only work with anonymised data" is one of the most common claims in privacy policies. It is also one of the most commonly wrong ones.
The new guidance clarifies the definition of anonymous data and provides a stricter framework for judging whether a dataset is genuinely stripped of identifiable information. This continues a long regulatory trajectory: European regulators have consistently held that identifiability must be assessed against all means reasonably likely to be used to re-identify someone — including combining your dataset with other available data. If re-identification remains reasonably possible, the data is pseudonymised, not anonymised, and every GDPR obligation still applies.
Two practical consequences follow. First, anonymisation is itself a processing operation — you need a lawful basis to anonymise personal data in the first place. Second, the burden is on you to demonstrate that your technique actually works, not merely to assert it.
Web Scraping for Generative AI
Generative AI models are trained on enormous datasets, much of it scraped from the public internet — and much of that includes personal data: names in news articles, photos, forum posts, professional profiles. The EDPB's guidance establishes standards for this practice and reinforces the principle regulators have repeated since the first wave of AI enforcement actions: the fact that data is publicly available does not exempt anyone from the GDPR.
For organizations training or fine-tuning models on scraped data, that means identifying a lawful basis (in practice, usually legitimate interest with a documented balancing test), honoring data subject rights even for scraped data, and being transparent about the practice. For the much larger group of companies that merely use AI vendors, it means your supplier's data practices are now part of your own due diligence — and your users increasingly expect your privacy policy to say what AI touches their data.
Blockchain Guidelines Finalized
The EDPB also adopted the final version of its guidelines on processing personal data via blockchain. The core tension is structural: blockchains are designed to be immutable, while the GDPR grants individuals rights to erasure and rectification. The finalized guidelines set out the regulator's view on how these can be reconciled — and the practical pattern that has emerged from the consultation process is consistent: architect systems so personal data lives off-chain wherever possible, treat on-chain storage of personal data as a design decision requiring justification, and document how rights requests are handled for anything the ledger touches.
Does This Affect You?
Work through this list — if any apply, this guidance is relevant to your compliance posture:
What to Update in Your Privacy Documentation
Your privacy policy's anonymisation language
Search your policy for the words "anonymous," "anonymised," and "aggregated." For each claim, ask: could a motivated party re-identify individuals by combining this data with other sources? If yes, the honest word is "pseudonymised" — and the section needs to describe the processing, lawful basis, and retention that apply to personal data. An AI-generated privacy policy built from your actual practices avoids the template trap of inheriting anonymisation claims you can't defend.
Your AI disclosures
If AI systems process personal data anywhere in your stack, your policy should say so: what data goes in, for what purpose, whether it is used for training, and what rights users have. This overlaps heavily with the transparency obligations arriving under the EU AI Act — our EU AI Act compliance guide covers how the two frameworks interact, and our privacy policy guide for AI companies shows what strong AI disclosures look like in practice.
Your rights-handling process
Scraped data, training datasets, and pseudonymised datasets are all still subject to access and erasure requests. If your data subject request process can't reach those datasets, it has a gap. A structured DSAR intake and tracking workflow makes the difference between a defensible process and an inbox.
Action Checklist
Frequently Asked Questions
Is pseudonymised data still personal data under the GDPR?
Yes. Pseudonymisation (replacing identifiers with tokens or hashes) is a security measure, not an exit from the GDPR. Only data that cannot reasonably be re-identified — by anyone, using any means reasonably likely to be used — is anonymous and out of scope.
We only use publicly available data. Do we still need a lawful basis?
Yes. This is the central point the EDPB keeps repeating: public availability does not change data's status as personal data. Scraping, storing, and processing it are processing operations that need a lawful basis, transparency, and respect for data subject rights.
We use AI vendors but don't train models ourselves. Are we affected?
Yes, in two ways. As a controller you are responsible for what your processors do with your users' data — so vendor AI practices belong in your due diligence and your data processing agreements. And your own privacy policy should disclose the AI processing your users are actually subject to.
Does this apply to companies outside the EU?
If you offer goods or services to people in the EU or monitor their behavior, the GDPR applies to that processing regardless of where you're established — the same extraterritorial scope that has always applied.
The Bottom Line
This guidance doesn't create new law — it removes the ambiguity that let loose claims about anonymisation, scraping, and blockchain survive unexamined. The organizations that will feel it most are the ones whose privacy policies were written once and never revisited. PolicyForge exists for exactly this failure mode: it generates policies from your actual data practices, flags when regulatory changes affect your documents through its GDPR compliance tooling, and keeps AI disclosures current as your stack changes.
This article is for general information and is not legal advice.Recommended for You
Related Posts
Canada OPC Issues PIPEDA Guidance for Financial Entities
The Office of the Privacy Commissioner of Canada has published new guidance for financial reporting entities on submitting privacy codes of practice for regulatory review.
EDPB and AMLA to Develop Joint Guidelines on Information Sharing
The EDPB and AMLA are developing Joint Guidelines to clarify how organizations can share information to combat financial crime while complying with the GDPR and the upcoming AML Regulation.
Ready to generate your legal policies?
Create compliant privacy policies, terms of service, and more with AI assistance.