Back to Blog
    Regulatory Updates

    EDPB Issues GDPR Guidelines on AI Web Scraping and Anonymisation

    The EDPB has adopted new guidelines on data anonymisation, web scraping for generative AI, and blockchain. Learn what this GDPR guidance means for your privacy compliance.

    PolicyForge Legal Team
    July 25, 2026
    8 min read
    gdpr
    edpb
    artificial intelligence
    web scraping
    blockchain
    privacy policy
    Share:

    The European Data Protection Board (EDPB) has adopted new guidance on three of the hardest problems in modern data protection: when data is truly anonymous, whether scraping the public web to train generative AI is lawful, and how blockchain projects can coexist with the GDPR. If your business claims to anonymise data, uses or builds AI, or touches distributed-ledger technology, this is guidance you will be measured against.

    TL;DR — Key Takeaways
  1. The EDPB has clarified the definition of anonymous data, raising the bar for claiming your datasets fall outside the GDPR.
  2. Web scraping for generative AI is squarely in scope: publicly available data is still personal data, and scraping it requires a lawful basis.
  3. The final version of the EDPB's blockchain guidelines has been adopted, addressing the clash between immutable ledgers and GDPR rights like erasure.
  4. Practical impact: audit every place your policies say "anonymised," document your AI training data practices, and keep personal data off-chain where you can.
  5. Key FactsDetails
    What happenedEDPB adopted guidance on anonymisation and web scraping for generative AI; finalized blockchain guidelines
    Who issued itEuropean Data Protection Board
    RegulationGDPR (EU)
    AnnouncedJuly 8, 2026
    Impact levelHigh
    Most affectedAI developers and deployers, analytics and data companies, blockchain projects, anyone claiming "anonymised" data

    What the EDPB Announced

    According to the official EDPB announcement, the Board has issued guidance covering three connected areas. Each one closes a gray zone that businesses have been operating in for years.

    Anonymisation: A Higher Bar for "Not Personal Data"

    Under Recital 26 of the GDPR, truly anonymous data falls entirely outside data protection law — which is why "we only work with anonymised data" is one of the most common claims in privacy policies. It is also one of the most commonly wrong ones.

    The new guidance clarifies the definition of anonymous data and provides a stricter framework for judging whether a dataset is genuinely stripped of identifiable information. This continues a long regulatory trajectory: European regulators have consistently held that identifiability must be assessed against all means reasonably likely to be used to re-identify someone — including combining your dataset with other available data. If re-identification remains reasonably possible, the data is pseudonymised, not anonymised, and every GDPR obligation still applies.

    Two practical consequences follow. First, anonymisation is itself a processing operation — you need a lawful basis to anonymise personal data in the first place. Second, the burden is on you to demonstrate that your technique actually works, not merely to assert it.

    Web Scraping for Generative AI

    Generative AI models are trained on enormous datasets, much of it scraped from the public internet — and much of that includes personal data: names in news articles, photos, forum posts, professional profiles. The EDPB's guidance establishes standards for this practice and reinforces the principle regulators have repeated since the first wave of AI enforcement actions: the fact that data is publicly available does not exempt anyone from the GDPR.

    For organizations training or fine-tuning models on scraped data, that means identifying a lawful basis (in practice, usually legitimate interest with a documented balancing test), honoring data subject rights even for scraped data, and being transparent about the practice. For the much larger group of companies that merely use AI vendors, it means your supplier's data practices are now part of your own due diligence — and your users increasingly expect your privacy policy to say what AI touches their data.

    Blockchain Guidelines Finalized

    The EDPB also adopted the final version of its guidelines on processing personal data via blockchain. The core tension is structural: blockchains are designed to be immutable, while the GDPR grants individuals rights to erasure and rectification. The finalized guidelines set out the regulator's view on how these can be reconciled — and the practical pattern that has emerged from the consultation process is consistent: architect systems so personal data lives off-chain wherever possible, treat on-chain storage of personal data as a design decision requiring justification, and document how rights requests are handled for anything the ledger touches.

    Does This Affect You?

    Work through this list — if any apply, this guidance is relevant to your compliance posture:

  6. You describe data as "anonymised" or "aggregated" in your privacy policy, sales materials, or data processing agreements.
  7. You build, fine-tune, or train AI systems on data collected from the web or from your users.
  8. You use third-party AI tools (chatbots, analytics, recommendation engines) that process your customers' personal data.
  9. You sell, share, or license datasets derived from personal data.
  10. You operate or build on blockchain infrastructure where any personal data — including wallet-linked identifiers — is involved.
  11. You serve EU or UK users in any of the above scenarios, regardless of where your company is based.
  12. What to Update in Your Privacy Documentation

    Your privacy policy's anonymisation language

    Search your policy for the words "anonymous," "anonymised," and "aggregated." For each claim, ask: could a motivated party re-identify individuals by combining this data with other sources? If yes, the honest word is "pseudonymised" — and the section needs to describe the processing, lawful basis, and retention that apply to personal data. An AI-generated privacy policy built from your actual practices avoids the template trap of inheriting anonymisation claims you can't defend.

    Your AI disclosures

    If AI systems process personal data anywhere in your stack, your policy should say so: what data goes in, for what purpose, whether it is used for training, and what rights users have. This overlaps heavily with the transparency obligations arriving under the EU AI Act — our EU AI Act compliance guide covers how the two frameworks interact, and our privacy policy guide for AI companies shows what strong AI disclosures look like in practice.

    Your rights-handling process

    Scraped data, training datasets, and pseudonymised datasets are all still subject to access and erasure requests. If your data subject request process can't reach those datasets, it has a gap. A structured DSAR intake and tracking workflow makes the difference between a defensible process and an inbox.

    Action Checklist

  13. Read the source. Start with the EDPB's announcement and the underlying guidelines relevant to your business.
  14. Inventory anonymisation claims. List every document — public and contractual — that asserts data is anonymous, and pressure-test each against the clarified definition.
  15. Map your AI data flows. Document which systems (yours and your vendors') train on or process personal data, and the lawful basis for each.
  16. Fix the vocabulary. Replace "anonymised" with "pseudonymised" wherever re-identification is reasonably possible, and update the attached obligations.
  17. Review blockchain architecture. Confirm personal data is stored off-chain where feasible and rights-request handling is documented for the rest.
  18. Refresh your privacy policy so it reflects today's practices — then diarize a review whenever you add an AI tool or data product.
  19. Frequently Asked Questions

    Is pseudonymised data still personal data under the GDPR?

    Yes. Pseudonymisation (replacing identifiers with tokens or hashes) is a security measure, not an exit from the GDPR. Only data that cannot reasonably be re-identified — by anyone, using any means reasonably likely to be used — is anonymous and out of scope.

    We only use publicly available data. Do we still need a lawful basis?

    Yes. This is the central point the EDPB keeps repeating: public availability does not change data's status as personal data. Scraping, storing, and processing it are processing operations that need a lawful basis, transparency, and respect for data subject rights.

    We use AI vendors but don't train models ourselves. Are we affected?

    Yes, in two ways. As a controller you are responsible for what your processors do with your users' data — so vendor AI practices belong in your due diligence and your data processing agreements. And your own privacy policy should disclose the AI processing your users are actually subject to.

    Does this apply to companies outside the EU?

    If you offer goods or services to people in the EU or monitor their behavior, the GDPR applies to that processing regardless of where you're established — the same extraterritorial scope that has always applied.

    The Bottom Line

    This guidance doesn't create new law — it removes the ambiguity that let loose claims about anonymisation, scraping, and blockchain survive unexamined. The organizations that will feel it most are the ones whose privacy policies were written once and never revisited. PolicyForge exists for exactly this failure mode: it generates policies from your actual data practices, flags when regulatory changes affect your documents through its GDPR compliance tooling, and keeps AI disclosures current as your stack changes.

    This article is for general information and is not legal advice.
    PLT

    PolicyForge Legal Team

    Our expert legal team combines decades of compliance experience with cutting-edge AI technology to deliver accurate, up-to-date legal guidance.

    GDPR Compliance
    Data Protection
    Privacy Law
    Business Regulations

    Related Posts

    Regulatory Updates

    Canada OPC Issues PIPEDA Guidance for Financial Entities

    The Office of the Privacy Commissioner of Canada has published new guidance for financial reporting entities on submitting privacy codes of practice for regulatory review.

    7/25/20268 min read
    Regulatory Updates

    EDPB and AMLA to Develop Joint Guidelines on Information Sharing

    The EDPB and AMLA are developing Joint Guidelines to clarify how organizations can share information to combat financial crime while complying with the GDPR and the upcoming AML Regulation.

    7/25/20268 min read

    Ready to generate your legal policies?

    Create compliant privacy policies, terms of service, and more with AI assistance.