The Guidelines 03/2026 on web scraping in the context of generative AI, adopted by the European Data Protection Board “EDPB” for public consultation on 7 July 2026, are notable not only for what they require but for what they acknowledge. The document is unusually candid about three limitations: an epistemic one (the controller may not always know what it has collected), a technical one (what a model has learned cannot, today, be easily unlearned), and an institutional one (some of the assessments woven into the GDPR analysis sit, at least in part, with other authorities and courts). These acknowledgements are welcome, and they distinguish the text from more declaratory guidance. The tension is that the requirements built on top of them are not always adjusted accordingly, and that gap, between what the EDPB admits and what it nonetheless requires, is where the most interesting questions of the consultation lie.
Briefly, the Guidelines cover scraping performed by private entities, whether carried out in-house, commissioned from a third party or effected through the acquisition of pre-scraped datasets. They work through the familiar sequence: allocation of controller and processor roles, the core principles of Article 5 GDPR (purpose limitation, transparency, minimization, accuracy), the choice of legal basis, with legitimate interest under Article 6(1)(f) GDPR treated as the realistic candidate and consent all but discarded, and the treatment of special categories of data incidentally swept up in the collection, for which the EDPB adapts the CJEU’s GC & Others framework. Little of this structure will surprise anyone who has followed the Board’s recent work on AI. What rewards attention is how each of these familiar steps is made to function once the three limitations above enter the analysis.
Continue Reading Regulating the Irreversible: The EDPB’S Web Scraping Guidelines and the Limits of GDPR Orthodoxy








