EDPB Issues Draft Guidelines on Web Scraping for Generative AI Training
Summary
- Publicly accessible personal data is not automatically available for AI training use: In Guidelines 03/2026, published for public comment on July 7, 2026, the EDPB clarifies that the mere fact that personal data is publicly accessible online does not, on its own, allow it to be collected and reused for AI training under the GDPR.
- Web scraping for AI training will face closer GDPR scrutiny: The Guidelines emphasize that companies that rely on scraped data must carefully assess their legal basis, transparency obligations, data minimization measures, and treatment of special categories of personal data. Broad or untargeted scraping, and scraping carried out despite clear website restrictions, will be more difficult to justify.
- The Guidelines signal a significant shift in regulatory direction: Companies that use scraped or third-party internet data for AI training should begin reassessing their legal basis, collection criteria, privacy notices, technical safeguards, and role allocation across the AI supply chain.
On July 7, 2026, the European Data Protection Board (EDPB) published for public comment Guidelines 03/2026 on Web Scraping in the Context of Generative AI. The Guidelines set out the EDPB’s position on scraping publicly available personal data to train and fine-tune generative AI models, and make clear that such use is subject to the GDPR.
The Guidelines apply to organizations that scrape data themselves, instruct third parties to do so, or acquire datasets that have already been scraped. This may include the training of new generative AI models or the fine-tuning of existing ones.
The EDPB makes clear that the GDPR principles of transparency, data minimization, accuracy, and lawful processing apply in full to the use of scraped data for AI training. The fact that data is publicly available does not, by itself, provide a legal basis for using it in AI training.
Key Takeaways
- GDPR fully applies to AI training on public data. Extracting, cleaning, structuring, and storing personal data through web scraping for AI training are all processing operations under the GDPR, even where the data is publicly available.
- Legitimate interest as a legal basis. To rely on Article 6(1)(f) of the GDPR, a controller must identify a real and specific legitimate interest, show that scraping is genuinely necessary for AI training, and demonstrate that the interest is not outweighed by individuals’ rights and reasonable expectations.
- Data minimization throughout the AI training lifecycle. Data minimization is not simply a post-collection clean-up exercise. Controllers must define clear collection criteria before scraping begins, apply filters during collection, and anonymize, pseudonymize, or substitute synthetic data where appropriate before training begins.
- Transparency. Where individual notice is impossible or would involve a disproportionate effort, controllers must publish comprehensive public notices describing their scraping activities, the categories of data collected, the data sources, the legal basis relied upon, and the GDPR rights available to individuals.
- Special categories of personal data: heightened safeguards required. If scraping involves special category data, controllers need both a lawful basis under Article 6 of the GDPR and a separate condition under Article 9. They must also implement technical and organizational measures to prevent the collection of such data, and, if incidental collection occurs, delete it promptly. Safeguards should also be implemented throughout model development and deployment to prevent data disclosure.
Practical Implications
Increased regulatory scrutiny of AI training data. The Guidelines reflect a clear regulatory focus on AI training datasets. Organizations using publicly accessible personal data should identify a valid justification for that use beyond the fact that the data is available online.
Controller vs. processor dynamics. Where an AI developer instructs a third party to scrape data and determines the purposes and means of collection, the developer is considered the controller. If both parties jointly determine the purposes and means of processing, they will generally be considered joint controllers. This may affect how data processing agreements are structured across the supply chain.
Untargeted scraping is much harder to defend. Broad, untargeted web crawling without strict collection criteria now carries significant regulatory risk. The Guidelines make clear that data minimization must be addressed before collection begins and reinforced during and after collection, requiring organizations to apply precise technical targeting before data is extracted.
Recommendations for Businesses
- Map all training and fine-tuning datasets that include scraped or third-party data, and assess the risks associated with them, by reference to data sensitivity, anti-scraping signals, and data protection roles.
- Validate the legitimacy of data sources and ensure transparency in the collection process, including through the use of timestamps and accuracy checks.
- Conduct legitimate interest assessments for existing and planned scraping processes, with particular attention to necessity, reasonable expectations, minors’ data, and compliance with website restrictions.
- Update privacy notices to reflect AI training purposes, including the relevant data categories, sources, legal basis, available rights, and channels for objection.
In light of the EDPB’s Guidelines, organizations can no longer rely on the public availability of data as the sole justification for AI training. We recommend that our clients review their data collection mechanisms, audit their AI training datasets, update privacy notices, and re-evaluate vendor agreements to ensure close alignment with these new regulatory expectations.
Our Privacy, AI, and Cyber Department would be pleased to provide further guidance on compliance with these regulatory developments and related changes in the generative AI landscape.
***
Dr. Avishay Klein is a partner and head of the firm’s Privacy, Cyber and AI Department.
Adv. Masha Yudashkin is an associate in the firm’s Privacy, Cyber and AI Department.
Adv. Eviatar Rich is an associate in the firm’s Privacy, Cyber and AI Department.

