A US federal judge allows Reddit's core DMCA claims against Perplexity AI and SerpApi to proceed, in a significant test of AI scraping liability under copyright law
A Manhattan federal judge has allowed Reddit to proceed with its core claims that Perplexity AI and search-data provider SerpApi violated the Digital Millennium Copyright Act (DMCA) by using an automated scraping operation to bypass Google's anti-bot protections and obtain Reddit content at scale. The ruling means that Reddit's claim, that the defendants improperly circumvented technological protection measures to access and reproduce its content, survived a challenge to have those claims dismissed at an early stage of the proceedings. The DMCA, a US federal statute, prohibits the circumvention of technological measures that control access to copyrighted works, and Reddit's case turns on whether its anti-bot infrastructure constitutes such a measure and whether scraping through it constitutes a violation. The case is being closely watched by AI developers, content platforms, and technology lawyers globally because it squarely addresses whether AI companies can lawfully harvest large-scale training and retrieval data from platforms that have technical barriers in place to prevent automated access. A ruling against Perplexity AI and SerpApi on the merits would significantly increase the legal risk profile of web-scraping as a data acquisition method for AI systems, with direct consequences for how AI firms structure their data licensing strategies. While the case proceeds in US federal court, its implications extend to UK and EU practitioners advising AI companies on data sourcing, particularly given parallel debates about the scope of text and data mining exceptions under UK and EU copyright law.
Why this matters
The survival of Reddit's DMCA claims is significant because it signals that US courts are willing to treat AI-driven circumvention of anti-bot protections as a potentially actionable copyright violation, rather than dismissing such claims at the pleadings stage. If Reddit ultimately succeeds on the merits, it would establish a precedent that AI retrieval and search companies must either license content from platforms or avoid scraping through technical barriers, reshaping the economics of AI data acquisition materially. For UK practitioners, the case provides a real-time reference point as UK copyright law debates the scope of the text and data mining exception for commercial AI training, a question that remains unresolved in UK legislative policy.
On the Ground
This case generates work across IP (intellectual property) litigation, technology transactions, and AI governance advisory. IP litigators would be advising AI companies and content platforms on their exposure under equivalent provisions in UK and EU copyright law, while technology transactions teams would be reviewing and renegotiating data licensing agreements to reduce scraping-related risk. A trainee would assist with disclosure review and categorisation of technical evidence relating to the scraping operation, prepare witness statement bundles, and research the scope of the DMCA's anti-circumvention provisions and their analogues in UK law for use in skeleton argument drafting.
Interview prep
Question you might get
“What are the legal implications of Reddit's DMCA claims surviving the motion to dismiss for AI companies that rely on web scraping to build their datasets or retrieval systems?”
Sign up free to see the full answer
A model answer you can lift into an interview — how to frame this story for a partner.
Sign up freeMy notes
saved