US newspaper publishers sue OpenAI and Microsoft for allegedly scraping hundreds of thousands of articles to train ChatGPT and Copilot, as Anthropic separately accuses Alibaba of using 25,000 fraudulent accounts to distil its Claude model
A coalition of publishers collectively owning close to 400 local and regional newspapers across the United States has filed suit against OpenAI and Microsoft in the US District Court for the Southern District of New York, alleging the 'systematic and willful theft of hundreds of thousands of articles' scraped from the internet to train ChatGPT and Copilot. The lawsuit asserts that these products have generated 'hundreds of billions of dollars in market value' for the defendants without any payment to the publishers whose work made it possible. Separately, Anthropic has sent a letter to US Senate Banking Committee chair and ranking member — Senators Tim Scott and Elizabeth Warren — accusing Chinese technology firm Alibaba of using nearly 25,000 fraudulent Claude accounts between late April and early June to conduct tens of millions of exchanges with the chatbot, which were then used as raw training data for Alibaba's AI system. This process, known as adversarial distillation (training a new AI model by using its interactions with an existing model as data), allegedly violates Anthropic's terms of service. Anthropic has previously levelled similar accusations against , , and ; has separately accused DeepSeek of the same practice. The AI companies' primary legal defence against copyright claims is that scraping publicly available online content is permissible under doctrine in US copyright law — an argument that courts have not yet definitively resolved. The White House published a memo in April stating the Trump administration would partner with private companies to fight 'industrial-scale campaigns to distil US frontier AI systems', drawing a distinction between that practice and ordinary small-scale distillation. The converging wave of litigation — from publishers, between AI companies, and now escalating to Senate-level lobbying — signals that AI training data liability is moving from theoretical to acutely contested.