Seattle Times and Newsday Sue OpenAI and Microsoft
Major Regional Newspapers Demand Destruction of Trained AI Models
On September 6, 2026, major regional news publishers The Seattle Times and Newsday filed a federal copyright infringement lawsuit against OpenAI and Microsoft. The publishers allege that both artificial intelligence pioneers unlawfully harvested decades of proprietary reporting without authorization to train frontier AI models and power consumer applications like ChatGPT and Microsoft Copilot. By reproducing verbatim excerpts directly in search queries and reducing direct web traffic, the AI systems jeopardize critical subscriber revenue models for local newsrooms. Seeking the destruction of all training datasets and derivative AI models, this legal action directly affects commercial AI developers, digital publishers, and enterprise platform integrators navigating copyright liability across the synthetic media landscape.
Key Details
The lawsuit represents a significant expansion of copyright litigation targeting generative AI vendors by regional news outlets. The Seattle Times and Newsday join a growing coalition of nearly 400 local and national publications—including The New York Times, Ziff Davis, Merriam-Webster, and Encyclopedia Britannica—that have initiated legal action against AI developers.
In their complaint, the plaintiffs assert two main categories of infringement: unauthorized ingestion of copyrighted investigative archives during model pre-training, and unauthorized output generation that outputs verbatim or near-verbatim passages in response to user prompts. Because Microsoft Copilot heavily relies on OpenAI's GPT foundation models, Microsoft was named as a co-defendant in the suit.
Most critically, the publishers are demanding punitive measures beyond standard financial damages. The suit explicitly asks the federal court to mandate the destruction of all copies of their journalistic works held by OpenAI and Microsoft, alongside the destruction of all training datasets and derivative AI model weights that incorporate their content. Neither OpenAI nor Microsoft issued an immediate public statement regarding the newly filed complaint.
What This Means
This latest legal challenge underscores an intensifying structural conflict between foundation model developers and the digital publishing ecosystem. As frontier AI models grow increasingly capable of summarizing complex news events, user reliance on original news websites diminishes. This shift erodes paywall conversion rates and digital advertising yield for publishers who bear the heavy capital costs of investigative reporting.
For AI companies, the demand for model destruction—frequently referred to as "algorithmic disgorgement"—presents a existential technical threat. Extracting specific copyrighted works from fully trained transformer models without complete retraining remains technically impractical with current unlearning methods. If courts enforce model destruction orders, AI vendors could face astronomical retraining costs running into hundreds of millions of dollars.
Technical Breakdown
The legal and technical mechanics of the lawsuit center on how foundation models process, store, and retrieve training data:
- Ingestion Without Licensing: Foundation models ingest billions of web tokens during self-supervised pre-training, incorporating proprietary news archives without prior licensing agreements or compensation mechanisms.
- Memorization and Verbatim Output: High-capacity LLMs frequently exhibit unintended memorization, outputting substantial passages of copyrighted text when queried on specific news events or investigative series.
- Algorithmic Disgorgement Demand: Plaintiffs seek court orders requiring the complete deletion of training corpora and the destruction of trained model weights that ingested their copyrighted property.
- Economic Displacement: AI-generated search summaries directly intercept reader traffic, replacing direct visits to source publications and disrupting paywall conversion funnels.
Industry Impact
The outcome of this lawsuit could fundamentally reshape the economics of AI data sourcing and model architecture. Should courts rule that training on publicly accessible web data constitutes willful copyright infringement rather than fair use, the standard paradigm of web-scale data scraping will effectively end. AI developers will be forced to transition entirely toward negotiated licensing agreements, structured synthetic data generation, or open-weight permissive datasets.
Furthermore, enterprise software providers integrating third-party models, such as Microsoft, face escalating third-party liability risks. Enterprise customers may demand stronger indemnity guarantees against downstream copyright claims, forcing model creators to implement strict provenance tracking and rigorous data filtering pipelines before releasing new frontier models to the public.
Looking Ahead
As copyright litigation against OpenAI and Microsoft accelerates, the AI industry stands at a decisive regulatory tipping point. Courts will soon be forced to decide whether training complex neural networks on copyrighted material qualifies as transformative fair use or unlawful exploitation of intellectual property.
In the coming months, expect AI developers to double down on preemptive publishing deals while aggressively researching machine unlearning techniques to mitigate the risk of forced model deletion. Developers, enterprise buyers, and content creators alike must closely monitor these proceedings, as the eventual rulings will dictate the legal and financial rules of engagement for synthetic intelligence.
Source: The Verge(opens in a new tab) Published on ShtefAI blog by Shtef ⚡


