Microsoft Called AI Scraping Largest Theft in History While Using NYT Content, Filings Reveal

TL;DR
- Newly unsealed filings in the New York Times copyright lawsuit claim Microsoft executives privately warned that AI web scraping could amount to the largest theft of labor in human history and gut the publishing industry.
- The Times alleges Microsoft and OpenAI knowingly built training datasets from millions of paywalled Times articles, bypassing paywalls and technical protections to feed ChatGPT and Copilot.
- The revelations strike at the heart of the fair-use defense and could reshape damages, licensing talks, and the broader legal battle over generative AI and copyright.
Inside Microsoft, A Dire Warning About AI Scraping
According to newly unsealed court documents in the New York Times lawsuit against Microsoft and OpenAI, some of Microsoft's own people saw the copyright storm coming long before ChatGPT went mainstream.
The filings, made public this week in federal court in Manhattan, include internal emails, chat logs, and strategy memos in which Microsoft executives and researchers allegedly debated the ethics and legality of scraping the open — and not-so-open — web to train large language models.
In one of the most quoted passages, a Microsoft executive is said to have privately described unchecked AI scraping as the largest theft of labor in human history, warning colleagues that if AI systems could simply ingest and regurgitate the work of journalists, authors, and artists for free, there would be little incentive left to create original work.
Another internal discussion flagged the specific risk to publishers, with employees warning that AI search and chatbots could gut publishers by answering user queries directly instead of sending traffic back to the source sites that produced the information.
For the Times, those private anxieties are now public ammunition. Its lawyers argue they prove Microsoft and OpenAI knew exactly what they were doing — and what it would cost creators.
Paywalled Times Content Allegedly Used Anyway
At the core of the unsealed material is a detailed account of how, the Times alleges, its journalism ended up inside OpenAI's models.
The complaint has long claimed that millions of Times articles — including investigations, reviews, and breaking news published behind a hard paywall — were swept into massive web-scale training datasets used for GPT-3.5, GPT-4, and Microsoft's Copilot systems.
The new filings go further, laying out how that allegedly happened. Prosecutors of the press case point to datasets like WebText2, Common Crawl derivatives, and other curated corpora that OpenAI built from links shared on Reddit, scraped web pages, and automated crawlers.
The Times alleges those crawlers did not respect paywalls, subscription gates, or robots.txt signals meant to block automated copying. Instead, copies of paywalled articles pulled from syndicated copies, aggregators, browser workarounds, and text extracted before the paywall loaded were allegedly retained and used for training.
Internal OpenAI and Microsoft communications cited in the filings suggest engineers were aware that high-quality publisher content — especially from outlets like the Times, the Wall Street Journal, and major book publishers — dramatically improved model performance on factuality, writing style, and reasoning. One memo described publisher data as premium fuel that was essential to making chatbots sound authoritative.
Microsoft, which invested more than $13 billion into OpenAI and integrated its models deeply into Bing, Copilot, and Azure, is accused of not just funding that effort but actively participating in the copying, storage, and commercial deployment of models it knew were trained on unlicensed copyrighted work.
Why The Gutting Publishers Memo Matters
The line about gutting publishers is doing heavy lifting for the Times' legal strategy.
Copyright cases against AI companies will likely turn on fair use — whether training on copyrighted text and then generating competing summaries is transformative enough to be allowed without permission or payment.
By surfacing internal warnings that AI answers could replace the original articles and siphon away subscriptions, advertising, and affiliate revenue, the Times hopes to undercut that defense. Under U.S. copyright law, harm to the market for the original work is a key factor courts weigh.
The filings include examples where, the Times says, Bing Chat and ChatGPT produced near-verbatim excerpts of Times stories or highly detailed summaries that allowed users to skip visiting NYTimes.com altogether. In some tests cited by the paper, the bots reproduced the opening paragraphs of paywalled investigations word-for-word when prompted in certain ways.
Microsoft and OpenAI have previously argued that any such regurgitation is a rare bug, not a feature, and that training is protected fair use similar to how a human learns from reading. They have also pointed to opt-out tools and partnerships with other publishers as proof they are acting in good faith.
The Times' lawyers say the new documents tell a different story: one of deliberate scale first, permission later.
What Microsoft and OpenAI Are Saying Now
Neither company has conceded wrongdoing in response to the unsealed filings.
In prior public statements and court responses, OpenAI has maintained that training on publicly available internet material is fair use, that it respects publisher opt-outs, and that it offers takedown and licensing programs. Microsoft has similarly argued it is entitled to innovate, that it has built guardrails to prevent infringement, and that copyright law should not stand in the way of a transformative technology.
Both companies have also struck licensing deals with other major publishers, including the Associated Press, Axel Springer, News Corp, and Condé Nast — deals the Times points to as an admission that licenses are in fact needed.
Microsoft is expected to argue in upcoming briefs that cherry-picked internal Slack messages and emails reflect vigorous internal debate, not corporate policy, and that colorful phrases about theft of labor were personal opinions about competitors scraping Microsoft content, not admissions about its own conduct.
OpenAI is expected to argue that the datasets in question contained only a tiny fraction of Times content relative to trillions of training tokens, and that the Times failed to prove the models were built to memorize rather than learn general language patterns.
What It Means For The High-Stakes Copyright Battle
The Times v. Microsoft and OpenAI case, filed in December 2023, is widely viewed as the bellwether for dozens of related lawsuits from authors, newspapers, record labels, and visual artists.
Legal experts say the unsealed filings could have three major impacts.
First, willfulness. If a jury finds Microsoft and OpenAI knew they were using paywalled work and foresaw harm to publishers, statutory damages could soar to $150,000 per infringed work — a potentially existential number when multiplied across millions of articles.
Second, injunctions and product changes. The Times is seeking not just money but an order to destroy models trained on its work. While courts rarely order model deletion, detailed evidence of paywall bypassing could push Judge Sidney Stein to impose stricter filtering, attribution, or traffic-sharing requirements.
Third, leverage for licensing. The entire news industry is watching. If internal warnings make a fair-use dismissal less likely, AI companies may be forced to negotiate broader, more expensive licensing deals for training data and real-time grounding, similar to how music streaming services were forced to license catalogs.
The next key milestones are summary judgment briefing later this fall and a potential trial in 2027. With discovery now unsealed, expect more internal emails, dataset inventories, and prompt logs to become public in the coming weeks — and more uncomfortable quotes for AI executives to explain.
Get All The Latest Updates Delivered Straight To Your Inbox For Free!