Newly unredacted material in The New York Times' copyright suit against OpenAI and Microsoft shows senior people at both companies privately describing their AI practices in terms that closely track the plaintiffs' claims, according to TechCrunch, which reviewed the filings. In a January 2023 internal memo quoted in the case, Brent Hecht, Microsoft's director of Applied Science, called the mass copying of news content "the largest theft of labor in human history."
TechCrunch notes that much of the new detail comes from the Times' legal brief rather than the underlying exhibits, which remain sealed. The quoted lines therefore appear without their original context. Ars Technica and The Wall Street Journal published reports on the unsealed documents around the same day, but the details below come from TechCrunch's account.
The filings contain internal Microsoft data showing that Copilot's answer engine cut click-through rates to the Times' domain by as much as 93 percent compared with traditional Bing search. In a January 2024 presentation, Hecht described that decline as a feedback spiral that would damage Microsoft's models and the wider web at once. A company document quoted in the brief observed how unusual it is for an end product to endanger the economic base of the suppliers its own content pipeline depends on.
Other passages address whether AI products harm the market for copyrighted work, one factor in a fair-use analysis. Nick Turley, OpenAI's head of ChatGPT, wrote internally that publishers face a survival-level threat from chatbot products, which he said largely substitute for their work and will do so more as the models improve. OpenAI president Greg Brockman described the models as excellent at news. Microsoft CEO Satya Nadella, deposed earlier this year, agreed that conversations with chatbots have substituted for visits to the underlying source. He testified that companies should license paywalled material used for grounding or training. Had he known OpenAI trained on paywalled content, he said, he would have invoked Microsoft's right to require OpenAI to retrain its models.
The brief also quantifies the copying for the first time. OpenAI's mid-training datasets contain more than 91,692 copies of works from the Times, the Daily News and the Center for Investigative Reporting, and one Common Crawl-derived dataset held over 2 million documents from nytimes.com alone. A dataset assembled under an effort called Project Mango contains at least 160,903 unique works from the news plaintiffs. Per the filing, OpenAI handed Microsoft the full GPT-3 training set, while Microsoft passed training data back through initiatives called Project Taxi and Project Mango.
On conduct, the plaintiffs allege that OpenAI employees devised a way around the Times' paywall without detection (when researcher Nick Ryder shared the workaround, Brockman replied approvingly), scraped content from the Bing index, built news-heavy datasets such as WebText and WebText2, and stripped copyright notices from training data so the models would not show them to users.
The disclosures arrive after several court decisions favoring AI developers' arguments that model training qualifies as fair use. The Trump administration also filed a brief this month defending OpenAI's unlicensed training. Courts consider harm to the market for the original work when evaluating fair use, and the internal statements address that issue. "The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong," Steven Lieberman, counsel for the New York Daily News, told TechCrunch. OpenAI and Microsoft did not respond to the publication's requests for comment.
The 93% click-through decline gives marketing and search teams a benchmark from inside an AI platform rather than from third-party analytics. Microsoft measured the effect of its answer engine on referral traffic to one publisher, so the figure should not be treated as a market-wide average. The filings may increase pressure from regulators and publishers over crawling, licensing and attribution, though that outcome remains uncertain. Changes to those rules would affect which sources answer engines can use and how they credit them. Brands that rely on citations from AI systems should expect the available source pool to change as licensing disputes continue.



