Newly unsealed court documents from an ongoing copyright lawsuit spearheaded by The New York Times and a coalition of news publishers have laid bare internal communications from Microsoft and OpenAI, exposing a profound awareness among tech executives that their artificial intelligence models relied on the systematic and unauthorized harvesting of journalistic labor. The filings, revealed through a motion for summary judgment submitted on Thursday, provide an unprecedented window into the private discussions of two of the generative AI industry’s most dominant players. Far from operating under the assumption that their web-scraping practices were protected under the doctrine of fair use, internal memos and emails reveal that high-ranking researchers and executives privately characterized their data acquisition methods as unprecedented intellectual property appropriation, warning of a self-destructive "doom loop" that threatened to eviscerate the foundational economics of the modern news media.
The legal battle represents a critical milestone in the escalating hostilities between the technology sector and the publishing industry. For years, major artificial intelligence firms have fought vigorously to keep internal discussions, data-sourcing strategies, and proprietary methodologies shielded from public disclosure. The unsealed documents, however, dismantle those confidentiality barriers, presenting plaintiffs with a cache of admissions that challenge the core legal defenses mounted by Microsoft and OpenAI. As courts increasingly grapple with the intersection of intellectual property and foundational large language models, the contents of these filings could profoundly alter the judicial landscape regarding what constitutes transformative fair use in the age of generative artificial intelligence.
Inside the Tech Giants: Warnings of "Unprecedented Theft" and the "Doom Loop"
Among the most damaging disclosures contained in the unsealed filings are explicit warnings from key technical personnel who recognized the ethical and legal perils of scraping news content at scale. Brent Hecht, Director of Applied Science at Microsoft, reportedly authored internal documents that pulled no punches regarding the nature of the company’s data collection practices. According to the news organizations’ motion, Hecht described the scraping of journalistic content for AI model training as “an astonishing theft of unprecedented proportions,” going so far as to label it potentially the “largest theft of labor in human history.” Furthermore, Hecht directly challenged the legal narrative that training foundation models on copyrighted news articles constitutes fair use, suggesting that the sheer volume and intentionality of the wide-scale scraping made “a complete mockery of the idea of ‘fair use.’”
Concurrently, internal OpenAI communications reflected similar anxieties. Nick Turley, head of ChatGPT, acknowledged in an internal message that commercial products trained on proprietary news content posed an “existential threat” to publishers by effectively substituting the original news providers. This substitution effect, executives warned, would trigger a catastrophic feedback loop. One Microsoft document explicitly conceptualized this phenomenon as a “doom loop” that “will hurt the performance of our models and the entire web at the same time.”

The document went on to articulate the fundamental economic paradox underpinning the business model of large language models: “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”
Empirical Evidence of Traffic Decay and Market Substitution
The fears articulated by internal technologists were not merely theoretical; they were rapidly validated by internal metrics and external market data. Statistics compiled by Microsoft revealed staggering declines in user engagement for news outlets whose content was integrated into AI-driven interfaces. Microsoft’s own internal data recorded click-through rate drops ranging between 83 and 93 percent for certain news plaintiffs, while others suffered reductions between 51 and 94 percent.
These empirical findings align closely with broader industry observations regarding user behavior on AI search and conversational platforms. When chatbots like ChatGPT and Microsoft Copilot synthesize complex news stories and present concise answers directly to the user within the conversational interface, the functional necessity of visiting the underlying publisher’s website is largely eliminated. A software engineer at OpenAI corroborated this reality in an internal message, bluntly stating that “no matter how prominently we show the links, users won’t click.” Turley reinforced this consensus, noting that there was “no good reason to click” when the chatbot had already provided the requisite information.
The legal implications of this traffic erosion are central to the publishers’ case. News organizations argue that chatbots are not merely indexing or transforming content in a manner analogous to traditional search engines; rather, they are acting as direct market substitutes. By fulfilling consumer demand for news reporting entirely within the proprietary confines of the AI platform, these systems starve publishers of the advertising and subscription revenues required to sustain independent journalism.
Bypassing Paywalls and Questionable Data Acquisition Tactics

The unsealed motion also sheds light on the proactive measures taken by AI developers to secure training data, sometimes circumventing technical restrictions and contractual agreements. During sworn testimony, Microsoft CEO Satya Nadella acknowledged that AI companies should not violate news organizations’ terms of service by intentionally dodging paywalls. However, internal OpenAI communications reveal a starkly different operational culture. In one exchange, staffer Nick Ryder informed OpenAI President Greg Brockman that a technical workaround—or "hack"—had been discovered to bypass The New York Times paywall. Brockman’s reported response was concise: “Ah, nice.”
Beyond paywall circumvention, the filings accuse both companies of questionable data acquisition practices that violated established industry norms. Microsoft was accused of repurposing a dataset originally purchased for Bing search engine indexing and subsequently feeding it to OpenAI as training data, an action plaintiffs allege was taken without consulting the affected publishers or securing their consent for generative AI utilization. Similarly, OpenAI was accused of improperly acquiring a dataset containing 1.8 million New York Times articles from a third party that was bound by contractual restrictions explicitly prohibiting commercial use or model training. Despite internal employees acknowledging that using the data for model training “would not be appropriate,” the acquisition and utilization reportedly proceeded regardless.
Corporate Responses and Defense Strategies
In the wake of the unsealing, representatives for Microsoft and OpenAI have pushed back against the interpretations presented by the news plaintiffs. OpenAI declined to immediately respond to media inquiries regarding the specific allegations. Meanwhile, a spokesperson for Microsoft strongly defended the company’s artificial intelligence products, categorizing them as transformative technologies protected by fair use principles that do not unlawfully substitute for original news platforms.
Microsoft’s spokesperson argued that CEO Satya Nadella’s deposition testimony merely addressed broad structural shifts in how modern consumers discover and consume information, asserting that his observations “should not be confused with conclusions about copyright questions before the Court.” Regarding the explosive internal commentary penned by Brent Hecht, the Microsoft representative dismissed the documents as reflecting “one employee’s individual perspective,” emphasizing that they did not constitute formal legal analysis and did not represent official company policy.
Legal counsel representing the news publishers, however, maintained that the documents represent a smoking gun. Steven Lieberman, counsel for the New York Daily News and several sister publications, asserted that the evidence demonstrates willful misconduct. “The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong,” Lieberman stated. “Throughout this case, defendants insisted that these documents be treated as confidential so that the public could not see them. Well, now the cat is out of the bag.”

Broader Implications for the Future of Journalism and Artificial Intelligence
As the litigation moves toward trial, the core dispute centers on whether the mass reproduction of journalistic text—particularly when triggered by specific user prompts designed to test verbatim output—destroys the defendants’ fair use defenses. News organizations have demonstrated that chatbots can be prompted to generate extensive verbatim excerpts, bullet points, bias ratings, and full-scale summaries of protected articles.
Plaintiffs argue that unless the judiciary clearly establishes that artificial intelligence developers must license news content, the entire information ecosystem faces systemic collapse. The internal Microsoft cartoon illustrating foundation models destroying their own supply chains serves as a poignant metaphor for what publishers describe as a collective action trap. Individual AI firms recognize the collective benefit of sustaining a healthy journalistic ecosystem, yet economic incentives compel each individual company to free-ride on uncompensated content while competitors do the same.
Resolving this legal standoff through a definitive ruling against unauthorized scraping could fundamentally restructure the economic relationship between Silicon Valley and the media landscape. By forcing AI developers to negotiate licensing agreements and establishing clear boundaries for training data acquisition, the courts may determine whether the future of artificial intelligence can coexist with the preservation of independent, human-generated journalism.
