Unsealed Filings Reveal Internal Concerns at Microsoft and OpenAI

Recently unsealed court documents have brought to light internal communications from Microsoft and OpenAI, indicating that executives within both companies privately acknowledged the controversial data practices underpinning their large language models (LLMs). These filings reveal that while both entities were engaged in scraping paywalled content, including from The New York Times, and building datasets from it, a Microsoft executive reportedly labeled this activity as “the largest theft of labor in human history” and “staggering in its scale.”

Implications for Publishers and the Internet

The internal discussions also showed awareness of the severe implications for publishers. Microsoft’s internal warnings highlighted that generative AI products created from such scraped content could lead to a “doom loop” that would ultimately gut publishers and potentially “kill the entire internet.” This underscores a significant internal understanding of the destructive potential and ethical quandaries associated with the rapid development and deployment of AI technologies.