@jon-nyc
I've started taking Jane Friedman's newsletter, because I trust her reporting. Today's newsletter spoke to your question about whether other big AI companies trained on pirated work:
OpenAI’s piracy looks to be worse than Anthropic’s piracy
One of the key AI lawsuits I’ve been following is this one in the Southern District of New York—a different district from the one where the Anthropic case was settled, and the home of New York publishers and landmark copyright cases. It involves both the Authors Guild and the New York Times, since the court consolidated the cases. Over the last couple of weeks, a lot of information has come to light from unsealed documents. The New York Times is probably taking great pleasure in publishing articles about what’s in these documents, revealing things like:
A Microsoft director calling AI scraping the “largest theft of labor in human history”
OpenAI’s head of ChatGPT writing that AI posed an “existential threat” to publishers
OpenAI developing a hack to get around the New York Times paywall
The disclosures related to books are equally bad. OpenAI, like Anthropic, created an internal library of books from the pirate website LibGen and even shared those works with outsiders. One of OpenAI’s employees testified, “The better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon.” And an OpenAI researcher described his mission as creating a machine to supplant human authors."