Anthropic Stole This From Me
By Charles Graeber
Mr. Graeber is the author of “The Good Nurse” and “The Breakthrough” and is one of the named plaintiffs in Bartz v. Anthropic.
We need to talk about a potential A.I. famine.
If we do nothing, starving A.I. chatbots may be left with nothing to eat but their own responses. Think of a photocopy of a photocopy of a photocopy (if you remember photocopies). Quality degrades; mistakes amplify and dominate.
The result is what experts describe as model collapse. Imagine a world where A.I. slop is even worse. It’s not a future we want.
I know this because I did my part to head off where we are now.
Kirk Wallace Johnston, Andrea Bartz and I brought a class-action lawsuit against Anthropic. Our complaint was piracy and copyright violation; Anthropic had stolen our books because it needed them to create a commercial product, a generative A.I. chatbot designed specifically to write like us. But they didn’t ask us or pay us. That seemed unfair, and illegal.
Piracy is a crime as old as gold — or in the modern digital context, Napster. And intellectual property theft is pretty much baked into the A.I. development story. Most, if not all, A.I. companies built their tech on piracy — though they may call it “fair use.”
Anthropic’s story offered a twist. While many A.I. companies scraped the internet (including The New York Times, which is why OpenAI, Perplexity and Microsoft are currently being sued by the company for copyright infringement) for material, Anthropic also scanned, digitized and later destroyed physical books through a secret and expensive legal workaround.
If this was Anthropic’s legal fig leaf, it worked. The judge found that training the A.I. on physical books that Anthropic purchased was fair use, but illegally downloading and storing millions of copyrighted books, which Anthropic also did, was copyright infringement.
At $1.5 billion, ours was the largest copyright settlement in American history. Nearly half a million works were represented in our class action. After fees and lawyers (who deserve full credit for this case even existing), that works out to about $3,000 per book, which the author has to split with the publisher.
For a writer like me, who spent about 14 years researching and writing two books, that money is hardly life-changing. But in legal circles, it was considered a big win, four times the most common per-work penalty for illegal downloads.
For Anthropic, stealing was a smart business expense.
In 2021, Anthropic was a late but ambitious entrant to the A.I. race. Its goal, according to the fair use order in our case, was to create an artificial intelligence capable of writing that “an editor would approve of,” and that consumers would pay for.
Such a chatbot couldn’t be trained on mere tweets and trolls. It would need to mimic quality human writing at an industrial scale.
The Anthropic founder and chief executive Dario Amodei wanted to avoid the “legal/practice/business slog” of buying and licensing the books the company needed. Stealing was faster and cheaper. The price of virtue was deemed too high. Instead, the company is expected to be valued at $2 trillion in its looming I.P.O., which would make it the largest in history.
But the problem of stolen material won’t go away.
For many authors, it’s potentially ruinous. And this is where an A.I. famine could have its roots.
We had hoped our case would help confirm the protections of copyright in the age of A.I. and the rights of artists and creators to control their own work. And we hoped a judge would rule that our life’s work, in whatever form, was not a free natural resource that could be taken against our will and used to compete with us, if not put us out of work.
That did not happen.
Another judge in the same California district court disagreed, pointedly callingthe analogy between for-profit mimic machines and ambitious schoolchildren “inapt” and casting doubt on the ruling.
There will be other cases, and other judges with other opinions. This question is far from settled. It will be kicking around the courts for a long time.
Claude was built and may still be training on our copyrighted books. I’m not aware of any major A.I. companies doing otherwise.
For several years, authors have had a sense that the harm caused by unlicensed training on books had a negative impact on their livelihoods. A working paper published this month looked at more than 14,000 books sold on Amazon between 2023 and March of this year, putting that harm into hard numbers.
The paper found that the book market has been increasingly flooded with A.I. slop that has financial implications for publishers and authors. The slop confuses consumers and may mess up best-seller lists. Increasingly, even “human-authored” books are contaminated with A.I.-produced material.
For authors, “market dilution” by A.I. slop will result in loss of sales and income. It’s bad and getting worse. The median income of a full-time author is about $20,000. There’s not much further to fall.
But sadly, it’s the A.I.s that may suffer the most. A.I.s can’t keep improving by training only on A.I. output. The model collapse scenario threatens to wipe out whatever gains chatbots like Claude enjoyed from being trained on books in the first place.
And fewer working authors mean fewer original, A.I.-free books. For a ravenous A.I chatbot, this is a potential death spiral. For both authors and Anthropic, it’s a future we should work together to avoid.
One day, hopefully copyright protections will be extended to prevent the unlicensed training of A.I.s.
Until then, we live in a legal gray area, in which giant cash-rich corporations gorge freely on the rights of individual creators. The big ones will continue to eat the little ones unless the little ones can band together.
It’s up to us to demand legislation extending copyright protections in the 21st century. Many members of Congress appear willing; property protections have appeal for both sides of the aisle.
Affirming those copyright protections to include A.I. might even force A.I. companies to hire human writers to create original works for their L.L.M.s to train on. Or perhaps they’ll start spending their treasure like modern Medicis, funding the quality work A.I.s need to stay sharp. At the very least, they might invest in publishing houses and writing programs, fending off model collapse while ushering in a Silicon Era of artistic enlightenment.
I worry for this generation of artists coming of age in yet another technological adolescence, on the brink of so many cultural and economic disruptions.
But I do not worry for the future of art itself. We will always need human mediators to translate the human experience and make the flesh word.
No comments:
Post a Comment