The Value of Books in the Era of AI
The statements that have surfaced in the OpenAI copyright litigation shows a disdain for books that is hard for a book lover to understand. In 2019, I was childishly drawing my about my love for books,
and how I wanted to read them all.
Meanwhile, a new love for books was growing in another part of the country. But it wasn’t the same kind of love. It was more like the love of a colonizer spotting fertile land. They didn’t plan on reading the books. The machines they were building couldn’t even read the books because reading and learning from a book is a human activity. (Whatever the computers are doing is not that.) In the world of AI, technologists were ogling over books as data sources.
Instead of opening doors to new worlds or teaching new ways to think, books were just awesome datasets.
The craft of writing was just another set of patterns. Well, it wasn’t just any set of patterns. It was an especially good one. In books, the storytelling actually made sense!
Books were the golden treasure. Instead of buying the treasure (who does that?!), they grabbed it.
OpenAI downloaded tons of books from a pirated library of books called Libgen to train its early models, seemingly to improve them enough to win more funding.
Employees seemed to understand this was legally dubious.
They also seemed to grasp the impact this would have on creatives.
But they did it anyway. They downloaded pirated, copyrighted books, used copyrighted work to train their AI models, and distributed the book data to Microsoft.
The authors and newspapers suing OpenAI see the downloading, training, and distribution as three different moments of copyright infringement. At some point in the next year, a the federal judge in this case will have to decide whether the law allows this kind of copying.
Thanks for being here!
I’m a lawyer (Michigan Law alum) and artist, and I
simplify legal and technology concepts for lawyers, technologists, and academics with my human-made infographics.
run workshops for professionals that help build their creative brain muscles in a world of AI
provide analysis on the Supreme Court and rule of law using words and illustrations.
If you’d like me to run a workshop for your organization, translate your big idea, or break down the latest legal issues, I'd love to connect!
Sources
All of these quotes appearing in the illustrations (and written in text form below) come from the In Re Open Ai, Copyright Infringement Litigation.
In Re OpenAI., Copyright Infringement Litigation, Class Plaintiffs’ Statement of Undisputed Facts
In Re OpenAI., Copyright Infringement Litigation, Class Plaintiffs’ Brief for Summary Judgment
The docket is neverending, but it’s here.
“Ben Mann, one of the OpenAl employees who originally downloaded LibGen, agreed that " ‘books are invaluable for long-range context modeling research and coherent storytelling.’ "
“A separate brainstorming document prepared by Nicholas Ryder listed ‘[train on all the books in the world’ under the heading ‘New data sources.’ “
“Testifying as OpenAI's corporate designee, Nicholas Ryder stated that ‘[blooks are an example of a dataset that ... contain very long, uninterrupted, correlated sequences of words,’and that ‘in the early part of OpenAI, people were very preoccupied about this feature of books.’ “
“Jack Clark, an early OpenAl employee who later co-founded Anthropic, wrote: ‘[o]ur work on Al and Creativity is going to increasingly lead to us creating systems that substitute for the labor of the people that define the 'culture' of society ... The better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon... Our work in this area will make people unemployed...There will be a point where a bunch of artists express worry about what we're doing here and we'll likely ignore their concerns and release anyway...’ “
“Dario Amodei, OpenAI's then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ “
“We trained GPT-3 on pirated stuff! No sharing that!"
“OpenAI employee Nat McAleese then wrote: ‘Train on libgen, raise money, get new data, delete the old models, be clean forever more. Temporal regulatory arbitrage.’ “