The Value of Books in the Era of AI

The statements that have surfaced in the OpenAI copyright litigation shows a disdain for books that is hard for a book lover to understand. In 2019, I was childishly drawing my about my love for books,

Illustration of a girl reading books.



and how I wanted to read them all.

Illustration of a girl who is obsessed with books, pointing to millions and saying she wants to read them all.

Meanwhile, a new love for books was growing in another part of the country. But it wasn’t the same kind of love. It was more like the love of a colonizer spotting fertile land. They didn’t plan on reading the books. The machines they were building couldn’t even read the books because reading and learning from a book is a human activity. (Whatever the computers are doing is not that.) In the world of AI, technologists were ogling over books as data sources.

Drawing of a document that says "new data sources," "train on all the books of the world."

Instead of opening doors to new worlds or teaching new ways to think, books were just awesome datasets.

A robot drawing, where the robot is discussing the "value" of books as data.


The craft of writing was just another set of patterns. Well, it wasn’t just any set of patterns. It was an especially good one. In books, the storytelling actually made sense!

Drawing of statement OpenAI employee made about the value OpenAI saw in books, as data.







Books were the golden treasure. Instead of buying the treasure (who does that?!), they grabbed it.

Drawing of quote about how OpenAI trained on pirated books. Drawing includes a pirate ship.








OpenAI downloaded tons of books from a pirated library of books called Libgen to train its early models, seemingly to improve them enough to win more funding.

Drawing of a guy at a computer talking about training on libgen.


Employees seemed to understand this was legally dubious.

Quote from Dario Amodei about libgen being sketchy.


They also seemed to grasp the impact this would have on creatives.

A long quote in word bubbles about the impact that AI could have on creatives.


But they did it anyway. They downloaded pirated, copyrighted books, used copyrighted work to train their AI models, and distributed the book data to Microsoft.

The authors and newspapers suing OpenAI see the downloading, training, and distribution as three different moments of copyright infringement. At some point in the next year, a the federal judge in this case will have to decide whether the law allows this kind of copying.


Thanks for being here!



I’m a lawyer (Michigan Law alum) and artist, and I

If you’d like me to run a workshop for your organization, translate your big idea, or break down the latest legal issues, I'd love to connect!


Sources

All of these quotes appearing in the illustrations (and written in text form below) come from the In Re Open Ai, Copyright Infringement Litigation.

In Re OpenAI., Copyright Infringement Litigation, Class Plaintiffs’ Statement of Undisputed Facts

In Re OpenAI., Copyright Infringement Litigation, Class Plaintiffs’ Brief for Summary Judgment

The docket is neverending, but it’s here.

  • “Ben Mann, one of the OpenAl employees who originally downloaded LibGen, agreed that " ‘books are invaluable for long-range context modeling research and coherent storytelling.’ "

  • “A separate brainstorming document prepared by Nicholas Ryder listed ‘[train on all the books in the world’ under the heading ‘New data sources.’ “

  • “Testifying as OpenAI's corporate designee, Nicholas Ryder stated that ‘[blooks are an example of a dataset that ... contain very long, uninterrupted, correlated sequences of words,’and that ‘in the early part of OpenAI, people were very preoccupied about this feature of books.’ “

  • “Jack Clark, an early OpenAl employee who later co-founded Anthropic, wrote: ‘[o]ur work on Al and Creativity is going to increasingly lead to us creating systems that substitute for the labor of the people that define the 'culture' of society ... The better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon... Our work in this area will make people unemployed...There will be a point where a bunch of artists express worry about what we're doing here and we'll likely ignore their concerns and release anyway...’ “

  • “Dario Amodei, OpenAI's then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ “

  • “We trained GPT-3 on pirated stuff! No sharing that!"

  • “OpenAI employee Nat McAleese then wrote: ‘Train on libgen, raise money, get new data, delete the old models, be clean forever more. Temporal regulatory arbitrage.’ “

Next
Next

Authors Sue Open AI And Their Filings Read Like a Book