Between June and September 2023, book authors filed five class action lawsuits—four against OpenAI, Inc. and one against Meta Platforms, Inc.—claiming their books were copied without their permission to train artificial intelligence large language models (LLMs). The lawsuits target OpenAI’s ubiquitous product, ChatGPT, and Meta’s similar offering, LLaMA. ChatGPT and LLaMA are text-based generative AI products: upon a human user’s entry of a text prompt, they will simulate a human response based on information extracted from “training” text input into the language model. In the OpenAI cases, the plaintiffs claim ChatGPT’s ability to generate lengthy summaries and, in some instances, verbatim quotations of plaintiffs’ copyrighted texts as evidence of wholesale copying.

Like many cases before them featuring a clash of technology and copyright, these cases will likely include a showdown on fair use. OpenAI and Meta will likely seek to create a favourable comparison between their generative AI products and other technology products that have survived copyright claims under the doctrine of fair use. These cases include Perfect 10, Inc. v. Amazon.com, Inc., 508 F.3d 1146 (9th Cir. 2007), in which the Ninth Circuit Court of Appeals held the framing and hyperlinking of copyrighted images for use in a search engine to be a fair use, and Authors Guild v. Google, Inc., 804 F.3d 202 (2d Cir. 2015), in which the Second Circuit held Google Books’ searching function and display of “snippets” of copyrighted text also to be fair uses. However, the U.S. Supreme Court has since handed down the widely publicized decision in Andy Warhol Found. for the Visual Arts, Inc. v. Goldsmith, 598 U.S. 508 (2023), which found against the existence of fair use based on reasoning thought by some to reflect a reining in of federal jurisprudence on the bounds of “transformativeness.” Accordingly, the OpenAI and Meta lawsuits will be the most significant tests of the impact of the Supreme Court’s Goldsmith decision on the landscape of the law of fair use.

In its complaint against OpenAI, the Authors Guild reveals several other breadcrumbs of fair use arguments to come. In a nod to the third factor, the plaintiffs note OpenAI could have trained their LLMs on works in the public domain rather than copyrighted text. In a nod to the fourth factor, they allege OpenAI has previously admitted to licensing works for use as training text, thus suggesting OpenAI knows there is a viable licensing market for this purpose but in this instance has chosen to deprive the plaintiffs of licensing revenue by copying their works without a licence.

The judge in the case against Meta recently dismissed the majority of the plaintiffs’ claims and provided the plaintiffs with an opportunity to file an amended complaint. During oral argument on the motion to dismiss, the judge suggested the plaintiffs’ vicarious liability claim, which alleges “every output” from LLaMA constitutes vicarious copyright infringement, needed to be narrowed to survive dismissal. A different judge expressed a similar view in a case filed by visual artists against the creators of visual art generative AI products, dismissing the artists’ claim that every output image generated by DeviantArt’s Stable Diffusion product is a derivative work of a copyrighted image. These judges’ views demonstrate the challenge plaintiffs face in crafting sufficiently specific allegations about an emerging technology that is perhaps not yet fully understood by all.