In the past few years, a wave of authors has taken AI companies to court, accusing them of quietly feeding copyrighted books into the training pipelines of large language models (LLMs). The core allegation is simple: you used our books without permission. The legal response from LLM developers has been equally straightforward—they’ve leaned heavily on the doctrine of fair use, the long-standing safety valve in U.S. copyright law that allows certain unlicensed uses of copyrighted works.

Fair use, codified in 17 U.S.C. § 107, permits “the fair use of a copyrighted work … for purposes such as criticism, comment, news reporting, teaching …, scholarship, or research.” Courts evaluate fair use through four familiar but flexible non-exclusive factors:

  1. The purpose and character of the use, including whether it is commercial or nonprofit;
  2. The nature of the copyrighted work;
  3. The amount and substantiality of the portion used; and
  4. The effect of the use on the potential market for the work.

In 2025, the Northern District of California issued the first two substantive opinions applying these factors to LLM training: Bartz v. Anthropic1 and Kadrey v. Meta Platforms. Both judges ultimately held that using copyrighted books to train LLMs qualifies as fair use. But beneath that shared outcome lies a striking divergence in reasoning—an early judicial split that will shape how courts, policymakers, and the AI industry think about training data for years to come.

The tension between the decisions begins with a deceptively simple question: What exactly is the “use” the court should be evaluating? From there, the judges parted ways on how much weight to give each statutory factor and how to conceptualize the relationship between copying, machine learning, and market harm.

The factual backdrop in both cases was similar. Each company had amassed large collections of books through a mix of authorized and unauthorized channels. Anthropic downloaded unlicensed copies from online libraries and later bought physical books to digitize. Meta purchased digital copies but also supplemented its dataset with unauthorized downloads.2 These shared facts only sharpened the contrast in how the two courts approached the same legal problem.

Factor 1 — Purpose and Character of the Use

The sharpest divide deriving from in-depth analyses of Bartz and Kadrey begins with a deceptively simple question: What exactly is the “use” the court should be evaluating?

Bartz: A Disaggregated, Step by Step Approach

In Bartz, Judge Alsup broke Anthropic’s conduct into three distinct uses, each analyzed independently:

  1. Creating “training copies” from subsets of its internal library
  2. Digitizing lawfully purchased print books
  3. Acquiring and storing unauthorized copies from online libraries

Judge Alsup, by isolating each step, prevented the transformative purpose of training from “laundering” or justifying improper copying upstream. In his view, the fact that training itself is transformative does not retroactively sanitize the acquisition or retention of unauthorized copies.

Kadrey: A Holistic, Purpose Driven Lens

Judge Chhabria took the opposite tack. In Kadrey, he treated Meta’s downloading of unauthorized books as part of the overall training process, which he deemed highly transformative. Although he acknowledged that downloading is “a different use” from the copying that occurs during training, he insisted that it must be evaluated “in light of its ultimate, highly transformative purpose: training Llama.”

From that vantage point, he concluded that because the end use was transformative, the intermediate downloading was transformative as well. He rejected the idea that each step required its own independent justification.

Common Ground—and a Critical Divergence

Both judges agreed that using copyrighted works to train LLMs is highly transformative. Judge Alsup even called Anthropic’s training “spectacularly transformative,” likening it to a student learning to write by reading widely—absorbing patterns, not reproducing text.

But the agreement ended there.

Judge Alsup drew a bright line at Anthropic’s downloading of “pirated library copies,” calling it “inherently, irredeemably infringing” regardless of whether the copies were later used for training. Judge Chhabria, by contrast, treated the use of shadow libraries as relevant but not dispositive. He noted that plaintiffs had shown no evidence that Meta’s downloads benefitted the unlawful sites (e.g., through ad revenue), and he declined to treat “bad faith” as outcome determinative.

Factor 2 — Nature of the Copyrighted Work

Here, the judges aligned. Both recognized that the works at issue—novels, memoirs, and other expressive texts—sit at the heart of copyright protection. As a result, Factor 2 favored the plaintiffs in both cases, though neither court treated this factor as especially weighty.

Factor 3 — Amount and Substantiality of the Portion Used

Both courts reached the same bottom line for training: copying entire works was permissible.

Training Requires Full Text Copying

Both judges accepted that full text ingestion is reasonably necessary for LLM training and that the amount of copyrighted material that could be output to the public was insubstantial. On this point, the decisions are strikingly aligned.

Where They Split: Unauthorized Acquisition

Judge Alsup again drew a line that Judge Chhabria declined to draw.

  • For Anthropic’s permanent retention of pirated copies, Judge Alsup found Factor 3 weighed against fair use, emphasizing that Anthropic “lacked any entitlement to hold copies of the books at all.”
  • Judge Chhabria, by contrast, did not analyze Meta’s acquisition separately. He treated full text copying as necessary for training and did not carve out a separate inquiry for how Meta obtained the books.

Factor 4 — Effect on the Market

Factor 4 proved to be the most conceptually challenging—and the most consequential—of the four factors. Both judges considered three theories of potential market harm:

  1. Direct substitution (LLM outputs replacing the books themselves)
  2. Harm to a licensing market for training data
  3. Indirect substitution / market dilution (AI generated works flooding the market and devaluing human authorship)

The judges agreed on the first two theories but diverged sharply on the third.

Direct Substitution: No Harm

Both judges rejected the idea that LLM outputs displace demand for the books, noting that the models do not reproduce or distribute plaintiffs’ works.

Licensing Market: Circular and Unsupported

Both judges also rejected the argument that training harms an emerging licensing market for training data, calling the theory circular. As Judge Alsup put it, this is not a market the Copyright Act “entitles Authors to exploit.”

Indirect Substitution: The Major Split

This is where the decisions part ways.

Bartz: No Cognizable Harm

Judge Alsup dismissed concerns about market dilution, comparing them to the idea that teaching students to write well unlawfully competes with authors. He concluded that such competition is not the kind of harm the Copyright Act recognizes.

Kadrey: A Serious and Novel Risk

Judge Chhabria took the opposite view. He criticized Bartz for “brushing aside” the risk of market flooding and called the analogy to schoolchildren “inapt.” He emphasized that LLMs can generate millions of works instantly, creating a risk of unprecedented market saturation. He warned that such flooding could reduce incentives for authors to create—precisely the harm copyright law seeks to prevent.

Yet despite this concern, he ultimately found that the plaintiffs had not produced evidence of actual or likely market harm. As a result, Meta prevailed on Factor 4.

Still, his analysis leaves a clear roadmap for future plaintiffs: Factor 4 could become the battleground for challenging AI training as fair use.

Combining the Factors

Bartz v. Anthropic

Judge Alsup held that all factors except Factor 2 favored fair use for training. He granted summary judgment to Anthropic on that issue.

But he denied summary judgment on the separate issue of downloading unauthorized copies, finding that all four factors weighed against fair use. That denial—and the resulting exposure to liability—ultimately drove the parties to settle.

Kadrey v. Meta Platforms

Judge Chhabria granted summary judgment to Meta on its fair use defense for both the downloading and the training, concluding that the copying was transformative and that plaintiffs had not shown market harm sufficient to reach a jury. The case continues, but the fair use ruling stands as the most expansive judicial endorsement to date of LLM training as fair use.

Summary of Analysis

IssueBartz v. Anthropic (Alsup, J.)Kadrey v. Meta (Chhabria, J.)
How the “use” is definedDisaggregated into separate acts: (1) training copies, (2) digitizing purchased books, (3) downloading unauthorized copies. Each step must independently satisfy fair use.Holistic view: downloading unauthorized copies is evaluated in light of the ultimate transformative purpose of training Llama. Intermediate steps need not be justified separately.
Transformative purposeTraining is “spectacularly transformative.” But transformative purpose does not justify acquiring or retaining pirated copies.Training is “highly transformative,” and this transformative purpose extends to the acquisition of books, including unauthorized downloads.
Treatment of unauthorized copiesSharp line: downloading “pirated library copies” is “inherently, irredeemably infringing,” regardless of later use.Relevant but not dispositive. No evidence that Meta’s downloads benefitted shadow libraries; “bad faith” alone does not defeat fair use.
Factor 2 — Nature of the workFavors plaintiffs (creative works), but given little weight.Same: favors plaintiffs but not significant in the overall analysis.
Factor 3 — Amount usedFull text copying is necessary for training and favors fair use. But retaining pirated copies weighs against fair use.Full text copying is necessary for training; does not separately analyze acquisition. Factor favors Meta.
Factor 4 — Market effectRejects all three theories of harm except for pirated copies. No cognizable harm from training; licensing market theory is circular; indirect substitution is speculative.Treats Factor 4 as most important. Rejects direct substitution and licensing market theories, but takes indirect substitution seriously. Ultimately finds insufficient evidence of harm.
Outcome on trainingFair use; summary judgment for Anthropic.Fair use; summary judgment granted.
Outcome on unauthorized copiesNot fair use; summary judgment denied; case settles.Fair use; summary judgment for Meta.
Analytical philosophyFormalist, step specific, skeptical of using transformative purpose to justify upstream copying.Purpose driven, flexible, emphasizes fair use’s adaptability to technological change.

Conclusion: Implications for Future Litigation

Taken together, Bartz and Kadrey mark the first major judicial attempt to map 20th century fair use doctrine onto 21st century AI training practices — and they reveal a fault line that future courts will have to confront.

Three themes stand out for consideration by Leason Ellis clients, existing and prospective:

  1. The definition of “use” will shape the entire fair use analysis.

    Whether courts adopt Bartz’s step by step approach or Kadrey’s holistic framing will determine how much scrutiny is applied to the acquisition and handling of training data. Plaintiffs will likely push for the Bartz model, which isolates and challenges each act of copying. Defendants will favor Kadrey’s broader lens, which treats training as a single transformative enterprise.

  2. Factor 4 is emerging as the battleground.

    Both judges rejected direct substitution and licensing market theories, but Kadrey opened the door to a more sophisticated argument: indirect substitution through market flooding. That theory is now the most promising path for plaintiffs — but only if they can produce empirical evidence of actual or likely market harm. Expect future cases to feature expert testimony, economic modeling, and data on AI generated content’s impact on book markets.

  3. The treatment of unauthorized sources remains unsettled.

    Judge Alsup’s categorical rejection of pirated copy retention stands in stark contrast to Judge Chhabria’s more forgiving approach. This split creates uncertainty for developers who rely on mixed source datasets. Appellate courts will likely need to clarify whether the legality of acquisition is a threshold issue or merely one factor among many.

  4. The Ninth Circuit, and likely the U.S. Supreme Court, will have to intercede.

    The doctrinal tension between these decisions is too significant to remain unresolved. As more cases are filed (and they will be), courts will be forced to decide:

    • Is LLM training more like reading, copying, or data analysis?
    • Can transformative purpose justify intermediate copying?
    • Does the Copyright Act recognize a market for training data licenses?
    • How should courts measure indirect market harm in an AI saturated ecosystem?

The answers to these questions will determine not only the legality of current AI training practices but also the shape of the creative economy for decades to come.

Leason Ellis professionals are closely monitoring these and other AI-related issues impacting the legal landscape.


1 Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007 (N.D. Cal. 2025) and Kadrey v. Meta Platforms, Inc., 788 F. Supp. 3d 1026 (N.D. Cal. 2025).

2 The plaintiff Authors’ books alleged to be infringed by Meta were all downloaded by Meta from online libraries of unauthorized copies of copyrighted books.