When Training Data Becomes a Legal Claim
- Aişe Gül Akkoyun
- Jan 16
- 3 min read
Publishers move to join a proposed class action against Google over AI training — and why that procedural step matters more than it sounds.
Background
On 15 January 2026, publishers Hachette Book Group and Cengage Learning filed a motion to intervene in an existing, author-led proposed class action against Google in the Northern District of California, arising out of allegations that Google misused copyrighted books to train its Gemini AI models. The publishers' filing specifically sought class representation for the publisher side of the alleged harm, distinct from the authors already named in the pending suit. A hearing on the motion was subsequently set for 6 May 2026; Google opposed the bid, arguing it was untimely and duplicative of the existing action.
Why intervention, not a fresh suit — for now
Filing a motion to intervene in an existing proposed class action, rather than bringing an independent claim, is the lower-cost path procedurally: it lets a rights-holder attach itself to a case's existing class definition and discovery record rather than starting from zero. That is the calculation Hachette and Cengage appear to have made in January 2026. Whether it was the right one is a separate question — and one this case itself would go on to test, since intervention motions are frequently contested, narrowed, or withdrawn as a case develops, especially where limitation periods and class-definition boundaries are in play.
Where this sits in the wider AI-training litigation landscape
The filing did not arise in isolation. It followed a wave of AI-training copyright litigation across creative industries — including a widely reported case brought by Universal Music Group against Anthropic over the alleged unauthorised use of song lyrics to train the Claude model — and reflects a broader pattern: rights-holders across music and publishing are converging on a small number of recurring legal theories against a small number of frontier AI developers, chiefly unauthorised copying and the outer boundaries of fair use.
Why this belongs in a legal-data frame, not just a copyright one
Each of these cases produces a growing public record of exactly what a plaintiff alleges was used for training, without requiring access to the defendant's actual training pipeline. That is unusual: it means the litigation record itself is becoming one of the better public sources of information about what content categories large AI developers have trained on, assembled adversarially rather than through voluntary disclosure.
What would be worth measuring
How the legal theories argued across these AI-training suits (fair use, licensing, DMCA copyright-management-information claims) cluster or diverge by plaintiff industry
Whether motions to intervene in existing class actions are becoming the dominant procedural strategy for rights-holders, relative to filing independent suits
Whether settlement patterns in earlier AI-training litigation (where they exist) are shaping the relief sought in later-filed cases
Whether courts are converging on a common evidentiary standard for what counts as adequate proof that specific copyrighted works were used in training
The open question
If training-data litigation continues to consolidate around a handful of large proceedings rather than fragmenting, the resulting case law may end up shaping AI training practices for the industry as a whole well before any legislature acts. Whether that is a desirable way to set data-governance policy — through adversarial litigation between a few large rights-holders and a few large developers — is a separate question from whether it is, in practice, the mechanism actually doing the work.
Written as of mid-January 2026, at the point the intervention motion was filed and before the scheduled May hearing. Later developments in this litigation, including whether the intervention motion was granted, narrowed, or withdrawn, are not covered here.
Related on DLS
Reform as a Corpus — another example of treating a legal process's documentary record as data in its own right
Sources
ChatGPT Is Eating the World, "Cengage Learning, Inc., Hachette Book Group, Inc. ask to intervene in copyright suit v. Google" (15 January 2026) — https://chatgptiseatingtheworld.com/2026/01/15/cengage-learning-inc-hachette-book-group-inc-ask-to-intervene-in-copyright-suit-v-google/
NISO, "Cengage and Hachette File Motion to Join Class-Action Lawsuit Against Google" (February 2026) — https://www.niso.org/niso-io/2026/02/cengage-and-hachette-file-motion-join-class-action-lawsuit-against-google
Courthouse News Service, "Authors, illustrators push for copyright owner class in case against Google AI" — https://www.courthousenews.com/authors-illustrators-push-for-copyright-owner-class-in-case-against-google-ai/
Publishers Weekly, "Publishers Strike Back Against Google in Infringement Suit" — https://www.publishersweekly.com/pw/by-topic/digital/copyright/article/99650-publishers-strike-back-against-google-in-infringement-suit.html
A note on later developments: in July 2026, Cengage and Hachette in fact withdrew from this intervention effort and, together with Elsevier and author Scott Turow, filed an independent class action against Google over its Gemini models — reportedly because Google could have asserted a three-year statute of limitations against claims falling outside the original class definition. That reversal is a real and separate development, sitting outside the January 2026 window this post covers, but it is worth knowing if you are tracking the case going forward.



Comments