Adobe Faces Proposed Class-Action Over Alleged Use of Authors’ Work in AI Training
Adobe, like many major tech companies, has aggressively expanded into artificial intelligence in recent years. Since 2023, the company has rolled out a variety of AI services, including Firefly, its AI-driven media-generation suite. However, this deep dive into AI may now be creating legal challenges. A newly filed lawsuit alleges that Adobe used pirated books to train one of its AI models.
Lawsuit Filed by Oregon Author
The proposed class-action, brought on behalf of Elizabeth Lyon, an Oregon-based author, claims that Adobe incorporated pirated versions of numerous books — including her own — into training its SlimLM program. Adobe markets SlimLM as a “small language model series that can be optimized for document assistance tasks on mobile devices.”
According to Adobe, SlimLM was pre-trained on SlimPajama-627B, described as a “deduplicated, multi-corpora, open-source dataset” released by Cerebras in June 2023. Lyon, who has published several non-fiction guidebooks, alleges that some of her works were included in the pretraining dataset used by Adobe.
Details of the Alleged Copyright Infringement
As reported by Reuters, the lawsuit claims:
“The SlimPajama dataset was created by copying and manipulating the RedPajama dataset (including copying Books3). Thus, because it is a derivative copy of the RedPajama dataset, SlimPajama contains the Books3 dataset, including the copyrighted works of Plaintiff and the Class members.”
Books3 is a vast collection of 191,000 books frequently used to train generative AI systems and has been a recurring source of legal disputes. Similarly, RedPajama has been cited in multiple lawsuits within the tech industry.
For instance, in September 2025, a lawsuit against Apple accused the company of using copyrighted material to train its Apple Intelligence model, citing RedPajama as the source and alleging the works were used “without consent and without credit or compensation.” In October, a separate lawsuit against Salesforce claimed the company had also relied on RedPajama to train its AI models.
Industry-Wide Legal Implications
These cases highlight a growing challenge for the tech sector. AI models require vast datasets for training, and some of these datasets have allegedly included pirated content. In a landmark case in September, Anthropic agreed to pay $1.5 billion to authors who accused the company of using pirated works to train its chatbot, Claude. This settlement was viewed as a potential turning point in ongoing copyright disputes surrounding AI training data — a legal battleground that is only intensifying.
Conclusion
As AI continues to reshape the software and creative industries, tech companies are increasingly under scrutiny for how they acquire and use content for training models. Adobe’s legal challenges are part of a broader industry debate over copyright, consent, and the ethical use of authors’ works in AI development.






0 Comments