
WHAT WAS YOUR AI TRAINED ON?
The dataset your model learned from might not have been yours to use. “The AI did it” isn't a defence.
The legal basis for training AI on copyrighted material is being decided in real time, and the early results cut both ways. In Bartz v. Anthropic, the court held that training on lawfully acquired books can be fair use, but that using pirated copies to build the training set is not. Anthropic settled for $1.5 billion in September 2025 and agreed to destroy the pirated dataset. In Thomson Reuters v. Ross Intelligence, a federal court rejected a fair use defence outright, on facts closer to commercial, non-transformative use competing directly with the plaintiff's product — the first ruling of its kind against an AI company. The New York Times' case against OpenAI and Microsoft tests yet another theory: that AI outputs compete directly with the content used to train the model. There is no single rule yet, only a fast-growing set of fact patterns separating lawful training from infringing training.
OUR APPROACH
For any client building or fine-tuning models, we document the legal basis for every training dataset before it goes into the model — owned, properly licensed, or genuinely public domain — rather than after a regulator or litigant asks. Where clients are on the other side of this, licensing their own content for AI training, we structure those deals to reflect the value and retain control, rather than the vague blanket terms early movers in this market have signed away.
TAKE
If you can't say where your training data came from, neither can your investor's counsel during diligence — and that's not a technical problem, it's a legal one.