Anthropic's Music Torrent Problem Is Worse Than It Looks
Sony and Warner Chappell just sued Anthropic over torrented song files used to train Claude. The company already admitted to torrenting books, but this lawsuit shows the scope of its data sourcing has always been bigger than disclosed.
On Friday, Sony Music Publishing and Warner Chappell Music filed a copyright suit against Anthropic, alleging the company used BitTorrent to download songbooks and then fed them to Claude. The key claim isn't just that music was copied. It's that Anthropic has never acknowledged that those torrented files contained anything beyond books.
Here's the thing: Anthropic admitted to torrenting books months ago, during a separate class action over copyrighted prose. That admission was framed almost as a routine matter, a messy but ultimately explainable data collection step. The music lawsuit blows that framing apart.
Timeline of a Data Sourcing Problem
Let's walk through the sequence. In 2023, Anthropic was assembling training corpora for Claude. Somewhere in that process, the company used BitTorrent, a peer-to-peer protocol, to grab large collections of text files. Those files, according to the new complaint, included song lyrics from at least several thousand works owned by Sony and Warner Chappell.
Anthropic's legal team has already acknowledged using torrented datasets in earlier court filings. That admission came with a caveat: they said they didn't know what was inside every file. But the companies' lawyers aren't buying the ignorance defense, and honestly, from a compliance standpoint, they shouldn't.
So now we're at a moment where the public record shows a clear pattern. Anthropic used Torrents, it extracted data, and it never fully inventoried what it grabbed. The question isn't whether music ended up in Claude's training set. The question is what else is sitting in those files. That's the hole the plaintiffs are driving through.
Impact: A Precedent for Every AI Company
This case matters well beyond Anthropic. The recording industry has been watching AI training pipelines for years, waiting to see how courts treat mass copying of copyrighted material. Sony and Warner could have gone after any model maker, but they picked a company that already admitted to torrenting. Smart legal strategy, actually.
If the case moves forward, discovery could force Anthropic to open up its entire data acquisition log. That's a nightmare for a company that markets itself as a safety-first, ethically conscious AI lab. And it's a gift to copyright plaintiffs everywhere, because it shows that even well-funded AI labs cut corners.
But let's be clear about what isn't happening here. This isn't a case about AI generating music that sounds like copyrighted songs. It's about the training data itself, the raw material, and specifically the method of obtaining it. Torrenting isn't inherently illegal, but using it to grab songbooks you don't have rights to is a straightforward copyright problem. The music industry doesn't have to prove Claude outputs lyrics verbatim. They just have to prove the unlicensed copies entered the model's training pipeline.
Outlook: What to Watch Next
The immediate next step is procedural. Expect Anthropic to file a motion to dismiss in the next few weeks, likely arguing that training on copyrighted material constitutes fair use. That argument will face a hostile reception given the torrenting admissions, and the court will probably push for discovery before making any calls.
Beyond the courtroom, watch for other music publishers to file their own suits. Universal Music Group hasn't joined this one yet, but the complaint names two of the three majors. If UMG follows, Anthropic will be facing a coordinated legal front across the entire industry.
The other thing to watch is whether Anthropic settles. The company has deep pockets and a brand built on transparency. A settlement would let it avoid discovery, but it would also set a price for every other AI lab that used scraped lyrics. That's not a small consideration. The precedent here's important for anyone building foundation models in the next few years.
So here's where I land: Anthropic's lawyers are going to have a hard time explaining why torrented books were acceptable but torrented music wasn't, because they were in the same files. The real story isn't just the music. It's that the company's data sourcing practices were sloppy, and the copyright system is finally catching up to that fact.
What regulators are really signaling: AI companies can't just download the internet and hope nobody checks the file inventory. Those files have a way of ending up in court.