The AI cases are about the download, not the model
Three courts have now split the same question the same way — what you do with a file may be fair use, but how you got it, and whether you handed it on, is the liability. That makes transport the whole question, and it is the most flattering thing that has ever happened to Usenet architecture.
The largest documented piracy operation in the public record right now is not a scene group. It is a set of companies with a combined market capitalisation in the trillions, and the part they are actually paying for is not the product they shipped. It’s the download.
On 4 September the Authors Guild filed for summary judgment against OpenAI before Judge Sidney Stein in the Southern District of New York, arguing that ChatGPT was built on “concealed mass piracy”: that OpenAI torrented books from Library Genesis, relabelled the corpora “Libgen1” and “Libgen2” as the innocuous “Books1” and “Books2” in the GPT-3 paper, and then deleted the LibGen files in the summer of 2022 for legal reasons — the only training corpora it has ever deleted. Whatever the court makes of it, the filing is worth reading for a reason that has nothing to do with AI. Across three cases, judges have converged on a distinction that anyone who moves files should have opinions about.
The split: acquisition, not use
Judge Alsup drew the line first, in Bartz v. Anthropic in June 2025. Training a model on books was “exceedingly transformative” and therefore fair use. Downloading those books from a pirate library and keeping them in a permanent, general-purpose digital collection was not — that part was, in his words, inherently and irredeemably infringing, and no amount of transformation downstream cured it.
Then look at what that cost. Anthropic settled for $1.5 billion across a works list of 482,460 titles — roughly $3,100 a book — with final approval entered on 20 July 2026. Not one dollar of that was for the model. It was for the acquisition and the retention.
The folk rule most people carry around is the opposite: that downloading is a grey area and distribution is where the real exposure starts. The courts have been quietly inverting it. The download is the tort. What you built afterwards is a separate question you might well win.
What made these cases winnable: the swarm
Now the part that should interest a Usenet audience, because it is architectural rather than legal.
Kadrey v. Meta is the clearest illustration. Meta torrented roughly 80.6 TB from LibGen, including the entire SciMag section — about 81 million scientific articles. Its own engineers understood the problem in advance: internal messages weighed torrenting as “much faster” against the con that Meta would have to seed the files from its own cluster servers. They did it anyway. Plaintiffs later ran forensic analysis of the swarm and reconstructed what those servers had made available to other peers.
Meta actually won the fair-use ruling on training. It is still in court over the transport. In March 2026 the judge allowed plaintiffs to amend and add contributory-infringement claims built entirely on the seeding, and summary judgment on that distribution claim will not be heard until 25 February 2027. A claim that survives the loss of the main claim, generated purely by protocol behaviour.
That claim does not exist without BitTorrent. The swarm is a public register of who held what and when, maintained by the protocol itself, readable by anyone with a client and some patience — and every participant is a distributor by construction, whether or not they intended to be. Usenet’s client–server model produces nothing comparable. An NNTP session is a private transaction between one reader and one provider. There are no peers to enumerate, no “made available” to infer from the protocol, and downloading never makes you an uploader as a side effect of downloading. This is the same structural point behind DMCA vs NTD: the interesting variable is rarely the content, it’s the plumbing.
Don’t over-read it
The evidence doesn’t vanish. It relocates.
In a BitTorrent case the record is public and free — anyone motivated can build it from the outside. In a Usenet case the record is private and sits with your counterparties: connection logs and payment history at your provider, account and API-key activity at your indexer. That’s a smaller and better-defended surface, but it’s a real one, and it is reachable by subpoena or by a provider deciding it would rather cooperate.
It also explains a pattern people misread as technical immunity. BREIN and its counterparts have gone after uploaders, indexers and providers rather than downloaders for years. That is target selection driven by where the evidence and the leverage are — not proof that the downloader is unreachable. Usenet removes the free evidence. It does not remove the evidence.
The finding that should bother hoarders
Read Alsup’s reasoning once more. He did not condemn the transient copy made in the course of a transformative use. He condemned the permanent, general-purpose library retained afterwards.
A 200 TB archive assembled over years, kept because it might be useful later, is a permanent general-purpose library. The analogy is imperfect in the direction that matters — these were commercial actors shipping derived products at scale, and courts weigh that heavily — but the doctrinal shape is the one now being tested, and “I only ever downloaded it, I never shared it” is a defence aimed at the claim the courts have been deprioritising.
What follows for your setup
If liability tracks acquisition and record-keeping rather than sharing, the leverage is at the endpoints that keep records:
- Provider logging policy is the load-bearing choice, not the marketing copy around it. Read what “no logs” actually excludes, and how long connection metadata lives.
- The payment rail is the cleanest identifier most people hand over voluntarily. The no-KYC provider guide exists for exactly this reason.
- VPN placement matters less than either — the honest version of that argument is that it moves your IP off the provider’s log and does nothing about the account.
- Indexer accounts are the least-defended piece of the chain and the one people think least about; the privacy stack piece puts them in order.
None of that is new advice. What’s new is the reason: the theory of liability being validated in court right now is about how the file arrived and what you kept, and both of those live in the same records.
There is an asymmetry here worth naming plainly. Meta and OpenAI will pay in the hundreds of millions or the billions, book the number as a cost of doing business, and continue shipping the models. The operators of the libraries they downloaded from face criminal exposure. Whatever those cases establish, it isn’t that the download was harmless — it’s that scale plus counsel converts it into a line item.