The Fair Use Fallacy
A defense you raise after you are sued is not a permission you hold in advance.
On July 31, 2026, a German court applied American copyright law to an American company’s training conduct, weighed the American fair use defense on its merits, and rejected it.
In other words, Germany is assisting American courts with their interpretation of copyright law since some of our courts seem to be struggling with a basic idea: Stealing at scale is still stealing.
The case was GEMA v. Suno (42 O 763/25). GEMA, the German collecting society, sued the AI music company Suno over six specific compositions. The 42nd Civil Chamber of the Munich Regional Court and presiding judge Elke Schwager found for GEMA. Suno was ordered to stop reproducing those works, to stop training on them, to disclose the connected revenue, and to pay damages still to be set. The judgment is not final, and Suno has said it is evaluating its options, including an appeal.
One caution governs every statement here about the chamber’s reasoning: the written judgment is not published. Yet, the results and remedies are safe to state flatly. The reasoning comes from the oral pronouncement and from consistent specialist reporting, which is still reporting, and many people are already giving their hot takes on the action.
For years now, a large part of the AI industry has operated on a studied misunderstanding: training on copyrighted material is fair use. The claim shows up in filings, in investor decks, and in the confident tone of people who have never argued it in front of a judge. It stopped working like a legal argument and started working like a foundational strategic move.
This shouldn’t be a foundation for your business plan because it’s a legal defense, not a money-saving tool. In Munich it traveled, got its hearing, and did not hold.
Fair Use and Unfair Use
Fair Use is designed as a tenet to allow for reviews and news reporting, truly educational work, and parody, along with the troublesome reference to ‘transforming’ the work. Public domain work is a different matter entirely: no permission is needed and no license is owed, so fair use never enters into it. Everything in this piece is about work someone owns.
Some AI companies are still trying to use fair use as a means of avoiding paying license fees to creative people for their work by assuming that one problematic ruling in a San Francisco courtroom and a few other muddled judgments amount to a blanket okay. The same assumption led Meta to argue it can use anyone’s work without compensation, and led another model company to buy and destroy millions of books rather than license a single one.
The big fallacy is the blanket assumption. It is the claim that fair use covers AI training as a category, in advance, regardless of how the material was obtained, which market the output enters, and which work was used. This is enough of a problem that one key point seems to need stating:
Fair use is not a business strategy.
It is an affirmative defense. You raise it after you have been sued. It gets decided on your specific facts, in a courtroom you did not choose, under the law of whichever country reaches your conduct.
Building a capital structure on it means building on an outcome you do not control.
None of this is an accusation of bad faith. Most of the people making the blanket claim believe it, and some were told it by lawyers answering a narrower question than the one that got repeated. Some of these legal experts are also basically counting on the fact that this will take so long that AI companies will be able to get ahead of it. I think that’s a fundamental reason why the German court was so clear: regardless of whatever lawfare people might launch and efforts to delay and obscure, fair use as a business strategy will never work in the EU.
Four places the argument goes wrong
1. The category error. Load-bearing, and the rest follows from it.
Fair use is a four-factor balancing test, applied to a particular use, of a particular work, in a particular market. A judge weighs purpose, the nature of the work, how much was taken, and market effect, together, on a record, after both sides have argued.
Nobody says “photocopying is fair use.” People say “this photocopying, of this book, for this classroom, was fair use,” after a court has decided. The sentence “AI training is fair use” performs a quiet conversion: a case-by-case defense raised after the fact becomes a permission granted before the fact, covering an entire category of activity by everyone in it.
No such permission exists. No statute grants it, no ruling confers it, no mechanism in the doctrine could produce it. A defense that has to be won cannot be held in advance.
2. Equivocation on “transformative.” In Campbell v. Acuff-Rose Music (1994), the Supreme Court gave the word its legal meaning: a use is transformative when it adds “new expression, meaning, or message.” In Andy Warhol Foundation v. Goldsmith (2023), the Court narrowed it, refocusing the first factor on the purpose of the specific use and on whether that use substitutes commercially for the original. Newness alone does not carry the factor.
Engineering uses the same word for something else. Training transforms text into weights. Nothing survives in its original shape. That is transformation in the ordinary technical sense, and it is real.
Then the two meanings get swapped. The engineering fact (the input was transformed) is offered as the legal conclusion (the use was transformative). Same word, two meanings, and the ambiguity does the argumentative work. Warhol is the Court saying the inquiry does not run that way: what matters is the purpose of this use and whether it substitutes for the original where the original earns.
3. Assuming the conclusion on factor four. The fourth factor asks what happens to the market for the original work if this use becomes widespread.
The blanket claim treats it as satisfied by assumption. But the output frequently enters the same market the input came from. Models trained on production music produce production music.
Two theories sit underneath that factor, and the courts have already split. In Thomson Reuters v. Ross Intelligence, Judge Stephanos Bibas held that the effect on a potential market for AI training data was itself enough to carry factor four against the developer. In Kadrey v. Meta, Judge Vince Chhabria rejected that same licensing theory as circular, reasoning that whether you hold a right to license for a use depends on whether that use is fair. What he called far more promising was market dilution: AI output flooding and devaluing the market for the originals. He then granted Meta partial summary judgment because the plaintiffs had not built a record on dilution.
It is a ruling handing the other side a map to exploit.
4. Conflating acquisition with use. How a copy was obtained is a separate question from what was done with it. Buying a book and reading it is not the same act as downloading a pirated library and reading it. The blanket formulation erases the distinction, because a category-level permission has no room for facts about sourcing. And sourcing is where the American cases have actually turned.
The case everyone cites
On June 23, 2025, Judge William Alsup of the Northern District of California held that training large language models on lawfully obtained books was fair use, as was format-shifting purchased print books into digital form. Building a library from pirated sources was not. Bartz v. Anthropic is the citation reached for most often, and it repays a closer reading than the coverage gave it.
A version of the criticism going around says Alsup did not understand the technology. That charge does not survive contact with the record. Alsup is a longtime hobbyist programmer who learned Java in connection with Oracle v. Google. He did the technical homework, and it shows.
He did something else, though. He reached for an analogy, and the analogy did the legal work. The training, he wrote, was like “any reader aspiring to be a writer.” A person reads many books, absorbs them, and writes something new.
That is an intuition about how learning works, and in that sentence it is doing the work a four-factor test is supposed to do. Transformativeness gets settled before the factors have been properly worked through.
The reason the analogy fails is not that a model is not a mind. That argument is unwinnable and, honestly, uninteresting. The reason it fails is narrower: a human reader does not produce a durable, transferable, commercially licensable artifact that can serve the same market as the works it read, at a scale that changes the market’s structure. A person who reads a thousand novels produces, at most, one novelist. A training run produces an asset that can be copied, sold, licensed, and deployed against the same readership indefinitely.
Factor four is where the analogy dies. And the analogy was doing the work that kept the opinion from having to arrive there under full weight.
One more detail, and it is the concrete one. In practice, that lawful format-shift meant buying physical books, cutting the bindings off, scanning the pages, and discarding what was left. The reasoning turned on the destruction: keeping the paper would have made two copies, and shredding it left one. The opinion’s clean path around the sourcing problem runs through the destruction of the book.
That is the ruling. Draw from it what you like.
On the settlement, and on the number. The pirated-library claims settled rather than going to trial. The settlement was large, and the figure has circulated widely.
I am not going to quote it, and the reason matters. A settlement is the price of getting caught, not the price of a license. Treating a liability number as a rate card is a mistake with consequences: the figure travels, lands in someone’s head as a benchmark, and shows up as a ceiling in a negotiation where it has no business being. What the settlement establishes is that unlicensed sourcing now carries a priced, enforceable liability. It says nothing about what licensed, consented, documented content is worth.
Valuing content for the AI training and citation market is part of my work at Credtent, and I am particularly concerned by how often people mix up a penalty payment for past infringement with a negotiated licensing payment, which should be ongoing, not a one-time payment for the past.
And the precedent is thinner than the coverage suggests. What exists is one district judge’s summary-judgment order, never reviewed by a circuit court, because the piece that would have been appealed settled instead. Claimants who opted out kept their claims, and three separate actions are live.
An unappealed trial-court order is being treated as settled law across an entire industry. That is the same fallacy again, running one level up.
One correction to the way all of this gets described, including in an earlier draft of my own: “circuit split” is imprecise. Ross, Bartz, and Kadrey are district-court decisions. No federal appellate court has ruled on fair use for generative AI training. The first to take up AI training at all, the Third Circuit, heard argument in the Ross appeal on June 11, 2026 and had not ruled as of August 3, 2026, though Ross was a non-generative search tool, which is a real distinction.
Which is not “the courts have spoken.” It is that the outcome depends on which facts you present, to which judge, on which record. That is what a case-by-case defense looks like from the inside.
What the models retain
One factual point deserves care, because it is easy to overstate and just as easy to wave away.
Models do not store their training corpus in any ordinary sense. Anyone claiming a model is a compressed archive of everything it read is overreaching, and a technically literate reader will discard the rest of the argument on the spot.
The defensible statement is narrower and holds up: models demonstrably retain and can be induced to reproduce substantial verbatim passages from their training data, disproportionately for text that appeared many times in the corpus. This is not seriously contested in the research literature. Extraction research going back to Carlini and colleagues in 2021 recovers training examples from deployed models, and later work distinguishes what can be pulled out with an arbitrary prompt from what emerges when you supply the original prefix. Memorization scales with model capacity, with duplication in the corpus, and with prompt context.
That is enough to defeat “we do not retain anything.” It is not enough to support “the corpus is in there.” The truth in between is where the legal argument actually lives, and it is the ground the Munich chamber stood on.
The defense that traveled
Germany has no fair use of its own. It has a closed list of enumerated statutory exceptions, plus the text-and-data-mining carve-out in Article 4 of the EU’s DSM Directive, which comes with a machine-readable opt-out rights holders can exercise. On the German side of the case, the chamber found the TDM exception did not apply, because the works were not merely analyzed during training but retained in the models in a reproducible form.
Then it went further than the German-law question required. Suno’s training happened in the United States. The court took jurisdiction over that conduct through Germany’s Collecting Societies Act, applied American copyright law to it, and held the copies were not covered by 17 U.S.C. 107. Reporting indicates the chamber worked the Warhol factors and found they ran against Suno. It distinguished Bartz and Kadrey directly: there, the training material was not reproduced to users in the outputs, whereas simple prompts to Suno produced outputs substantially similar to the originals.
So the American defense traveled. It got its hearing, on American law, in a German courtroom, and it lost, because the outputs reproduced the inputs.
This is a big deal. The industry has been running a global business on a jurisdiction-specific American affirmative defense as though it were a worldwide operating license. The defense does not stop existing at the border. It stops being reliable there, in front of a court that owes it nothing, working from a different exception regime and from the narrower reading of transformativeness the Supreme Court adopted in Warhol.
Whether a German court should be applying American copyright law to American conduct is a real question, and territoriality is likely to be central if this is appealed. But the shape is visible: a defense you have to win separately in every country you sell into is not a plan for those markets. It is a list of the places you have not been tested yet.
The objection that runs only one direction
Since the July 31 ruling, a criticism has been circulating that deserves a fair statement before an answer. It goes like this: Germany should not decide how other countries’ laws are enforced, and once courts start interpreting foreign law across borders, the slope runs somewhere bad for everyone.
Part of that is a real concern, and I have already conceded it. Whether a German chamber should apply American copyright law to American conduct is a legitimate question, territoriality is likely to be central on appeal, and Munich may not survive it.
But a second argument accompanies this question: Germany does not get to decide anything about how American law is implemented. That is not a jurisdictional concern. It is a supremacy claim, and the way to test it is to run it in reverse.
Suppose the Supreme Court eventually holds that fair use is a good defense for extracting the value of copyrighted work into a model. What follows is not hard to picture. Companies would carry that ruling into the EU and say: the way we built our models is legal, so you cannot stop us. And everyone on every side of this knows a US ruling would not bind the EU for a minute. The supremacy claim runs in one direction only. A doctrine expected to stop at the border when it loses and cross it when it wins is not a theory of jurisdiction. It is a preference.
The slippery slope is the other half, and it earns the name fallacy on its own terms: it asks you to fear a distant collapse instead of examining the step in front of you. The step in front of you was narrow. A chamber weighed the defense on its own terms, on six specific works, and found it did not cover these facts. If the reasoning is wrong, that is what the appeal is for.
Plus, the objection leaves out the context it sits in. The prevailing American posture in 2026 is to resist regulation that protects creatives, rights holders, and the people who supply the data. Against that posture, a foreign court saying in advance that fair use cannot be exported as an excuse is, I think, the right move, not overreach. My read of Munich, for whatever a read of an unpublished judgment is worth, is a chamber speeding the question along rather than building anything grander: saying clearly, before years of parallel litigation grind toward the same place, that this is the situation under any reasonable reading of the law it was handed.
It is the fallacy this piece is named for, translated into foreign policy: a defense treated as a permission, now with a flag on it. You cannot treat fair use as a business plan to save money, then carry it across borders as if it were a passport.
August 2, 2026
Two things landed on the same day, and both treat provenance as the compliance mechanism.
In the European Union, the AI Act’s Article 50 transparency obligations took effect on August 2, 2026: disclosure that content is AI-generated or manipulated, and machine-readable marking of synthetic content. On the same date, the Commission and the AI Office gained enforcement powers over general-purpose AI providers. Those providers’ underlying obligations, including the copyright policy and the public summary of training content, have applied since August 2, 2025. What changed is that a regulator can now act on them.
In California, the AI Transparency Act (SB 942, as amended by AB 853) became operative for covered providers on August 2, 2026, a date moved from January 1, 2026 to align with the European timeline (and dismissal of some court challenges). Those providers must offer detection tooling and attach both visible and latent disclosures to AI-generated content.
Neither settles the fair use question. What they do is quieter and, I believe, more durable. They make documentation of origin a legal obligation rather than a courtesy, which means the question stops being “can we argue our way out of this later” and becomes “can we show our work now.”
What I am certain about, and what I am not
It would be dishonest to name a fallacy and then commit a version of it in the other direction, so it is worth saying where my confidence sits and where it does not.
I am certain the categorical claim is wrong. “AI training is fair use” is not a statement the doctrine is capable of supporting, in the same way “photocopying is fair use” is not.
I am not certain how any particular case comes out. American courts will go on finding some training uses fair, and I think every such finding that blesses unconsented, uncompensated use of work someone owns is wrongly decided, for the reasons laid out above. Munich may read differently once the written judgment issues, and it may not survive appeal. The Third Circuit may disagree with things written here. New facts and better records will change outcomes, and they should. No one can predict exactly how this plays out, but I have faith that the people will support artists and creative people, including their right to own their work.
What to build instead
There is a version of this argument that ends in a lawsuit and a version that ends in a market. I would rather have the market that sets standards that people can trust.
Credtent exists to protect creative people without slowing down AI. We are a neutral third-party designed to bridge this gap, not one of the players trying to set up a marketplace that favors their model or content.
An AI industry that cannot get credible, high-quality content is not a good outcome, and neither is a creative economy strip-mined to feed it. The gap between those two is where I spend my time, and Credtent is closing it with infrastructure rather than arguments: documented provenance, certified origin, consent recorded at the source, and licensing that reaches the person who made the thing.
Creative Origin does that for a creative work, mapping onto the disclosure question Article 50 asks. Creative Consent does it for a platform’s treatment of the people who supply it. Commercial Safety does it for an AI company that wants to show where its training data came from, before someone requires it.
Credtent works with over a dozen enterprise content partners and hundreds of individual creatives, including best-selling authors and musicians. I did not build this because I predicted a ruling in Munich. I built it because documentation was always going to matter eventually.
Fair use will still be there when you need it. That is the point of a defense. Just do not confuse it with a business plan to save billions of dollars using other people’s content to build a trillion-dollar technology solution.
What do you think? Please share your thoughts below.






