The Justice Department filed a statement of interest in the sprawling OpenAI copyright case, declaring the administration’s official position that training a large language model on copyrighted text amounts to fair use. From the DOJ’s perspective, it is indeed a matter of national security that authors and publications receive nothing as reams and reams of their otherwise protected material gets fed into the maw of ChatGPT to build out the weighted text-generation engine to help a high school student finish their book report.
The brief is bad, though it stumbles toward the correct legal conclusion. Assuming OpenAI acquired the material legally — and that hasn’t always been the case with AI training — it should be fair use to train a model, with caveats for making sure the model isn’t spitting back the exact text on the back end like a copying machine. But, since we’re talking about this Department of Justice, this is more a case of even a corrupt clock being right twice a day.
Remember how the administration and OpenAI have reportedly discussed handing the federal government a 5 percent equity stake in the company? That’s roughly $42.6 billion against the company’s $852 billion valuation. Seems pretty significant in light of the Justice Department swooping into a potentially existential legal battle. The brief opens with “The Interest Of The United States” and it runs three pages. It declines to mention the prospects of ownership.
It does, however, get to national security real fast.
Rules of law that make it significantly more difficult to develop a robust AI industry in the United States therefore threaten national security and give a competitive advantage to foreign adversaries who are not so encumbered.
Every AI booster loves to take a hammer to the “in case of emergency, say it’s a matter of national security” glass. There’s no denying the role AI will play in cybersecurity, but the defense of our technology infrastructure does not turn on whether a model is producing a passable tight five. This case started with stand-up comics suing over OpenAI ingesting their sets. If software developers want to fight about training on copyrighted coding that’s one thing, and it could raise genuine issues as to how “transformative” the output could possibly be given the constraints of programming languages. But to invoke national security in a case where the New York Times is hopping mad about purloined restaurant reviews is a joke of a stretch.
The brief cites the White House cybersecurity order twice about frontier model benchmarking and CISA directives. The DOJ also commits the administration to protecting “American ingenuity and intellectual property from exploitation and theft.” It will say nothing about protecting artists from exploitation and theft for the rest of the brief.
For its claim that LLMs have delivered major research breakthroughs, the United States cites a blog post from OpenAI and a blog post from Anthropic. With all the resources of the federal government behind it, the DOJ could not be bothered to look up any independent verification of the technology’s accomplishments and went with a pair of press releases from the companies with the highest incentive to fudge on this point.
This is eye-rollingly frustrating because the government doesn’t need to stoop to these canards to make this point. Courts have spent years taking corporate-approved bats to the fair use piñata, but copyright law includes a fair use exception for the purpose of facilitating greater creative expression. Every dispute like this loves to foreground the creative artist, but these cases rarely reward the artist as much as the mega business holding that artist’s rights. Large corporate rights-holders prefer to lock away the intellectual property interests they manage to extract more and more wealth — a drive that has reached a zenith in modern digital licensing enshittification, where no matter how much a customer spends on a book or song, the rights-holder will claim it’s merely a license that can be revoked without refund at any moment. But the framers of America’s copyright laws recognized that encouraging ingenuity requires some freedom to copy.
To enjoy the protection of the fair use doctrine, the law considers: (1) the purpose and character of the use, including whether the use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and, (4) the effect of the use upon the potential market for or value of the copyrighted work. No factor is dispositive, but they all inform a decision on whether the copying is intended to undermine the value of copyrighted work. OpenAI is a commercial enterprise — non-profit status notwithstanding — and it’s certainly copying whole works. But it’s doing so to teach an algorithm to guess at the most likely next word in the English language.
As long as the model has guardrails against “please regurgitate one specific training text verbatim” — which is admittedly a question in this litigation — training should, in the abstract, be fair use.
The Copyright Office reached the opposite conclusion on fair use, which the brief handles in a footnote by observing that the administration if trying to fire the head of the Copyright Office. But Trump hasn’t succeeded on that point, having lost in the D.C. Circuit, and losing again at the Supreme Court in June when the justices declined to stay her reinstatement.
The crux of the DOJ statement is that the Copyright Office and everyone arrayed against the AI companies have misunderstood the risk of substantial similarity by “conflating training (which requires copying of entire works, but no public access) with outputs (which the public may access, but which will often if not always lack substantial similarity).” Which is frustratingly correct. If the plaintiffs’ can show OpenAI producing substantially similar outputs — outputs that directly copy substantial portions of copyrighted work in a way meant to undermine the market — then they should lose. But training by itself should be fair use.
The statement’s economic argument is sharper, and half right:
It is not in the public’s interest for the largest technology companies to have an oligopoly on LLM training due to licensing entry barriers that function primarily as large subsidies for old mainstream media companies.
It is not in the public’s interest for the largest technology companies to have an oligopoly… period.
The risk of paying licensing fees isn’t the issue. We’ve got an oligopoly because of compute and capital expenditures. Stripping content costs out of the model doesn’t open the door to mom-and-pop LLM labs, it removes a line item from the balance sheets of the existing oligopoly.
Reading should be fair use. Publishers don’t like this because they’ve spent the last several years trying to squeeze more and more out of controlling access to their content, but if a company acquires it legally, it should be able to absorb the information for the purpose of teaching a computer how words work. Again. provided the company acquired it legitimately in the first place.
Going after AI for training rests on the worldview that all art is on loan. No one can buy a book, they just rent a revocable right to read it. That approach is doing more to stifle the creative arts than any algorithmic training exercise.
And it’s already generating absurdities. There’s a social media uproar over AI companies destroying rare books in the training process. Except the only reason they’re doing this is that, under Bartz, the tech companies were basically told that gingerly copying a book and returning it to circulation could not be fair use, but that buying a print book, shearing off the spine, scanning it, and pulping the pages would be fair use because one copy replaced one copy. Keeping the book would have been two copies, after all! So Anthropic bought millions of used books and ran them through hydraulic cutters, out-of-print and hard-to-find titles included. Destroying rare books should not be the cost of protecting the intellectual property rights of rent-seekers.
What sucks about this DOJ statement is not that it’s wrong, but that it’s a cynical, selective intervention for corrupt purposes. The government should be seeking to beef up fair use generally and forging statutory responses to end burdensome digital licensing regimes. Instead, they parachute in to one case solely to offer a financial giveaway to a multibillion dollar company that they intend to take a financial stake in.
Again, corrupt clocks are right twice a day.
(Statement on the next page…)
Earlier: Two Judges, Same District, Opposite Conclusions: The Messy Reality Of AI Training Copyright Cases
Trump AI Regulation Order Hallucinates More Fake Law Than Any AI
Joe Patrice is a senior editor at Above the Law and co-host of Thinking Like A Lawyer. Feel free to email any tips, questions, or comments. Follow him on Twitter or Bluesky if you’re interested in law, politics, and a healthy dose of college sports news.