Open-Weight AI Models: Why Nvidia, Microsoft, and Meta Are Warning Against Restrictions
On July 24, 2026, a coalition of 25 major technology companies — including **Nvidia, Microsoft, Meta, Palantir, Google, IBM, and Hugging Fac...
Record-Breaking Settlement: At $1.5 billion, this is the largest copyright class action settlement in US history, with roughly **$3,100 per author** across more than **480,000 works**.
Where the Liability Lies: The court ruled that training AI models on copyrighted books was "exceedingly transformative" (fair use), but *downloading and storing pirated copies from shadow libraries* was infringement.
Massive Claim Rate: Over **90% of eligible works** (440,000+) had claims filed — far above the ~10% typical for class actions — showing strong author engagement.
Attorney Fees: Class counsel trimmed their fee request to ~12.5% of the fund ($122 million), below the typical ~30% in such cases.
Narrow Release: The settlement covers only how Anthropic *acquired* books in the past — it does not license future training, and authors retain the right to sue over what Claude *generates*.
Why This Matters: This case creates a clear market signal: AI companies must license or lawfully acquire training data. Sourcing from pirate sites now carries proven financial consequences, setting a benchmark for dozens of similar pending lawsuits against OpenAI, Meta, and Google.
The case originated when a group of authors sued Anthropic for using pirated copies of their books from Library Genesis and Pirate Library Mirror — two notorious "shadow libraries" — to train the Claude model line. U.S. District Judge William Alsup (now retired) issued a split ruling in June 2025 that has since shaped how the industry views its legal exposure.
Judge Alsup found that training Claude on books was transformative and likely fair use — comparing it to "any reader aspiring to be a writer" learning from existing works. However, downloading and storing over seven million pirated books in a central repository was not protected. Since US copyright law allows statutory damages of up to $150,000 per work for willful infringement, Anthropic faced theoretical liability in the hundreds of billions of dollars. Rather than test those numbers at trial (set for December 2025), the company settled.
Under the final terms approved by Judge Araceli Martínez-Olguín:
$1.5 billion plus interest: paid into a fund covering more than 480,000 books
Anthropic must destroy the original pirated files and any copies derived from them
Payments are structured in four installments through September 2027, with an initial $300 million already in escrow
Money will not reach authors until any appeals are resolved
The settlement's narrow scope is critical for the industry:
It grants no license to continue training on copyrighted books
Authors keep the right to sue over model outputs (generated content)
Any book not on the settlement list can still be litigated
A handful of authors and publishers have already opted out to pursue individual claims
Class counsel Justin Nelson called this "the first of its kind in the AI era." While a settlement sets no binding legal precedent, it establishes a market price for unlicensed training data. Defending similar cases to trial can cost millions, and the statutory damage exposure remains catastrophic. This creates strong incentives for AI developers to:
License or lawfully purchase training materials
Scrub data provenance to remove infringing content
Negotiate publisher agreements rather than rely on web-scraped or pirated sources
This is a US case, but the warning about American legal risk travels globally. Any lab training on English-language books — regardless of where it's based — must now consider US copyright exposure. The ruling reinforces that while fair use defenses may protect the *act of training*, the *source* of training data remains fully subject to infringement claims.
Compiled by Yanuki using the latest trends and data: The distinction between transformative training and infringing data sourcing is likely to be the defining legal framework for AI copyright disputes going forward. It rewards companies that invest in legitimate data acquisition and penalizes shortcuts — a dynamic that will reshape the multi-billion-dollar market for training data.
Why did Anthropic settle if the court said training was fair use?
While training was deemed transformative, the court found that downloading pirated books from shadow libraries was direct infringement. With statutory damages of up to $150,000 per work across millions of books, Anthropic faced potential liability exceeding hundreds of billions of dollars — making settlement the pragmatic choice.
How much will individual authors receive?
The settlement provides roughly $3,100 per author from a fund covering more than 480,000 works. However, authors will not receive payments until after any appeals are resolved, and attorney fees (~$122 million) and litigation costs are deducted first.
Does this set a legal precedent for other AI copyright cases?
Legally, no — a settlement binds only the parties involved. However, it establishes a market benchmark for what training data is worth and signals to other AI companies (OpenAI, Meta, Google) the financial risks of unlicensed data sourcing. It may encourage settlements in the dozens of similar pending cases.
Can authors still sue over what AI models generate?
Yes. The settlement covers only how Anthropic *acquired* the books in the past. Authors retain full rights to sue over any content Claude *generates* that infringes their copyrights, and over any books not included in the settlement list.
What must Anthropic do beyond paying money?
Anthropic must destroy the original files it downloaded from Library Genesis and Pirate Library Mirror, along with any copies made from them. The company admitted no wrongdoing as part of the settlement.
Your copyright has leverage.: The 90%+ claim rate shows that collective action works. If your work is used in AI training without permission, you may have recourse.
Register your works.: Some objections noted that non-US-registered works were excluded — ensure your copyrights are properly registered to maximize protection.
Clean your data.: The key takeaway is that sourcing training data matters. Even if training is fair use, where you get the data is not. Conduct thorough provenance audits.
License early.: The cost of licensing data upfront is far lower than the risk of class-action exposure. Negotiate publisher agreements now.
Why this matters to you:: Every time you use an AI tool, it was trained on data. This settlement ensures that creators are compensated — which ultimately leads to better, more ethically sourced AI products.
Watch the next wave:: The unsettled question of whether scraping the open web is fair use will define the next chapter of AI copyright law.
This settlement marks a historic moment in the relationship between AI development and copyright law. The core question — is it fair use to train AI on copyrighted content? — remains unresolved. What is now clear is that *how* companies source that data carries serious financial consequences.
What do you think? Will this settlement push AI companies toward licensed data, or will they continue to test the boundaries of fair use? Do the $3,100-per-author payouts adequately compensate creators, or is the amount too low given the scale of Anthropic's valuation?
Share this with others who need to stay ahead of this trend! Let us know your thoughts in the comments below.
<div style="margin-top: 20px;">
</div>
---
Sources:
On July 24, 2026, a coalition of 25 major technology companies — including **Nvidia, Microsoft, Meta, Palantir, Google, IBM, and Hugging Fac...
On Saturday, July 25, 2026, OpenAI experienced a significant global outage that left thousands of users unable to access ChatGPT, the compan...
A 55-year-old Florida pastor, Scott Winters, has filed a lawsuit against OpenAI and its founder Sam Altman, alleging that medical advice fro...
Artificial intelligence has shifted from a software race to an infrastructure arms race, and no company illustrates this better than Google....
⚠ Disclaimer: Yanuki provides article summaries and links for reference only. Yanuki does not endorse, verify, or guarantee the accuracy of third-party sources. Please review original sources and verify information independently. Managed by the Yanuki Data Engine. Full Disclaimer