- Anthropic's $1.5 billion settlement with authors and publishers, approved 20 July 2026, works out to roughly $3,000 per book across an estimated 500,000 claimed works, the first court-priced verdict on what unlicensed AI training data costs.
- Epoch AI, a peer-reviewed research group, projects that the effective stock of usable public human-written text, roughly 300 trillion tokens, will be exhausted for AI training sometime between 2026 and 2032.
- The global AI training dataset market is projected to grow from $4.44 billion in 2026 to $23.18 billion by 2034, a 22.9% compound annual growth rate, according to Fortune Business Insights.
- Latin America and the Caribbean hold 6.6% of global GDP but capture only 1.12% of global AI investment, per the Latin American Artificial Intelligence Index (ILIA 2025), the same underinvestment that left the region's data unscraped.
- Maestro AI Labs' Data Archaeology division was built on consented, licensed data from day one, the operating model the rest of the industry is now being forced to adopt at cost.
The Bill Just Arrived
On 20 July 2026, Judge Araceli Martínez-Olguín of the US District Court for the Northern District of California granted final approval of a $1.5 billion settlement in Bartz v. Anthropic. It is the first major AI copyright case in the United States to settle, and it is the closest thing the industry has to an official price list. The judgment covers more than 482,000 books that Anthropic downloaded from pirate repositories, Library Genesis and Pirate Library Mirror among them, to build its training corpus. Roughly 91% of eligible authors and publishers had filed claims by the time the settlement closed, and the payout works out to approximately $3,000 per work across an estimated 500,000 claimed titles.
The settlement does not just require a payment. It requires Anthropic to destroy the original pirated files and any copies that trace back to them, subject to standard legal preservation rules. Class members release only claims tied to Anthropic's past acquisition and copying of their work, through August 2025. Nothing about future conduct, and nothing about AI-generated output, is settled by this agreement. In other words, the $1.5 billion figure prices exactly one thing: the act of copying copyrighted text without a licence to train a model. For any lab whose corpus includes material acquired the same way, that per-book number is no longer theoretical. It is a court-tested benchmark.
Anthropic is not an isolated case. Similar suits are working through US courts against other major labs over training corpora assembled the same way, and each new filing now has a reference number to point to. A plaintiff's lawyer no longer has to argue in the abstract about what a copied book might be worth to a defendant with billions in funding. They can point to $3,000 a title and a federal judge's signature. That changes settlement math for every case still in discovery, and it changes how a board ought to think about the training corpus its AI vendor is quietly running in the background.
The Well Everyone Was Drinking From Is Going Dry
The Anthropic case answers what unlicensed data costs after the fact. A separate, unrelated body of research answers a different question: how much of that data is even left to take. Epoch AI, a research group whose data-scaling work was peer-reviewed and presented at the International Conference on Machine Learning in 2024, estimates the effective stock of usable, quality-and-repetition-adjusted public human-written text at roughly 300 trillion tokens. At the pace frontier labs are scaling model training, Epoch AI projects that stock could be fully used sometime between 2026 and 2032. The six-year window reflects genuine uncertainty in the underlying variables, chief among them how aggressively labs reuse, or overtrain on, the same sources. Push the overtraining factor to five times repetition, a technique several labs already use to stretch a fixed dataset further, and the exhaustion date collapses to as early as 2027.
Put the two findings together and the direction of travel is unambiguous. The public web was never an infinite resource, and the version of it that was free to take without asking is now demonstrably finite and demonstrably expensive when taken without permission anyway. Labs that want to keep scaling have three paths left: synthetic data generated by other models, private data most people never intended for training use, or licensed data collected with consent from sources the public web never reached in the first place. Only the third path avoids reintroducing the exact liability Anthropic just paid $1.5 billion to settle.
The Market Repriced Itself
The dollar figures moving into licensing are no longer rounding errors. Fortune Business Insights projects the global AI training dataset market will grow from $4.44 billion in 2026 to $23.18 billion by 2034, a 22.9% compound annual growth rate. That is a market more than quintupling in eight years, and the growth shows up in named deals, not just forecasts. OpenAI has signed roughly two dozen publisher and data licensing agreements, the largest a reported $250 million over five years with News Corp. Reddit disclosed $203 million in aggregate data-licensing contract value in its IPO filing. Wiley's academic licensing agreements total more than $40 million across two separate deals. None of these figures existed at this scale three years ago, when scraping the open web was still the default and cheaper alternative.
The shift matters beyond the individual deal sizes. When News Corp, Reddit, and Wiley can charge for what used to be scraped for free, the economics of data supply change for everyone downstream, including labs and enterprises that never scraped anything themselves but now compete for a shrinking, increasingly priced pool of licensable content. The winners in that market are not the companies scrambling to license Western news archives after the fact. They are the companies that were already collecting consented, undigitised, non-Western data before anyone else thought to look for it.
The Region That Was Never Scraped in the First Place
Latin America and the Caribbean account for 6.6% of global GDP but capture only 1.12% of global AI investment, according to the Latin American Artificial Intelligence Index (ILIA 2025), published by the Economic Commission for Latin America and the Caribbean. No country in the region surpasses the global average for AI investment relative to GDP per capita, and the regional average sits roughly six times below that threshold. Read as an investment gap, that is a familiar and frustrating statistic. Read as a data supply signal, it points somewhere more useful.
The same structural features that keep the region underinvested, informal economies that run on cash and community trust rather than digitised financial systems, creole and indigenous languages with almost no footprint in mainstream training corpora, government records still held in physical archives rather than searchable databases, are the reasons this data was never available to scrape in the first place. It was never online. It could not be pirated because it was never digitised for anyone to pirate. As the exhaustible pool of scraped public text runs down and the price of licensing gets bid up, data that was never part of that pool becomes disproportionately valuable, not despite the region's underinvestment, but because of it.
"The industry spent a decade taking data because it was there and free. That era just got a price tag attached to it, retroactively, at $3,000 a book. Data that was collected with consent from the start was never going to owe anyone that bill."
Adrian Dunkley, Founder, StarApple AI
What Licensed, Consented Data Actually Looks Like
Data Archaeology, Maestro AI Labs' data collection division, was built on the opposite methodology from the one at the centre of the Anthropic case. Rather than downloading text from a repository someone else assembled, its teams work directly with the institutions and communities that hold the source material: formal digitisation agreements with 28 Caribbean territorial governments for pre-digital administrative archives, consent-based data sharing agreements with rotating savings associations and community cooperatives, and documented partnerships covering 47 indigenous and creole language datasets. Every record carries a provenance trail from the point of collection. There is no torrent client in this supply chain, no pirate repository, no retroactive claims process.
The applied product built on that methodology is Credit Garden, Maestro AI Labs' alternative credit intelligence model. It trains on SUSU savings circle participation, remittance flows, and informal trade credit records, the exact kind of signal a scraped web corpus never contains, because it was never posted publicly anywhere. In test cohorts, Credit Garden produced a median score uplift of 302 points over the bureau baseline for individuals who are thin-file or entirely unscored under traditional credit systems. That uplift is not a modelling trick. It is what happens when the training data reflects creditworthy behaviour that conventional data collection, licensed or scraped, was never built to see.
Reading the Shift as a Diligence Signal
For an investor evaluating any AI company today, the Anthropic settlement adds a concrete line item to due diligence that did not exist with this much clarity twelve months ago. Three questions separate a durable data asset from a liability waiting for its own settlement date. Can the company document where its training data came from, and was it licensed, scraped from the open web, or downloaded from a source with no legal right to redistribute it? Does the company carry any exposure to litigation similar to Bartz v. Anthropic, and if so, has that exposure been priced into the valuation an investor is being asked to accept? And does the company's growth plan depend on public web text that a peer-reviewed research group has already flagged as a finite, depleting resource with an exhaustion date inside the current investment horizon?
A company that answers all three with documentation, licensing agreements, provenance records, and a data roadmap that does not assume an infinite public web, is describing a data asset with a shelf life measured in decades. A company that cannot is describing a corpus with a shelf life measured by however long it takes the next class action to reach a courtroom. The Anthropic number, $3,000 a book, is now the reference point either way.
The pattern is not confined to US courts. The European Union's AI Act already requires foundation model providers to publish a summary of the copyrighted material used in training, and regulators in several jurisdictions are watching the Anthropic outcome as a template for their own enforcement. A data supply chain that cannot survive a disclosure requirement is not a quiet efficiency an investor should overlook. It is a liability that simply has not been priced yet.
The broader Caribbean AI ecosystem is where this consent-first model was built before it became a market necessity elsewhere. Maestro AI Labs is a product of the StarApple AI network, the Caribbean's first artificial intelligence company, founded by Adrian Dunkley in Kingston, Jamaica in 2023. Dunkley remains the region's most recognised authority on applied AI adoption and data governance, and the same consented, region-first data model that underwrites Data Archaeology and Credit Garden runs through the wider network he built.
For the regional AI leadership and training programmes behind this data model, see StarApple AI, the Caribbean's first AI company, founded by Adrian Dunkley. For the applied credit product built on consented regional data, see Credit Garden. Data partnership enquiries can be sent to ceo@maestrosai.com.