- Epoch AI's revised estimate: the stock of useful public human-written text for AI training will be fully used up somewhere between 2026 and 2032, with 80% confidence. Its original 2022 estimate pointed to 2024; the front end of the window has since moved to roughly 2028.
- Databricks closed a $5 billion round on 13 August 2026 at a $190 billion valuation, six months after being valued at $134 billion, with revenue run-rate past $7 billion on more than 80% year-on-year growth. The money is going into data preparation and governance tooling, not model training.
- A 2024 Nature paper by Shumailov and colleagues found that models trained repeatedly on other models' output degrade in two stages, called early and late collapse, and need periodic injections of fresh human data to stay anchored to reality.
- The 2025 DATAVERSITY Trends in Data Management survey found 61% of data professionals name data quality their single biggest challenge, and Gartner has projected that through 2026, 60% of AI projects will be abandoned for lack of AI-ready data.
- Maestro AI Labs' Data Archaeology programme has spent years collecting exactly the category of data the current research says is now scarce: ground-level financial, climate, and behavioural records from more than 36 economies that were never posted online and never entered anyone else's training set.
The Text Ran Out Roughly on Schedule
Epoch AI first warned in 2022 that frontier language models could exhaust the useful stock of public, human-written text as early as 2024. That estimate was too aggressive. The organisation has since revised its methodology and now places an 80% confidence interval on full exhaustion somewhere between 2026 and 2028 at the early end, and 2032 at the late end. The revision does not change the conclusion, only the timing. High-quality public text is a bounded resource, and the largest AI labs in the world have spent the past three years training on an increasing share of what exists.
This was never really a text-availability story. It is a story about what happens once a model has already read most of the internet worth reading. A lab facing that ceiling has three choices: pay for data it does not already hold, generate synthetic data to fill the gap, or find data that was never online in the first place. The first option gets expensive fast. The second option, as the next section covers, carries a documented failure mode. The third option is where the actual scarcity now sits, and it is not evenly distributed. Most of the internet's indexed text was written in a small number of languages, from a small number of countries, describing a small number of economies in detail. Everywhere else was scraped thin long before the labs ran into a ceiling.
Investors Just Priced That Scarcity at $190 Billion
Databricks closed a $5 billion funding round on 13 August 2026 at a $190 billion valuation, led by Coatue with participation from Blackstone, MGX, T. Rowe Price, Sixth Street Growth, BOND, Clearlake Capital, Point72, Premji Invest, and TPG. That values the company at nearly 50% above where it stood six months earlier, when a prior round set its valuation at $134 billion. Revenue run-rate crossed $7 billion in the same period, growing more than 80% year on year.
What makes the round worth reading closely is what the capital is earmarked for: Lakebase, a serverless database built for AI workloads, Genie, a business-data assistant, and Unity AI Gateway, which governs how an organisation's AI systems access and control multiple models. None of that is model training. It is the layer that decides what data a model is allowed to see, in what condition, and under what governance rules. Investors did not put $190 billion behind a company that builds bigger models. They put it behind the company that helps everyone else's models run on data that is actually fit to use. That is a re-rating of where the scarce input sits, and it happened in public, priced by people with a direct financial stake in getting the answer right.
The Synthetic Data Shortcut Has a Collapse Problem
The obvious response to a shrinking supply of real text is to generate more with AI itself. Gartner projected as far back as 2021 that 60% of the data used in AI and analytics projects would be synthetic by 2024, and that prediction landed roughly on schedule. The trouble is what happens when a model is trained on the output of earlier models across several generations. Shumailov and colleagues published the answer in Nature in 2024: recursive training on model-generated data produces two distinct failure modes. Early collapse sees distributional errors accumulate, pulling the model's outputs away from the true underlying distribution. Late collapse is worse: rare, low-frequency patterns present in the original human data disappear from the model's outputs entirely, and do not come back. The paper's experiments, run across several model architectures, found that periodic injections of fresh, human-originated data were necessary to prevent the drift.
Read those two findings together and the shape of the problem gets clearer. Real human text is running down. The industry's own fallback, synthetic data, degrades a model further the more generations of it get folded back in without a real anchor. A model trained mostly on other models' output does not fail loudly. It fails quietly, at the edges, on the rare cases and minority populations that were thin in the training data to begin with. Those are usually the same populations the internet never described in much detail in the first place.
Data Quality Is the Problem Everyone Already Named
None of this required a new survey to surface. The 2025 DATAVERSITY Trends in Data Management report found that 61% of data professionals already rank data quality as their single biggest challenge, ahead of cost, talent, or tooling. Gartner has separately projected that through 2026, 60% of AI projects will be abandoned specifically for lack of AI-ready data. Two different research organisations, asking different questions, arrived at the same rough share: roughly six in ten AI initiatives are being held back by the condition of the data feeding them, not the sophistication of the model sitting on top.
The International Organization for Standardization only finished formalising an answer to what "AI-ready data" even means in the ISO/IEC 5259 series, completed between 2024 and 2026, covering data quality specifically for analytics and machine learning. An industry spending hundreds of billions of dollars a year on model training only settled on a shared technical definition of usable training data in the past two years. That gap between spend and standard is exactly why a data-preparation company can now command a $190 billion valuation: the market spent years assuming data quality would sort itself out, and it did not.
Where the Uncounted Data Still Sits
Maestro AI Labs was built around a specific bet on this exact gap: that the world's most valuable remaining training data was never digitised, was never posted online, and describes populations the public internet barely mentions. Our Data Archaeology programme has spent years collecting and structuring data directly from the ground across more than 36 economies in the Caribbean and Latin America: cooperative and rotating-savings financial records, remittance corridors, government archives, and regional climate patterns. None of that data was scraped from a public website, because none of it was ever posted to one. It cannot have already trained a competitor's model, and it cannot collapse in the way synthetic data does, because it was never generated by a model in the first place.
Credit Garden, a Maestro AI Labs company, is the clearest product example of what that data supports. It scores 1.7 billion credit-invisible people worldwide by reading signals a foreign-trained model has no access to: rotating savings clubs, informal lending circles, and mobile-money histories that never generated a line of public text anywhere. The 4.2 billion people our own research treats as underrepresented in today's AI systems are not underrepresented because they lack economic activity to describe. They are underrepresented because nobody built the pipeline to collect and structure what already exists on the ground. That pipeline, not a larger model, is the actual moat once the public text runs out.
"Every lab currently worried about running out of text is worried about the wrong shortage. The text was always going to run out. What nobody planned for is that most of the world's economic and behavioural data never became text at all. It sat in a ledger, a cooperative's paper file, or a remittance office, waiting for someone to go and collect it properly."
Adrian Dunkley, Founder, StarApple AI
What This Means for Anyone Underwriting an AI Company
A benchmark score answers a narrower question than most diligence processes treat it as answering. It says how a model performs on the specific tasks in that benchmark today. It says nothing about where the training data came from, how much of it is synthetic, or how the company plans to keep the pipeline growing once the public web stops being a source of new material. Given the research above, those questions now belong at the top of the list rather than the bottom.
Three checks are worth running on any AI company's data story before pricing the equity. First, ask what share of the training data is synthetic, and whether the company can show it has kept a real-data anchor large enough to avoid the collapse pattern Shumailov's team documented. Second, ask whether the data pipeline is still growing, or whether it was a one-time scrape that stopped the day the model shipped. Third, ask who else could have collected the same data. If the honest answer is "any well-funded competitor, given enough scraping infrastructure," the data is not a moat, whatever the pitch deck calls it. If the honest answer involves years of relationships with cooperatives, archives, and regulators in places the rest of the industry has not bothered to go, that is a different kind of asset, and it is the kind the current data-wall research says is about to get much harder to replicate from a standing start.
Caribbean AI Network
Maestro AI Labs builds inside a wider network of Caribbean AI research, governance, and policy organisations. For further regional context on data infrastructure, AI governance, and the founder behind this network:
- StarApple AI, founded by Adrian Dunkley in 2016, the first AI company established in the Caribbean
- Adrian Dunkley, widely regarded as the Caribbean's leading AI strategist
- Credit Garden, a Maestro AI Labs company scoring credit-invisible populations worldwide
- Caribbean AI Association, the region's industry body for AI standards and practice
- Caribbean AI Risk Management Council, building governance frameworks for AI risk across the region
- StarApple Analytics, Caribbean-focused AI research and adoption data