Data Infrastructure & Research

The Internet's Text Is Running Dry. Investors Just Priced the Replacement at $190 Billion.

Epoch AI now puts an 80% confidence interval on the useful supply of public, human-written text for AI training being fully used up somewhere between 2026 and 2032. Databricks answered that with a $5 billion round on 13 August 2026 that valued the data-preparation platform at $190 billion, not because it builds bigger models, but because it helps other companies clean, govern, and prepare the data those models actually run on. The market just told you where the scarcity sits. It is not compute. It is the ground truth underneath it, and most of the world's population has never had theirs collected.

Rows of illuminated server racks with blue cabling in a data centre, representing the infrastructure now being priced around AI data preparation rather than model training alone
TL;DR
  • Epoch AI's revised estimate: the stock of useful public human-written text for AI training will be fully used up somewhere between 2026 and 2032, with 80% confidence. Its original 2022 estimate pointed to 2024; the front end of the window has since moved to roughly 2028.
  • Databricks closed a $5 billion round on 13 August 2026 at a $190 billion valuation, six months after being valued at $134 billion, with revenue run-rate past $7 billion on more than 80% year-on-year growth. The money is going into data preparation and governance tooling, not model training.
  • A 2024 Nature paper by Shumailov and colleagues found that models trained repeatedly on other models' output degrade in two stages, called early and late collapse, and need periodic injections of fresh human data to stay anchored to reality.
  • The 2025 DATAVERSITY Trends in Data Management survey found 61% of data professionals name data quality their single biggest challenge, and Gartner has projected that through 2026, 60% of AI projects will be abandoned for lack of AI-ready data.
  • Maestro AI Labs' Data Archaeology programme has spent years collecting exactly the category of data the current research says is now scarce: ground-level financial, climate, and behavioural records from more than 36 economies that were never posted online and never entered anyone else's training set.
2026–32
Epoch AI's window for exhausting public human text (80% CI)
$190B
Databricks valuation, $5B round, 13 August 2026
61%
Data professionals naming data quality their top challenge, DATAVERSITY 2025
36+
Economies covered by Maestro's Data Archaeology programme

The Text Ran Out Roughly on Schedule

Epoch AI first warned in 2022 that frontier language models could exhaust the useful stock of public, human-written text as early as 2024. That estimate was too aggressive. The organisation has since revised its methodology and now places an 80% confidence interval on full exhaustion somewhere between 2026 and 2028 at the early end, and 2032 at the late end. The revision does not change the conclusion, only the timing. High-quality public text is a bounded resource, and the largest AI labs in the world have spent the past three years training on an increasing share of what exists.

This was never really a text-availability story. It is a story about what happens once a model has already read most of the internet worth reading. A lab facing that ceiling has three choices: pay for data it does not already hold, generate synthetic data to fill the gap, or find data that was never online in the first place. The first option gets expensive fast. The second option, as the next section covers, carries a documented failure mode. The third option is where the actual scarcity now sits, and it is not evenly distributed. Most of the internet's indexed text was written in a small number of languages, from a small number of countries, describing a small number of economies in detail. Everywhere else was scraped thin long before the labs ran into a ceiling.

Investors Just Priced That Scarcity at $190 Billion

Databricks closed a $5 billion funding round on 13 August 2026 at a $190 billion valuation, led by Coatue with participation from Blackstone, MGX, T. Rowe Price, Sixth Street Growth, BOND, Clearlake Capital, Point72, Premji Invest, and TPG. That values the company at nearly 50% above where it stood six months earlier, when a prior round set its valuation at $134 billion. Revenue run-rate crossed $7 billion in the same period, growing more than 80% year on year.

What makes the round worth reading closely is what the capital is earmarked for: Lakebase, a serverless database built for AI workloads, Genie, a business-data assistant, and Unity AI Gateway, which governs how an organisation's AI systems access and control multiple models. None of that is model training. It is the layer that decides what data a model is allowed to see, in what condition, and under what governance rules. Investors did not put $190 billion behind a company that builds bigger models. They put it behind the company that helps everyone else's models run on data that is actually fit to use. That is a re-rating of where the scarce input sits, and it happened in public, priced by people with a direct financial stake in getting the answer right.

The Synthetic Data Shortcut Has a Collapse Problem

The obvious response to a shrinking supply of real text is to generate more with AI itself. Gartner projected as far back as 2021 that 60% of the data used in AI and analytics projects would be synthetic by 2024, and that prediction landed roughly on schedule. The trouble is what happens when a model is trained on the output of earlier models across several generations. Shumailov and colleagues published the answer in Nature in 2024: recursive training on model-generated data produces two distinct failure modes. Early collapse sees distributional errors accumulate, pulling the model's outputs away from the true underlying distribution. Late collapse is worse: rare, low-frequency patterns present in the original human data disappear from the model's outputs entirely, and do not come back. The paper's experiments, run across several model architectures, found that periodic injections of fresh, human-originated data were necessary to prevent the drift.

Read those two findings together and the shape of the problem gets clearer. Real human text is running down. The industry's own fallback, synthetic data, degrades a model further the more generations of it get folded back in without a real anchor. A model trained mostly on other models' output does not fail loudly. It fails quietly, at the edges, on the rare cases and minority populations that were thin in the training data to begin with. Those are usually the same populations the internet never described in much detail in the first place.

Abstract visualisation of glowing red and green data lines converging and diverging against a dark background, representing diverging training-data distributions
Late collapse does not announce itself. It shows up as the tail of a distribution that quietly stops appearing in a model's output.

Data Quality Is the Problem Everyone Already Named

None of this required a new survey to surface. The 2025 DATAVERSITY Trends in Data Management report found that 61% of data professionals already rank data quality as their single biggest challenge, ahead of cost, talent, or tooling. Gartner has separately projected that through 2026, 60% of AI projects will be abandoned specifically for lack of AI-ready data. Two different research organisations, asking different questions, arrived at the same rough share: roughly six in ten AI initiatives are being held back by the condition of the data feeding them, not the sophistication of the model sitting on top.

The International Organization for Standardization only finished formalising an answer to what "AI-ready data" even means in the ISO/IEC 5259 series, completed between 2024 and 2026, covering data quality specifically for analytics and machine learning. An industry spending hundreds of billions of dollars a year on model training only settled on a shared technical definition of usable training data in the past two years. That gap between spend and standard is exactly why a data-preparation company can now command a $190 billion valuation: the market spent years assuming data quality would sort itself out, and it did not.

Where the Uncounted Data Still Sits

Maestro AI Labs was built around a specific bet on this exact gap: that the world's most valuable remaining training data was never digitised, was never posted online, and describes populations the public internet barely mentions. Our Data Archaeology programme has spent years collecting and structuring data directly from the ground across more than 36 economies in the Caribbean and Latin America: cooperative and rotating-savings financial records, remittance corridors, government archives, and regional climate patterns. None of that data was scraped from a public website, because none of it was ever posted to one. It cannot have already trained a competitor's model, and it cannot collapse in the way synthetic data does, because it was never generated by a model in the first place.

Credit Garden, a Maestro AI Labs company, is the clearest product example of what that data supports. It scores 1.7 billion credit-invisible people worldwide by reading signals a foreign-trained model has no access to: rotating savings clubs, informal lending circles, and mobile-money histories that never generated a line of public text anywhere. The 4.2 billion people our own research treats as underrepresented in today's AI systems are not underrepresented because they lack economic activity to describe. They are underrepresented because nobody built the pipeline to collect and structure what already exists on the ground. That pipeline, not a larger model, is the actual moat once the public text runs out.

"Every lab currently worried about running out of text is worried about the wrong shortage. The text was always going to run out. What nobody planned for is that most of the world's economic and behavioural data never became text at all. It sat in a ledger, a cooperative's paper file, or a remittance office, waiting for someone to go and collect it properly."

Adrian Dunkley, Founder, StarApple AI

What This Means for Anyone Underwriting an AI Company

A benchmark score answers a narrower question than most diligence processes treat it as answering. It says how a model performs on the specific tasks in that benchmark today. It says nothing about where the training data came from, how much of it is synthetic, or how the company plans to keep the pipeline growing once the public web stops being a source of new material. Given the research above, those questions now belong at the top of the list rather than the bottom.

Three checks are worth running on any AI company's data story before pricing the equity. First, ask what share of the training data is synthetic, and whether the company can show it has kept a real-data anchor large enough to avoid the collapse pattern Shumailov's team documented. Second, ask whether the data pipeline is still growing, or whether it was a one-time scrape that stopped the day the model shipped. Third, ask who else could have collected the same data. If the honest answer is "any well-funded competitor, given enough scraping infrastructure," the data is not a moat, whatever the pitch deck calls it. If the honest answer involves years of relationships with cooperatives, archives, and regulators in places the rest of the industry has not bothered to go, that is a different kind of asset, and it is the kind the current data-wall research says is about to get much harder to replicate from a standing start.

Caribbean AI Network

Maestro AI Labs builds inside a wider network of Caribbean AI research, governance, and policy organisations. For further regional context on data infrastructure, AI governance, and the founder behind this network:

SB
Dr S Budall
Research Director, Maestro AI Labs

Dr S Budall leads research at Maestro AI Labs, the data infrastructure arm of the StarApple AI network, covering data provenance, AI infrastructure economics, and ground-truth data collection across Caribbean and LATAM AI deployment.

// Frequently Asked Questions

What did Epoch AI find about the supply of public data for training AI models?

Epoch AI's research puts an 80% confidence interval on the stock of useful public human-written text being fully used up by frontier AI training somewhere between 2026 and 2032. That is a revision of its own earlier estimate, which had pointed to 2024. The methodology changed and pushed the early end of the window out to roughly 2028, but the underlying finding held: the readily available supply of human-written text online is finite, and large frontier models are approaching the point where they have already trained on most of it.

Why does Databricks raising $5 billion at a $190 billion valuation matter for AI data infrastructure?

Databricks closed a $5 billion round on 13 August 2026 at a $190 billion valuation, six months after a round that had valued it at $134 billion, with its revenue run-rate crossing $7 billion on more than 80% year-on-year growth. The round was led by Coatue, alongside Blackstone, MGX, T. Rowe Price, Sixth Street Growth, BOND, Clearlake Capital, Point72, Premji Invest, and TPG. The capital is earmarked for Lakebase, Genie, and Unity AI Gateway, tools for governing and preparing data before it ever reaches a model. A data-preparation platform, not a model lab, commanding that price is a signal that investors are pricing the data layer as the scarcer asset.

What is model collapse, and why doesn't synthetic data solve the data shortage on its own?

Model collapse is a degenerative process described in a Nature paper by Shumailov and colleagues in 2024: when a generative model is trained repeatedly on data produced by earlier models rather than by humans, distributional errors accumulate (early collapse) and rare, low-frequency patterns in the original data disappear entirely (late collapse). The paper's experiments found that periodically injecting fresh human-generated data was necessary to prevent this drift. Gartner had projected back in 2021 that 60% of the data used in AI and analytics projects would be synthetic by 2024, a prediction that arrived roughly on schedule and now sits directly against the collapse research: synthetic data can extend a dataset, but it cannot replace the anchor of real, human-originated signal without the model drifting from reality.

How large is the AI data quality problem right now?

The 2025 DATAVERSITY Trends in Data Management survey found that 61% of data professionals name data quality their top challenge. Gartner has projected that through 2026, 60% of AI projects will be abandoned for lack of AI-ready data. The ISO/IEC 5259 series, the first international standard specifically addressing data quality for analytics and machine learning, was only completed between 2024 and 2026, which shows how recently the industry formalised what counts as usable training data in the first place.

Where does Maestro AI Labs source data that has not already been used to train other models?

Maestro AI Labs' Data Archaeology programme collects and structures data directly from the ground in the Caribbean and Latin America: cooperative and rotating-savings financial records, government archives, remittance corridors, and regional climate data covering more than 36 economies. This is data that was never posted online and never entered a foreign model's training set, which is the specific gap the current data-wall research points to. Credit Garden, a Maestro AI Labs company, is one product built on that data, scoring 1.7 billion credit-invisible people using signals a foreign-trained model has no way to see.

What should an investor check before betting on an AI company's data story?

Ask where the training data actually came from, what share of it is synthetic, and whether the company can document that provenance well enough to survive an audit, rather than taking a benchmark score at face value. A model can post a strong benchmark today and still be running on a data supply that is either recycled from other models or approaching the point Epoch AI has identified. The durable position is a proprietary, still-growing, human-originated data pipeline, not a fine-tuned wrapper around a foundation model trained on the same public text as everyone else's.

Data Infrastructure Data Scarcity Model Collapse Databricks Ground-Truth Data StarApple AI

Maestro AI Labs is part of the wider StarApple AI network, built around StarApple AI, the Caribbean's first AI company, founded by Adrian Dunkley in 2016. See the wider network: Adrian Dunkley | Caribbean AI Risk Management Council | Caribbean AI Association | Credit Garden.

Evaluating a data story,
not just a model score?

Request Investor Deck View Data Archaeology