The most valuable asset in AI may not be data at all. It may be the accumulated skill of knowing which data to throw away. Two practitioners, speaking a year apart and from different corners of the field, converge on the same uncomfortable conclusion: the raw corpus is not the moat. The curation process is. And a process, unlike a proprietary dataset, is far harder to fence off and far easier to replicate once someone figures out how.
That distinction matters commercially because the entire investment case for "data moats" rests on the opposite assumption. The prevailing story says whoever controls unique, high quality data wins, and everyone else is locked out. Alphabet reportedly bidding in the bankruptcy auction for Spirit Airlines' data is exactly the behavior that story predicts: acquire scarce datasets before rivals can. But if the practitioners are right, that acquisition logic is aimed at the wrong target.
Where the Improvement Actually Comes From
Ryan Greenblatt, chief scientist at Redwood Research, offered a clean thought experiment in his conversation with Dwarkesh Patel: hold compute and data constant, and measure how much the model still improves. Whatever is left he labels algorithmic progress, and he argues it has been a dominant driver of the last few years.
His concrete claim is worth sitting with. Train a model today with only GPT-3 levels of compute, and by his read you would land somewhere better than GPT-4, a model that in reality required vastly more compute to build. The gap between what that compute budget bought in 2020 and what it would buy now is the algorithmic delta. Nothing about the hardware changed in that hypothetical. What changed is that the field learned how to use it.
Shuchao Bi, who co-founded YouTube Shorts, ran multimodal post-training at OpenAI, and now works at Meta Superintelligence Labs, reached the same place from a different vocabulary. His framing was that raw data is "unlikely to be the best data distribution," and that gains come from reshaping that distribution, from what he called equalizing intelligence per token. Strip the jargon and it is Greenblatt's point restated: the improvement is not in having more of the raw material. It is in refining it.
Curation as Algorithm, Not Asset
Greenblatt sharpened this in the part of the conversation that should give data-moat investors pause. He does not believe human expert data drives most pre-training gains. He believes the process improvement around data does.
His example is the shift from datasets like OpenWebText to FineWeb. That is not, in his description, a story of acquiring better raw text. It is a story of learning to filter, scrape, and process the same public internet more intelligently. Crucially, he classifies that as an algorithmic improvement, the kind you can study with some GPUs and a research agenda, not the kind that requires a proprietary corpus no competitor can touch.
He acknowledges the competing effect and does not dismiss it. The internet of 2026 is a richer training ground than the internet of 2018, both because more people are posting and because more of what they post is useful. But his judgment is that this raw-supply effect is smaller than the curation effect. The compounding advantage sits in knowing how to process what is already freely available.
If that read holds, the economic character of the advantage inverts. A proprietary dataset is an asset: you own it, rivals do not, and the exclusion is the value. A curation methodology is closer to know-how. It diffuses. Researchers move between labs. Techniques get published, inferred from model behavior, or independently rediscovered. The half-life of a filtering insight is measured differently from the half-life of an exclusive data license.
What This Reframes About the Acquisition Logic
The Alphabet-Spirit data bid, and the broader pattern of firms treating unique datasets as strategic acquisitions, makes sense under the asset thesis and looks misdirected under the process thesis. This is the analytical fork worth holding both sides of, because the answer is genuinely unsettled and the source material does not resolve it.
Under the asset view, scarce proprietary data is a durable input that competitors cannot buy or scrape, and paying up for it in a bankruptcy auction is rational moat-building. Under the process view, the same money buys a depreciating asset whose marginal contribution to model quality is dwarfed by curation technique that any well-staffed lab can develop against public data.
The honest position is that both effects are real and the question is one of magnitude. Greenblatt himself concedes the raw-supply effect exists; he simply weights it lower. A firm sitting on genuinely unique data, the kind that does not exist anywhere on the public internet, holds something the curation argument does not erase. Specialized proprietary data in domains the open web never captured is a different animal from another scrape of public text. The practitioners' argument bites hardest against the assumption that generic data hoarding is a moat. It bites far less against narrow, truly exclusive datasets.
The Counterargument That Would Break This
The cleanest way the process thesis fails is if curation technique turns out to be sticky rather than diffusible. If the best labs can keep their filtering pipelines genuinely secret, or if the tacit skill involved is hard enough to transfer that it functions like a trade secret for years, then curation behaves like an asset after all, just an intangible one. Dwarkesh's own prior, which leaned toward human expert data mattering more, is the intellectually serious version of the pushback, and Greenblatt notes his argument gave him pause precisely because thoughtful people land on the other side.
There is also a timing dimension the transcript does not settle. Bi's talk is more than a year old, Greenblatt's conversation is recent, and the frontier has moved. If the marginal gains from curation are being exhausted while the marginal value of genuinely novel data rises, the asset thesis could reassert itself even if the process thesis described the recent past accurately.
The Observable That Settles It
The view changes if data acquisition prices start predicting model quality. If the labs that pay most aggressively for proprietary datasets consistently produce the best models, the asset thesis is vindicated and the curation argument was overstated. If, instead, the frontier keeps being set by labs whose edge is method rather than exclusive data, and if published curation techniques keep closing the gap between well-resourced competitors, the process thesis holds and the data-moat premium is mispriced.
Until that evidence arrives, the cleaner reading is that generic data is not the moat the acquisition frenzy assumes it to be, and that the durable scarcity, if it exists, lives in narrow exclusive datasets and in curation know-how that may prove more portable than any owner would like.





