A new schizophrenia study identified 766 genes associated with the disorder — 641 of them never seen in previous transcriptomic analyses. The researchers used AI to reconstruct coordinated genetic activity across six brain regions, drawing on samples from hundreds of donors and data from over 102,000 people. It is, by any measure, a significant advance in understanding a disease that has resisted clean genetic explanations for decades. It was published in Nature Genetics and involved teams from the Lieber Institute for Brain Development, the University of Bari, and dozens of psychiatric centers worldwide.
The surface story is straightforward: AI is cracking biology open. That's the headline you'll see. And it's not wrong — but it skips the part that actually matters.
The datasets and public biological archives are the load-bearing walls. The models are the furniture. Everyone is investing in furniture. - The Systems Bastard
The Hidden Foundation
Here's the mechanism everyone is glossing over. The reason AI works on problems like protein folding or schizophrenia genetics is not mainly the model architecture. It's the dataset underneath. According to recent analysis, AlphaFold's primary condition for success was the Protein Data Bank, a dataset of roughly 170,000 experimentally validated protein structures that took 53 years of international scientific cooperation and an estimated $21 billion worth of experimental work to assemble. The Protein Data Bank itself operates on federal funding. That's the entire operational budget for a resource that enabled a Nobel Prize.
Now look at what's happening on the other side of the ledger. Major technology companies have announced substantial capital expenditures for infrastructure, almost all of it poured into data centers, GPUs, and cooling systems. That's not R&D. That's infrastructure to run and serve models. Meanwhile, the U.S. federal government maintains a core AI research budget that is substantially smaller. The ratio between private infrastructure investment and public research funding is significant. Industry is spending considerably more than the government spends on AI research, and virtually none of that industry spending goes toward the slow, painstaking, publicly shared dataset construction that made the breakthroughs possible in the first place.
This is the structural problem. The datasets and public biological archives are the load-bearing walls. The models are the furniture. Everyone is investing in furniture. According to academic researchers, being an AI academic today presents challenges similar to being a biologist in a world in which private companies had exclusive control over gene-editing tools. Universities can study how major AI systems behave, but they can't do detailed research on their design and training. They can't even afford to query the APIs rigorously — the cost of repeated calls to frontier models is prohibitive for most labs. Federal scientific funding in the U.S. is contracting. Research funding programs have faced reductions and pauses in new applications. The pipeline is freezing at exactly the moment it should be expanding.
According to recent commentary, AI agents — systems that model the research process itself, not just pattern-match against existing data — are a promising path to accelerating science. Fine. But agents still need something to reason about. Efforts like the Protein Data Bank took decades because they required international coordination, experimental labor, and sustained public funding — all things the current incentive structure is actively undermining. The schizophrenia study analyzed tissue samples from six brain regions across hundreds of donors. That kind of work is expensive, slow, and unglamorous. No venture fund is writing checks for it.
The fix is specific, boring, and would annoy every AI company currently claiming to advance science: mandate that any company using publicly funded datasets in commercial AI products pay a licensing royalty back into the funding of those datasets. Not a voluntary partnership. Not a PR initiative. A structural revenue loop that makes the damn load-bearing walls self-sustaining, instead of relying on the goodwill of congressional appropriations committees that have already demonstrated exactly how much goodwill they have.