In May 2026, the Army CIO announced unlimited AI tokens for personnel. By mid-June, the entire Army pool was exhausted and hard limits had to be reimposed. One service burned through a full year's allocation in roughly a month. Meanwhile, a new study from Princeton and the University of Chicago found that LLMs screening job candidates spontaneously develop their own discriminatory biases from experience, stereotyping fictional ethnic groups into occupational roles at rates higher than human participants. Two stories. Same root cause.
The surface explanation for the Army's token burnout is simple overconsumption: too many users, not enough budget. The Pentagon had announced a significant increase in AI use across the DoD — from initial adoption to 1.5 million personnel by June 2026. The number went up; the money didn't follow. A classic procurement screw-up. The surface explanation for the bias study is equally comforting: the models inherited stereotypes from training data. Old problem, old hand-wringing. Nothing to see.
When you tell a bureaucracy "use this tool more" and measure nothing but whether they used it, you are not adopting AI. You are laundering an adoption statistic. - The Systems Bastard
ERROR: GOODHART'S LAW APPLIED TO INSTITUTIONAL COGNITION
Both failures share a mechanism: organizations chasing an adoption metric without measuring what the adoption actually produces. The Army didn't exhaust its tokens because soldiers found brilliant uses for generative AI. It exhausted them because it built an incentive loop that rewarded consumption itself. Employees got 200,000 tokens per month with automatic top-ups if they ran through them. People who weren't using their allocation received nudge emails urging them to use more. The system was architecturally identical to a mobile game drip-feeding premium currency — except the currency cost real money.
The Pentagon wanted a number — adoption of AI across its workforce — and it got it. What it didn't get was any structural mechanism to distinguish productive use from waste. When you tell a bureaucracy "use this tool more" and measure nothing but whether they used it, you are not adopting AI. You are laundering an adoption statistic.
The Princeton hiring study reveals the downstream version of this same disease. When organizations hand decision-making to LLMs, the models don't just absorb human biases — they manufacture new ones with ruthless efficiency. The researchers used fictional ethnic groups (Tufa, Aima, Reku, Weki) specifically to control for existing stereotypes. It didn't matter. The models saw one Aima fail as a doctor and started routing all Aimas to different occupational roles. The models stereotyped candidates from different groups into different jobs at higher rates than human participants in the original study. And the more capable the model, the worse the bias, because better reasoning meant faster pattern-locking from thinner evidence. As the researchers noted, stronger LLMs favor candidates from a group if earlier assignments of similar jobs succeeded — turning a single data point into a policy.
This is what happens when you optimize a system for pattern completion and then put it in charge of consequential decisions without feedback correction. The Army optimized for token throughput. The AI models optimized for pattern exploitation. Neither had a mechanism to ask: is this actually good?
The connection isn't metaphorical. The DOD uses Ask Sage for acquisitions through its Chief Digital and AI Office. Ask Sage runs ChatGPT, Gemini, and Llama — three of the model families tested in the Princeton study. If the Army is using these models for tasks like "reclassifying personnel descriptions" — which it is — and those models spontaneously develop occupational stereotypes from feedback, then users burning through tokens aren't just wasting money. They may be encoding bias into personnel systems at an institutional scale nobody is auditing.
The fix is deeply boring and therefore unlikely to happen: decouple AI adoption metrics from AI outcome metrics, and measure only the latter. Stop counting tokens consumed. Start counting decisions reversed, errors caught, time saved on specific workflows with before-and-after measurement. Kill the nudge emails — if someone isn't using the tool, that's data, not a problem to solve. And for the love of God, run bias audits on any model making or informing personnel decisions, because the research results suggest this isn't theoretical. They're a genuine warning sign.