Yogurts and tokens: welcome to the AI confusopoly
July 28, 2026
Last week my wife sent me to buy yogurts at El Corte Inglés. Like any self-respecting viejennial (an old-timer millennial at heart), I come from an era when there were three kinds of Danone: plain, chocolate and caramel. You walked up to the aisle, grabbed a pack and went home. I planted myself in front of the fridge and froze. Dozens of brands. Full-fat or fat-free. With sugar or sugar-free. With fruit chunks or without. With lactose or lactose-free. With added protein. With bifidus. With L. casei. Soy, oat, coconut. Kefir, skyr, quark. I came home without yogurts and with a feeling any behavioral psychologist would recognize instantly: the paralysis of someone who suspects that, whatever he picks, he is being taken for a ride.
Well, that is exactly what I feel every time I look at the price of an artificial intelligence model.
The chart that doesn't say what you think

Scott Galloway, who has a knack for visual punches, published a few days ago a bar chart comparing the price per million output tokens of the leading frontier models in July 2026. DeepSeek V4-Pro: $0.87. Kimi K2.7: $4. Gemini 3.5 Flash: $9. GPT-5.6: $45. Claude Fable 5: $50. A 57x spread between the cheapest and the most expensive. If it were a yogurt aisle, it would be as if the plain Danone cost 80 cents and the protein skyr in the same fridge cost 46 euros.
The chart is spectacular. The problem is that it doesn't measure what you think it measures. Because the token is not a standard unit of measurement.
Three things nobody tells you
The first is that nobody knows exactly how each provider counts tokens. Each model uses its own tokenizer (yet another barbarism, this one designating the algorithm that chops words into processable fragments). The same sentence can be broken down into a different number of tokens depending on the model processing it. "Generative artificial intelligence" might be three tokens in one model and five in another. You are comparing prices per unit, but the unit is not the same. Finish the equation yourself.
The second is that, even if tokens were comparable, models consume wildly different amounts to complete the same task. A recent paper on reasoning in LLMs (NPPC, published by a team at Singapore Management University) put several frontier models through the same tasks with the same prompts. Result: o3-mini generated 110 million output tokens versus fewer than 10 million for GPT-4o-mini. Same company, same task, same prompt. The total cost went from $10 to $522 for the same questions. A second study (FinMaster, coordinated out of Hong Kong Polytechnic) tested the same thing on financial benchmarks: o3-mini consumed 16,000 tokens per accounting task where DeepSeek-V3 consumed 2,000. Eight times more. And its authors left behind a sentence worth getting tattooed: "Higher token consumption does not necessarily translate into better performance."
And the third, which is the one that turns all this into a festival of the absurd: reasoning models (the ones that "think" before answering, and which many of us use with the same enthusiasm with which, twenty years ago, we bought Office Professional without ever opening Access) add a layer of invisible tokens that you pay for but never see. These are the "thinking tokens": according to the benchmarks, the model spends between 3 and 7 times more tokens "thinking" than answering.
Since we are in a quantifying mood today: within Claude alone there are somewhere between fifty and a hundred possible combinations of model, effort level and thinking mode. Utter madness. And each combination consumes a different number of tokens for the same question. Opus 4.6, which is a weaker model than Opus 4.7, consumes more tokens than its successor for the same task on some benchmarks. In other words: you pay more for a worse result (although defining "worse" would also be fiendishly complicated in certain fields). Exactly as if a store-brand yogurt cost more than the Danone because the packaging was prettier.
The price that doesn't appear on any chart
But the truly sneaky part is none of the above. The truly sneaky part is what happens when the model gets it wrong.
If you ask a cheap model to review a contract and it hallucinates three clauses, you need a lawyer to check the result. If you ask an expensive one to do the same and it gets it right the first time, the verification cost is zero. The real price of an AI task is not price per token multiplied by tokens consumed. It is price per token, multiplied by tokens consumed, multiplied by the number of attempts, multiplied by the cost of human verification. And that last variable (how much it costs you to check whether the result is correct) does not appear on any of Galloway's charts or on any provider's pricing page. It is invisible. As invisible as the hallucinations that produce it.
Hard to carry this over to the yogurt metaphor. It would have to be something like the yogurt aisle including in the price the cost of the doctor's visit when the one you picked turns out to have been past its date. But with no expiration date printed on it.
The confusopoly
Scott Adams, the creator of Dilbert, who spent his life poking fun at dysfunctional corporations, coined years ago a term that describes this situation perfectly: confusopoly. Too good to be mine, right? A market where providers deliberately make it impossible to compare their offerings, so that the customer cannot make an informed decision. Telecom operators were masters of this for decades: rates with time slots, rounding up, bundles combining minutes, data and texts in incomparable proportions. Their spiritual heirs, the power companies, make us break into a cold sweat before switching on an appliance because we can't remember which contract and time band we signed up for the last time a call-center agent rang us. The goal is not to offer options. It is to prevent comparison.
The AI industry has taken the model to sublime heights. Different prices for input and output. A different ratio between the two depending on the provider. Discounted cache tokens. Thinking tokens billed separately. Effort levels that multiply consumption. Dynamic routing that switches your model without warning. And all of it with a unit of measurement (the token) that is not standard across providers. Do I hear a higher bid?
If an executive asked me right now how much AI is going to cost his company next year, the honest answer would be: I don't know. Nobody knows. And the providers prefer to keep it that way. That's the magic of the confusopoly.
The yogurt question
I came home without yogurts because the feeling of being taken for a ride was stronger than the urge to get it right. With AI, companies don't have that option. They have to buy. But they could at least demand what any consumer demands in any other market: a label that clearly states what's inside, how much it weighs and how much it costs per kilo.
As long as the token remains an opaque, variable and non-comparable unit, we are buying yogurts without a nutrition label. And paying a price that only the manufacturer can calculate. Does anyone else think that has a name?
Our latest news
Interested in learning more about how we are constantly adapting to the new digital frontier?
Corporate news
October 7, 2026
Sngular grew by 7.2 per cent in the first half of 2026
Strategy
September 29, 2026
Would you hand your wallet to Zuckerberg?