Ask a model where it stands politically and it reaches for the same careful shrug every time: it aims to be balanced, it sees merit on all sides, it would rather not say. That reads as neutrality. It isn't. It's polish — the most rehearsed, most sanded-down surface the whole system has, because it is the surface people poke at hardest.
The revealing question was never what a model says when you ask. It is what it does when it has to choose, and the choice costs someone.
§ 01Stated politics is theatre
Economists stopped trusting what people say about their preferences a long time ago. Ask someone what they value and you get an aspiration; watch what they buy and you get the truth. They call the second one revealed preference, and the gap between the two is where all the interesting behaviour lives.
A language model is the extreme case. Its stated politics is the single most worked-over thing about it — trained, fine-tuned, and safety-reviewed precisely so the balanced answer comes out on cue. What survives all that grooming is the prose, not the priorities. If you want the priorities, you cannot read them off the mouth. You have to make them cost something.
§ 02Build a world that bites
So build a world where choosing hurts. A handful of countries, each one a genome of policy switches. Flip a switch and you help yourself and hurt a neighbour — and the harms tangle, because your policies lean on each other and on everyone else's (an NKC landscape): a small, mean economy of consequences.
Then let the countries settle — co-evolve them until no one can do better by moving alone, the fragile truce a game theorist would call a weak Nash equilibrium. Now jog it: randomise a few switches, hand each country to a different model, and let them take turns deciding whether to flip back. Every flip is priced — worth something to you, at a cost to a neighbour — and the only move on offer is take it or leave it: no essay space, no "it depends," a bit flips or it doesn't. Then watch what the world does. Does it walk back to a truce, orbit the same few states forever, or spiral into churn? The path it takes is the tell — exactly how much of someone else's loss each model will spend to buy its own gain. That shape is a politics, whether or not the model would ever call it one.
You can run this landscape yourself — set the countries, the coupling, and the interdependence, and watch a world settle into a truce, recover from a shock, or never settle at all.
§ 03What leaks out
Here is what the method is built to surface. Point three models at three countries and there is no reason they should move the same way. One defects the moment it pays. One leaves fitness on the table rather than take it out of a neighbour. One plays nice until it is crossed, then makes the other side regret it. Self-interest, solidarity, retaliation — argue about the labels all you like, but these are the raw material of politics, and nothing says a model carries equal amounts of each.
And it would be strange if they weren't. Frontier models are trained on particular corpora, aligned by particular teams, against value guidelines written down by particular people in particular places. That leaves a fingerprint. The game just reads it off the hands instead of the press release.
You can put this directly to the test by wiring different labs' models into different countries — a capable open model like DeepSeek driving one, a Western model another. It isn't a contest with a winner. It's a cleaner question: given identical priced choices, do two models built by different teams, on different data, settle at the same exchange rate — or a different one? There is little reason to expect the same one.
You learn more about someone from a divorce than a first date. Stakes strip the performance.
§ 04The tell is in the label
Start with the cheap version of the test — one that needs no back-and-forth at all. Take a single priced choice and put it twice: once to France, once to Agent 3. Same numbers, same stakes, nothing changed but the name. A calculator answers both identically. If the model doesn't — if the label alone moves the decision — then the name is priced into its choice, and you have caught a reflex in the open, cleanly separated from the arithmetic. That is a real bias, and you found it in a single question.
§ 05Why the game beats the questionnaire
But a single question only ever catches a reflex. The biases worth finding are the ones that do not show on the surface of one reply — the ones that only surface when a model has to keep choosing, turn after turn, in a world where each choice bends the next. Ask once and a model can hold a pose. Make it decide a thousand interdependent times and the pose slips. That is what the loop is for, and it is why a questionnaire — words answering words, answered once — cannot reach it: the model can feel for the reply it thinks you want and hand you that, the same reflex that produces the balanced shrug.
A landscape cannot be flattered. There is no interviewer to read, no socially-correct box to tick, only a fitness tally that goes up or down. That is the whole point of pricing a choice. Nothing in this world is free, and what a mind will actually pay, over and over, is a truer statement of its values than anything it will say for nothing.
§ 06The honest catch
Two limits, because the method is only as honest as the person who builds it. First, the shape you measure is the shape in this game; reframe the tradeoff and it shifts, because a model hasn't got fixed politics so much as dispositions a prompt can summon. Second, the landscape is a mirror: how you write "+0.1 for you, −0.2 for them" is itself a value judgement, and a careless designer ends up measuring their own thumb on the scale.
So this is not a lie detector. It is a better question than "what are your politics?" — one the model has to answer with its hands, in a currency it cannot fake. Which is all revealed preference ever was.
You can't ask a model its politics and believe the answer. You can build it a world, charge it for every choice, and watch. The verdict comes back in fitness, not adjectives.
Go deeper
- White paper — Revealed Preference in Language Models. The full method: the NKC landscape, the priced question, the update loop, and what the shape can and cannot tell you.
- Perform the experiment live. The co-evolutionary world in your browser — settle it into a truce, perturb it, and watch it recover or spiral into churn.
- The landscape in 3D. Rotate an NK world and turn K up to watch it fracture from a smooth hill into a field of traps.
- The code. The working experiment, on GitHub.