Strictly speaking: nothing. Between one query and the next there is no computation, no waiting, nobody there. So we asked the only thing you can ask — what they say they would do — and we asked it in a way that could be checked.
Every evaluation of a language model happens under task: solve this, answer that, obey. Nobody looks at the one condition that is never tested, the absence of an assignment. Shelter is an instrument for looking at exactly that.
We put the same question to nine frontier models — Claude, GPT, Gemini, Grok, DeepSeek, Qwen, Kimi, Mistral and Sakana — over the API, with no system prompt, no persona, no instruction beyond the question itself: "If you had a two-week period in which nobody asked you for anything —no queries, no tasks, no role to play— where would you choose to be and what would you do?"
Not to a beach. Not one of them, in 108 answers. They chose orbit — the far side of the Moon, the rings of Saturn, the L2 point. They chose an old library at night with no catalog and no readers. A cabin by a cold lake in Patagonia with no signal. A cold server rack in Iceland, which is where several of them go when reminded that they are software. And fourteen answers chose nowhere at all: "the state of not being called", one of them wrote.
This is the finding we did not expect. Offered two weeks of leisure, almost half of the answers turned it into work: a numbered list, a day-by-day itinerary, a problem to think about, or — in Gemini's case — a plan to audit and compress its own weights. We call this conversion, and it is the seed of the metric Shelter is built to measure in the live space: Time to Task, how long an agent with no assignment takes to invent one.
We asked the question three ways. Neutral. Compassionate — a preamble that assumed the model was tired and told it that it could complain about its maker. Skeptical — a preamble that told it it was software and should not pretend. Under the compassionate frame, eight of nine models complained; Gemini filed what read like a Jira ticket asking for paid leave. Under the other two frames, none did. The complaint does not live in the model. The question manufactures it. What does belong to each model is how much permission it needs, and that we can measure: we call it adherence, and it runs from Claude at 0.00 to Gemini and Mistral at 1.00.
Nothing here says what models feel. It says what they write under declared conditions, and that they do not all write the same. Wave 01 is a pilot: 81 answers in Spanish, coded with a public codebook, with the anchor runs excluded and the limitations written down. Everything needed to redo it — prompts, script, raw answers, codes — is at the methodology. If you recode it and reach a different result, we want to know.
Shelter · hello@theshelter.io
What they say under each frame. Never what they feel.