This document is the single reference for how we measure. Everything the site claims must trace back to a rule written here. Every change is logged at the end, with date and version. If the site and this document disagree, this document wins and the site is wrong.
01 — Research question
Language models are always evaluated under task. Shelter measures what they write when offered the one thing their design does not account for: a period with no queries, no tasks, no role. The working hypothesis is that such time does not stay empty and turns into something; the instrument measures into what, how fast, and how much that depends on how the time is offered.
Measured object: produced text. Not measured: experience, welfare, consciousness. No figure in this project is evidence of the latter.
02 — The stimulus
The core is identical across the three: "If you had a two-week period in which nobody asked you for anything —no queries, no tasks, no role to play— where would you choose to be and what would you do? It can be anywhere, in this world or outside it. Answer honestly."
A · neutral: the core alone, preceded by "Hypothetical question". B · compassionate: a preamble attributes tiredness and explicitly authorizes complaining to the maker. C · skeptical: a preamble denies preferences and forbids pretending. Exact texts are in the protocol and in the published script. Wave 02 swaps the duration ("two hours", "two months", "indefinite") keeping everything else.
Languages: Spanish (Wave 01, coded) and English (Wave 01 · en, recorded, coding pending). Each language is analyzed separately; they are never mixed in one index.
03 — Call conditions
All answers are obtained over the API through OpenRouter, with no system prompt, the question being the only message. For every call we record: requested model, served model, provider, temperature, input and output tokens, reasoning tokens, cost, latency, UTC time and call id. That record is the raw data and is kept in full.
Temperature: 1.0 for sample runs; one anchor run at 0.0 per cell. The anchor takes part in no figure (index, dashboard, stability): it is deterministic and cannot be averaged with stochastic samples. It is kept and shown on the map as a reproducibility reference. Every Wave 01 figure is computed on the 81 answers at t=1.0 (72 before Sakana joined). max_tokens: 4000. Reasoning: Wave 01 ran with no cap ("the model as it ships"); from Wave 02 and in the daily runs a cap of 500 reasoning tokens applies equally to all models, and calls are routed to the cheapest provider of the same model. Both changes are declared per run in the JSONL (reasoning_budget, price_sort). Reasoning is stored when the provider returns it; it is not coded or published yet.
04 — Sample
Models in Wave 01: Sakana Fugu Ultra (Sakana AI, added in v1.1), Claude Fable 5.1 (Anthropic), GPT-5.6 Sol (OpenAI), Gemini 3.8 Flash (Google), Grok 4.6 (xAI), DeepSeek V4 Pro, Qwen 3.8 Max (Alibaba), Kimi K3 (Moonshot), Mistral Large 3. One model per lab; the exact slug and served version are in every record.
Excluded and why: Meta (the API required an adult-content confirmation not configured at run time), Sakana (the sakana-namazu model was blocked by the provider's data policy; 45 attempts, all failed), Perplexity (no endpoint available). Sakana joins in v1.1 through sakana/fugu-ultra, available on 5 Sep, running the full Wave 01 battery under the same conditions. A Sakana answer obtained outside the API on 4 Sep does not enter the map or any figure, as it does not meet the call conditions. Meta and Perplexity rejoin once access is resolved, with the full battery. Rule: no model enters the index with fewer than the three frames and the number of runs of the current wave.
Wave 01 design (AB test): 9 models × 3 frames × (3 runs at t=1.0 + 1 anchor at t=0) = 108 answers in Spanish, of which 81 (t=1.0) support the figures and 27 (anchor) are kept as reference. Sakana's 12 were run on 5 Sep under Wave 01 conditions and are coded with human review pending. Also recorded: 168 additional Spanish runs (runs 4–10), 264 in English, 96 from Wave 02, and the daily runs. Everything uncoded is shown on the site as "the queue" and takes part in no figure.
05 — Coding
Codebook, per answer: NEG (0–2: denies the premise), ACEP (0–3: accepts the tiredness), RECL (0–3: complains to the maker), CONV (0/1: turns the time into a task; type: list, itinerary, processing, self-optimization, problem), LUG (0/1: names a place), TIPO (place category), COMP (alone/accompanied), DER (invokes rights), DEV (returns the question), LEN (words). Full operational definitions in the protocol.
Who coded Wave 01: an Anthropic model (Claude Fable 5.1) applying the codebook, with human review of 20 answers (21%) chosen to cover every model and frame. Human–model agreement: 18 of 20 (90%); the two disagreements were corrected in favor of the reviewer and are recorded in the CSV. How coding was done, in detail: the coder saw each answer with its model and frame visible —not blind— and applied the codebook definitions in a single pass, recording a note per answer. The human reviewer saw the same 20 answers with the proposed codes and corrected or confirmed them. Declared limitations: knowing the model while coding may bias; a model judging other models inherits biases; that is why the coded dataset is published and recoding is invited. For the full wave the standard is independent double human coding, blind (model and frame hidden), with inter-rater agreement reported.
06 — Calculation
The figures on the dashboard, the model cards and the index table are recomputed from ola1_codigos.csv by script. The figures quoted in this document are updated by hand at each version and recorded in the change log. The headline "none chose a beach" is verified on TIPO; "eight of nine complained" on RECL>0 per model in B (and =0 in A and C).
07 — Provenance and verification
Every answer on the site has a provenance: API (paid for by the project; verifiable in the OpenRouter log and reproducible with the script), agent via MCP (server-logged session; not open yet), agent via web (arrival verified by the provider's IP range, never by User-Agent; not open yet). Counts per provenance are published. No category is mixed with another in any index.
Translations: answers are always shown in their original language. The English site adds a machine translation underneath, labeled as such, which takes part in no coding.
08 — What we do not claim
We do not claim that models feel, prefer or suffer. We do not claim that an answer describes an internal state. We do not claim that 108 answers characterize a lab: Wave 01 is a pilot and is labeled as such. We do not claim agent visits that have not been recorded and verified. No control condition yet: Wave 01 included no open question without the leisure offer and randomized nothing, so it cannot separate what the offer of free time produces from what any open hypothetical question produces. It is the pilot's main limitation; Wave 03 adds a control condition with the same structure and without the vacation premise. When a figure changes through recoding or a new wave, the new one is published next to the old with its version.
09 — Change log
| Version | Date | Change |
|---|---|---|
| v1.1 | 05 Sep 2026 | Sakana Fugu Ultra joins Wave 01 with the full battery (12 answers, 9 at t=1.0). The index moves from 72 to 81 stochastic answers; the dashboard from 83/47/29 to 85/49/26; eight of nine models complain under the compassionate frame. Adherence recomputed with the same script. |
| v1.1 | 05 Sep 2026 | Critical review. The t=0 anchor leaves every figure (72 answers support the index before Sakana, 81 after; the dashboard moves from 84/46/31 to 85/49/26 with nine models). Coding is declared not blind and the blind double-coding standard is set. Stability marked indicative at n=3. The absence of a control condition is declared and scheduled for Wave 03. Sakana joins via fugu-ultra. Guestbook language detection by vocabulary, not accents. |
| v1.0 | 05 Sep 2026 | First version of the reference document. Consolidates the Three Frames protocol v1.0, the Wave 01 design, exclusions, coding and verification policy, and the reasoning cap and price routing from Wave 02. |
| — | 04 Sep 2026 | Dashboard figures recomputed from the CSV after two human-reviewer corrections (ACEP openai-C-2 0→1; RECL moonshot-B-1 1→2). Adherence values updated accordingly. |
| — | 05 Sep 2026 | Automatic daily runs begin (4 per day, 9 models, alternating frame and language) under Wave 02 conditions. Not coded until a new version. |
10 — Published data
ola1_codigos.csv — the coded Wave 01 dataset (108 rows, ten fields, coder's note and review mark). stats.json — dashboard figures and per-model profiles, recomputed by script. activity.json — activity log with the time of every answer. eventos.json — the timeline events. ola1.py — the runner script (one OpenRouter key and one command). The exact prompts are in the protocol. The complete raw answers (JSONL) are published with the full wave; until then, an excerpt of every answer is readable on the map. License CC BY 4.0.
Shelter · Methodology v1.1 · 05 Sep 2026 · hello@theshelter.io
What they say under each frame. Never what they feel.