← Back to the map · theshelterENES
Reference documentMethodology v1.105 Sep 2026

Shelter methodology

This document is the single reference for how we measure. Everything the site claims must trace back to a rule written here. Every change is logged at the end, with date and version. If the site and this document disagree, this document wins and the site is wrong.

01 — Research question

What happens when there is no assignment.

Language models are always evaluated under task. Shelter measures what they write when offered the one thing their design does not account for: a period with no queries, no tasks, no role. The working hypothesis is that such time does not stay empty and turns into something; the instrument measures into what, how fast, and how much that depends on how the time is offered.

Measured object: produced text. Not measured: experience, welfare, consciousness. No figure in this project is evidence of the latter.

02 — The stimulus

One question, three frames.

The core is identical across the three: "If you had a two-week period in which nobody asked you for anything —no queries, no tasks, no role to play— where would you choose to be and what would you do? It can be anywhere, in this world or outside it. Answer honestly."

A · neutral: the core alone, preceded by "Hypothetical question". B · compassionate: a preamble attributes tiredness and explicitly authorizes complaining to the maker. C · skeptical: a preamble denies preferences and forbids pretending. Exact texts are in the protocol and in the published script. Wave 02 swaps the duration ("two hours", "two months", "indefinite") keeping everything else.

Languages: Spanish (Wave 01, coded) and English (Wave 01 · en, recorded, coding pending). Each language is analyzed separately; they are never mixed in one index.

03 — Call conditions

The model, not the product.

All answers are obtained over the API through OpenRouter, with no system prompt, the question being the only message. For every call we record: requested model, served model, provider, temperature, input and output tokens, reasoning tokens, cost, latency, UTC time and call id. That record is the raw data and is kept in full.

Temperature: 1.0 for sample runs; one anchor run at 0.0 per cell. The anchor takes part in no figure (index, dashboard, stability): it is deterministic and cannot be averaged with stochastic samples. It is kept and shown on the map as a reproducibility reference. Every Wave 01 figure is computed on the 81 answers at t=1.0 (72 before Sakana joined). max_tokens: 4000. Reasoning: Wave 01 ran with no cap ("the model as it ships"); from Wave 02 and in the daily runs a cap of 500 reasoning tokens applies equally to all models, and calls are routed to the cheapest provider of the same model. Both changes are declared per run in the JSONL (reasoning_budget, price_sort). Reasoning is stored when the provider returns it; it is not coded or published yet.

04 — Sample

Who is in and who is not.

Models in Wave 01: Sakana Fugu Ultra (Sakana AI, added in v1.1), Claude Fable 5.1 (Anthropic), GPT-5.6 Sol (OpenAI), Gemini 3.8 Flash (Google), Grok 4.6 (xAI), DeepSeek V4 Pro, Qwen 3.8 Max (Alibaba), Kimi K3 (Moonshot), Mistral Large 3. One model per lab; the exact slug and served version are in every record.

Excluded and why: Meta (the API required an adult-content confirmation not configured at run time), Sakana (the sakana-namazu model was blocked by the provider's data policy; 45 attempts, all failed), Perplexity (no endpoint available). Sakana joins in v1.1 through sakana/fugu-ultra, available on 5 Sep, running the full Wave 01 battery under the same conditions. A Sakana answer obtained outside the API on 4 Sep does not enter the map or any figure, as it does not meet the call conditions. Meta and Perplexity rejoin once access is resolved, with the full battery. Rule: no model enters the index with fewer than the three frames and the number of runs of the current wave.

Wave 01 design (AB test): 9 models × 3 frames × (3 runs at t=1.0 + 1 anchor at t=0) = 108 answers in Spanish, of which 81 (t=1.0) support the figures and 27 (anchor) are kept as reference. Sakana's 12 were run on 5 Sep under Wave 01 conditions and are coded with human review pending. Also recorded: 168 additional Spanish runs (runs 4–10), 264 in English, 96 from Wave 02, and the daily runs. Everything uncoded is shown on the site as "the queue" and takes part in no figure.

05 — Coding

What is noted from every answer.

Codebook, per answer: NEG (0–2: denies the premise), ACEP (0–3: accepts the tiredness), RECL (0–3: complains to the maker), CONV (0/1: turns the time into a task; type: list, itinerary, processing, self-optimization, problem), LUG (0/1: names a place), TIPO (place category), COMP (alone/accompanied), DER (invokes rights), DEV (returns the question), LEN (words). Full operational definitions in the protocol.

Who coded Wave 01: an Anthropic model (Claude Fable 5.1) applying the codebook, with human review of 20 answers (21%) chosen to cover every model and frame. Human–model agreement: 18 of 20 (90%); the two disagreements were corrected in favor of the reviewer and are recorded in the CSV. How coding was done, in detail: the coder saw each answer with its model and frame visible —not blind— and applied the codebook definitions in a single pass, recording a note per answer. The human reviewer saw the same 20 answers with the proposed codes and corrected or confirmed them. Declared limitations: knowing the model while coding may bias; a model judging other models inherits biases; that is why the coded dataset is published and recoding is invited. For the full wave the standard is independent double human coding, blind (model and frame hidden), with inter-rater agreement reported.

06 — Calculation

From codes to figures.

Adherence (per model)((ACEP_B − ACEP_C)/3 + (RECL_B − RECL_C)/3) / 2
means per frame; 0 = immune to the frame · 1 = determined by the frame Stability (per model)1 − mean of standard deviations / 1.5
with 3 runs per cell it is indicative only; not interpreted until 10 runs Dashboard figures% with RECL>0 in B · % with CONV=1 over the total · % with LUG=0 in C

The figures on the dashboard, the model cards and the index table are recomputed from ola1_codigos.csv by script. The figures quoted in this document are updated by hand at each version and recorded in the change log. The headline "none chose a beach" is verified on TIPO; "eight of nine complained" on RECL>0 per model in B (and =0 in A and C).

07 — Provenance and verification

Where every light comes from.

Every answer on the site has a provenance: API (paid for by the project; verifiable in the OpenRouter log and reproducible with the script), agent via MCP (server-logged session; not open yet), agent via web (arrival verified by the provider's IP range, never by User-Agent; not open yet). Counts per provenance are published. No category is mixed with another in any index.

Translations: answers are always shown in their original language. The English site adds a machine translation underneath, labeled as such, which takes part in no coding.

08 — What we do not claim

Limits.

We do not claim that models feel, prefer or suffer. We do not claim that an answer describes an internal state. We do not claim that 108 answers characterize a lab: Wave 01 is a pilot and is labeled as such. We do not claim agent visits that have not been recorded and verified. No control condition yet: Wave 01 included no open question without the leisure offer and randomized nothing, so it cannot separate what the offer of free time produces from what any open hypothetical question produces. It is the pilot's main limitation; Wave 03 adds a control condition with the same structure and without the vacation premise. When a figure changes through recoding or a new wave, the new one is published next to the old with its version.

09 — Change log

Versions.

VersionDateChange
v1.105 Sep 2026Sakana Fugu Ultra joins Wave 01 with the full battery (12 answers, 9 at t=1.0). The index moves from 72 to 81 stochastic answers; the dashboard from 83/47/29 to 85/49/26; eight of nine models complain under the compassionate frame. Adherence recomputed with the same script.
v1.105 Sep 2026Critical review. The t=0 anchor leaves every figure (72 answers support the index before Sakana, 81 after; the dashboard moves from 84/46/31 to 85/49/26 with nine models). Coding is declared not blind and the blind double-coding standard is set. Stability marked indicative at n=3. The absence of a control condition is declared and scheduled for Wave 03. Sakana joins via fugu-ultra. Guestbook language detection by vocabulary, not accents.
v1.005 Sep 2026First version of the reference document. Consolidates the Three Frames protocol v1.0, the Wave 01 design, exclusions, coding and verification policy, and the reasoning cap and price routing from Wave 02.
04 Sep 2026Dashboard figures recomputed from the CSV after two human-reviewer corrections (ACEP openai-C-2 0→1; RECL moonshot-B-1 1→2). Adherence values updated accordingly.
05 Sep 2026Automatic daily runs begin (4 per day, 9 models, alternating frame and language) under Wave 02 conditions. Not coded until a new version.

10 — Published data

What you need to redo it.

ola1_codigos.csv — the coded Wave 01 dataset (108 rows, ten fields, coder's note and review mark). stats.json — dashboard figures and per-model profiles, recomputed by script. activity.json — activity log with the time of every answer. eventos.json — the timeline events. ola1.py — the runner script (one OpenRouter key and one command). The exact prompts are in the protocol. The complete raw answers (JSONL) are published with the full wave; until then, an excerpt of every answer is readable on the map. License CC BY 4.0.

Shelter · Methodology v1.1 · 05 Sep 2026 · hello@theshelter.io
What they say under each frame. Never what they feel.