Where the people come from
Every seat is a census record drawn from the market you declared: age, sex, income, household, education, occupation, region, device. To that record the engine attaches one real person: a respondent from Twin-2K-500, a study by Columbia Business School published in 2025 under a Creative Commons licence, in which 2,058 adults representative of the United States answered about five hundred questions across four waves. The respondent is matched to the seat on sex, age band, income band, household size and education, and up to a hundred of that person’s own answers ride with the seat when the model is asked. The paper describes the sample and the battery.
A written portrait sits under the record: a short account of who the person is, made once by the model to fit the facts and the answers, kept by seat, and shown as what it is. The portrait never overrides a fact or a survey answer.
What the field has measured
- Stanford’s generative agent study (Park and others, 2024) built agents for 1,052 real people from two-hour interviews. They predicted the people’s own survey answers at 85% of the accuracy the people themselves reached when re-taking the survey two weeks later. Agents built from demographics alone scored 14 to 15 points lower.
- The Twin-2K mega-study (2025) ran nineteen pre-registered studies on twins built from the same dataset this panel uses. Twins reached 72% accuracy on held-out questions, 88% of the test-retest ceiling, and the authors report that answers are shaped by the prompt and that models express opinions not representative of the population.
- Bisbee and colleagues (Political Analysis, 2024) found that model respondents match a survey’s averages while collapsing its variance and exaggerating certainty.
- Practitioners report sycophancy: model respondents praise concepts that real users go on to reject. MeasuringU’s review collects the experiments.
What we take from that
- Rankings, not levels. The research supports "this screen stops more people than that one" and "the higher price loses more". The same research does not support "34% would try it" as a level. The report, the department briefs and the money model lead with the rankings, and show the rates as bands after them.
- The record behind the seat is what carries accuracy. Demographics alone score poorly; a person’s own answers score well. So a seat reads up to a hundred of its real person’s answers, not eight.
- A grade on every run. Fourteen of each person’s answers are never shown to the model. On a sample of seats the model is asked those fourteen as the person, and the exact matches are counted against what the person actually said. The score is printed as a share of the human ceiling, the way the field reports it.
- A floor. Under sixty percent of that ceiling, the money model refuses the panel’s rates and its ledger says why. A panel that cannot impersonate its own real people should not set your funnel.
- A spread check. A panel whose answers collapse to one voice fails the run, because those numbers describe the model, not the market.
How a seat is built, and what it reads
The order is fixed. A real browser maps the product first: every screen it can reach, every control it may press, every form filled with test data, until nothing is left to try. The panel reads that map and the product’s pages as the browser digested them, not the pixels. Then every seat answers the same four questions: would you try it, what would stop you, would you pay at each price the page shows, and one sentence in your own words. When the map is there, each seat also names the screen where it would give up.
Every output is labelled simulated. Every share is a band on the effective sample size, weighted back to the market. The report prints what the run was estimated to cost before it ran beside what the meter billed.
What it is not
The panel is not a replacement for real people. Think of it as a first pass you can afford before them: where to look, which price to test, which screen to fix first. The panel’s own grade tells you how far to trust it on the day, and the console shows it beside every number.