Trust

How accurate the panel is

The panel grades itself on every run, against the real people behind its seats. The figure is on the report and on the run screen. This page says what it means. Our own number goes here once the first keyed runs exist, with the run it came from. Nothing is quoted before it is measured.

The grade

Behind every seat is a real survey respondent. Fourteen of that person’s answers are held out: the model never sees them. On a sample of seats each run, the model is asked those fourteen questions as the person, from the record and the answers it was shown, and each answer is scored as an exact match or not against what the person actually said. The share of matches is the panel’s accuracy on that run.

The ceiling

People do not agree with themselves perfectly. Asked the same question twice a week apart, they change their answer some of the time. That re-test agreement is the ceiling any twin can reach, and the field reports twin accuracy as a share of it. Until the shelf carries the study’s fourth wave, the ceiling is the published figure from the Twin-2K mega-study: twins at 72% accuracy were 88% of test-retest, which puts the ceiling near 82%. Once the fourth wave is on the shelf the ceiling is measured from it, per question, and the report says which it used.

The floor

Under sixty percent of the ceiling, the Lab does not use the panel’s rates. Its ledger says so, with the score. Every department brief carries the grade beside its source, and Strategy prints it in words. A panel that was never graded is used as before, and its rows say it was not graded.

Our figure

Not measured yet on a live model. The engine is built and guarded, and every run with a model key prints its grade. The first keyed runs will put the number here, with the model, the seats, the date and the report it came from, and this page will be updated the same day. Until then the only accuracy figures on this page are the field’s, linked above.

How to read a grade

  • A grade is about the run in front of you: the model, the seats, the day. Read it next to the numbers, not as a property of the product.
  • A high grade says the model impersonated its people well on questions it had not seen. A high grade does not say the people are your market; the declared market and the census frame decide that, and the report names both.
  • Rankings survive a middling grade better than levels do. Where people stop and which price loses more are the outputs to act on first.