Open research

Do people know how much they know?

Who's Bluffing is a game and an open study. Every completed round is an anonymous data point about calibration: whether people who say they are 80% sure are right 80% of the time, and whether a little feedback every day changes it.

The question

Decades of research say that people are overconfident: they feel surer than their answers turn out to be. Who's Bluffing measures this at scale, every day, with questions whose answers have a public source, and adds a second question: does playing regularly, with feedback after every answer, make people better calibrated?

Why honest confidence scores best

Each answer in a round scores 100 − 400 × (c − y)² points, where c is the confidence you chose (from 0.5 to 1) and y is 1 if you were right and 0 if not. That is the quadratic score Brier (1950) proposed for weather forecasts, rescaled to points. It is a strictly proper scoring rule: your expected points are highest when the confidence you state is your real chance of being right, so bluffing up or hedging down can only cost you on average (Gneiting and Raftery, 2007). A 50% answer scores 0 whether it is right or wrong, because a coin flip says nothing; 100% wins 100 points when right and loses 300 when wrong, so certainty pays only when you really are certain.

Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3. Gneiting, T., and Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378.

Hypotheses

Written down before launch and tested as written.

Numbers we publish

The public counts are computed once a day by a scheduled job and use these definitions, word for word from the pre-registration:

MAU: anonymous ids with ≥ 1 play — a completed round of 10 (ranked or quick), a completed full assessment, or an in-channel Slack or Discord answer to the daily question — in the trailing 30 days, summed across web, Slack and Discord without cross-surface deduplication (a person who plays on two surfaces counts twice; stated wherever MAU is reported). DAU likewise for one UTC day.

Communities: Slack workspaces, Discord servers and rooms with ≥ 1 play in the trailing 30 days, plus classrooms with ≥ 5 finished assessments; reported per platform and summed.

The current numbers are on the live stats page.

Open data

Anonymous row-level data, with every exclusion flag kept rather than deleted, will be released on OSF, Hugging Face and Kaggle with a data card describing how it was collected, its sample bias and what it should not be used for. Classroom sessions are pooled without class codes. The code for every analysis is in the analysis folder.

Pre-registration

The full pre-registration (design, measures, hypotheses, exclusions, sample sizes and inference) is prereg/PREREG.md on GitHub. It was frozen on 2026-10-04, before the public launch on Monday 2026-10-05; any change since is logged in the changelog with a date and a reason.

How to cite

Until the first dataset has its own identifier, please cite the project and the pre-registration:

Who's Bluffing (2026). Who's Bluffing? A daily calibration game and open dataset.
https://whosbluffing.com. Pre-registration: https://github.com/rongtnt/whos-bluffing/blob/main/prereg/PREREG.md