Loading today's questions…

Why calibration matters

How Sure? helps you practice calibration: how closely your confidence matches your accuracy. If you call many answers “80% sure,” about 80% of them should be right. You can measure that by recording each probability and result. The scoring rule comes from Glenn W. Brier's 1950 paper Verification of forecasts expressed in terms of probability (Monthly Weather Review, vol. 78, no. 1). It penalizes the squared gap between your stated probability and the result. Your expected score is highest when you report what you actually believe, which makes this a proper scoring rule. Regular, prompt feedback can improve calibration, so the game gives you five scored judgments a day and a running profile of your stated confidence versus your results.

These games measure related skills: Word Tactics scores Best Move, Beat the Solver and Split the Pack decisions against a measured best, and Last Row tests endgame deduction. How Sure? tests calibration on factual questions.

Scoring and data

How confidence scoring works

The headline result is your confidence score out of 100. A score of 74.5 is the baseline for choosing 50% on every question. Each round uses this Brier-style formula: points = 20 × ((1 − (c/100 − o)²) − 0.0199) / 0.98, where c is your stated confidence and o is 1 if you were right, 0 if not. Your confidence becomes a probability, and the squared difference between that probability and the result is the penalty. The underlying scale remains compatible with earlier results up to 99%; normal play now offers four clearer choices: 55%, 65%, 80% and 95%. This is a proper scoring rule: honest confidence has the highest expected score.

At 50%, you earn about 14.9 points whether you are right or wrong. Five answers at 50% total exactly 74.5 points. On average, you score most by reporting what you actually believe.

Daily total bands and what they mean
Daily totalProfileWhat it means
90+ Strong confidence score this run Most answers were correct, and your confidence matched the results.
80+ Confidence helped this run Most high-confidence answers were correct.
75+ Above the 50% baseline You scored above the 74.5 points earned by choosing 50% every round.
60+ Below the 50% baseline this run Several confident wrong answers lowered your score.
0+ Confident mistakes were costly this run Several high-confidence answers were wrong. Use lower percentages when the evidence is weak.
Where the questions and figures come from

Method notes. Scoring: each round is worth 20.0 points under the formula above; the daily total sums five rounds (0–100). The share strip encodes each round with the thresholds cautious < 75% / confident ≥ 75%: 🟩 confident and right, 🟨 cautious and right, 🟧 cautious and wrong, 🟥 confident and wrong. Your picks and the questions themselves are never included. Data: every question compares two entities of the same class using values from Wikidata (available under CC0 1.0), fetched 2026-08-09, keeping only entities with at least 25 Wikipedia language editions and showing an “as of” year wherever the figure is dated (populations and GDP are skipped entirely without one). Pairs keep a value ratio between 1.15× and 20.0×, and difficulty follows the Weber fraction: people discriminate magnitudes by ratio, so values close together make hard questions and far-apart values make easy ones. Each day ramps easy to hard with at most two questions per class. These are Wikidata's recorded figures, not fresh surveys. An “as of” date tells you how old each one is.