Loading today's questions…

What this game trains

Most quizzes score what you know. How Sure? scores whether you know how well you know it — a skill called calibration. A person is well calibrated when the things they call “80% sure” turn out true about 80% of the time. Decades of forecasting research say two useful things about it. First, calibration is measurable: state a probability, record the outcome, and the gap between stated confidence and realized accuracy is a number, not an impression. The scoring rule that makes this honest dates to Glenn W. Brier's 1950 paper Verification of forecasts expressed in terms of probability (Monthly Weather Review, vol. 78, no. 1), written for weather forecasters: penalize the squared gap between stated probability and what happened. Under a quadratic rule of this kind, your expected score is highest when you report exactly what you believe — overclaiming and hedging both cost you, which is what makes it a proper scoring rule rather than a bravado contest. Second, calibration responds to practice: forecasters who state probabilities and get prompt, scored feedback tend to become better calibrated at that task than those who never see their record. That feedback loop is the whole design here — eight scored judgments a day, and a running profile that tells you what “80% sure” has actually meant in your hands. The facts are the furniture; the slider is the game.

The same discipline runs through the site's word games: Solver Says and Probe score decisions against a measured best, and Last Row removes luck from deduction entirely. How Sure? is the same idea aimed at what you know rather than what you spell.

How scoring works

Every round pays out under a quadratic (Brier-style) proper scoring rule: points = 12.5 × ((1 − (c/100 − o)²) − 0.0199) / 0.98, where c is your stated confidence and o is 1 if you were right, 0 if not. In words: your stated confidence c becomes a probability, the squared miss between it and the outcome is your penalty, and the result is rescaled so the worst possible round (99% and wrong) scores 0 and the best (99% and right) scores 12.5. Because the rescaling is a straight linear map, the rule stays proper: reporting exactly what you believe maximizes your expected points, so the slider has no bluffing strategy.

At 50% the payoff is the same whether you are right or wrong — about 9.3 points — so saying “I don't know” is always safe. A full day of 50% answers scores exactly 74.5: that is the shrug line, and any total above it means your knowledge plus your honesty added real value.

Daily total bands and what they mean
Daily totalBandWhat it means
90+ Sharp and honest You were right on nearly every round and said so. This band is unreachable on knowledge alone — it also needs the slider to match it.
80+ Calibrated Your confidence tracked your accuracy: high stakes mostly landed, misses mostly came on cautious rounds.
75+ Above the shrug line Setting every slider to 50% scores exactly 74.5. You beat that, so your knowledge plus honesty added value — barely.
60+ The slider cost you You would have scored more by admitting 50% everywhere. The points went to confident misses — stake high only when you would be surprised to be wrong.
0+ Confidently wrong Several high-confidence misses. That is the exact habit this game trains away: the fix is lowering the slider, not knowing more facts.

Method notes. Scoring: each round is worth 12.5 points under the formula above; the daily total sums eight rounds (0–100). The share strip encodes each round with the thresholds cautious < 75% / confident ≥ 75%: 🟩 confident and right, 🟨 cautious and right, 🟧 cautious and wrong, 🟥 confident and wrong — your picks and the questions themselves are never included. Data: every question compares two entities of the same class using values from Wikidata (available under CC0 1.0), fetched 2026-08-09, keeping only entities with at least 25 Wikipedia language editions and showing an “as of” year wherever the figure is dated (populations and GDP are skipped entirely without one). Pairs keep a value ratio between 1.05× and 20.0×, and difficulty follows the Weber fraction: people discriminate magnitudes by ratio, so values close together make hard questions and far-apart values make easy ones. Each day ramps easy to hard with at most two questions per class. These are Wikidata's recorded figures, not fresh surveys — an “as of” date tells you how old each one is.