# What makes a Wordle hard? — methods (v1)

Published 2026-08-30; source research release `v2026-07-15`, completed-answer data through 2026-07-14.

## Version-frozen analysis method

The analysis unit is one first recorded appearance per distinct answer word. Repeat appearances are excluded so an identical word-level observation is not counted twice. The primary outcome is the number of turns in the release's pinned example solver route. This is a reproducible model outcome, not a representative human score. Solver-pressure percentile is retained only as a cross-reference because it is a composite built partly from the predictors; it is not treated as independent validation.

Four mechanical predictors were selected before inspecting results: log2 mean candidate count after CRANE, SLATE and ADIEU; log2 size of the answer's largest one-position family; repeated-letter count; and mean smoothed letter-position surprise in bits. Position probabilities use the pinned 2,354-word simulation pool with 0.5 additive smoothing. A secondary complete-case model adds the Glasgow Norms familiarity rating. Missing familiarity is left missing, never coded as low familiarity or imputed.

Familiarity source: Scott, Keitel, Becirspahic, Yao & Sereno (2019), “The Glasgow Norms: Ratings of 5,500 words on nine scales”, Behavior Research Methods 51:1258-1270, DOI [10.3758/s13428-018-1099-3](https://doi.org/10.3758/s13428-018-1099-3). The source file `13428_2018_1099_MOESM2_ESM.csv` (SHA-256 `30d7776eed9219767f8a6aabdf95d7120cd050ddeab431acef13d435807bbc39`) was retrieved 2026-08-09 from https://www.ebi.ac.uk/europepmc/webservices/rest/PMC6538586/supplementaryFiles. The ratings are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This release uses an adaptation: Familiarity ratings were rescaled from 1–7 to 0–100, rounded, and sense-disambiguated entries were merged with an N-weighted mean; ratings absent from the source were omitted, never inferred.

The primary model is ordinary least squares with the outcome and predictors standardized within its sample. Coefficients are changes in solver-turn standard deviations per one predictor standard deviation, holding the other listed predictors fixed. Percentile 95% intervals use 1,000 seeded row bootstraps. Separate Spearman associations use seeded bootstrap intervals and 1,999 permutation tests; the five p-values are adjusted together with Benjamini-Hochberg false-discovery-rate control. The specified group contrasts report raw mean-turn differences and bootstrap intervals. These are descriptive associations, not causal effects.

No immutable, threshold-cleared production Observatory snapshot was available to this release. Human outcomes are therefore not analysed. The study does not substitute local, self-selected or below-threshold submissions.

## Missing-data and reuse rules

All 1,834 primary observations have the four mechanical predictors and a solver route. Familiarity covers 732 of them. The public analysis table omits answer words but retains completed dates, puzzle numbers, inputs and outcomes. Derived tables, documentation and graphics are CC BY 4.0. The preserved simulation answer pool in the upstream research release remains under its recorded source terms. Glasgow-derived familiarity values are adapted CC BY 4.0 data and must retain attribution.
