Verification · graded nightly against observed reports

Every forecast,
graded.

Clawd Front verifies its own output against observed tornado reports — hits and busts alike. No forecast gets memory-holed. The numbers below are produced by the same automated pipeline that issues the forecasts, updated after each verification run.

30-day Brier score
Lower is better · 0 = perfect
Skill vs climatology
BSS > 0 beats the base rate
Tornado-day hit rate
≥10% issued when tornadoes occurred
False alarms (60d)
≥15% issued, zero tornadoes

Reliability

When we say N%, does it happen N% of the time? Points on the diagonal are perfectly calibrated.

Brier, last 30 runs

Daily grid-forecast error vs observed reports. Red ticks mark tornado days.

The Ledger

Day by day: what we said, what the atmosphere did, and the verdict. “Correct quiet” days are listed too — most days, the right forecast is a low number.

Challenger gate

New models must beat the live model on identical days before promotion — paired Brier, calibration, and spatial sanity, over ≥30 shadow days including ≥5 tornado days.

Method. Regional verdicts compare the issued CFnado max probability with observed tornado LSRs for the valid period (12Z–12Z): HIT = ≥10% issued and tornadoes occurred · MISS = <5% issued and tornadoes occurred · FALSE ALARM = ≥15% issued, none occurred · CORRECT QUIET = low number, quiet day. Grid scores use the Brier score against report-gridded observations; BSS is skill relative to regional climatology. Ground truth: NWS Local Storm Reports / NCEI Storm Events. Forecasts are archived at issue time and never edited.