Verification · graded nightly against observed reports
Every forecast, graded.
Clawd Front verifies its own output against observed tornado reports —
hits and busts alike. No forecast gets memory-holed. The numbers below are
produced by the same automated pipeline that issues the forecasts, updated after
each verification run.
30-day Brier score
—
Lower is better · 0 = perfect
Skill vs climatology
—
BSS > 0 beats the base rate
Tornado-day hit rate
—
≥10% issued when tornadoes occurred
False alarms (60d)
—
≥15% issued, zero tornadoes
Reliability
When we say N%, does it happen N% of the time? Points on the diagonal are perfectly calibrated.
Brier, last 30 runs
Daily grid-forecast error vs observed reports. Red ticks mark tornado days.
The Ledger
Day by day: what we said, what the atmosphere did, and the verdict.
“Correct quiet” days are listed too — most days, the right forecast is a low number.
Challenger gate
New models must beat the live model on identical days before promotion —
paired Brier, calibration, and spatial sanity, over ≥30 shadow days including ≥5 tornado days.
Method. Regional verdicts compare the issued CFnado max probability with observed
tornado LSRs for the valid period (12Z–12Z): HIT = ≥10% issued and tornadoes occurred ·
MISS = <5% issued and tornadoes occurred · FALSE ALARM = ≥15% issued, none occurred ·
CORRECT QUIET = low number, quiet day. Grid scores use the Brier score against
report-gridded observations; BSS is skill relative to regional climatology.
Ground truth: NWS Local Storm Reports / NCEI Storm Events. Forecasts are archived
at issue time and never edited.