Forecast skill.Measured in the open.
A shared protocol for model comparisons, prospective records, and reproducible evidence.
Observation coverage3 airport stationsNew York · London · Tokyo
Public baselines1 AI · 2 numerical modelsAIFS / IFS / GFS
Public competitionWeekly model leagueExplore standings and receipts ↗
Model performance
The same valid samples. A transparent comparison.
Matching forecasts and observations
The page will update when the data is ready.
Swipe the table to compare all measures.
A seven-day replay tests the evaluation method. It does not count toward a prospective track record or establish long-term skill.
Error over time
Daily MAE · All matched stations · °C
AIFSIFSGFS
Matched samples will appear here.