Published scoring
Forecast accuracy
Every forecast we issue is verified against ERA5 reanalysis and scored against the ECMWF IFS baseline. Wins and losses both appear here. A scoreboard that only shows wins is marketing, and nobody believes it.
Illustrative figures
The numbers on this page are worked examples showing the format this scoreboard will take. They are not measured results. Live verification goes in at Phase 3, at which point this page is generated from /v1/skill and updates automatically.
Gain against the IFS baseline, by variable
Positive means our ensemble beat the operational baseline. Negative means it did not. Precipitation is where AI ensembles currently struggle, and beyond day seven we lose.
2 m temperature
day 1–5
100 m wind (u)
day 1–5
100 m wind (v)
day 1–5
Surface solar radiation
day 1–5
MSLP
day 1–5
Total precipitation
day 1–5
Precipitation, day 7+
day 7–10
Forecast skill by lead time
Normalised skill, where 1.0 is a perfect forecast. Ensembling holds its advantage furthest out, which is exactly where the constituent models start to disagree with each other.
| Model | Day 1 | Day 2 | Day 3 | Day 4 | Day 5 | Day 7 | Day 10 |
|---|---|---|---|---|---|---|---|
| Ensemble Weather AI | 0.97 | 0.94 | 0.90 | 0.85 | 0.79 | 0.66 | 0.48 |
| GraphCast | 0.96 | 0.93 | 0.88 | 0.82 | 0.75 | 0.60 | 0.41 |
| Aurora | 0.96 | 0.93 | 0.89 | 0.83 | 0.77 | 0.63 | 0.45 |
| AIFS | 0.95 | 0.91 | 0.86 | 0.80 | 0.73 | 0.58 | 0.40 |
Methodology
- Ground truth
- ERA5
- Copernicus reanalysis, the reference dataset the field verifies against.
- Baseline
- ECMWF IFS
- The operational physics-based system every AI model is measured against.
- Metric
- RMSE / CRPS
- Root mean squared error for deterministic fields; CRPS for probabilistic ones.
- Grid
- 0.25°
- Global latitude/longitude grid, matching the resolution of the models scored.
Scores are computed on the same forecasts we serve to customers — not on a separate research configuration. The verification runs on a rolling thirty-day window and is recomputed daily as ERA5 becomes available.
Model weights are fitted per variable, region and lead time, because no single model wins everywhere. That is the entire reason an ensemble beats its members, and it is also why the weights change over time.
