Ensemble Weather AIGet API access

About

No single model wins everywhere

That sentence is the whole company. It is also a measurable fact about the current generation of AI weather models, and it is why an ensemble beats any of its members.

Why ensembling works

AI weather models now match or beat the physics-based systems they were trained to imitate. Published evaluations put Microsoft’s Aurora ahead of ECMWF’s operational IFS-HRES on roughly 92% of verification targets at 0.1° resolution, and ahead of every operational tropical cyclone system at one-to-five-day leads.

But the picture is not one model sweeping the board. AIFS beats GenCast and GraphCast on state-level precipitation. GraphCast leads at 0.25° out to five days. They win on different variables, different regions, different lead times and different weather regimes — and, crucially, they fail differently.

That is precisely the condition under which combining models produces a real improvement rather than an average of averages. When errors are uncorrelated, the combination is better than any input. When they are correlated, it is not. Knowing which situation you are in — and weighting accordingly — is the actual work.

Why we publish losses

Most weather vendors publish no verification at all. You are asked to trust a marketing claim about accuracy, with no baseline, no dataset and no way to check.

We publish continuously against ERA5 reanalysis and the ECMWF IFS baseline, and we show where we lose. Right now that is precipitation, and it gets worse beyond day seven. Putting that on a public page costs us something in the short term and is the only thing that makes the rest of the numbers believable.

A scoreboard that never shows a loss is read as marketing — and correctly so.

What we took from open science

The approach owes a debt to open forecasting competitions, and to Zeus (Subnet 18 on Bittensor) in particular, where forecasters are scored continuously against ERA5 with weighted RMSE adjusted for task difficulty. The discipline of that scoring — a named baseline, a public dataset, results that update whether or not you like them — is what we adopted.

What we did not adopt is the token mechanics. We sell an API to organisations that need an invoice, an SLA and someone to call at 3am.

What we will not claim

We forecast environments favourable to tornado formation. We do not forecast individual tornadoes days in advance, because nobody can — tornado-scale predictability is measured in hours.

We sell calibrated probability and quantified uncertainty, not certainty. That is what an ensemble actually provides, and it is what someone managing real risk actually wants to buy.

Many models. One forecast.

We are hiring the team that builds this — data infrastructure first.