Checking a forecast against reality
How accuracy is measured, and a simple routine to learn how a model behaves in your own waters.
The measures
Forecasters score models by comparing each forecast with what was observed. Four numbers cover most needs:
| Measure | In plain words | Tells you |
|---|---|---|
| Error | Forecast minus observation, for one time | Whether this one forecast was too high or too low |
| Bias | The average of the errors, with their signs | Which way the model leans: negative means it tends to under-forecast |
| Mean absolute error | The average size of the errors, ignoring sign | How far off it typically is |
| RMSE | Like the mean absolute error, but large misses count more | How bad the worst misses are |
| Skill | Whether the forecast beats a simple reference such as “tomorrow is like today” | Whether the forecast adds value at all |
A worked example
Five afternoons of 10 m wind at a harbour, forecast the day before and then measured:
| Day | Forecast (kn) | Observed (kn) | Error (kn) |
|---|---|---|---|
| 1 | 12 | 14 | −2 |
| 2 | 18 | 20 | −2 |
| 3 | 15 | 16 | −1 |
| 4 | 22 | 25 | −3 |
| 5 | 10 | 11 | −1 |
Bias: −1.8 kn (the model leaned low by almost 2 kn). Mean absolute error: 1.8 kn. RMSE: 1.95 kn. Because every error has the same sign, this is a consistent bias you can allow for.
A simple routine
- Pick a reference. The nearest coastal station or buoy, a harbour or coast guard report, or your own instruments and logbook.
- Record the forecast before the event, with the model, the run time and the forecast hour. Write it down or take a screenshot: memory remembers misses and forgets hits.
- Record what happened at the same time and for the same quantity.
- Repeat for ten to twenty cases, and split them by situation: onshore and offshore winds, sea-breeze days, fronts.
- Look for a pattern. A consistent lean or a consistent timing error is something you can plan for.
Common pitfalls
- Comparing different things. Mean wind against a gust, 10 m wind against an instrument at the masthead, or a model wind in a grid box against a station on a hill.
- Station quality and exposure. A sensor in the lee of a building or on a headland does not measure what the model represents.
- Too few cases. One bad forecast is not a bias.
- Time errors. UTC against local time, or a value that means the previous hour.
- Grid-box versus point. A fine model should be checked against a nearby point; a coarse one cannot be expected to match a single spot: What km and NM mean on a model.
Related: Why models disagree and How far ahead can you trust a forecast?