Draft

Checking a forecast against reality

How accuracy is measured, and a simple routine to learn how a model behaves in your own waters.

The measures

Forecasters score models by comparing each forecast with what was observed. Four numbers cover most needs:

MeasureIn plain wordsTells you
ErrorForecast minus observation, for one timeWhether this one forecast was too high or too low
BiasThe average of the errors, with their signsWhich way the model leans: negative means it tends to under-forecast
Mean absolute errorThe average size of the errors, ignoring signHow far off it typically is
RMSELike the mean absolute error, but large misses count moreHow bad the worst misses are
SkillWhether the forecast beats a simple reference such as “tomorrow is like today”Whether the forecast adds value at all

A worked example

Five afternoons of 10 m wind at a harbour, forecast the day before and then measured:

DayForecast (kn)Observed (kn)Error (kn)
11214−2
21820−2
31516−1
42225−3
51011−1

Bias: −1.8 kn (the model leaned low by almost 2 kn). Mean absolute error: 1.8 kn. RMSE: 1.95 kn. Because every error has the same sign, this is a consistent bias you can allow for.

A simple routine

  1. Pick a reference. The nearest coastal station or buoy, a harbour or coast guard report, or your own instruments and logbook.
  2. Record the forecast before the event, with the model, the run time and the forecast hour. Write it down or take a screenshot: memory remembers misses and forgets hits.
  3. Record what happened at the same time and for the same quantity.
  4. Repeat for ten to twenty cases, and split them by situation: onshore and offshore winds, sea-breeze days, fronts.
  5. Look for a pattern. A consistent lean or a consistent timing error is something you can plan for.

Common pitfalls

  • Comparing different things. Mean wind against a gust, 10 m wind against an instrument at the masthead, or a model wind in a grid box against a station on a hill.
  • Station quality and exposure. A sensor in the lee of a building or on a headland does not measure what the model represents.
  • Too few cases. One bad forecast is not a bias.
  • Time errors. UTC against local time, or a value that means the previous hour.
  • Grid-box versus point. A fine model should be checked against a nearby point; a coarse one cannot be expected to match a single spot: What km and NM mean on a model.

Related: Why models disagree and How far ahead can you trust a forecast?