Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

METAR Comparison to Gridded Reanalysis Product (AORC)

Boise State University

Overview

This notebook explores how we might compare METAR products to gridded reanalysis data sets, like ERA-5, WRF, or AORC. We compare single a single site in a time series and look at the comparison in some interactive map-based visualizations.

For this example, we compare METAR with AORC data, a high-resolution reanalysis product produced by NOAA for the continental U.S. and Alaska from 1979-2023. This dataset has only 8 meteorological variables so is relatively simple to understand.

  1. Accessing METAR data

  2. Accessing AORC data

  3. Comparing time series

  4. Interactive map-based visualizations

Prerequisites

ConceptsImportanceNotes
PandasNecessaryFamiliarity with tabular data
S3 file accessHelpfulUnderstanding data access from a cloud-hosted bucket
  • Time to learn: 25 minutes


Imports

Time series comparison

Get METAR data

Data is hosted on Dynamical. We choose a single station, a time range, and select only the variables that we want to compare. Since this example is written by a snow scientist from Boulder, CO, our examples will focus on the winter time in the Colorado Rocky Mountains.

The METAR column names are weird, so we define some more descriptive names here. The full documentation for column names can be found here.

The Eagle-Vail airport is located just off i70 to the West of Denver. If you want help finding your nearest METAR station, the Local statistics notebook shows how to do this.

Source
Loading...
/home/runner/micromamba/envs/METAR_archive-cookbook/lib/python3.14/site-packages/geopandas/explore.py:383: UserWarning: CartoDB tiles now require an API key. Please provide one to continue using the tiles. You can request the key at https://carto.com/basemaps/apikey/.
  tiles = tiles.build_url(scale_factor="{r}")
Loading...

We’ll need different files if the time range spans multiple years.

['https://data.source.coop/dynamical/asos-parquet/year=2021/data.parquet']
Index(['valid', 'longitude', 'latitude', 'station', 'name', 'country', 'tmpc', 'relh', 'drct', 'sknt'], dtype='str')
Loading...

Get the AORC data for the target time and location.

AORC data has 8 variables:

  1. APCP_surface: Accumulated precipitation at surface (mm)

  2. TMP_2maboveground: Temperature at 2m above ground (K)

  3. SPFH_2maboveground: Specific humidity at 2m above ground (g/g)

  4. DLWRF_surface: Downward longwave radiation (W/m^2)

  5. DSWRF_surface: Downward shortwave radiation (W/m^2)

  6. PRES_surface: Pressure at surface (Pa)

  7. UGRD_10maboveground: East-West wind speed (m/s)

  8. VGRD_10maboveground: North-South wind speed (m/s)

We’ll use xarray and s3fs to open the cloud-hosted dataset. The advantage of this is that we can quickly access very large amounts of data. df_aorc has a full year of data over the Western U.S.

['s3://noaa-nws-aorc-v1-1-1km/2021.zarr']

AORC data is hosted in Zarr format by NOAA on a public S3 bucket.

Loading...
Single variable size: 287.93 GB

Before doing anything else, we want to get rid of all the extra data! We select the single grid cell that contains the target point and select out time range.

Single variable size: 0.0000033751 GB

In order to compare this to the METAR data, we’ll need to do some conversions. Since the METAR units are generally more understandable, we’ll convert to those.

VariableMETAR unitAORC unit
Temperaturedegrees Cdegrees K
HumidityRelative humiditySpecific humidity (g/g)
WindWind direction (degrees) and wind speed (knots)E-W/N-S wind speeds (m/s)

We run the conversions then ditch the extra columns.

Converting TMP_2maboveground in K to tmpc in C
Converting SPFH_2maboveground to relh (%)
Converting UGRD_10maboveground and VGRD_10maboveground to wind direction (deg)
Converting UGRD_10maboveground and VGRD_10maboveground to wind speed (knots)

Now that the units are harmonized, we need to deal with the timestamps. The METAR data has datetime\[us, UTC\] format (microsecond precision in UTC timezone), while the AORC uses datetime\[ns\] (nanosecond precision in local timezones). We’ll move everything to the local timezone.

Loading...

Compare!

Now that everything is matched up and in the same format, we can make some plots to compare the two datasets.

Here, we just plot each variable against the other.

Source
<Figure size 1200x800 with 4 Axes>

That’s a little messy. Let’s try plotting with a median filter.

A kernel size of 25 for the median filter means that the window size will be about 1 day.

Source
<Figure size 1200x800 with 4 Axes>

Now, lets plot just the different between each the values. The METAR data is subtracted from the AORC, so this difference is positive (blue) where the AORC is overestimating the METAR data and negative (red) where it is underestimating.

Source
<Figure size 1200x800 with 4 Axes>

What’s next?

What else might be interesting to compare?

  • Bias by time of day

  • Precipitation (not reported by all METAR stations)

Spatial comparison

We might want to look at whether performance changes over a wider range. For this, we’ll expand our area of interest to the full state of Colorado and make some interactive map visualizations.

Get AORC data for all of Colorado in the given time range. For space considerations, we are going to limit this down to just the month of January.

Single variable size: 0.2813384263 GB
Converting TMP_2maboveground in K to tmpc in C
Loading...
Single variable size: 0.2813384263 GB
Loading...

Get METAR data for all stations in Colorado

Loading...
Loading...
Loading...

Let’s look at where all of the METAR stations that we found are. The size of the bubble corresponds to how many observations each location has.

Source
Loading...

Click on a point for more details about any station.

Source
/home/runner/micromamba/envs/METAR_archive-cookbook/lib/python3.14/site-packages/geopandas/explore.py:383: UserWarning: CartoDB tiles now require an API key. Please provide one to continue using the tiles. You can request the key at https://carto.com/basemaps/apikey/.
  tiles = tiles.build_url(scale_factor="{r}")
Loading...

Next, we match the stations to their nearest AORC gridcell and create a dataframe that unifies this information.

Loading...

This dataframe can be aggregated by station to calculate the mean bias and mean root mean squared error (RMSE) throughout the month.

Loading...

Now, we average the AORC across the time axis so that we can plot it spatially.

Done averaging AORC
Source
Looks like you are using a tranform that doesn't support FancyArrowPatch, using ax.annotate instead. The arrows might strike through texts. Increasing shrinkA in arrowprops might help.
<Figure size 2000x1600 with 2 Axes>

This map shows the average modeled temperature for CO for Jan-Mar 2021, and the station biases and error for each station. The stations are colored based on the AORC bias and sized based on the mean RMSE over the whole time period.

Does bias change with elevation? Lets define some bins and see whether RMSE and bias are very different.

Source
<Figure size 1600x600 with 2 Axes>

Does model performance change with time of day?

Source
<Figure size 1600x600 with 2 Axes>

Summary and things to try

There are many other things we might want to look at to evaluate how well the AORC reanalysis data performs against these METAR observations. Hopefully this notebook starts to show you how to harmonize different datasets to compare them (timezones, datatypes, units, etc!).

What’s next?

In the next notebook, we will explore restructuring the METAR dataset into a more optimal analysis-ready, cloud optimized (ARCO) format.