Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Creating a More Cloud-Native METAR Format

CIRES University of Colorado Boulder

We will convert the dynamical Meteorological Aerodrome Report from Parquet to Xarray’s DataTree representation.

Overview

Within this notebook, we will create an interactive visualization of the latest METAR data across all stations. We will use the following libraries for our visualizations:

  1. Xarray

Prerequisites

ConceptsImportanceNotes
ParquetRequiredCloud Computing
ZarrRequiredCloud Computing
  • Time to learn: 10 minutes


Loading...
Loading...
Loading...
Loading...

We will just be working with the past year’s worth of data, so the parquet file can be found here.

Get the past three days of data

This query will request just the latest hours worth of data across all stations.

end time: 2026-10-01 05:28:02.008277
start time: 2026-09-28 05:28:02.008414

Query the database for data between timestamps requires SQL-like queries

Loading...

Print the first couple columns of data to get an understanding of the structure

Loading...

Understanding the data columns

The documentation can be found here for the individual columns:

https://github.com/dynamical-org/asos-parquet#key-fields

Print the columns as follows:

Index(['station', 'valid', 'longitude', 'latitude', 'tmpf', 'tmpc', 'dwpf', 'dwpc', 'relh', 'drct', 'sknt', 'gust', 'alti', 'mslp', 'vsby', 'p01i', 'p01m', 'state', 'geometry', 'name', 'elevation', 'country', 'county', 'wfo', 'tzname', 'bbox', 'year'], dtype='str')

Now create bind descriptions to each of the columns

Plot Timeseries data from a station local to Boulder, Colorado

Select all the data from the Boulder station, “BDU” and plot it.

Loading...
Loading...

Iterate through some variables and compose an interactive plot of the past three days

Loading...

METAR Parquet Antipatterns

There are a couple antipatterns that we are currently using in this cloud computing scheme:

  • partitioning --> the data is currently not accessed through a single catalog

    • we need to derive the start & end time to query specific files

  • compression --> similar data is not stored adjacent across rows of the parquet file

  • expressions --> data is not being accessed in a pythonic manner, we are instead relying on SQL-like queries

  • geospatial information is not changing, station attributes such as “lat/lon” are stored in a redundant manner

We would like to access the data using a more pythonic approach. The underlying data is timeseries, so the goal is to store it as timeseries DataArrays.

All of these problems can be fixed by Xarray’s DataTree implementation.

Loading...
Loading...
Latitude: 40.0394, Longitude: -105.2258, Elevation: 1612.0
Station: BDU, Country: US and time zone name: America/Denver

Create all of the DataArrays with associated labels and units

Create and Xarray Dataset

Combine all of the DataArray variables together and include attributes that don’t change for the DataSet such as the GPS location and the elevation.

Loading...

Create an Xarray DataTree

The solution to having disperate datasets will be to join them together using Xarray’s DataTree implementation. It is a graph like structure, so it uses nodes and edges to connect dataset in a hierarchical like manner.

Here is an example of an Xarray DataTree with a family:

<xarray.DataTree 'Homer'>
Group: /
├── Group: /Bart
└── Group: /Lisa

To access a child I just denote them by name:

Loading...

First example: Define the METAR station for Boulder, ‘BDU’ and add it as a child to the parent

So “metar” now has the child dataset, “bdu”

And note that there are attributes specific to that station tracked with that

Frozen({'bdu': <xarray.DataTree 'bdu'> Group: /bdu Dimensions: (time: 206) Coordinates: * time (time) datetime64[us] 2kB 2026-09-28T05:35:00 ... 2026... Data variables: temperature (time) float64 2kB 60.8 60.8 60.8 62.6 ... 59.0 59.0 59.0 dew_point (time) float64 2kB 50.0 50.0 50.0 46.4 ... 41.0 42.8 42.8 relative_humidity (time) float64 2kB 67.55 67.55 67.55 ... 54.86 54.86 wind_direction (time) float64 2kB 220.0 0.0 260.0 0.0 ... 0.0 0.0 250.0 wind_speed (time) float64 2kB 4.0 0.0 3.0 0.0 ... 0.0 0.0 0.0 3.0 wind_gust (time) float64 2kB nan nan nan nan ... nan nan nan nan pressure (time) float64 2kB 30.1 30.1 30.09 ... 30.06 30.05 30.05 visibility (time) float64 2kB 10.0 10.0 10.0 10.0 ... 10.0 10.0 10.0 precipitation (time) float64 2kB 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 Attributes: station: BDU country: US timezone: America/Denver latitude: 40.0394 longitude: -105.2258 elevation: 1612.0})

We can store the data in a single cloud-native dataset now, no need to partition because the data is automatically chunked

The dataset can be saved as follows:

<xarray.backends.zarr.ZarrStore at 0x7f2336ebf1c0>

Create a datatree of some Colorado stations

Pick a couple example stations to save:

Iterate through all selected stations and add them to the parent

3

Create Consolidated DataTree of the Station Data

Frozen({'BDU': <xarray.DataTree 'BDU'> Group: /BDU Dimensions: (time: 206) Coordinates: * time (time) datetime64[us] 2kB 2026-09-28T05:35:00 ... 2026... Data variables: temperature (time) float64 2kB 60.8 60.8 60.8 62.6 ... 59.0 59.0 59.0 dew_point (time) float64 2kB 50.0 50.0 50.0 46.4 ... 41.0 42.8 42.8 relative_humidity (time) float64 2kB 67.55 67.55 67.55 ... 54.86 54.86 wind_direction (time) float64 2kB 220.0 0.0 260.0 0.0 ... 0.0 0.0 250.0 wind_speed (time) float64 2kB 4.0 0.0 3.0 0.0 ... 0.0 0.0 0.0 3.0 wind_gust (time) float64 2kB nan nan nan nan ... nan nan nan nan pressure (time) float64 2kB 30.1 30.1 30.09 ... 30.06 30.05 30.05 visibility (time) float64 2kB 10.0 10.0 10.0 10.0 ... 10.0 10.0 10.0 precipitation (time) float64 2kB 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 Attributes: station: BDU country: US timezone: America/Denver latitude: 40.0394 longitude: -105.2258 elevation: 1612.0, 'DEN': <xarray.DataTree 'DEN'> Group: /DEN Dimensions: (time: 73) Coordinates: * time (time) datetime64[us] 584B 2026-09-28T05:53:00 ... 202... Data variables: temperature (time) float64 584B 64.0 65.0 65.0 ... 59.0 59.0 56.0 dew_point (time) float64 584B 47.0 49.0 50.0 ... 48.0 46.0 46.0 relative_humidity (time) float64 584B 53.94 56.15 58.28 ... 61.99 69.05 wind_direction (time) float64 584B 180.0 190.0 180.0 ... 0.0 80.0 130.0 wind_speed (time) float64 584B 7.0 12.0 15.0 15.0 ... 0.0 4.0 4.0 wind_gust (time) float64 584B nan nan nan nan ... nan nan nan nan pressure (time) float64 584B 30.11 30.1 30.1 ... 29.99 30.01 30.04 visibility (time) float64 584B 10.0 10.0 10.0 ... 10.0 10.0 10.0 precipitation (time) float64 584B 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 Attributes: station: DEN country: US timezone: America/Denver latitude: 39.8328 longitude: -104.6575 elevation: 1656.0, 'DRO': <xarray.DataTree 'DRO'> Group: /DRO Dimensions: (time: 137) Coordinates: * time (time) datetime64[us] 1kB 2026-09-28T05:53:00 ... 2026... Data variables: temperature (time) float64 1kB 59.0 58.0 59.0 60.0 ... 51.0 50.0 50.0 dew_point (time) float64 1kB 48.0 48.0 46.0 43.0 ... 50.0 50.0 50.0 relative_humidity (time) float64 1kB 66.84 69.28 61.99 ... 100.0 100.0 wind_direction (time) float64 1kB 0.0 70.0 90.0 nan ... 0.0 0.0 0.0 wind_speed (time) float64 1kB 0.0 6.0 4.0 3.0 ... 3.0 0.0 0.0 0.0 wind_gust (time) float64 1kB nan nan nan nan ... nan nan nan nan pressure (time) float64 1kB 30.21 30.21 30.22 ... 30.07 30.07 visibility (time) float64 1kB 10.0 10.0 10.0 10.0 ... 0.5 0.25 0.25 precipitation (time) float64 1kB 0.0 0.0 0.0 0.0001 ... 0.0 0.0 0.0 0.0 Attributes: station: DRO country: US timezone: America/Denver latitude: 37.1515 longitude: -107.7538 elevation: 2038.0})

Write all of the data as a single dataset

<xarray.backends.zarr.ZarrStore at 0x7f2336ebd870>

Here is how you would open the data

Loading...

To get a slice of 1 hours data you could do something like:

Loading...

And then to get the temperature data:

Loading...

And the relative humidity:

Loading...

References

Future work

Using the above approach, a next step could be to expand the DataTree implementation to the full dataset.

What’s next?

In the next and final notebook, we will translate METAR observations into plain-language, actionable weather guidance for non-expert users, as a first step towards developing an AI agent.