Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Access ERA5 data from NCAR’s Geoscience Data Exchange (GDEX) and compute anomaly

Overview

ERA5 data is available from NCAR’s GDEX in Analysis Ready Cloud Optimized (ARCO) zarr or virtual zarr formats (kerchunk). To learn how to create the annual zarr store used in this notebook, please see this notebook.

In this notebook,

  • We will read data zarr stores from NCAR’s GDEX endpoint

  • Compute temperature anomaly for the years 1940-2023

Prerequisites

ConceptsImportanceNotes
Intro to XarrayNecessary
Intro to IntakeNecessary
Understanding of ZarrHelpful
  • Time to learn: 30 minutes

Imports

Specify global variables

Create a Dask cluster

Dask Introduction

Dask is a solution that enables the scaling of Python libraries. It mimics popular scientific libraries such as numpy, pandas, and xarray that enables an easier path to parallel processing without having to refactor code.

There are 3 components to parallel processing with Dask: the client, the scheduler, and the workers.

The Client is best envisioned as the application that sends information to the Dask cluster. In Python applications this is handled when the client is defined with client = Client(CLUSTER_TYPE). A Dask cluster comprises of a single scheduler that manages the execution of tasks on workers. The CLUSTER_TYPE can be defined in a number of different ways.

  • There is LocalCluster, a cluster running on the same hardware as the application and sharing the available resources, directly in Python with dask.distributed.

  • In certain JupyterHubs Dask Gateway may be available and a dedicated dask cluster with its own resources can be created dynamically with dask.gateway.

  • On HPC systems dask_jobqueue is used to connect to the HPC Slurm and PBS job schedulers to provision resources.

The dask.distributed client python module can also be used to connect to existing clusters. A Dask Scheduler and Workers can be deployed in containers, or on Kubernetes, without using a Python function to create a dask cluster. The dask.distributed Client is configured to connect to the scheduler either by container name, or by the Kubernetes service name.

Select the Dask cluster type

The default will be LocalCluster as that can run on any system.

If running on a HPC computer with a PBS Scheduler, set to True. Otherwise, set to False.

If running on Jupyter server with Dask Gateway configured, set to True. Otherwise, set to False.

Python function for a PBS Cluster

Python function for a Gateway Cluster

Python function for a Local Cluster

Python logic to select the Dask Cluster type

This uses True/False boolean logic based on the variables set in the previous cells

/home/runner/micromamba/envs/ERA5_interactive/lib/python3.14/site-packages/distributed/node.py:195: UserWarning: Port 8787 is already in use.
Perhaps you already have a cluster running?
Hosting the HTTP server on port 35273 instead
  warnings.warn(
Loading...

Open ERA5 annual means file

Fetching long content....

Save anomaly

  • Plot the temperature anomaly

Close the Dask Cluster

It’s best practice to close the Dask cluster when it’s no longer needed to free up the compute resources used.