This tutorial will demonstrate how to download Daymet: Daily Surface Weather Data on a 1-km Grid for North America, Version 4 for a bounding box. Daymet dataset is available for North America (NA), including Hawaii (HI) and Puerto Rico (PR).
This tutorial will demonstrate how to directly access and subset Daymet dataset using NASA OPeNDAP Hyrax. Hyrax allows direct access to selected variables for the dataset within the spatial-temporal bounds without having to download the whole data files.
Requirements¶
import earthaccess
import geopandas as gpd
import xarray as xr
from pathlib import Path
from shapely.geometry import boxAccess to NASA OPeNDAP requires NASA Earthdata authentication. We recommend authenticating your Earthdata Login (EDL) information using the earthaccess python library as follows:
# works if the EDL login already been persisted to a netrc
try:
earthaccess.login(strategy="netrc")
except FileNotFoundError:
# ask for EDL credentials and persist them in a .netrc file
earthaccess.login(strategy="interactive", persist=True)Searching for Daymet Files¶
Daymet data files (or granules) are natively in netCDF4 format, and each file has one year’s worth of data. Data files are organized by variables (each for dayl, prcp, tmin, tmax, srad, swe, vp) and regions (each for US, PR, HI). For this tutorial, we will download data for the year 2024. We will first search for all the granules in the area of interest within the bounds of 170 W, 67.5N and 136.5W, 71.5N degrees.
# bound
BOUND = (-170, 67.5, -136.5, 71.5)
gdf = gpd.GeoDataFrame(geometry=[box(*BOUND)], crs="EPSG:4326")
# plot
gdf.explore(color="red", alpha=0.3)We will now search for all the Daymet granules overlapping the above bounding box using the Earthaccess Python library. We will also limit the period of our search to the year 2024 (Jan 1 to Dec 31, 2024). The earthaccess’s search_data allows searching dataset granules using NASA’s CMR API. It requires the dataset digital object identifier (DOI), time, and the bounding box.
# Daymet DOI
DOI = '10.3334/ORNLDAAC/2129'
# dates
DATES = ("2024-01-01", "2024-12-31")
# search data
granules = earthaccess.search_data(
doi=DOI,
bounding_box=BOUND,
temporal=DATES,
)
print(f"Total granules found: {len(granules)}")Total granules found: 7
Now, the day granule information, including data URLs, is stored in granules. The OPeNDAP URLs are in the RelatedUrls fields of the granule metadata. Let’s extract the OPeNDAP URLS.
# opendap urls
hyrax = [g._filter_related_links('USE SERVICE API')[0].replace("https", "dap4") for g in granules]
# print first 3
hyrax [:3]['dap4://opendap.earthdata.nasa.gov/collections/C2532426483-ORNL_CLOUD/granules/Daymet_Daily_V4R1.daymet_v4_daily_na_tmin_2024.nc',
'dap4://opendap.earthdata.nasa.gov/collections/C2532426483-ORNL_CLOUD/granules/Daymet_Daily_V4R1.daymet_v4_daily_na_vp_2024.nc',
'dap4://opendap.earthdata.nasa.gov/collections/C2532426483-ORNL_CLOUD/granules/Daymet_Daily_V4R1.daymet_v4_daily_na_prcp_2024.nc']Project the bounds to Daymet CRS¶
Daymet is natively in lambert_conformal_conic (LCC) coordinate reference system. We can now reproject the bounding coordinates of the above geopandas object (gdf) to bounds_lcc for requesting to OPeNDAP later.
#Daymet LCC proj
DAYMET_PROJ = (
"+proj=lcc +ellps=WGS84 +a=6378137 +b=6356752.314245 "
"+lat_1=25 +lat_2=60 +lon_0=-100 +lat_0=42.5 "
"+x_0=0 +y_0=0 +units=m +no_defs"
)
# convert to LCC proj and retrieve bounds
gdf_lcc = gdf.to_crs(DAYMET_PROJ)
bounds_lcc = gdf_lcc.total_bounds
# print the four corners in LCC
bounds_lccarray([-2823222.26105783, 3101471.12476104, -1401567.56219731,
4328450.66251338])Retrieve Daymet subset with OPeNDAP¶
Access to NASA OPeNDAP requires NASA Earthdata authentication. Please set up a new account if you do not already have a NASA Earthdata account. We will use earthaccess module to set up the authentication and return a session instance with an authorized bearer token.
session = earthaccess.get_requests_https_session()OPeNDAP allows access to the dataset of interest, so you can only download the variables you are interested in and, by doing so, save the download bandwidth.
Now, we will retrieve the dataset from the OPeNDAP Hyrax server. The following step will take time to run.
# loop over each URL
def subset_hyrax(url, b):
# open the dataset using xarray pydap engine
ds = xr.open_dataset(url,
session=session,
engine="pydap",
decode_coords="all")
# file basename without extension
f_name = Path(url).stem
# get daymet variables from filename
daymet_var = f_name.split("_")[-2]
# select area within the bounding coordinates
dummy=ds[daymet_var].sel(
x=slice(b[0], b[2]),
y=slice(b[3], b[1]),
)
# append each time step. This step is necessary so OPeNDAP
# requests are made one timestep at a time and it does not
# exceed the server's memory limit
subsets = []
for i in range(dummy.time.size):
subsets.append(dummy.isel(time=i))
# concatenate all the subsets data along time dimension
# NOTE: THIS STEP WILL TAKE TIME AS IT MAKES A TOTAL OF 365
# OPENDAP REQUESTS AND RETRIEVES SUBSET DATA
concat = xr.concat(subsets,
dim="time",
join='override',
coords='minimal',
compat='override')
# SAVE SUBSET DATA AS NETCDF FILE
f_sub = f"{f_name}_subset.nc"
concat.to_netcdf(f_sub)
return f_sub
# loop over the OPeNDAP urls and download the subset as netCDF file
daymet_f = [subset_hyrax(g, bounds_lcc) for g in hyrax]- Thornton, M. M., Shrestha, R., Wei, Y., Thornton, P. E., & Kao, S.-C. (2022). Daymet: Daily Surface Weather Data on a 1-km Grid for North America, Version 4 R1. ORNL Distributed Active Archive Center. 10.3334/ORNLDAAC/2129