Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Exploring GEDI L4A data structure

This tutorial will explore data structure, variables, and quality flags of the Global Ecosystem Dynamics Investigation (GEDI) L4A Footprint Level Aboveground Biomass Density (AGBD)Dubayah et al. (2022) dataset. GEDI L4A dataset is currently available for the period starting 2019-04-17 and covers 52 North to 52 South latitudes.

Details about the GEDI L4A dataset are available in the User Guide and the Algorithm Theoretical Basis Document (ATBD)] Kellner et al. (2022). Please refer to other tutorials for searching and downloading the GEDI L4A dataset and subsetting GEDI L4A footprints by area of interest.

Requirements

Additional prerequisites are provided here.

2. Example file

In this tutorial, we will use one GEDI L4A file, GEDI04_A_2020207182449_O09168_03_T03028_02_002_02_V002.h5. Let’s first search the file using the earthaccess module. Please refer to the previous tutorial for more details on downloading programmatically.

Loading...

The file is ~359 MB in size. We will go ahead and use earthaccess python module to get the file directly and save it to the full_orbits directory.

Loading...
Loading...
Loading...

GEDI footprints data, including the Level 4A, are natively in HDF5 format, and each file represents one of the four International Space Station (ISS). The file naming convention of the GEDI L4A is as GEDI04_A_YYYYDDDHHMMSS_O[orbit_number]_[granule_number]_T[track_number]_[PPDS_type]_ [release_number]_[production_version]_V[version_number].h5:

  • GEDI04_A: product short name

  • YYYYDDDHHMMSS: date and time of acquisition in Julian day of year, hours, minutes, and seconds format

  • [orbit_number]: orbit number

  • [granule_number]: sub-orbit granule (or file) number

  • [track_number]: track number

  • [PPDS_type]: positioning and pointing determination system (PPDS) type (00: predict, 01: rapid, > 02: final)

  • [release_number]: release number 002, representing the SOC SDS (software) release

  • [production_version]: granule production version

  • [version_number]: dataset version

The GEDI L4A file GEDI04_A_2020207182449_O09168_03_T03028_02_002_02_V002.h5 represents orbit number 09168, sub-orbit 03, and track number 03028 for the starting date-time in YYYYDDDHHMMSS format 2020207182449 or 2020-07-25 18:24:49 UTC. The file has the PPDS type of 02, which means it is final quality data. The file has a granule production version of 02 and Version 002 (Release 002).

3. Root Groups

Let’s open this file using the h5py python library and print the root-level groups of the file.

['ANCILLARY', 'BEAM0000', 'BEAM0001', 'BEAM0010', 'BEAM0011', 'BEAM0101', 'BEAM0110', 'BEAM1000', 'BEAM1011', 'METADATA']

All science variables are organized by eight beams of the GEDI. Please refer to the GEDI L4A user guide and GEDI L4A data dictionary for details. The root-level group also includes ANCILLARY and METADATA groups.

3. METADATA group

The METADATA group contains information about the file metadata. All the metadata attributes are provided in the DatasetIdentification group within METADATA. Let’s print its content.

Loading...

4. ANCILLARY group

The ANCILLARY group has three sub-groups - model_data, pft_lut, and region_lut.

['model_data', 'pft_lut', 'region_lut']

The model_data subgroup contains information about parameters and variables from the L4A models used to generate predictions. The details about model_data variables are available in ‘Table 6’ of the user guide. Let’s print the model_data variables.

dtype([('predict_stratum', 'O'), ('model_group', 'u1'), ('model_name', 'O'), ('model_id', 'u1'), ('x_transform', 'O'), ('y_transform', 'O'), ('bias_correction_name', 'O'), ('fit_stratum', 'O'), ('rh_index', 'u1', (8,)), ('predictor_id', 'u1', (8,)), ('predictor_max_value', '<f4', (8,)), ('vcov', '<f8', (5, 5)), ('par', '<f8', (5,)), ('rse', '<f4'), ('dof', '<u4'), ('response_max_value', '<f4'), ('bias_correction_value', '<f4'), ('npar', 'u1')])

Some model_data parameters have more than one dimension (e.g., vcov, which is the variance-covariance matrix of model parameters, has a dimension of 5 x 5). First, let’s read all one-dimensional variables into a pandas dataframe.

Loading...

The predict_stratum names are divided into two parts - the first part corresponding to PFTs and the second part to world regions.

PFTs (corresponding MCD12Q1 V006 classes are in parenthesis)
DBT: deciduous broadleaf trees (Class 4), DNT: deciduous needleleaf trees (Class 3), EBT: evergreen broadleaf trees (Class 2), ENT: evergreen needleleaf trees (Class 1), GSW: grasses, shrubs and woodlands (Classes 5, 6, 11)

World Regions
Af: Africa, Au: Australia and Oceania, Eu: Europe, NAm : North America, NAs: North Asia, SA: South America, SAs : South Asia

The information on PFTs and World Regions are also provided in pft_lut and region_lut groups within the ANCILLARY group.

Loading...
Loading...

In Table 2 above, the model_group column represents model group:

Model Group
1: all predictors considered, 2: no RH metrics below RH50, 3: forced inclusion of RH98, 4: forced inclusion of RH98 and no RH metrics below RH50

x_transform and y_transform are predictor and response transform functions (e.g., sqrt, log). The column bias_correction_name specifies the bias correction methods used - Snowdon (1991) or Baskerville (1972). For details, please refer to Section 4.3 of the ATBD.

To retrieve values from multidimensional variables, we can directly slice using the index from pandas dataframe model_data_df. For example, to print the variance-covariance matrix of the DBT_NAm prediction stratum (deciduous broadleaf trees North America, index=6), we can use:

array([[ 4.19623613, -0.30473042, -0.09762442, 0. , 0. ], [-0.30473042, 0.05645894, -0.02532849, 0. , 0. ], [-0.09762442, -0.02532849, 0.03291459, 0. , 0. ], [ 0. , 0. , 0. , 0. , 0. ], [ 0. , 0. , 0. , 0. , 0. ]])

Let’s also print out the variables predictor_id, rh_index, and par for the DBT_NAm prediction stratum.

predictor_id: [1 2 0 0 0 0 0 0]
rh_index: [50 98  0  0  0  0  0  0]
par: [-120.77709198    5.50771856    6.80801821    0.            0.        ]

The predictor_id variable gives the identifier of the predictor variable. The same consecutive predictor values indicate that the product of two metrics was used in the linear model. The variable rh_index gives the index of the relative height (RH) metrics used as predictors. For DBT_NAm stratum, two predictors RH50 and RH98 were used. The variable par gives model coefficients - the first element of the list is always the intercept of the linear model.

We can use the values from these three variables (predictor_id, rh_index, and par) to construct the linear AGBD model for each prediction strata. Refer to GEDI_L4A_Common_Queries for more information on how to assign coefficients and predictor variables to the linear models.

Loading...

As we see in the table above, some predict strata share the same AGBD model. There are a total of 12 AGBD models representing 35 prediction strata.

5. BEAM groups

GEDI instrument has two full-power lasers, each producing two beams and one coverage laser producing four beams, generating a total of 8 beams. GEDI L4A data are grouped by these beams (group name starting with BEAMXXXX). Let’s print the name and description of BEAM groups. Each of the beams is assigned beam and channel numbers.

Loading...

Let’s look into the data structure of one of the beams, BEAM0110, a full-powered beam. All of the 8 beams have the same set of variables and variable groups.

Loading...

‘L2A’ in the ‘source’ column indicates that the variable originates from the GEDI Level 2A dataset, which provides footprint level elevation and canopy height metrics. Let’s look into each of the variables in more detail. The AGBD variables include agbd, agbd_pi_lower, agbd_pi_upper, agbd_prediction, agbd_se. The variable delta_time provides time for each shot. shot_number is a unique number assigned to each GEDI shot. Quality flags are included in l2_quality_flag and l4_quality_flag, among others. Geolocation of shots is provided in lat_lowestmode, lon_lowestmode, and elev_lowestmode. Let’s use this information to plot the path of each of the eight beams. First, let’s read the latitudes, longitudes, and time variables into a pandas dataframe.

Loading...

We can now plot each of the eight beams. The total number of shots (n) in each beam is shown in the parenthesis.

<Figure size 1400x600 with 8 Axes>

As shown in Figure 1 above, there are different numbers of shots in each beam group. The beams are usually always turned on over the land surface. This particular orbit scanned from west to east, as indicated by the delta_time variable (blue being the minimum and yellow the maximum).

The variable delta_time provides time for each shot in seconds since Jan 1 00:00 2018. We can convert the time to actual date and time using datetime module.

Start time: 2020-07-25 19:10:52.328367
End time: 2020-07-25 19:34:02.323636
Total period: 0:23:09.995269
Total shots: 1336839

As shown above, the GEDI instrument acquired the first shot of this sub-orbit at 2020-07-25 19:10:52 and the last shot at 2020-07-25 19:34:02. In this 23 minutes period, all eight beams recorded a total of ~1.3 million shots!

To look at the spatial distribution of the GEDI shots, let’s zoom in to an area in the Mount Leuser National Park in Indonesia, where the orbit overpasses, and plot all the beams together.

<Figure size 700x700 with 1 Axes>

GEDI beams are ~600m apart across the track. The distance between the two outer beams is ~4.2 km.

Now, let’s look at variables within one of the beam groups, BEAM0110, in more detail. We will read all the base variables within this beam group into a pandas dataframe.

Now, the beam_df dataframe contains all the variables within the BEAM0110 group. Let’s convert it to a geopandas dataframe to assign spatial context.

5.1 Prediction Stratum

Let’s compute percent shots by the predict_stratum variable for the BEAM0110.

<Figure size 2000x400 with 1 Axes>

This particular suborbit passes through 10 different prediction strata, as shown in Figure 3, covering several forest regions. About 23% of the total shots have no prediction stratum assigned and are not in one of the 35 GEDI predict strata (e.g., water).

There are three sub-groups within each beam group indicated by ‘GROUP’ in Table 5 above - agbd_prediction, geolocation, and land_cover_data.

6. Land Cover Group

The group land_cover_data within the beam group contains land cover information for the GEDI shots. Let’s read all the variables within the \BEAM0110\land_cover_data group.

Loading...

Let’s look at some key variables in more detail.

6.1 Plant Functional Types

The pft_class variable provides the plant functional types for the shots, which are derived from the MODIS land cover type product (MCD12Q1). It uses LC_Type5 or Annual PFT classification. Please refer to Table 7 of the MCD12Q1 user guide for the PFT type legend. Let’s plot distribution of shots by PFT class for the BEAM0110.

<Figure size 1200x700 with 1 Axes>

As we noted earlier (Figures 3 and 4), water bodies constitute ~23% of the total shots.

6.2 Urban Proportion

Another variable of interest is urban_proportion, which provides the proportion of urban land cover within the search area defined by urban_focal_window_size variable. The window size is defined as the number of pixels in the 25 m global urban mask developed by the GEDI science team using the TerraSAR-X and TanDEMX global urban footprint data products.

<Figure size 1200x700 with 2 Axes>

As shown in Figure 5, 0.9% of the total GEDI BEAM0110 shots falls over the urban areas (>50% urban cover). The shots with <=50% urban land cover is shown in the grey color.

6.3 Phenology

GEDI L4A footprints include phenology information derived from VIIRS/S-NPP Land Surface Phenology VNP22Q2 in the variables leaf_on_doy, leaf_off_doy and leaf_on_cycle. These variables together with the PFT are used to indicate whether the GEDI shots were recorded during the leaf-off or leaf-on conditions, provided in the leaf_off_flag. When the shot is collected after the onset of maximum greenness and before the midpoint of the senescence phase, they are marked with a flag of 1 (leaf-off).

Let’s plot the leaf_off_flag on a map.

<Figure size 1200x700 with 1 Axes>

In Figure 6 above, the leaf-off shots are indicated with red. The GEDI orbit was collected on July 25, 2020. There were only 1263 shots with leaf-off conditions, all within the Deciduous Broadleaf Trees PFT.

6.4 Tree Cover

GEDI uses Landsat-derived tree cover, which is defined as canopy closure for all vegetation taller than 5 m in height (Hansen et al., 2013). The following Figure 7 shows the percent tree cover.

<Figure size 1200x700 with 2 Axes>

7. Geolocation Group

The geolocation subgroup within the BEAM group contains elevation, latitude, and longitude of the lowest mode for each of the seven algorithm settings groups. Canopy height metrics are computed relative to the elevation of the lowest mode or the last mode of the GEDI waveform. GEDI uses six different algorithm settings (a1 to a6) for webform processing (refer to the GEDI L1/L2 Waveform Processing ATBD, Table 5 for more details). Additional algorithm settings of a10 in GEDI L4A indicate that a5 has been used, but the lowest detected mode is likely noise. When this occurs, a higher mode was used to calculate RH metrics (see GEDI L4A ATBD for details).

The geolocation subgroup also provides a beam sensitivity metric for each of the seven algorithm settings. The sensitivity indicates signal strength of the GEDI shot to penetrate canopies and detect the true ground level. The subgroup also includes stale_return_flag (0=good) to indicate the signal quality.

Let’s read all the variables within the \BEAM0110\geolocation group.

Loading...

8. AGBD variables

The agbd variable within the BEAM group contains estimated aboveground biomass density. Before looking further into agbd variables, let’s go and read in the variables agbd_prediction subgroup. The agbd_prediction subgroup contains AGBD variables for all 7 algorithm settings groups (see Section 7 above).

The agbd_prediction subgroup contains additional important prediction parameters used. Let’s print the attributes and the values.

Loading...

In the table above, predictor_offset and response_offset are the offset appplied to predictors and response variables before fitting the AGBD model. The variable alpha is the alpha value used for calculation of prediction intervals. The AGBD models currently have maximum number of 4 predictors (max_nvar) and 7 algorithm settings groups (l2a_alg_count).

Now, let’s plot the agbd variable for the BEAM0110 along the longitude axis.

<Figure size 2000x700 with 2 Axes>

As shown in Figure 8 above, ~56% of the shots in BEAM0110 have agbd values. The shots over water surface and areas with no vegetation have no agbd values and are indicated with fill values of -9999.

Let’s also plot the agbd on a map. In the map below, the shots with no agbd values are indicated with grey dots.

<Figure size 1200x700 with 2 Axes>

Plotting thousands of shots from the BEAM0110 in one figure may not show the variability in details. Let’s subset the data into a smaller area for a closer look. We will subset an area in the Mount Leuser National Park in Indonesia, where the GEDI orbit overpasses. We will clip around 6 km from Mt. Leuser.

Loading...

The geopandas dataframe leuser_gdf contains all the variables within BEAM0110 group for an area of 6 km around Mt. Leuser. Let’s plot the shots on a map.

<Figure size 800x800 with 1 Axes>

The start and end shot numbers are also shown on the map. GEDI shot numbers are 18 characters and in the format OOOOOBBFFFNNNNNNNN, where:

  • OOOOO: Orbit number

  • BB: Beam number

  • RR: Reserved for future use

  • G: Sub-orbit granule number

  • NNNNNNNN: Shot index

In Figure 10 above, the start shot has a value of 091680600300633705, which means it is from the orbit number 09168, the beam number 06, the sub-orbit granule number of 3 and the shot index of 00633705 within the orbit. There are a total of 169 shots from beam 06 (or BEAM0110) within 6km of Mt. Leuser.

Now, let’s plot the AGBD variables.

<Figure size 2000x800 with 1 Axes>

The agbd shows the selected aboveground biomass density, shown by black dots in Figure 11 above. The upper and lower prediction intervals are shown with a green-filled area.

GEDI shots are ~60m apart along the track. In the figure above, there are ~16-17 shots (shown by dots) within a distance of 1000 m. Note that a total of 25 shots with agbd fill values (-9999) are not plotted.

The above plot shows the AGBD estimates at alpha=0.1 (90% confidence interval), the default confidence interval used for these variables. To compute the estimates at a different confidence interval (e.g., 95%), we can use agbd_t and agbd_t_se variables, which provides the prediction in transform space. For more details, please refer to GEDI L4A Common Queries.

<Figure size 2000x800 with 1 Axes>

As noted above in Section 7, the subgroup agbd_prediction provides AGBD estimates for all seven algorithm settings groups (a1-6 and a10). Let’s plot all of them together. For clarity, we will plot selected shots in the following figure.

<Figure size 2000x800 with 1 Axes>

Figure 13 above shows the variability of the predicted biomass among the different algorithm setting groups.

9. Predictor variables

The xvar variable provides predictor variables. Let’s print the number of shots per predict_stratum in the leuser_gdf subset.

predict_stratum EBT_SAs 169 Name: count, dtype: int64

As shown above, all the GEDI shots near Mt. Leuser is classified as EBT_SAs prediction stratum (evergreen broadleaf trees PFT and South Asia region). The AGBD prediction model for EBT_SAs is given in Table 3 above as: AGBD = -104.9654541015625 + 6.802174091339111 x RH_50 + 3.9553122520446777 x RH_98.

The xvar variable has 4 columns, which we read into the dataframe above as xvar_1, xvar_2, xvar_3, and xvar_4. Let’s print the columns.

Loading...

Since the linear model for EBT_SAs prediction stratum has only two predictor variables, RH_50 and RH_98, the columns xvar_3 and xvar_4 in the above table are filled with zero values.

The xvar predictors are transformed and as well as offset applied. Table 3 above shows that square-root (sqrt) transformation is applied to the EBT_SAs prediction stratum both x- and y- variables, and a predictor offset (predictor_offset) of 100 are applied. We will need the offset and transformation to derive the true RH metric, and plot them along with agbd value.

<Figure size 2000x800 with 2 Axes>

10. Flag variables

Different flag variables are included for each shot. Most of these flags originate from lower-level products such as GEDI L2A, while some of the flags are specific to the GEDI L4A dataset. Let’s print all the flag variables’ names and descriptions.

Loading...

Let’s plot algorithm_run_flag, l2_quality_flag, and l4_quality_flag into a plot.

<Figure size 2200x800 with 2 Axes>

The algorithm_run_flag set is set to 1 when the shot quality is good to run the GEDIL4_A algorithm. When the algorithm is not run for the shot (indicated by algorithm_run_flag=0), all the AGBD variables will be set to the fill values (-9999). The l2_quality_flag identifies shots with good GEDI L2A quality metrics for the prediction of AGBD. The l4_quality_flag identifies shots that are representative of the applied AGBD models. Let’s apply both l2_quality_flag and l4_quality_flag to the dataframe.

Now, the dataframe leuser_gdf contains all the rows where both l2_quality_flag and l4_quality_flag filters are applied. Refer to the GEDI L4A ATBD and GEDI L4A Common Queries documents for more details about the quality flags.

11. Final Note

As we saw in this tutorial, in addition to AGBD estimates, a GEDI L4A file has information about predictor variables from the lower level GEDI products and other associated variables derived from multiple sensors such as MODIS, VIIRS, TanDEM-X, and Landsat, among others. GEDI generates ~15-16 orbits each day. GEDI is still collecting data every day and will do so for the next few years, which is an immense amount of data at our hands.

References
  1. Dubayah, R. O., Armston, J., Kellner, J. R., Duncanson, L., Healey, S. P., Patterson, P. L., Hancock, S., Tang, H., Bruening, J., Hofton, M. A., Blair, J. B., & Luthcke, S. B. (2022). GEDI L4A Footprint Level Aboveground Biomass Density, Version 2.1. 10.3334/ORNLDAAC/2056
  2. Kellner, J. R., Armston, J., & Duncanson, L. (2022). Algorithm theoretical basis document for GEDI footprint aboveground biomass density. Earth and Space Science, 9, e2022EA002516. 10.1029/2022EA002516
  3. Dubayah, R., Hofton, M., Blair, J., Armston, J., Tang, H., & Luthcke, S. (2021). GEDI L2A Elevation and Height Metrics Data Global Footprint Level V002. 10.5067/GEDI/GEDI02_A.002
  4. Friedl, M., & Sulla-Menashe, D. (2019). MCD12Q1 MODIS/Terra+Aqua Land Cover Type Yearly L3 Global 500m SIN Grid V006. NASA Land Processes Distributed Active Archive Center. 10.5067/MODIS/MCD12Q1.006
  5. Zhang, X., Friedl, M., & Henebry, G. (2020). VIIRS/NPP Land Cover Dynamics Yearly L3 Global 500m SIN Grid V001. NASA Land Processes Distributed Active Archive Center. 10.5067/VIIRS/VNP22Q2.001
  6. Dubayah, R., Luthcke, S., Hofton, M., & Blair, J. (2020). GEDI Waveform ATBD V001. 10.5067/DOC/GEDI/GEDI_WF_ATBD.001