Random Forest SPEI Dataset Reconstruction for 1 km Drought Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing SPEI datasets have low spatial resolution and spatiotemporal discontinuity, limiting their ability to accurately quantify drought events and providing excessive errors when used for quantitative analysis.

Innovation Solution

A high-resolution SPEI dataset development method based on a random forest regression model, combining meteorological station data, remote sensing data, and reanalysis data, with a spatial resolution of 1 km, using the FAO Penman-Monteith formula for potential evapotranspiration calculation and bicubic interpolation for data resampling, to construct a standardized precipitation evapotranspiration index (SPEI) dataset.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing SPEI datasets are used, then drought event identification is possible, but spatial resolution is low and spatiotemporal discontinuity exists

Engineering Contradiction:
Improvespatial resolutionVSAvoidspatiotemporal continuity
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent divides the study area into multiple grid cells at 1km resolution and processes data separately for each grid cell, enabling high-resolution spatial analysis while maintaining temporal continuity through systematic processing of monthly data across multiple years

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a random forest regression model as an intermediary to interpolate and reconstruct SPEI values across the study area, filling spatial gaps and ensuring continuous spatiotemporal coverage where direct measurements are unavailable

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If existing SPEI datasets are used, then qualitative analysis of drought events is possible, but quantitative analysis produces excessive errors

Engineering Contradiction:
Improvequantification accuracyVSAvoidanalysis error
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent replaces traditional mechanical/statistical interpolation methods with a random forest regression model that uses machine learning algorithms, enabling more accurate quantitative analysis by capturing non-linear relationships and reducing calculation errors

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the resolution parameter from coarse to 1km fine resolution and transforms the data processing approach to include monthly aggregation and random forest modeling, thereby improving quantification accuracy while reducing analysis errors

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If high-resolution data is generated, then spatial distribution features are refined, but data processing complexity increases

Engineering Contradiction:
Improvespatial resolutionVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent employs the random forest model to automatically interpolate and reconstruct SPEI values across the entire study area without requiring manual processing of each grid cell, thereby reducing processing complexity while maintaining high resolution

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates a unified processing framework that handles multiple tasks (data collection, cleaning, interpolation, reconstruction, and validation) within a single integrated system, reducing overall processing complexity despite the high resolution requirements

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The method achieves higher operation speed, prediction accuracy, and resistance to overfitting, enabling precise identification and quantitative research of drought events with refined spatial distribution features, guiding deeper drought monitoring and identification research.

Implementation Method 1

constructing the random forest regression model according to the sample points obtained in the step 7; wherein 80% of the sample points are randomly selected as training samples, and 20% of the sample points are used as testing samples

Methodology Applied
Scientific EffectRandom forest regression:

Implementation Method 2

calculating monthly potential evapotranspiration (PET) information on a station according to a FAO Penman-Monteith formula

Methodology Applied
Scientific EffectEvaporation: Evaporation

Implementation Method 3

resampling spatial resolutions of the precipitation data, the land surface temperature data, the shortwave radiation data and the elevation data to 1 kilometer (km) through a bicubic interpolation algorithm

Methodology Applied
Scientific EffectInterpolation:

Data Source

PatentUS20240094436A1High-resolution standardized precipitation evapotranspiration index dataset development method based on random forest regression model
Publication Date: 2024.03.21 HENAN UNIVERSITY
  • US20240094436A1 patent drawing
  • US20240094436A1 patent drawing
  • US20240094436A1 patent drawing

AI summary

A high-resolution SPEI dataset development method based on a random forest regression model is provided. In the method, meteorological station data, GPM remote sensing precipitation data, MODIS land surface temperature data, ERA5-Land shortwave radiation data and SRTM digital elevation model data are combined; and a spatial pattern of SPEI index at different time scales of a target area is predicted by constructing a spatiotemporal relationship between the SPEI index and the precipitation, land surface temperature, shortwave radiation and elevation data. The method fully utilizes advantages that the random forest is high in precision and avoids overfitting in model prediction, and inputs station data and remote sensing and reanalysis data simultaneously into the model for training, which can solve problems of mismatch of an existing SPEI dataset with the station data and low spatial resolution, and the spatial resolution of SPEI dataset is effectively improved.