Random Forest SPEI Dataset Reconstruction for 1 km Drought Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing SPEI datasets have low spatial resolution and spatiotemporal discontinuity, limiting their ability to accurately quantify drought events and providing excessive errors when used for quantitative analysis.
Innovation Solution
A high-resolution SPEI dataset development method based on a random forest regression model, combining meteorological station data, remote sensing data, and reanalysis data, with a spatial resolution of 1 km, using the FAO Penman-Monteith formula for potential evapotranspiration calculation and bicubic interpolation for data resampling, to construct a standardized precipitation evapotranspiration index (SPEI) dataset.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing SPEI datasets are used, then drought event identification is possible, but spatial resolution is low and spatiotemporal discontinuity exists
Solution Approach 1:
The patent divides the study area into multiple grid cells at 1km resolution and processes data separately for each grid cell, enabling high-resolution spatial analysis while maintaining temporal continuity through systematic processing of monthly data across multiple years
Solution Approach 2:
The patent introduces a random forest regression model as an intermediary to interpolate and reconstruct SPEI values across the study area, filling spatial gaps and ensuring continuous spatiotemporal coverage where direct measurements are unavailable
2Measurement precision
If existing SPEI datasets are used, then qualitative analysis of drought events is possible, but quantitative analysis produces excessive errors
Solution Approach 1:
The patent replaces traditional mechanical/statistical interpolation methods with a random forest regression model that uses machine learning algorithms, enabling more accurate quantitative analysis by capturing non-linear relationships and reducing calculation errors
Solution Approach 2:
The patent changes the resolution parameter from coarse to 1km fine resolution and transforms the data processing approach to include monthly aggregation and random forest modeling, thereby improving quantification accuracy while reducing analysis errors
3Measurement precision
If high-resolution data is generated, then spatial distribution features are refined, but data processing complexity increases
Solution Approach 1:
The patent employs the random forest model to automatically interpolate and reconstruct SPEI values across the entire study area without requiring manual processing of each grid cell, thereby reducing processing complexity while maintaining high resolution
Solution Approach 2:
The patent creates a unified processing framework that handles multiple tasks (data collection, cleaning, interpolation, reconstruction, and validation) within a single integrated system, reducing overall processing complexity despite the high resolution requirements
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The method achieves higher operation speed, prediction accuracy, and resistance to overfitting, enabling precise identification and quantitative research of drought events with refined spatial distribution features, guiding deeper drought monitoring and identification research.
Implementation Method 1
constructing the random forest regression model according to the sample points obtained in the step 7; wherein 80% of the sample points are randomly selected as training samples, and 20% of the sample points are used as testing samples
Implementation Method 2
calculating monthly potential evapotranspiration (PET) information on a station according to a FAO Penman-Monteith formula
Implementation Method 3
resampling spatial resolutions of the precipitation data, the land surface temperature data, the shortwave radiation data and the elevation data to 1 kilometer (km) through a bicubic interpolation algorithm
Data Source
AI summary
A high-resolution SPEI dataset development method based on a random forest regression model is provided. In the method, meteorological station data, GPM remote sensing precipitation data, MODIS land surface temperature data, ERA5-Land shortwave radiation data and SRTM digital elevation model data are combined; and a spatial pattern of SPEI index at different time scales of a target area is predicted by constructing a spatiotemporal relationship between the SPEI index and the precipitation, land surface temperature, shortwave radiation and elevation data. The method fully utilizes advantages that the random forest is high in precision and avoids overfitting in model prediction, and inputs station data and remote sensing and reanalysis data simultaneously into the model for training, which can solve problems of mismatch of an existing SPEI dataset with the station data and low spatial resolution, and the spatial resolution of SPEI dataset is effectively improved.


