Multi-source rainfall fusion correction method for regional rainfall type landslide threshold
Through the multi-source precipitation fusion framework, the accuracy and spatial matching problems of determining rainfall-induced landslide thresholds in existing technologies are solved, high-precision regional landslide threshold calculation is achieved, and the development of regional early warning systems is supported.
Patent Information
- Application Number
- CN202510863918.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies are unable to accurately determine the regional rainfall-induced landslide threshold in areas without measured precipitation data or with discontinuous data. In addition, the low spatial resolution of reanalysis precipitation data makes it impossible to accurately define the landslide threshold, which limits the development of regional early warning systems.
A multi-source precipitation fusion framework was adopted. Statistical downscaling was performed and the number of dry and wet days was corrected based on the characteristics of rainfall events that triggered landslides. A daily precipitation conversion function was established. The meteorological station data were corrected using the Penman-Monteith method. Representative sequences were generated using the Thiessen polygon method and the nearest neighbor method. The cumulative distribution function was fitted using the Gumbel distribution and other methods. A regional raster conversion function was established. Finally, the threshold was determined using the power law method and the two-dimensional Bayesian method.
The generation of long-term, high-resolution, and high-precision fused precipitation data provides a technical means to establish regional landslide thresholds in areas with missing or discontinuous data, thereby improving the accuracy and coverage of the regional landslide early warning system.
Smart Images

Figure CN120744751A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of hydrological meteorology and water disaster technology, in particular to a fusion correction method for multi-source precipitation data (including observed precipitation, satellite precipitation and reanalysis precipitation) and its application in regional rainfall-induced landslide thresholds. Background Art
[0002] Landslides are one of the most dangerous and frequent natural geological hazards worldwide, posing a serious threat to human life, health, and property. Precipitation, a primary dynamic trigger of landslides, infiltrates rocks and soils, altering their inherent mechanical equilibrium, increasing soil bulk density and pore water pressure, reducing shear strength, and ultimately leading to slope instability. Therefore, determining regional rainfall-induced landslide threshold curves based on precipitation indicators is crucial for establishing regional early warning systems, conducting risk assessments, and conducting other disaster prevention and mitigation efforts.
[0003] The existing technical solution proposes a statistical method for calculating the threshold of regional rainfall-induced landslides. This method has been widely used due to its simple calculation, ease of use and low cost. However, the accuracy of this method is directly affected by the quality of the input data from the rain gauge. Therefore, it is impossible to establish an effective threshold in areas without measured precipitation data or where the data is discontinuous, which limits the development of regional early warning systems. In addition, rain gauges can only reflect the precipitation conditions in a limited area near the station, while landslides often occur in remote mountainous areas with high altitudes, steep slopes and few people. When the rain gauge is far away from the location where the landslide occurs, it cannot represent the actual triggering rainfall intensity for the landslide.
[0004] In recent years, major weather forecast centers around the world have used numerical prediction models and assimilation techniques to generate a series of global reanalysis precipitation datasets, such as the European Centre for Medium Range Weather Forecasts (ECMWF)'s ERA5 dataset and the Japan Meteorological Agency (JMA)'s JRA-55 and CRA datasets. This type of data offers advantages such as a long time span, wide coverage, and global temporal and spatial uniformity. It serves as an important reference for regions where observational data is scarce or unavailable, and has been widely used in water resource management and flood disaster management. However, unlike the point-like distribution of rainfall station data, reanalysis precipitation data are stored in raster format. Their low precision and spatial resolution make them inadequate for accurately defining landslide thresholds.
[0005] Therefore, a multi-source precipitation data fusion framework has been developed to address landslide threshold characteristics. This framework overcomes the low precision, short time series, and spatiotemporal mismatch between rainfall intensity and landslide location in existing data. This approach expands the application scale from single points to grids, and from individual landslides to regional clusters of rainfall-induced landslides. This approach will not only directly promote the development of regional landslide early warning systems and serve the national comprehensive disaster prevention and mitigation plan, but also improve the effectiveness of disaster prevention and mitigation efforts, providing key scientific and technological support for building a climate-resilient disaster prevention and control system. Summary of the Invention
[0006] In order to address the defects of the above-mentioned existing technologies in that they cannot provide rainfall landslide thresholds in areas with continuous or scarce data, the present invention aims to propose a multi-source precipitation fusion framework for landslide forecasting by means of statistical downscaling, correcting the number of dry and wet days based on the characteristics of rainfall events that trigger landslides, and establishing a daily-scale precipitation conversion function. This framework overcomes the defects of low precision, short time series, and spatiotemporal mismatch between rainfall intensity and landslide location in existing data, and upgrades the application scale from single point / single slope to grid / region. This provides an effective technical means for establishing an early landslide warning system for large-scale regions and areas with missing / discontinuous precipitation observation records.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A multi-source precipitation fusion correction method for regional rainfall-induced landslide thresholds includes:
[0009] Step 1: Data collection. This includes downloading 30m-resolution DEM data for the study area, historical daily precipitation data from the China Meteorological Network, satellite precipitation data, and global reanalysis precipitation data (e.g., CRA, ERA5, and ERA5-Land). Historical landslide catalog data includes downloading the NASA Global Landslide Catalog, landslide records from the China Geological Survey, and interpretations of historical satellite imagery (e.g., Landsat).
[0010] Step 2: Data preprocessing. This involves clipping the .shp file based on the study area boundaries and extracting precipitation data for the target study period. Daily precipitation data from the China Meteorological Network is topographically corrected using the Penman-Monteith method. This step ultimately integrates precipitation data from various sources into data of the same time period and spatial resolution.
[0011] Step 3: Statistical downscaling. Based on statistical interpolation methods (ordinary kriging or cokriging), the large-scale, low-resolution reanalysis product raster precipitation values (if necessary, combined with geographic elevation data) that have been processed and clipped are downscaled to the target spatial resolution to generate a new raster precipitation dataset.
[0012] Step 4: Correct and identify dry and wet days. First, an appropriate threshold is selected to distinguish between rainy and rainless states. Specifically, the optimal threshold is selected based on the overlap of rainfall event characteristics (such as the start date of the rainfall event, peak flood time, and the number of rainfall events per year).
[0013] Step 5: Fitting the precipitation probability distribution. First, the rainfall series of the station and the raster data are paired according to the spatial coordinates. Based on the number of stations contained in the raster area, the Thiessen polygon method or the nearest neighbor method is selected to generate a representative series. The discriminant value obtained by the above method is used to filter out wet days with small daily precipitation. The cumulative distribution functions (CDFs) of the representative station and the interpolated reanalysis precipitation are fitted.
[0014] Step 6: Bias correction. The cumulative distribution function is fitted with six functions: Gumbel distribution, gamma distribution, Gaussian distribution, and Weibull distribution to obtain shape and scale parameters, and a regional grid conversion function is established to achieve accuracy correction of the fused daily precipitation.
[0015] Step 7: Threshold establishment and method evaluation. The power law method and two-dimensional Bayesian method were used to establish regional rainfall-induced landslide thresholds based on the fused data and observed data, respectively. The accuracy of the method was evaluated using the confusion matrix and ROC curve.
[0016] The beneficial effects of the present invention are:
[0017] The present invention provides a technical means for establishing regional landslide thresholds in data-missing or discontinuous areas.
[0018] The present invention generates long-term (>30 years), high-resolution (≤1km), and high-precision fused precipitation data.
[0019] The conversion function relationship of the present invention can be applied to the correction of historical and future precipitation data on the same grid. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 This is a framework diagram of a multi-source precipitation fusion correction method for regional landslide threshold determination according to an embodiment of the present invention;
[0022] Figure 2Flowchart for the identification and correction of three methods for determining dry / wet days based on rainfall event characteristics according to an embodiment of the present invention; (a) is a schematic diagram of the overlap of the number of annual precipitation events, (b) is a schematic diagram of the overlap of the start dates of rainfall events, and (c) is a schematic diagram of determining the optimal discrimination value for the overlap of false peak values of precipitation. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0024] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] like Figure 1 As shown, this embodiment discloses a multi-source precipitation fusion correction method for regional landslide threshold determination, including: S1: obtaining a global reanalysis precipitation dataset / satellite precipitation dataset and a ground meteorological station observation dataset, obtaining a landslide catalog set, and obtaining a high-precision terrain and boundary vector file of the study area; preprocessing meteorological station data according to terrain characteristics; S2: statistically downscaling the reanalysis data; S3: discrimination correction based on dry days / wet days of landslide events; S4: fitting the daily precipitation probability distribution; S5: quantile mapping deviation correction; S6: determining the rainfall-landslide critical threshold curve and accuracy evaluation based on the conditional Bayesian method.
[0026] This embodiment discloses a multi-source precipitation fusion correction method for regional landslide threshold determination, including: step 1, data collection step. The collection step includes downloading 30m resolution DEM data of the study area, historical daily precipitation data from the domestic relevant department network contained in the study area, satellite precipitation data and global reanalysis precipitation data (such as CRA, ERA5, ERA5-Land). Landslide historical cataloging data includes downloading the Global Landslide Catalog (NASA Global Landslide Catalog), landslide records formed by domestic relevant departments, and historical satellite image (such as Landsat) interpretation results. Step 2, data preprocessing step. The preprocessing step includes clipping the shp file according to the boundary of the study area and extracting precipitation data for the target study period. Among them, the daily precipitation of the domestic relevant departments needs to be terrain corrected by the Penman-Monteith method. The final result of this step is to integrate precipitation data from different sources into data with the same time period and spatial resolution. Step 3, statistical downscaling. Based on statistical interpolation methods (ordinary kriging or cokriging), the large-scale, low-resolution reanalysis product data raster precipitation values (if necessary, combined with geographic elevation data) that have been processed and clipped are downscaled to the target spatial resolution to generate a new raster precipitation dataset. Step 4: Correct and identify the number of dry and wet days. First, an appropriate threshold needs to be selected to distinguish between rainy and rainless states. Specifically, the optimal threshold is selected based on the degree of overlap of rainfall event characteristics (such as the start date of the rainfall event, the peak time, and the number of rainfall events per year). Step 5: Fit the precipitation probability distribution. First, the rainfall series of the station is paired with the raster data according to the spatial coordinates. Based on the number of stations contained in the raster area, the Thiessen polygon method or the nearest neighbor method is selected to generate a representative series. The discriminant value obtained by the above method is used to filter out wet days with small daily precipitation. The cumulative distribution functions (CDFs) of the representative station and the reanalysis precipitation interpolation are fitted. Step 6: Bias correction. The cumulative distribution function was fitted using six functions: Gumbel, Gamma, Gaussian, and Weibull distributions to obtain shape and scale parameters. A regional raster conversion function was then established to achieve accuracy correction of the fused daily precipitation. Step 7: Threshold establishment and method evaluation. Regional rainfall-induced landslide thresholds were established using the power law method and the two-dimensional Bayesian method based on the fused and observed data, respectively. The accuracy of the methods was evaluated using confusion matrices and receiver operating characteristic (ROC) curves.
[0027] The input data of step S1 include: daily precipitation observation data of the site, satellite remote sensing precipitation data, global reanalysis precipitation data and landslide historical catalog data; wherein, the sources of the landslide historical catalog data include: global landslide catalog, landslide archive records of the Geological Survey and historical satellite image interpretation results.
[0028] The preprocessing of precipitation data in step S1 includes: converting the format of the original precipitation data, extracting the daily precipitation sequence of the target time series according to the boundary file of the target study area, and extracting the precipitation data within the target period. For the precipitation data of the meteorological station, terrain correction is performed according to the Penman-Monteith formula. Since meteorological stations are often located in relatively flat areas, the radiation R in the Penman-Monteith formula is used to correct the terrain. n , temperature T and aerodynamic parameter v2 and other terrain sensitive parameters correction, the precipitation observed at the station can better serve the early warning of rainfall-induced landslides in mountainous areas;
[0029]
[0030] Among them, ET o Reference crop evapotranspiration, R n is the net radiation, T is the temperature, G is the soil heat flux density, γ is the psychrometric constant, u2 is the wind speed at 2 meters, e s is the saturated water vapor pressure, e a is the actual water vapor pressure, Δ is the slope of the saturated water vapor pressure curve, ET o is the reference crop evapotranspiration.
[0031] The global landslide database described in step S1 is processed with the landslide catalog of the study area, which has the following features: spatial matching based on geographic coordinate positions, setting a buffer zone of 1 km, and considering two landslide points as representing the same landslide event when they are located within each other's buffer zones; deleting duplicate record entries and records of landslides induced by earthquakes and engineering excavation.
[0032] Specifically, the radiation R in the Penman-Monteith formula is used. n , temperature T and aerodynamic parameter v2 and other terrain-sensitive parameters to correct the rainfall data of the observation station, so that it can better represent the actual precipitation values in mountainous areas and areas with larger slopes;
[0033]
[0034] The preprocessing of the landslide historical catalog data includes: spatial matching based on geographic coordinate positions, setting the buffer zone as the target distance, when two landslide points are located in each other's buffer zone, they represent the same landslide event, and deleting duplicate record entries and non-rainfall-induced landslide records.
[0035] Spatial matching was performed based on geographic coordinates, and the buffer zone was set to 1 km. When two landslide points were located within each other's buffer zones, they were considered to represent the same landslide event. Duplicate records and non-rainfall-induced landslide records were deleted.
[0036] Step S2 statistically downscaling the reanalysis data includes: optional module 1 is for downscaling the raster reanalysis data by coupling the ordinary Kriging interpolation method with the terrain factor for different terrain features such as mountainous areas, hills, plains and coastal areas, and improving its resolution from 0.25° to 0.1°; optional module 2 is for downscaling the raster reanalysis data by coupling the ordinary Kriging interpolation method with the rainfall intensity factor for different rainfall intensities such as light rain, moderate rain and heavy rain, and improving its resolution from 0.25° to 0.1°; optional module 3 is for improving its resolution from 0.25° to 0.1° by superimposing the ordinary Kriging interpolation method with the time scale conversion method for different time scale formats such as annual scale, seasonal scale and monthly scale.
[0037] Step S2 statistically downscaling the reanalysis data includes:
[0038] In view of the different terrain characteristics of mountainous areas, hilly areas, plains and coastal areas, the reanalysis data were downscaled using the ordinary kriging interpolation method coupled with terrain factors, and the resolution was increased from 0.25° to 0.1°.
[0039] For different rainfall intensities, such as light rain, moderate rain, and heavy rain, the reanalysis data were downscaled using the ordinary Kriging interpolation method coupled with the rainfall intensity factor, increasing its resolution from 0.25° to 0.1°.
[0040] For different time scale formats such as annual scale, seasonal scale and monthly scale, the ordinary Kriging interpolation method is superimposed with the time scale conversion method to increase the resolution from 0.25° to 0.1°.
[0041] Specifically, the Ordinary Kriging (OK) or Co-Kriging (Co-Kriging) geostatistical algorithm uses spatial data correlation to improve the coarse resolution to high resolution. Taking the Ordinary Kriging interpolation method as an example, under the assumption of second-order stationarity, the linear estimation algorithm uses the random function Z(x i ) and the variogram γ(h) is the estimated point x q Provides the best unbiased estimate The expression is as follows:
[0042]
[0043] Where, is the weight coefficient. This coefficient is determined by minimizing the error variance under the unbiased constraint, and the generation formula is as follows:
[0044]
[0045] The variogram γ describes the spatial dependence of the data and is defined as the semivariance measure with respect to the distance h:
[0046]
[0047] where γ(h) is the semivariance; N(h) is the total number of pairs separated by distance h; Z(x i ) and Z(x i+h ) are samples measured at point i and (i+h) respectively.
[0048] According to the characteristics of the study area, the terrain (such as mountains, hills, plains and coastal areas) module, time scale (annual scale, seasonal scale) module and rainfall intensity module (light rain, moderate rain, heavy rain and rainstorm) can be selected and downscaling can be performed using the co-kriging method.
[0049] Step S3 is based on the discrimination and correction of dry days / wet days of landslide events, including:
[0050] identifying representative reanalysis data grids or representative sites using the geographic coordinates of landslide occurrences in the landslide history catalog data, and pairing the representative precipitation series of the reanalysis data grids or the representative sites with the downscaled precipitation series;
[0051] Based on the pairing results, the optimal discriminant value is determined based on the overlap of the number of annual precipitation events, the overlap of the start date of rainfall events, and the overlap of the peak value of precipitation events. The discriminant value is selected within the range of 0 mm to 50 mm. When the rainfall is lower than the selected discriminant value, the condition is corrected to no rain, and the day is marked as a dry day. The number of consecutive dry days that separate the two rainfall events is determined based on the rainy and dry seasons, such as 4 days in the rainy season and 2 days in the dry season, to generate the number of rainfall events. The number of annual rainfall events is calculated based on the overlap of the number of annual precipitation events, and the discriminant value is determined by comparing the number of rainfall events. When the multi-year mean absolute error (MAE) reaches the minimum value, the corresponding selected discriminant value is considered to be optimal, that is, the number of dry / wet days determined under the condition of using this discriminant value is considered to be consistent with the actual state. The expression is as follows:
[0052]
[0053] Among them E i is the number of rainfall events in the i-th year after the dry and wet day discrimination value of rainfall characteristics is determined, O i is the number of rainfall events in year i recorded for the landslide;
[0054] The method for determining the optimal discriminant value based on the coincidence of the starting date of the rainfall event is as follows: when the rainfall event determined by the selected discriminant value and the representative precipitation sequence of the station have the same starting date, the same precipitation event is captured. When the multi-year average absolute error (MAE) reaches the minimum value, the selected discriminant value is considered to be the optimal value; the method for determining the optimal discriminant value based on the coincidence of the peak value of the precipitation event is as follows: when the rainfall event determined by the selected discriminant value and the representative precipitation sequence of the station have the same maximum precipitation date, the same precipitation event is captured. When the multi-year average absolute error (MAE) reaches the minimum value, the selected discriminant value is considered to be the optimal value.
[0055] Step S4 is to fit the daily precipitation probability distribution, and the specific steps include:
[0056] Step S4-1 is the pairing process, which pairs the precipitation time series of the station with the downscaled precipitation data. When there are fewer than three stations, the nearest neighbor method is used to determine the grid area controlled by the corresponding station to generate a representative precipitation series. When there are more than three stations, the Thiessen polygon method based on Delaunay triangulation is used to determine the grid area controlled by the corresponding station to generate a representative precipitation series.
[0057] Step S4-2 is to fit the daily precipitation series. This involves drawing the empirical cumulative distribution function (ECDF) based on the observed precipitation data and the representative grid precipitation data after the dry and wet day correction described in step 1. The ECDF is used to calculate the proportion of daily precipitation that is less than or equal to a certain daily amount and is defined as:
[0058] Among them F n (x) is the value of the empirical cumulative distribution function, which represents the cumulative probability that precipitation is ≤ x; n is the total number of historical precipitation data samples, x i is the ith actual precipitation observation value, I(x i ≤x) is the indicator function, where x is the value of precipitation. Six probability functions including Gumbel distribution, gamma distribution, Pearson type III distribution, Gaussian distribution and Weibull distribution are used to fit the cumulative distribution functions (CDFs).
[0059] Furthermore, step S5 of quantile mapping deviation correction includes:
[0060] The cumulative distribution function is fitted based on the six probability functions, and the optimal probability fitting function is determined. The positional and morphological parameters in the fitting formula are extracted to form a transfer function (CF). This is then combined with a quantile map (QQ plot) to achieve correction. That is, the corrected fused raster data and the observed data are plotted in the same QQ plot to check whether they are closer to the 45° diagonal. The quantile map deviation correction formula is:
[0061]
[0062] in, and F m,h They correspond to the observation site data x o,h The inverse function of the cumulative distribution function and the fused raster data x before correction m,h The historical cumulative distribution function, x m,p (t) is the data before bias correction, The data are bias-corrected.
[0063] Specifically, identifying landslide events triggered by precipitation and extracting the correct values of cumulative precipitation-precipitation duration pairs (E-Dpairs) are prerequisites for establishing empirical rainfall thresholds. To this end, the present invention needs to reconstruct individual rainfall events from continuous rainfall time series, and separate two rainfall events by rainless periods, such as Figure 2 As shown, rainfall events 1 and 2 are separated by a period of consecutive dry days. For traditional rain gauges, a rainless state is defined as 0 mm. Due to measurement and random errors, the interpolated reanalysis precipitation data contain many small values close to 0, which can severely alter the number of rainy days and significantly affect the threshold formula. To obtain accurate ED pairs for rainfall events, the present invention proposes the following mechanism to determine the rain / no-rain state discriminant value, thereby adjusting for the difference in dry-wet days between the interpolated results and the observed data.
[0064] The proposed mechanism includes the following steps: 1) Generate precipitation time series pairs: Use the geographic coordinates of the landslide to identify representative reanalysis data grids or rain gauges. For rain gauges, the representative grid is the one closest to the landslide site. For ERA5 or ERA5-Land, the representative grid is the grid point containing the landslide site. For each landslide record, the matching rainfall time series from the rain gauge is paired with the ERA5 rainfall time series for the corresponding grid point.
[0065] 2) Reconstruct rainfall events using discriminant values: with an interval of 0.1 mm, a range of discriminant values from 0 to 50 mm were selected for input testing. When the rainfall was lower than the selected discriminant value, the condition was corrected to no rain, and the day was marked as a dry day. Since the climate of the study area is characterized by distinct dry and wet seasons, based on the typical rainfall patterns observed in these seasons, the minimum interval between two consecutive rainfall events was defined as two days from May to September (wet season) and four days from October to April (dry season). The dry season is usually characterized by short rainfall periods and long intervals.
[0066] 3) Determine the discriminant value by comparing the number of rainfall events: Based on observational data, the rainfall series is segmented year by year, and the number of rainfall events is counted each year. By calculating and comparing the number of rainfall events, the rain / no rain detection capabilities of different discriminant values are evaluated. When the mean absolute error (MAE) reaches the minimum value, the discriminant value is considered optimal, indicating that the dry / wet days determined using this discriminant value are consistent with the actual state.
[0067]
[0068] This paper uses the Quantile Mapping (QM) method to improve the accuracy of ERA5 daily precipitation estimates. The QM method assumes that climate model biases are stable from history to the future and is a method for adjusting climate model outputs. Specifically, the QM method uses a conversion function between the cumulative distribution functions (CDFs) of observed and model data over the historical period to correct for model biases using a quantile-quantile plot. The specific expression is as follows:
[0069]
[0070] Among them, F o,h and F m,h They correspond to the observed data x o,h and model data x m,h function.
[0071] The bias correction module consists of three processes: a pairing process, a transfer function process, and a bias correction process. The pairing process pairs the rainfall series from a station with the raster reanalysis dataset. Depending on the number of available stations across the study area, this process has two scenarios: When there are fewer than three stations, a nearest neighbor method is used to locate the rainfall time series from the station and the corresponding grid. If there are more than three stations, a Delaunay triangulation function using the Thiessen polygon method is used to determine the grid area controlled by each station.
[0072] Transfer function process: First, the discriminant value obtained by the aforementioned method is used to filter out wet days with low daily precipitation. Then, the cumulative distribution functions (CDFs) of the interpolated precipitation values of the representative stations and the reanalysis are simultaneously calculated and fitted.
[0073] Furthermore, step S6 determines the rainfall-landslide critical threshold curve and accuracy assessment based on the conditional Bayesian method, including:
[0074] According to the optimal dry-wet day discrimination threshold obtained above, the independent rainfall events are segmented according to the interruption day criterion, that is, the dry days in the rainy season are greater than or equal to 4 days, and the dry days in the dry season are greater than or equal to 2 days. The accumulated rainfall E and the duration D are calculated, and the landslide records are traversed to generate the trigger type (E L -D L ) and non-trigger type (E N -D N ) Landslide rainfall event set;
[0075] The critical rainfall landslide threshold line is defined using two-dimensional Bayesian method. The specific formula is as follows:
[0076]
[0077] Here, "P(B,C│A)" represents the joint conditional probability of observing two specific values or ranges of cumulative rainfall and rainfall duration, given that a landslide (A) has occurred; "P(A)" represents the prior probability of a landslide event, and "P(B,C)" is the joint probability of observing two specific value ranges of cumulative rainfall and rainfall duration. The discriminant efficacy of thresholds was evaluated using confusion matrices (true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN), as well as receiver operating characteristic (ROC) curves. The optimal rainfall characteristics for identifying dry and wet days were determined.
[0078] This embodiment discloses a multi-source precipitation fusion correction method for regional landslide threshold determination, including: obtaining a global reanalysis precipitation dataset / satellite precipitation dataset and a ground meteorological station observation dataset, obtaining a landslide catalog, and obtaining a high-precision terrain and boundary vector file for the study area; preprocessing meteorological station data according to terrain characteristics; statistical downscaling; correction based on the discrimination of dry days / wet days of landslide events; fitting the daily precipitation probability distribution; quantile mapping bias correction; threshold establishment and accuracy assessment.
[0079] Based on the collected initial precipitation data and the target time period (such as the time series daily precipitation data from 2000 to 2020);
[0080] For the precipitation data of meteorological stations, terrain correction is performed according to the Penman-Monteith formula. Since meteorological stations are often located in relatively flat areas, the radiation R in the Penman-Monteith formula is used to correct the terrain. n , temperature T and aerodynamic parameter v2 correction of terrain sensitive parameters, precipitation observation can better serve the early warning of rainfall-induced landslides in mountainous areas;
[0081]
[0082] The global landslide database was further processed with the landslide catalog of the study area. Its features include: spatial matching based on geographic coordinate positions, setting a buffer zone of 1 km, and considering two landslide points as representing the same landslide event when they are located within each other's buffer zone; and deleting duplicate record entries and non-rainfall-induced landslide records.
[0083] According to statistical downscaling (combined with the attached Figure 1 Examples are described), including:
[0084] Ordinary Kriging is used to perform statistical interpolation based on different terrain regions (mountainous areas, hilly areas, plains, or coastal areas) to generate raster data with target accuracy (e.g., spatial resolution of 0.1°≈9km);
[0085] According to different rainfall intensities (light rain, moderate rain, heavy rain, etc.), ordinary kriging method is used for statistical interpolation to generate raster data with target accuracy (e.g. spatial resolution of 0.1°≈9km);
[0086] Ordinary Kriging method is used for statistical interpolation according to different time scales (such as annual scale, seasonal scale and monthly scale) to generate raster data with target accuracy (such as spatial resolution of 0.1°≈9km).
[0087] According to the daily precipitation probability fitting (combined with the attached Figure 1 Examples are described), including:
[0088] Pairing process: Pair the precipitation time series of the station with the raster data. When the number of stations is less than three, the nearest neighbor method is used to determine; when the total number of stations exceeds three, the Thiessen polygon method based on Delaunay triangulation is used to determine the grid area controlled by the corresponding station to generate a representative precipitation series;
[0089] Daily precipitation series fitting: Based on the observed precipitation data and representative grid precipitation data, after the dry and wet days correction in step 1, the empirical cumulative distribution function is drawn; the cumulative distribution function is based on the proportion of daily precipitation less than or equal to a certain amount, and is defined as:
[0090]
[0091] Where x is the given precipitation;
[0092] Six probability functions including Gumbel distribution, gamma distribution, Pearson type III distribution, Gaussian distribution and Weibull distribution are used to fit cumulative distribution functions (CDFs).
[0093] According to the quantile mapping bias correction (combined with the attached Figure 1 Examples are described), including:
[0094] Compare and determine the optimal probability fitting cumulative distribution functions (CDFs), determine the position parameters and morphological parameters after fitting to determine the transfer function (CF), combine with the quantile map (QQ map) to achieve correction, plot the corrected fused raster data and observation data in the same QQ map, and check whether it is closer to the 45° diagonal line. The specific expression is as follows:
[0095]
[0096] Among them, F o,h and F m,h They correspond to the observed data x o,h and model data x m,h function.
[0097] Correction of dry / wet day discrimination of landslide events (combined with the attached Figure 2 Examples are described), including:
[0098] Step (1) pairing rainfall sequences; step (2) optimal threshold for distinguishing dry and wet days; step (3) reconstructing cumulative rainfall-duration pairs (ED); step (4) determining the critical rainfall threshold formula for triggering landslide events in the study area and step (5) using confusion matrix and ROC curve for correction steps;
[0099] In step (2), there are three identification and correction methods for determining dry / wet days based on the characteristics of rainfall events, namely, the method based on the overlap of the number of annual precipitation events ( Figure 2 a), coincidence of rainfall event start date ( Figure 2 b) and the coincidence degree of false peak of precipitation to determine the optimal discrimination value ( Figure 2 c). When the rainfall is lower than the selected discrimination value, the condition is corrected to no rain, and the day is marked as a dry day. The number of consecutive dry days that separate the two rainfall events is determined based on the rainy and dry seasons (e.g., 4 days in the rainy season and 2 days in the dry season). Three methods are used to calculate the number of annual rainfall events. The discrimination value is determined by comparing the number of rainfall events. When the mean absolute error (MAE) reaches the minimum value, the discrimination value is optimal, that is, the dry / wet days determined under the condition of using this discrimination value are considered to be consistent with the actual state. The expression is as follows:
[0100]
[0101] Where E is the number of rainfall events in the i-th year determined based on the dry-wet day discrimination value of rainfall characteristics, and O is the number of rainfall events in the i-th year recorded by the landslide;
[0102] The dry and wet day discrimination threshold determined by the optimal rainfall characteristic method can be applied to Figure 1 Step 2 is shown.
[0103] According to the threshold-based establishment and accuracy evaluation (combined with the attached Figure 2 Examples are described), including:
[0104] According to the interruption day criterion (dry days ≥ 4 days in rainy season; dry days ≥ 2 days in dry season), the independent rainfall events are divided, the cumulative rainfall E and the duration D are calculated, and the landslide records are traversed to generate the trigger type (E L -D L ) and non-trigger type (E N -D N ) Landslide rainfall event set;
[0105] Using the power law method E = α·D β Or the two-dimensional Bayesian definition of the critical rainfall landslide threshold line, the specific formula is as follows:
[0106]
[0107] Here, "P(B,C│A)" represents the joint conditional probability of observing two specific values or ranges of cumulative rainfall and rainfall duration, given that a landslide (A) has occurred; "P(A)" represents the prior probability of a landslide event, and "P(B,C)" is the joint probability of observing two specific value ranges of cumulative rainfall and rainfall duration. The discriminant efficacy of thresholds was evaluated using confusion matrix true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN), as well as receiver operating characteristic (ROC) curves, to determine the optimal rainfall characteristics for identifying dry and wet days.
[0108] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A multi-source precipitation fusion correction method for regional rainfall-induced landslide threshold, characterized by: include: S1: Preprocess meteorological station data using the Penman-Monteith formula according to terrain characteristics; S2: statistical downscaling of reanalysis data; S3: Correction for discrimination based on dry / wet days of landslide events; S4: daily precipitation probability distribution fitting; S5: Quantile mapping bias correction; S6: Determination of rainfall-landslide critical threshold curve and accuracy evaluation based on conditional Bayesian method.
2. The multi-source precipitation fusion correction method for regional rainfall-type landslide threshold according to claim 1 is characterized in that: The input data of step S1 include: daily precipitation observation data of the site, satellite remote sensing precipitation data, global reanalysis precipitation data and landslide historical catalog data; wherein, the sources of the landslide historical catalog data include: global landslide catalog, landslide archive records of the Geological Survey and historical satellite image interpretation results.
3. The multi-source precipitation fusion correction method for regional rainfall-type landslide threshold according to claim 2 is characterized in that: The preprocessing of precipitation data in step S1 includes: The original precipitation data was formatted, and the daily precipitation sequence of the target time series was extracted according to the boundary file of the target study area, and the precipitation data within the target period was extracted. For the precipitation data of the meteorological station, the terrain correction was performed according to the Penman-Monteith formula. Since meteorological stations are often located in relatively flat areas, the radiation R in the Penman-Monteith formula was used to correct the terrain. n , temperature T and aerodynamic parameter v2 correction of terrain sensitive parameters, station observation of precipitation can better serve the early warning of rainfall-induced landslides in mountainous areas; Among them, ET o Reference crop evapotranspiration, R n is the net radiation, T is the temperature, G is the soil heat flux density, γ is the psychrometric constant, u2 is the wind speed at 2 meters, e s is the saturated water vapor pressure, e a is the actual water vapor pressure, Δ is the slope of the saturated water vapor pressure curve, ET o is the reference crop evapotranspiration; The global landslide catalog and the landslide archive records of the Geological Survey described in step S1 are processed, and its features include: spatial matching based on geographic coordinate positions, setting a buffer zone of 1 km, and when two landslide points are located within each other's buffer zones, they are considered to represent the same landslide event; duplicate record entries and records of landslides induced by earthquakes and engineering excavation are deleted.
4. The multi-source precipitation fusion correction method for regional rainfall-type landslide threshold according to claim 1 is characterized in that: Step S2 statistically downscaling the reanalysis data includes: In view of the different terrain characteristics of mountainous areas, hilly areas, plains and coastal areas, the reanalysis data were downscaled using the ordinary kriging interpolation method coupled with terrain factors, and the resolution was increased from 0.25° to 0.1°. For different rainfall intensities, such as light rain, moderate rain, and heavy rain, the reanalysis data were downscaled using the ordinary Kriging interpolation method coupled with the rainfall intensity factor, increasing its resolution from 0.25° to 0.1°. For different time scale formats such as annual scale, seasonal scale and monthly scale, the ordinary Kriging interpolation method is superimposed with the time scale conversion method to increase the resolution from 0.25° to 0.1°.
5. The multi-source precipitation fusion correction method for regional rainfall-type landslide threshold according to claim 1 is characterized in that: Step S4 is to fit the daily precipitation probability distribution, and the specific steps include: Step S4-1 is the pairing process, which pairs the precipitation time series of the station with the downscaled precipitation data. When there are fewer than three stations, the nearest neighbor method is used to determine the grid area controlled by the corresponding station to generate a representative precipitation series. When there are more than three stations, the Thiessen polygon method based on Delaunay triangulation is used to determine the grid area controlled by the corresponding station to generate a representative precipitation series.
6. The multi-source precipitation fusion correction method for regional rainfall-type landslide threshold according to claim 5 is characterized in that: Step S3 is based on the discrimination and correction of dry days / wet days of landslide events, including: identifying representative reanalysis data grids or representative sites using the geographic coordinates of landslide occurrences in the landslide history catalog data, and pairing the representative precipitation series of the reanalysis data grids or the representative sites with the downscaled precipitation series; Based on the pairing results, the optimal discriminant value is determined based on the overlap of the number of annual precipitation events, the overlap of the start date of rainfall events, and the overlap of the peak value of precipitation events. The discriminant value is selected within the range of 0 mm to 50 mm. When the rainfall is lower than the selected discriminant value, the condition is corrected to no rain, and the day is marked as a dry day. The number of consecutive dry days that separate the two rainfall events is determined based on the rainy and dry seasons, such as 4 days in the rainy season and 2 days in the dry season, to generate the number of rainfall events. The number of annual rainfall events is calculated based on the overlap of the number of annual precipitation events, and the discriminant value is determined by comparing the number of rainfall events. When the multi-year mean absolute error (MAE) reaches the minimum value, the corresponding selected discriminant value is considered to be optimal, that is, the number of dry / wet days determined under the condition of using this discriminant value is considered to be consistent with the actual state. The expression is as follows: Among them E i is the number of rainfall events in the i-th year after the dry and wet day discrimination value of rainfall characteristics is determined, O i is the number of rainfall events in year i recorded for the landslide; The method for determining the optimal discriminant value based on the coincidence of the starting date of the rainfall event is as follows: when the rainfall event determined by the selected discriminant value and the representative precipitation sequence of the station have the same starting date, the same precipitation event is captured. When the multi-year average absolute error (MAE) reaches the minimum value, the selected discriminant value is considered to be the optimal value; the method for determining the optimal discriminant value based on the coincidence of the peak value of the precipitation event is as follows: when the rainfall event determined by the selected discriminant value and the representative precipitation sequence of the station have the same maximum precipitation date, the same precipitation event is captured. When the multi-year average absolute error (MAE) reaches the minimum value, the selected discriminant value is considered to be the optimal value.
7. The multi-source precipitation fusion correction method for regional rainfall-type landslide threshold according to claim 5 is characterized in that: Daily precipitation series fitting, that is, based on observed precipitation data and representative grid precipitation data, draw the empirical cumulative distribution function ECDF; the cumulative distribution function is used to calculate the proportion of daily precipitation less than or equal to a certain day, and its expression is: Among them F n (x) is the value of the empirical cumulative distribution function, which represents the cumulative probability that precipitation is ≤ x; n is the total number of historical precipitation data samples, x i is the ith actual precipitation observation value, I(x i ≤x) is the indicator function, where x is for a given precipitation value; six probability functions including Gumbel distribution, gamma distribution, Pearson type III distribution, Gaussian distribution and Weibull distribution are used to fit the cumulative distribution functions (CDFs).
8. The multi-source precipitation fusion correction method for regional rainfall-type landslide threshold according to claim 1 is characterized in that: Step S5, quantile mapping bias correction, includes: The cumulative distribution function is fitted according to the six probability functions, and the optimal probability fitting function is determined. The location parameters and morphological parameters in the fitting formula are extracted to form a transfer function, and the correction is achieved by combining the quantile map. The corrected fused raster data and the observed data are plotted on the same quantile map. The quantile map deviation correction expression is: in, and F m,h They correspond to the observation site data x o,h The inverse function of the cumulative distribution function and the fused raster data x before correction m,h The historical cumulative distribution function, x m,p (t) is the data before bias correction, The data are bias-corrected.
9. The multi-source precipitation fusion correction method for regional rainfall-type landslide threshold according to claim 1 is characterized in that: Step S6 determines the rainfall-landslide critical threshold curve and accuracy assessment based on the conditional Bayesian method, and its characteristics include: Based on the optimal dry-wet day discrimination threshold obtained above, the independent rainfall events are segmented according to the interruption day criterion to generate the triggering and non-triggering landslide rainfall event sets; The critical rainfall landslide threshold line is defined using two-dimensional Bayesian method. The specific formula is as follows: Among them, P(B,C│A) represents the joint conditional probability of observing two specific values or ranges of cumulative rainfall and rainfall duration simultaneously, given the premise that landslide A has occurred; P(A) represents the prior probability of a landslide event, and P(B,C) is the joint probability of observing two specific value ranges of cumulative rainfall and rainfall duration; the confusion matrix (true positive TP, false positive FP, true negative TN, false negative FN) and the receiver operating characteristic curve (ROC) are used to evaluate the quality of the threshold.
Citation Information
Cited By
Event-level rainfall sample construction method and device for rainfall type landslide
CN121834357A