A high-resolution mountain flood-prone mapping method based on a mixture model

By combining the Top-SSF model and the random forest algorithm, and using DEM and meteorological data to assess the flood susceptibility of mountainous areas, the problem of high-precision flood susceptibility mapping in mountainous watersheds with no available data was solved, and high-precision flood risk identification and prediction were achieved.

CN120850095BActive Publication Date: 2026-04-07SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies are insufficient for conducting high-precision flood susceptibility assessments in mountainous watersheds lacking data. Traditional hydrological models rely on multi-source data and expert experience, leading to uncertainties in the results and making it impossible to effectively identify flood-risk areas.

Method used

A hybrid model approach was adopted, combining the Top-SSF model and the random forest algorithm. Flood susceptibility mapping was carried out using DEM data and meteorological parameters. The model parameters were optimized using the SCE-UA algorithm, key factors were screened, and an RF model was constructed to predict flood susceptibility.

Benefits of technology

It enables the generation of 15m×15m ultra-fine resolution flood susceptibility maps with only basic topographic data and limited meteorological parameters, improving spatial accuracy by three orders of magnitude, reducing factor weight uncertainty, and enhancing the spatial matching degree and prediction accuracy of risk zoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850095B_ABST
    Figure CN120850095B_ABST
Patent Text Reader

Abstract

The application discloses a high-resolution mountain flood-prone mapping method based on a mixed model, and the method comprises the following steps: S1, calibrating a Top-SSF model based on a physical process in a plurality of mountain basins respectively; S2, presetting a period of flood peak flow calculation; S3, making a flood inventory map of the basins; S4, screening key flood-inducing factors; S5, constructing a flood-prone model based on RF; and S6, high-resolution flood-prone mapping in the mountains. The application can be widely applied to a mountain flood disaster early warning system and flood control management, and provides scientific basis and technical support for disaster prevention and control. Through identification of high-risk areas and prediction of potential flood threats, the method significantly improves the efficiency and accuracy of flood risk management, and provides an innovative solution for flood-prone assessment of basins without data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of mountain flood disaster prevention, and in particular to a high-resolution mountain flood-prone mapping method based on a hybrid model. BACKGROUND

[0002] Global mountain flood disaster prevention faces severe challenges, and its strong suddenness and destructive power pose a major threat to mountain ecological safety and economic and social development. According to statistics, flood disasters account for one-third of the total amount of global natural disasters, and mountainous areas have become high-risk areas for mountain flood disasters due to their complex terrain and concentrated rainfall.

[0003] The current mainstream flood-prone assessment technology system has multidimensional defects and cannot meet the precise disaster prevention and mitigation needs of mountainous areas. First, the application bottleneck of traditional hydrological physical models. Traditional hydrological models based on Saint-Venant equations (such as HEC-HMS, SWMM) simulate the basin runoff process through mathematical equations, but their operation is highly dependent on multi-source data support such as terrain, rainfall, and flow. In mountainous basins lacking hydrological observation stations, the difficulty in obtaining basic data makes model parameter calibration extremely difficult, and the high computational cost required to solve partial differential equations severely restricts the production of flood inventory maps in data-poor mountainous areas. Second, the application limitations of traditional multi-criteria decision methods for flood-prone mapping. Multi-criteria decision methods based on the analytic hierarchy process have long relied on expert experience to assign factor weights in flood-prone assessment, and their inherent defects significantly restrict the objectivity and scientificity of the results. Such methods determine the weights of each factor by constructing a hierarchical structure and a pairwise comparison matrix, but the introduction of expert subjective judgment in the core link leads to significant uncertainty and poor repeatability in weight allocation. For example, in complex mountainous terrain conditions, the relative importance of key factors such as slope and terrain wetness index is difficult to accurately quantify through qualitative experience, and expert cognitive bias can easily cause the risk zoning results to deviate from the actual disaster distribution. In addition, the analytic hierarchy process has weak modeling capabilities for nonlinear relationships and cannot effectively reveal the dynamic coupling mechanism of the rainstorm-runoff-concentration process, especially in dealing with multi-factor interactions such as the storage effect of vegetation cover on soil infiltration. Therefore, there is an urgent need to develop a high-resolution mountain flood-prone mapping method based on a hybrid model to provide scientific basis and technical support for mountain flood disaster prevention in data-poor mountainous areas. SUMMARY

[0004] In order to overcome the shortcomings and deficiencies of the prior art, the present application provides a high-resolution mountain flood-prone mapping method based on a hybrid model.

[0005] The technical solution adopted by the present application is a high-resolution mountain flood-prone mapping method based on a hybrid model, comprising the following steps:

[0006] Step S1: Calibration of Top-SSF model based on physical processes in multiple mountainous river basins; Top-SSF model combines canopy interception, soil infiltration and storm runoff hydrological processes, and optimizes model parameters through SCE-UA algorithm, uses historical flood data for calibration, calibration process uses historical flood event data, and optimizes model parameters through SCE-UA algorithm, target function is Kling-Gupta efficiency coefficient, Nash-Sutcliffe coefficient, relative error of peak flow and peak time error as evaluation indicators;

[0007] Step S2: Calculation of preset period peak flow; Gumbel extreme value distribution model is constructed in multiple mountainous river basins through historical rainfall data, and the preset period rainfall intensity and value of the basin are calculated, combined with ERA5-Land reanalysis data to obtain meteorological parameters at the same period, input into the calibrated Top-SSF model, simulate the preset period flood hydrograph of multiple basins, and extract the peak flow value of each basin;

[0008] Step S3: Production of basin flood list map; use DEM data for flow direction analysis, calculate terrain humidity index HAND by tracking the elevation difference of slope grid cells to the nearest river drainage point; combined with the peak flow of the basin output by the Top-SSF model, use Manning formula to inverse river cross section water depth, identify potential flooded area by spatial comparison of HAND value and river water depth, calculate flooded water depth for each pixel, and form basin flood spatial distribution list map;

[0009] Step S4: Screening of key flood inducing factors; 20 flood inducing factors are determined through preliminary screening, then variance inflation factor analysis is used to exclude variables with significant multicollinearity, and combined with autocorrelation and significance test to remove insignificant factors;

[0010] Step S5: Construction of RF-based flood prone model; based on flood list data and screened key flood inducing factors, construct random forest model, use 70% data for training, 30% data for testing, and perform independent verification through receiver operating characteristic curve and statistical indicators;

[0011] Step S6: High-resolution flood prone mapping in mountainous areas: through spatial matching of each grid cell in mountainous river basin with 16 key flood inducing factors, apply trained random forest classification model for flood prone prediction, model output probability value F(X) quantifies the flood prone degree of unit cell; combined with 21 typical mountain flood events recorded in the Flood and Drought Bulletin, carry out independent verification, through confusion matrix analysis and receiver operating characteristic curve evaluation, verify the spatial recognition accuracy and prediction reliability of the model for high-resolution flood sensitive area.

[0012] Further, the step S2 comprises:

[0013] S21: Estimation of rainfall events in a preset period, obtain historical rainfall data of the study area from the weather station, use the extreme value distribution method to analyze the frequency of rainfall data, and calculate the rainfall intensity and rainfall amount in the preset period;

[0014] S22: Weather data preparation, extract the weather data of the preset period from the ERA5-Land reanalysis data, including temperature, relative humidity, wind speed, and solar radiation;

[0015] S23: Input the rainfall events in the preset period and their corresponding weather data into the calibrated Top-SSF model, simulate the flood hydrograph of multiple basins in the preset period, and extract the peak flow.

[0016] Further, the step S3 comprises:

[0017] S31: Calculate the flow direction of each grid cell according to the DEM;

[0018] S32: Identification of the nearest drainage point, mark all slope grid cells with the same river flow direction, identify the nearest drainage point, and for each grid cell, track to the nearest river grid cell along the flow direction and record its elevation value ;

[0019] S33: HAND value calculation, subtract the elevation of the nearest drainage point from the elevation of the slope grid cell to obtain the HAND value,

[0020]

[0021] wherein, is the HAND value; is the elevation of the slope grid cell; is the elevation of the nearest drainage point;

[0022] S34: Calculation of the flooded area, first calculate the river water level according to the river flow;

[0023] According to the peak flow extracted by the Top-SSF model Combined with the geometric and hydraulic relationship of the river cross section, the Manning formula is used to estimate the river water depth;

[0024]

[0025] wherein, is the peak flow, is the Manning coefficient, is the river cross section area; is the hydraulic radius; is the river slope;

[0026] S35: Submerged area determination, compare HAND value with river water level to determine potential submerged areas, if a grid cell satisfies the following conditions, it is considered to be submerged:

[0027]

[0028] S36: Water depth calculation, for submerged areas, calculate the water depth d of each grid cell;

[0029]

[0030] wherein, is the water depth.

[0031] Further, the step S4 comprises:

[0032] S41: Preliminary screening, preliminarily determine 20 flood inducing factors, including topography, vegetation, meteorological factors;

[0033] S42: VIF analysis, exclude factors with high multicollinearity through VIF analysis, and retain factors that have significant impact on flood;

[0034] S43: Correlation analysis, analyze the autocorrelation between factors and the correlation with flood, and eliminate statistically insignificant factors.

[0035] Further, the step S5 comprises:

[0036] S51: Data preparation, randomly sample from the flood list map of 80 river basins to generate 320,000 data points, of which 70% are used for training and 30% are used for testing, each data point contains 16 key flood inducing factors;

[0037] S52: Train RF binary classification model to distinguish flood points and non-flood points;

[0038] S53: Performance evaluation, evaluate the performance of the model on the test set, and use the receiver operating characteristic curve and statistical indicators, specificity, accuracy for verification;

[0039]

[0040]

[0041]

[0042]

[0043] wherein, TP is the number of true positives, actual flood points and correctly predicted as flood points; TN is the number of true negatives, actual non-flood points and correctly predicted as non-flood points; FP is the number of false positives, actual non-flood points but wrongly predicted as flood points; FN is the number of false negatives, actual flood points but wrongly predicted as non-flood points.

[0044] Further, the step S6 comprises:

[0045] S61: input data, match all grid cells in the study area with key flood inducing factors;

[0046] S62: model prediction, use the trained RF model to make classification prediction for each grid cell, judge whether it belongs to flood prone area, the probability value F(X) output by the RF model represents whether the grid cell belongs to flood prone area;

[0047] S63: independent verification, use 21 mountain flood events in the flood and drought bulletin to independently verify the accuracy of the model prediction.

[0048] Further, the step S1 takes KGE as the objective function, and Nash-Sutcliffe coefficient flood peak flow relative error and peak time error as evaluation indexes, which are calculated as follows:

[0049] Set as the objective function:

[0050]

[0051] wherein, is the Pearson correlation coefficient; is the bias term, i.e. the mean ratio; is the variable term, i.e. the coefficient of variation ratio;

[0052] Set the upper and lower limits and initial values of the model parameters; set the number of iterations, the maximum number of evolutionary cycles, the percentage change allowed in the cycle and the number of evolutionary steps for each complex before shuffling in the optimization algorithm, run the optimization algorithm, save the optimal parameters; use the collected mountain basin flood events to evaluate the model performance with , and as evaluation indexes;

[0053]

[0054]

[0055]

[0056] wherein, is the measured flood value at a certain time; is the simulated flood value at a certain time; is the measured flood mean value; is the relative error of flood peak; is the measured flood peak flow; is the simulated flood peak flow; is the error of peak occurrence time; is the measured peak occurrence time; is the simulated peak occurrence time.

[0057] Further, the step S21, the extreme value distribution method, the probability density function:

[0058]

[0059] wherein, is a random variable; is a location parameter, reflecting the central position of the distribution; is a scale parameter, reflecting the dispersion degree of the distribution, is a probability density function, is a natural constant;

[0060] Cumulative distribution function:

[0061]

[0062] Quantile function:

[0063]

[0064] wherein, is a non-transcendental probability.

[0065] Further, the step S31, according to the DEM, the flow direction of each grid cell is calculated, and the calculation is as follows:

[0066] The water flow direction of each grid cell points to the lowest elevation point in its adjacent 8 grid cells;

[0067]

[0068] wherein, is the elevation of the adjacent grid cell, m; is the flow direction of the grid cell.

[0069] Further, the step S52 trains the RF binary classification model to distinguish flood points and non-flood points, and calculates as follows:

[0070] According to the input feature X, the output label Y is predicted, wherein, : represents 16 key flood-inducing factors of each data point; : represents a binary classification label, represents a flood point, represents a non-flood point;

[0071] A decision number classification rule is generated by each decision tree by splitting the feature space, and the classification rule of the tth decision tree is assumed to be:

[0072]

[0073] That is, if the input feature X satisfies the splitting condition of the tree, it is predicted as a flood point , otherwise it is predicted as a non-flood point .

[0074] The integration rule of the random forest is that the random forest integrates the results of multiple decision trees, and adopts majority voting to make the final classification:

[0075]

[0076] wherein, is the total number of decision trees in the random forest, is the classification result of the tth decision tree on the input X, and F(X) is the final classification result of the random forest model.

[0077] Compared with the prior art, the present application has the advantages of:

[0078] (1) Breaking through the data dependence bottleneck of traditional physical models, realizing accurate mapping of "dataless basins": In view of the high dependence of traditional hydrological models (such as HEC-HMS) on topography, rainfall, flow and other multi-source observation data, the present application innovatively adopts a coupling framework of DEM data driven Top-SSF runoff model and HAND hydrodynamic model, combined with the strong generalization ability of random forest, to generate a 15m x 15m ultra-fine resolution flood-prone map under the condition of only basic topographic data and limited meteorological parameters. Compared with the traditional 0.1°x0.1° (about 10km x 10km) resolution physical model, the spatial accuracy is improved by 3 orders of magnitude, and the "dataless basin" modeling problem caused by the lack of monitoring sites in mountainous areas is successfully overcome.

[0079] (2) Reconstruct the multi-factor weight assignment paradigm to eliminate the subjective bias of expert experience: In view of the weight uncertainty problem caused by the dependence of the analytic hierarchy process on expert experience, the invention establishes a "mechanism-oriented-data-driven" double-cycle factor screening system: First, 12 key inducements such as slope (Slope), topographic position index (TPI), and topographic wetness index (TWI) are screened out from 20 candidate factors; Then through random forest iterative training, the contribution of each factor to the flood response is quantified, realizing the paradigm shift from "qualitative experience" to "quantitative data" in weight distribution. This method reduces the uncertainty of factor weight by 74%, and the spatial matching degree of risk zoning and historical disaster points is improved to 85.71%. BRIEF DESCRIPTION OF DRAWINGS

[0080] Figure 1 The method flowchart of the invention. DETAILED DESCRIPTION

[0081] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict, and the present application will be further described in detail below in combination with the drawings and specific embodiments.

[0082] As Figure 1 shown, a high-resolution mountain flood-prone mapping method based on a hybrid model, the application cases of the method are as follows:

[0083] The southwest mountainous area of China includes Yunnan Province, Guizhou Province, Sichuan Province, Guangxi Zhuang Autonomous Region and Chongqing City, with a total area of about 1,373,100 square kilometers. More than 80% of the area is steep mountainous area, with an average elevation of 2,130 meters and an average slope of 8.8°. The forest coverage rate is 51%. Influenced by the Pacific and Indian Ocean monsoon, this area belongs to subtropical monsoon climate, with an average annual temperature of 20°C and an annual precipitation of 1,400 mm, concentrated from May to September. The complex terrain and seasonal heavy rainfall make this area a high-incidence area of mountain flood disasters. The dense river network and high soil permeability further aggravate the flood risk. From 1960 to 2015, a total of 7,447 mountain flood events were recorded, resulting in 18,257 deaths, posing a serious threat to human life and property and infrastructure.

[0084] The establishment of the method includes the following steps:

[0085] Step S1: Calibration of the Top-SSF model based on physical processes in 80 southwest mountainous basins; the Top-SSF model considers key hydrological processes such as canopy interception, soil infiltration, and storm runoff, and optimizes model parameters through the SCE-UA algorithm. The calibration process uses historical flood data in the southwest region, and optimizes model parameters through the SCE-UA algorithm. The objective function is the Kling-Gupta efficiency coefficient (KGE), and the Nash-Sutcliffe coefficient (NSE), relative error of flood peak flow (REPF), and peak time error (RPT) are used as evaluation indicators. The results show that in the calibration period, about 50% of the basins have NSE values > 0.78, < 12%, and < 3 hours (50th percentile). In the validation period, about 50% of the basins have NSE values > 0.75 (maximum NSE = 0.96), < 11%, and < 3 hours (50th percentile). The error is within 3 hours. These results indicate that the model is reliable in capturing the dynamic characteristics of floods in the study area under complex topography and climate conditions.

[0086] Step S2: Calculation of preset period flood peak flow; taking 80 mountainous basins in the southwest region as an example, Gumbel extreme value distribution model is constructed based on historical rainfall data (1991-2020), and the preset period rainfall intensity and value of 80 basins in the southwest region are calculated. Combined with ERA5-Land reanalysis data to obtain meteorological parameters (temperature, humidity, etc.) during the same period, the calibrated Top-SSF model is input, and the preset period flood hydrograph of 80 basins is simulated. Finally, the flood peak flow value of each basin is obtained.

[0087] Step S3: Production of basin flood list map; based on high-precision DEM data, flow direction analysis is carried out, and terrain humidity index HAND is calculated by tracking the elevation difference of slope grid cells to the nearest river drainage point; combined with the flood peak flow of the basin output by the Top-SSF model, the Manning formula is used to inverse the water depth of the river cross section, and finally the potential inundation area is identified by spatial comparison of HAND value and river water depth, and the inundation water depth is calculated pixel by pixel to form a high-resolution flood spatial distribution list map in the southwest mountainous region.

[0088] ​​​​​​​​​Step S4: Screening key flood inducing factors; 20 flood inducing factors (including slope, terrain position index TPI, terrain wetness index TWI, vegetation parameters such as forest coverage, and meteorological elements such as rainfall and temperature) are determined through preliminary screening, followed by variance inflation factor (VIF) analysis to exclude variables with significant multicollinearity (such as NDVI, soil type), and combined with autocorrelation and significance test (P>0.05) to remove insignificant factors. Finally, 16 key flood driving factors are selected, mainly including core topographic feature parameters such as slope, TPI, TWI, and dominant meteorological elements such as rainfall intensity.

[0089] Step S5: Constructing RF-based flood susceptibility model; based on the flood list data and the screened key flood inducing factors, a random forest (RF) model is constructed, 70% of the data is used for training, 30% of the data is used for testing, and independent verification is carried out through the receiver operating characteristic curve (ROC) and statistical indicators (such as AUC, sensitivity, specificity).

[0090] Step S6: High-resolution flood susceptibility mapping in mountainous areas: Through the spatial matching data of each grid cell in the mountainous watershed with 16 key flood inducing factors (such as slope, terrain position index TPI, terrain wetness index TWI, etc.), the trained random forest (RF) classification model is applied for flood susceptibility prediction, and the model output probability value F(X) quantifies the flood susceptibility of the cell; combined with the records of 21 typical mountain flood events in China's flood and drought bulletins from 2010 to 2024, independent verification is carried out, and through confusion matrix analysis and ROC curve evaluation, the spatial recognition accuracy and prediction reliability of the model for high-resolution flood sensitive areas are verified.

[0091] Further, the step S2 comprises:

[0092] S21: Estimation of rainfall events in a preset period, obtain historical rainfall data (hourly rainfall from 1991 to 2020) of the study area from meteorological stations. Ensure that the time span of the data is long enough (usually at least 30 years) to improve statistical reliability. Use extreme value distribution method (such as Gumbel distribution) to analyze the frequency of rainfall data. Calculate the rainfall intensity and rainfall in the preset period.

[0093] S22: Meteorological data preparation, extract the meteorological data of the preset period corresponding to the 80 basins in the Southwest mountainous area from the ERA5-Land reanalysis data, including temperature, relative humidity, wind speed, solar radiation, etc.

[0094] S23: Input the rainfall events in the preset period and their corresponding meteorological data into the calibrated Top-SSF model to simulate the flood hydrograph of the 80 basins in the preset period, and extract the peak flow.

[0095] Further, the step S3 comprises:

[0096] S31: Calculate the flow direction of each grid cell according to DEM.

[0097] S32: Identify the nearest drainage point, mark all slope grid cells that have the same flow direction as the river grid cell, and identify the nearest drainage point. For each grid cell, track along the flow direction to the nearest river grid cell and record its elevation value .

[0098] S33: Calculate the HAND value, subtract the elevation of the nearest drainage point from the elevation of the slope grid cell to obtain the HAND value.

[0099]

[0100] wherein, HAND is the HAND value, m; z is the elevation of the slope grid cell, m; z0 is the elevation of the nearest drainage point, m.

[0101] S34: Calculate the inundation area, first calculate the river water level according to the river flow.

[0102] Peak flow extracted according to the Top-SSF model Combined with the geometric and hydraulic relationship of the river section, the Manning formula is used to estimate the river water depth.

[0103]

[0104] wherein, Q is the peak flow, m3 / s, n is the Manning coefficient, A is the cross-sectional area of the river section, m3; R is the hydraulic radius, m; I is the river slope.

[0105] S35: Determine the inundation area, compare the HAND value with the river water level to determine the potential inundation area. If the grid cell meets the following conditions, it is considered to be inundated:

[0106]

[0107] S36: Calculate the water depth, for the inundation area, calculate the water depth d

[0108]

[0109] wherein, Water depth, m.

[0110] Further, the step S4 comprises:

[0111] S41: Preliminary screening, 20 flood-inducing factors are initially identified, including topography (slope, terrain position index TPI, terrain wetness index TWI, etc.), vegetation (forest coverage), meteorology (rainfall, temperature), etc.

[0112] S42: VIF analysis, factors with high multicollinearity are excluded through VIF analysis. In the southwest mountainous area, the VIF values of elevation (Elevation), normalized difference vegetation index (NDVI), precipitation (Prec), air temperature (Tem), and soil type (Soil) exceed the threshold of 10, indicating significant multicollinearity problems, and thus are excluded.

[0113] S43: Correlation analysis, analyze the autocorrelation between factors and the correlation with floods, and eliminate statistically insignificant factors (P>0.05). The terrain relief index (TRI) and the terrain position index (TPI) show a significant negative correlation, but TPI is retained due to its effectiveness in complex terrain analysis and its ability to characterize flow velocity and flow path. Similarly, the river sediment transport index (STI) and the river dynamic index (SPI) show a significant positive correlation, but SPI can directly quantify the flow energy and is more suitable for identifying high flood risk areas and riverbank erosion prone areas, so STI is excluded. In addition, the influence of STI on the flooded area is not statistically significant (P>0.05). Finally, 16 key flood-inducing factors are selected, including slope, TPI, TWI, etc.

[0114] Further, the step S5 comprises:

[0115] S51: Data preparation, 320,000 data points are generated by randomly sampling from the flood inventory maps of 80 basins, of which 70% are used for training and 30% are used for testing. Each data point contains 16 key flood-inducing factors.

[0116] S52: Train RF binary classification model, distinguish flood points and non-flood points.

[0117] S53: Performance evaluation, evaluate the model performance on the test set, use ROC curve and statistical indicators (such as AUC, sensitivity (Sensitivity), specificity (Specificity), accuracy (Accuracy) for verification.

[0118]

[0119]

[0120]

[0121]

[0122] in, The number of True Positives, which are actually flood points and were correctly predicted as flood points; A true negative is the number of instances that are actually non-flood points but were correctly predicted as non-flood points. False positives are the number of points that are not actually flood points but are incorrectly predicted as flood points. False negatives are the number of instances that were actually flooded but incorrectly predicted as non-flooded. Validation metrics further confirmed the accuracy of the invention: Sensitivity = 0.998, Specificity = 0.997, AUC = 0.994 (159,807 true positives, 159,506 true negatives). The high AUC and low false positive rate indicate that the model can effectively distinguish between inundated and non-inundated areas, validating its reliability as a flood risk assessment tool in data-scarce mountainous areas.

[0123] Further, step S6 includes:

[0124] S61: Input data to match all grid cells within the study area with 16 key flood-inducing factors.

[0125] S62: Model prediction. The trained RF model is used to classify and predict whether each grid cell belongs to a flood-prone area. The probability value F(X) output by the RF model represents whether the grid cell belongs to a flood-prone area.

[0126] S63: Independent validation was conducted using 21 flash flood events from the 2010-2024 China Flood and Drought Disaster Bulletin to verify the accuracy of the model's predictions. 85.71% (18 events) occurred in high-risk areas, confirming the predictive effectiveness of the susceptibility map. For example, flash flood events in Wushan County, Guizhou Province, and Guannan County, Chongqing Province, were accurately identified, demonstrating the framework's applicability to real-world disaster scenarios.

[0127] Further, in step S1, the objective function is KGE, and the Nash-Sutcliffe coefficient (NSE) and peak flow relative error (NSE) are used. ) and peak occurrence time error ( As an evaluation indicator, it is calculated as follows:

[0128] set up The objective function is:

[0129]

[0130] wherein, is the Pearson correlation coefficient; is the bias term (i.e. mean ratio); is the variability term (i.e. coefficient of variation ratio);

[0131] Then, the model parameters upper and lower limits and initial values are set; the iteration number, the maximum evolution cycle number, the percentage change allowed in the cycle and the number of evolution steps of each complex before shuffling in the optimization algorithm are set. The optimization algorithm is run, and the optimal parameters are saved; the collected flood events in the southwest mountainous river basin are used to , and as evaluation indexes to evaluate the model performance.

[0132]

[0133]

[0134]

[0135] wherein, is the observed flood value at a certain moment, m3 / s; is the simulated flood value at a certain moment, m3 / s; is the observed flood mean value, m3 / s; is the relative error of flood peak, m3 / s; is the observed flood peak flow, m3 / s; is the simulated flood peak flow, m3 / s; is the error of peak occurrence time, h; is the observed peak occurrence time, h; is the simulated peak occurrence time, h.

[0136] Further, in step S21, the extreme value distribution method (such as Gumbel distribution) is calculated as follows:

[0137] Probability density function (PDF):

[0138]

[0139] wherein, is a random variable; is a location parameter (reflecting the central position of the distribution); is a scale parameter (reflecting the dispersion degree of the distribution), is the probability density function.

[0140] Cumulative distribution function:

[0141]

[0142] Quantile function (inverse CDF):

[0143]

[0144] where, is the non-transcendental probability (e.g., , denotes the preset period event).

[0145] Further, step S31, the flow direction of each grid cell is calculated according to DEM, and the calculation is as follows:

[0146] The water flow direction of each grid cell points to the lowest elevation point in its adjacent 8 grid cells.

[0147]

[0148] where, is the elevation of the th adjacent grid cell, m; is the flow direction of the grid cell.

[0149] Further, step S52 trains the RF binary classification model to distinguish between flood points and non-flood points, and the calculation is as follows:

[0150] Random forest is an ensemble learning method based on decision trees. For binary classification problems, the goal is to predict the output label Y from the input feature X, where,

[0151] : represents 16 key flood-inducing factors for each data point.

[0152] : represents the binary classification label, represents a flood point, represents a non-flood point.

[0153] Decision number classification rule, each decision tree generates a classification rule by dividing the feature space. Suppose the classification rule of the tth decision tree is:

[0154]

[0155] That is, if the input feature X satisfies the splitting condition of the tree, it is predicted to be a flood point ( ), otherwise it is predicted to be a non-flood point ( ).

[0156] The ensemble rule of random forest is to integrate the results of multiple decision trees and make the final classification by majority voting.

[0157]

[0158] where, is the total number of decision trees in the random forest. is the classification result of the t-th decision tree for the input X. F(X) is the final classification result of the random forest model.

[0159] The hybrid model-based high-resolution mountain flood susceptibility mapping method performs satisfactorily in the Southwest China mountainous region. First, it shows high precision in flood susceptibility prediction, with an accuracy of up to 0.98 and an AUC value of up to 0.994 on the test set. Independent verification using 21 mountain flood event data from 2010 to 2024 shows that 85.71% of the mountain flood events were accurately predicted, further proving the reliability and accuracy of the model. In addition, this method realizes a 15m x 15m resolution flood susceptibility map covering a total area of 1,373,100 square kilometers in the Southwest China mountainous region. Compared with the traditional flood inventory map, this method identifies 40.16% additional potential flood susceptible areas, especially in areas that may be affected under extreme scenarios such as dam breaches or storm surges.

[0160] The application solves the dual difficulties of strong data dependence and large subjective factor weight in mountain flood susceptibility assessment by constructing a "physical mechanism constraint + machine learning driven" integrated modeling paradigm. In view of the high dependence of traditional hydrological models (such as HEC-HMS) on topography, rainfall, flow and other multi-source observation data, the Top-SSF runoff model and the HAND water power model framework are innovatively integrated, and the basic topographic parameters are automatically extracted by driving 15-meter high-precision DEM data, and combined with the strong nonlinear fitting ability of the random forest algorithm, a super-fine resolution (15m x 15m) flood susceptibility map can be generated under the condition of only basic topographic data and limited meteorological parameters, which is 3 orders of magnitude higher in spatial accuracy than the traditional 0.1° x 0.1° (about 10km x 10km) physical model. The technical system effectively replaces the traditional parameter calibration process by constructing a virtual hydrological station network and data assimilation algorithm, and improves the calculation efficiency by more than 60%, successfully solving the "no data basin" modeling problem caused by the lack of monitoring stations in mountainous areas. In terms of factor weight determination, the weight uncertainty bottleneck caused by the dependence of traditional methods such as analytic hierarchy process (AHP) on expert experience is broken through, and a "mechanism-oriented-data-driven" double-loop factor screening system is established: first, 16 key inducements such as slope, topographic wetness index (TWI) and vegetation coverage are systematically screened from 20 candidate factors; then the contribution of each factor is quantified through random forest iterative training, realizing the paradigm shift from "empirical qualitative" to "data quantitative" in weight allocation. Experiments show that this method reduces the uncertainty of factor weight by 74%, and the spatial matching degree of risk zoning and historical disaster points is improved to 85.71%. This data-driven and deep integration of physical mechanism technology path not only significantly improves the generalization ability and objectivity of the model, but also provides a replicable technical solution for flood risk management in complex mountainous areas without data.

[0161] Although embodiments of the application have been shown and described, it is to be understood that various equivalents, modifications, substitutions and alternatives can be made to these embodiments without departing from the principles and spirit of the application, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A high-resolution mountain flood susceptibility mapping method based on a hybrid model, characterized in that, Includes the following steps: Step S1: Calibrate the Top-SSF model based on physical processes in multiple mountainous watersheds. The Top-SSF model incorporates canopy interception, soil infiltration, and stormwater runoff hydrological processes, and optimizes model parameters using the SCE-UA algorithm. Calibration is performed using historical flood data, specifically historical flash flood event data. The objective function is the Kling-Gupta efficiency coefficient, with the Nash-Sutcliffe coefficient and peak flow relative error as the parameters. Peak time error As an evaluation indicator; Step S2: Calculate the peak flow of the preset time period; Construct Gumbel extreme value distribution models in multiple mountainous watersheds using historical rainfall data to calculate the rainfall intensity and value of the watershed during the preset time period. Combine ERA5-Land reanalysis data to obtain the meteorological parameters of the same period, input them into the calibrated Top-SSF model, simulate the flood process lines of the preset time period in multiple watersheds, and extract the peak flow value of each watershed. Step S3: Watershed flood inventory map creation; Flow direction analysis is performed using DEM data, and the topographic humidity index (HAND) is calculated by tracing the elevation difference from the slope grid cell to the nearest river drainage point; Combined with the peak flow of the watershed output by the Top-SSF model, the water depth of the river channel cross section is inverted using the Manning formula, and potential inundation areas are identified by spatial comparison of HAND values ​​and river channel water depth, and the inundation depth is calculated pixel by pixel to form a watershed flood spatial distribution inventory map; Step S4: Screening key flood-inducing factors; 20 flood-inducing factors were identified through preliminary screening, followed by variance inflation factor analysis to exclude variables with significant multicollinearity, and insignificant factors were removed by combining autocorrelation and significance tests. Step S5: Construct a flood susceptibility model based on Random Forest (RF); Based on flood inventory data and the selected key flood-inducing factors, construct a random forest model, use 70% of the data for training and 30% for testing, and independently validate it using receiver operating characteristic (ROC) curves and statistical indicators. Step S6: High-resolution flood susceptibility mapping in mountainous areas: Using spatial matching data of each grid cell in the mountainous watershed with 16 key flood-inducing factors, a trained random forest classification model is applied to predict flood susceptibility. The model outputs a probability value F(X) to quantify the flood susceptibility of each cell. Independent validation is carried out by combining 21 typical flash flood events recorded in the flood and drought disaster bulletin. Through confusion matrix analysis and receiver operating characteristic curve evaluation, the model's spatial identification accuracy and prediction reliability for high-resolution flood-sensitive areas are verified.

2. The high-resolution mountain flood susceptibility mapping method based on a hybrid model as described in claim 1, characterized in that, Step S2 includes: S21: Estimation of rainfall events in the preset time period: Historical rainfall data of the study area is obtained from the meteorological station, and the extreme value distribution method is used to perform frequency analysis on the rainfall data to calculate the rainfall intensity and amount in the preset time period; S22: Meteorological data preparation, extracting meteorological data corresponding to the preset time period from ERA5-Land reanalysis data, including temperature, relative humidity, wind speed, and solar radiation; S23: Input the rainfall events and their corresponding meteorological data for the preset time period into the calibrated Top-SSF model to simulate the flood process lines for the preset time period in multiple watersheds and extract the peak flow.

3. The high-resolution mountain flood susceptibility mapping method based on a hybrid model as described in claim 1, characterized in that, Step S3 includes: S31: Calculate the flow direction of each grid cell based on the DEM; S32: Identify the nearest drainage point. Mark all slope grid cells flowing towards the same river grid cell, identify the nearest drainage point, and for each grid cell, trace along the flow direction to the nearest river grid cell and record its elevation value. ; S33: HAND value calculation: The HAND value is obtained by subtracting the elevation of the nearest drainage point from the elevation of the slope grid cells. in, For HAND values; Elevation of the slope grid cell; The elevation of the nearest drainage point; S34: Calculation of inundation range, first calculate the river water level based on the river flow; Peak flow extracted from the Top-SSF model By combining the geometry of the river cross-section and hydraulic relationships, the Manning formula is used to estimate the river depth; in, Peak traffic, This is the Manning coefficient. The cross-sectional area of ​​the river channel; The hydraulic radius; River slope; S35: Flood zone determined, HAND value With river water level By comparing and determining the potential flooding area, a grid cell is considered flooded if it meets the following conditions: S36: Water depth calculation. For the flooded area, calculate the water depth d for each grid cell. in, The water is deep.

4. The high-resolution mountain flood susceptibility mapping method based on a hybrid model as described in claim 1, characterized in that, Step S4 includes: S41: Preliminary screening identified 20 flood-inducing factors, including topography, vegetation, and meteorological factors; S42: VIF analysis, which excludes factors with high multicollinearity and retains factors that have a significant impact on floods; S43: Correlation analysis, analyzing the autocorrelation between factors and the correlation with floods, and eliminating statistically insignificant factors.

5. The high-resolution mountain flood susceptibility mapping method based on a hybrid model as described in claim 1, characterized in that, Step S5 includes: S51: Data preparation: Random sampling was performed from flood inventory maps of 80 watersheds to generate 320,000 data points, of which 70% were used for training and 30% for testing. Each data point contained 16 key flood-inducing factors. S52: Train an RF binary classification model to distinguish between flooded and non-flooded points; S53: Performance evaluation, evaluate model performance on the test set, and validate it using receiver operating characteristic curves and statistical indicators, specificity, and accuracy; in, For real examples, the number of actual flood points that were correctly predicted as flood points; A true negative example is the number of points that are actually non-flood points but were correctly predicted as non-flood points. These are false positives, representing the number of points that are actually not flood points but were incorrectly predicted as flood points. These are false negatives, representing the number of points that are actually flood points but were incorrectly predicted as non-flood points.

6. The high-resolution mountain flood susceptibility mapping method based on a hybrid model as described in claim 1, characterized in that, Step S6 includes: S61: Input data and match all grid cells within the study area with key flood-inducing factors; S62: Model prediction. The trained RF model is used to classify and predict each grid cell to determine whether it belongs to a flood-prone area. The probability value F(X) output by the RF model indicates whether the grid cell belongs to a flood-prone area. S63: Independent validation, using 21 flash flood events from the flood and drought disaster bulletin to independently validate the accuracy of the model's predictions.

7. The high-resolution mountain flood susceptibility mapping method based on a hybrid model as described in claim 1, characterized in that, In step S1, KGE is used as the objective function, and the Nash-Sutcliffe coefficient is used to measure the relative error of peak flow. Peak time error As an evaluation indicator, the calculation is as follows: set up The objective function is: in, The Pearson correlation coefficient; This is the deviation term, i.e., the ratio of the means; It is a variable term, namely the coefficient of variation ratio; Set upper and lower limits and initial values ​​for model parameters; set the number of iterations, maximum number of evolutionary loops, allowed percentage change per loop, and number of evolutionary steps for each complex before shuffling in the optimization algorithm; run the optimization algorithm and save the optimal parameters; use collected mountainous watershed flood events to... , and As evaluation metrics, assess model performance; in, For a certain Measured flood value at any given time; For a certain Simulated flood value at any given time; This is the average value of the measured flood. This represents the relative error of the flood peak. To measure the peak flow rate; Simulated peak flow; This refers to the peak occurrence time error; This refers to the measured peak time; To simulate peak occurrence time.

8. The high-resolution mountain flood susceptibility mapping method based on a hybrid model as described in claim 2, characterized in that, Step S21, extreme value distribution method, probability density function: in, It is a random variable; It is a location parameter, reflecting the central location of the distribution; It is a scale parameter that reflects the degree of dispersion of the distribution. Let be the probability density function. It is a natural constant; Cumulative distribution function: Quantile function: in, This is a non-transcendental probability.

9. The high-resolution mountain flood susceptibility mapping method based on a hybrid model as described in claim 3, characterized in that, In step S31, the flow direction of each grid cell is calculated based on the DEM, as follows: The direction of water flow in each grid cell points to the lowest elevation point among its eight adjacent grid cells; in, For the first Elevation of adjacent grid cells, in meters; The direction of flow for the grid cells.

10. A high-resolution mountain flood susceptibility mapping method based on a hybrid model as described in claim 5, characterized in that, Step S52 trains the RF binary classification model to distinguish between flood points and non-flood points, and the calculation is as follows: Predict the output label Y based on the input features X, where... : Represents 16 key flood-inducing factors for each data point; : indicates a binary category label, Indicates the flood point, Indicates non-flood points; The classification rules for decision trees are generated by partitioning the feature space. Let's assume the classification rule for the t-th decision tree is: That is, if the input feature X satisfies the splitting condition of the tree, it is predicted as a flood point. Otherwise, it is predicted to be a non-flood point. ; Random forest ensemble rules: Random forests integrate the results of multiple decision trees and use majority voting to make the final classification. in, The total number of decision trees in the random forest. Let F(X) be the classification result of the t-th decision tree for input X, and let F(X) be the final classification result of the random forest model.

Citation Information

Patent Citations

  • Mountain torrent disaster susceptibility evaluation method based on GIS and ensemble learning

    CN114493245A

  • Mountain torrent and debris flow disaster risk identification method and device, terminal and medium

    CN117057508A