Method for determining debris flow susceptible area driven by multiple initiation mechanisms in cold and cold mountainous area
By combining the debris flow formation mechanism and the geographical characteristics of high-altitude mountainous areas, and using a hybrid machine learning model to process data, the prediction accuracy problem of multiple types of debris flow areas in high-altitude mountainous areas is solved, and high-precision determination of debris flow prone areas is achieved.
Patent Information
- Application Number
- CN202510455602.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
AI Technical Summary
When traditional methods are applied separately to multiple types of mudslide areas, the prediction accuracy is limited, especially when different types of mudslides coexist in high-altitude mountainous areas, it is difficult to achieve accurate determination of mudslide prone areas.
Combining the mudslide formation mechanism and the geographical characteristics of high-altitude mountainous areas, the disaster-causing factors are selected, and data is processed using the ArcGIS platform. The support vector machine model, random forest model, and extreme gradient enhancement model are coupled with Alibaba and the top 40 thieves hyperparameter optimization algorithm, a hybrid machine learning model is constructed, iterative calculation is performed, and the prone spatial distribution map of different types of mudslides is fused to determine the prone area of mudslideslides.
High-precision prediction of multiple types of mudslide areas is realized. Through multi-model coupling to explore the relationship between different types of data, the problem of limited prediction accuracy in multiple types of mudslide areas is solved, and the mudslide prone mapping is provided under the multiple trigger mechanism.
Smart Images

Figure CN120296582A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of debris flow assessment, and particularly to a method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas. Background Art
[0002] Rainstorm-induced debris flow is abbreviated as RTDF, glacier debris flow is abbreviated as GDF, and glacial lake outburst debris flow is abbreviated as GLODF. Support vector machine is abbreviated as SVC, random forest is abbreviated as RF, and neural network is abbreviated as CNN. Debris flow is a unique large-scale surface movement in mountainous areas under the interaction of hydrological, geological, and geomorphic factors. Debris flow can usually be triggered by different water sources and can be divided into three main types: RTDF, GDF, and GLODF. RTDF is a type of debris flow triggered by heavy precipitation, which will erode and entrain loose materials such as rockfalls and landslides on both sides of the gully, and is the most widely distributed. Heavy rainfall usually causes soil saturation and slope instability, thus generating debris flow. They usually occur in mid- and low-altitude areas, with a high occurrence frequency, and can quickly change the surface morphology and have a wide range of influence. GDF develops in the marginal areas of alpine glaciers or snow cover, with moraine soil, ice avalanche deposits as the main solid material supply source, and forms debris flow under the excitation of snowmelt water, ice avalanche melt water, etc., usually with a large scale. The loose materials accumulated by glaciers and avalanches slide rapidly under the push of snowmelt water, with strong destructive power. LODF is a debris flow triggered by glacial lake outburst. After the glacial lake outburst, the released flood continuously erodes and entrains the moraine deposits left by glacial retreat in the gully and the loose materials generated by unstable slope collapses, and then transforms into debris flow. The formation process of GLODF is relatively complex and has a low occurrence frequency, but they have the characteristics of high potential energy, high explosiveness, suddenness, and long propagation distance. Due to the huge influence of glacial lake outburst, GLODF often causes serious disasters. The triggering mechanisms of the above three types of debris flows are different, but in the same river basin in alpine mountainous areas, the three types of debris flows may occur simultaneously, bringing greater catastrophic consequences. The determination of debris flow prone areas is a key step in disaster prevention and mitigation, and debris flow susceptibility mapping is determined to be an effective way for disaster risk management. Currently, many scholars have drawn debris flow prone areas around the world from the formation mechanism of debris flow using different methods, including the expert knowledge method, physical method, and data-driven method. The expert knowledge method relies on field investigations and understanding of the debris flow formation mechanism, which is simple and easy to operate, but has strong subjectivity. The physical method can evaluate debris flow prone areas through numerical simulation, with high accuracy, but requires complex input parameters and is difficult to apply to large areas. The data-driven method fits the relationship between disaster-causing factors and debris flow through disaster data, which is relatively objective and applicable to large areas, but has high requirements for the accuracy of disaster data.
[0003] However, when the above method is applied alone to multi-type debris flow areas, the prediction accuracy is limited, especially ignoring the coexistence of different types of debris flows. With the progress of artificial intelligence technology, machine learning methods such as SVC, RF, and CNN have been widely used in debris flow susceptibility mapping, which can more effectively explore the non-linear relationship between disasters and disaster-causing factors. However, traditional machine learning methods are difficult to discover the hidden relationships between data, limiting the further improvement of classification accuracy. Summary of the Invention
[0004] To overcome the deficiencies of the prior art, the purpose of the present invention is to provide a method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas, so as to solve the problem that the prediction accuracy of traditional methods is limited when applied alone to multi-type debris flow areas.
[0005] To achieve the above purpose, the present invention provides the following solutions:
[0006] A method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas, comprising:
[0007] Select disaster-causing factors in combination with debris flow formation mechanisms and geographical characteristics of alpine mountainous areas; the geographical characteristics of alpine mountainous areas include: topographic and geomorphic characteristics, geological characteristics, land cover characteristics, meteorological characteristics, and glacier and ice lake characteristics;
[0008] Collect debris flow data of the study area, project the debris flow data onto WGS_1984_UTM using the ArcGIS platform and resample it at a spatial resolution of 30 meters to obtain mapped data;
[0009] Perform bilinear interpolation on the continuous data in the mapped data and nearest neighbor interpolation on the categorical data in the mapped data to obtain interpolated data;
[0010] Use ArcGIS to calculate the disaster-causing factors of the interpolated data to obtain index statistical data;
[0011] Use collinearity analysis to eliminate the multi-collinear data in the index statistical data to obtain preprocessed data;
[0012] Individually couple a support vector machine model, a random forest model, an extreme gradient boosting model with the Alibaba and the Forty Thieves hyperparameter optimization algorithm to obtain a hybrid machine learning model;
[0013] Use the pre-trained hybrid machine learning model to perform iterative calculations on the preprocessed data to obtain RTDF susceptibility spatial distribution maps, GDF susceptibility spatial distribution maps, and GLODF susceptibility spatial distribution maps;
[0014] Fuse the RTDF susceptibility spatial distribution map, the GDF susceptibility spatial distribution map, and the GLODF susceptibility spatial distribution map to obtain a multi-type debris flow probability map, and determine the debris flow prone area based on the multi-type debris flow probability map.
[0015] Preferably, the disaster-causing factors include: channel gradient, channel connectivity, excess volume, glacier slope, dam slope, rock strength index, fault distance, loose material volume, normalized difference vegetation index, number of days with daily precipitation greater than 25 mm within a year, cumulative maximum precipitation for three consecutive days within a year, daily maximum temperature, annual accumulated temperature, overall glacier loss rate, overall glacier volume, overall glacier area, upstream glacier area, upstream glacier volume, upstream glacier loss rate, upstream catchment area of the ice lake, moraine volume, ice lake volume, and ice lake area.
[0016] Preferably, the elimination criterion for the collinearity analysis is that the Pearson correlation coefficient is greater than 0.8.
[0017] Preferably, the calculation expression for the channel connectivity is:
[0018]
[0019] where IC k is the channel connectivity of the kth one; D up,k is the uphill component of the kth one; D dn,k is the downhill component of the kth one; is the average weight of the uphill contribution area; S is the average weight of the uphill slope; A1 is the uphill area; d e is the river length extracted according to the downstream steepest downhill flow direction of the eth one; W e is the weight value of the eth grid point; S e is the slope value of the eth grid point; RI is the surface roughness; RI MAX is the maximum value of the surface roughness.
[0020] Preferably, the calculation expression for the excess volume is:
[0021] V E =Z E ×A2;
[0022] where Z E =Z - Z gully +S t ×dis; V E represents the excess volume; Z E is the remaining height; A2 is the bottom area; Z is the true elevation; Z gully is the true height of the channel; S tis the threshold angle; dis is the distance from the river course.
[0023] Preferably, the calculation expression of the Alibaba and the Forty Thieves hyperparameter optimization algorithm includes:
[0024] x i = l j + r × (u j - l j ),
[0025]
[0026] x' t+1 = Td t [(u j - l j ) + l j ; r3 ≥ 0.5, r4 ≤ Pp t And
[0027]
[0028] where x i is the position of the i-th thief; l j is the lower limit of the j-th dimension; r is a basic random number uniformly distributed between 0 and 1; u j is the upper limit of the j-th dimension; is the position of the i-th thief in the (t + 1)-th iteration; gbest t represents the optimal position reached by all thieves in the t-th iteration; Td t is the tracking distance of all thieves in the t-th iteration; is the optimal position reached by the i-th thief in the t-th iteration; represents the relative position between Alibaba and the i-th thief in the t-th iteration; r1, r2, r3, r4, rand are the first random number, the second random number, the third random number, the fourth random number, and the fifth random number uniformly distributed between 0 and 1, respectively; is the intelligence level at which Marjaneh deceives the i-th thief in the t-th iteration; x' t+1 is the new position of the thief.
[0029] Preferably, the expression of the optimization objective of the support vector machine model is:
[0030]
[0031] where w is the parameter vector; b is the intercept; x DFI is the disaster-causing factor of the debris flow sample; N is the total number of debris flow samples.
[0032] Preferably, the expression of the loss function of the support vector machine model is:
[0033]
[0034] Among them, L(w, b, a) is the calculated value of the loss function; a DFI is the DFI-th Lagrange multiplier; y DFI is the debris flow sample identifier (0 or 1); N is the total number of debris flow samples.
[0035] Preferably, the expression of the decision function of the support vector machine model is:
[0036] f(x test ) = sign(α DFI y DFI K + b);
[0037] Among them, f(x test ) is the calculated value of the decision function; α DFI is the DFI-th Lagrange multiplier; K is the kernel function.
[0038] The present invention discloses the following technical effects:
[0039] The present invention provides a method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine regions. Through multi-model coupling, it solves the problem of limited prediction accuracy when traditional methods are applied alone in multi-type debris flow regions, and realizes the mining of relationships between different types of data; by selecting disaster-causing factors in combination with the debris flow formation mechanism and the geographical characteristics of alpine regions, it solves the problem of incomplete debris flow formation mechanism in traditional methods, and realizes debris flow susceptibility mapping under multiple triggering mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 is a schematic flow chart of determining debris flow prone areas driven by multiple triggering mechanisms in alpine regions provided by an embodiment of the present invention;
[0042] Figure 2 is a schematic diagram of the SVC principle provided by an embodiment of the present invention;
[0043] Figure 3 is a schematic diagram of the RF principle provided by an embodiment of the present invention;
[0044] Figure 4Distribution map of multiple types of debris flows in the mountainous area of the research region provided by the embodiment of the present invention;
[0045] Figure 5 PPC diagram of candidate indicators for multiple types of debris flows provided by the embodiment of the present invention, Figure 5 (a) PPC diagram of candidate indicators for RTDF type debris flows, Figure 5 (b) PPC diagram of candidate indicators for GDF type debris flows, Figure 5 (c) PPC diagram of candidate indicators for GLODF type debris flows;
[0046] Figure 6 Susceptibility map of multiple types of debris flows provided by the embodiment of the present invention;
[0047] Figure 7 Provided by the embodiment of the present invention, Figure 7 (a) Schematic diagram of the ROC curve for evaluating the susceptibility of RTDF by three models, Figure 7 (b) Schematic diagram of the ROC curve for evaluating the susceptibility of GDF by three models, Figure 7 (c) Schematic diagram of the ROC curve for evaluating the susceptibility of GLODF by three models, Figure 7 (d) Schematic diagram of the relative importance of susceptibility evaluation indicators for RTDF, Figure 7 (e) Schematic diagram of the relative importance of susceptibility evaluation indicators for GDF, Figure 7 (f) Schematic diagram of the relative importance of susceptibility evaluation indicators for GLODF. Detailed implementation manners
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0049] The purpose of the present invention is to provide a method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas, and solve the problem of limited prediction accuracy when traditional methods are applied alone to areas with multiple types of debris flows.
[0050] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0051] Figure 1 Schematic diagram of the process for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas provided by the embodiment of the present invention, as shown in Figure 1As shown in the figure, the present invention provides a method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas, including:
[0052] Step 100: Select hazard factors in combination with debris flow formation mechanisms and geographical characteristics of alpine mountainous areas; the geographical characteristics of alpine mountainous areas include: topographic and geomorphic characteristics, geological characteristics, land cover characteristics, meteorological characteristics, and glacier and ice lake characteristics;
[0053] Step 200: Collect debris flow data of the study area, project the debris flow data onto WGS_1984_UTM using the ArcGIS platform and resample it at a spatial resolution of 30 meters to obtain mapped data;
[0054] Step 300: Perform bilinear interpolation on continuous data in the mapped data and nearest neighbor interpolation on categorical data in the mapped data to obtain interpolated data;
[0055] Step 400: Calculate the hazard factors of the interpolated data using ArcGIS to obtain index statistical data;
[0056] Step 500: Use collinearity analysis to eliminate multicollinear data in the index statistical data to obtain preprocessed data;
[0057] Step 600: Individually couple the support vector machine model, random forest model, extreme gradient boosting model with the Alibaba and the Forty Thieves hyperparameter optimization algorithm to obtain a hybrid machine learning model;
[0058] Step 700: Use the hybrid machine learning model to perform iterative calculations on the preprocessed data to obtain the RTDF susceptibility spatial distribution map, GDF susceptibility spatial distribution map, and GLODF susceptibility spatial distribution map;
[0059] Step 800: Fuse the RTDF susceptibility spatial distribution map, GDF susceptibility spatial distribution map, and GLODF susceptibility spatial distribution map to obtain a multi-type debris flow probability map, and determine the debris flow prone areas based on the multi-type debris flow probability map.
[0060] Specifically, the hazard factors include: channel gradient, channel connectivity, excess volume, glacier slope, dam slope, rock strength index, fault distance, loose material volume, normalized difference vegetation index, number of days with daily precipitation greater than 25 mm within a year, cumulative maximum precipitation for three consecutive days within a year, daily maximum temperature, annual accumulated temperature (ACT), overall glacier loss rate, overall glacier volume, overall glacier area, upstream glacier area, upstream glacier volume, upstream glacier loss rate, upstream catchment area of ice lake, moraine volume (MV), ice lake volume, and ice lake area.
[0061] Optionally, the elimination criterion for collinearity analysis is that the Pearson correlation coefficient is greater than 0.8.
[0062] Specifically, the calculation expression for channel connectivity is:
[0063]
[0064] Where, IC k is the k-th channel connectivity; D up,k is the k-th uphill component; D dn,k is the k-th downhill component; is the average weight of the uphill contribution area; S is the average weight of the uphill slope; A1 is the uphill area; d e is the river length extracted according to the downstream steepest downhill flow direction for the e-th; W e is the weight value of the e-th grid point; S e is the slope value of the e-th grid point; RI is the surface roughness; RI MAX is the maximum value of the surface roughness.
[0065] Preferably, the calculation expression for the excess volume is:
[0066] V E =Z E ×A2;
[0067] Where, Z E =Z - Z gully +S t ×dis; V E represents the excess volume; Z E is the remaining height; A2 is the bottom area; Z is the true elevation; Z gully is the true height of the channel; S t is the threshold angle; dis is the distance from the river channel.
[0068] Furthermore, the calculation expression of the Alibaba and the Forty Thieves hyperparameter optimization algorithm includes:
[0069] x i =l j +r × (u j -l j );
[0070]
[0071] x′ t+1 =Td t [(u j -l j ) + l j ; r3 ≥ 0.5, r4 ≤ Pp tand
[0072]
[0073] where x i is the position of the i-th thief; l j is the lower bound of the j-th dimension; r is a basic random number uniformly distributed between 0 and 1; u j is the upper bound of the j-th dimension; is the position of the i-th thief at the (t + 1)-th iteration; gbest t represents the optimal position reached by all thieves at the t-th iteration; Td t is the tracking distance of all thieves at the t-th iteration; is the optimal position reached by the i-th thief at the t-th iteration; represents the relative position between Alibaba and the i-th thief at the t-th iteration; r1, r2, r3, r4, rand are the first random number, the second random number, the third random number, the fourth random number, and the fifth random number uniformly distributed between 0 and 1 respectively; is the intelligence level at which Marjaneh deceives the i-th thief at the t-th iteration; x′ t+1 is the new position of the thief.
[0074] Specifically, the expression of the optimization objective of the support vector machine model is:
[0075]
[0076] where w is the parameter vector; b is the intercept; x DFI is the disaster-causing factor of the debris flow sample; N is the total number of debris flow samples.
[0077] Furthermore, the expression of the loss function of the support vector machine model is:
[0078]
[0079] where L(w, b, a) is the calculated value of the loss function; a DFI is the DFI-th Lagrange multiplier; y DFI is the debris flow sample identifier (0 or 1); N is the total number of debris flow samples.
[0080] Even further, the expression of the decision function of the support vector machine model is:
[0081] f(x test ) = sign(α DFI y DFI K + b);
[0082] where f(x test) is the calculated value of the decision function; α DFI is the DFI-th Lagrange multiplier; K is the kernel function.
[0083] Specifically, in this embodiment, considering the formation process of debris flow and the geographical characteristics of alpine regions, 23 indicators are selected (refer to Table 1). Among them, 9 indicators are used to evaluate the RTDF probability, 14 indicators are used to evaluate the GDF probability, and 18 indicators are used to evaluate the GLODF probability. These indicators can be divided into five categories: topography, meteorology, geology, glaciers and glacial lakes, and land cover. A brief description of each indicator is given below. To ensure consistency and comparability, all data are re-projected to WGS_1984_UTM using the ArcGIS platform. And the spatial resolution is unified to 30 meters through resampling, where bilinear interpolation is applied to continuous data and nearest neighbor interpolation is applied to categorical data. This method minimizes the differences caused by resolution and projection changes while maintaining the accuracy and integrity of the original dataset. To meet the input requirements of the machine learning model, the zonal statistics tool in ArcGIS is used to calculate the statistical values of each evaluation indicator for each catchment unit, such as the mean and maximum values.
[0084] Table 1
[0085]
[0086]
[0087] Furthermore, topography and geomorphology. The channel gradient (CG) is a characterization of the terrain situation of each catchment area, which can reflect the energy conditions for the formation of debris flow and can be calculated by dividing the elevation difference by the main channel length. The channel connectivity (IC) is used to characterize the potential connection between the watershed outlet and various parts of the watershed, that is, the probability of debris flow spreading to the gully mouth, and can be calculated by using the ArcGIS modeling tool. The surface roughness can be used as a weighting factor for the channel connectivity. The calculation formula of IC is as follows:
[0088]
[0089] The surface roughness is calculated by the standard deviation of all grid points within a radius of 3 grid points around each grid point.
[0090] Specifically, the materials of debris flow are sourced from loose accumulations along the way, such as landslides and collapses (erosion and entrainment). The loose accumulations in the gully are mainly sourced from the unstable slopes on both sides of the landslides. However, it is difficult to obtain the landslide volume over a large area through means such as remote sensing image interpretation. Therefore, the excess volume is selected to characterize the potential source volume (unstable rock mass volume) along the downstream. The excess volume (EV) has been proven to have a close correlation with the distribution of the study area, the Hengduan Mountains, and landslides. The excess volume, that is, the volume of the rock mass on the hillslope exceeding the threshold, is calculated by the following formula:
[0091] V E =Z E ×A2
[0092] Z E (x,y)=Z(x,y)-Z gully +S t ×dis(x,y)
[0093] Among them, (x,y) represents the position of a certain point on the gully hillslope.
[0094] Furthermore, the glacier slope (GS) refers to the average slope of the glacier-covered area and can reflect the stability of the ice mass. The greater the glacier slope, the greater the likelihood of ice avalanches and ice landslides, and the greater the likelihood of triggering glacier debris flows and glacial lake outburst debris flows. The dam slope (DS) is an important indicator reflecting the stability of the glacial lake dam and is obtained by calculating the average slope in the potential area of the dam. This indicator is used to evaluate the stability of glacial lakes in the High Asia and the Qinghai-Tibet Plateau.
[0095] Specifically, meteorology. Although current precipitation controls the occurrence of debris flows directly, the precipitation in the early stage plays an important role in the formation of debris flows. Therefore, the effective cumulative precipitation is a key factor in evaluating the susceptibility of debris flows and can be simply parameterized by the maximum cumulative precipitation over three consecutive days within a year (AMCP3D). At the same time, the number of days with ≥25mm (AND25) is selected to represent the intensity and frequency of precipitation. Temperature controls glacier behavior. Extreme high temperatures (parameterized by the daily maximum temperature (DMT)) may cause a large amount of loss of glaciers in a short time, generating a large amount of glacial meltwater and ice avalanches / landslides, promoting the formation of glacier debris flows. If there is a glacial lake below the glacier, it may also trigger glacial lake outburst debris flows.
[0096] Furthermore, geology. Rocks provide the material basis for debris flows. Rock strength represents the potential for rock strain softening and potential liquefaction ability. According to rock hardness, lithology is divided into five categories: extremely soft, soft, medium, hard, and extremely hard. Using spatial statistical tools, the average rock strength of each catchment is calculated as the Rock Strength Index (RSI). Frequent earthquakes may occur near fault zones, which not only cut rock masses to form joints and fractures and produce loose materials for debris flows, but may also trigger landslides and avalanches, thus leading to GLODF. The distance from the fault (DFF) is used to quantify the impact of tectonic movement on debris flows, and the average distance between each catchment and the fault is extracted through spatial analysis tools.
[0097] Specifically, glacial ice lakes. The reserves of glaciers and the rate of glacier loss in a basin jointly determine the volume of glacial meltwater, providing water source conditions for the formation of glacial debris flows. It can be characterized by the overall glacier mass loss rate (GMLR), glacier volume (GV), and glacier area (GA). Glacial activities upstream may produce solids and fluids flowing into the lake, triggering outburst floods from ice lakes. Therefore, the upstream glacier mass loss rate (UGMLR), upstream glacier area (UGA), and upstream glacier volume (UGV) are incorporated into the index system for the susceptibility of outburst floods from ice lakes. The upstream watershed area of the ice lake (UWA) is also an index to quantify the susceptibility of fluids formed by precipitation and snowmelt flowing into the lake. The ice lake volume (GLV) and ice lake area (GLA) are directly related to the storage capacity of the ice lake and are the main water source conditions for the formation of outburst floods from ice lakes. The ice lake volume is calculated through the area-volume empirical formula, which is fitted based on the measured data of 25 ice lakes in the mountainous area of the study area.
[0098] Furthermore, surface cover. Debris flows continuously erode and entrain loose materials during the flow process. Therefore, the volume of loose materials (LMV) is a key indicator for the formation of mudflows. The Normalized Difference Vegetation Index (NDVI) is used to describe the vegetation density based on remote sensing images.
[0099] Preferably, collinearity analysis and parameter standardization. Collinearity analysis is a necessary step in parameter preprocessing, which can diagnose multicollinearity in the dataset. Debris flow indicators with strong correlation, that is, those with multicollinearity, may contain redundant information, leading to model instability and possible deviation in the output results. In this embodiment, the Pearson correlation coefficient (PPC) is calculated to diagnose the correlation between indicators. Indicator pairs with a PPC greater than 0.8 are strong correlation indicator pairs, and only one of them will be retained.
[0100] Specifically, a hybrid machine learning model. Machine learning methods belong to data-driven methods and are more objective than traditional methods. They are widely used to fit the relationship between triggering factors and debris flows. Using machine learning methods such as Support Vector Machine (SVM), Random Forest (RF), and Extreme Gradient Boosting (XGB), a hybrid machine learning model is constructed by coupling the hyperparameter optimization algorithm of Alibaba and the Forty Thieves to conduct multi-type debris flow susceptibility mapping. The machine learning models used in this embodiment are as follows:
[0101] 1) Support Vector Machine
[0102] Reference Figure 2 , B1 is hyperplane 1, and b11 is the maximum separation gap of hyperplane B1; B2 is hyperplane 2, and b21 is the maximum separation gap of hyperplane B2. Support Vector Machine (SVM) is a supervised learning model, including Support Vector Classification (SVC), Linear Classification Support Vector Machine (Linear SVC), and Nu Support Vector Machine (Nu-SVC), etc., and has been widely used in the assessment of debris flows. In this embodiment, the Support Vector Classification (SVC) model is selected to fit the relationship between glacial lake outburst debris flows and samples. Based on the principle of structural risk minimization, Support Vector Classification classifies the training set by solving the hyperplane with the largest separation gap in the vector space. The larger the interval, the smaller the total error of the classifier. The optimal decision boundary is solved by minimizing the loss function. Taking binary classification as an example below, the principle of SVC is briefly introduced:
[0103] Suppose the training set is N debris flow samples (x DFI , y DFI ), and each sample has n features (x 1DFI , x 2DFI , …, x nDFI ). w·x DFI + b = 0 represents the decision boundary. The SVC loss function can be expressed as:
[0104]
[0105] s.t.y DFI (w·x DFI + b) ≥ 1, DFI = 1, 2, …, N
[0106] Using the Lagrange multiplier, the loss function is rewritten in a form considering the constraint conditions:
[0107]
[0108] Under the Lagrange dual function, the values of w and b can be deduced, and the dual function is expressed as follows:
[0109]
[0110] Furthermore, the optimal solution of the objective function can be obtained, and the final decision function can be obtained:
[0111] f(x test ) = sign(α DFI y DFI K + b)
[0112] where α DFI y DFI K represents the weighted contribution of the support vector to the new sample.
[0113] Preferably, in order to find the linear decision boundary existing therein when dealing with non-linear data, it is often necessary to project the data in the original data space into a high-dimensional space. However, this projection operation may cause a problem of huge computational amount. To solve this problem, the kernel trick is usually used to achieve efficient calculation. In SVC, there are four optional common kernel functions. Since the Gaussian radial basis kernel function (rbf) has the characteristics of fast calculation speed and high accuracy, the rbf kernel function is selected in this embodiment.
[0114] 2) Random Forest
[0115] Referring to Figure 3 , the random forest (Random Forest) is a typical bagging algorithm in ensemble learning, consisting of a series of decision trees, which is suitable for the study of debris flow susceptibility. The basic principle of the random forest is to randomly extract some samples from the original sample set, build decision trees (D1,..., D k ) according to the decision rules, and then construct the random forest. The final result is determined by voting. The random forest overcomes overfitting by generating a group of classifiers, improving the prediction accuracy and stability.
[0116] 3) Extreme Gradient Boosting (XGB)
[0117] XGB is an effective algorithm that enhances the gradient boosting decision tree method. By using a technique called tree boosting, XGB significantly improves the model performance. The core concept of XGB is to gradually construct a series of continuous decision trees, and each tree aims to correct the errors of the previous tree. This method usually improves the model accuracy while reducing the risk of overfitting.
[0118] 4) Alibaba and the Forty Thieves Hyperparameter Optimization Algorithm
[0119] Ali Baba and the Forty Thieves (AFT) is a new meta-heuristic algorithm for solving global optimization problems. The algorithm simulates the social behavior of humans looking for other humans. In order to improve the convergence speed and fitting accuracy of the machine learning algorithm, this embodiment uses the Alibaba and the Forty Thieves hyperparameter optimization algorithm to achieve automatic optimization of hyperparameters when SVC, RF and XGB are used to predict the susceptibility of debris flows, avoiding overfitting and subjectivity. For example, the parameters of SVC include kernel function, penalty parameter (C), Gaussian kernel function parameter (gamma), etc., the parameters of RF include evaluation criteria (criterion), maximum depth (max_depth), the number of trees (n_estimators), etc., and the parameters of XGB include learning rate (learning_rate), maximum depth (max_depth), the number of trees (n_estimators), etc. When the forty thieves are looking for Alibaba, there are the following assumptions:
[0120] ① Forty thieves formed a group and got guidance from someone or one of the thieves to find Alibaba's house. This information may or may not be correct.
[0121] ②The forty thieves will travel a distance from the initial point until they find Alibaba’s house.
[0122] ③Marjaneh could deceive the thieves many times and use shrewd methods to protect Alibaba in a certain way, making them immune to a certain percentage of attacks.
[0123] The above search behavior can be mathematically modeled as follows:
[0124] The AFT algorithm first randomly initializes the positions of n thieves. The positions of the thieves are:
[0125] x i = l j +r×(u j -l j )
[0126] Optionally, model validation. In this embodiment, the ROC curve and the confusion matrix are selected to evaluate the accuracy and performance of the debris flow disaster prediction model. The ROC depicts the function between the true positive rate and the false positive rate and is widely used to judge the performance of classifiers. The larger the area under the ROC curve, the higher the accuracy of the model. The confusion matrix (CM) is another useful method to measure the prediction accuracy of the model, where the columns and rows represent the results of the predicted classes and the actual classes respectively. The CM describes a series of performance metrics of algorithms, including True Positive Class, Recall, True Negative Class, and Accuracy. The relevant calculation formulas are as follows:
[0127]
[0128] TP and FN refer to the number of samples correctly predicted by the model; TN and FP refer to the number of samples mispredicted by the model.
[0129] Specifically, case analysis. In a mountainous area of the study area, the regional area is about 66.22×104 km 2 , which belongs to the junction of the monsoon and plateau climate zones. Affected by the activities of two plates, multiple fault zones such as the Main Frontal Thrust (MFT), Main Boundary Thrust (MBT), and Main Central Thrust (MCT) have developed in the study area. The exposed rock strata in this mountain range are rich and diverse, including Precambrian metamorphic rocks, Ordovician to Neoproterozoic marine strata, and Phanerozoic sedimentary strata. Affected by factors such as strong uplift, glacial activity, and river incision, the study area has become a region with drastic topographic changes, with high mountains and deep valleys, and the altitude ranges from 100 m in the south to 8848 m in the middle. The annual average temperature in the study area is higher in the south and lower in the north, with a range of about -24 to 30 °C. The precipitation range is about 128 to 3000 mm, showing a decreasing trend from south to north. This mountain range has 20,436 glaciers (with a total area of about 19,899.22 km 2 ) and 8,204 glacial lakes (635.13 km 2 ). In short, the study area provides an ideal environment for predicting the susceptibility of debris flows driven by multiple mechanisms. The complex topography, active geological structure, and hydrological conditions in this area are conducive to the formation of various types of debris flows.
[0130] Furthermore, debris flow samples. A multi-type debris flow inventory is a prerequisite for debris flow susceptibility mapping and is crucial for fitting the relationship between disasters and inducing factors. The spatial locations of various types of debris flows are identified through documented historical disasters, field surveys, and remote sensing interpretation of Google Maps. Finally, 1043 debris flows in this mountainous area are identified and mapped, including 538 RTDFs, 454 GDFs, and 51 GLODFs. Each type of debris flow has a different mechanism, and separate studies are required for the spatial prediction of each type of debris flow. The samples of each type of debris flow are randomly divided into two parts, with 70% of the samples used for model training and the remaining samples used to verify the model accuracy. Watershed units are more suitable for predicting the occurrence of debris flows because these disaster processes, such as heavy precipitation, glacial activity, glacial lake outburst, debris flow dynamic movement, erosion entrainment, and deposition along the debris flow path, occur within the watershed units. Therefore, watershed units are used as mapping units for evaluating the probability of debris flow occurrence. Using the hydrological tools in the ArcGIS platform and a 30m digital elevation model (DEM), the study area is divided into 12,031 watersheds.
[0131] Specifically, correlation analysis. The parameter correlation matrix reveals pairs of indicators that show strong correlations between different types of debris flows. The specific pairs are as follows (refer to Figure 5 (a) to Figure 5 (c)): AND25 and AMCP3D in RTDF (PPC is 0.88); GA and GMLR (-0.80) and GA and GV (0.95) in GDF; UGA and UGMLR (-0.93), UGA and UGV (0.80), UGA and UWA (0.86), UWA and UGMLR (-0.82), and GLA and GLV (0.87) in GLODF. To eliminate multicollinearity in the indicator system, AND25 in RTDF, GA in GDF, and UGA, UWA, and GLA in GLODF are deleted. The remaining indicators are normalized to improve calculation efficiency and model stability.
[0132] Furthermore, multi-type debris flow susceptibility. In this embodiment, the RTDF, GDF, and GLODF probability maps generated by the most accurate model among the three hybrid machine learning models are selected and combined to create a comprehensive multi-type debris flow probability map. To consider a conservative case, in this embodiment, catchment areas with probability levels of VL and L are classified as not prone to debris flows, while catchment areas with probability levels of M, H, and VH are considered prone to debris flows. Therefore, the probability maps of each debris flow type are integrated to generate a multi-type debris flow probability map, as shown in Figure 6As shown. The analysis shows that 37.84% of the catchments are not prone to debris flows. On the contrary, 53.21% of the catchments are prone to single types of debris flows: 50.12% are RTDF, 2.87% are GDF (mainly in the northern part of the western mountain range), and 0.22% are GLODF (mainly along the ridges of the central and eastern mountain ranges). In addition, 0.75% of the catchments are vulnerable to at least two types of debris flows: GDF and GLODF, RTDF and GLODF, and RTDF and GDF. In addition, 1.73% of the catchments are vulnerable to all three types of debris flows, mainly located in the central and southern slopes of the mountain range.
[0133] Optionally, model performance comparison. In this case, according to the different initiation processes of different types of debris flows, a probability index system driven by a multiple-trigger mechanism was constructed. The ATF hyperparameter optimization algorithm was combined with three classic machine learning algorithms, and the constructed probability index system was used to generate probability maps of single-type and multi-type debris flows in the Himalayan region. This case solved the key problem of evaluating the probability of debris flows driven by a multiple-trigger mechanism in areas with diverse climate types and high altitudes. Compared with the regional assessment only for single-type debris flows, the proposed method significantly improved the accuracy of debris flow probability assessment by incorporating multiple debris flow types, thus providing valuable support for regional disaster prevention and mitigation.
[0134] Preferably, this case adopted three novel hybrid machine learning models to evaluate the probabilities of RTDF, GDF, and GLODF. Various performance indicators were used to evaluate the effectiveness of these models, and the results are shown in Table 2 and Figure 7 (a) to Figure 7 (c). Compared with ATF-RF and ATF-XGB, the ATF-SVC model performed well in evaluating the probabilities of RTDF and GLODF, with an AUC greater than 0.95 and lower MSE and RMSE values. On the contrary, the ATF-RF model was proven to be more suitable for evaluating the probability of GDF because its ACC and AUC values were higher than those of other models, and its MSE and RMSE were lower.
[0135] Table 2
[0136]
[0137]
[0138] Furthermore, this embodiment applied SHAP (SHapley Additive exPlanations) to quantify the relative importance of the indicators affecting the formation of various types of debris flows (refer to Figure 7 (d) to Figure 7(f)). The results show that RSI, AMCP3D, and IC are the main contributing factors to the formation of RTDF. Specifically, RSI reflects the provenance conditions of RTDF. Vulnerable rocks within the catchment area facilitate the transformation of floods into debris flows, thus promoting the formation of debris flows. AMCP3D represents the triggering conditions of RTDF, and continuous heavy rainfall is a key factor. IC characterizes the channel conditions necessary for the development of debris flows. For GDF, the most important factors are DMT, IC, GMLR, NDVI, and RSI. Similar to RTDF, IC and RSI represent the channel and source conditions respectively. In addition, DMT and GMLR are closely related to glacial activities. Extreme temperatures and rapid glacial retreat generate glacial meltwater, providing the necessary water source for the formation of GDF. For GLODF, the key factors include UGMLR, UGV, DS, GLV, and MV. UGMLR represents the intensity of glacial activities upstream of the glacial lake. Intense activities may trigger large-scale mass movements and subsequent glacial lake outburst floods. UGV represents the upstream glacial volume, indirectly reflecting the contribution of glaciers to the lake. DS and GLV affect the stability of the glacial lake, while MV represents the source conditions for the formation of GLODF.
[0139] The beneficial effects of the present invention are as follows:
[0140] Through multi-model coupling, the present invention realizes the mining of relationships between different types of data and improves the classification accuracy of the model; by combining the debris flow formation mechanism and the geographical characteristics of alpine regions to select disaster-causing factors, it realizes debris flow susceptibility mapping under multiple triggering mechanisms and improves the reliability of the model.
[0141] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0142] In this embodiment, specific examples are used to elaborate on the principle and implementation manner of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas, characterized in that, Including: Select hazard factors by combining debris flow formation mechanism and geographical characteristics of alpine regions; The geographical characteristics of the alpine regions include: topographic and geomorphic characteristics, geological characteristics, land cover characteristics, meteorological characteristics, and glacier and ice lake characteristics; Collect debris flow data of the study area, project the debris flow data onto WGS_1984_UTM using the ArcGIS platform and resample it at a spatial resolution of 30 meters to obtain mapped data; Perform bilinear interpolation on the continuous data in the mapped data and nearest neighbor interpolation on the categorical data in the mapped data to obtain interpolated data; Use ArcGIS to calculate the hazard factors of the interpolated data to obtain index statistical data; Use collinearity analysis to eliminate the multicollinear data in the index statistical data to obtain preprocessed data; Individually couple the support vector machine model, random forest model, extreme gradient boosting model with the Alibaba and the Forty Thieves hyperparameter optimization algorithm to obtain a hybrid machine learning model; Use the pre-trained hybrid machine learning model to perform iterative calculations on the preprocessed data to obtain RTDF susceptibility spatial distribution map, GDF susceptibility spatial distribution map, and GLODF susceptibility spatial distribution map; Fuse the RTDF susceptibility spatial distribution map, the GDF susceptibility spatial distribution map, and the GLODF susceptibility spatial distribution map to obtain a multi-type debris flow probability map, and determine the debris flow prone areas based on the multi-type debris flow probability map.
2. The method for determining debris flow-prone areas driven by multiple triggering mechanisms in alpine mountainous areas according to claim 1, wherein The hazard factors include: channel gradient, channel connectivity, excess volume, glacier slope, dam slope, rock strength index, fault distance, loose material volume, normalized difference vegetation index, number of days with daily precipitation greater than 25 mm in a year, cumulative maximum precipitation for three consecutive days in a year, daily maximum temperature, annual accumulated temperature, overall glacier loss rate, overall glacier volume, overall glacier area, upstream glacier area, upstream glacier volume, upstream glacier loss rate, upstream catchment area of ice lake, moraine volume, ice lake volume, and ice lake area.
3. A method for determining debris flow-prone areas driven by multiple triggering mechanisms in alpine mountainous areas according to claim 1, characterized in that, The elimination criterion for the collinearity analysis is that the Pearson correlation coefficient is greater than 0.
8.
4. The method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas according to claim 1, wherein The calculation expression for the channel connectivity is: Among them, IC k is the k-th channel connectivity; D up,k is the k-th uphill component; D dn,k is the k-th downhill component; is the average weight of the uphill contribution area; is the average weight of the uphill slope; A1 is the uphill area; d e is the river length of the e-th extracted according to the downstream steepest downhill flow direction; W e is the weight value of the e-th grid point; S e is the slope value of the e-th grid point; RI is the surface roughness; RI MAX is the maximum value of the surface roughness.
5. A method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas according to claim 1, characterized in that The calculation expression for the excess volume is: V E = Z E × A2; wherein, Z E = Z - Z gully + S t × dis; V E represents the excess volume; Z E is the remaining height; A2 is the bottom area; Z is the true elevation; Z gully is the true height of the channel; S t is the threshold angle; dis is the distance from the river channel.
6. The method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas according to claim 1, characterized in that The calculation expression of the Alibaba and the Forty Thieves hyperparameter optimization algorithm includes: x i = l j + r × (u j - l j )、 x t ' +1 = Td t [(u j - l j ) + l j ; r3 ≥ 0.5, r4 ≤ Pp t and r3<0.5; Among them, x i is the position of the i-th thief; l j is the lower limit of the j-th dimension; r is a basic random number uniformly distributed between 0 and 1; u j is the upper limit of the j-th dimension; is the position of the i-th thief in the (t + 1)-th iteration; gbest t represents the optimal position reached by all thieves in the t-th iteration; Td t is the tracking distance of all thieves in the t-th iteration; is the optimal position reached by the i-th thief in the t-th iteration; represents the relative position between Alibaba and the i-th thief in the t-th iteration; r1, r2, r3, r4, rand are the first random number, the second random number, the third random number, the fourth random number, and the fifth random number uniformly distributed between 0 and 1 respectively; is the intelligence level at which Marjaneh deceives the i-th thief in the t-th iteration; x t ' +1 is the new position of the thief.
7. A method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas according to claim 1, characterized in that, The expression of the optimization objective of the support vector machine model is: where, w is the parameter vector; b is the intercept; x DFI is the disaster-causing factor of the debris flow sample; N is the total number of debris flow samples.
8. A method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas according to claim 7, characterized in that, The expression of the loss function of the support vector machine model is: Among them, L(w, b, a) is the calculated value of the loss function; a DFI is the DFI-th Lagrange multiplier; y DFI is the debris flow sample identifier (0 or 1); N is the total number of debris flow samples.
9. A method for determining debris flow prone areas driven by multiple triggering mechanisms in alpine mountainous areas according to claim 8, characterized in that, The expression of the decision function of the support vector machine model is: f(x test ) = sign(α DFI y DFI K + b); where f(x test ) is the calculated value of the decision function; α DFI is the DFI-th Lagrange multiplier; K is the kernel function.