Landslide risk assessment method based on feature screening and differential evolution algorithm optimization

CN115482138BActive Publication Date: 2026-08-28FUZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211161265.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-22
Publication Date
2026-08-28
Estimated Expiration
2042-09-22

AI Technical Summary

Technical Problem

由于机器学习模型最优超参数难以搜索确定,在实际危险性评估中难以真正发挥效用,模型构建时一般使用默认参数或通过网格搜索获取较优值,部分研究者也尝试使用了遗传算法与粒子群算法[21-22]等进化算法进行超参数的率定并取得一定效果,但总体上优化效率较低

Benefits of technology

[0014]本发明主要应用于区域滑坡等类似灾情的危险性评估研究,从而获得更加准确的灾害评估结果。本发明在构建因子敏感性指数开展敏感性分析的基础上,结合重要性分析、相关性分析、共线性分析构建4D特征筛选法用于评估因子的综合定量筛选,并引入DE算法对推广能力较强的SVM与MLP的主要超参数进行全局搜索优化,具有两个优点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115482138B_ABST
    Figure CN115482138B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of landslide risk assessment method based on feature screening and differential evolution algorithm optimization.Based on the quantitative sensitivity analysis of constructing factor sensitivity index, four-dimensional feature screening method is used for evaluating factor comprehensive optimization by combining importance analysis, correlation analysis and collinearity analysis;To overcome the problem that model is difficult to optimize, differential evolution algorithm is introduced to optimize two kinds of machine learning models with stronger generalization ability, such as support vector machine and multilayer perceptron.The case evaluation results show that: four-dimensional feature screening method can more objectively and comprehensively select higher suitable risk assessment factors, thereby reducing data dimension, reducing information redundancy to improve the performance of evaluation model;Differential evolution algorithm has significant optimization effect on support vector machine and multilayer perceptron, which is beneficial to enhance the accuracy of model landslide risk assessment.The present application has important significance for objective selection of influencing factors and machine learning model optimization in landslide risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a landslide hazard assessment method based on feature selection and differential evolution algorithm optimization. Background Technology

[0002] Driven by both dramatic global climate change and accelerated economic development, the situation regarding disaster governance and risk management is becoming increasingly severe. Disaster hazard assessment research is not only a crucial foundation for disaster risk assessment but also a necessary reference for regional development. In the context of the first national comprehensive survey of natural disaster risks, a comprehensive assessment of the hazard of sudden geological disasters can better measure disaster risks, thereby providing more objective and accurate information support for planning and decision-making.

[0003] The essence of the study on the risk assessment of sudden geological disasters is to measure the degree of danger of a specific disaster. The process of implementation is generally to combine typical influencing factors to construct an assessment index system and perform a composite calculation of the disaster risk index [1]. Since the geographical environment, geological characteristics and social conditions of different regions are different, even the same type of disaster will have different influencing factors in different regions. Therefore, the premise of effectively carrying out disaster risk assessment is to effectively select influencing factors. According to the source of influence, risk assessment factors usually include three categories: geographical influencing factors, geological influencing factors such as fault structure and stratum lithology, and social influencing factors such as transportation roads and land use. Geographical influencing factors can be further divided into micro-topographic factors such as slope, aspect and curvature, macro-topographic factors such as topographic relief and surface roughness, hydrological environmental factors such as topographic humidity index and runoff intensity index, and natural geographical factors such as rainfall distribution and vegetation cover [1-4]. Currently, most of the influencing factors used in risk assessment are selected based on existing research results or experience. Although some studies have screened assessment factors by combining sensitivity analysis[5] or importance analysis[6] with correlation or collinearity analysis, there is still a lack of objective and comprehensive assessment factor selection methods in actual risk assessment.

[0004] At present, the commonly used models for landslide hazard assessment at home and abroad are mainly divided into three categories: empirical assessment models, statistical analysis models, and machine learning models [7-8]. Empirical assessment models such as the analytic hierarchy process [9] are simple in principle and easy to calculate, but they are highly subjective. Statistical analysis models such as the information content method

[10] can overcome the influence of factor quantification and error in the assessment process to a certain extent, but they are also difficult to reveal the relationship between each assessment factor and landslide disaster in high-dimensional space. Machine learning models are widely used in the field of sudden geological disaster research due to their excellent ability to handle high-dimensional nonlinear problems. Commonly used models include logistic regression model

[11] , random forest

[12] , decision tree

[13] , support vector machine

[14] , artificial neural network [15-17], and composite model [8, 18] or other complex model [19-20]. Since the optimal hyperparameters of machine learning models are difficult to search and determine, they are difficult to truly play a role in actual risk assessment. When building models, default parameters are generally used or better values ​​are obtained through grid search. Some researchers have also tried to use evolutionary algorithms such as genetic algorithms and particle swarm algorithms [21-22] to calibrate hyperparameters and have achieved certain results, but the overall optimization efficiency is low. Compared with evolutionary algorithms such as genetic algorithms and particle swarm algorithms, differential evolutionary algorithms have the advantages of fewer control parameters, higher optimization efficiency, better parallel performance and faster convergence speed

[23] . They have received widespread attention from scholars at home and abroad and have been widely used in signal processing, satellite communication, image processing and other fields

[24] . However, they are rarely used in landslide risk assessment research. Summary of the Invention

[0005] The purpose of this invention is to provide a landslide hazard assessment method based on feature selection and differential evolution algorithm optimization. A relatively comprehensive assessment factor database is initially constructed by combining regional disaster background and research results. Addressing the difficulty in selecting assessment factors, a Factor Sensitivity Index (FSI) method is proposed for quantitative sensitivity assessment. A four-dimensional (4D) feature selection method is constructed by combining importance analysis, correlation analysis, and collinearity analysis to select assessment factors with stronger suitability and effectiveness for model construction. To address the difficulty in optimizing machine learning models, a Differential Evolution (DE) algorithm is introduced to optimize two assessment models with strong generalization capabilities: Support Vector Machine (SVM) and Multilayer Perceptron (MLP). The optimized assessment models are evaluated, and the zoning effect and assessment accuracy of the models are analyzed to provide decision support for disaster governance and risk management.

[0006] To achieve the above objectives, the technical solution of the present invention is: a landslide hazard assessment method based on feature selection and differential evolution algorithm optimization, comprising the following steps:

[0007] Step S1: Landslide Background Analysis and Preliminary Selection of Assessment Factors

[0008] (1) Analysis of regional disaster background; (2) Preliminary selection of regional disaster impact factors to be screened; (3) Data acquisition and spatial consistency processing of impact factors; (4) Construction of spatial database for regional disaster risk assessment factors;

[0009] Step S2: Construction of a four-dimensional feature screening method for evaluating factor selection:

[0010] (1) Status classification and determination coefficient calculation of each preliminary selection factor; (2) Sensitivity analysis of the sensitivity index of the preliminary selection factors; (3) Importance analysis of the preliminary selection factors based on the integrated algorithm of bagging and boosting strategies; (4) Correlation analysis of each preliminary selection factor; (5) Collinearity analysis of each preliminary selection factor; (6) Comprehensive optimization of risk assessment factors.

[0011] Step S3: Landslide hazard assessment based on differential evolution algorithm optimization:

[0012] (1) Normalization of landslide hazard assessment factors; (2) Extraction of non-landslide samples based on unsupervised clustering; (3) Global search optimization of main hyperparameters based on DE algorithm; (4) Construction of landslide hazard assessment model; (5) Calculation of regional landslide normalized hazard index; (6) Mapping of regional landslide hazard.

[0013] Compared with the prior art, the present invention has the following beneficial effects:

[0014] This invention is primarily applied to hazard assessment research of regional landslides and similar disasters, thereby obtaining more accurate disaster assessment results. Based on sensitivity analysis using a constructed factor sensitivity index, this invention combines importance analysis, correlation analysis, and collinearity analysis to construct a 4D feature screening method for comprehensive quantitative screening of assessment factors. Furthermore, it introduces the DE algorithm to globally search and optimize the main hyperparameters of SVM and MLP, which have strong generalization capabilities. This has two advantages:

[0015] (1) The 4D feature screening method, which follows the principle of “ensuring sensitivity, retaining importance, eliminating correlation and avoiding collinearity”, can select risk assessment factors with higher suitability more objectively and comprehensively, thereby reducing feature dimensions and information redundancy to improve the performance of the assessment model.

[0016] (2) The DE algorithm has a significant optimization effect on SVM and MLP, which is beneficial to improving the accuracy of landslide hazard assessment. The DE algorithm can obtain better hyperparameters from global search. The AUC values ​​of the SVM and MLP models optimized by the DE algorithm are significantly improved compared with the evaluation accuracy of the unoptimized models. Attached Figure Description

[0017] Figure 1 This invention presents the technical approach for a landslide hazard assessment method based on feature selection and differential evolution algorithm optimization.

[0018] Figure 2 The implementation process of the differential evolution algorithm. Detailed Implementation

[0019] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.

[0020] This invention discloses a landslide hazard assessment method based on feature selection and differential evolution algorithm optimization, which mainly includes three parts: landslide background analysis and preliminary selection of assessment factors, comprehensive optimization of assessment factors, and construction of a machine learning-based landslide hazard assessment model. The specific technical route is as follows: Figure 1 As shown.

[0021] 1. Landslide Background Analysis and Preliminary Selection of Assessment Factors

[0022] Based on the background of regional landslide disasters and existing research findings, and considering the ease of data acquisition, relevant factors affecting regional landslides were initially selected from three categories of influencing factors: geographical influencing factors (such as micro-topographic factors like slope, macro-topographic factors like topographic relief, hydrological environmental factors like topographic humidity index, and natural geographical factors like rainfall), geological influencing factors (such as lithology and faults), and social influencing factors (such as land use type and roads). A regional landslide hazard assessment factor database was then constructed based on the initially selected factors.

[0023] 2. Construction of a four-dimensional feature screening method for evaluating factor selection

[0024] To address the difficulty in selecting evaluation factors, an FSI was constructed and calculated based on the initial selection of evaluation factors for factor sensitivity analysis. Combining factor importance analysis, correlation and collinearity analysis, a four-dimensional feature screening method was constructed according to the principles of "ensuring sensitivity, retaining importance, eliminating correlation, and avoiding collinearity". The evaluation factors for the final input model were comprehensively selected based on the quantitative analysis results of the four dimensions.

[0025] 2.1 Quantitative Analysis Based on Factor Sensitivity Index

[0026] There are currently three main quantitative analysis models for assessing the suitability of factors in disaster research: ① using the combined value of the quantitative index corresponding to each state level of the factor [5]; ② using the average value of the quantitative index corresponding to each state level of the factor

[25] ; ③ using the range of the quantitative index corresponding to each state level of the factor

[26] . Using the combined value and the average value to represent the influence intensity of the factor on the disaster will weaken the influence of the category with a large impact on the disaster in the factor state level and strengthen the influence of the category with a small impact on the disaster, resulting in the positive and negative effects canceling each other out, weakening the actual influence of the factor on the disaster, while the range can better reveal the overall intensity of the factor's influence on the disaster.

[0027] Therefore, based on existing research, this paper collectively refers to the calculation results based on quantitative indices such as Frequency Ratio (FR), Information Value (I), and Certainty Factor (CF) as the Determination Coefficient (DC), and proposes to use the Factor Sensitivity Index (FSI) to reflect the overall influence of a certain assessment factor on disaster occurrence. The calculation method is as follows:

[0028] ESI i =DC (i,max) -DC (i,min) (1)

[0029] In the formula: i represents the i-th evaluation factor; FSI i DC represents the sensitivity index of the i-th evaluation factor; (i,max) DC represents the maximum value of the coefficient of determination among all state levels of the i-th evaluation factor; (i,min) This represents the minimum value of the determination coefficient among all state levels of the i-th evaluation factor. Furthermore, a judgment threshold μ is introduced in the calculation; when FSI... i When the threshold value is greater than μ, the assessment factor i is considered to have a strong influence on landslide occurrence and is suitable for inclusion in the study. The judgment threshold in large-scale studies is generally set to 1; however, based on the study scale and regional characteristics, the judgment threshold μ = 0.5 is set here.

[0030] To avoid the randomness of a single algorithm, this paper calculates the frequency ratio coefficient, information index, and determinism coefficient for each factor at each state level based on the state classification results. The average of the three determinism coefficients is taken as the final sensitivity index for each factor. Quantitative sensitivity analysis is then conducted based on the calculation results.

[0031] 2.2 Factor Feature Importance Analysis

[0032] Importance analysis is the process of learning from an effective dataset to obtain the relative importance of each input feature [6, 27]. To avoid the randomness of a single model, a random forest with a bagging ensemble strategy

[13] and a gradient boosting tree with a boosting ensemble strategy

[28] are used together to conduct factor importance analysis. Finally, the feature importance representation values ​​are weighted according to the classification accuracy to obtain the composite importance feature value of each factor.

[0033] The feature importance value of an evaluation factor is a specific numerical value. Currently, there is no clear standard to measure the importance of a factor; its contribution to the model results can only be judged based on the relative magnitude of this value. Therefore, this paper sets a feature importance threshold d to remove factors that contribute too little to the model. Let d = 1 / 2n (n is the number of evaluation factors). When the feature importance value > d, the corresponding factor is considered to contribute significantly to the model results and should be retained; when the feature importance value < d, the corresponding factor is considered to contribute relatively little to the model results and should be removed.

[0034] 2.3 Factor Correlation and Collinearity Analysis

[0035] In order to reduce factor correlation and improve model stability and accuracy, and considering that collinearity cannot be ruled out when factor correlation is low, this paper combines correlation[5] and collinearity[6] analysis results for comprehensive consideration. This paper uses Pearson correlation coefficient to analyze the degree of correlation between factors. When the absolute value of the correlation coefficient between two factors is greater than 0.40, it is considered that there is obvious correlation and they should be removed. In the collinearity analysis, when TOL is less than 0.10 and VIF is greater than 10, it is considered that the factor is seriously collinear and should be removed[6].

[0036] 3. Landslide hazard assessment based on differential evolution algorithm optimization

[0037] To eliminate the influence of dimensions, the hazard assessment factors obtained from the final screening were normalized: discrete data were processed using one-hot encoding, and continuous data were processed using linear normalization. In the final model construction, a Self-Organizing Map (SOM) neural network and a Gaussian Mixed Model (GMM) were first used to perform unsupervised landslide hazard clustering on all grid cells in the study area. Then, negative samples were extracted from the common areas of the clusters with the lowest landslide hazard, and these were combined with known landslide samples to construct a machine learning dataset. Finally, the DE algorithm was used to globally search and optimize the main hyperparameters of two representative models, SVM and MLP, and the obtained hyperparameter optimization results were substituted into the model for landslide hazard assessment.

[0038] 3.1 Support Vector Machine (SVM)

[0039] Support Vector Machine [7, 29] is a classic machine learning algorithm based on structural risk minimization and VC dimension theory. It has unique advantages in the nonlinear high-dimensional pattern recognition problem of landslide hazard assessment, and has better assessment and generalization capabilities than other traditional methods.

[0040] When there is a set of samples X in the landslide dataset M i (i = 1, 2, ..., n), X i Let y represent an input vector containing i landslide assessment factors. i =1 indicates a landslide has occurred, y i =-1 indicates no landslide occurs, and n is the number of samples in the landslide dataset. When the problem is binary classification, the ultimate goal of SVM is to find an optimal hyperplane to separate the training data into two classes. The hyperplane equation can then be expressed as:

[0041] wx+b=0 (2)

[0042] In the formula: w represents the normal vector; x represents a point on the hyperplane; b represents a constant. When w and b are optimal, the found optimal hyperplane will maximize the distance between positive and negative samples. In nonlinear problems, kernel functions, slack variables, and penalty factors are also introduced, and the formula for finding the optimal hyperplane is as follows:

[0043]

[0044] In the formula: W represents the weight vector that determines the hyperplane direction, used to determine the direction of the segmentation plane; b is a constant term representing the displacement; h represents the number of support vector points; ξ i Let represent the slack variable, indicating the distance of the sample point from the outlier; C (C > 0) represents the penalty parameter, indicating the degree of penalty for misclassified samples. Equation (3) is subject to:

[0045] y i (w T x+b)≥1-ξ i (4)

[0046] After being constrained by Lagrange multipliers, the optimization problem is transformed into a quadratic programming algorithm for solution.

[0047] 3.2 Multilayer Perceptron (MLP)

[0048] Multilayer perceptron [30-31] is a multilayer feedforward network model that propagates errors in one direction. It is one of the most widely used and fundamentally studied network models. MLP can identify different datasets in the entire dataset without requiring pre-existing knowledge or experience, nor does it require pre-existing statistical models to train the data. It can effectively solve binary classification problems of nonlinear data, and therefore is suitable for landslide hazard assessment. MLP consists of an input layer, a hidden layer, and an output layer composed of similar neurons. The input layer neurons and the hidden layer neurons form an association matrix of the processed objects through connections. The connections between the hidden layer neurons and the output layer neurons form a decision matrix of the processed objects. The three layers of neurons are connected by certain weights to form a stable network structure with decision-making capabilities.

[0049] Many studies have shown that a model with only one hidden layer can usually achieve good results. However, when there are too many hidden layers, problems such as excessive computational consumption and overfitting are likely to occur. Therefore, this paper fixes the number of hidden layers to one when constructing the MLP and focuses on optimizing the number of hidden layer nodes. The empirical value of the number of hidden layer neurons in a neural network can be calculated according to the following formula

[32] :

[0050]

[0051] In the formula: S is the number of hidden layer neurons; m is the number of input layer neurons; n is the number of output layer neurons; a is an adjustment constant from 1 to 10.

[0052] 3.3 Differential Evolution Algorithm

[0053] Differential Evolutionary Algorithm [23, 33] is a heuristic, efficient, random search, globally parallel optimization algorithm based on population differences. It has excellent characteristics such as simple principle, fewer parameters, faster computation, and good robustness, and is an important branch of evolutionary algorithm research. There are few application cases of the DE algorithm in landslide hazard assessment. Therefore, this paper introduces the DE algorithm to optimize machine learning models.

[0054] The novel population generation scheme of the DE algorithm differs from other evolutionary algorithms. Essentially, it is a greedy genetic algorithm based on real-number encoding with a merit-preserving strategy. The DE / rand / 1 / bin mutation strategy ensures that mutated individuals are generated from three distinct random individuals, exhibiting strong global search capabilities and thus better maintaining population diversity. Let N be the number of individuals used in the optimization iteration process. p D-dimensional vectors For each generation, g = 0, 1, 2, ..., G max The population, with a total number of iterations, is set to G. max If every individual in the population is a candidate solution, then the DE algorithm implementation process is as follows (see...). Figure 2 ):

[0055] (1) Population initialization. When the initial generation number g = 0, let x max,j and x min,j To find the upper and lower bounds of the j-th dimension (j = 1, 2, ..., D) in the solution space, the rand() function is used to generate random numbers in the range [0, 1]. The initial population generation expression is then:

[0056]

[0057] (2) Mutation operation. For individuals in the g-th generation population. Randomly select 3 different individuals (r1, r2, r3 ∈ [1, N) p If r1, r2, r3 ≠ i, then the expression for the newly generated mutated individual is:

[0058]

[0059] In the formula: F is the variation factor (0≤F≤2), which controls the difference vector. For individuals The impact.

[0060] (3) Crossover operation. Parent individual and mutated individuals The crossover operation is performed by equation (8). This operation method ensures population diversity in the generation of experimental individuals. At the same time, ensure that at least one variable is taken from

[0061]

[0062] In the formula: the rand(1, n) function is used to generate random integers in the range [1, n], thereby ensuring At least for Provide a decision variable to maintain the diversity distribution of the population; the crossover factor CR (0≤F≤1) affects the variation of individuals. Replace the target individual The probability plays a decisive role. The larger the crossover factor, the faster the convergence rate; the smaller the crossover factor, the better the population diversity is maintained.

[0063] (4) Selection operation. A "greedy" strategy is adopted, starting from... and Individuals with better fitness are selected as the next generation of the population. The specific operational equation is as follows:

[0064]

[0065] Application Cases

[0066] This invention takes landslides in Fujian Province as an example to verify the effectiveness of a landslide hazard assessment method based on feature screening and differential evolution algorithm optimization.

[0067] (I) Construction and Implementation of Four-Dimensional Feature Filtering Method

[0068] (1) Preliminary selection and processing of evaluation factors

[0069] A background analysis of landslide disasters in Fujian Province was conducted, and 21 hazard assessment factors were initially selected (see Table 1). The state classification scheme for the influencing factors is as follows: For discrete factors, continuous values ​​were directly reassigned for differentiation; for continuous factors, the distance to roads and rivers were divided into 6 levels based on 400, 800, 1200, 1600, and 2000 m, the distance to fault structures were divided into 6 levels based on 600, 1200, 1800, 2400, and 3000 m, the slope aspect was divided into 9 categories based on flat land and 8 slope aspect orientations, and the surface curvature was divided into 3 levels based on <0, =0, and >0. For the remaining factors, continuous assessment factors were sampled based on known landslide sample points, and the natural discontinuity method was used to classify the factors into 9 categories based on the sampling results, so that the state classification results are more consistent with the landslide distribution characteristics in the study area.

[0070] To facilitate calculation and analysis, the study area was divided into 30m×30m grid units based on the area size and data conditions, resulting in a total of 134,288,470 evaluation units.

[0071] Table 1. Preliminary Selection Results of Risk Assessment Factors

[0072]

[0073] (2) Calculation and analysis of sensitivity index of evaluation factors

[0074] The sensitivity index calculation results for each preliminary evaluation factor are shown in Table 2. Table 2 shows that land use type has the greatest impact on landslides, while topographic curvature has the least impact. The average sensitivity index of six factors—profile curvature, planar curvature, slope variability, aspect variability, fault structure, and topographic curvature—is below the judgment threshold of 0.5, indicating a relatively small impact on landslides in Fujian Province. The average sensitivity index of land use type and surface incision depth is greater than 2, indicating a more significant impact on landslides in Fujian Province. Overall, macro-topographic factors have a significantly greater impact on landslides than micro-topographic factors, suggesting that macro-topographic features are more conducive to revealing the landslide background or influencing landslide formation.

[0075] Table 2. Calculation results of sensitivity analysis of risk assessment factors.

[0076]

[0077] (3) Importance analysis of evaluation factor characteristics

[0078] Based on the model construction, the composite importance eigenvalues ​​of each initially selected evaluation factor in this paper were calculated, as shown in Table 3. Table 3 shows that the evaluation factor with the highest eigenvalue is annual average rainfall, which is significantly higher than other factors. This indicates that rainfall has significant value in predicting landslide risk in Fujian Province and is consistent with actual disaster occurrences. With an eigenvalue threshold d of 0.02, only topographic curvature has excessively low importance.

[0079] Table 3. Calculation results of importance analysis of risk assessment factors.

[0080]

[0081] (4) Evaluation Factors Comprehensive Optimization and Result Analysis

[0082] After completing sensitivity and importance analyses, correlation and collinearity analyses are performed on the 21 initially selected evaluation factors. Finally, based on the construction principles of the 4D feature screening method and combined with the quantitative calculation results of sensitivity, importance, correlation, and collinearity, a comprehensive optimization of landslide influencing factors in Fujian Province is conducted. The screening process is as follows:

[0083] (1) Based on the results of sensitivity analysis and importance analysis, six factors were directly eliminated: profile curvature, plane curvature, slope variability, slope aspect variability, distance to fault structure, and topographic curvature.

[0084] (2) Rainfall has a significant impact on landslides in Fujian, and the annual average rainfall has the highest characteristic importance, so it is retained; considering the correlation of factors, the annual average normalized vegetation index and elevation are removed.

[0085] (3) The surface cutting depth is the most sensitive to revealing the characteristics of landslides in Fujian, so it is retained; considering the correlation and collinearity of factors, the three factors of slope, surface roughness and topographic relief are removed.

[0086] To verify the effectiveness of the 4D feature screening method, comparative experiments and analyses were conducted using a control group with 21 initially selected evaluation factors and 10 selected evaluation factors. Model accuracy was evaluated using ROC curves (Table 4). Table 4 shows that after removing some factors based on sensitivity analysis, importance analysis, and sensitivity and importance analysis, the AUC values ​​of the ROC curves for both SVM and MLP did not change significantly. This indicates that the factors selected by importance and sensitivity analysis are more typical, achieving almost consistent model accuracy with fewer factors. The evaluation accuracy of retaining average annual rainfall and surface incision depth is generally better than retaining other highly correlated factors, indicating that the evaluation factors retained according to the principles are relatively superior. Therefore, after removing the above 11 factors, the remaining factors all have high sensitivity and importance and are relatively independent. The final 10 suitability evaluation factors used for modeling are: aspect, elevation variation coefficient, land use type, average annual rainfall, surface incision depth, distance to river, distance to road, engineering geological rock group, topographic humidity index, and runoff intensity index.

[0087] Table 4. Comparison Calculation Results of Models Based on 4D Feature Selection Method

[0088]

[0089] (II) Optimization and Construction of Landslide Risk Assessment Model

[0090] The DE algorithm for hyperparameter optimization is implemented in Python. The population size is set to 20 and the maximum number of evolutions is set to 50. The other parameters are all default values. Finally, the landslide occurrence probability of all evaluated grid samples is used as the risk index (value is 0 to 1). They are sorted from high to low and divided into five zoning levels according to the area ratio of 1:2:4:2:1: extremely high landslide risk area, relatively high landslide risk area, medium landslide risk area, relatively low landslide risk area and extremely low landslide risk area [8] for landslide zoning research.

[0091] (1) Construction and optimization of support vector machine evaluation model

[0092] In the application of the method of this invention, the radial basis function is selected as the kernel function. The DE algorithm is used to optimize the main hyperparameters of SVM, the penalty factor C and gamma. The penalty factor C and gamma are set to be in the interval [0.1, 10]. The optimal hyperparameters are found by continuous search. Finally, the optimal hyperparameter combination is obtained as C = 10.00 and gamma = 0.2318, which is then substituted into SVM for optimization.

[0093] (2) Construction and optimization of multilayer perceptron evaluation model

[0094] In the application of the method of this invention, the number of hidden layer nodes S is set in the interval [6, 16] by the DE algorithm to find the optimal hyperparameters through discrete search, and the initial learning rate ILR is found in the interval [0.001, 0.1] by continuous search. The optimal hyperparameter combination obtained in the end is S = 14 and ILR = 0.0095, which is then substituted into MLP for optimization.

[0095] (3) Evaluation and analysis of model evaluation results

[0096] This invention calculates the classification accuracy and ROC values ​​of the ROC curves for the training set, test set, and complete set data (Table 5), and obtains the confusion matrix of each model (Table 6). Based on these results, a preliminary evaluation of the model accuracy is conducted. Table 5 shows that the classification accuracy of each model built on the three datasets is higher than 0.85 and is basically consistent, indicating that the model fitting effect is reasonable. The classification accuracy and AUC value of the DE-SVM built on the three datasets are significantly higher than those of SVM. The classification accuracy and AUC value of the DE-MLP built on the three datasets are basically consistent with those of MLP, indicating that the DE algorithm significantly improves the classification accuracy of SVM, while the improvement in classification performance of MLP is not significant. As shown in Table 6, the recall rates of both classes of samples in DE-SVM are higher than those of SVM, and the overall accuracy of DE-SVM is 4% higher than that of SVM. This further demonstrates that the DE algorithm is beneficial to improving the classification accuracy of SVM. The overall accuracy of DE-MLP is consistent with that of MLP, but the recall rate of landslide samples in DE-MLP is higher than that in MLP, indicating that MLP optimized by the DE algorithm can also identify landslides more effectively.

[0097] To better compare and evaluate the actual assessment improvement effects of the two landslide hazard assessment models, this invention mainly uses success rate curves to analyze the fitting degree between the hazard index calculated using training set, test set, and complete set data and all known landslide samples [34-36]. The results are shown in Table 7, and the results of the integrated power curve of the complete data are shown in Table 8. As shown in Table 7, the AUC values ​​of the integrated power curves of the three types of data for the same model are basically consistent, indicating that the assessment effect of the model constructed using the above datasets is relatively stable. The AUC values ​​of the success rate curves of SVM and MLP after optimization by the DE algorithm based on the three types of datasets are significantly improved. According to the calculation results of the success rate curve of the complete dataset, the AUC value of DE-SVM is 0.8208, which is 4.43% higher than that of SVM without DE algorithm optimization; the AUC value of DE-MLP is 0.7750, which is 4.37% higher than that of MLP without DE algorithm optimization. The model evaluation results show that the DE algorithm has strong global search capabilities, and using the DE algorithm to optimize SVM and MLP can significantly improve the accuracy of landslide hazard assessment.

[0098] Table 5. Classification accuracy and ROC curve calculation results for various datasets.

[0099]

[0100] Table 6 shows the calculation results of the confusion matrix for each model.

[0101]

[0102]

[0103] Table 7 Calculation results of power curves for various data integration methods

[0104]

[0105] Table 8 Success Rate Curves and AUC Values ​​of Landslide Risk Assessment Model in Fujian Province

[0106]

[0107] References:

[0108] [1] Zhang HJ, Zhao Z, Chen JH, et al. A deep one-dimensional convolutional neural network method for landslide risk assessment: A case study in Lushan, Sichuan, China[J]. Journal of Natural Disasters, 2021, 30(3): 191-198. DOI: 10.13577 / j.jnd.2021.0321

[0109] [2]Dou J, Yunus AP, Bui DT, et al. Improved landslide assessment using support vector machine with bagging, boosting, and stacking ensemble machine learning framework in a mountainous watershed, Japan[J]. Landslides, 2020, 17(3): 641-658. DOI: 10.1007 / s10346-019-01286-5

[0110] [3]Wang Y, Fang ZC, Hong H Y. Comparison of convolutional neural networks for landslide susceptibility mapping in Yanshan County, China[J]. The Science of the Total Environment, 2019, 666: 975-993. DOI: 10.1016 / j.scitotenv.2019.02.263

[0111] [4] Yang C, Lin GF, Zhang MF, et al. Soillandslide susceptibility assessment based on DEM[J]. Journal of Geo-Information Science, 2016, 18(12): 1624-1633. DOI: 10.3724 / SP.J.1047.2016.01624

[0112] [5] Niu QF, Feng ZB, Dang XH, et al. Suitability analysis of topographic factors in loess landslide research[J]. Journal of Geo-Information Science, 2017, 19(12): 1584-1592. DOI: 10.3724 / SP.J.1047.2017.01584

[0113] [6] Wang Y, Fang ZC, Niu RQ, et al. Landslide susceptibility analysis based on deep learning[J]. Journal of Geo-Information Science, 2021, 23(12): 2244-2260. DOI: 10.12082 / dqxxkx.2021.210057

[0114] [7] Li YY, Mei HB, Ren XJ, et al. Geological disaster susceptibility evaluation based on certainty factor and support vector machine[J]. Journal of Geo-Information Science, 2018, 20(12): 1699-1709. DOI: 10.12082 / dqxxkx.2018.180349

[0115] [8] Huang FM, Yin KL, Jiang SH, et al. Landslide susceptibility assessment based on clustering analysis and support vector machine[J]. Chinese Journal of Rock Mechanics and Engineering, 2018, 37(1): 156-167. DOI: 10.13722 / j.cnki.jrme.2017.0824

[0116] [9] Huang JL, Zhao JG, Zhang T, et al. AHP-based hazard degree assessment of high-speed landslide of reservoir bank[J]. Journal of Natural Disasters, 2011, 20(5): 95-99. DOI: 10.13577 / j.jnd.2011.0515

[0117]

[10] Zhang ZY, Deng MG, Xu SG, et al. Comparison of landslide susceptibility assessment models in Zhenkang County, Yunnan Province, China[J]. Chinese Journal of Rock Mechanics and Engineering, 2022, 41(1): 157-171. DOI: 10.13722 / j.cnki.jrme.2021.0360

[0118]

[11] Aditian A,Kubota T,Shinohara Y.Comparison of GIS-based landslidesusceptibility models using frequency ratio,logistic regression,andartificial neural network in a tertiary region of Ambon, Indonesia[J].Geomorphology,2018,318:101-111.DOI:10.1016 / j.geomorph.2018.06.006

[0119]

[12] Liang Z,Wang C M,Zhang Z M,et al.A comparison of statistical andmachine learning methods for debris flow susceptibility mapping[J].StochasticEnvironmental Research and Risk Assessment,2020,34(11):1887-1907.DOI:10.1007 / s00477-020-01851-8

[0120]

[13] Chang Z L,Du Z,Zhang F,et al.Landslide susceptibility predictionbased on remote sensing images and GIS:Comparisons of supervised andunsupervised machine learning models[J].Remote Sensing,2020,12(3):502.DOI:10.3390 / rs12030502

[0121]

[14] U, M, Z,et al.Machine learning based landslideassessment of the Belgrade metropolitan area:Pixel resolution effects and across-scaling concept[J].Engineering Geology,2019,256:23-38.DOI:10.1016 / j.enggeo.2019.05.007

[0122]

[15] Wang Y,Fang Z C,Wang M,et al.Comparative study of landslidesusceptibility mapping with different recurrent neural networks[J].Computers&Geosciences,2020,138:104445. DOI:10.1016 / j.cageo.2020.104445

[0123]

[16] Bragagnolo L,da Silva R V,Grzybowski J M V.Landslidesusceptibility mapping with r.landslide:A free open-source GIS-integratedtool based on Artificial Neural Networks[J].Environmental Modelling&Software,2020,123:104565.DOI:10.1016 / j.envsoft.2019.104565

[0124]

[17] Thi Ngo P T,Panahi M,Khosravi K,et al.Evaluation of deep learningalgorithms for national scale landslide susceptibility mapping of Iran[J].Geoscience Frontiers,2021,12(2):505-519. DOI:10.1016 / j.gsf.2020.06.013

[0125]

[18] Zhu L,Huang L H,Fan L Y,et al.Landslide susceptibility predictionmodeling based on remote sensing and a novel deep learning algorithm of acascade-parallel recurrent neural network[J].Sensors,2020,20(6):1576.DOI:10.3390 / s20061576

[0126]

[19] Huang F M,Zhang J,Zhou C B,et al.A deep learning algorithm usinga fully connected sparse autoencoder neural network for landslidesuseeptibility prediction[J].Landslides, 2020,17(1):217-229.DOI:10.1007 / s10346-019-01274-9

[0127]

[20] Huang F M,Yin K L,Zhang G R,et al.Landslide displacementprediction using discrete wavelet transform and extreme learning machinebased on chaos theory[J].Environmental Earth Sciences,2016,75(20):1-18.DOI:10.1007 / s12665-016-6133-0

[0128]

[21] Hu AL, Wang KW, Li JL, et al. Prediction of landslide stability based on SVM model optimized by intelligent algorithm[J]. Journal of Natural Disasters, 2016, 25(5): 46-54. DOI: 10.13577 / j.jnd.2016.0506

[0129]

[22] Zhu CH, Zhang JJ, Liu Y, et al. Comparison of GA-BP and PSO-BPneural network models with initial BP model for rainfall-induced landslidesrisk assessment in regional scale: A case study in Sichuan, China[J].NaturalHazards, 2020, 100(1): 173-204.DOI: 10.1007 / s11069-019-03806-x

[0130]

[23] Zou Q. Study on the theory and method of comprehensive analysis and intelligent assessment of flood disaster risk [D]. Wuhan: Huazhong University of Science and Technology, 2013. DOI: 10.7666 / d.D409168

[0131]

[24] Ding QF, Yin X Y. Research survey of differential evolution algorithms[J]. CAAI Transactions on Intelligent Systems, 2017, 12(4): 431-442. DOI: 10.11992 / tis.201605015

[0132]

[25] Tang C. Characteristics of landslides and its hazard assessment in Boon area, Germany[J]. Journal of Soil and Water Conservation, 2000, 14(1): 48-53, 81. DOI: 10.13870 / j.cnki.stbcxb.2000.01.010

[0133]

[26] Li YM, Xie YY, Jiang DM, et al. Study onsensitivity in disaster-pregnant environmental factors of slope geological hazards in Nujiang prefecture[J]. Research of Soil and Water Conservation, 2018, 25(5): 300-305. DOI: 10.13869 / j.cnki.rswc.2018.05.043

[0134]

[27] Lin RF, Liu JP, Xu SH, et al. Evaluation method of landslide susceptibility based on random forest weighted information[J]. Science of Surveying and Mapping, 2020, 45(12): 131-138. DOI: 10.16251 / j.cnki.1009-2307.2020.12.020

[0135]

[28] Song Y X. Dynamic evaluation of landslide risk based on integrated monitoring of satellite, airborne and ground-based data [D]. Wuhan: China University of Geosciences, 2019. DOI: 10.27492 / d.cnki.gzdzu.2019.000169.

[0136]

[29] Xu SH, Liu JP, Wang XH, et al. Landslide susceptibility assessment method incorporating index of entropy based on support vector machine: A case study of Shanxi Province[J]. Geomatics and Information Science of Wuhan University, 2020, 45(8): 1214-1222. DOI: 10.13203 / j.whugis20200109

[0137]

[30] Wang ZH, Hu ZW, Zhao WJ, et al. Research on regional landslide susceptibility assessment based on multiple layer perceptron-Taking the hilly area in Sichuan as example[J]. Journal of Disaster Prevention and Mitigation Engineering, 2015, 35(5): 691-698. DOI: 10.13409 / j.cnki.jdpme.2015.05.021

[0138]

[31] Liu J, Wu Z. Landslide risk assessment of the Zhouqu-Wudu section of bailong river basin based on geographic informationsystem[J]. China Earthquake Engineering Journal, 2020, 42(6): 1723-1734. DOI: 10.3969 / j.issn.1000-0844.2020.06.1723

[0139]

[32] Shi ZY, Zhu HY, Wang JJ, et al. Analysis on susceptibility assessment of soil landslide in Xiangxi Prefecture from the perspective of coupling model[J]. Research of Soil and Water Conservation, 2021, 28(3): 377-383. DOI: 10.13869 / j.cnki.rswc.20210111.001

[0140]

[33] Storn R, Price K.Differential evolution-A simple and efficientheuristic for global optimization over continuous spaces[J].Journal of GlobalOptimization, 1997, 11(4): 341-359. DOI: 10.1023 / A: 1008202821328

[0141]

[34] Pourghasemi HR, Jirandeh AG, Pradhan B, et al.Landslidesusceptibility mapping using support vector machine and GIS at the Golestan Province, Iran[J].Journal of Earth System Science, 2013, 122(2): 349-369.DOI: 10.1007 / s12040-013-0282-2

[0142]

[35] Pradhan BA comparative study on the predictive ability of thedecision tree, support vector machine and neuro-fuzzy models in landslidesusceptibility mapping using GIS[J].Computers&Geosciences, 2013, 51: 350-365.DOI: 10.1016 / j.cageo.2012.08.023

[0143]

[36] Chung CJ, Fabbri A G. Predicting landslides for risk analysis-Spatial models tested by a cross-validation technique[J]. Geomorphology, 2008, 94(3 / 4): 438-452. DOI: 10.1016 / j.geomorph.2006.12.036.

[0144] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.

Claims

1. A landslide hazard assessment method based on feature selection and differential evolution algorithm optimization, characterized in that, Comprising the following steps: Step S1, Preliminary selection of landslide background analysis and assessment factors: (1) Analysis of regional disaster background; (2) Preliminary selection of regional disaster influencing factors to be screened; (3) Acquisition of influencing factor data and spatial consistency processing; (4) Construction of spatial database for regional disaster risk assessment factors; Step S2, Construction of four-dimensional feature screening method for optimal selection of assessment factors: (1) Status classification and determination coefficient calculation of each primary selected influencing factor; (2) Sensitivity analysis performed by constructing sensitivity indexes for primary selected influencing factors; (3) Importance analysis of primary selected influencing factors based on the ensemble algorithm integrating bagging strategy and boosting strategy; (4) Correlation analysis of each primary selected influencing factor; (5) Collinearity analysis of each primary selected influencing factor; (6) Comprehensive optimal selection of risk assessment factors; wherein, in step (2), sensitivity analysis is performed by constructing sensitivity indexes for primary selected influencing factors, that is, the sensitivity index of a factor is used to generally reflect the influence degree of a certain assessment factor on disaster occurrence, and the calculation method is as follows: In the formula: i represents the i-th evaluation factor; FSI i This represents the sensitivity index of the i-th evaluation factor; This represents the maximum value of the coefficient of determination in each state classification of the i-th evaluation factor; This represents the minimum value of the coefficient of determination among all state classifications of the i-th evaluation factor; In addition, a judgment threshold is introduced into the calculation. ,when At that time, the evaluation factors were considered It has a strong influence on the occurrence of landslides; The sensitivity index of a factor is obtained by respectively calculating the frequency ratio coefficient, information volume index and certainty coefficient of each status classification of each factor based on the status classification result, and taking the average value of the three types of certainty coefficients as the final sensitivity index of each factor; Step S3, Landslide risk assessment optimized based on differential evolution algorithm: (1) Normalization of landslide risk assessment factors; (2) Non-landslide sample extraction based on unsupervised clustering; (3) Global search optimization of main hyperparameters based on DE algorithm; (4) Construction of landslide risk assessment model; (5) Calculation of normalized landslide risk index for the region; (6) Regional landslide risk mapping.

2. The landslide hazard assessment method based on feature selection and differential evolution algorithm optimization according to claim 1, characterized in that, In step (2) of step S1, the regional disaster influencing factors include three categories of influencing factors: geographical influencing factors, geological influencing factors and social influencing factors; wherein, the geographical influencing factors include micro topographic factors, macro topographic factors, hydrological environmental factors and physical geographical factors, the geological influencing factors include lithology and faults, and the social influencing factors include land use type and roads.

3. The landslide hazard assessment method based on feature selection and differential evolution algorithm optimization according to claim 1, characterized in that, In step (3) of step S2, the importance analysis of primary selected influencing factors based on the ensemble algorithm integrating bagging strategy and boosting strategy is: adopting the random forest with bagging ensemble strategy and the gradient boosting decision tree with boosting ensemble strategy to jointly carry out factor importance analysis; performing weighted calculation on the characteristic importance characterization values according to the classification accuracy to obtain the composite importance characteristic values of each factor; A judgment threshold d for feature importance is set to eliminate factors with too small contribution to the model; Let d=1 / 2n, where n is the number of assessment factors, when the characterization value > d, it is considered that the corresponding factor has a large contribution to the model result and is retained; when the characterization value < d, it is considered that the corresponding factor has a small contribution to the model result and is eliminated.

4. The landslide hazard assessment method based on feature selection and differential evolution algorithm optimization according to claim 1, characterized in that, In step (4) of step S2, Pearson correlation coefficient is used to analyze the correlation degree between factors, when the absolute value of the correlation coefficient of two factors is greater than 0.40, it is considered that there is obvious correlation, and the factor is eliminated.

5. The landslide hazard assessment method based on feature selection and differential evolution algorithm optimization according to claim 1, characterized in that, In step (5) of step S2, if the tolerance TOL is less than 0.10 and the variance inflation factor VIF is greater than 10, the factor is considered to be severely collinear and is removed.

6. The landslide hazard assessment method based on feature selection and differential evolution algorithm optimization according to claim 1, characterized in that, In step (1) of step S3, the landslide hazard assessment factors are normalized, that is: the hazard assessment factors obtained by the final screening are normalized: discrete data are processed by one-hot coding, and continuous data are processed by linear normalization.

7. The landslide hazard assessment method based on feature selection and differential evolution algorithm optimization according to claim 1, characterized in that, In step (2) of step S3, non-landslide samples are extracted based on unsupervised clustering, namely: first, the self-organizing feature map neural network and Gaussian mixture model are used to perform unsupervised landslide hazard clustering on all grid units in the study area, and then negative samples are extracted from the common area of ​​the cluster with the lowest landslide hazard, and a machine learning dataset is constructed with known landslide samples.

Citation Information

Patent Citations

  • Bagging landslide forecasting method

    CN112381115A