Landslide susceptibility evaluation method based on optimized non-landslide data selection and SHAP value
By optimizing the selection of non-landslide data and the SHAP value, a machine learning model was constructed and the importance of environmental factors was quantified. This solved the problems of lack of interpretability and negative sample uncertainty in the assessment of landslide susceptibility by machine learning models, and achieved a more accurate assessment of landslide susceptibility and decision support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-04-07
AI Technical Summary
Existing machine learning models lack interpretability in landslide susceptibility assessment, making it difficult to provide decision-making information. Furthermore, the high uncertainty in negative sample selection leads to prediction bias.
By optimizing the selection of non-landslide data and SHAP values, three negative sample datasets (Random, Buffer, and IV-Buffer) were obtained using ArcMap software. LR, SVM, DT, and XGBoost models were constructed, and the performance of the models was evaluated using ROC curves and performance indicators. The importance of environmental factors was quantified using the SHAP method to reveal the contribution mechanism of factors to landslide occurrence.
It improves the interpretability and predictive accuracy of the model, provides decision support, ensures sample balance, and enhances the interpretability and predictive performance of the model.
Smart Images

Figure CN121809228A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of geological disaster prevention and control engineering technology, specifically to a landslide susceptibility evaluation method based on optimized selection of non-landslide data and SHAP values. Background Technology
[0002] The goal of landslide susceptibility (LSM) research is to calculate the spatial distribution probability of a slope evolving into a landslide under the nonlinear interaction of multiple disaster-causing environmental factors in a certain area.
[0003] Currently, LSM methods are categorized into four types: physical models, heuristic models, statistical models, and machine learning models. Each model has its own advantages, disadvantages, and limitations. Physical models simulate site characteristics and have high accuracy, but this method is only suitable for small-scale analysis of local areas, and large-scale analysis is computationally expensive. Heuristic models perform parametric analysis of landslide condition factors, but this method relies on expert knowledge and has a high degree of subjectivity. Statistical models simplify the relationships between data through mathematical concepts, such as information content models, evidence weight models, and frequency ratio models, but they have certain limitations in revealing the nonlinear relationships between various environmental factors and landslides. With the development of artificial intelligence technology, machine learning models, which can achieve high accuracy without relying on too much prior knowledge, have been widely applied to LSM research. These models can effectively capture and simulate the complex nonlinear relationships between environmental factors and landslide states, thereby achieving a comprehensive assessment of landslide risk. These models include LR, ANN, BP, RF, DT, and XGBoost models. Studies have shown that tree models have better generalization ability than other models, but the results vary under different geological conditions. Therefore, exploring highly accurate models for a specific region has significant practical implications.
[0004] Machine learning models are highly dependent on the quality of the data samples for classification problems. Research has shown that in landslide susceptibility studies, machine learning models can only achieve superior performance when the sample sizes are balanced. Landslide susceptibility datasets consist of landslide hazard points (positive samples) and non-landslide points where no landslides have been detected (negative samples). Landslide hazard points are selected based on historical landslide records or remote sensing imagery and aerial photographs, resulting in relatively low uncertainty. Negative landslide samples are usually not directly available, and most literature uses random selection, leading to greater uncertainty in negative samples and causing model prediction bias. Furthermore, most machine learning models are considered black-box models, lacking interpretability of results. Although the predictions are relatively accurate, they do not provide much decision-making information for disaster prevention and mitigation. Summary of the Invention
[0005] The purpose of this invention is to provide a landslide susceptibility assessment method based on optimized selection of non-landslide data and SHAP values, aiming to solve the technical problem that existing landslide prediction methods using machine learning models lack interpretability and are difficult to provide decision-making information.
[0006] To achieve the above objectives, this invention provides a landslide susceptibility assessment method based on optimized selection of non-landslide data and SHAP values, comprising the following steps:
[0007] Step 1: Obtain landslide logging information for the study area, collect data on 12 environmental factors, and calculate the values of each environmental factor using frequency ratios;
[0008] Step 2: Based on the selected environmental factors and the IV model, the susceptibility region is partitioned. ArcMap 10.8 software is used to obtain the datasets of three types of negative samples: Random, Buffer, and IV-Buffer, and these are combined with the landslide samples to construct the overall dataset.
[0009] Step 3: Test the correlation between environmental factors and landslide status using Person correlation coefficient, and eliminate highly correlated factors; at the same time, construct LR, SVM, DT and XGBoost models, simulate on three types of datasets respectively, and use ROC curves and four performance indicators to judge the model performance.
[0010] Step 4: Select the best model from the four machine learning models in the results of Step 3, perform susceptibility mapping, and divide it into 5 susceptibility levels. Compare the advantages and disadvantages of each model.
[0011] Step 5: The importance of the 12 environmental factors selected in Step 1 was quantified using the SHAP method, and the contribution mechanism of each factor to the occurrence of landslides was revealed by dependency plots and partial dependency plots.
[0012] Optionally, the environmental factors collected in step 1 are 12 landslide-related factors, specifically including: elevation, slope, aspect, plane curvature, profile curvature, stratigraphic lithology, distance from fault, distance from river, annual average rainfall, normalized difference vegetation index (NDVI), land use type, and distance from road.
[0013] The formula for calculating the frequency ratio (FR) is:
[0014]
[0015] Where Npix(Si) is the number of landslide pixels in the i-th factor, and Npix(Ni) is the total number of pixels in the i-th factor.
[0016] Optionally, in step 2, the formula for calculating the information value I of the IV model is:
[0017]
[0018] The study area was divided into different susceptibility levels by summing the information values of each factor.
[0019] The three negative sample selection strategies are: Random strategy, which randomly selects non-slide points throughout the study area; Buffer strategy, which selects non-slide points outside the landslide point buffer zone; and IV-Buffer strategy, which selects non-slide points in low-susceptibility areas in combination with buffer constraints.
[0020] Use geographic information system software to sample negative samples, ensuring that the ratio of negative samples to positive samples is 1:1.
[0021] Optionally, in step 3, a Pearson correlation coefficient threshold is set, and redundant factors are removed when the absolute value of the correlation coefficient between two environmental factors exceeds the threshold.
[0022] The LR model uses the sigmoid function, the SVM model uses the radial basis function (RBF), the DT model uses the Gini index or information gain, and the XGBoost model uses the gradient boosting framework.
[0023] Model performance was evaluated using the AUC value of the ROC curve, as well as accuracy, precision, recall, and F1 score.
[0024] Optionally, in step 4, the model with the highest AUC value and the best overall performance is selected as the final evaluation model.
[0025] The susceptibility index is divided into five levels: extremely low, low, medium, high, and extremely high, using either the natural breakpoint method or the equal interval method.
[0026] Optionally, in step 5, the SHAP method quantifies the importance of environmental factors based on the Shapley value theory. A positive SHAP value indicates an increase in the probability of landslides, while a negative value indicates a decrease in the probability of landslides. The SHAP dependency plot shows the relationship between factor values and SHAP values, revealing nonlinear relationships and threshold effects. The partial dependency plot shows the average impact of factors on the probability of landslides under different values.
[0027] This invention provides a landslide susceptibility assessment method based on optimized selection of non-landslide data and SHAP values. First, corresponding environmental factors are selected according to the geological characteristics of the study area. Then, the values of each factor are calculated using frequency ratios. Next, negative landslide samples are screened using Acrmap 10.8 software, thus constructing three types of sample datasets: Random, Buffer, and IV-Buffer. After obtaining the sample datasets, LR, SVM, DT, and XGBoost classification models are implemented using Python to conduct susceptibility assessment analysis of the region. Simulations are performed on the three types of datasets, and ROC curves and four performance indicators are used to judge model performance. Subsequently, the optimal model among the four machine learning models is selected for susceptibility mapping, divided into five susceptibility levels. The advantages and disadvantages of each model are compared. Finally, the SHAP method is used to quantify the importance of influencing factors, and dependency and partial dependency plots reveal the contribution mechanism of each factor to landslide occurrence. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a schematic diagram illustrating the specific execution process of a landslide susceptibility evaluation method based on optimized selection of non-landslide data and SHAP values according to the present invention.
[0030] Figure 2 This is a frequency ratio distribution diagram of various environmental factors in an embodiment of the present invention.
[0031] Figure 3 This is a distribution map of basic environmental factors of landslides in an embodiment of the present invention.
[0032] Figure 4 This is a diagram showing the results of the landslide negative sample selection strategy in an embodiment of the present invention.
[0033] Figure 5 This is a graph showing the correlation analysis results among various influencing factors in an embodiment of the present invention.
[0034] Figure 6 This is the ROC curve of different negative landslide sample selection methods in the embodiments of the present invention.
[0035] Figure 7 This is a landslide susceptibility result diagram of the Hybrid IV-Buffer Sampling method in this embodiment of the invention.
[0036] Figure 8 This is a dependency graph between the dominant factor and the model prediction results in an embodiment of the present invention. Detailed Implementation
[0037] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0038] This invention provides a landslide susceptibility assessment method based on optimized selection of non-landslide data and SHAP values, comprising the following steps:
[0039] Step 1: Obtain landslide logging information for the study area, collect data on 12 environmental factors, and calculate the values of each environmental factor using frequency ratios;
[0040] Step 2: Based on the selected environmental factors and the IV model, the susceptibility region is partitioned. ArcMap 10.8 software is used to obtain the datasets of three types of negative samples: Random, Buffer, and IV-Buffer, and these are combined with the landslide samples to construct the overall dataset.
[0041] Step 3: Test the correlation between environmental factors and landslide status using Person correlation coefficient, and eliminate highly correlated factors; at the same time, construct LR, SVM, DT and XGBoost models, simulate on three types of datasets respectively, and use ROC curves and four performance indicators to judge the model performance.
[0042] Step 4: Select the best model from the four machine learning models in the results of Step 3, perform susceptibility mapping, and divide it into 5 susceptibility levels. Compare the advantages and disadvantages of each model.
[0043] Step 5: The importance of the 12 environmental factors selected in Step 1 was quantified using the SHAP method, and the contribution mechanism of each factor to the occurrence of landslides was revealed by dependency plots and partial dependency plots.
[0044] The environmental factors collected in step 1 are 12 landslide-related factors, specifically including: elevation, slope, aspect, plane curvature, profile curvature, stratigraphic lithology, distance from fault, distance from river, annual average rainfall, normalized difference vegetation index (NDVI), land use type, and distance from road.
[0045] The formula for calculating the frequency ratio (FR) is:
[0046]
[0047] Where Npix(Si) is the number of landslide pixels in the i-th factor, and Npix(Ni) is the total number of pixels in the i-th factor.
[0048] In step 2, the formula for calculating the information value I of the IV model is:
[0049]
[0050] The study area was divided into different susceptibility levels by summing the information values of each factor.
[0051] The three negative sample selection strategies are: Random strategy, which randomly selects non-slide points throughout the study area; Buffer strategy, which selects non-slide points outside the landslide point buffer zone; and IV-Buffer strategy, which selects non-slide points in low-susceptibility areas in combination with buffer constraints.
[0052] Use geographic information system software to sample negative samples, ensuring that the ratio of negative samples to positive samples is 1:1.
[0053] In step 3, a Pearson correlation coefficient threshold is set, and redundant factors are removed when the absolute value of the correlation coefficient between two environmental factors exceeds the threshold.
[0054] The LR model uses the sigmoid function, the SVM model uses the radial basis function (RBF), the DT model uses the Gini index or information gain, and the XGBoost model uses the gradient boosting framework.
[0055] Model performance was evaluated using the AUC value of the ROC curve, as well as accuracy, precision, recall, and F1 score.
[0056] In step 4, the model with the highest AUC value and the best overall performance is selected as the final evaluation model;
[0057] The susceptibility index is divided into five levels: extremely low, low, medium, high, and extremely high, using either the natural breakpoint method or the equal interval method.
[0058] In step 5, the SHAP method quantifies the importance of environmental factors based on the Shapley value theory. A positive SHAP value indicates an increase in the probability of landslides, while a negative value indicates a decrease in the probability of landslides. The SHAP dependency plot shows the relationship between factor values and SHAP values, revealing nonlinear relationships and threshold effects. The partial dependency plot shows the average impact of factors on the probability of landslides under different values.
[0059] The specific execution process is as follows: Figure 1 As shown, please refer to Figures 2 to 8 The following is a further explanation with reference to specific embodiments:
[0060] This embodiment uses Changzhou District, Wanxiu District, and Longxu District of Wuzhou City, Guangxi Zhuang Autonomous Region as the study area. The geographical location of the area is 110°0′~111°30′ east longitude and 23°0′~23°40′ north latitude, with a total area of approximately 1809.28 km². The study area has typical landforms, belonging to the South China hilly terrain, with a terrain distribution that is high in the northwest and low in the southeast, with significant differences in altitude and obvious changes in slope. Demonstrating the method in this area yields significant results.
[0061] In reality, the area of non-landslide regions is much larger than that of landslide regions, resulting in a severe imbalance in the distribution of data categories. If training is directly based on this type of data, the model often tends to predict non-landslide samples. Therefore, this embodiment adopts a balanced sampling strategy to select samples, maintaining a 1:1 ratio between the two types of data.
[0062] The operation steps in the embodiment are as follows (operation step three in the embodiment corresponds to step 2 of the method of the present invention, and step four corresponds to the fusion of steps 3 to 5):
[0063] Step 1:
[0064] The landslide logging information for the region was obtained. A total of 141 landslide records were used in Wuzhou City for this method demonstration.
[0065] Step Two:
[0066] Influencing factors can be selected based on local geological characteristics, and an evaluation factor FR is introduced. When FR > 1, it indicates that the factor has a positive driving effect on landslides; when FR < 1, it has a negative driving effect. This study selected 12 environmental factors across five categories: topographic factors, geological factors, hydrological factors, vegetation factors, and human activity factors, for landslide susceptibility modeling in the study area, using a 30 m spatial resolution raster as the evaluation unit. The characteristics of each factor are as follows.
[0067] 1. Geological factors: The main geological features in this area are soil and rock types, including sandstone, Quaternary rocks, clastic rocks, and granite. The FR (friction radius) of the Quaternary rocks and clastic rocks is greater than 1, while the FR of the sandstone and granite is significantly lower.
[0068] 2. Hydrological factors: The FR within 800 m of the river channel in this area is greater than 1, and the probability of landslides is highest in the 200–500 m range.
[0069] 3. Vegetation Factors: The effects of vegetation factors are reflected by NDVI, which indicates the regional vegetation cover and health status. A lower NDVI value indicates a stronger inhibition of shallow landslides. When NDVI is less than 0.2, there are no landslides in the area; when NDVI is greater than 0.2, the landslide frequency gradually decreases as the NDVI value increases.
[0070] 4. Human activity factors:
[0071] In this area, the FR (Fluid Rate) within the 0–900 m buffer zone of the road is greater than 1, while the probability of landslides is lower in areas more than 900 m away from the road. All selected human activity factors are significantly positively correlated with the occurrence of landslides.
[0072] With the selection of influencing factors now complete, the Frequency Ratio (FR) model is used to analyze these factors, mapping their impact on landslide susceptibility under different conditions. Discrete variables are kept in their fixed categories, while continuous variables are divided into sub-intervals using ArcMap 10.8's natural break method. The analysis results are as follows: Figure 2 and Figure 3 As shown.
[0073] Step 3: Select three types of negative samples
[0074] 1: Selection of negative samples for the Random term:
[0075] Negative landslide samples were randomly selected within the study area using the Create Random Points tool in ArcMap 10.8 software.
[0076] 2: Buffer-based term negative sample selection
[0077] A buffer was created using the Buffer tool in ArcMap 10.8, with a buffer distance set to 500m. The resulting area was used as the selection region for negative landslide samples. Negative landslide samples were then randomly selected using the Create Random Points tool. The results are as follows. Figure 4 (a)
[0078] 3: Negative Sample Selection for Hybrid IV-Buffer Term
[0079] The information value (IV) method was used to calculate the information content of each environmental factor, and the total information content was obtained by overlaying the values using the Raster Calculator tool in ArcMap 10.8 software. Figure 4 (b) Using the natural breakpoint method for partitioning, the greater the information content, the more prone the area is to landslides. Since the information content method cannot perfectly avoid all landslide points, in order to enhance the reliability of the negative landslide samples, extremely low and low-prone areas are extracted, and a buffer is further established in these areas using the Buffer tool with a buffer distance of 500m. Finally, negative landslide samples are constructed in these areas.
[0080] The three negative sample datasets have now been selected, and the final result is as follows: Figure 4 (c);
[0081] Step 4: Conduct landslide susceptibility assessment and analyze the results.
[0082] 1. Correlation analysis of influencing factors
[0083] In landslide susceptibility studies, the data is typically conducted in grid units, with millions of data points. To improve the computational efficiency and interpretability of the model, it is necessary to eliminate multicollinearity among the influencing factors in advance to avoid introducing redundant information. This makes it easier to assess the contribution of each factor to the landslide state. Therefore, the Pearson correlation coefficient is used to calculate the linear relationship between the factors, and the 12 factors selected in step 1 are sequentially labeled as a1.
[0084] The correlation coefficients between the various factors selected in this method demonstration are as follows: Figure 5 When the absolute value of the correlation coefficient is greater than 0.5, it indicates that there is a strong correlation between the two variables; conversely, it indicates that the correlation is weak or not significant. Figure 4 The lower triangular area displays the scatter distribution of each variable pairwise and its linear regression line, through annotation. The indicators directly reflect the linear relationship between two variables. The larger the value, the stronger the correlation. The upper triangular number represents the Pearson correlation value between the two variables. It can be seen that there is no obvious collinearity problem among the various influencing factors, thus verifying that the selection of 12 influencing factors in this demonstration is reasonable.
[0085] 2. Validation and Comparison of Landslide Susceptibility Models
[0086] After completing the selection and evaluation of influencing factors and constructing three negative sample datasets in the above steps, the final model validation and analysis will begin.
[0087] First, ArcMap 10.8 software was used to divide the study area into 1,973,041 raster cells at a 30m resolution, and the CGCS2000_GK_Zone_19 coordinate system was uniformly adopted to ensure accurate description of the selected area. Datasets were constructed for the three types of negative landslide samples obtained in the above steps (Randomsampling, Buffer-based sampling, and HybridIV-Buffersampling), and all environmental factors were introduced as input variables. Four types of machine learning models (LR, SVM, DT, and XGBoost) were built using Python. Positive and negative landslide samples were randomly divided in a 1:1 ratio, with 50% used for training and 50% for validation. Model performance was evaluated using ROC curves and relevant indicators, and the results are shown below. Figure 6 As shown in Table 1.
[0088] Table 1 Results of the model under different landslide negative sample selection methods
[0089]
[0090] To further verify the application effect of the model in this demonstration, the best model trained using Hybrid IV-Buffer Sampling among the four models was saved. The susceptibility index of the study area was calculated and imported into ArcMap 10.8 for susceptibility partitioning. The study area was divided into five levels—very low, low, moderate, high, and very high—using the natural breakpoint method. The final results are as follows: Figure 7 As shown in Table 2.
[0091] Table 2. Evaluation results of the susceptibility level of the Hybrid IV-Buffer Sampling method.
[0092]
[0093] 3. The model is interpretable.
[0094] After obtaining the landslide susceptibility prediction results for Wuzhou area from the above steps, it is necessary to analyze the contribution and role of influencing factors in the model to clarify the contribution of each influencing factor. This is of crucial value for disaster prevention and mitigation. Therefore, this experiment uses the SHAP method to perform feature importance analysis on the XGBoost model under Hybrid IV-Buffer sampling. The final results are as follows: Figure 8 As shown.
[0095] The above description discloses only one preferred embodiment of the present invention, and should not be construed as limiting the scope of the present invention. Those skilled in the art will understand that all or part of the processes of the above embodiments can be implemented, and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.
Claims
1. A landslide susceptibility assessment method based on optimized selection of non-landslide data and SHAP values, characterized in that, Includes the following steps: Step 1: Obtain landslide logging information for the study area, collect data on 12 environmental factors, and calculate the values of each environmental factor using frequency ratios; Step 2: Based on the selected environmental factors and the IV model, the susceptibility region is partitioned. ArcMap 10.8 software is used to obtain the datasets of three types of negative samples: Random, Buffer, and IV-Buffer, and these are combined with the landslide samples to construct the overall dataset. Step 3: Test the correlation between environmental factors and landslide status using Person correlation coefficient, and eliminate highly correlated factors; at the same time, construct LR, SVM, DT and XGBoost models, simulate on three types of datasets respectively, and use ROC curves and four performance indicators to judge the model performance. Step 4: Select the best model from the four machine learning models in the results of Step 3, perform susceptibility mapping, and divide it into 5 susceptibility levels. Compare the advantages and disadvantages of each model. Step 5: The importance of the 12 environmental factors selected in Step 1 was quantified using the SHAP method, and the contribution mechanism of each factor to the occurrence of landslides was revealed by dependency plots and partial dependency plots.
2. The landslide susceptibility assessment method based on optimized non-landslide data selection and SHAP values as described in claim 1, characterized in that, The environmental factors collected in step 1 are 12 landslide-related factors, specifically including: elevation, slope, aspect, plane curvature, profile curvature, stratigraphic lithology, distance from fault, distance from river, annual average rainfall, normalized difference vegetation index (NDVI), land use type, and distance from road. The formula for calculating the frequency ratio (FR) is: ; Where Npix(Si) is the number of landslide pixels in the i-th factor, and Npix(Ni) is the total number of pixels in the i-th factor.
3. The landslide susceptibility assessment method based on optimized non-landslide data selection and SHAP values as described in claim 2, characterized in that... In step 2, the formula for calculating the information value I of the IV model is: ; The study area was divided into different susceptibility levels by summing the information values of each factor. The three negative sample selection strategies are: Random strategy, which randomly selects non-slide points throughout the study area; Buffer strategy, which selects non-slide points outside the landslide point buffer zone; and IV-Buffer strategy, which selects non-slide points in low-susceptibility areas in combination with buffer constraints. Use geographic information system software to sample negative samples, ensuring that the ratio of negative samples to positive samples is 1:
1.
4. The landslide susceptibility assessment method based on optimized non-landslide data selection and SHAP values as described in claim 3, characterized in that... In step 3, a Pearson correlation coefficient threshold is set, and redundant factors are removed when the absolute value of the correlation coefficient between two environmental factors exceeds the threshold. The LR model uses the sigmoid function, the SVM model uses the radial basis function (RBF), the DT model uses the Gini index or information gain, and the XGBoost model uses the gradient boosting framework. Model performance was evaluated using the AUC value of the ROC curve, as well as accuracy, precision, recall, and F1 score.
5. The landslide susceptibility assessment method based on optimized non-landslide data selection and SHAP values as described in claim 4, characterized in that... In step 4, the model with the highest AUC value and the best overall performance is selected as the final evaluation model; The susceptibility index is divided into five levels: extremely low, low, medium, high, and extremely high, using either the natural breakpoint method or the equal interval method.
6. The landslide susceptibility assessment method based on optimized non-landslide data selection and SHAP values as described in claim 5, characterized in that, In step 5, the SHAP method quantifies the importance of environmental factors based on the Shapley value theory. A positive SHAP value indicates an increased probability of landslides, while a negative value indicates a decreased probability of landslides. The SHAP dependency plot shows the relationship between factor values and SHAP values, revealing nonlinear relationships and threshold effects; the partial dependency plot shows the average impact of factors on landslide probability under different values.