Reservoir bank landslide deformation key control factor identification method based on interpretable machine learning

Through an interpretable machine learning method, XGBoost and SHAP value analysis and combined with SVM verification, the problem of lack of explanatory nature of traditional landslide deformation models is solved, and the accurate identification and prediction of key control factors for landslide deformation is achieved.

CN120372571APending Publication Date: 2025-07-25ZHENGZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510519621.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Traditional landslide identification methods rely on expert experience and are difficult to generalize, statistical analysis cannot quantify the interaction between factors, existing machine learning models lack explanatory nature, and it is difficult to clarify the specific contribution of each factor to landslide deformation.

Method used

Using an interpretable machine learning method, XGBoost gradient enhancement algorithm and SHAP value analysis are used, combined with SVM verification, landslide deformation prediction model is constructed to clarify the importance and interaction of control factors.

Benefits of technology

The identification accuracy and reliability of key control factors for landslide deformation are improved, the contribution of each factor to landslide deformation is clarified, and the prediction accuracy of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372571A_ABST
    Figure CN120372571A_ABST
Patent Text Reader

Abstract

The invention relates to a reservoir bank landslide deformation key control factor identification method based on interpretable machine learning. The method specifically comprises the following steps: S1, acquiring and preprocessing landslide deformation data and environmental data; s2, extracting potential control factors from reservoir water level data and rainfall data in the environment data; s3, establishing a prediction model between the potential control factors and the deformation data by using an XGBoost gradient lifting algorithm; s4, calculating an SHAP value of each potential control factor in the XGBoost prediction model, and carrying out importance sorting and visualization on the potential control factors based on the SHAP values; and S5, establishing a prediction model between the screened key control factors and the deformation data by using an SVM algorithm, and verifying and evaluating the key control factors in combination with the performance of the prediction model. According to the method, an interpretability mechanism of the XGBoost-SHAP method is utilized, model interpretability is introduced, the importance of each control factor to landslide deformation and interaction among different control factors are mastered, and the accuracy of key control factor identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of landslide deformation evolution mechanism, and particularly relates to a method for identifying key control factors of landslide deformation based on interpretable machine learning Background Technique

[0002] As a common geological disaster, landslides are characterized by strong suddenness, great destructive power, and wide influence range, seriously threatening people's lives and property safety and the ecological environment. The deformation evolution of landslide geological disasters is a complex non-linear process, affected by the combined action of multiple factors, including topography and geomorphology, geological structure, properties of rock and soil masses, hydrogeological conditions, and human activities. To effectively prevent and control landslide disasters, it is crucial to accurately identify the key control factors of landslide deformation

[0003] Traditional methods for identifying key factors of landslides mainly rely on expert experience and statistical analysis. Methods based on expert experience rely on the understanding of landslides in specific regions and are difficult to generalize to other regions; while methods based on statistical surveys usually assume that various factors are independent of each other and are difficult to capture the interaction between factors. In addition, these methods can only identify significant influencing factors and cannot quantify the specific contribution degree of each factor to landslide deformation. In recent years, with the rapid development of machine learning, data mining, and multi-field key information monitoring technologies, data-driven methods for identifying key factors of landslides have gradually become a research hotspot. Machine learning algorithms such as random forest, support vector machine, and neural network have been more widely applied to landslide susceptibility assessment and deformation prediction. However, most existing methods focus on the prediction accuracy of the model and ignore the interpretability of the model, that is, most prediction algorithms are black-box models and it is difficult to reveal the specific contribution degree of each factor to landslide deformation Summary of the Invention

[0004] The purpose of the present invention is to propose a method for identifying key control factors of reservoir bank landslide deformation based on interpretable machine learning, to solve the problem that traditional prediction algorithms lack interpretability and it is difficult to clarify the interaction between various control factors

[0005] A method for identifying key control factors of reservoir bank landslide deformation based on interpretable machine learning provided by the present invention includes the following steps

[0006] S1. Data acquisition and preprocessing: Acquire historical landslide deformation data and environmental data, and remove and fill outliers and missing values

[0007] S2. Screening of potential control factors: Extract potential control factors from reservoir water level data and rainfall data in environmental data

[0008] S3. Construction of XGBoost prediction model: Use the XGBoost gradient boosting algorithm to establish a prediction model between the preprocessed deformation data and potential control factor data

[0009] S4. Identification of key control factors based on the SHAP value method: Calculate the SHAP values of each potential control factor, identify the importance of the control factors, and visualize them;

[0010] S5. Verification and evaluation of key factors based on SVM: Use the SVM algorithm to establish a prediction model between the screened key control factors and the deformation data, and verify and evaluate the key control factors in combination with the performance of the prediction model.

[0011] Preferably, step S1 is specifically as follows:

[0012] S11. Obtain landslide geological data, deformation history data, and environmental data;

[0013] S12. Preprocess the collected data; the preprocessing includes removing missing values and handling outliers.

[0014] Preferably, step S2 is specifically as follows:

[0015] S21. Define the reservoir water level height and reservoir water level fluctuation at different time intervals from the reservoir water level data as potential control factors;

[0016] S22. Define the maximum rainfall and cumulative rainfall at different time intervals from the rainfall data as potential control factors.

[0017] Preferably, step S3 is specifically as follows:

[0018] S31. Establish a prediction model using the XGBoost gradient boosting algorithm, specifically In the formula, is the prediction result of sample i after t iterations, is the prediction result of the first t - 1 decision trees, f t (x i ) is the t-th decision tree model;

[0019] S32. Load the deformation data set, separate the features and labels, and divide the deformation data set into a test set and a training set;

[0020] S33. Train the XGBoost model using the training set and perform prediction evaluation on the model performance using the test set.

[0021] Preferably, step S4 is specifically as follows:

[0022] S41. Calculate the SHAP values of each factor in the XGBoost prediction model, specifically: In the formula, Shapley(X j ) is the potential feature X jThe contribution (i.e., SHAP value), n is the number of potential features, N\{j} is the set of all possible features excluding X j from all possible feature sets, S is a feature set in N\{j}, f(S) is the predicted value of a model in set S, and f(S∪{j}) is the predicted value of the model after adding feature X j to set S;

[0023] S42. Initialize the SHAP interpreter object using the trained XGBoost model and the deformed dataset, and calculate the SHAP values of all instances in the deformed dataset;

[0024] S43. Sort the features according to their average absolute SHAP values in all instances, and visualize the importance of control factors in combination with the summary graph;

[0025] S44. Key control factor screening. Convert the SHAP value of each factor into a proportion of the total SHAP values of all factors to compare the key control factors in different time periods and different data types. Specifically: In the formula, Val j is the transformed SHAP value of feature X j and n is the number of features.

[0026] 6. A method for identifying key control factors of landslide deformation based on interpretable machine learning according to claim 1, wherein step S5 is specifically as follows:

[0027] S51. Use the SVM method to establish a prediction model for the key control factors after screening and the displacement data. Specifically: f(x) = sign(w·x + b), where w is the weight vector, x is the input feature vector, and b is the bias;

[0028] S52. Divide the deformed dataset. 75% is the training set data and 25% is the test set data. Use the key control factors after screening as the input of the prediction model and the deformed data as the output of the prediction model;

[0029] S53. Use the SVM prediction model to calculate R 2 to evaluate the model and verify the key control factors. Specifically In the formula, y i is the true value, is the average value of the true values, and is the predicted value.

[0030] The beneficial effects provided by the present invention are: overcoming the limitations of conventional black-box prediction algorithms, identifying the key control factors of landslide deformation by introducing interpretable machine learning methods, clarifying the importance of control factors for landslide deformation and the interaction between control factors, and improving the reliability of key control factors and the accuracy of model prediction. Brief Description of the Drawings

[0031] Figure 1 It is a schematic flow chart of the method of the present invention.

[0032] Figure 2 It is a graph of monthly monitoring data and daily monitoring data of landslide deformation in an embodiment of the present invention.

[0033] Figure 3 It is a summary graph of SHAP of control factors of monthly monitoring data in an embodiment of the present invention.

[0034] Figure 4 It is a summary graph of SHAP of control factors of daily monitoring data in an embodiment of the present invention.

[0035] Figure 5 It is a comparison graph of screening thresholds of key factors in an embodiment of the present invention.

[0036] Figure 6 It is a graph of verification and evaluation results of key factors in an embodiment of the present invention. Specific Embodiment Method

[0038] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described below in conjunction with the accompanying drawings. Before formally elaborating on the present invention, a general description of the solution of the present invention will be given first for easy understanding.

[0039] A method for identifying key control factors of reservoir bank landslide deformation based on interpretable machine learning, as Figure 1-6 shown, includes the following steps:

[0040] S1. Data acquisition and preprocessing; It should be noted that step S1 is specifically as follows: S11. Obtain deformation history data and environmental data; As an embodiment, collect deformation history data, where the deformation history data includes landslide surface displacement, deep displacement, and crack development information. Collect environmental data, where the environmental data includes key information monitoring data such as groundwater fluctuation data, rainfall, and reservoir water level. S12. Perform preprocessing on the collected data; The preprocessing sequentially includes: removing missing values and dealing with outliers. As an embodiment, specifically removing missing values means that for data with a missing value ratio lower than 5%, directly delete the missing value data therein. Dealing with outliers specifically means that for numerical outlier data, remove the original outliers and interpolate and fill the data before and after the missing values (see Figure 2 ).

[0041] S2. Screening of potential control factors: Extract potential control factors from environmental data. As an example, the types and processes of extracting environmental control factors in the present invention are as follows: The height and fluctuation of the reservoir water level. Define the height value and fluctuation value of the reservoir water level within different time intervals based on the monitoring accuracy as potential control factors; The cumulative rainfall and the maximum rainfall. Define the cumulative rainfall and the maximum rainfall within different time intervals based on the monitoring accuracy as potential control factors. The final potential control factors are shown in Table 1 and Table 2:

[0042] Table 1 Potential control factors of long-term monthly monitoring data

[0043]

[0044] Table 2 Potential control factors of short-term daily monitoring data

[0045]

[0046]

[0047] S3. Construction of XGBoost prediction model: Use the XGBoost gradient boosting algorithm to establish a prediction model between potential control factors and deformation data. It should be noted that step S3 is specifically as follows: S31. Use XGBoost to establish a prediction model, where XGBoost obtains the final prediction model through the superposition of multiple decision trees. Specifically: In the formula, is the prediction result of sample i after t iterations, is the prediction result of the first t - 1 decision trees, and f t (x i ) is the t-th decision tree model. S32. Load the deformation data set, separate the features and labels, and prepare for model training; Divide the deformation data set into a training set and a test set. In the present invention, the division ratio of the training set and the test set is 75% and 25%. S33. Use the training set to train the XGBoost model, and use a random seed to ensure the repeatability of the results; Use the test set for prediction and evaluate the model performance.

[0048] S4. Identification of key factors based on the SHAP method: Determine the main control factors of landslide deformation by calculating the SHAP values of each factor. It should be noted that step S4 is specifically as follows: S41. Calculate the SHAP value of each factor. The SHAP value is the marginal contribution of each potential factor in the prediction model. Specifically: In the formula, Shapley(X j ) is the contribution of the potential feature X j (i.e., the SHAP value), n is the number of potential features, and N\{j} is the set excluding X jAll possible feature sets, S is a feature set in N\{j}, f(S) is a model prediction value in set S, and f(S∪{j}) is the model prediction value after adding feature X to set S j . S42. Initialize the SHAP interpreter object using the trained XGBoost model and the transformed dataset, and calculate the SHAP values of all instances in the transformed dataset. S43. Feature importance ranking and visualization. Rank the features according to their average absolute SHAP values across all instances, and use summary plots and other visualization techniques to illustrate the impact of each feature on the model prediction for that instance (see Figure 3 , Figure 4 ). Specifically, each point on the SHAP summary plot is the SHAP value of a feature and an instance. The feature importance decreases gradually from top to bottom. The scatter point color changes from blue to red, representing that the value of the feature increases from small to large. Each point represents the SHAP value of a sample, which represents the contribution of this feature to a single prediction, and the set of points represents the direction and magnitude of the overall impact of the feature on the prediction result. S44. Key control factor screening. Convert the SHAP value of each factor into a proportion of the total SHAP values of all factors to compare the key control factors in different time periods and different data types. Specifically: In the formula, Val j is the transformed SHAP value of feature X j , and n is the number of features. For the transformed SHAP values, key control factors can be screened by setting thresholds (1%, 5%, 10%, 15%) and performing trial calculations. As an example, the key factor threshold is selected as 5% (see Figure 5 ).

[0049] S5: Key control factor verification: Use the SVM method to establish a prediction model between the key control factors and the transformed data, and verify the key control factors through model evaluation. It should be noted that step S5 is specifically as follows: S51. Use the SVM method to establish a prediction model for the screened key control factors and displacement data. Specifically: f(x) = sign(w·x + b), where f(x) is the predicted value, w is the weight vector, x is the input feature vector, and b is the bias. Nonlinear regression is obtained by introducing a kernel function. S52. Divide the transformed dataset, with 75% as the training set data and 25% as the test set data. S53. Call the SVM prediction model to calculate R 2 to evaluate the model and verify the key control factors. Specifically In the formula, y i is the true value, is the average value of the true values, is the predicted value. In the embodiments of the present invention, the verification results R 2 of each monitoring point are all close to 0.9 (see Figure 6) verified the effectiveness of screening out the main controlling factors of landslide deformation.

[0050] The beneficial effects of the present invention are as follows: It overcomes the limitations of the black-box models of conventional machine learning and data mining algorithms, introduces interpretable machine learning to increase the interpretability of the prediction model, and clarifies the importance of each control factor for landslide deformation and the interaction between control factors. In addition, the SHAP values of each factor are transformed to facilitate the comparison of different types of landslide deformation data and the changes of key control factors at different times, improving the accuracy of identifying the main controlling factors of landslide deformation.

[0051] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for identifying key control factors of reservoir bank landslide deformation based on interpretable machine learning, characterized in that It includes the following steps: S1. Data acquisition and preprocessing: Acquire the landslide deformation historical data and environmental data, and remove and fill the outliers and missing values; S2. Screening of potential control factors: Extract the potential control factors from the reservoir water level data and rainfall data in the environmental data; S3. Construction of XGBoost prediction model: Use the XGBoost gradient boosting algorithm to establish a prediction model between the preprocessed deformation data and the potential control factor data; S4. Identification of key control factors based on the SHAP value method: Calculate the SHAP values of each potential control factor, identify the importance of the control factors and visualize them; S5. Verification and evaluation of key factors based on SVM: Use the SVM algorithm to establish a prediction model between the screened key control factors and the deformation data, and verify and evaluate the key control factors in combination with the performance of the prediction model.

2. The method for identifying key control factors of landslide deformation on the reservoir bank based on interpretable machine learning according to claim 1, wherein: The specific steps of S1 are as follows: S11. Acquire the landslide geological data, deformation historical data and environmental data; S12. Preprocess the collected data; the preprocessing includes removing the missing values and dealing with the outliers.

3. The identification method of key control factors for landslide deformation of reservoir bank based on interpretable machine learning according to claim 1, characterized in that: The specific steps of S2 are as follows: S21. Define the reservoir water level height and reservoir water level fluctuation at different time intervals from the reservoir water level data as potential control factors; S22. Define the maximum rainfall and cumulative rainfall at different time intervals from the rainfall data as potential control factors.

4. The identification method of key control factors for the deformation of reservoir bank landslides based on interpretable machine learning according to claim 1, wherein: The specific steps of S3 are as follows: S31. Establish a prediction model using the XGBoost gradient boosting algorithm, specifically as follows In the formula is the prediction result of sample i after t iterations, is the prediction result of the first t - 1 decision trees, and f t (x i ) is the t-th decision tree model; S32. Load the deformation data set, separate the features and labels, and divide the deformation data set into a test set and a training set; S33. Use the training set to train the XGBoost model, and use the test set to predict and evaluate the performance of the model.

5. The identification method of key control factors for landslide deformation of reservoir banks based on interpretable machine learning according to claim 1, characterized in that: The specific steps of S4 are as follows: S41. Calculate the SHAP value of each factor in the XGBoost prediction model, specifically: In the formula, Shapley(X j ) is the contribution of the potential feature X j (i.e., the SHAP value), n is the number of potential features, N\{j} is all possible feature sets excluding X j , S is a feature set in N\{j}, f(S) is a model prediction value in the set S, and f(S∪{j}) is the model prediction value after adding the feature X j to the set S; S42. Initialize the SHAP interpreter object with the trained XGBoost model and the deformation data set, and calculate the SHAP values of all instances in the deformation data set; S43. Sort the features according to the average absolute SHAP values of the features in all instances, and visualize the importance of the control factors in combination with the summary diagram; S44. Screening of key control factors. The SHAP value of each factor is converted into a proportion of the total SHAP values of all factors to compare the key control factors in different time periods and different data types. Specifically: In the formula, Val j is the feature X j The converted SHAP value, and n is the number of features.

6. The identification method of key control factors for landslide deformation of reservoir bank based on interpretable machine learning according to claim 1, characterized in that: The specific steps of S5 are as follows: S51. Use the SVM method to establish a prediction model for the screened key control factors and the displacement data, specifically: f(x)=sign(w·x + b), where f(x) is the predicted value, w is the weight vector, x is the input feature vector, and b is the bias; S52. Divide the deformation data set, 75% is the training set data and 25% is the test set data. Use the screened key control factors as the input of the prediction model and the deformation data as the output of the prediction model. S53. Calculate R using the SVM prediction model 2 Evaluate the model and verify the key control factors, specifically where y i is the true value, is the average value of the true values, is the predicted value.

Citation Information

Cited By

  • Tailing pond emergency supervision internet-of-things large model system and method

    CN121541508A

  • Classification and early warning method of reservoir water type landslide based on multi-source monitoring data

    CN122511069A