An interpretable landslide surface displacement prediction method based on LightGBM and SHAP
By combining LightGBM and SHAP, the feature engineering and parameter optimization of the landslide surface displacement prediction model are solved, and the existing model's long training time, time-consuming and laborious parameter adjustment and poor interpretability are achieved, and efficient, accurate and interpretable landslide displacement prediction is achieved.
Patent Information
- Application Number
- CN202410589223.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-13
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-05-13
AI Technical Summary
The existing landslide displacement prediction model has the problems of long training time, time-consuming and labor-intensive adjustment of parameters, and poor interpretability, making it difficult to achieve a balance between prediction accuracy, training speed and interpretability.
The landslide surface displacement prediction method based on LightGBM and SHAP is adopted, and the model parameters are optimized through the combination of feature engineering, LightGBM model training and tree structure probability density estimation algorithm, and the SHAP value is used to calculate the contribution of the features to the prediction results to achieve the interpretability of the model.
It improves the stability and accuracy of the prediction results, enhances the interpretability of the model, makes the prediction results more reliable and can be used for landslide risk management and decision-making more effectively.
Smart Images

Figure CN118536032B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of landslide surface displacement prediction, and particularly relates to an interpretable landslide surface displacement prediction method based on LightGBM and SHAP. Background Art
[0002] Landslides are common geological disasters, characterized by numerous disaster-causing factors and complex coupling disaster-causing mechanisms. Generally speaking, landslide displacement is considered the most direct early warning indicator for identifying instability and describing different failure stages, and can intuitively reflect the deformation characteristics of landslides. Therefore, it is of great significance to make a reliable prediction of the short-term changes in surface displacement quickly and accurately to avoid life and property losses caused by landslides. Landslide deformation prediction can be mainly divided into three prediction methods according to different principles: physical model method, mathematical statistics method, and intelligent model method. The physical model method combines the landslide process with mechanics and mathematical theories, and can be effectively applied to the prediction and early warning of disasters such as slope displacement deformation and landslide collapse, with strong interpretability. However, the above prediction models are only applicable to specific situations, with poor generalization, and when dealing with complex environments or many difficult-to-quantify uncertainty factors, it is difficult to construct the model and the prediction accuracy is insufficient. The analysis model based on mathematical statistics has poor flexibility and cannot deeply extract the internal characteristics of data. It is only applicable to the prediction of landslide monitoring data with linear or approximately linear deformation trends, ignoring the uncertainties existing in the landslide process, and cannot achieve good prediction results for landslide processes with strong nonlinear changes. With the rapid development of computer technology and artificial intelligence, more and more machine learning and deep learning models have been applied to landslide displacement prediction at home and abroad.
[0003] However, in existing research, although certain achievements have been made in the application of machine learning methods to landslide displacement prediction, the problem of poor model interpretability caused by the "black box" drawback has not been overcome, that is, in the prediction model, only the input and output are known, and it is impossible to estimate the importance of each feature to the model prediction result, let alone understand the interaction relationship between different features.
[0004] Therefore, it is of great significance to construct an efficient and interpretable landslide displacement prediction analysis model to achieve the balance among prediction accuracy, model training speed, and interpretability, which can provide new ideas for landslide risk management and decision-making. Summary of the Invention
[0005] To solve the defects and deficiencies of the above-mentioned prior art, the present invention provides an interpretable landslide surface displacement prediction method based on LightGBM and SHAP, which is used to solve the problems of long model training time, time-consuming and laborious manual adjustment of model parameters, and poor model interpretability existing in the prior art.
[0006] To achieve the above object, the present invention adopts the following technical solutions to solve the problem:
[0007] An interpretable landslide surface displacement prediction method based on LightGBM and SHAP, specifically including the following steps:
[0008] Step 1: Obtain landslide monitoring data, including displacement data, precipitation, mud water level, and water content at different positions and depths of the slope body.
[0009] Step 2: Perform data preprocessing of outlier removal and missing value filling on the multivariate data.
[0010] Step 3: According to the known attributes of the data processed in Step 2, perform feature engineering, including feature screening and feature construction.
[0011] Step 4: Construct a model using multiple input items obtained in Step 3 as influencing factors; specifically including the following steps:
[0012] Step 4-1: Divide the training data set and the test set, and use the training data set to construct a LightGBM landslide surface displacement prediction model.
[0013] Step 4-2: Use the tree structure probability density estimation algorithm to train the LightGBM landslide surface displacement prediction model constructed in Step 4-1 to optimize the model parameters.
[0014] Step 4-3: Select the model with the optimal accuracy as the final prediction model and output the prediction result.
[0015] Step 5: Use three evaluation indicators, RMSE, MAE, and R2, to evaluate the prediction model.
[0016] Step 6: Calculate the marginal contribution of each input feature to the model result through the SHAP method, so as to be able to explain the influence degree of different input features on the surface displacement.
[0017] The method of the present invention can effectively improve the stability and accuracy of the prediction result and improve the interpretability of the prediction result.
[0018] Further, in Step 1, the displacement data in three directions are respectively named x, y, z, the rainfall is named tyl, the mud water level is named nsw, and the water content at different positions and depths of the slope body are respectively named hs1, hs2, and hs3.
[0019] Further, in Step 2, the data selected in Step 1 is processed for outliers, and the removed outliers are regarded as missing values. At the same time, due to the defects of the acquisition instrument or signal transmission, there are some missing values in the data itself. All missing values are filled using the nearest neighbor interpolation method.
[0020] Further, step 3 specifically includes the following steps:
[0021] Step 3-1: Use the Spearman correlation coefficient method to calculate the correlation coefficient between the multivariate monitoring data X and the surface displacement Y, and eliminate the monitoring items with a correlation coefficient lower than 0.2 to reduce data redundancy and achieve feature screening;
[0022] The Spearman correlation coefficient is a method to measure the linear correlation degree between random variables X and Y. The value range of the correlation coefficient is [-1, 1]. The larger the value of the correlation coefficient, the higher the correlation degree between X and Y, and vice versa. The calculation formula is as follows:
[0023]
[0024] where, x i and y i are the values of the random variables X and Y when the observation value is i, and are the average values of the random variables X and Y.
[0025] Step 3-2: For the features retained after screening, calculate their short-term features and long-term features respectively as the model input items. Specifically, for the surface displacement, take the average displacement in the previous 6 / 12 hours, for the cumulative rainfall, take the average cumulative rainfall in the previous 6 / 12 / 24 hours, and for the soil moisture content, take the average soil moisture content in the previous 6 / 12 / 24 hours. Add the calculation results as new features to the dataset.
[0026] Further, in step 4-1, during the training process, divide the training set and the test set in a ratio of 4:1, and further divide the training set into a training set and a validation set using 5-fold forward validation during the training process. Take the average MSE value of the 5-fold forward validation as the optimization objective function for model parameter adjustment.
[0027] Further, in step 4-1, use the training dataset to construct a LightGBM landslide surface displacement prediction model. LightGBM uses histogram optimization to segment continuous feature values, and each decision tree finds the leaf with the largest split gain for splitting, and prevents overfitting by restricting the depth of the tree. The core idea of LightGBM is to use decision trees as base learners to iteratively train the data to obtain the optimal model. The calculation is shown in Equation (2):
[0028]
[0029] In Equation 2, F T (x i ) is the predicted value of the i-th base learner; x iFor the i-th sample, υ is the set of learners.
[0030] Furthermore, in the step 4-2, the optimized parameters are the number of leaves num_leaves, the number of base learners n_estimators, the learning rate learning_rate, and the maximum tree depth max_depth, a total of 4 parameters; the tree structure probability density estimation algorithm used is to establish a surrogate model by classification method. Instead of directly estimating the output of each sample point, the sample points are vaguely divided into two types: good and bad, and the expected improvement is used as its acquisition function to generate new sampling points. The calculation formula of the expected improvement is:
[0031]
[0032] where M is the surrogate model constructed using the kernel density estimation method, y* represents a certain threshold, y represents the current optimization objective function value, x is the recommended hyperparameter value, and p M (y|x) represents the surrogate probability model of y calculated using the model M when x takes values.
[0033] Furthermore, in the step 4-3, selecting the model with the optimal accuracy means comparing the prediction accuracies among the multiple linear regression model, the random forest regression model, and the light gradient boosting model to obtain the model.
[0034] Furthermore, in the step 5, the mean absolute error (MAE), the root mean squared error (RMSE), and the coefficient of
[0035] determination (R2) are used for evaluation, and the calculation formulas are shown as follows:
[0036]
[0037] In the formula, i is the sample number; n is the number of samples, is the predicted surface displacement value, y i is the actual value of the surface displacement, is the actual average value of the surface displacement.
[0038] Among them, the closer MAE and RMSE are to 0, the better the model fitting effect, and the closer R2 is to 1, the better the model fitting effect.
[0039] Furthermore, in the step 6, the calculation formula of the SHAP value is:
[0040]
[0041] Where: f(x) is a machine learning model, which is the LightGBM model in this patent; z* = {0, 1}, when feature i is observed, z* = 1, otherwise 0; if i participates in the prediction process, M is the number of features; φ i is the contribution of feature i, and the expression is
[0042]
[0043] Where: φ i represents the SHAP value of feature i; N is the set of sample features to be explained, S is the set containing non-zero indices in z*, M is the total number of different input features, and f(S∪{i}) - f(S) is the contribution of feature i to the prediction result.
[0044] Furthermore, in step 6, the marginal contributions of each input feature to the model result are calculated by the SHAP method, so as to be able to explain the influence degree of different input features on the ground surface displacement. The specific steps are as follows:
[0045] Step 6-1: For the selected input feature, use SHAP to calculate the SHAP values of all samples of this feature, and take their average value as the global feature importance of the selected input feature, so as to obtain the global interpretation of the input feature;
[0046] Step 6-2: Use the positive and negative effects of SHAP values to analyze the interaction of different input features on the ground surface displacement prediction result, and improve the interpretability of the model.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0048] (1) The present invention combines the LightGBM model and the SHAP model to realize the prediction of landslide ground surface displacement, improves the problem of poor interpretability in the short-term prediction of ground surface displacement by a single model, and compared with the random forest regression model and the multiple linear regression model, the model constructed in this patent has higher prediction accuracy and better interpretability.
[0049] (2) The parameter optimization method combining the tree structure probability density estimation algorithm and forward verification used in step 4-2 of the present invention can find the optimal hyperparameters of the model with fewer iteration times compared with random search and genetic algorithm search, and improves the parameter tuning efficiency.
[0050] (3) The LightGBM-SHAP model constructed in the present invention quantifies the influence of features on the short-term prediction of surface displacement in a consistent manner. By making full use of the interpretability mechanism of the SHAP model, the key features affecting surface displacement prediction are mined through global interpretability, and the interactive influence of different features on the prediction results is analyzed through feature interaction analysis, improving the credibility of the prediction results and effectively enhancing the interpretability of the model. Description of the Drawings
[0051] Figure 1 Flow chart of interpretable LightGBM-SHAP surface displacement prediction;
[0052] Figure 2 Graph of the relationship between displacement in the Y direction, soil moisture content, and rainfall;
[0053] Figure 3 Heat map of feature correlation based on spearman;
[0054] Figure 4 Process diagram of hyperparameter optimization for parameter optimization using the tree structure probability density estimation algorithm;
[0055] Figure 5 Comparison graph of typical day prediction results;
[0056] Figure 6 Comparison graph of evaluation indicators of different models;
[0057] Figure 7 Global feature density scatter plot based on SHAP values;
[0058] Figure 8 Feature interaction graph between 12-hour average rainfall and historical displacement in the y direction;
[0059] Figure 9 Feature interaction graph between displacement in the x direction and historical displacement values in the y direction. Detailed Implementation Manner
[0060] The core of the present invention is to provide a method for predicting landslide surface displacement with interpretability. Specifically, it includes the following steps:
[0061] Step 1: Obtain landslide monitoring data, including displacement data, precipitation, mud level, and moisture content at different depths at different positions of the slope body;
[0062] Step 2: Perform data preprocessing on multivariate data, including outlier removal and missing value filling;
[0063] Step 3: According to the known attributes of the data processed in Step 2, perform feature engineering, including feature screening and feature construction;
[0064] Step 4: Construct a model using the multiple input items obtained in Step 3 as influencing factors; specifically, it includes the following steps:
[0065] Step 4-1: Divide the training data set and the test set, and use the training data set to construct a LightGBM landslide surface displacement prediction model;
[0066] Step 4-2: Use the tree structure probability density estimation algorithm to train the constructed LightGBM landslide surface displacement prediction model and optimize the model parameters;
[0067] Step 4-3: Select the model with the optimal accuracy as the final prediction model and output the prediction result;
[0068] Step 5: Evaluate the prediction model using three evaluation indicators of MSE, MAE, and R2;
[0069] Step 6: Calculate the marginal contribution of each input feature to the model result through the SHAP method, so as to explain the influence degree of different input features on the surface displacement.
[0070] The related process is as Figure 1 shown. To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0071] Example:
[0072] Adopt an interpretable LightGBM landslide surface displacement short-term prediction model based on SHAP to conduct short-term prediction of the surface displacement of a landslide monitoring site in Shennongjia, Hubei Province.
[0073] As described in Step 1, first collect the measured data of a certain unstable slope from February 27th to July 22nd of a certain year based on a GNSS displacement monitor, a soil moisture content meter, a rain gauge, and a mud water level gauge, with a time resolution of 1 hour.
[0074] As described in Step 2, conduct outlier inspection and missing value filling on the collected data. Among them, the y-direction displacement, the soil moisture content at a depth of 30 cm in the middle of the landslide, and the rainfall data after outlier inspection and missing value filling are as Figure 2 shown.
[0075] As described in Step 3, use the Spearman correlation coefficient method to calculate the correlation coefficient between the monitoring data and the surface displacement value. The Spearman correlation heat map between each monitoring data and the surface displacement amount is as Figure 3 shown, Figure 3Among them, the larger the value, the darker the color indicates a higher correlation between variables. x, y, and z represent the ground displacement in three directions; tyl represents the rainfall; hs1, hs2, and hs3 represent the soil moisture content at different depths and different positions of the slope body; nsw represents the mud level. From Figure 3 it can be seen that the correlation between hs1, nsw and the ground displacement in the y direction is very weak (less than 0.2). To avoid data redundancy, they are not used as input features.
[0076] Then, considering the lag effect of the ground displacement response to rainfall and the autocorrelation of time series data, feature construction is performed on the selected input data, and recent features and long-term features are calculated through historical feature data. The specific operation is to calculate the mean value every 6h, every 12h, and every 24h for each existing feature, and add the calculation results as features to the data set, so as to increase the number of features and improve the model training effect. The final model input features are shown in Table 1 below:
[0077] Table 1 Model input feature
[0078] Table 1 Model input feature
[0079]
[0080] As described in Step 4-1, in this example, the model uses a 4:1 ratio to divide the training set and the test set, and adopts 5-fold forward validation during the training process, that is, the training set is further divided into a validation set. After the constructed LightGBM landslide ground displacement prediction model fits and trains the training set and obtains a certain number of predicted values in the validation set, the true observed values from the validation set are added to the original training set to make it a new training set, and then a certain number of predictions are made on the new validation set, and the model is readjusted. This process is repeated on the entire training set, and finally the optimal model is trained.
[0081] Table 2 LightGBM parameter and optimization value
[0082] Table 2 LightGBM parameter and optimization value
[0083]
[0084] As shown in Table 2, in this example, the LightGBM algorithm includes four parameters: the number of leaves num_leaves, the number of base learners n_estimators, the learning rate learning_rate, and the maximum tree depth max_depth. Within the set parameter range, TPE optimization, random search optimization, and genetic algorithm (GA) search optimization are respectively used. Taking the mean squared error (MSE) of forward validation as the optimization objective function, training is iterated 50 times to obtain the optimized parameter values of LightGBM.
[0085] As described in step 4-2, the MSE change curve of the training process is plotted as Figure 4 shown. After 50 iterations, the TPE optimization takes 15.7 s and finds the optimal parameters at the 26th time, the random search takes 20.5 s and finds the optimal parameters at the 40th time, the genetic algorithm search takes 24.5 s and finds the optimal parameters at the 43rd time. Moreover, the MSE values of the random search (0.187 mm) and the genetic algorithm search (0.196 mm) are both greater than the MSE value of the TPE optimization (0.172 mm). The experimental results show that the TPE optimization can obtain the optimal parameter combination in less time, performs better than the random search and the genetic algorithm search in this experiment, and has more advantages in the parameter optimization of LightGBM.
[0086] As described in step 4-3, in order to test the model prediction effect, the data from July 5th to July 8th, 2023 during the rapid deformation period is selected for further analysis. The LightGBM-SHAP model established in this paper is used to realize the prediction of landslide surface displacement, and the predicted values are compared with the predicted values of random forest regression (RFR) and multiple linear regression (MLR). From Figure 5 it can be seen that the prediction results of the LightGBM-SHAP model fit well with the actual values. Combining the rainfall monitoring data, it is found that the landslide deformation is significantly affected by rainfall during this period. After the rapid deformation caused by rainfall ends, the deformation will gradually return to stability. The attention of the model to the displacement peak helps to grasp the overall displacement trend.
[0087] As described in step 5, Figure 6 The comparison of the average R2, RMSE, and MAE calculated by running the three models multiple times using the same dataset is shown. From Figure 6 it can be seen that the prediction ability of the lightweight gradient multiple regression model is significantly stronger than that of the RFR model and the MLR model. In summary, the multi-factor landslide surface displacement prediction model based on LightGBM-SHAP in this paper has both strong generalization ability and high accuracy.
[0088] As described in step 6-1, Figure 7It is a scatter plot of feature density calculated by the SHAP method. In the figure, the feature importance gradually decreases from top to bottom. The scatter color from blue to red represents the value of the feature increasing from small to large. Each point represents the SHAP value of a sample, which represents the contribution of this feature to a single prediction, and the set of points represents the direction and magnitude of the overall influence of the feature on the prediction result. From Figure 7 It can be seen that among the 16 input features used in this example, historical surface displacement and displacements in the x and z directions are the most important features affecting the model prediction result, reflecting the autocorrelation of surface displacement data. Secondly, the average historical rainfall is considered, and the influence of the remaining variables on surface displacement is relatively small. Moreover, the feature contribution degrees of the average rainfall and average moisture content features after feature construction to displacement prediction are higher than those of the original features before feature construction.
[0089] As described in step 6-2, Figure 8 、 Figure 9 shows the feature interaction diagram of the 12-hour average rainfall, displacement in the x direction, and historical displacement value in the y direction. In the figure, the abscissa represents the magnitude of the selected feature value, the left axis represents the SHAP value calculated through this feature, and the right vertical axis represents the magnitude of the feature value with the largest interaction with this feature. The color change from blue to red in the figure indicates the change of the feature value from small to large. From Figure 8 It can be seen that for samples where the value of tylp12 is relatively large and the corresponding historical displacement value in the y direction is also relatively large (negative direction), the smaller their SHAP values are, the larger (negative direction) the predicted surface displacement value is, and the greater the contribution to the prediction of the y displacement value. From Figure 9 It can be seen that the displacement in the x direction shows a non-linear influence on the prediction result of the surface displacement in the y direction. When the absolute value of x is larger, the absolute value of its SHAP is larger, and the greater the contribution of this sample to the final prediction result of the model.
[0090] It can be seen that the LightGBM-SHAP model used in the present invention can fully reflect the characteristics of the landslide itself and the influence of external factors in predicting landslide displacement. It has significant advantages in aspects such as dataset division and controlling model overfitting, can accurately predict future landslide deformation values with relatively few data in terms of time span, and at the same time makes the prediction interpretable.
Claims
1. An interpretable landslide surface displacement prediction method based on LightGBM and SHAP, characterized in that: The steps include: Step 1: Obtain landslide monitoring data, including displacement data, precipitation, infrasound, and moisture content at different locations and depths of the slope; Step 2: Data preprocessing for multivariate data by removing outliers and filling missing values; Step 3: Based on the known attributes of the data processed in step 2, perform feature screening and feature construction; Step 4: Build a model based on the multiple input items obtained in step 3 as influencing factors; specifically, the following steps are included: Step 4-1: Divide the training data set and the test set, and use the training data set to build the LightGBM landslide surface displacement prediction model; Step 4-2: Use the tree structure probability density estimation algorithm to train the LightGBM landslide surface displacement prediction model constructed in step 4-1 and optimize the model parameters; Step 4-3: Select the model with the best accuracy as the final prediction model and output the prediction results; Step 5: Use RMSE\MAE\R2 three evaluation indicators to evaluate the prediction model; Step 6: The marginal contribution of each input feature to the model results is calculated using the SHAP method, which can explain the degree of influence of different input features on surface displacement.
2. The method for predicting landslide surface displacement with interpretability based on LightGBM and SHAP according to claim 1 is characterized in that: The input features in step 1 include infrasound, soil moisture content at different locations and depths of the slope, and mud water level.
3. The method for predicting landslide surface displacement with interpretability based on LightGBM and SHAP according to claim 1 is characterized in that: In step 3, feature screening and feature construction are performed on multiple input items, specifically including the following steps: Step 3-1: Use the Spearman correlation coefficient method to calculate the correlation coefficient between the multivariate monitoring data and the landslide surface displacement, and remove the monitoring items with a correlation coefficient lower than 0.2 to reduce data redundancy and achieve feature screening. The calculation formula of the Spearman correlation coefficient is as follows: Among them, x i and i is the value of random variables X and Y when the observed value is i, and is the average of the random variables X and Y; Step 3-2: For the features retained after screening, perform feature construction and calculate their recent features and long-term features respectively as model input items.
4. The method for predicting landslide surface displacement with interpretability based on LightGBM and SHAP according to claim 3 is characterized in that: In step 3-2, feature construction is performed based on the screened features. The specific operations for constructing recent features and long-term features are as follows: the surface displacement takes the average displacement of the previous 6\12 hours, the cumulative rainfall takes the average cumulative rainfall of the previous 6\12\24 hours, and the soil moisture takes the average soil moisture of the previous 6\12\24 hours. The calculation results are added to the data set as new features.
5. The method for predicting landslide surface displacement with interpretability based on LightGBM and SHAP according to claim 1 is characterized in that: In step 4-1, the training data set and the test data set are divided; During the training process, the training set and the test set were divided into a training set and a test set according to a ratio of 4:
1. The training set was further divided into a training set and a validation set using a 5-fold forward validation. The average MSE value of the 5-fold forward validation was taken as the optimization objective function for model parameter adjustment.
6. The method for predicting landslide surface displacement with interpretability based on LightGBM and SHAP according to claim 1 is characterized in that: In step 4-1, the training data set is used to construct the LightGBM landslide surface displacement prediction model. LightGBM uses histogram optimization to segment continuous eigenvalues, and its decision tree searches for the leaf with the largest split gain for splitting each time, and prevents overfitting by limiting the depth of the tree. The core idea of LightGBM is to use the decision tree as a base learner to iteratively train the data to obtain the optimal model; the calculation is shown in formula (2): In formula 2, F T (x i ) is the predicted value of the i-th base learner; x i is the i-th sample, and υ is the set of learners.
7. The method for predicting landslide surface displacement with interpretability based on LightGBM and SHAP according to claim 1 is characterized in that: In the step 4-2, the constructed LightGBM landslide surface displacement prediction model is trained using a tree structure probability density estimation algorithm to optimize model parameters, wherein the optimized parameters are 4 parameters, namely, the number of cotyledons num_leaves, the number of base learners n_estimators, the learning rate learning_rate, and the maximum value of the tree depth max_depth; The tree structure probability density estimation algorithm used is to establish a proxy model by classification method. It does not directly estimate the output of each sample point, but fuzzily divides the sample points into two types: good and bad, and uses expected improvement as its acquisition function to generate new sampling points. The expected improvement calculation formula is: Where M is the proxy model constructed using the kernel density estimation method, y* represents a threshold, y represents the current optimization objective function value, and p M (y|x) represents the conditional probability of y calculated using model M under the value of x.
8. The method for predicting landslide surface displacement with interpretability based on LightGBM and SHAP according to claim 1 is characterized in that: In step 6, the calculation formula of the SHAP value is: Where: f(x) is the machine learning model, which is the LightGBM model in this patent; z*={0,1}, when feature i is observed, z*=1, otherwise it is 0; if i participates in the prediction process, M is the number of features; φ i is the contribution of feature i, expressed as: Where: φ i represents the SHAP value of feature i; N is the set of sample features that need to be explained, S is the set containing non-zero indexes in z*, M is the total number of different input features, and f(S∪{i})-f(S) is the contribution of feature i to the prediction result.
9. The method for predicting landslide surface displacement with interpretability based on LightGBM and SHAP according to claim 1 is characterized in that: In step 6, the marginal contribution of each input feature to the model result is calculated by the SHAP method, so as to explain the influence of different input features on the surface displacement, which specifically includes the following steps: Step 6-1: For the selected input feature, use SHAP to calculate the SHAP value of all samples of the feature, and use its average value as the global feature importance of the selected input feature, so as to obtain the global explanation of each input feature; Step 6-2: Use the positive and negative effects of SHAP values to analyze the interaction of different input features on the surface displacement prediction results to improve the interpretability of the model.
Citation Information
Patent Citations
Soil water content influence factor sensitive interval judgment method based on interpretable ensemble learning model
CN116205310A
Data-driven dynamic landslide susceptibility evaluation and disaster-inducing factor change inference method
CN117709583A