Power transmission line iron tower periphery landslide susceptibility prediction method
By optimizing the data sampling method and model hyperparameters, combining key factor screening and multiple machine learning models, the data sampling randomness and factor weight allocation problems in the prediction of landslides around the transmission line tower are solved, improving the accuracy and reliability of the prediction, and providing a scientific basis for geological disaster prevention and control.
Patent Information
- Application Number
- CN202510606981.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the method of predicting landslides around the tower of the transmission line has problems such as strong randomness in data sampling and difficult allocation of key factor weights, which affects insufficient prediction accuracy.
By collecting historical landslide data around the transmission line tower, determining the influencing factors and preprocessing, filtering key factors, building a variety of machine learning models, and using the capuchin monkey search algorithm to optimize the model hyperparameters, combining multiple data sampling methods and multicollinear analysis, optimizing data sampling and model hyperparameters, and building a landslide susceptibility prediction model.
It significantly improves the accuracy and reliability of landslide prone prediction, reduces prediction bias caused by improper factor selection, enhances the representativeness and generalization ability of the model, and provides scientific prediction basis and visual analysis.
Smart Images

Figure CN120471446A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of dynamic landslide assessment, and in particular relates to a method for predicting the susceptibility of landslides around a transmission line tower. Background Art
[0002] Nowadays, with improved living standards, people's expectations for the range and amount of power they can transmit have soared. In complex mountainous areas, transmission towers are often placed on hillside crests. However, with global warming and increased human engineering activities, slope stability is a crucial consideration. Landslides caused by slope instability are characterized by suddenness and widespread impact. Therefore, efficient and high-quality landslide susceptibility prediction around transmission towers is of great practical significance.
[0003] Among them, sample data collection methods and prediction model construction are key nodes in landslide susceptibility research. Data sampling methods are often highly subjective and random, such as random sampling and spatial analysis based on geographic information systems (GIS), which make it difficult to fully capture the complex mechanisms of landslide formation. At the same time, there are many landslide-causing factors, and how to improve the weight of key factors in the prediction process is an issue worthy of attention. On the other hand, the performance of the prediction model is strongly correlated with its own hyperparameter allocation. Among them, population-based heuristic optimization algorithms have become one of the best methods for rationally allocating hyperparameters due to their ability to simulate the habits of plants and animals to solve for the global optimal solution.
[0004] One existing method for studying landslide susceptibility under the combined effects of transmission lines and environmental factors involves extracting landslide data, dividing the grid, screening influencing factors, enhancing negative samples, and constructing a prediction model consisting of a base learner using a CART decision tree, a support vector machine, and a BP neural network, along with a logistic regression meta-learner. This model is then trained and applied to achieve landslide susceptibility prediction, addressing the uncertainty associated with negative samples in traditional methods. In this method, the selection of positive samples is equally critical in landslide susceptibility prediction, as these represent known landslide areas and their characteristics are crucial for the model to learn landslide occurrence patterns. Focusing solely on optimizing negative samples while ignoring the quality and representativeness of positive samples can lead to an inadequate understanding of the characteristics of landslide areas, thus affecting overall prediction accuracy.
[0005] Another existing approach constructs a standardized set of raster cells by obtaining hazard factor vectors or raster files from environmental factors. A PU-bagging landslide susceptibility prediction network and a similarity metric loss landslide susceptibility prediction network are then constructed. These networks are then used to generate pre-correction and corrected prediction probability scores. These scores are then combined to generate the final prediction probability score, and the landslide susceptibility prediction results are obtained using the natural breakpoint method. However, this approach has shortcomings in terms of diversity and optimization of data sampling methods, factor importance analysis, model optimization, and comprehensiveness of validation methods. Specifically, first, this approach only utilizes PU-bagging and similarity metric loss networks for landslide susceptibility prediction, failing to fully consider the impact of different sampling methods on prediction results. Second, this approach lacks a factor importance analysis process or a focus on key factors. Third, while this approach constructs PU-bagging and similarity metric loss landslide susceptibility prediction networks and improves prediction accuracy by correcting the prediction scores, there may still be room for improvement in model optimization. Summary of the Invention
[0006] In response to the above-mentioned deficiencies in the prior art, the method for predicting landslide susceptibility around transmission line towers provided by the present invention solves the problems of randomness in data sampling in existing landslide susceptibility prediction methods and difficulty in allocating weights of key factors in constructing the model, which in turn affects the accuracy of prediction.
[0007] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: a method for predicting the susceptibility of landslides around transmission line towers, comprising the following steps:
[0008] Collect historical landslide data around transmission line towers, identify influencing factors, and pre-process them;
[0009] Screen the influencing factors after pretreatment and determine the key influencing factors;
[0010] Sampling historical landslide data and constructing a sample data set based on the selected key influencing factors; the data sampling methods include sampling at 0 meters, sampling at 10 meters, and surface sampling;
[0011] Build at least three machine learning models, train them using various sample datasets, and optimize model hyperparameters using the Capuchin Search algorithm during training to obtain several preliminary prediction models.
[0012] Evaluate the performance of each preliminary prediction model, determine the optimal data sampling method and the combination of preliminary prediction models, and obtain a landslide susceptibility prediction model;
[0013] The sampling data of the area to be predicted obtained by the optimal data sampling method are input into the landslide susceptibility prediction model to obtain the landslide susceptibility prediction results, and the landslide susceptibility map is drawn by visual analysis.
[0014] Furthermore, the influencing factors include elevation, aspect, slope, plan curvature, water flow intensity index, profile curvature, sediment transport index, terrain roughness index, terrain wetness index, lithology distribution map, land type, vegetation cover index NDVI, rainfall, distance to rivers and distance to roads;
[0015] The preprocessing includes data standardization and normalization.
[0016] Furthermore, the method for screening the impact factors is specifically as follows:
[0017] Calculate the variance inflation factor of each influencing factor, determine the correlation between the influencing factors, and screen the available influencing factors through multicollinearity between the influencing factors;
[0018] Calculate the difference between each available influencing factor in landslide data samples and non-landslide data samples, and determine the contribution of each influencing factor to landslide susceptibility prediction;
[0019] According to the contribution of each available influencing factor to the landslide susceptibility prediction, the key influencing factors are screened and obtained.
[0020] Furthermore, the variance inflation factor of the impact factor is expressed as:
[0021]
[0022] The contribution of influencing factors to landslide susceptibility prediction is expressed as:
[0023]
[0024] Where W'(A) represents the contribution of influencing factor A to landslide susceptibility prediction, W(a) represents the weight of influencing factor A, R represents a randomly selected sample, H represents the nearest neighbor sample of the same type as R, M represents the nearest neighbor sample of a different type from R, diff(A, R, H) represents the difference between samples R and H in influencing factor A, diff(A, R, M) represents the difference between samples R and M in influencing factor A, and m represents the number of iterations.
[0025] VIF i represents the equation expansion factor of the i-th impact factor, It represents the coefficient of determination obtained by linear regression with the i-th influencing factor as the dependent variable and other influencing factors as independent variables.
[0026] Furthermore, the method of constructing the sample data set is:
[0027] The boundaries of the historical landslide data range are proportionally reduced inward by 0 meters and 10 meters respectively; 0 meters and 10 meters represent the vertical distances between the two boundary lines before and after scaling;
[0028] Sampling was performed at 0 m from the boundary between the landslide and non-landslide areas, and at 10 m from the boundary between the landslide and non-landslide areas to obtain the corresponding 0 m sample data set and 10 m sample data set;
[0029] In the 0-meter sample data set and the 10-meter sample data set, the number of minimum surface area units and the number of sample points in each sampling sample in the sample data set to which it belongs are calculated, and then random sampling is performed in the entire surface area of the landslide and non-landslide areas to obtain the corresponding surface area sample data set.
[0030] Furthermore, the method for obtaining the preliminary prediction model is as follows:
[0031] Determine the type of machine learning model and its hyperparameters that need to be optimized;
[0032] Train machine learning models using sample datasets;
[0033] During the training process, the Capuchin search algorithm is used to optimize the hyperparameters of the machine learning model, and the fitness function is used to evaluate the performance of the hyperparameters, thereby dynamically adjusting the search strategy;
[0034] Repeat the model training and hyperparameter optimization process until the maximum number of iterations is reached, determine the optimal hyperparameters of the model, and then obtain a preliminary prediction model;
[0035] Among them, each sample data set and each machine learning model are trained in any combination to obtain a corresponding preliminary prediction model.
[0036] Furthermore, the formula for optimizing model hyperparameters is:
[0037]
[0038] Where, represents the position of the i-th capuchin monkey in the j-th dimension at the t-th iteration, represents the velocity of the jth dimension of the i-th capuchin monkey at the t+1th iteration, w represents the inertia weight, c1 and c2 represent the learning factors, r1 and r2 represent random numbers, represents the best historical position of the i-th capuchin monkey in the j-th dimension, Represents the global optimal position; where the ordinal number i of the capuchin represents the model hyperparameter combination, and j represents the ordinal number of the model hyperparameter combination.
[0039] Furthermore, the machine learning model includes a KNN model, a random forest model and an XGB model;
[0040] The optimization hyperparameter of the KNN model is the k value, the optimization hyperparameter of the random forest model is the number of trees, and the optimization hyperparameter of the XGB model is the learning rate.
[0041] Furthermore, the performance evaluation indicators of the preliminary prediction model include accuracy, precision, recall, F1 value and Kappa coefficient.
[0042] The beneficial effects of the present invention are:
[0043] By optimizing data sampling, key factor selection, model hyperparameter optimization, and comprehensive model validation, this paper significantly improves the accuracy and reliability of landslide susceptibility prediction, and has important practical application value, as embodied in the following aspects:
[0044] (1) Through standardization and normalization, the data dimension is unified, the error caused by inconsistent data dimension is reduced, and the stability and prediction accuracy of subsequent model training are improved.
[0045] (2) By calculating the contribution of influencing factors to landslide susceptibility prediction, the importance of key influencing factors was screened and determined, focusing on factors such as slope, aspect, rainfall, normalized vegetation index, elevation, and distance to rivers. This improved the accuracy of the model in predicting landslide susceptibility and avoided prediction bias caused by inappropriate factor selection.
[0046] (3) The sample datasets were constructed by using multiple data sampling methods, which reduced the errors caused by sample bias and spatial heterogeneity, significantly improved the representativeness of the landslide and non-landslide sample datasets, and were able to more comprehensively capture the complex mechanisms of landslide formation compared to traditional random sampling and GIS-based spatial analysis.
[0047] (4) The Capuchin Search Algorithm (Capsa) was combined with the hyperparameters of the K-nearest neighbor (KNN), random forest (RF), and extreme gradient boosting tree (XGB) models to significantly improve the generalization ability and prediction performance of the models. The Capsa algorithm simulates the dynamic behavior of capuchin monkeys during foraging and can balance global and local search in the solution space, thereby improving the prediction accuracy of the models.
[0048] (5) The model was comprehensively verified and evaluated through ten-fold cross-validation and multiple evaluation indicators (such as Accuracy, F1, Kappa, ROC curve, and PR curve), ensuring the stability and reliability of the model on different data sets and providing a scientific basis for landslide susceptibility prediction;
[0049] (6) Through visual analysis, a landslide susceptibility map was drawn, which intuitively displayed the prediction results of the model under different sampling methods, providing a scientific basis for geological disaster prevention and control and land use planning, and providing operational technical support for practical engineering applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 This is a flow chart of the method for predicting the susceptibility of landslides around transmission line towers in the present invention.
[0051] Figure 2 Schematic diagram of data sampling and sample data set construction in the present invention. DETAILED DESCRIPTION
[0052] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0053] The present invention aims at predicting the susceptibility of landslides under the combined effects of transmission lines and environmental factors, and realizes high-precision prediction of landslide susceptibility. Figure 1 As shown, the following steps are included:
[0054] Collect historical landslide data around transmission line towers, identify influencing factors, and pre-process them;
[0055] Screen the influencing factors after pretreatment and determine the key influencing factors;
[0056] Sampling historical landslide data and constructing a sample data set based on the selected key influencing factors; the data sampling methods include sampling at 0 meters, sampling at 10 meters, and surface sampling;
[0057] Build at least three machine learning models, train them using various sample datasets, and optimize model hyperparameters using the Capuchin Search algorithm during training to obtain several preliminary prediction models.
[0058] Evaluate the performance of each preliminary prediction model, determine the optimal data sampling method and the combination of preliminary prediction models, and obtain a landslide susceptibility prediction model;
[0059] The sampling data of the area to be predicted obtained by the optimal data sampling method are input into the landslide susceptibility prediction model to obtain the landslide susceptibility prediction results, and the landslide susceptibility map is drawn by visual analysis.
[0060] The present invention significantly improves the accuracy and reliability of landslide susceptibility prediction by optimizing data sampling methods, key factor selection, model hyperparameter optimization and comprehensive model verification, and has important practical application value.
[0061] The implementation of the above landslide susceptibility prediction method will be further described below in conjunction with Examples 1 to 6.
[0062] Example 1:
[0063] In this embodiment, the basic data required for landslide susceptibility prediction and its preprocessing process are described.
[0064] In this embodiment, historical landslide data are collected, and the influencing factors include elevation, slope aspect, slope gradient, plan curvature, water flow intensity index, profile curvature, sediment transport index, terrain roughness index, terrain wetness index, lithology distribution map, land type, vegetation cover index NDVI, rainfall, distance to the river, and distance to the road.
[0065] In this embodiment, the preprocessing of the collected data includes data standardization and normalization processing; specifically, the formulas for standardization and normalization processing are as follows:
[0066]
[0067]
[0068] In the formula, z represents the standardized value, x represents the original data value, μ represents the mean of the data, σ represents the standard deviation of the data, and x represents the standard deviation of the data. norm Represents the normalized value, x min Indicates the minimum value of the data, x max Indicates the maximum value of the data.
[0069] In this embodiment, by standardizing and normalizing the collected data, the uniformity of the data dimension is ensured, the error caused by inconsistent data dimension is reduced, and the stability and prediction accuracy of subsequent model training are improved.
[0070] Example 2:
[0071] In this embodiment, the process of screening the determined influencing factors based on the collected landslide data to determine the key influencing factors is described in detail.
[0072] In this embodiment, the method for screening the impact factors is specifically as follows:
[0073] Calculate the variance inflation factor of each influencing factor, determine the correlation between the influencing factors, and screen the available influencing factors through multicollinearity between the influencing factors;
[0074] Calculate the difference between each available influencing factor in landslide data samples and non-landslide data samples, and determine the contribution of each influencing factor to landslide susceptibility prediction;
[0075] According to the contribution of each available influencing factor to the landslide susceptibility prediction, the key influencing factors are screened and obtained.
[0076] In this example, the variance inflation factor (VIF) of the impact factor is used to quantify the multicollinearity problem between features (influence factors). Multicollinearity refers to the high correlation between features, which may lead to model instability or reduced explanatory power. By calculating the VIF value, the linear correlation between each impact factor and other impact factors can be evaluated, thereby screening out impact factors with strong independence. If the multicollinearity between the impact factors is below the preset threshold, it is proved that all impact factors can be used, and the available impact factors are screened.
[0077] Among them, the variance inflation factor of the impact factor is expressed as:
[0078]
[0079] Where VIF i represents the equation expansion factor of the i-th impact factor, It represents the coefficient of determination obtained by linear regression with the i-th influencing factor as the dependent variable and other influencing factors as independent variables.
[0080] In this embodiment, based on the available influencing factors obtained through correlation screening, the importance of each influencing factor to the landslide susceptibility prediction is evaluated by calculating the contribution of each influencing factor to the landslide susceptibility prediction. The core idea is to calculate the difference between the influencing factors in landslide samples and non-landslide samples. The relevant calculation formula is:
[0081]
[0082] where W′(A) represents the contribution of influencing factor A to landslide susceptibility prediction, W(A) represents the weight of influencing factor A, R represents a randomly selected sample, H represents the nearest neighbor sample of the same type as R, M represents the nearest neighbor sample of a different type from R, diff(A, R, H) represents the difference between samples R and H in influencing factor A, diff(A, R, M) represents the difference between samples R and M in influencing factor A, and m represents the number of iterations.
[0083] In this embodiment, the influencing factors are screened and the importance of key influencing factors is determined by calculating the variance inflation factor and contribution of the influencing factors. The key influencing factors obtained by screening include slope, aspect, rainfall, normalized vegetation index, elevation, distance to the river and other factors; through the above-mentioned screening process of influencing factors, the accuracy of the model in predicting landslide susceptibility is improved, and the prediction deviation caused by improper selection of influencing factors is avoided.
[0084] Example 3:
[0085] In this embodiment, a method for constructing a sample data set by data sampling is described in detail.
[0086] In this embodiment, Figure 2 As shown, the method to construct the sample data set is:
[0087] The boundaries of the historical landslide data range are proportionally reduced inward by 0 meters and 10 meters respectively; 0 meters and 10 meters represent the vertical distances between the two boundary lines before and after scaling;
[0088] Sampling was performed at 0 m from the boundary between the landslide and non-landslide areas, and at 10 m from the boundary between the landslide and non-landslide areas to obtain the corresponding 0 m sample data set and 10 m sample data set;
[0089] In the 0-meter sample data set and the 10-meter sample data set, the number of minimum surface area units and the number of sample points in each sampling sample in the sample data set to which it belongs are calculated, and then random sampling is performed in the entire surface area of the landslide and non-landslide areas to obtain the corresponding surface area sample data set.
[0090] This embodiment provides a specific example of area sampling. During the area sampling process in the landslide area and the non-landslide area, the smallest area in the landslide and non-landslide samples is determined respectively. Based on this, the number of areas that can be sampled in the remaining areas can be calculated. For example, there are three areas with sizes of 2, 6, and 8, respectively. The corresponding sampling points of the three areas are 1 (2 / 2), 3 (6 / 2), and 4 (8 / 2), respectively.
[0091] In this example, data sampling was performed using the three aforementioned methods. When constructing a sample dataset, the sampled data and influencing factors were combined using a GIS tool through sample point attribute extraction to construct the corresponding sample dataset. Combined with the prediction results of the subsequent model construction, the advantages and disadvantages of different data sampling methods, such as linear versus planar sampling, can be determined, as well as the importance of selecting the location of baselines (e.g., 0 meters and 10 meters).
[0092] In this embodiment, for the sample data set constructed by each data sampling method, the data set is divided into a training set, a test set, and a validation set in a ratio of 6:2:2 to participate in subsequent model training.
[0093] In this embodiment, multiple data sampling methods are provided to construct corresponding sample data sets. Among them, the area sampling method reduces the errors caused by sample bias and spatial heterogeneity, significantly improves the representativeness of landslide and non-landslide sample data sets, and can more comprehensively capture the complex mechanism of landslide formation compared with traditional random sampling and GIS-based spatial analysis.
[0094] Example 4:
[0095] This embodiment focuses on the process of constructing and training a preliminary prediction model for landslide susceptibility prediction.
[0096] In this embodiment, the method of constructing a machine learning model, training and obtaining a preliminary prediction model is specifically as follows:
[0097] Determine the type of machine learning model and its hyperparameters that need to be optimized;
[0098] Train machine learning models using sample datasets;
[0099] During the training process, the Capuchin search algorithm is used to optimize the hyperparameters of the machine learning model, and the fitness function is used to evaluate the performance of the hyperparameters, thereby dynamically adjusting the search strategy;
[0100] Repeat the model training and hyperparameter optimization process until the maximum number of iterations is reached, determine the optimal hyperparameters of the model, and then obtain a preliminary prediction model;
[0101] Among them, each sample data set and each machine learning model are trained in any combination to obtain a corresponding preliminary prediction model.
[0102] In this embodiment, the machine learning models include KNN model, random forest model and XGB model;
[0103] The optimization hyperparameter of the KNN model is the k value, the optimization hyperparameter of the random forest model is the number of trees, and the optimization hyperparameter of the XGB model is the learning rate.
[0104] In the process of training the machine learning model constructed using the sample data set, the capuchin search algorithm is used to optimize the model's hyperparameters. The capuchin search algorithm simulates the dynamic behaviors of capuchins such as jumping, swinging, and climbing during foraging, balances global search and local search in the solution space, and thus efficiently finds the optimal solution; and evaluates the performance of each set of hyperparameters through the fitness function, and dynamically adjusts the search strategy, ultimately finding the appropriate hyperparameter combination, significantly improving the model's prediction accuracy and generalization ability.
[0105] Among them, the formula for optimizing the hyperparameters of the Capuchin search algorithm model is:
[0106]
[0107] Where, represents the position of the i-th capuchin monkey in the j-th dimension at the t-th iteration, represents the velocity of the jth dimension of the i-th capuchin monkey at the t+1th iteration, w represents the inertia weight, c1 and c2 represent the learning factors, r1 and r2 represent random numbers, represents the best historical position of the i-th capuchin monkey in the j-th dimension, Represents the global optimal position; where the ordinal number i of the capuchin represents the model hyperparameter combination, and j represents the ordinal number of the model hyperparameter combination.
[0108] In this embodiment, a corresponding preliminary prediction model is trained for each machine learning model constructed for each sample data set, and preliminary prediction models under different sampling methods and different machine learning model combinations are obtained, so that the optimal combination method can be selected through subsequent model prediction for the actual landslide susceptibility prediction process.
[0109] In this example, the capuchin search algorithm (Capsa) is combined to optimize the hyperparameters of the K-nearest neighbor (KNN), random forest (RF), and extreme gradient boosting tree (XGB) models, significantly improving the generalization ability and predictive performance of the models. The Capsa algorithm simulates the dynamic behavior of capuchins during foraging and can balance global and local searches in the solution space, thereby improving the prediction accuracy of the models.
[0110] Example 5:
[0111] In this embodiment, the performance of the constructed preliminary prediction model is evaluated to select the optimal model as the landslide susceptibility prediction model.
[0112] In this embodiment, the performance evaluation indicators for the preliminary prediction model include accuracy, precision, recall, F1 value and Kappa coefficient.
[0113] Accuracy represents the proportion of samples correctly predicted by the model to the total number of samples. Precision represents the proportion of samples predicted by the model to be positive that are actually positive. It focuses on the accuracy of the prediction results, that is, how many samples predicted by the model to be positive are actually positive. Recall represents the proportion of samples correctly predicted to be positive by the model to be positive among samples that are actually positive. It focuses on the model's coverage of positive samples, that is, how many true positive samples the model can find. The F1 value is the harmonic mean of precision and recall, and is used to measure the balanced performance of the model. The Kappa coefficient is used to measure the consistency of the model's prediction results with random predictions.
[0114] In this embodiment, the calculation formulas for accuracy, precision, recall, F1 value and Kappa coefficient are:
[0115]
[0116] In this embodiment, the model was comprehensively verified and evaluated through multiple evaluation indicators (such as Accuracy, F1, Kappa, ROC curve and PR curve), ensuring the stability and reliability of the model on different data sets, and providing a scientific basis for landslide susceptibility prediction.
[0117] Example 6:
[0118] In this embodiment, based on the landslide susceptibility prediction model obtained by combining the optimal data sampling method and the preliminary prediction model determined by the above-mentioned performance evaluation, the sampling data of the area to be predicted obtained by sampling the optimal data sampling method is input into it to obtain the landslide susceptibility prediction result, and the landslide susceptibility prediction result is visualized and analyzed to draw a landslide susceptibility map.
[0119] In this embodiment, a landslide susceptibility map of the area to be predicted is drawn through visual analysis, and the prediction results of the model under different sampling methods are intuitively displayed, which provides a scientific basis for geological disaster prevention and control and land use planning, and provides operational technical support for actual engineering applications.
[0120] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
[0121] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.
Claims
1. A method for predicting landslide susceptibility around transmission line towers, characterized in that: The following steps are involved: Collect historical landslide data around transmission line towers, identify influencing factors, and pre-process them; Screen the influencing factors after pretreatment and determine the key influencing factors; Sampling historical landslide data and constructing a sample data set based on the selected key influencing factors; the data sampling methods include sampling at 0 meters, sampling at 10 meters, and surface sampling; Build at least three machine learning models, train them using various sample datasets, and optimize model hyperparameters using the Capuchin Search algorithm during training to obtain several preliminary prediction models. Evaluate the performance of each preliminary prediction model, determine the optimal data sampling method and the combination of preliminary prediction models, and obtain a landslide susceptibility prediction model; The sampling data of the area to be predicted obtained by the optimal data sampling method are input into the landslide susceptibility prediction model to obtain the landslide susceptibility prediction results, and the landslide susceptibility map is drawn by visual analysis.
2. The method according to claim 1, characterized in that The influencing factors include elevation, aspect, slope, plan curvature, water flow intensity index, profile curvature, sediment transport index, terrain roughness index, terrain wetness index, lithology distribution map, land type, vegetation cover index NDVI, rainfall, distance to rivers and distance to roads; The preprocessing includes data standardization and normalization.
3. The method according to claim 1, characterized in that The method for screening the impact factors is specifically as follows: Calculate the variance inflation factor of each influencing factor, determine the correlation between the influencing factors, and screen the available influencing factors through multicollinearity between the influencing factors; Calculate the difference between each available influencing factor in landslide data samples and non-landslide data samples, and determine the contribution of each influencing factor to landslide susceptibility prediction; According to the contribution of each available influencing factor to the landslide susceptibility prediction, the key influencing factors are screened and obtained.
4. The method according to claim 3, characterized in that The variance inflation factor of the impact factor is expressed as: The contribution of influencing factors to landslide susceptibility prediction is expressed as: Where W'(A) represents the contribution of influencing factor A to landslide susceptibility prediction, W(A) represents the weight of influencing factor A, R represents a randomly selected sample, H represents the nearest neighbor sample of the same type as R, M represents the nearest neighbor sample of a different type from R, diff(A, R, H) represents the difference between samples R and H in influencing factor A, diff(A, R, M) represents the difference between samples R and M in influencing factor A, and m represents the number of iterations. VIF i represents the equation expansion factor of the i-th impact factor, It represents the coefficient of determination obtained by linear regression with the i-th influencing factor as the dependent variable and other influencing factors as independent variables.
5. The method according to claim 1, wherein The method to construct the sample dataset is: The boundaries of the historical landslide data range are proportionally reduced inward by 0 meters and 10 meters respectively; 0 meters and 10 meters represent the vertical distances between the two boundary lines before and after scaling; Sampling was performed at 0 m from the boundary between the landslide and non-landslide areas, and at 10 m from the boundary between the landslide and non-landslide areas to obtain the corresponding 0 m sample data set and 10 m sample data set; In the 0-meter sample data set and the 10-meter sample data set, the number of minimum surface area units and the number of sample points in each sampling sample in the sample data set to which it belongs are calculated, and then random sampling is performed in the entire surface area of the landslide and non-landslide areas to obtain the corresponding surface area sample data set.
6. The method according to claim 1, characterized in that The method for obtaining the preliminary prediction model is as follows: Determine the type of machine learning model and its hyperparameters that need to be optimized; Train machine learning models using sample datasets; During the training process, the Capuchin search algorithm is used to optimize the hyperparameters of the machine learning model, and the fitness function is used to evaluate the performance of the hyperparameters, thereby dynamically adjusting the search strategy; Repeat the model training and hyperparameter optimization process until the maximum number of iterations is reached, determine the optimal hyperparameters of the model, and then obtain a preliminary prediction model; Among them, each sample data set and each machine learning model are trained in any combination to obtain a corresponding preliminary prediction model.
7. The method according to claim 6, characterized in that The formula for optimizing model hyperparameters is: Where, represents the position of the i-th capuchin monkey in the j-th dimension at the t-th iteration, represents the velocity of the jth dimension of the i-th capuchin monkey at the t+1th iteration, w represents the inertia weight, c1 and c2 represent the learning factors, r1 and r2 represent random numbers, represents the best historical position of the i-th capuchin monkey in the j-th dimension, Represents the global optimal position; where the ordinal number i of the capuchin represents the model hyperparameter combination, and j represents the ordinal number of the model hyperparameter combination.
8. The method according to claim 1, characterized in that The machine learning models include KNN model, random forest model and XGB model; The optimization hyperparameter of the KNN model is the k value, the optimization hyperparameter of the random forest model is the number of trees, and the optimization hyperparameter of the XGB model is the learning rate.
9. The method according to claim 1, characterized in that The performance evaluation indicators of the preliminary prediction model include accuracy, precision, recall, F1 value and Kappa coefficient.