Landslide susceptibility evaluation method based on support vector machine model
By using the support vector machine model, the problem of insufficient consideration of factor weight relationships in landslide susceptibility assessment was solved, resulting in a more accurate assessment of landslide susceptibility. The support vector machine model improves the scientificity and reliability of geological disaster prevention and control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LANZHOU UNIV
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies fail to adequately consider the weighting relationships between various evaluation factors in landslide susceptibility assessments, leading to inaccurate results.
A support vector machine (SVM) model was adopted. By collecting basic data, constructing an indicator system, training the SVM model, and using the best model, landslide susceptibility was predicted, and a landslide susceptibility zoning map was generated.
It improves the accuracy of landslide susceptibility assessment and provides a reliable and accurate reference for geological disaster prevention and control.
Smart Images

Figure CN121997156A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of geological hazard assessment technology, specifically relating to a landslide susceptibility assessment method based on a support vector machine model. Background Technology
[0002] my country's terrain slopes from west to east, with diverse landforms. In areas prone to sandstorms, loess deposits, and karst topography, various geological disasters are frequent. my country's vast territory and varied terrain create complex natural geological environments conducive to geological disasters. The nature and intensity of human engineering activities also differ across regions, resulting in significant variations in the types, intensity, and severity of geological disasters. my country has become one of the countries most prone to geological disasters in the world. Landslides are widespread in China, making them a key focus of our geological disaster research. Geological disaster susceptibility assessment studies the likelihood of geological disasters occurring in a given area, typically using gradations of extremely high, high, medium, low, and extremely low susceptibility. Accurate prediction of potential landslides before they occur allows for timely prevention and control measures, thus avoiding or mitigating potential losses and harm.
[0003] In recent years, with the development of remote sensing technology, geographic information systems, and computer technology, the accuracy of landslide susceptibility assessment results has been continuously improving. Currently, many practical production units use information-based models, which may not fully consider the weight relationships between various assessment factors, potentially leading to inaccurate susceptibility zoning results. Therefore, machine learning methods are being widely adopted. This study therefore selects a machine learning model for susceptibility assessment. Summary of the Invention
[0004] To address the above problems, this invention proposes a landslide susceptibility evaluation method based on a support vector machine model.
[0005] The technical solution of this invention is: a landslide susceptibility assessment method based on a support vector machine model, comprising the following steps:
[0006] S1. Collect basic data of the area to be studied, process and analyze the basic data, and construct an indicator system;
[0007] S2. Based on the known distribution of landslide points, randomly select an equal number of non-landslide samples in the non-landslide area;
[0008] S3. Based on the sample points and index system of non-slope samples, the support vector machine is trained to obtain the optimal support vector machine model;
[0009] S4. Use the optimal support vector machine model to predict the landslide susceptibility of the study area, and generate a landslide susceptibility zoning map based on the prediction results.
[0010] Furthermore, S1 includes the following sub-steps:
[0011] S11. Collect basic data of the area to be studied and standardize the basic data;
[0012] S12. Perform principal component analysis and Spearman correlation analysis on the standardized basic data to obtain several evaluation factors.
[0013] S13. Select stratigraphic lithology, landform type, peak ground acceleration, road density, fault density, normalized vegetation index, river network density, annual average precipitation, land type, slope aspect and elevation from several evaluation factors to construct an indicator system.
[0014] Furthermore, in S2, based on the known landslide sample points, ArcGIS's random point generation function is used to generate an equal number of non-landslide samples in the non-landslide area as the landslide points.
[0015] Furthermore, S3 includes the following sub-steps:
[0016] S31. Extract the attributes of several evaluation index factors of the index system to the landslide sample points.
[0017] S32. Construct a support vector machine and determine the optimization problem of the support vector machine;
[0018] S33. Process the optimization problem of the support vector machine to obtain the latest optimization problem;
[0019] S34. Based on the latest optimization problem, the support vector machine algorithm is used to train the sample points to obtain the optimal support vector machine model.
[0020] Furthermore, in S31, the expression for the support vector machine is:
[0021] ;
[0022] in, Describes the normal vector of the hyperplane. Indicates the offset top. This represents the feature vector of a sample.
[0023] Furthermore, the optimization problem of support vector machines is expressed as follows:
[0024] 2 1, ;
[0025] in, Describes the normal vector of the hyperplane. Indicates the offset top. Indicates the first The category label of each sample, Indicates the first The feature vector of each sample This represents the index of a sample in the dataset.
[0026] Furthermore, in S33, the expression for the latest optimization problem is:
[0027] ;
[0028] in, Describes the normal vector of the hyperplane. Indicates the offset top. Represents the regularization parameter. This represents the total number of training samples. Indicates the first Slack variables for each sample, Represents slack variables. Indicates the first The category label of each sample, Indicates the index of a sample in the dataset. Indicates the first The feature vector of each sample.
[0029] The beneficial effects of this invention are: the landslide susceptibility assessment method provided by this invention improves the accuracy of landslide susceptibility assessment results and can provide a reference for geological disaster prevention and control work. Therefore, this invention is highly reliable, accurate, and objective and scientific. Attached Figure Description
[0030] Figure 1 The flowchart shows a landslide susceptibility assessment method based on a support vector machine model.
[0031] Figure 2 This is a schematic diagram illustrating the principle of the landslide susceptibility assessment method according to an embodiment of the present invention.
[0032] Figure 3 This is a flowchart of the slope unit division process;
[0033] Figure 4 This is a zoning map showing the susceptibility to landslides in Sichuan Province. Detailed Implementation
[0034] The embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0035] like Figure 1As shown, this invention provides a landslide susceptibility assessment method based on a support vector machine model, comprising the following steps:
[0036] S1. Collect basic data of the area to be studied, process and analyze the basic data, and construct an indicator system;
[0037] S2. Based on the known distribution of landslide points, randomly select an equal number of non-landslide samples in the non-landslide area;
[0038] S3. Based on the sample points and index system of non-slope samples, the support vector machine is trained to obtain the optimal support vector machine model;
[0039] S4. Use the optimal support vector machine model to predict the landslide susceptibility of the study area, and generate a landslide susceptibility zoning map based on the prediction results.
[0040] In this embodiment of the invention, S1 includes the following sub-steps:
[0041] S11. Collect basic data of the area to be studied and standardize the basic data;
[0042] S12. Perform principal component analysis and Spearman correlation analysis on the standardized basic data to obtain several evaluation factors.
[0043] S13. Select stratigraphic lithology, landform type, peak ground acceleration, road density, fault density, normalized vegetation index, river network density, annual average precipitation, land type, slope aspect and elevation from several evaluation factors to construct an indicator system.
[0044] In this embodiment of the invention, in S2, based on known landslide sample points, the random point generation function of ArcGIS is used to generate an equal number of non-landslide samples in the non-landslide area as the landslide points.
[0045] In this embodiment of the invention, S3 includes the following sub-steps:
[0046] S31. Extract the attributes of several evaluation index factors of the index system to the landslide sample points.
[0047] S32. Construct a support vector machine and determine the optimization problem of the support vector machine;
[0048] S33. Process the optimization problem of the support vector machine to obtain the latest optimization problem;
[0049] S34. Based on the latest optimization problem, the support vector machine algorithm is used to train the sample points to obtain the optimal support vector machine model.
[0050] In this embodiment of the invention, in S31, the expression for the support vector machine is:
[0051] ;
[0052] in, Describes the normal vector of the hyperplane. Indicates the offset top. This represents the feature vector of a sample.
[0053] In this embodiment of the invention, the optimization problem of the support vector machine is expressed as:
[0054] 2 1, ;
[0055] in, Describes the normal vector of the hyperplane. Indicates the offset top. Indicates the first The category label of each sample (usually +1 or -1), where +1 represents a "landslide" sample and "-1" represents a "non-landslide" sample in this landslide evaluation method; Indicates the first The feature vector of each sample represents a grid cell in this landslide evaluation method; This represents the index of a sample in the dataset, used to iterate through all landslide and non-landslide sample points.
[0056] In this embodiment of the invention, in S33, the expression for the latest optimization problem is:
[0057] ;
[0058] in, Describes the normal vector of the hyperplane. Indicates the offset top. Represents the regularization parameter. This represents the total number of training samples. Indicates the first Slack variables for each sample, Represents slack variables. Indicates the first The category label of each sample, Indicates the index of a sample in the dataset. Indicates the first The feature vector of each sample.
[0059] Disaster-prone factors are selected based on the basic data of the area to be studied. The basic data includes geological data, remote sensing data, historical disaster data, road data, human activity data, land use type data, precipitation data, and water system distribution data.
[0060] Based on the obtained data and drawing on previous research, appropriate evaluation indicators were selected, and the data of each indicator were quantified using ArcGIS software. The quantified indicator data were then overlaid with slope units, and the original attribute values of 16 initial evaluation indicators were extracted to landslide hazard points. The original data were standardized using SPSS software. Principal component analysis and Spearman correlation analysis were performed on the standardized data to remove factors with high correlation. Finally, 11 factors were selected as evaluation indicators for this landslide susceptibility assessment: stratigraphic lithology, geomorphological type, peak ground acceleration, road density, fault density, normalized difference vegetation index, river network density, average annual precipitation, land type, slope aspect, and elevation. An indicator system was constructed based on these 11 evaluation factors.
[0061] Select landslide samples, and generate the same number of non-landslide samples using the random point generation function in ArcGIS. Reclassify the obtained raster data in ArcGIS, dividing it into five levels. Finally, extract the attributes of 11 evaluation index factors to the sample points. Install Python libraries in VS Code and construct a support vector machine model. The basic form of a support vector machine is as follows:
[0062] The goal of SVM is to find a hyperplane that separates data points of different classes. The margin maximization principle of SVM is to find a hyperplane that maximizes the distance from the nearest data point (support vector) to the hyperplane. This distance is called the margin, and can be represented as: Margin = .
[0063] To maximize the margin, SVM needs to minimize Simultaneously, the hyperplane must correctly classify all training data points. To solve the optimization problem, SVM uses the Lagrange multiplier method. The Lagrange function is defined as follows:
[0064] L( )= 2 - [ ];
[0065] in It is a Lagrange daily number.
[0066] Solve for the Lagrange dual function: L( ).
[0067] First of all, b. Take the partial derivative and set it to zero:
[0068] = =0;
[0069] ;
[0070] ;
[0071] ;
[0072] Will Substituting the Lagrange function, we obtain the dual problem:
[0073] ;
[0074] To handle non-linearly separable data, SVM introduces a kernel function K( , Common kernel functions include:
[0075] Linear kernel: K( , )= ;
[0076] Polynomial kernel: K( , )= ;
[0077] Radial basis function (RBF) kernel: K( , )=exp(-γ 2 );
[0078] After using a kernel function, the dual problem becomes:
[0079] 0, ;
[0080] In practical applications, data may not be perfectly linearly separable. Therefore, SVM introduces soft margins and a regularization parameter C, allowing some data points to violate the margin rule. The formulas and mathematical expressions of SVM are the core of its theoretical foundation. By maximizing the margin, using kernel functions, and introducing soft margins, SVM can effectively handle both linear and nonlinear classification problems.
[0081] In Python, the Scikit-Learn library can be used to implement SVM; the support vector machine algorithm is used to train the training data; the model parameters are determined, and the SVM parameters, such as kernel function type, penalty coefficient C, and kernel function parameter γ, are continuously adjusted to obtain the best model performance.
[0082] The model was validated and evaluated using test set data and receiver operating characteristic (ROC) curves.
[0083] The landslide susceptibility of the entire study area was predicted using a trained SVM model, and a landslide susceptibility zoning map was generated based on the prediction results.
[0084] The model's predictions were analyzed and compared with the actual landslide distribution to verify the model's accuracy.
[0085] The following description is based on specific embodiments.
[0086] Step 101: Select disaster-prone factors based on the basic data of the area to be studied. The basic data includes geological data, remote sensing data, historical disaster data, road data, human activity data, land use type data, precipitation data, and water system distribution data.
[0087] The evaluation factors include physical geography, basic geology, meteorology and hydrology, and human engineering factors. Physical geography includes nine evaluation factors: elevation, slope, aspect, curvature, topographic humidity index, topographic relief, surface roughness, NDVI, and landform type. From geological conditions, three evaluation factors are selected: stratigraphic lithology, peak ground acceleration, and distance from faults. From meteorology and hydrology, two evaluation factors are selected: average annual precipitation and distance from water systems. From human engineering activities, two evaluation factors are selected: distance from roads and land use type. A total of 16 evaluation factors are selected.
[0088] Step 102: Quantify the obtained indicator data using ArcGIS software.
[0089] Data such as elevation, slope, and aspect were extracted from DEM data; stratigraphic lithology and fault zones were obtained from basic geological data; and river vectors, settlement data, and vegetation cover data were extracted from high-resolution optical imagery, such as Gaofen satellite imagery and Google imagery. Rainfall data was also extracted using data provided by the meteorological bureau or from statistical yearbooks. Since some factors are vector data, the conversion tools in ArcGIS software were used to convert these factors into raster data.
[0090] Step 103: Establish slope units and overlay the quantified index data onto the slope units;
[0091] The division of slope units mainly involves topographic analysis and the application of GIS technology. Common methods for slope unit division include hydrological analysis, curvature methods, multi-scale segmentation methods, and quantitative segmentation methods. This study primarily utilizes ArcGIS and hydrological analysis methods, mainly dividing slope units by analyzing topographic flow direction, discharge, and catchment points. Figure 3This is a flowchart of the slope unit division process.
[0092] Step 104: Standardize the original data using SPSS software, and perform principal component analysis and Spearman correlation analysis on the standardized data to remove factors with high correlation.
[0093] Landslides form under many conditions, so there are many evaluation factors needed to study landslide geological hazards. If the correlation between evaluation factors is high, the evaluation results will be inaccurate. Therefore, after selecting the evaluation factors, it is necessary to calculate the correlation between each factor and eliminate highly correlated factors based on the calculation results, thereby avoiding mutual interference between factors.
[0094] Principal component analysis (PCA) is mainly used for dimensionality reduction and feature extraction. It transforms a set of potentially correlated variables into a set of linearly uncorrelated variables through orthogonal transformation; these new variables are called principal components.
[0095] Perform factor analysis on the standardized data. In SPSS software, select Analyze - Dimensionality Reduction - Factor. In the factor analysis interface, set the required parameters: in the descriptive statistics interface, check the significance level of the correlation matrix and the KMO and Bartlett's tests of sphericity to determine whether the factors can be used for principal component analysis.
[0096] The test results show that the KMO value is 0.819, which is greater than 0.8, making it very suitable for principal component analysis; the Bartlett sphericity test results show that the significance value is 0.000, which is less than 0.05, indicating that the selected index is suitable for principal component analysis.
[0097] The Spearman correlation coefficient reflects the direction and strength of the association between the trends of two random variables. It is a rank correlation analysis method commonly used in data analysis. It is calculated based on the rank of the data sample, rather than directly using the raw data.
[0098] When the absolute value of the correlation coefficient between two evaluation factors is greater than 0.5, it indicates that they have a high correlation; less than 0.3 indicates that the two factors are not related; 0.3 to 0.5 indicates a low correlation; 0.5 to 0.8 indicates a moderate correlation; and greater than or equal to 0.8 indicates a high correlation between the factors.
[0099] Step 105: Eleven evaluation factors were selected, and an indicator system was constructed based on these eleven evaluation factors.
[0100] Analysis showed that the five factors of profile curvature, slope, topographic humidity index, topographic relief, and ground roughness were highly correlated with other factors, so they were removed.
[0101] This invention ultimately selected 11 factors, including stratigraphic lithology, landform type, peak ground acceleration, road density, fault density, normalized difference vegetation index, river network density, average annual precipitation, land type, slope aspect, and elevation, as evaluation indicators for this landslide susceptibility assessment.
[0102] Step 106: Select landslide samples and generate the same number of non-landslide samples as the landslide samples using the random point generation function in ArcGIS software.
[0103] Step 107: Establish an SVM model and use the trained SVM model to predict landslide susceptibility across the entire study area. Generate a landslide susceptibility zoning map based on the prediction results, such as... Figure 4 The model's predictions were analyzed and compared with the actual landslide distribution to verify the model's accuracy.
[0104] Step 108: In ArcGIS software, the obtained raster data information is reclassified into five levels, and finally the attributes of 11 evaluation index factors are extracted to the sample points.
[0105] Install Python libraries in VS Code, construct a random forest model and a support vector machine model, and plot the ROC curve.
[0106] The obtained susceptibility evaluation result points are connected with the sample point attribute table. The Kriging method in ArcGIS is used to generate polygons. The polygon data obtained by the difference is then clipped using the administrative divisions of Sichuan Province.
[0107] Step 109: After obtaining the landslide susceptibility raster data layer, select the classification method in ArcGIS to classify the landslide susceptibility of the study area into levels. The results are divided into 5 levels: extremely low susceptibility area, low susceptibility area, medium susceptibility area, high susceptibility area, and extremely high susceptibility area.
[0108] After obtaining the corresponding susceptibility zoning map, the zoning results were analyzed according to the zoning map. Landslides in Sichuan Province are mainly distributed in the contact zone between the eastern edge of the Qinghai-Tibet Plateau and the Sichuan Basin, the high mountain and canyon areas in the western part of the study area, the mountainous areas around the basin in the northeast and southern Sichuan, and the Panxi region.
[0109] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A landslide susceptibility assessment method based on a support vector machine model, characterized in that, Includes the following steps: S1. Collect basic data of the area to be studied, process and analyze the basic data, and construct an indicator system; S2. Based on the known distribution of landslide points, randomly select an equal number of non-landslide samples in the non-landslide area; S3. Based on the sample points and index system of non-slope samples, the support vector machine is trained to obtain the optimal support vector machine model; S4. Use the optimal support vector machine model to predict the landslide susceptibility of the study area, and generate a landslide susceptibility zoning map based on the prediction results.
2. The landslide susceptibility assessment method based on the support vector machine model according to claim 1, characterized in that, S1 includes the following sub-steps: S11. Collect basic data of the area to be studied and standardize the basic data; S12. Perform principal component analysis and Spearman correlation analysis on the standardized basic data to obtain several evaluation factors. S13. Select stratigraphic lithology, landform type, peak ground acceleration, road density, fault density, normalized vegetation index, river network density, annual average precipitation, land type, slope aspect and elevation from several evaluation factors to construct an indicator system.
3. The landslide susceptibility evaluation method based on the support vector machine model according to claim 1, characterized in that, In step S2, based on known landslide sample points, ArcGIS's random point generation function is used to generate an equal number of non-landslide samples in the non-landslide area.
4. The landslide susceptibility assessment method based on the support vector machine model according to claim 1, characterized in that, S3 includes the following sub-steps: S31. Extract the attributes of several evaluation index factors of the index system to the landslide sample points. S32. Construct a support vector machine and determine the optimization problem of the support vector machine; S33. Process the optimization problem of the support vector machine to obtain the latest optimization problem; S34. Based on the latest optimization problem, the support vector machine algorithm is used to train the sample points to obtain the optimal support vector machine model.
5. The landslide susceptibility assessment method based on the support vector machine model according to claim 4, characterized in that, In step S31, the expression for the support vector machine is: ; in, Describes the normal vector of the hyperplane. Indicates the offset top. This represents the feature vector of a sample.
6. The landslide susceptibility assessment method based on the support vector machine model according to claim 4, characterized in that, The optimization problem of the support vector machine is expressed as follows: 2 1, ; in, Describes the normal vector of the hyperplane. Indicates the offset top. Indicates the first The category label of each sample, Indicates the first The feature vector of each sample Indicates the index of a sample in the dataset.
7. The landslide susceptibility assessment method based on the support vector machine model according to claim 4, characterized in that, In S33, the expression for the latest optimization problem is: ; in, Describes the normal vector of the hyperplane. Indicates the offset top. Represents the regularization parameter. This represents the total number of training samples. Indicates the first Slack variables for each sample, Represents slack variables. Indicates the first The category label of each sample, Indicates the index of a sample in the dataset. Indicates the first The feature vector of each sample.