Landslide susceptibility prediction method based on maximum slope ratio gradient generalization
Through the generalization of maximum slope rate slope and the random forest model, the problem of slope characteristics difference in landslide susceptibility prediction is solved, and higher prediction accuracy and model performance are achieved.
Patent Information
- Application Number
- CN202510744039.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
In the prior art, the use of landslide terrain slope after landslide proneness prediction results in large differences between the slope characteristics and the landslide terrain slope before landslide, resulting in inaccurate prediction.
The maximum slope slope generalization method is used, combined with the random forest model, and the landslide susceptibility prediction model is established by selecting the maximum slope slope and other geological factors, and the Gini coefficient is used for feature selection and decision tree construction, and random forests are generated for prediction.
The accuracy of landslide proneness prediction is improved. Through accuracy, recall, accuracy and ROC curve verification, the maximum slope slope generalization method significantly improves the prediction accuracy and performance of the model.
Smart Images

Figure CN120256809A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of landslide susceptibility prediction, and in particular to a landslide susceptibility prediction method based on the generalization of the maximum slope rate and slope. Background Art
[0002] When conducting a landslide disaster susceptibility assessment, the selection of evaluation indicators has a great impact on the accuracy of model prediction. The evaluation indicators are often selected around the topographic conditions, stratigraphic lithology conditions, geological structure conditions, meteorological conditions, etc. for landslide occurrence. In the selection of evaluation indicators for topographic conditions, the slope gradient is the most important evaluation indicator considered in landslide susceptibility assessment, and a suitable slope range plays an important role in the development of landslides.
[0003] The current evaluation units for landslide susceptibility assessment mainly include grid units and slope units, etc. The slope unit is widely used in medium and large-scale assessments because it can better reflect the characteristics of a landslide as a slope disaster. In the evaluation of slope units, it is necessary to generalize some non-single characteristic values. Currently, the main evaluation indicator considered for the influence of slope gradient is the average slope, and usually, the average slope within the slope unit is obtained as the slope generalization value and added to the slope unit. When selecting landslide samples, the terrain of the evaluation unit obtained is the terrain after sliding, which results in a steeper slope at the rear edge of the landslide and a smaller slope in the front accumulation area, with a large difference from the terrain slope before the landslide. It is inappropriate to use the average slope after sliding to characterize the terrain slope condition before the slope slides, and it cannot accurately represent the slope characteristic before the landslide slides. Summary of the Invention
[0004] The present invention provides a landslide susceptibility prediction method based on the generalization of the maximum slope rate and slope to solve one or more of the above problems.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A landslide susceptibility prediction method based on the generalization of the maximum slope rate and slope, comprising: S1. Randomly select an equal number of non-landslide samples around the landslide; S2. Randomly divide the landslide and non-landslide samples in a ratio of 7:3, where 70% is the training set and 30% is the test set; S3. Take samples with replacement for the row samples in the training set, take column samples for the characteristic factors, and randomly select samples from the characteristic factors to obtain a sample set containing row samples and column samples; S4. Traverse the column samples, select the most suitable characteristic for splitting based on the Gini coefficient, and complete the construction of the decision tree; S5. Repeat S3 and S4 to generate multiple decision trees to form a random forest; S6. Input the test set into the random forest for prediction.
[0006] In this specification, the slope unit is selected as the evaluation unit, and the evaluation factors of the evaluation unit are obtained. The evaluation factors include average elevation, average slope direction, formation lithology, distance from the fault, average historical earthquake point density, distance from the water system, and annual average rainfall. The average slope and the maximum slope rate of the evaluation unit are respectively combined with the evaluation factors to form a data set, and the training and testing of the landslide susceptibility evaluation model are respectively carried out.
[0007] In this specification, 133 landslides are selected, and an equal number of non-landslide samples are randomly selected around the landslides. The slopes where landslides have occurred are marked as 1, and the slopes where no landslides have occurred are marked as 0, obtaining 266 input samples.
[0008] In this specification, 186 row samples in the training set are sampled 186 times with replacement. Sampling with replacement will result in duplicate samples in the collected row sample data, so that when training, the input samples of each decision tree are not all samples, making it relatively less likely to overfit. After the row sampling is completed, column sampling is performed on the feature factors, and m samples are randomly selected from 8 feature factors, where m ≤ 8, to form an input sample set containing 186 row samples and m column samples. The 8 feature factors are average slope, average elevation, average slope direction, formation lithology, distance from the fault, average historical earthquake point density, distance from the water system, and annual average rainfall, or the 8 feature factors are maximum slope rate, average elevation, average slope direction, formation lithology, distance from the fault, average historical earthquake point density, distance from the water system, and annual average rainfall.
[0009] In this specification, traverse the m feature vectors, select the most suitable feature for splitting. The method of feature selection is the Gini coefficient. Feature selection is carried out according to the Gini coefficient. The value range of the Gini coefficient is from 0 to 1, where 0 indicates the highest purity of the data set, and 1 indicates the lowest purity of the data set. For each possible division threshold, the samples in the data set whose values of this feature are less than or equal to the division threshold are assigned to one side l, and the samples greater than the division threshold are assigned to the other side r. Their Gini coefficients are respectively: ; ; Among them, is the Gini coefficient corresponding to node l, is the Gini coefficient corresponding to node r, is the probability of the first type of data points in node l, is the probability of the second type of data points in node l, the probability of the first type of data points in node r, is the probability of the second type of data points in node r; For each partitioning threshold, the Gini coefficients are calculated for the samples on both sides respectively, and the final Gini coefficient of the partitioning threshold is obtained by weighted averaging according to the respective sample sizes. The partitioning threshold with the minimum Gini coefficient is selected as the best partitioning threshold for the continuous feature. Then, the Gini coefficients of the m features are compared, and the feature vector with the minimum Gini coefficient is selected as the best splitting feature, and its best partitioning threshold is used as the splitting point. The above process of node classification selection is repeated until the minimum number of samples for internal node splitting reaches 2 or the depth of the tree reaches 10, at which point the splitting stops and the decision tree construction is completed.
[0010] In this specification, the non - continuous factors (average slope aspect, formation lithology, distance to fault, distance to water system) are standardized to be in the range of [-1, 1].
[0011] In this specification, the accuracy, recall rate, precision, F1 - score, ROC curve and AUC value are used to evaluate the model accuracy.
[0012] In this specification, two landslide susceptibility evaluation models are established based on the two experimental variables of the maximum slope rate and the average slope in the evaluation unit. After the landslide susceptibility evaluation model based on the random forest runs normally, the susceptibility index p of each evaluation unit in the study area is calculated, that is, the probability of landslide occurrence predicted by the random forest model. According to the susceptibility index p of the evaluation unit, the landslide susceptibility is divided into extremely high susceptibility , high susceptibility , medium susceptibility , low susceptibility , extremely low susceptibility , and the landslide susceptibility evaluation maps of the study area under different generalized methods of slope gradient are generated respectively.
[0013] In summary, the present invention has at least the following beneficial effects: The present invention conducts a landslide susceptibility evaluation on the selected study area through a random forest model. Other parameters in the susceptibility evaluation system refer to the commonly used parameters in previous studies. The landslide sample slope generalization parameters are calculated using the above - mentioned two methods respectively, and then the accuracy of different slope gradient generalization methods is evaluated. The accuracy of the susceptibility evaluation results is verified by using the accuracy, recall rate, precision, F1 - score and ROC curve methods of the evaluation results. The present invention effectively improves the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0015] Figure 1 It is a schematic diagram of the landslide susceptibility prediction method based on the generalization of the maximum slope ratio and slope in the present invention.
[0016] Figure 2 It is a schematic diagram of the maximum slope ratio and slope calculation method in the present invention.
[0017] Figure 3 It is a schematic diagram of the decision tree model in the present invention.
[0018] Figure 4 It is a schematic diagram of the model construction process in the present invention.
[0019] Figure 5 It is a schematic diagram of the ROC curve of the susceptibility evaluation model in the present invention. Specific Embodiments
[0020] In the following text, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the embodiments of the present invention. Therefore, the accompanying drawings and the description are considered to be exemplary in nature rather than restrictive.
[0021] The following disclosure provides many different embodiments or examples for implementing different structures of the embodiments of the present invention. To simplify the disclosure of the embodiments of the present invention, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the embodiments of the present invention. In addition, the embodiments of the present invention may repeat reference numerals and / or reference letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0022] The following will describe the embodiments of the present invention in detail with reference to the accompanying drawings.
[0023] As Figure 1 shown, this embodiment provides a landslide susceptibility prediction method based on the generalization of the maximum slope ratio and slope, including: S1. Randomly select an equal number of non-landslide samples around the landslide; S2. Randomly divide the landslide and non-landslide samples in a ratio of 7:3, where 70% is the training set and 30% is the test set; S3. Take sampling with replacement for the row samples in the training set, take column sampling for the feature factors, randomly select samples from the feature factors, and obtain a sample set containing row samples and column samples; S4. Traverse the column samples, select the most suitable feature for splitting according to the Gini coefficient, and complete the construction of the decision tree; S5. Repeat S3 and S4 to generate multiple decision trees to form a random forest; S6. Input the test set into the random forest for prediction.
[0024] The present invention proposes the maximum slope rate gradient, which is used to characterize and generalize the slope gradient index in landslide samples. The maximum slope rate gradient comprehensively considers the terrain slope difference before and after the landslide of the landslide sample, and is defined as the maximum gradient value of the line connecting the landslide rear wall (A) and the landslide front edge (B) in the sliding direction, denoted by Max(S), where A is the determined landslide rear wall line and B is the determined landslide accumulation front edge line. As Figure 2 shown, the green line is the landslide rear wall line A and the landslide front edge line B; the blue line is the slope rate line corresponding to the segmentation point of the landslide rear wall and the front edge, the red line is the maximum slope rate line, Max(S) is the maximum slope rate gradient corresponding to the maximum slope rate line, and based on the point generation tool along the line of the ArcGIS10.8 platform, fixed points A are generated at equal intervals on A and B i、 B i (default is 10), and the maximum slope rate gradient value Max(S) is obtained by finding the maximum value corresponding to the top and bottom pairwise. Compared with the traditional method, it can more accurately reflect the slope characteristics of the slope, which is expressed as: ;
[0025] Where: is the point elevation difference between; L is the point distance difference between.
[0026] To verify the effectiveness and advancement of this method, an experimental study was conducted on the landslide-developed area in Chaya Region, Qamdo City, Tibet. The study area is located in Chaya Region, Qamdo City, Tibet Autonomous Region, and is a rectangular area in the area of Qamdo Karuo District - Chaya County, with an area of 4213.36 km 2The area is generally located in the eastern margin of the Qinghai-Tibet Plateau and the northern section of the Sanjiang area, belonging to the tectonic-erosion and erosional high mountain geomorphic area. The topographic and geomorphic differences are large, the geological structure is complex, and structures such as faults and folds are distributed in NW-trending bands. The main exposed strata are Quaternary alluvial-proluvial deposits, Paleogene, Jurassic, Triassic, Permian, Devonian sandstone, mudstone, limestone, etc., and local magmatic rocks are exposed. The overall area is high in the west and low in the east, with the lowest elevation of 3605 m and the highest elevation of 5615 m. The Lancang River runs through the western part of the study area from south to north, and the valleys on both sides are crisscrossed. The slope gradients are mostly between 20° and 40°, and locally greater than 40°.
[0027] The susceptibility assessment of landslides in the selected study area is carried out through a random forest model. For other parameters in the susceptibility assessment system, the commonly used parameters from previous studies are referred to. The slope generalization parameters of landslide samples are calculated using the above two methods respectively, and then the accuracy of different slope gradient estimation methods is evaluated. The accuracy of the susceptibility assessment results is verified using methods such as accuracy rate, recall rate, precision rate, F1 score, and ROC curve of the assessment results.
[0028] Data sources: Landslide data comes from field investigations and remote sensing interpretations. There are 133 slopes with developed landslides in the study area. Fault and stratigraphic data come from 1:250,000 geological maps; topographic data is based on 1:50,000 topographic maps; rainfall data comes from the monthly average rainfall data from 1958 to 2022 obtained from the National Data Center of the Qinghai-Tibet Plateau; historical earthquake point data comes from the National Earthquake Science Data Center.
[0029] Selection of evaluation indicators: The formation of landslides is mainly affected by factors such as topography, geological structure, and lithology. In addition to using two calculation methods for the above-mentioned slope factor, in order to more accurately predict the model, the present invention combines the conditions required for landslide occurrence and selects the following evaluation indicators considering the availability and computability of comprehensive evaluation factors: Factors related to topographic and geomorphic conditions (1) Average elevation: It refers to the average value of all elevation grids within the evaluation unit; (2) Average slope aspect: It refers to the average value of all slope aspect grids within the evaluation unit; Factors related to geological conditions (1) Main stratigraphic lithology: It refers to the main lithology combination within the study area; according to the characteristics of the stratigraphic lithology in the study area, the main stratigraphic lithology within the evaluation unit is extracted, which is divided into 15 categories, namely ① C1k.; ② D3z; ③ E g ; ④ J1w ⑤ J2d; ⑥ J3x; ⑦ P2j; ⑧ Pt 1-2 K.\Pt 1-2 J.\Pt3Y; ⑨ Qpal; ⑩ T 1-2m; ⑪T3a; ⑫T3b; ⑬T3d; ⑭T3j; ⑮η\γ\δ\ψ l \; (2)Distance from the fault: Divide it according to the intervals of 500m, 1000m, 2000m, 4000m, >4000m, and extract the main distance from the fault values within the evaluation unit; (3)Average historical earthquake point density: It refers to the average value of the point density within the evaluation unit obtained by using the point density tool in ArcGIS based on the national historical earthquake record data (2022 version); Hydrological condition-related factors (1)Distance from the water system: Divide it according to the intervals of 500m, 1000m, 1500m, 2000m, 2500m, >2500m, and extract the main distance from the water system values within the evaluation unit; (2)Annual average rainfall: Obtained by averaging the monthly average rainfall from 1958 to 2022. The calculation formula is as follows: Annual average rainfall m = Total rainfall S / Number of years n; Random forest model Random forest is an ensemble learning method. It consists of multiple decision trees, each of which is independently trained, and a subset of its input features is randomly selected. By voting or taking the average of multiple decision trees, random forest can effectively avoid the overfitting problem and has good generalization ability and robustness.
[0030] First, randomly select a part of the samples from the training dataset to construct each decision tree. This process is called "bootstrap sampling" or "out-of-bag sampling", where the probability of each sample being selected is 1 / n, and n is the total number of samples. When dividing the nodes of each decision tree, randomly select a subset of features and choose the best feature for division from them. This process can reduce the correlation between features and enhance the stability and generalization ability of the model. Then, based on the above steps, construct multiple decision trees, each of which is generated by continuously bisecting the features until the stopping condition is met. Usually, random forest uses the CART algorithm to construct decision trees. Finally, for classification tasks, random forest uses voting to predict the category of the target variable; for regression tasks, random forest uses the average value to predict the value of the target variable.
[0031] Model accuracy evaluation: Common model accuracy evaluation methods for random forest models include accuracy, recall, F1-score, confusion matrix, ROC curve, AUC value, and cross-validation, etc. In this invention, accuracy, recall, precision, and F1-score are selected for model accuracy evaluation. At the same time, by plotting the ROC curve, the relationship between the true positive rate (TPR) and false positive rate (FPR) under different thresholds is shown, and the area under the curve (AUC value) is calculated to measure the classification performance of the random forest model. The closer the AUC value is to 1, the better the performance of the classifier; conversely, the closer the AUC value is to 0, the worse the performance of the classifier. Theoretically, the AUC value of a perfect classifier is 1, while the AUC value of a random classifier is 0.5.
[0032] Model construction: The evaluation unit selected in this invention is the slope unit. Taking 133 landslides in the study area as the research object, an equal number of non-landslide samples are randomly selected around the landslides. The slopes where landslides have occurred are marked as "1", and the slopes where no landslides have occurred are marked as "0". The average slope and maximum slope rate of the evaluation unit are combined with the data of other 7 evaluation factors respectively to form a data set, and the landslide susceptibility evaluation model is trained and tested respectively.
[0033] As Figure 4 shown, the model construction process is as follows: (1)Data processing First, non-continuous factors (average slope direction, formation lithology, distance from fault, distance from water system) are standardized to be in the interval [-1,1]; (2)Data division The 266 input samples are randomly divided in a ratio of 7:3, where 70% is the training set T (186 samples) and 30% is the test set C (80 samples).
[0034] (3)Random sampling The model will sample the row samples in the training set T (i.e., 186 samples) 186 times with replacement. Sampling with replacement will result in duplicate samples in the collected row sample data. In this way, when training, the input samples of each decision tree are not all the samples, making it relatively less likely to overfit. After the row sampling is completed, the model will sample the feature factors column-wise, randomly selecting m samples (m ≤ 8) from the above 8 feature factors, thus forming an input sample set containing 186 row samples and m column samples.
[0035] (4)Construct decision trees Traverse m feature vectors and select the most suitable feature for splitting. The method for feature selection is the Gini impurity. The model selects features based on the Gini impurity, which is a metric for measuring the purity of a dataset. It measures the probability that two randomly selected samples from the dataset have different class labels. The value range of the Gini impurity is from 0 to 1, where 0 indicates the highest purity of the dataset and 1 indicates the lowest purity. For example, in a random sampling of m column samples, since non-continuous features have been standardized to form continuous data, for continuous features, the model will first sort all possible values. For each possible splitting threshold, the samples in the dataset whose feature values are less than or equal to the splitting threshold are assigned to one side l, and the samples greater than the splitting threshold are assigned to the other side r. Their Gini impurities are respectively: ; ; where, is the Gini impurity corresponding to node l, is the Gini impurity corresponding to node r, is the probability of the first type of data points in node l, is the probability of the second type of data points in node l, the probability of the first type of data points in node r, is the probability of the second type of data points in node r; The evaluation unit where a landslide has occurred is denoted as "1", and the evaluation unit where no landslide has occurred is denoted as "0".
[0036] For each splitting threshold, calculate the Gini impurities of the samples on both sides respectively, and obtain the final Gini impurity of the splitting threshold by weighted averaging according to the respective sample numbers. Select the splitting threshold with the minimum Gini impurity as the best splitting threshold for continuous features. Then compare the Gini impurities of m features, and select the feature vector with the minimum Gini impurity as the best splitting feature, and its best splitting threshold as the splitting point. Repeat the above node classification and selection process until the minimum number of samples for internal node splitting reaches 2 or the depth of the tree reaches 10 to stop splitting. The decision tree is constructed. As Figure 3 shown, blue represents the class with a label value of 1 (corresponding to landslide), and orange represents the class with a label of 0 (corresponding to non-landslide). The darker the color, the higher the classification purity. The closer the Gini impurity is to 0, the higher the purity. The first line in English is the abbreviation of the evaluation factor. The four parameters respectively represent the Gini impurity (gini), the number of samples (samples), the class count (value), and the final classification (class).
[0037] (5)Construct a random forest Repeat the processes (3) and (4) 100 times to generate 100 decision trees. Combine the 100 constructed decision trees into a random forest. The mode of the data of the prediction results of the 100 decision trees is statistically calculated as the output result of the prediction model of the final random forest.
[0038] (6) Prediction Bring the test set C randomly divided in (2) into the random forest prediction model constructed in (5) for result prediction.
[0039] (7) Evaluate the model performance Use accuracy, recall, precision, F1-score, ROC curve and AUC value to evaluate the model accuracy.
[0040] Evaluation results: Based on the two experimental variables of the maximum slope rate and the average slope in the evaluation unit, two landslide susceptibility evaluation models are established. After the landslide susceptibility evaluation model based on the random forest runs normally, calculate the susceptibility index p of each evaluation unit in the study area, that is, the landslide occurrence probability predicted by the random forest model. According to the susceptibility index p of the evaluation unit, the landslide susceptibility is divided into five categories: extremely high susceptibility ( ), high susceptibility ( ), medium susceptibility ( ), low susceptibility ( ), and extremely low susceptibility ( ).
[0041] Model accuracy evaluation: In order to evaluate the modeling impact of two slope unit slope generalization methods on the random forest landslide susceptibility model, use accuracy, recall, precision, F1-score, ROC curve and AUC value to evaluate the model accuracy of the two models. The model accuracy evaluation of the two models is shown in Table 1.
[0042] Table 1 Model accuracy evaluation table
[0043] The above table shows the classification evaluation indicators of the test set, and the classification effect of the random forest on the test data is measured by quantitative indicators. Among them, accuracy is the proportion of correctly predicted samples in the total samples, and the larger the accuracy, the better. Recall is the proportion of samples predicted as positive samples among the results that are actually positive samples, and the larger the recall, the better. Precision is the proportion of samples that are actually positive samples among the results predicted as positive samples, and the larger the precision, the better. The F1-score is the harmonic mean of precision and recall.
[0044] In the evaluation of model accuracy, the accuracy of the test set is usually used to evaluate the accuracy of the model. As can be seen from the above table, the accuracy, recall, precision, and F1 score of the random forest landslide susceptibility evaluation model with the maximum slope rate (Max(S)) as the generalization of the slope gradient are 91.25%, 91.54%, 91.22%, and 91.22% respectively, which are 7.50%, 7.07%, 6.90%, and 7.47% higher than those of the generalization method of the average slope respectively, proving that the generalization method of the maximum slope rate can significantly improve the accuracy of the model.
[0045] The ROC curve is a curve commonly used in machine learning to evaluate the performance of classifiers. Since the ROC curve is obtained by gradually decreasing the landslide susceptibility threshold to get the coordinates corresponding to each point on the curve, the closer the curve is to the upper left corner, the more landslides are concentrated in a high-grade and small-range interval, and the larger the area under the curve (AUC value), so the better the prediction effect. The ROC curves of the random forest model under the two slope generalization methods of the present invention are as Figure 5 shown. As can be seen from the figure, the AUC values of the two models of maximum slope rate generalization (Max(S)) and average slope (AS) generalization are 0.952 and 0.933 respectively. The AUC value of the maximum slope rate generalization model is significantly higher than the other one, proving that the slope data processing method of maximum slope rate generalization can improve the model performance of random forest landslide susceptibility prediction.
[0046] The present invention proposes a slope generalization calculation method for landslide susceptibility evaluation and compares and verifies this method with the commonly used methods in the past. The results show that: in the evaluation of model accuracy, the accuracy of the test set is usually used to evaluate the accuracy of the model. The accuracy, recall, precision, and F1 score of the random forest landslide susceptibility evaluation model with the maximum slope rate (Max(S)) as the generalization of the slope gradient are 91.25%, 91.54%, 91.22%, and 91.22% respectively, which are 7.50%, 7.07%, 6.90%, and 7.47% higher than those of the generalization method of the average slope respectively, proving that the generalization method of the maximum slope rate can significantly improve the accuracy of the model.
[0047] The results of the ROC curve show that the AUC values of the two models of maximum slope rate generalization and average slope generalization are 0.952 and 0.933 respectively. The AUC value of the maximum slope rate generalization model is significantly higher than the other one, proving that the slope gradient generalization data processing method of the maximum slope rate can improve the model performance of random forest landslide susceptibility prediction.
[0048] The above-described embodiments are used to illustrate the present invention and are not used to limit the present invention. Therefore, changes in the example numerical values or replacement of equivalent elements still belong to the scope of the present invention.
[0049] From the above detailed description, those of ordinary skill in the art can clearly understand that the present invention can indeed achieve the aforementioned objectives, and it actually complies with the provisions of the Patent Law.
[0050] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
[0051] It should be noted that the above description of the process is only for illustration and explanation, and does not limit the scope of application of this specification. For those skilled in the art, various corrections and changes can be made to the process under the guidance of this specification. However, these corrections and changes are still within the scope of this specification.
[0052] The basic concept has been described above. Obviously, for those of ordinary skill in the art after reading this application, the above invention disclosure is only for illustration and does not constitute a limitation to this application. Although not explicitly stated here, those of ordinary skill in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, so such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of this application.
[0053] In addition, unless explicitly stated in the claims, the order of the processing elements and sequences, the use of numerical letters, or the use of other names described in this application are not used to limit the order of the processes and methods of this application. Although some currently considered useful invention embodiments are discussed through various examples in the above disclosure, it should be understood that such details are only for the purpose of illustration, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all corrections and equivalent combinations that conform to the essence and scope of the embodiments of this application. For example, although the implementation of the above various components can be embodied in a hardware device, it can also be implemented as a pure software solution, for example, installed on an existing server or mobile device.
[0054] Similarly, it should be noted that, in order to simplify the description disclosed in the present application and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present application, sometimes multiple features are incorporated into one embodiment, drawing or description thereof. However, this method of the present application should not be construed as reflecting an intention that the claimed subject matter requires more features than are expressly recited in each claim. On the contrary, the subject matter of the invention should have fewer features than the above single embodiment.
Claims
1. A landslide susceptibility prediction method based on the generalization of the maximum slope rate and gradient, characterized in that Including: S1. Randomly select an equal number of non-landslide samples around the landslide; S2. Randomly divide the landslide and non-landslide samples in a ratio of 7:3, where 70% is the training set and 30% is the test set; S3. Take samples with replacement for the row samples in the training set, take column samples for the feature factors, and randomly select samples from the feature factors to obtain a sample set containing row samples and column samples; S4. Traverse the column samples, select the most suitable feature for splitting according to the Gini coefficient, and complete the construction of the decision tree; S5. Repeat S3 and S4 to generate multiple decision trees to form a random forest; S6. Input the test set into the random forest for prediction.
2. The landslide susceptibility prediction method based on maximum slope rate and gradient generalization according to claim 1, wherein Select the slope unit as the evaluation unit, obtain the evaluation factors of the evaluation unit. The evaluation factors include average elevation, average slope aspect, formation lithology, distance from the fault, average historical earthquake point density, distance from the water system, and annual average rainfall. Combine the average slope and the maximum slope ratio of the evaluation unit with the evaluation factors respectively to form a data set, and conduct the training and testing of the landslide susceptibility evaluation model respectively.
3. The landslide susceptibility prediction method based on maximum slope rate and slope generalization according to claim 2, characterized in that, Select 133 landslides, and at the same time randomly select an equal number of non-landslide samples around the landslides. Mark the slopes where landslides have occurred as 1, and mark the slopes where no landslides have occurred as 0 to obtain 266 input samples.
4. The landslide susceptibility prediction method based on maximum slope rate and gradient generalization according to claim 3, characterized in that Take samples with replacement 186 times for the 186 row samples in the training set. Samples may be repeated in the collected row sample data during sampling with replacement. After the row sampling is completed, take column samples for the feature factors, and randomly select m samples from 8 feature factors, where m ≤ 8, to form an input sample set containing 186 row samples and m column samples. The 8 feature factors are average slope, average elevation, average slope aspect, formation lithology, distance from the fault, average historical earthquake point density, distance from the water system, and annual average rainfall, or the 8 feature factors are maximum slope ratio, average elevation, average slope aspect, formation lithology, distance from the fault, average historical earthquake point density, distance from the water system, and annual average rainfall.
5. The landslide susceptibility prediction method based on maximum slope rate and slope generalization according to claim 4, wherein Traverse m feature vectors, select the most suitable feature for splitting. The method for feature selection is the Gini coefficient. Feature selection is carried out based on the Gini coefficient. The value range of the Gini coefficient is from 0 to 1, where 0 indicates the highest purity of the data set, and 1 indicates the lowest purity of the data set. For each possible splitting threshold, the samples in the data set with the value of this feature less than or equal to the splitting threshold are assigned to one side l, and the samples greater than the splitting threshold are assigned to the other side r. Their Gini coefficients are respectively: ; ; Among them, is the Gini coefficient corresponding to node l, is the Gini coefficient corresponding to node r, is the probability of the first type of data points in node l, is the probability of the second type of data points in node l, the probability of the first type of data points in node r, is the probability of the second type of data points in node r; For each division threshold, calculate the Gini coefficient for the samples on both sides respectively, and obtain the final division threshold Gini coefficient by weighted average according to the respective sample numbers. Select the division threshold with the minimum Gini coefficient as the best division threshold for continuous features. Then compare the Gini coefficients of m features, select the feature vector with the minimum Gini coefficient as the best splitting feature, and its best division threshold as the division point. Repeat the above node classification selection process until the minimum number of samples for internal node splitting reaches 2 or the depth of the tree reaches 10 and stop splitting, and the decision tree construction is completed.
6. The landslide susceptibility prediction method based on maximum slope ratio generalization according to claim 5, characterized in that Standardize the non-continuous factors so that they are in the interval [-1, 1].
7. The landslide susceptibility prediction method based on the generalization of the maximum slope rate and slope as claimed in claim 6, wherein Use accuracy, recall rate, precision, F1 score, and ROC curve and AUC value to evaluate the model accuracy.
8. The landslide susceptibility prediction method based on maximum slope rate and gradient generalization according to claim 7, characterized in that Based on the two test variables of the maximum slope rate and the average slope in the evaluation unit, two landslide susceptibility evaluation models are established. After the landslide susceptibility evaluation model based on random forest runs normally, the susceptibility index p of each evaluation unit in the study area is calculated, that is, the probability of landslide occurrence predicted by the random forest model. According to the susceptibility index p of the evaluation unit, the landslide susceptibility is divided into extremely high susceptibility , high susceptibility , medium susceptibility , low susceptibility , and extremely low susceptibility , and the landslide susceptibility evaluation maps of the study area under different generalized methods of slope gradient are generated respectively.
Citation Information
Patent Citations
Slope stability evaluation method
CN114676550A
Debris flow susceptibility prediction method and device based on fault distribution index, and medium
CN117540830A
Assessment method and device for seismic landslide hazard based on landslide-density-newmark (LS-d-newmark) model, and processing device
US20240020441A1