A landslide susceptibility prediction method considering geospatial constrained sampling
By adopting a sampling method that takes into account geospatial constraints in landslide proneness prediction, the problems of data sampling uncertainty and subjective arbitrary in the prior art are solved, the prediction accuracy and reliability are improved, and the global representation of non-landslide samples are ensured.
Patent Information
- Application Number
- CN202410367603.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-03-28
AI Technical Summary
In the existing landslide prone prediction methods, there is uncertainty and subjective arbitrary data sampling, resulting in insufficient accuracy and reliability of landslide prone prediction.
The sampling method that takes into account geospatial constraints is adopted. By selecting the landslide suspense evaluation factor, pre-processing and weight calculation, low-prone areas are extracted as feature space sampling areas, and combining the digital elevation model to extract the slope unit as the geospatial sampling area, finally obtaining the sampling area through superposition, calculating the sampling weight and random sampling is performed.
The data sampling accuracy and reliability of landslide prone prediction are improved, the global representation of non-landslide samples is ensured, spatial aggregation is avoided, and the reliability and accuracy of prediction results are improved.
Smart Images

Figure CN118194129B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of landslide susceptibility prediction, and in particular to a landslide susceptibility prediction method taking into account geographic space constraint sampling. Background Art
[0002] Every year, landslide disasters cause huge casualties and economic and property losses around the world. The causes of landslides are complex. They are nonlinear deformation and destruction that develop and evolve under the coupling of multiple factors. Existing technologies make it difficult to provide timely and accurate early warning and forecasting. Therefore, it is an inevitable requirement for geological disaster prevention and control management at this stage to grasp the high-risk areas of landslide disasters in advance. As a method for predicting the probability of regional landslide disasters, landslide susceptibility assessment mainly answers the difficult question of "where landslides are likely to occur". It has now become a powerful tool for land planning, project site selection and disaster risk management.
[0003] Data sampling is a necessary step in landslide susceptibility prediction, which mainly includes landslide data sampling (positive samples), non-landslide data sampling (negative samples) and factor data sampling. The quality of sampling data is crucial to landslide susceptibility modeling. Unreliable or even erroneous sampling data will seriously interfere with the nonlinear modeling of landslides and disaster-predisposing factors, resulting in overestimation or underestimation of the evaluation results. Overestimation of landslide susceptibility will limit the proper function and value of the land and hinder the economic development of the region; underestimation of landslide susceptibility is likely to cause great property losses and casualties in the future.
[0004] However, in the current landslide susceptibility prediction research, data sampling is not given enough attention, and data sampling is mainly obtained by random sampling. For landslide sample sampling, most studies use sampling methods based on landslide points, such as the center point of the landslide, the center point of the back wall of the landslide, etc. The sampling results based on landslide points are not completely consistent with the actual range of the landslide, and the location selection of landslide sampling points has certain ambiguity and uncertainty. For non-landslide sample sampling, most studies use a buffer zone with the landslide as the center, and perform random sampling of non-landslide samples outside the buffer zone. This sampling method has a large degree of subjective arbitrariness. Moreover, the non-landslide sampling area determined by the buffer zone is very likely to contain potential landslide hazard areas. Obviously, unreasonable or even biased data samples can easily lead to the phenomenon that the landslide susceptibility prediction accuracy is too high but the prediction results are unreliable.
[0005] At present, there are two main methods for sampling landslide samples in the existing technology: sampling based on landslide points and sampling based on landslide surfaces. The sampling method based on landslide points mainly uses the center point of the landslide or the center point of the back wall of the landslide; the sampling method based on landslide surfaces mainly uses the actual delineated landslide surface as the sampling area for landslide samples. In addition, some studies have proposed to use a buffer zone to delineate the landslide sampling area, and sample landslide samples in the buffer zone. For non-landslide sample sampling, the point format is mainly used. The commonly used sampling method is to use the landslide as the center as a buffer zone, and perform random sampling outside the buffer zone. Some studies have also proposed random sampling of non-landslide samples in terrain areas such as rivers or low slopes where landslides are not prone to occur. Some studies have also used certain methods to preliminarily delineate low-prone areas for landslides, and perform random sampling in low-prone areas.
[0006] However, for landslide sample sampling, both the sampling methods based on landslide points and landslide surfaces have certain defects. For example, for the sampling method based on landslide points, the location selection of landslide sampling points has certain ambiguity and uncertainty, and the sampling results based on landslide points are not completely consistent with the actual range of landslides. For the sampling method based on landslide surfaces, the amount of sampled data will become larger, increasing the amount of model calculation. For the sampling method based on buffer zones, how to determine the buffer zone distance varies from person to person, and the sampling results have a large degree of subjective arbitrariness. For non-landslide sample sampling, the sampling strategy based on buffer zones has a large degree of subjective arbitrariness in determining the buffer zone distance, and the non-landslide sampling area determined by the buffer zone is likely to contain potential landslide hazard areas. However, the sampling of non-landslide samples in low-prone areas of landslides does not take into account the global spatial representativeness of the samples, and biased samples can easily lead to the phenomenon that the evaluation accuracy is too high and the evaluation results are unreliable. Summary of the invention
[0007] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a landslide susceptibility prediction method taking into account geographic space constraint sampling, which can overcome the deficiencies of existing data sample sampling and improve the accuracy and reliability of landslide susceptibility prediction.
[0008] To achieve the above object, the present invention provides the following solutions:
[0009] A landslide susceptibility prediction method taking into account geospatial constrained sampling, comprising:
[0010] Select landslide susceptibility evaluation factors based on the actual situation of the target area and the collected geographic basic data;
[0011] Preprocessing the landslide susceptibility evaluation factor, and using the frequency ratio method to calculate the relative influence of each attribute interval of the preprocessed landslide susceptibility evaluation factor on the occurrence of landslide;
[0012] The landslide susceptibility evaluation factors after pretreatment were ranked and sorted by using the analytic hierarchy process, and the relative weights of the landslide susceptibility evaluation factors were determined by pairwise comparison.
[0013] Performing weighted calculation according to the relative weight and the relative influence degree to obtain a susceptibility index of the target area, and extracting a low-susceptibility area as a feature space sampling area;
[0014] Based on the digital elevation model, ArcGIS software was used to extract the slope units in the target area as the geospatial sampling area;
[0015] Superimposing the feature space sampling area and the geographic space sampling area to obtain a sampling area;
[0016] Calculate the ratio of the number of low-susceptibility sample points in the slope unit to the number of slope unit grids, and normalize the ratio to obtain the sampling weight;
[0017] According to the sampling weight, the number of non-landslide samples in each slope unit is determined, and low-susceptibility samples in each slope unit are randomly sampled to obtain the final non-landslide sample data;
[0018] A landslide susceptibility prediction is performed based on the non-landslide sample data to obtain a landslide susceptibility prediction result for the target area.
[0019] Preferably, the basic geographic data includes topographic data, landform data, geological data, river network data, road data and annual average rainfall data.
[0020] Preferably, the landslide susceptibility assessment factors include: any one or more of slope, slope aspect, elevation, land use, stratum lithology, distance to a river, distance to a road, distance to a fault and average annual rainfall.
[0021] Preferably, the calculation formula for the relative impact degree is: Among them, F i Indicates the relative influence of the i-th attribute interval of the evaluation factor, L i represents the landslide area within the i-th attribute interval of the evaluation factor, L represents the total landslide area in the target area, and A i It represents the area of the i-th attribute interval of the evaluation factor, and A represents the total area of the target area.
[0022] Preferably, the calculation formula of the susceptibility index is: Where LSI represents the landslide susceptibility index, Q i represents the relative weight of the evaluation factors calculated by the hierarchical analysis method, F i It represents the relative influence of the evaluation factors calculated by the frequency ratio method, and m represents the total number of evaluation factors.
[0023] Preferably, the formula for calculating the ratio of the number of low-susceptibility sample points in a slope unit to the number of slope unit grids is: Among them, W i represents the sampling weight of the i-th slope unit, n represents the total number of slope units, N i represents the number of low-susceptibility samples in the slope unit, S i Indicates the number of slope cell grids.
[0024] Preferably, performing landslide susceptibility prediction based on the non-landslide sample data to obtain a landslide susceptibility prediction result for the target area includes:
[0025] Determine the landslide point or landslide surface sampling method based on the actual landslide data, and obtain landslide sample data by sampling;
[0026] The pre-processed landslide susceptibility evaluation factors were subjected to factor collinearity analysis to remove factors that were irrelevant to the modeling.
[0027] The landslide sample data, the non-landslide sample data and the removed evaluation factors are randomly divided into a training data set and a test data set according to a preset ratio;
[0028] The landslide susceptibility is modeled using machine learning algorithms, and the machine learning model is trained and tested using training and test data sets to obtain the optimal prediction model.
[0029] The evaluation factor data to be tested is input into the prediction model to calculate and obtain the landslide susceptibility prediction result.
[0030] Preferably, the factor collinearity analysis is performed using a variance inflation factor diagnostic method; the calculation formula of the variance inflation factor diagnostic method is: In which, assume that the independent variable X = [X 1 ,X 2 ,X 3 ,…,X N ], R 2 j is the jth independent variable X j VIF is the variance inflation factor, and N is the number of independent variables.
[0031] Preferably, the machine learning model is any one of a support vector machine model, a decision tree model and a random forest model.
[0032] Preferably, the preset ratio is 7:3.
[0033] The present invention discloses the following technical effects:
[0034] The non-landslide samples obtained by the present invention will be more globally representative, can effectively constrain the spatial distribution of non-landslide samples, avoid the spatial aggregation of non-landslide samples, and the landslide susceptibility modeling results will be more reliable and accurate. The susceptibility prediction results will be able to effectively guide regional land planning, project site selection, disaster risk management and other related work. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0036] Figure 1 A flow chart of a method provided by an embodiment of the present invention;
[0037] Figure 2 A schematic diagram of a technical route provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0039] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0040] Figure 1 A flow chart of a method provided by an embodiment of the present invention, such as Figure 1 As shown, the present invention provides a landslide susceptibility prediction method taking into account geospatial constrained sampling, comprising:
[0041] Step 100: selecting a landslide susceptibility evaluation factor according to the actual situation of the target area and the collected geographic basic data;
[0042] Step 200: preprocessing the landslide susceptibility evaluation factor, and using a frequency ratio method to calculate the relative influence of each attribute interval of the preprocessed landslide susceptibility evaluation factor on the occurrence of landslide;
[0043] Step 300: using the analytic hierarchy process to hierarchically sort the preprocessed landslide susceptibility evaluation factors, and determining the relative weights of the landslide susceptibility evaluation factors by pairwise comparison;
[0044] Step 400: performing weighted calculation according to the relative weight and the relative influence degree to obtain a susceptibility index of the target area, and extracting a low susceptibility area as a feature space sampling area;
[0045] Step 500: Based on the digital elevation model, using ArcGIS software to extract slope units in the target area as a geographic space sampling area;
[0046] Step 600: superimposing the feature space sampling area and the geographic space sampling area to obtain a sampling area;
[0047] Step 700: Calculate the ratio of the number of low-susceptibility sample points in the slope unit to the number of slope unit grids, and normalize the ratio to obtain a sampling weight;
[0048] Step 800: Determine the number of non-landslide samples in each slope unit according to the sampling weight, randomly sample low-susceptibility samples in each slope unit, and summarize to obtain final non-landslide sample data;
[0049] Step 900: Perform landslide susceptibility prediction based on the non-landslide sample data to obtain a landslide susceptibility prediction result for the target area.
[0050] like Figure 2 As shown, the technical route of this embodiment mainly includes two parts: non-landslide sample sampling and landslide susceptibility prediction.
[0051] Non-landslide sampling includes the following steps:
[0052] Step 1: Collect basic regional data, including topography, landforms, geology, river networks, roads, average annual rainfall, etc.;
[0053] Step 2: According to the actual situation of the study area and the basic data collected, select landslide susceptibility evaluation factors, such as slope, aspect, elevation, land use, stratum lithology, distance to the river, distance to the road, distance to the fault and average annual rainfall, among which the slope, aspect, distance to the river, distance to the road and distance to the fault can be calculated by ArcGIS software, and the stratum lithology and fault distribution can be obtained by vectorization of geological maps;
[0054] Exemplarily, the ArcGIS software mentioned in the above non-landslide sample sampling steps may also be replaced by other GIS software or tools.
[0055] Step 3: Preprocessing of landslide susceptibility evaluation factors, including projection transformation, resampling, and attribute classification;
[0056] Step 4: Use the frequency ratio method to calculate the relative influence of each attribute interval of the evaluation factor on the occurrence of landslides, that is, the frequency ratio. The calculation formula is as follows:
[0057]
[0058] In formula (1), F i represents the frequency ratio of the i-th attribute interval of the evaluation factor, L i represents the landslide area within the i-th attribute interval of the evaluation factor, L represents the total landslide area in the study area, and A i It represents the area of the i-th attribute interval of the evaluation factor, and A represents the total area of the study area.
[0059] Exemplarily, the frequency ratio method in step 3 may also be replaced by an information quantity method.
[0060] Step 5: Use the analytic hierarchy process to rank the evaluation factors and determine the relative weights of the evaluation factors through pairwise comparison;
[0061] Step 6: According to the relative weights of the evaluation factors and the factor frequency ratio, the susceptibility index of the study area is calculated by weighted calculation (the formula is as follows), and the low-susceptibility area is extracted as the characteristic space sampling area of the non-landslide sample;
[0062]
[0063] In formula (2), LSI represents the landslide susceptibility index, Q i represents the relative weight of the evaluation factors calculated by the hierarchical analysis method, F i It represents the relative influence of the evaluation factors calculated by the frequency ratio method, and m represents the total number of evaluation factors.
[0064] Step 7: Based on the Digital Elevation Model (DEM), ArcGIS software was used to extract the slope units in the study area as the geospatial sampling area for non-landslide samples;
[0065] Step 8: Overlay the characteristic spatial sampling area and geographic spatial sampling area of non-landslide samples to obtain the sampling area of non-landslide samples;
[0066] Step 9: Calculate the ratio of the number of low-susceptibility sample points in the slope unit to the number of slope unit grids, normalize the ratio, and use it as the sampling weight of non-landslide samples. The calculation formula is as follows:
[0067]
[0068] In formula (3), W i represents the sampling weight of the i-th slope unit, n represents the total number of slope units, N i represents the number of low-susceptibility samples in the slope unit, S i Indicates the number of slope cell grids.
[0069] Step 10: According to the sampling weight of non-landslide samples, determine the number of non-landslide samples in each slope unit, randomly sample low-susceptibility samples in each slope unit, and summarize them to obtain the final non-landslide samples.
[0070] Furthermore, landslide susceptibility prediction includes the following steps:
[0071] Step 1: According to the actual landslide data, determine the landslide point or landslide surface sampling method to obtain landslide sample data;
[0072] Step 2: Use the above non-landslide sample sampling method to obtain non-landslide sample data;
[0073] Step 3: Select landslide susceptibility evaluation factors and preprocess the landslide susceptibility evaluation factors, including projection transformation, resampling, and attribute classification;
[0074] Step 4: Perform factor collinearity analysis on the landslide susceptibility evaluation factors to remove factors that are irrelevant to the model building. The factor collinearity analysis can be performed using the Variance Inflation Factor (VIF) diagnostic method, and the calculation formula is as follows:
[0075]
[0076] In formula (4), it is assumed that the independent variable X = [X 1 ,X 2 ,X 3 ,…,X N ], R 2 j is the jth independent variable X j The coefficient of determination between the other independent variables. It is generally believed that when the variance inflation factor (VIF) value exceeds 10, it indicates that the variable factor has serious multicollinearity problems and needs to be eliminated.
[0077] Step 5: Randomly divide the landslide data, non-landslide data and evaluation factor data into training data set and test data set according to a certain ratio (such as 7:3);
[0078] Step 6: Select machine learning algorithms, such as support vector machine (SVM), decision tree, random forest, etc., as landslide susceptibility prediction methods;
[0079] Step 7: Use machine learning algorithms to model landslide susceptibility. Taking the SVM method as an example, use the training data set and the test data set to train and test the SVM model respectively to obtain the optimal SVM model.
[0080] Step 8: Input the evaluation factor data into the optimal SVM model and calculate the landslide susceptibility prediction results of the study area.
[0081] Exemplarily, the machine learning algorithm in step 6 and the SVM method mentioned in step 7 for landslide susceptibility prediction can be replaced by other similar data-driven algorithms, such as artificial neural network models and deep neural network models.
[0082] The present invention proposes a data sampling method that takes into account geographic space constraints for data sampling in landslide susceptibility prediction, especially data sampling of non-landslide samples. This method is combined with a machine learning algorithm to construct a landslide susceptibility prediction algorithm that takes into account geographic space constraints sampling. This research method will effectively improve the data sampling accuracy in landslide susceptibility prediction and improve the reliability and accuracy of landslide susceptibility prediction results.
[0083] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0084] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A landslide susceptibility prediction method taking into account geospatial constrained sampling, characterized in that: include: Select landslide susceptibility evaluation factors based on the actual situation of the target area and the collected geographic basic data; Preprocessing the landslide susceptibility evaluation factor, and using the frequency ratio method to calculate the relative influence of each attribute interval of the preprocessed landslide susceptibility evaluation factor on the occurrence of landslide; The landslide susceptibility evaluation factors after pretreatment were ranked and sorted by using the analytic hierarchy process, and the relative weights of the landslide susceptibility evaluation factors were determined by pairwise comparison. Performing weighted calculation according to the relative weight and the relative influence degree to obtain a susceptibility index of the target area, and extracting a low-susceptibility area as a feature space sampling area; Based on the digital elevation model, ArcGIS software was used to extract the slope units in the target area as the geospatial sampling area; Superimposing the feature space sampling area and the geographic space sampling area to obtain a sampling area; Calculate the ratio of the number of low-susceptibility sample points in the slope unit to the number of slope unit grids, and normalize the ratio to obtain the sampling weight; According to the sampling weight, the number of non-landslide samples in each slope unit is determined, and low-susceptibility samples in each slope unit are randomly sampled to obtain the final non-landslide sample data; Performing landslide susceptibility prediction based on the non-landslide sample data to obtain a landslide susceptibility prediction result for the target area; The calculation formula of the susceptibility index is: Where LSI represents the landslide susceptibility index, Q i represents the relative weight of the evaluation factors calculated by the hierarchical analysis method, F i It represents the relative influence of the evaluation factors calculated by the frequency ratio method, and m represents the total number of evaluation factors; The formula for calculating the ratio of the number of low-susceptibility sample points in a slope unit to the number of slope unit grids is: Among them, W i represents the sampling weight of the i-th slope unit, n represents the total number of slope units, N i represents the number of low-susceptibility samples in the slope unit, S i Indicates the number of slope cell grids.
2. The method for predicting landslide susceptibility taking into account geographic spatial constrained sampling according to claim 1, characterized in that: The basic geographic data include topographic data, landform data, geological data, river network data, road data and annual average rainfall data.
3. The method for predicting landslide susceptibility taking into account geographic spatial constrained sampling according to claim 1, characterized in that: The landslide susceptibility evaluation factors include: any one or several of the following: slope, slope aspect, elevation, land use, stratum lithology, distance to a river, distance to a road, distance to a fault, and annual average rainfall.
4. The method for predicting landslide susceptibility taking into account geographic spatial constrained sampling according to claim 1, characterized in that: The calculation formula for the relative impact degree is: Among them, F i Indicates the relative influence of the i-th attribute interval of the evaluation factor, L i represents the landslide area within the i-th attribute interval of the evaluation factor, L represents the total landslide area in the target area, and A i It represents the area of the i-th attribute interval of the evaluation factor, and A represents the total area of the target area.
5. The method for predicting landslide susceptibility taking into account geographic spatial constrained sampling according to claim 1, characterized in that: The landslide susceptibility prediction is performed according to the non-landslide sample data to obtain the landslide susceptibility prediction result of the target area, including: Determine the landslide point or landslide surface sampling method based on the actual landslide data, and obtain landslide sample data by sampling; The pre-processed landslide susceptibility evaluation factors were subjected to factor collinearity analysis to remove factors that were irrelevant to the modeling. The landslide sample data, the non-landslide sample data and the removed evaluation factors are randomly divided into a training data set and a test data set according to a preset ratio; The landslide susceptibility is modeled using machine learning algorithms, and the machine learning model is trained and tested using training and test data sets to obtain the optimal prediction model. The evaluation factor data to be tested is input into the prediction model to calculate and obtain the landslide susceptibility prediction result.
6. The method for predicting landslide susceptibility taking into account geographic spatial constrained sampling according to claim 5, characterized in that: The factor collinearity analysis is performed using the variance inflation factor diagnostic method; the calculation formula of the variance inflation factor diagnostic method is: In which, assume that the independent variable X = [X1, X2, X3, …, X N ], is the jth independent variable X j VIF is the variance inflation factor, and N is the number of independent variables.
7. The method for predicting landslide susceptibility taking into account geographic spatial constrained sampling according to claim 5, characterized in that: The machine learning model is any one of a support vector machine model, a decision tree model and a random forest model.
8. The method for predicting landslide susceptibility taking into account geographic spatial constrained sampling according to claim 5, characterized in that: The preset ratio is 7:3.
Citation Information
Patent Citations
Landslide susceptibility prediction method and system based on semi-supervised support vector machine model
CN114036841A
Landslide susceptibility evaluation method and system
CN114091274A
Landslide susceptibility evaluation method for areas with incomplete landslide sample data
CN114462835A