A non-landslide sample selection method for susceptibility modeling and computer equipment
By optimizing the selection of non-landslide samples through the calculation of non-landslide sample indices and ROC curves, and combining random functions and logistic regression models, the problem of insufficient representativeness of non-landslide samples was solved, improving the accuracy and stability of landslide susceptibility prediction and providing stronger evidence for landslide disaster prevention and control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies for landslide susceptibility assessment, the selection of non-landslide samples is not representative enough, resulting in insufficient accuracy and stability of machine learning models, which affects the accuracy of landslide susceptibility prediction.
By calculating the non-landslide sample index, the selection method of non-landslide samples is optimized using the information content model and ROC curve. By combining random functions to select non-landslide samples with an equal number of landslide samples, a more representative training set and validation set are formed. The logistic regression model is then applied to predict landslide susceptibility.
It improves the accuracy and stability of landslide susceptibility prediction models, reduces the uncertainty of non-landslide samples, provides a more accurate nonlinear functional relationship between landslides and evaluation indicators, and enhances the basis for landslide disaster prevention and control.
Smart Images

Figure CN115270598B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of geological disaster prediction, in particular to a non-landslide sample selection method for landslide susceptibility modeling and a computer device. BACKGROUND
[0002] As a major geological disaster, landslide has the characteristics of concealment, suddenness, uncertainty and strong destructiveness. Landslide susceptibility evaluation is a method to predict the spatial distribution of landslide susceptibility in the study area by combining the characteristics of topography and geomorphology. Through the landslide susceptibility zoning map, the staff can quickly identify the high landslide area. Therefore, scientific and accurate landslide susceptibility evaluation is not only the key technology of landslide monitoring and early warning, but also the basic work of disaster prevention and reduction.
[0003] The main research content of landslide susceptibility evaluation is the spatial probability of landslide occurrence at a certain location under certain conditions in a regional range. Generally, landslide susceptibility evaluation system includes two processes: quantitative evaluation of indicators and selection of evaluation model. Common quantitative methods include information content method, expert scoring method and frequency ratio method. Among them, the frequency ratio method, which starts from the structural information of the data itself, breaks the limitation of subjective scoring and realizes the objective quantification of the secondary attributes of landslide factors. The information content method selects the influence factors of geological disasters, analyzes the spatial distribution characteristics of each factor, calculates the information content of each factor in different intervals, and obtains the landslide comprehensive information content by using superposition analysis method. In the selection of evaluation model, the commonly used models include machine learning methods such as random forest and support vector machine.
[0004] Current research on landslide susceptibility evaluation mostly focuses on the selection of susceptibility model and the quantification method, and few studies are conducted on the selection of non-landslide samples. For machine learning models, the selection of samples usually determines the accuracy and stability of the entire model. If some points in the high susceptibility area are selected as non-landslide samples, it will have a certain impact on the modeling accuracy. In order to more accurately express the landslide susceptibility zoning, it is necessary to study the selection of samples in the modeling of machine learning algorithm to a certain extent.
[0005] When using machine learning method to conduct landslide susceptibility zoning, a major influencing factor is the selection of non-landslide samples. Currently, there are three main ways to select non-landslide samples when using machine learning method to conduct landslide susceptibility zoning: randomly selecting from non-landslide areas in the study area; randomly selecting outside a certain buffer distance from the landslide; and randomly selecting from the extremely low susceptibility area. However, all of them have certain limitations, and the representative of the selected non-landslide samples is not strong, and there is a certain error in the nonlinear function relationship between the trained model and the evaluation index. SUMMARY
[0006] The present application aims to solve the problems in the prior art, and provide a non-landslide sample selection method for susceptibility modeling, which provides comprehensive and representative non-landslide samples for susceptibility modeling based on a machine learning method.
[0007] According to one aspect of the present application, the present application provides a non-landslide sample selection method for susceptibility modeling, comprising the following steps:
[0008] S1: obtaining the spatial distribution position of landslides according to historical landslide catalog data; extracting landslide susceptibility evaluation indexes from multi-source data and dividing them into multiple evaluation units;
[0009] S2: calculating the frequency ratio of each evaluation index, and performing grading and normalization processing according to the frequency ratio;
[0010] S3: calculating the information content of each evaluation index of each grade after grading and normalization processing according to the information content model, and obtaining the information content of each evaluation unit in the study area through superposition analysis;
[0011] S4: dividing non-landslides into two categories according to the information content of each evaluation unit, i.e., information content less than or equal to 0 and information content greater than 0, simultaneously selecting non-landslide samples by using a non-landslide sample index, determining the optimal value of the non-landslide sample index by comparing the ROC curve of landslide susceptibility evaluation results, and selecting non-landslide units with the same number of landslide samples as the final non-landslide samples under the condition of the optimal value.
[0012] Preferably, in step S1, the landslide susceptibility evaluation indexes include a digital elevation model (DEM), a slope, a slope direction, a runoff intensity index, a humidity index, a ground roughness, a terrain relief degree, a curvature, a profile curvature, a plane curvature, a fault, an engineering geology rock group, a water system, a road, a land use, and a normalized vegetation index (NDVI).
[0013] Preferably, step S1 comprises:
[0014] S11: importing historical landslide catalog data into ArcGIS, determining the position of landslide points, and obtaining a landslide raster layer;
[0015] S12: extracting landslide susceptibility evaluation indexes from multi-source data according to landslide characteristics, and standardizing the raster layer of all evaluation indexes into a 30m*30m raster unit set.
[0016] Preferably, step S2 comprises:
[0017] S21: for the continuity evaluation index, firstly, preliminary discretization is carried out at certain intervals, the overall frequency ratio distribution curve of the evaluation index is obtained by statistics, and then the continuous evaluation index is discretized again by taking the mutation point of the distribution curve as the critical value, and the grading states with the same influence on the development of landslide are combined into the same class;
[0018] The calculation formula of the frequency ratio is as follows:
[0019]
[0020] In the formula, j represents the jth class, is the sum of the landslide unit areas in the jth class, A j is the total area of the evaluation unit in the jth class; S0 is the sum of the unit areas containing geological disasters, and S is the total area of the evaluation unit in the study area. The FR index represents the importance of the grading states of the index factor to the occurrence of landslide. FR>1 indicates that the state promotes the occurrence of landslide, and FR≤1 indicates that the state is not conducive to the occurrence of landslide.
[0021] S22: for the discrete evaluation index, the original index grading interval is directly used;
[0022] S23: the frequency ratio of each grade of each evaluation index is calculated, and each evaluation index is normalized according to the frequency ratio.
[0023] Preferably, step S3 comprises:
[0024] According to the information content model, the information content of each grade of each evaluation index is calculated, and the information content raster layer of each evaluation index is superimposed and analyzed to obtain the total information content raster layer. Each evaluation unit has a corresponding information content. The calculation formula of the information content is as follows:
[0025]
[0026] In the formula, I is the prediction value of the information content of the prediction area; m is the mth evaluation index, n is the number of evaluation indexes; N m is the area occupied by the factor x m . is the total area of the factor x containing geological disasters; S is the total area of the evaluation unit in the study area; S0 is the sum of the unit areas containing geological disasters.
[0027] Preferably, step S4 comprises:
[0028] S41: the number of landslide samples is determined according to the landslides in the historical landslide catalog data;
[0029] S42: non-landslide in the study area is bounded with information value 0, divided into two parts less than or equal to 0 and greater than 0; the non-landslide sample index is used to find the different selection range of non-landslide samples, wherein the calculation formula of the non-landslide sample index is:
[0030]
[0031] In the formula: N is the non-landslide sample index; X1 is the number of non-landslide samples with information less than 0; X2 is the number of non-landslide samples with information greater than 0;
[0032] S43: the ratio of non-landslide samples to landslide samples is 1:1, the non-landslide samples with different non-landslide sample index values are combined with landslide samples to form a data set, which is randomly divided into a training set and a validation set in a certain proportion, and the test set is all non-landslide and landslide samples in the study area, the training set is input into the logistic regression model to obtain the trained logistic regression model, the test set is input into the trained logistic regression model to obtain the non-landslide sample index when the ROC of the test set is the highest, the best value of the non-landslide sample index is determined, and the non-landslide samples are randomly extracted under the condition of the best value.
[0033] Preferably, in step S4:
[0034] Let P be the conditional probability of landslide occurrence, then the mathematical expression of the logistic regression model is:
[0035]
[0036] In the formula: β0, β1, β2, …, β i are logistic regression coefficients; X0, X1, X2, …, X i are independent variables; solving P of the formula, that is, obtaining the probability of landslide occurrence:
[0037]
[0038] In the formula: P is the probability of landslide occurrence in a single evaluation unit, the greater the P value, the more prone the unit is to landslide, and the higher the susceptibility grade, and the study area can be divided according to the P value.
[0039] According to another aspect of the present application, the present application provides a computer device, which comprises a memory and a processor, and a computer program is stored on the memory, and the computer program is executed by the processor to execute the non-landslide sample selection method.
[0040] The technical scheme provided by the present application has the following beneficial effects:
[0041] Randomly selected non-landslide sample as the training set and validation set of the machine learning model has strong uncertainty, resulting in a large number of errors in learning the characteristics of landslide samples and non-landslide samples, which reduces the modeling accuracy of landslide susceptibility prediction to some extent. The non-landslide sample based on the best value of the non-landslide sample index in the study area reduces the uncertainty of the non-landslide sample in the modeling process, improves the modeling accuracy, and makes the non-landslide sample more representative, so that the trained model can more accurately reflect the nonlinear function relationship between landslide and evaluation index. In summary, the method for selecting non-landslide samples according to the non-landslide sample index can provide comprehensive and representative non-landslide samples for modeling based on machine learning methods, and can be used as a basis for landslide disaster prevention. BRIEF DESCRIPTION OF DRAWINGS
[0042] The specific effects of the present application will be further described below in combination with the drawings and examples, and the drawings are as follows:
[0043] Figure 1 The flowchart of the non-landslide sample selection method for the present application;
[0044] Figure 2 The non-landslide sample distribution diagram of the present application; DETAILED DESCRIPTION
[0045] In order to have a clearer understanding of the technical features, objects and effects of the present application, the specific embodiments of the present application will be described in detail with reference to the drawings.
[0046] REFERENCE Figure 1 The embodiment provides a non-landslide sample selection method for susceptibility modeling, which comprises the following steps:
[0047] S1: obtaining the spatial distribution position of the landslide according to the historical landslide catalog data; extracting the landslide susceptibility evaluation index from the multi-source data and dividing it into multiple evaluation units;
[0048] Further, the specific steps of step S1 are as follows:
[0049] S11: determining the spatial distribution position of the landslide point according to the historical landslide catalog data to obtain the landslide grid layer;
[0050] S12: dividing the evaluation unit of the study area to be predicted for landslide susceptibility;
[0051] S13: extracting the evaluation index from the multi-source data according to the landslide characteristics, including topography, engineering geology, hydrological environment and surface covering, etc. All evaluation index grid maps are standardized to 30m*30m grid unit set.
[0052] S2: Calculate the frequency ratio of each evaluation index, and perform hierarchical and normalization processing according to the frequency ratio;
[0053] Further, step S2 specifically includes:
[0054] S21: For continuous evaluation indexes such as DEM and NDVI, first perform preliminary discretization at a very short interval, and then obtain the overall frequency ratio distribution curve of the index. Then, take the mutation point of the distribution curve as the critical value to perform secondary discretization of the continuous index, and merge the hierarchical states with the same influence on landslide development into the same class.
[0055] Taking the slope as an example: the slope is between 0° and 90°, which is first subdivided into 30 levels at an interval of 3° to obtain the frequency ratio distribution curve after preliminary discretization, and then the hierarchical states with the same influence on landslide development are merged into the same class for secondary discretization, and finally the slope index is divided into 5 levels.
[0056] S22: For discrete evaluation indexes such as land use type and slope direction, directly use the original index hierarchical interval.
[0057] Taking the slope direction as an example: the slope direction is divided into 9 types, i.e. plane, north, northwest, northeast, west, east, southeast, southwest, and south. The frequency ratio of each type can be directly obtained without the need for further discretization of the hierarchical interval.
[0058] S23: Calculate the frequency ratio of each level of each evaluation index, and perform normalization processing according to the frequency ratio.
[0059] S3: According to the information content model, calculate the information content of each level of each evaluation index after hierarchical and normalization processing, and obtain the information content of each evaluation unit in the study area through superposition analysis;
[0060] Further, step S3 specifically includes the following steps:
[0061] According to the information content model, calculate the information content of each level of each evaluation index, and perform superposition analysis on the information content raster layer of each evaluation index to obtain the total information content raster layer. Each evaluation unit has its corresponding information content.
[0062] S4: According to the information content of each evaluation unit, divide the non-landslide into two categories: information content less than or equal to 0 and information content greater than 0. At the same time, select non-landslide samples using the non-landslide sample index. By comparing the ROC curve of the landslide susceptibility evaluation results, determine the best value of the non-landslide sample index. Under the condition of the best value, use a random function to select the same number of non-landslide units as the landslide samples as the final non-landslide samples.
[0063] Further, step S4 specifically comprises the following steps:
[0064] S41: Determine the number of landslide samples according to the landslides in the historical landslide record data;
[0065] S42: Divide the non-landslide in the study area into two parts with information value 0 as the boundary, i.e., less than or equal to 0 and greater than 0; find different values of non-landslide samples by using non-landslide sample index, wherein the calculation formula of the non-landslide sample index is:
[0066]
[0067] In the formula, N is the non-landslide sample index; X1 is the number of non-landslide samples with information less than 0; X2 is the number of non-landslide samples with information greater than 0;
[0068] Taking N=2 as an example, when the non-landslide sample index N=2, it means that the number of landslide samples is 93687; the non-landslide units in the study area, i.e., the non-landslide units, are divided into two categories (information less than or equal to 0 and information greater than 0), 62458 are randomly selected from the part with information less than or equal to 0, and 31229 are randomly selected from the part with information greater than 0, which together form the non-landslide sample set.
[0069] S43: The ratio of non-landslide samples to landslide samples is 1:1; the non-landslide samples and landslide samples under different non-landslide sample index values are combined to form a data set, which is randomly divided into a training set and a validation set in a ratio of 7:3, and the test set is all non-landslide and landslide samples in the study area; the training set is input into the logistic regression model to obtain the trained logistic regression model; the test set is input into the trained logistic regression model to obtain the non-landslide sample index when the ROC of the test set is the highest, and the best value of the non-landslide sample index is determined; under the condition of the best value, non-landslide samples equal in number to landslide samples are randomly selected.
[0070] The logistic regression model is a generalized linear regression analysis model, which can form a multiple regression relationship between a dependent variable and multiple independent variables, thereby predicting the occurrence probability of an event in a certain area. The dependent variable of the logistic regression is a classification variable, and whether a landslide occurs is taken as the dependent variable, with the value "1" representing the occurrence of a landslide and the value "0" representing the non-occurrence of a landslide. Applying logistic regression to landslide susceptibility evaluation belongs to binary logistic regression analysis.
[0071] The application will be further described in conjunction with the drawings and specific embodiments, but not as a limitation of the application.
[0072] The research area of this embodiment is Fengjie County in Chongqing. Fengjie County is located in the northeast of Chongqing, straddling the upper reaches of the Yangtze River, and is located in the junction of Sichuan Basin, Upper Yangtze Fold Belt and Dabashan Fold Belt. The geological conditions in the county are complex, and it is a heavy disaster area of geological disasters. Unstable high slope, landslide, collapse, landslide and other geological disasters are widely distributed.
[0073] S1: According to the historical landslide catalog data, the spatial distribution position of the landslide is obtained; according to the multi-source data, the landslide susceptibility evaluation index is extracted and divided into multiple evaluation units, specifically:
[0074] According to the existing historical landslide catalog data, combined with the high-resolution image data of Google Map at the same period, the types, development and distribution and scale of geological disasters in Fengjie County are understood, and finally 1525 landslides are determined, the spatial distribution position of the landslide point is determined, and the landslide grid layer is obtained. The evaluation unit division is carried out in the research area which needs to be predicted for landslide susceptibility, that is, 30m*30m grid unit.
[0075] According to the characteristics of landslides, 16 evaluation indexes are extracted from multi-source data (Landsat OLI-8 remote sensing image data with a spatial resolution of 30m, digital elevation model DEM with a spatial resolution of 20m, geological structure data, 1:100,000 information map, land use type map with a spatial resolution of 10m), which are digital elevation model DEM, slope, slope direction, runoff intensity index, humidity index, ground roughness, terrain relief degree, curvature, profile curvature, plan curvature, fault, engineering geology rock group, water system, road, land use and normalized vegetation index NDVI. All evaluation index grid layers are standardized into 30m*30m grid unit set.
[0076] S2: Calculate the frequency ratio of each evaluation index, and grade and normalize according to the frequency ratio;
[0077] S3: According to the information content model, the information content of each grade of each evaluation index is calculated, and the information content of each evaluation unit in the research area is obtained through superposition analysis;
[0078] S4: All landslides are taken as samples, and non-landslide samples are selected according to non-landslide sample index. Through comparison experiment, it is determined that under the best value condition of non-landslide sample index, the modeling precision of landslide susceptibility prediction is the highest; specifically:
[0079] All landslides in the historical landslide catalog data in the research area are taken as samples, that is, 1525 landslides, which are converted into 93687 evaluation units. The ratio of non-landslide sample to landslide sample is 1:1;
[0080] The best value range of the non-landslide sample is found by using the non-landslide sample index. The ratio of the information amount less than or equal to 0 to the information amount greater than 0 is 50:1, 40:1, 30:1, 20:1, 10:1, 9:1, 8:1,..., 1:1, 1:2, 1:3,..., 1:10,..., 1:50, i.e. N=50, 40, 30, 20, 10, 9, 8,..., 1, 1 / 2, 1 / 3,..., 1 / 10,..., 1 / 50, a total of 27 groups of experiments. Taking the ratio of the information amount less than or equal to 0 to the information amount greater than 0 as 2:1 for example, the number of non-landslide samples is 93687, 62458 non-landslide samples are randomly selected from all non-landslide grids with the information amount less than or equal to 0 in the study area, and 31229 non-landslide samples are randomly selected from all non-landslide grids with the information amount greater than 0 in the study area, to form a non-landslide sample set.
[0081] The non-landslide samples with different values of the non-landslide sample index and the landslide samples are combined to form a data set, which is randomly divided into a training set and a validation set at a ratio of 7:3, and the test set is composed of all non-landslide samples and landslide samples in the study area.
[0082] The training set is input into a logistic regression model, the trained model is used for landslide susceptibility analysis, the susceptibility of the entire study area is predicted, and the landslide susceptibility probability value of each evaluation unit in the study area is obtained. According to the receiver operating characteristic curve ROC, the value of the non-landslide sample index of the ROC of the test set is N=4. From the non-landslide grids that meet the condition of the non-landslide sample index N=4, an equal amount of non-landslide samples and landslide samples are randomly selected as the final non-landslide sample set, and the distribution of the non-landslide samples is shown in FIG. 6. Figure 2
[0083] As an optional implementation, the embodiment provides a computer device, which includes a memory and a processor, and the memory stores a computer program. When the computer program is executed by the processor, each process of the non-landslide sample selection method embodiment described above is performed, and the same technical effect is achieved. To avoid repetition, details are not described here.
[0084] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that processes, methods, articles, or systems that include a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, methods, articles, or systems. Without more limitations, the element defined by the statement "including a..." does not exclude the presence of other identical elements in the process, method, article, or system including the element.
[0085] The above-mentioned embodiment serial numbers of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments. In the unit claims in which several devices are listed, several of the devices can be embodied by the same hardware item. The use of the words first, second, and third does not represent any order, and the words can be interpreted as identifiers.
[0086] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, which is made by using the content of the specification and drawings of the present application, is also included in the patent protection scope of the present application.
Claims
1. A non-slip sample selection method for susceptibility modeling, characterized by, Comprising the following steps: S1: obtaining the spatial distribution position of landslides according to historical landslide catalog data; extracting landslide susceptibility evaluation indexes from multi-source data and dividing them into multiple evaluation units; S2: calculating the frequency ratio of each evaluation index, grading and normalizing according to the frequency ratio; S3: according to the information content model, calculating the information content of each evaluation index of each evaluation index after grading and normalization, and obtaining the information content of each evaluation unit in the study area through superposition analysis; S4: according to the information content of each evaluation unit, dividing non-landslide into two categories with information content less than or equal to 0 and greater than 0, and selecting non-landslide samples by using non-landslide sample index; by comparing the ROC curve of landslide susceptibility evaluation results, the best value of non-landslide sample index is determined, and under the condition of the best value, the same number of non-landslide units as landslide samples are selected as the final non-landslide samples by using random function; The calculation formula of non-landslide sample index is: wherein: is the index of non-landslide samples; is the number of non-landslide samples with information content less than or equal to 0; is the number of non-landslide samples with information content greater than 0; Step S2 comprises: S21: for continuous evaluation indexes, firstly, preliminary discretization is carried out at a certain interval, and the overall frequency ratio distribution curve of the evaluation index is obtained, and then the continuous evaluation index is discretized again by taking the mutation point of the distribution curve as the critical value, and the same grading state with the same influence on landslide development is combined into the same category; The calculation formula of frequency ratio is as follows: In the formula: represents the jth category, is the sum of landslide unit areas in the jth category, is the total area of evaluation units in the jth category; is the sum of unit areas containing geological disasters, is the total area of evaluation units in the study area; FR The index indicates the importance of the state of the index factor at each level to the occurrence of landslides, FR > 1 indicates that the state promotes the occurrence of landslides, FR ≤ 1 indicates that the state is not conducive to the occurrence of landslides; S22: for discrete evaluation indexes, the original index grading interval is directly used; S23: calculate the frequency ratio of each level of each evaluation index, and normalize each evaluation index according to the frequency ratio; Step S4 comprises: S41: determining the number of landslide samples according to the landslides in the historical landslide catalog data; S42: dividing the non-landslide in the study area into two parts with information content value 0 as the boundary, i.e., less than or equal to 0 and greater than 0; find the different selection range of non-landslide samples by using non-landslide sample index; S43: the ratio of non-landslide samples to landslide samples is 1:1, the non-landslide samples under different non-landslide sample index values and landslide samples form a data set, which is randomly divided into training set and validation set in a certain proportion, and the test set is all non-landslide and landslide samples in the study area, the training set is input into the logistic regression model to obtain the trained logistic regression model, the test set is input into the trained logistic regression model to obtain the non-landslide sample index when the ROC of the test set is the highest, and the best value of the non-landslide sample index is determined, and under the condition of the best value, the same number of non-landslide samples as landslide samples are randomly selected.
2. The non-slip sample selection method of claim 1, wherein, In step S1, the landslide susceptibility evaluation indexes include: digital elevation model DEM, slope, slope direction, runoff intensity index, humidity index, ground roughness, terrain relief, curvature, profile curvature, plane curvature, fault, engineering geology rock group, water system, road, land use and normalized vegetation index NDVI.
3. The non-slip sample selection method of claim 1, wherein, Step S1 comprises: S11: import the historical landslide catalog data into ArcGIS, determine the landslide point position, and obtain the landslide raster layer; S12: Extract landslide susceptibility evaluation indexes from multi-source data according to landslide characteristics, and standardize all evaluation index raster layers into 30m 30m grid unit set, each grid unit is an evaluation unit.
4. The non-slip sample selection method of claim 1, wherein, Step S3 comprises: According to the information content model, the information content of each evaluation index at each level is calculated, and the information content grid layers of each evaluation index are superimposed and analyzed to obtain a total information content grid layer, each evaluation unit has a corresponding information content, and the calculation formula of the information content is: In the formula: is the prediction value of the information amount of the prediction area; m is the i-th evaluation index, m is the number of evaluation indexes; is the number of evaluation indexes; is the area occupied by the i-th factor; is the area occupied by the i-th factor; is the total area of the occurrence of geological disasters in the i-th factor; x is the total area of the occurrence of geological disasters in the i-th factor; is the total area of the evaluation unit in the study area; is the total area of the evaluation unit in the study area.
5. The non-slip sample selection method of claim 1, wherein, In step S4: Let a condition probability of landslide occurrence be P, and a mathematical expression of the logistic regression model is: In the formula: , , , ..., These are the logistic regression coefficients; , , , ..., As the independent variable; for the formula P Solving for the probability of a landslide, we obtain the following: In the formula: This represents the probability of a landslide occurring within a single evaluation unit. The higher the value, the more prone the unit is to landslides, and the higher the susceptibility level. This can be determined based on... P The study area was divided into landslide susceptibility zones.
6. A computer device, comprising: The computer device comprises a memory and a processor, and the memory stores a computer program. When the computer program is executed by the processor, the non-landslide sample selection method according to any one of claims 1-5 is executed.