Landslide susceptibility prediction method and system coupled with smoteenn and tabtransformer
Patent Information
- Application Number
- CN202310099090.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-01-29
AI Technical Summary
但传统的机器学习模型为浅层结构算法,在有限样本和计算单元的情况下对复杂函数的表示能力有限,无法充分挖掘滑坡的底层特征,导致易发性评估的可靠性较差
[0048]本发明构建了一种耦合Smoteenn和Tabtransformer的滑坡易发性预测方法,该方法充分考虑并表示了滑坡灾害的致灾因子与滑坡发育之间的非线性关系,对滑坡的致灾因子进行量化分析,为滑坡灾害的防治工作提供决策支持。
Smart Images

Figure CN116070762B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of geological disaster prevention and control, and in particular to a landslide susceptibility prediction method and system that couples Smoteenn and Tabtransformer. Background Technology
[0002] A landslide refers to the phenomenon where rock and soil on a slope slide downhill, either as a whole or in parts, under the influence of gravity, due to factors such as heavy rain, groundwater, earthquakes, or human activities.
[0003] Landslide susceptibility assessment can predict the locations where landslides are likely to occur. Selecting efficient and reliable quantitative assessment methods can prevent and reduce loss of life and property caused by landslides, and is of great significance for landslide disaster prevention. The accuracy of landslide susceptibility assessment is mainly affected by the sample size and the assessment method.
[0004] A common phenomenon in landslide susceptibility assessment is that the number of non-landslide samples far exceeds the number of landslide samples, a typical example of imbalanced data. Due to the lack of information on minority class samples (i.e., landslide samples), many classification algorithms struggle to accurately represent the inherent characteristics of landslides, significantly compressing the decision boundaries in the classification system. While detection models exhibit excellent overall accuracy, they cannot effectively detect the minority landslide samples. Related research indicates that the main difficulties and challenges in classifying imbalanced data are not caused by the imbalance itself, but rather by the inherent classification characteristics of the data. For example, the minority class samples are too few and unrepresentative; overlapping samples from different classes make it difficult to form clear dividing boundaries; and the discontinuous distribution of minority class samples complicates class segmentation. The imbalanced relationship between the number of landslide and non-landslide samples poses a challenge to improving the accuracy of landslide susceptibility assessment models.
[0005] Traditional machine learning methods can effectively reflect the nonlinear relationship between landslides and hazard-causing factors, and have been widely used in the field of landslide susceptibility assessment, such as random forests, support vector machines, and gradient boosting trees. However, traditional machine learning models are shallow-structured algorithms, and their ability to represent complex functions is limited with limited samples and computational units. They cannot fully explore the underlying features of landslides, resulting in poor reliability of susceptibility assessment. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a landslide susceptibility prediction method coupled with Smoteen and Tabtransformer, comprising the following steps:
[0007] S1: Obtain the spatial distribution of landslides and non-landslides and their causative factors in the study area, correlate the spatial distribution of landslides and non-landslides and their causative factors, and obtain a raster layer of landslide and non-landslide distributions.
[0008] S2: Calculate the information content of each disaster-causing factor, normalize the disaster-causing factors based on the information content, and obtain the processed disaster-causing factors.
[0009] S3: Construct training and validation sets using raster layers of landslide and non-landslide distributions, and resample the training set using the Smoteen algorithm to obtain a new training set;
[0010] S4: Construct a Tabtransformer model using the processed disaster-causing factors, train the Tabtransformer model using a new training set to obtain a trained Tabtransformer model, and obtain the landslide susceptibility assessment results for the study area using the trained Tabtransformer model.
[0011] S5: The study area is divided based on the landslide susceptibility assessment results to obtain a landslide disaster susceptibility prediction map.
[0012] Preferably, step S1 specifically includes:
[0013] S11: Based on historical landslide logging data and combined with high-resolution remote sensing images, the spatial distribution of landslides and non-landslides in the study area was obtained;
[0014] S12: Extract the disaster-causing factors that affect landslide development in the study area, map the disaster-causing factors to the raster cells of the spatial distribution of landslides and non-landslides, and obtain the raster layers of landslide and non-landslide distributions, wherein the raster layers contain the corresponding disaster-causing factors.
[0015] Preferably, step S2 specifically includes:
[0016] S21: Calculate the amount of information obtained for each disaster-causing factor. The calculation formula is shown in Formula 1:
[0017]
[0018] Where I represents the information content of the disaster-causing factor; i represents the i-th disaster-causing factor; and m represents the number of disaster-causing factors; x i Factors contributing to disaster i; N i For factor x i Area occupied; For factor x i The total area of the grid cells in the study area is S; S0 is the total area of the grid cells containing geological hazards.
[0019] S22: Disaster-causing factors are divided into discrete data and continuous data. For discrete data, the amount of information is calculated using Formula 1.
[0020] S23: For continuous data, the information abrupt change point is used as the critical value for multiple discretizations, and the graded states with the same impact on landslide development are merged into the same grade;
[0021] S24: Normalize each disaster-causing factor according to its information content to obtain the processed disaster-causing factors.
[0022] Preferably, step S3 specifically includes:
[0023] S31: In the raster layer of landslide and non-landslide distribution, landslide rasters are used as landslide samples and non-landslide rasters are used as non-landslide samples. The label of landslide samples is 1 and the label of non-landslide samples is 0. The landslide samples and non-landslide samples are divided into training set and validation set.
[0024] S32: For each landslide sample z in the training set a Using Euclidean distance as the standard, we find the landslide sample z. a The k most recent landslide samples;
[0025] S33: Determine the sampling ratio M based on the ratio of landslide samples to non-landslide samples in the study area, starting from the distance z. a M landslide samples z are randomly selected from the k most recent landslide samples. ab b = 1, 2, ..., M;
[0026] S34: In z a and z ab A new landslide sample z is randomly inserted between the two. new Obtain the second training set, landslide samples z new The formula is as follows:
[0027] z new =z a +rand(0,1)*|z a -z ab |
[0028] Where rand(0,1) represents any number in the interval (0,1);
[0029] S35: Use the K-nearest neighbor algorithm to predict each landslide sample in the second training set. If the predicted label of the landslide sample is inconsistent with the actual label, delete the landslide sample and obtain a new training set.
[0030] Preferably, step S4 specifically includes:
[0031] S41: Combine the processed disaster-causing factors corresponding to each raster layer in the raster layers of landslide and non-landslide distributions to construct a Tabtransformer model;
[0032] S42: Train the Tabtransformer model using a new training set, and adjust the model parameters using a trial-and-error algorithm based on the accuracy and loss value on the validation set to obtain a well-trained Tabtransformer model.
[0033] S43: Input the geological data of the study area into the trained Tabtransformer model to obtain the landslide susceptibility evaluation results, that is, the probability value of landslide occurrence corresponding to each grid cell.
[0034] Preferably, step S43 specifically includes:
[0035] S431: Extract landslide information from geological data using a trained Tabtransformer model. The calculation formula is as follows:
[0036]
[0037] Where L(x,y) is the loss function of the trained Tabtransformer model; x={x1,x2,…,x m} represents the landslide-causing factors, totaling m; y is the sample label, 1 for landslide and 0 for non-landslide; It is the set of all hazard-causing factors in the embedding layer; f θ A function for the sequence of Transformer layers; g φ For context-embedded functions;
[0038] S432: Yes Perform the operation and return the context embedding g. φ ={h1,h2,…,h m}, forming a vector of dimension d*m; input the vector into a multilayer perceptron to predict the sample label y; learn all the parameters of the trained Tabtransformer model through the loss function L(x,y), optimize the prediction results, and finally obtain the landslide susceptibility assessment results.
[0039] Preferably, step S5 specifically includes:
[0040] Based on the landslide susceptibility assessment results, the study area was divided into five levels: extremely high, high, medium, low, and extremely low susceptibility. The area proportions corresponding to each level were 45%, 25%, 15%, 10%, and 5%, respectively, and a landslide susceptibility prediction map was output.
[0041] A landslide susceptibility prediction system coupled with Smoteenn and Tabtransformer includes:
[0042] The raster layer acquisition module is used to acquire the spatial distribution of landslides and non-landslides and the disaster-causing factors in the study area, and to correlate the spatial distribution of landslides and non-landslides and the disaster-causing factors to obtain raster layers of landslide and non-landslide distributions.
[0043] The disaster-causing factor processing module is used to calculate the information content of each disaster-causing factor, and to normalize the disaster-causing factors based on the information content to obtain the processed disaster-causing factors.
[0044] The training set acquisition module is used to construct training and validation sets through raster layers of landslide and non-landslide distributions, and to resample the training set using the Smoteen algorithm to obtain a new training set.
[0045] The landslide susceptibility assessment result acquisition module is used to construct a Tabtransformer model through the processed disaster-causing factors, train the Tabtransformer model with a new training set, obtain a trained Tabtransformer model, and obtain the landslide susceptibility assessment result of the study area through the trained Tabtransformer model.
[0046] The landslide susceptibility prediction map acquisition module is used to divide the study area based on the landslide susceptibility assessment results and obtain a landslide susceptibility prediction map.
[0047] The present invention has the following beneficial effects:
[0048] This invention constructs a landslide susceptibility prediction method coupled with Smoteenn and Tabtransformer. This method fully considers and represents the nonlinear relationship between landslide hazard-causing factors and landslide development, and performs quantitative analysis on landslide hazard-causing factors to provide decision support for landslide disaster prevention and control. Attached Figure Description
[0049] Figure 1 This is a flowchart of a method according to an embodiment of the present invention;
[0050] Figure 2 Landslide susceptibility prediction map;
[0051] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0052] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0053] Reference Figure 1This invention provides a landslide susceptibility prediction method coupled with Smoteen and Tabtransformer, comprising the following steps:
[0054] S1: Obtain the spatial distribution of landslides and non-landslides and their causative factors in the study area, correlate the spatial distribution of landslides and non-landslides and their causative factors, and obtain a raster layer of landslide and non-landslide distributions.
[0055] S2: Calculate the information content of each disaster-causing factor, normalize the disaster-causing factors based on the information content, and obtain the processed disaster-causing factors.
[0056] S3: Construct training and validation sets using raster layers of landslide and non-landslide distributions, and resample the training set using the Smoteen algorithm to obtain a new training set;
[0057] S4: Construct a Tabtransformer model using the processed disaster-causing factors, train the Tabtransformer model using a new training set to obtain a trained Tabtransformer model, and obtain the landslide susceptibility assessment results for the study area using the trained Tabtransformer model.
[0058] S5: The study area is divided based on the landslide susceptibility assessment results to obtain a landslide disaster susceptibility prediction map.
[0059] In this embodiment, Fengjie County is selected as the study area. Fengjie County is located in the northeastern part of Chongqing Municipality, in the heart of the Three Gorges Reservoir area. The average elevation of the study area is 949m, with high terrain at both the north and south ends and low terrain in the middle, belonging to a typical high mountain and low valley landform. The county has a complex geological environment, frequent human engineering activities, and dynamic changes in the water level of the Three Gorges Reservoir, resulting in frequent geological disasters such as landslides and collapses.
[0060] Step S1 is as follows:
[0061] S11: Based on historical landslide logging data and combined with high-resolution remote sensing images, the spatial distribution of landslides and non-landslides in the study area was obtained;
[0062] S12: Extract the disaster-causing factors that affect landslide development in the study area, map the disaster-causing factors to the raster cells of the spatial distribution of landslides and non-landslides, and obtain the raster layers of landslide and non-landslide distributions, wherein the raster layers contain the corresponding disaster-causing factors.
[0063] Specifically, using software such as ArcGIS and ENVI, the disaster-causing factors affecting landslide development are extracted from the digital elevation model (DEM), Landsat-8 remote sensing images and geological maps, and spatial reference systems are established with the spatial distribution of landslides and non-landslides in S11, such as standard projection coordinate system and grid cell size.
[0064] In this embodiment, based on field survey data, historical landslide logging data, and high-resolution remote sensing images of the study area, 1,525 landslides were identified and imported into ArcGIS software to obtain the spatial distribution of landslides. Furthermore, the spatial locations of non-landslides can be obtained. The data format is raster image with a raster size of 30m.
[0065] Thirteen landslide-causing factors were extracted using data from digital elevation models (DEMs), including elevation, slope, aspect, plane curvature, profile curvature, humidity index, runoff intensity index, river distance, road distance, fault distance, stratigraphic lithology, land use, and normalized difference vegetation index (NDVI). Slope and aspect were extracted from the DEM; stratigraphic lithology and faults were extracted from 1:50,000 geological maps; NDVI was extracted from Landsat-8 remote sensing imagery; land use types were obtained from Tsinghua University's global 10-meter resolution land cover data; roads were extracted from the national road network data; and rivers were extracted from the DEM. River distance, fault distance, and road distance were obtained through buffer analysis of rivers, faults, and roads, respectively. The landslide-causing factors were spatially referenced to a raster layer representing the landslide spatial distribution, using a standardized projected coordinate system and raster cell size.
[0066] Furthermore, step S2 specifically involves:
[0067] S21: Calculate the amount of information obtained for each disaster-causing factor. The calculation formula is shown in Formula 1:
[0068]
[0069] Where I represents the information content of the disaster-causing factor; i represents the i-th disaster-causing factor; and m represents the number of disaster-causing factors; x i Factors contributing to disaster i; N i For factor x i Area occupied; For factor x i The total area of the grid cells in the study area is S; S0 is the total area of the grid cells containing geological hazards.
[0070] S22: Disaster-causing factors are divided into discrete data and continuous data. For discrete data, the amount of information is calculated using Formula 1.
[0071] Specifically, among the disaster-causing factors of landslides, stratigraphic lithology and land use are discrete data, while slope and NDVI are continuous data.
[0072] In this embodiment, there are three discrete data types: stratigraphy, land use, and slope aspect, which can be classified according to their inherent natural attributes; and ten continuous data types, including elevation, slope, profile curvature, plane curvature, humidity index, runoff intensity index, river distance, fault distance, normalized difference vegetation index (NDVI), and road distance.
[0073] Discrete data can be classified according to the inherent natural attributes of disaster-causing factors. For example, land use includes types such as arable land, water bodies, construction land, bare land, and construction land. Therefore, the disaster-causing factor—land use—can be divided into 5 levels.
[0074] S23: For continuous data, the information abrupt change point is used as the critical value for multiple discretizations, and the graded states with the same impact on landslide development are merged into the same grade;
[0075] S24: Normalize each disaster-causing factor according to its information content to obtain the processed disaster-causing factor; normalization can eliminate the dimensional influence between different disaster-causing factors.
[0076] Furthermore, step S3 specifically includes:
[0077] S31: In the raster layer of landslide and non-landslide distribution, landslide rasters are used as landslide samples and non-landslide rasters are used as non-landslide samples. The label of landslide samples is 1 and the label of non-landslide samples is 0. The landslide samples and non-landslide samples are divided into training set and validation set.
[0078] In this embodiment, the raster layer of landslide and non-landslide distributions in the study area contains a total of 4,575,207 raster cells. The labels of the cells are determined: the landslide samples (93,687) are labeled with 1, and the non-landslide samples (4,481,520) are labeled with 0. The number of non-landslide samples is approximately 48 times that of landslide samples. 70% of the landslide samples (65,581) and non-landslide samples (3,137,064) are randomly selected to form a training set, and the remaining 30% is used as a validation set.
[0079] S32: For each landslide sample z in the training set a Using Euclidean distance as the standard, we find the landslide sample z. a The k most recent landslide samples;
[0080] S33: Determine the sampling ratio M based on the ratio of landslide samples to non-landslide samples in the study area, starting from the distance z. a M landslide samples z are randomly selected from the k most recent landslide samples. ab b = 1, 2, ..., M; in this embodiment, the value of M is 48;
[0081] S34: In z a and zab A new landslide sample z is randomly inserted between the two. new Obtain the second training set, landslide samples z new The formula is as follows:
[0082] z new =z a +rand(0,1)*|z a -z ab |
[0083] Where rand(0,1) represents any number in the interval (0,1);
[0084] S35: Use the K-nearest neighbor algorithm to predict each landslide sample in the second training set. If the predicted label of the landslide sample is inconsistent with the actual label, delete the landslide sample and obtain a new training set.
[0085] Furthermore, step S4 specifically involves:
[0086] S41: Combine the processed disaster-causing factors corresponding to each raster layer in the raster layers of landslide and non-landslide distributions to construct a Tabtransformer model;
[0087] Specifically, the processed disaster-causing factors are combined and used as input data for the Tabtransformer model, which is then exported as a CSV file.
[0088] S42: Train the Tabtransformer model using a new training set, and adjust the model parameters using a trial-and-error algorithm based on the accuracy and loss value on the validation set to obtain a well-trained Tabtransformer model.
[0089] S43: Input the geological data of the study area into the trained Tabtransformer model to obtain the landslide susceptibility evaluation results, that is, the probability value of landslide occurrence corresponding to each grid cell.
[0090] Furthermore, step S43 specifically involves:
[0091] S431: Extract landslide information from geological data using a trained Tabtransformer model. The calculation formula is as follows:
[0092]
[0093] Where L(x,y) is the loss function of the trained Tabtransformer model; x={x1,x2,…,x m} represents the landslide-causing factors, totaling m; y is the sample label, 1 for landslide and 0 for non-landslide; It is the set of all hazard-causing factors in the embedding layer; f θ A function for the sequence of Transformer layers; g φ For context-embedded functions;
[0094] S432: Yes Perform the operation and return the context embedding g. φ ={h1,h2,…,h m}, forming a vector of dimension d*m; input the vector into a multilayer perceptron to predict the sample label y; learn all the parameters of the trained Tabtransformer model through the loss function L(x,y), optimize the prediction results, and finally obtain the landslide susceptibility assessment results.
[0095] Furthermore, step S5 specifically involves:
[0096] Based on the landslide susceptibility assessment results, the study area was divided into five levels: extremely high, high, medium, low, and extremely low susceptibility. The area proportions corresponding to each level were 45%, 25%, 15%, 10%, and 5%, respectively, and a landslide susceptibility prediction map was output.
[0097] In this embodiment, Figure 2 To predict landslide susceptibility, the accuracy of the prediction results is evaluated using Receiver Operating Characteristic Curves (ROC). To more clearly represent the evaluation effect, the Area Under the ROC Curve (AUC) is typically used as an indicator to measure the accuracy of the model's predictions. The AUC value ranges from 0 to 1; the closer the ROC curve is to the upper left corner, the larger the AUC value, indicating higher model accuracy. The AUC of the method in this invention is 85.9%, indicating that this model is an effective method for predicting landslide susceptibility.
[0098] This invention provides a landslide susceptibility prediction system coupled with Smoteenn and Tabtransformer, comprising:
[0099] The raster layer acquisition module is used to acquire the spatial distribution of landslides and non-landslides and the disaster-causing factors in the study area, and to correlate the spatial distribution of landslides and non-landslides and the disaster-causing factors to obtain raster layers of landslide and non-landslide distributions.
[0100] The disaster-causing factor processing module is used to calculate the information content of each disaster-causing factor, and to normalize the disaster-causing factors based on the information content to obtain the processed disaster-causing factors.
[0101] The training set acquisition module is used to construct training and validation sets through raster layers of landslide and non-landslide distributions, and to resample the training set using the Smoteen algorithm to obtain a new training set.
[0102] The landslide susceptibility assessment result acquisition module is used to construct a Tabtransformer model through the processed disaster-causing factors, train the Tabtransformer model with a new training set, obtain a trained Tabtransformer model, and obtain the landslide susceptibility assessment result of the study area through the trained Tabtransformer model.
[0103] The landslide susceptibility prediction map acquisition module is used to divide the study area based on the landslide susceptibility assessment results and obtain a landslide susceptibility prediction map.
[0104] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0105] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the unit claims listing several devices, several of these devices may be embodied by the same hardware item. The use of the terms first, second, and third, etc., does not indicate any order and can be interpreted as identifiers.
[0106] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A landslide susceptibility prediction method coupled with Smoteen and Tabtransformer, characterized in that, Includes the following steps: S1: Obtain the spatial distribution of landslides and non-landslides and their causative factors in the study area, correlate the spatial distribution of landslides and non-landslides and their causative factors, and obtain a raster layer of landslide and non-landslide distributions. S2: Calculate the information content of each disaster-causing factor, and normalize the disaster-causing factors based on the information content to obtain the processed disaster-causing factors, specifically: S21: Calculate the amount of information obtained for each disaster-causing factor. The calculation formula is shown in Formula 1: in, Let i be the information content of the disaster-causing factor; i is the i-th disaster-causing factor. The number of disaster-causing factors; Factors that cause disaster i; As factors Area occupied; As factors The total area affected by geological disasters in China; The total area of the grid cells in the study area; This is the sum of the areas of the grid cells containing geological hazards; S22: Disaster-causing factors are divided into discrete data and continuous data. For discrete data, the amount of information is calculated using Formula 1. S23: For continuous data, the information abrupt change point is used as the critical value for multiple discretizations, and the graded states with the same impact on landslide development are merged into the same grade; S24: Normalize each disaster-causing factor according to the amount of information it contains to obtain the processed disaster-causing factors. S3: Construct training and validation sets using raster layers representing landslide and non-landslide distributions. Resample the training set using the Smoteen algorithm to obtain a new training set. Specifically: S31: In the raster layer of landslide and non-landslide distribution, landslide rasters are used as landslide samples and non-landslide rasters are used as non-landslide samples. The label of landslide samples is 1 and the label of non-landslide samples is 0. The landslide samples and non-landslide samples are divided into training set and validation set. S32: For each landslide sample in the training set Using Euclidean distance as the standard, landslide samples were found. The k most recent landslide samples; S33: Determine the sampling ratio M based on the ratio of landslide samples to non-landslide samples in the study area, starting from the distance... M landslide samples are randomly selected from the k most recent landslide samples. b = 1, 2, ..., M; S34: In and A new landslide sample is randomly inserted between them. Obtain the second training set, landslide samples. The formula is as follows: in, Represents any number within the interval (0, 1); S35: Use the K-nearest neighbor algorithm to predict each landslide sample in the second training set. If the predicted label of the landslide sample is inconsistent with the actual label, delete the landslide sample and obtain a new training set. S4: Construct a Tabtransformer model using the processed disaster-causing factors, train the Tabtransformer model using a new training set to obtain a trained Tabtransformer model, and use the trained Tabtransformer model to obtain the landslide susceptibility assessment results for the study area, specifically: S41: Combine the processed disaster-causing factors corresponding to each raster layer in the raster layers of landslide and non-landslide distributions to construct a Tabtransformer model; S42: Train the Tabtransformer model using a new training set, and adjust the model parameters using a trial-and-error algorithm based on the accuracy and loss value of the validation set to obtain a well-trained Tabtransformer model. S43: Input the geological data of the study area into the trained Tabtransformer model to obtain the landslide susceptibility evaluation results, that is, the probability value of landslide occurrence corresponding to each grid cell. S5: The study area is divided based on the landslide susceptibility assessment results to obtain a landslide disaster susceptibility prediction map.
2. The landslide susceptibility prediction method coupled with Smoteen and Tabtransformer according to claim 1, characterized in that, Step S1 is as follows: S11: Based on historical landslide logging data and combined with high-resolution remote sensing images, the spatial distribution of landslides and non-landslides in the study area was obtained; S12: Extract the disaster-causing factors that affect landslide development in the study area, map the disaster-causing factors to the raster cells of the spatial distribution of landslides and non-landslides, and obtain the raster layers of landslide and non-landslide distributions, wherein the raster layers contain the corresponding disaster-causing factors.
3. The landslide susceptibility prediction method coupled with Smoteen and Tabtransformer according to claim 1, characterized in that, Step S43 is as follows: S431: Extract landslide information from geological data using a trained Tabtransformer model. The calculation formula is as follows: in, The loss function for the trained Tabtransformer model; There are m landslide-causing factors in total. For the sample labels, landslide is 1 and non-landslide is 0; It is the collection of all disaster-causing factors in the embedding layer; A function for the sequence of Transformer layers; For context-embedded functions; S432: Yes Perform the operation and return the context embedding. This forms a dimension d A vector of m; input the vector into the multilayer perceptron, and process the sample labels. Make predictions; using the loss function By learning all the parameters of the trained Tabtransformer model, optimizing the prediction results, and finally obtaining the landslide susceptibility assessment results.
4. The landslide susceptibility prediction method coupled with Smoteen and Tabtransformer according to claim 1, characterized in that, Step S5 is as follows: Based on the landslide susceptibility assessment results, the study area was divided into five levels: extremely high, high, medium, low, and extremely low susceptibility. The area proportions corresponding to each level were 45%, 25%, 15%, 10%, and 5%, respectively, and a landslide susceptibility prediction map was generated.
5. A landslide susceptibility prediction system coupled with Smoteen and Tabtransformer, characterized in that, include: The raster layer acquisition module is used to acquire the spatial distribution of landslides and non-landslides and the disaster-causing factors in the study area, and to correlate the spatial distribution of landslides and non-landslides and the disaster-causing factors to obtain raster layers of landslide and non-landslide distributions. The disaster-causing factor processing module is used to calculate the information content of each disaster-causing factor, and then normalize the disaster-causing factors based on the information content to obtain the processed disaster-causing factors. Specifically: The amount of information obtained for each disaster-causing factor is calculated using the formula shown in Formula 1: in, Let i be the information content of the disaster-causing factor; i is the i-th disaster-causing factor. The number of disaster-causing factors; Factors that cause disaster i; As factors Area occupied; As factors The total area affected by geological disasters in China; The total area of the grid cells in the study area; This is the sum of the areas of the grid cells containing geological hazards; Disaster-causing factors are divided into discrete data and continuous data. For discrete data, the amount of information is calculated using Formula 1. For continuous data, the information abrupt change point is used as the critical value for multiple discretizations, and the graded states with the same impact on landslide development are merged into the same grade; The disaster-causing factors are normalized according to the amount of information they contain to obtain the processed disaster-causing factors. The training set acquisition module is used to construct training and validation sets using raster layers representing landslide and non-landslide distributions. It then uses the Smoteen algorithm to resample the training set to obtain a new training set. Specifically: Landslide rasters in the raster layers of landslide and non-landslide distributions are used as landslide samples, and non-landslide rasters are used as non-landslide samples. The label of landslide samples is 1, and the label of non-landslide samples is 0. The landslide samples and non-landslide samples are divided into training set and validation set. For each landslide sample in the training set Using Euclidean distance as the standard, landslide samples were found. The k most recent landslide samples; The sampling ratio M was determined based on the ratio of landslide samples to non-landslide samples in the study area, starting from the distance. M landslide samples are randomly selected from the k most recent landslide samples. b = 1, 2, ..., M; exist and A new landslide sample is randomly inserted between the two. Obtain the second training set, landslide samples. The formula is as follows: in, Represents any number within the interval (0, 1); The K-nearest neighbors algorithm is used to predict each landslide sample in the second training set. If the predicted label of a landslide sample is inconsistent with the actual label, the landslide sample is deleted and a new training set is obtained. The landslide susceptibility assessment result acquisition module is used to construct a Tabtransformer model using processed disaster-causing factors, train the Tabtransformer model using a new training set, obtain a trained Tabtransformer model, and obtain the landslide susceptibility assessment results for the study area using the trained Tabtransformer model. Specifically: The processed disaster-causing factors corresponding to each raster layer in the raster layers of landslide and non-landslide distributions are combined to construct a Tabtransformer model; The Tabtransformer model is trained using a new training set. Based on the accuracy and loss value on the validation set, the model parameters are adjusted using a trial-and-error algorithm to obtain a well-trained Tabtransformer model. The geological data of the study area is input into the trained Tabtransformer model to obtain the landslide susceptibility evaluation results, that is, the probability value of landslide occurrence corresponding to each grid cell. The landslide susceptibility prediction map acquisition module is used to divide the study area based on the landslide susceptibility assessment results and obtain a landslide susceptibility prediction map.
Citation Information
Patent Citations
Geological disaster prediction method, device and equipment
CN111144651A
Landslide disaster risk regionalization map generation method
CN111858803A
Landslide susceptibility evaluation method based on sample automatic selection and earth surface deformation rate
CN113343563A