An extreme rainfall group landslide susceptibility assessment method fusing diffusion generation and explainable learning
By integrating diffusion generation and interpretable learning, this study addresses the challenges of sample scarcity and multi-factor coupling in landslide susceptibility assessment under extreme rainfall scenarios. It achieves high-precision landslide susceptibility assessment and risk classification, supporting regional early warning and disaster prevention decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-04-10
AI Technical Summary
Existing landslide susceptibility assessment methods suffer from problems such as scarce samples, difficulty in identifying boundary samples, and difficulty in interpreting multi-factor coupling relationships under extreme rainfall scenarios, resulting in insufficient prediction accuracy and limited usability of results.
By employing a fusion-diffusion generation and interpretable learning approach, and through multi-source factor data construction, feature response analysis, generative supplementation, and integrated discrimination, high-precision prediction and hierarchical representation of landslide susceptibility are achieved.
It significantly improves the stability and reliability of landslide probability prediction, provides engineering-traceable assessment results, and supports risk classification and disaster prevention decisions.
Smart Images

Figure CN121542895B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of extreme rainfall-induced group landslide susceptibility assessment, in particular to an extreme rainfall group landslide hierarchical susceptibility assessment method fusing diffusion generation and explainable learning. BACKGROUND
[0002] In recent years, high spatiotemporal concentration of extreme rainfall processes has significantly increased the triggering frequency of shallow group landslides in the southeast region. In particular, in areas with broken geological structure, intense topography, high rainfall intensity and short duration, landslide disasters have become the main disaster factors affecting traffic safety, residential stability and major engineering operation. Under the driving of heavy rainfall, such shallow landslides will undergo rapid shear softening and produce sliding and flowing integration along the slope, showing typical "flowing and sliding" catastrophic characteristics, further exacerbating the damage range and disaster effect. Existing landslide susceptibility assessment methods are mostly based on topography, geology, ecology, rainfall and engineering disturbance to construct statistical or empirical models, but still have the following technical bottlenecks: first, there is a significant nonlinear coupling relationship between multiple factors, and traditional methods are difficult to fully characterize the triggering mechanism; second, under the extreme rainfall scenario, landslide samples are scarce and unevenly distributed, resulting in insufficient model generalization ability and significantly reduced prediction accuracy.
[0003] In recent years, explainable machine learning methods have been introduced into landslide susceptibility research to improve prediction ability and reveal feature effects, but existing methods mostly rely on existing samples for fitting and do not address the "lack of key samples" and "fuzzy class boundary" problems. At the same time, existing generative models focus on general data augmentation and lack a constraint mechanism for landslide scenarios, making it impossible to generate high-credibility supplementary samples for extreme landslides or high-uncertainty areas, and still posing a risk of error accumulation and overestimation or underestimation of risk.
[0004] Based on the above deficiencies, there is an urgent need for a group landslide susceptibility assessment technology that can simultaneously realize multi-factor coupling cognition, key sample enhancement and high-precision explainability, to support regional early warning, risk classification and disaster prevention decision-making. SUMMARY
[0005] The present application proposes a hierarchical landslide susceptibility intelligent assessment method that fuses diffusion generative probabilistic modeling and explainable learning mechanism to address the problems of sample scarcity, boundary sample discrimination difficulty, multi-factor coupling relationship difficulty in explanation and insufficient spatial classification result usability in landslide disaster prediction under extreme rainfall conditions, and realizes high-precision prediction and hierarchical expression of landslide occurrence probability in the target area.
[0006] To achieve the above purpose, the present application provides the following technical solutions:
[0007] A method for extreme rainfall group landslide susceptibility assessment by fusing diffusion generation and interpretable learning, comprising the following steps:
[0008] S1: multi-source factor data construction and rasterization preprocessing; for multi-source data of terrain, geology, hydrology, vegetation, engineering disturbance and rainfall structure indexes, spatial projection unification, resolution resampling and numerical standardization processing are performed to form a multi-factor raster input matrix for model training.
[0009] S2: Construct a feature response analysis module for identifying the main control factor and eliminating redundant variables.
[0010] S3: Construct a generative sample generation module to expand the distribution range of training samples and enhance the boundary sample identification capability.
[0011] S4: Establish an integrated discriminant module to output rasterized probability results and grade expression.
[0012] S5: Output the layered landslide susceptibility product, and express the evaluation results in the form of risk level mapping for early warning linkage, risk control and engineering governance scheme.
[0013] Compared with the prior art, the present application has the following beneficial effects:
[0014] 1. The present application introduces a diffusion generative sample generation mechanism, which effectively makes up for the problem of insufficient training samples under extreme rainfall conditions by performing constraint generation on boundary samples and scarce class samples, significantly reduces the bias risk of the classifier in the high imbalance sample space, and improves the stability and reliability of landslide occurrence probability prediction under complex working conditions.
[0015] 2. The present application adopts an interpretable learning mechanism to build a factor contribution response system, which can realize the causal correlation analysis between input factors and output results without damaging the model discrimination performance, so that the landslide susceptibility result has engineering traceability.
[0016] 3. The present application constructs a layered landslide susceptibility grade output system, which converts continuous probability expression into executable risk classification results, and can be directly connected with spatial risk zoning, dynamic early warning scheduling and prevention and control resource allocation, realizes seamless transformation from model output to business execution, and improves the applicability of evaluation results in emergency management and traffic infrastructure disaster prevention.
[0017] 4. The present application overcomes the defects of traditional experience weight method such as strong subjectivity, pure data driven model uninterpretable and single output form difficult to land, and has engineering transformation and popularization value. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1A schematic diagram of a method according to an embodiment of the present application is shown in FIG. 1.
[0019] Figure 2 A schematic diagram of multi-source landslide hazard factor data construction according to an embodiment of the present application is shown in FIG. 2.
[0020] Figure 3 A schematic diagram of a generative sample supplement module according to an embodiment of the present application is shown in FIG. 3.
[0021] Figure 4 A schematic diagram of an integrated discriminant and grade classification module according to an embodiment of the present application is shown in FIG. 4.
[0022] Figure 5 A schematic diagram of a double-layer integrated strategy composite discriminant framework according to an embodiment of the present application is shown in FIG. 5.
[0023] Figure 6 A schematic diagram of landslide susceptibility grading results according to an embodiment of the present application is shown in FIG. 6. DETAILED DESCRIPTION
[0024] The technical solutions provided by the present application will be further described below in conjunction with specific embodiments and the accompanying drawings. The advantages and features of the present application will become clearer in conjunction with the following description.
[0025] As shown in FIG. 1, the present application proposes an extreme rainfall group landslide susceptibility evaluation method combining diffusion generation and interpretable learning, including five steps: Figure 1 S1: Multi-source factor data construction and rasterization preprocessing; For topography, geology, hydrology, vegetation, engineering disturbance and rainfall structure indicators and other multi-source data, spatial projection unification, resolution resampling, numerical standardization processing are performed to form a multi-factor raster input matrix for model training.
[0026] S2: Construct a feature response analysis module for identifying main control factors and removing redundant variables.
[0027] S3: Construct a generative sample supplement module for expanding the distribution range of training samples and enhancing the identification ability of boundary samples.
[0028] S4: Establish an integrated discriminant module for outputting rasterized probability results and grade expression.
[0029] S5: Output hierarchical landslide susceptibility product, express the evaluation results in risk grade mapping, for early warning linkage, risk control and engineering governance scheme formulation.
[0030] Through the combination of multi-source spatial data fusion, generative model sample supplement and interpretable machine learning method, the whole process calculation of landslide susceptibility from probability prediction to grade layering is realized.
[0031] EMBODIMENT
[0032] EMBODIMENT
[0033] As an example, the input data source of the present application comes from public remote sensing, surveying and mapping basic data and natural resource department measured geological data: the DEM data source is the PALSAR DEM (12.5m) product of ALOS satellite (used to calculate slope, slope direction, curvature, etc.); landslide information, lithology map, fault distribution are taken from regional geological map, geological survey database; NDVI is taken from Sentinel satellite remote sensing vegetation product; road vector is taken from public map vector data; rainfall data is taken from local rainfall station measured rainfall; the rest of the hydrological derived indicators (such as TWI) are calculated from the foregoing data.
[0034] S1: multi-source factor data construction and preprocessing:
[0035] As shown in Figure 2 , a multi-source factor data construction and preprocessing module is constructed for forming a multi-dimensional input factor matrix required for landslide susceptibility analysis. The input data includes topographic factors, geological factors, hydrological factors, ecological factors, rock-soil factors and human engineering disturbance factors. Among them, the topographic factors are calculated from digital elevation model (DEM), including slope, slope direction, curvature and terrain relief degree; the geological factors include lithology category and fault distance; the hydrological factors include terrain wetness index (TWI); the ecological factor adopts normalized vegetation index (NDVI); the rock-soil factor is a softening coefficient reflecting the strength and permeability characteristics of rock-soil body, which is obtained from exploration test data and then uniformly rasterized after being valued according to lithology zoning; the human engineering disturbance factor includes distance from road.
[0036] To ensure the consistency of spatial data, the present application adopts unified coordinate system and unified resolution (12.5 meters) for spatial projection transformation, and realizes data registration through resampling, interpolation and filtering operation. After standardization processing, all input data form a unified dimension feature matrix , wherein is the number of sample units, is the number of input factors, and the feature vector corresponding to the th sample unit is denoted as . The sample unit is a grid pixel under the unified coordinate system and 12.5m resolution. The feature matrix serves as the input basis for subsequent feature analysis and model training.
[0037] Through the processing of this module, it is ensured that the data of different sources are consistent in spatial scale, coordinate system and numerical domain, so as to avoid modeling deviation caused by data heterogeneity, and to provide accurate input for subsequent generative sampling and probability discrimination.
[0038] S2: feature response analysis module:
[0039] The characteristic response analysis module is used for identifying the main controlling factor of landslide occurrence and eliminating redundant variables, so as to realize dimension reduction and optimization of input factors.
[0040] The module processing process includes the following four sub-steps:
[0041] S2.1: Based on the information gain ratio (Information Gain Ratio, IGR), the discriminant contribution of each factor to the target variable (landslide / non-landslide) is calculated, and the factor importance ranking sequence is obtained;
[0042] The definition and calculation method of information gain ratio IGR is prior art. By suppressing the preference for factors with more values, the contribution of different factors is comparable. For continuous factors, first discretize by binning to obtain a finite value interval, then calculate the IGR value according to the standard definition and sort.
[0043] S2.2: Use an interpretable learning mechanism to quantify the causal response of the model output, and use a feature contribution explanation vector to measure the marginal influence of input factor changes on the prediction result ;
[0044] Wherein , represents the contribution intensity of the th factor to the prediction result of the sample, the symbol represents the action direction, and the absolute value represents the influence degree.
[0045] S2.3: Set the contribution threshold , according to the joint result of information gain ratio and , eliminate low-sensitive factors, and only keep the main control variables; the joint result refers to the reference of factor information gain ratio sorting and global contribution degree sorting: when the factor is lower than the same contribution threshold in the two types of sorting, it is determined to be low-sensitive or redundant, and is not included in the subsequent modeling input, otherwise it is retained.
[0046] The retained factors are normalized and scaled to form an optimized input set.
[0047] S2.4: The obtained optimized input set is used as the final input factor set, which provides a reduced and efficient input variable set for the S3 generation formula sample module as the boundary sample and the directional identification basis and generation condition variable of the scarce category sample. At the same time, it accepts the sample consistency feedback signal returned by S3, and according to this, the contribution threshold and the retained factor set are iteratively modified until the feedback index meets the preset consistency requirement.
[0048] After the above steps of the module are processed, the input dimension can be significantly reduced, the model calculation efficiency can be improved, and the interpretability and traceability of the subsequent model can be enhanced.
[0049] S3: Generating a sample module:
[0050] The generating sample module is constructed to expand the training sample distribution range and enhance the boundary sample recognition capability on the basis of the sample data optimized in S2.
[0051] The generating sample module includes a denoising diffusion probability model, a positive sample generator, and a negative sample generator, wherein the denoising diffusion probability model is used to perform a forward perturbation and a reverse denoising generation process. The positive sample generator is used to generate samples in a specific direction for landslide samples, and the negative sample generator is used to generate samples in a specific direction for non-landslide samples. Both of them take the factor set output by S2 as input. The module outputs a generated sample set for subsequent integrated discriminant module training.
[0052] S3.1: First, extract boundary samples and rare category samples according to sample distribution characteristics to form a sample set to be enhanced The method for extracting boundary samples and rare category samples is a prior art.
[0053] S3.2: Introduce a denoising diffusion probability model to perform a forward perturbation and a reverse generation process. The forward process forms a noise sequence by adding Gaussian noise to the original sample The reverse process uses a neural network model to gradually denoise the sample to approximate the true distribution.
[0054] S3.3: Set geological consistency constraints , and perform abnormal elimination and attribute consistency verification on the generated samples. The geological consistency constraints include the upper and lower limits of the value range of each factor of the generated sample, the legality of the geological category code, and the rule verification of the physical relationship between key factors, to eliminate unreasonable samples and ensure the physical reasonableness and geological interpretability of the generated samples.
[0055] If the generated sample deviates systematically in the consistency verification, output a consistency evaluation index as a sample consistency feedback signal, which is fed back to S2 to trigger adaptive adjustment of the reserved factor set and the threshold τ, and re-performs sample generation on the updated factor set. The adaptive adjustment refers to adjusting the threshold to change the number of reserved factors, and synchronously updating the reserved factor set, including increasing to eliminate more redundant factors, or decreasing to introduce more constraint factors.
[0056] S3.4: The generated samples are proportionally divided into a positive sample set and a negative sample set Mixing with original samples to form enhanced sample library The training objective is to minimize the Wasserstein distance between generated samples and real samples:
[0057]
[0058] where, represents the distribution of real samples, represents the distribution of generated samples; is the set of all feasible joint distributions; is a pair of samples sampled from the real sample distribution and the generated sample distribution, respectively; is the distance metric between samples, used to characterize the difference between the two distributions.
[0059] The generative sample augmentation module trains the positive sample generator and the negative sample generator independently, and performs generation sampling at multiple candidate time steps to obtain the corresponding generated sample set. For each , the Wasserstein distance between the generated sample distribution and the real sample distribution is calculated, and the time step that makes reach the minimum and tends to be stable within adjacent time steps is selected as the optimal generation time When the Wasserstein distance reaches the stable minimum value at the same time, the generation result is optimal.
[0060] The coverage of the generated sample space is significantly improved, providing balanced and complete data support for the subsequent integrated discriminant model.
[0061] As shown in Figure 3 , the distribution difference test results of real samples and generated samples in the sample generation process are given, where the horizontal axis 0-9 is the time step number of the diffusion model, and the bar proportion respectively represents the comparison of real sample distance and generated sample distance, which is used to evaluate the approximation degree of generated samples to real distribution at different time steps. The optimal generation time is the serial number corresponding to the time step that makes the distribution difference between real samples and generated samples minimum in the time step serial number 0-9, which corresponds to Figure 3 i.e. time step 7.
[0062] S4: Integrated discriminant module:
[0063] As shown in Figure 4 and Figure 5As shown, the integrated discrimination module of the present application is used to realize quantitative prediction of landslide occurrence probability. The model core is composed of a multi-algorithm fusion system, including two processing branches of random forest sub-model and XGBoost sub-model, and combining Bagging / Boosting double-layer integration strategy to form a composite discrimination framework.
[0064] The specific implementation is as follows:
[0065] S4.1: Read the training sample feature matrix output by S3 and the label vector , and construct fold cross-validation division. The landslide and non-landslide labels of the aforementioned sample units are obtained: if the sample unit is a landslide sample, it takes 1, otherwise it takes 0.
[0066] Output: fold data subset , is the number of cross-validation folds; the th fold represents the training subset, represents the validation subset.
[0067] S4.2: Random Forest Sub-model Training (Bagging Branch): Train the random forest sub-model based on the Bagging mechanism to obtain the probability output of the sample .
[0068] Output: Random Forest Probability Output .
[0069] S4.3: XGBoost Sub-model Training (Boosting Branch): Train the XGBoost sub-model with cross-entropy as the objective function:
[0070]
[0071] wherein, is the number of samples, is the label of the th sample, is the probability output of the XGBoost predicting the th sample as a landslide, is the internal parameter set of the XGBoost model.
[0072] Introduce a complexity penalty term to control the model size:
[0073]
[0074] wherein, The number of leaf nodes in the tree. The leaf node weight vector, for Regularization coefficient, This is the penalty coefficient for the number of leaf nodes, used to control model complexity.
[0075] The final optimization goal is:
[0076]
[0077] Gradient boosting iterative updates (calculation of optimal weights for leaf nodes) are employed:
[0078]
[0079] in, Indicates the first During the first round of gradient boosting, the... The output of each tree for a sample, i.e. the optimal leaf weight when the sample falls into the corresponding leaf node, is used as the incremental update of the model logit for that round. and The loss function is respectively paired with the first... The first and second gradients of the predicted values for each sample. for The regularization coefficient is summed over the set of samples contained in the current leaf node.
[0080] Output: XGBoost's logit output and probability output .
[0081] S4.4: Perform a two-level fusion of the outputs from the Bagging and Boosting branches to obtain the comprehensive logit:
[0082]
[0083] Output the probability of a landslide occurring:
[0084]
[0085] Output: Overall probability .
[0086] S4.5: Parameter Optimization and Model Determination. Using the cross-entropy loss output from S4.3. Using the average value under K-fold cross-validation as the objective function, Bayesian optimization is employed in the parameter space. Inner iterative search, To improve the learning rate of XGBoost iterations, leaf node weights regularization coefficient, penalty coefficient for the number of leaf nodes, maximum depth of a single tree, number of trees generated in each iteration of boosting. A set of candidate parameters is generated , that is, call S4.4 to complete model training and verification and obtain the corresponding loss and evaluation index, until the optimal parameters are obtained , and the final integrated discriminant model is obtained according to the solidification .
[0087] S4.6: Introduce SHAP mechanism to calculate the feature contribution degree vector of the sample under the final integrated discriminant model , quantifying the marginal influence of each factor on .
[0088] Output: single-sample contribution vector and factor rasterization contribution degree layer.
[0089] The final output of the module: continuous landslide probability layer of each grid unit and feature contribution degree result . Among them as the input of the S5 hierarchical susceptibility level output module, used for susceptibility level zoning.
[0090] Specifically, the feature contribution degree vector is calculated as follows: based on the Shapley value idea of game theory, the model output is decomposed into a reference value and the sum of the contributions of each feature. For a single sample , the change in model output caused by the addition or absence of each feature under different feature combinations is calculated, and the weighted expectation of all combination contributions is calculated to obtain the contribution value of each factor .
[0091] S5: Hierarchical susceptibility level output:
[0092] As shown in Figure 4 and Figure 6 , the hierarchical susceptibility level output module of the present application is used to convert the continuous probability result output by the final integrated discriminant model in S4 into discrete level distribution, realizing spatial hierarchical expression.
[0093] The natural breakpoint method is used to calculate the optimal classification threshold, and the core objective function is:
[0094]
[0095] Among them, represents the The landslide occurrence probability value corresponding to the The landslide occurrence probability value corresponding to the The landslide occurrence probability value corresponding to the The landslide occurrence probability value corresponding to the
[0096] The optimal classification threshold set is obtained by double-objective optimization of minimizing intra-class variance and maximizing inter-class variance . The natural breakpoint method is used to obtain the first The probability interval is divided into several risk levels.
[0097] In this embodiment, = 5, the probability interval is divided into five risk levels: extremely low ( ), low ( ), medium ( ), high ( ) and extremely high ( ).
[0098] To ensure the continuity of the spatial results, the present application further introduces a smoothing and spatial consistency checking mechanism to perform weighted average correction on the neighborhood probability difference:
[0099]
[0100] Among them, is the neighborhood set of the sample , is the spatial weight, is the neighborhood probability, is the normalization factor, is the smoothed probability. The spatial weight
[0101] is a distance-decaying function set according to the spatial distance, and is taken as , wherein is the distance between the sample and the neighborhood unit , and is the attenuation coefficient. The final layered landslide susceptibility level map is generated (as shown in
[0102] ), in which different risk levels are represented by five color gradients, from green (extremely low) to red (extremely high). The results can be directly imported into the geographic information system platform to realize spatial visualization display and risk zoning management. Figure 6
[0103] This example combines generative data augmentation with interpretable ensemble learning to improve sample utilization efficiency in the context of extreme rainfall-induced landslides. The output five-level landslide susceptibility ranking map provides quantitative support for regional disaster prevention and mitigation, achieving an integrated technical path from "data fusion-generative modeling-risk classification-outcome application", and has engineering promotion value and scientific guiding significance.
[0104] The above description is only a description of the preferred embodiments of the present application, and is not any limitation on the scope of the present application. Any modification or modification made by any person skilled in the art according to the above disclosed technical content shall be regarded as an equivalent effective embodiment, and shall fall within the scope of the technical scheme protected by the present application.
Claims
1. A method for extreme rainfall cluster landslide susceptibility assessment generated by fusion diffusion and interpretable learning, characterized in that, The method comprises the following steps: S1: multi-source factor data construction and grid preprocessing; for multi-source data of terrain, geology, hydrology, vegetation, engineering disturbance and rainfall structure index, spatial projection unification, resolution resampling, numerical standardization processing is performed to form a multi-factor grid input matrix for model training; S2: a feature response analysis module is constructed for identifying the main control factor and eliminating redundant variables; S3: a generative sample generation module is constructed for expanding the distribution range of training samples and enhancing the identification ability of boundary samples; S4: an integrated discriminant module is established for outputting grid probability results and grade expression; S5: output the layered landslide susceptibility product, the evaluation results are expressed by risk level mapping, which is used for early warning linkage, risk control and engineering management scheme; In step S2, the feature response analysis module is used to identify the main control factor of landslide occurrence and eliminate redundant variables, realizing dimension reduction and optimization of input factors; The processing process comprises the following steps: S2.1: calculate the discriminant contribution of each factor to the target variable based on information gain rate IGR to obtain a factor importance sorting sequence; By information gain suppression, the preference of factors with more values is suppressed, so that the contribution degrees of different factors are comparable; for continuous factors, first discretize by binning to obtain a limited value interval, then calculate the IGR value according to the standard definition and sort; S2.2: Quantify the causal response of the model output using an interpretable learning mechanism, employing feature contribution explanation vectors to measure the marginal impact of input factor changes on the predicted outcome ; wherein , represents the contribution intensity of the first factor to the prediction result of the sample, the symbol represents the action direction, and the absolute value represents the influence degree; S2.3: Set the contribution threshold , according to the joint results of information gain rate and , the low sensitive factors are removed, and only the main control variables are reserved; the joint results refer to the reference of factor information gain rate ranking and global contribution ranking: when the factor is lower than the same contribution threshold in the two types of ranking, it is determined as low sensitive or redundant, and is not included in the subsequent modeling input, otherwise it is reserved; Normalization and scale unification processing is performed on the reserved factors to form an optimized input set; S2.4: The obtained optimized input set is taken as the final input factor set, and the reduced efficient input variable set is provided for the formula supplement module of S3, as the directional identification basis and generation condition variable of boundary samples and scarce category samples; at the same time, the sample consistency feedback signal returned by S3 is accepted, and the contribution threshold and the reserved factor set are iteratively modified until the feedback index meets the preset consistency requirement.
2. The evaluation method according to claim 1, characterized in that Step S1 is specifically, A multi-source factor data construction and preprocessing module is constructed to form a multi-dimensional input factor matrix required for landslide susceptibility analysis; the input data includes terrain factors, geological factors, hydrological factors, ecological factors, rock-soil factors and human engineering disturbance factors; wherein the terrain factor is obtained by calculating the digital elevation model, including slope, slope direction, curvature and terrain relief degree; the geological factor includes lithology category and fault distance; the hydrological factor includes terrain humidity index; the ecological factor adopts normalized vegetation index; the rock-soil factor is a softening coefficient reflecting the strength and permeability of rock-soil mass, which is obtained by surveying test data, then assigned according to lithology zoning and unified raster processing; the human engineering disturbance factor includes distance from road; To ensure the consistency of spatial data, the unified coordinate system and the unified resolution are used for spatial projection transformation, and data registration is realized through resampling, interpolation and filtering operation; all input data are standardized to form a unified dimension feature matrix wherein is the number of sample units, is the number of input factors, and the feature vector corresponding to the th sample unit is denoted as ; the sample unit is a grid cell under the unified coordinate system and 12.5m resolution.
3. The evaluation method according to claim 2, characterized in that In step S3, The generative sample generation module comprises a denoising diffusion probability model, a positive sample generator and a negative sample generator, wherein the denoising diffusion probability model is used to perform forward disturbance and reverse denoising generation process; the positive sample generator generates positive samples in a specific direction, and the negative sample generator generates negative samples in a specific direction, both of which take the factor set output by S2 as input; the module output is a generated sample set for subsequent integrated discriminant module training; The processing process is as follows: S3.1: First, according to the sample distribution characteristics, the boundary samples and the rare category samples are extracted to form the to-be-enhanced sample set ; S3.2: Introducing a denoising diffusion probabilistic model to perform the forward perturbation and the backward generation process; the forward process forms a noisy sequence by adding Gaussian noise to the original sample , the backward process generates the sample step by step by denoising using a neural network model , so that it approximates the real distribution; S3.3: Set the geological consistency constraint condition The abnormal elimination and attribute consistency check are performed on the generated samples. The geological consistency constraint condition includes the upper and lower limits of the value range of each factor of the generated samples, the legality of the geological category code, and the rule check of the physical relationship between key factors, so as to eliminate unreasonable samples and ensure the physical reasonableness and geological interpretability of the generated samples. If the generated samples systematically deviate in the consistency check, output a consistency evaluation index as a sample consistency feedback signal back to S2 to trigger adaptive adjustment of the reserved factor set and threshold , and re-execute the sample completion generation on the updated factor set; the adaptive adjustment refers to adjusting the threshold to change the number of reserved factors, and synchronously updating the reserved factor set, including: increasing to eliminate more redundant factors, or decreasing to introduce more constraint factors; S3.4: Scale the generated samples and mix with the original samples to form the augmented sample pool ; the training objective is to minimize the Wasserstein distance between the generated samples and the real samples: wherein, represents a distribution of real samples, represents a distribution of generated samples; is a set of all possible joint distributions of both; is a pair of samples sampled from the real sample distribution and the generated sample distribution, respectively; is a distance measure between the samples, used to characterize the difference between the two distributions. The generative supplementation module is trained independently with positive and negative sample generators, and performs supplementation at multiple candidate time steps. The following steps generate samples, resulting in a corresponding set of generated samples; for each Calculate the generated sample distribution Compared with the true sample distribution Wasserstein distance between Select to make The optimal generation time is the time step that minimizes the time and whose change tends to stabilize within adjacent time steps. The optimal generation result is achieved when the Wasserstein distance simultaneously reaches a stable minimum.
4. The evaluation method according to claim 3, characterized in that In step S4, The integrated discriminant module is used to realize quantitative prediction of landslide occurrence probability, which is composed of a multi-algorithm fusion system, including a random forest sub-model and an XGBoost sub-model two processing branches, and a composite discriminant framework is formed by combining Bagging / Boosting double-layer integration strategy; Specifically, S4.1: read the training sample feature matrix output by S3 and label vector , and construct fold cross-validation division; The landslide and non-landslide labels of the sample unit are obtained: if the sample unit is a landslide sample, 1 is taken, otherwise 0 is taken. Output: Fold data subset , is the number of cross-validation folds; the fold denotes the training subset, denotes the validation subset; S4.2: Random Forest Submodel Training, i.e. Bagging Branch: Train a random forest submodel based on the Bagging mechanism to obtain the probability output for the sample ; Output: Random Forest probability output ; S4.3: XGBoost sub-model training, i.e., Boosting branch: Train XGBoos sub-model and take cross-entropy as the objective function: wherein, is the number of samples, is the label of the th sample, is the probability output of the th sample predicted by XGBoost as landslide, is the set of internal parameters of the XGBoost model; Introduce a complexity penalty term to control the model size: wherein, is the number of leaf nodes of the tree, is the leaf node weight vector, is the regularization coefficient, is the leaf node number penalty coefficient for controlling the model complexity; The final optimization objective is: And use gradient boosting iterative update to calculate the optimal weight of the leaf node: wherein, denotes the output of the th tree on the sample, i.e. the optimal leaf weight when the sample falls into the corresponding leaf node, is used as the incremental update of the logit of the model in this round, and are the first and second order gradients of the loss function with respect to the predicted value of the th sample, respectively, is a regularization coefficient, and the summation range is the sample set contained in the current leaf node. Output: XGBoost's logit output and probability output ; S4.4: Double-layer integration of the outputs of the Bagging branch and the Boosting branch to obtain the comprehensive logit: And output the landslide occurrence probability: Output: Combined probability ; S4.5: Parameter optimization and model determination; Cross-entropy loss output by S4.3 As the objective function, Bayesian optimization is adopted to iteratively search in the parameter space , where is the learning rate for XGBoost boosting iteration, is the regularization coefficient for leaf node weight, is the penalty coefficient for leaf node number, is the maximum depth of a single tree, is the number of trees generated for each boosting iteration; for each set of candidate parameters , i.e., S4.4 is called to complete model training and validation and obtain the corresponding loss, evaluation index, until the optimal parameters are obtained , and the final integrated discriminant model is obtained accordingly ; S4.6: In the final integrated discriminant model The SHAP mechanism is introduced below to calculate the feature contribution degree vector of the sample , quantifying the marginal influence of each factor on ; Output: single-sample contribution vector and each factor rasterized contribution map layer; Module final output: continuous probability layer of each grid unit of landslide , and feature contribution degree result ; wherein As the input of the S5 layering susceptibility grade output module, used for susceptibility grade zoning.
5. The evaluation method according to claim 4, characterized in that Feature contribution degree vector The calculation method is: decompose the model output into a reference value and the sum of the contributions of each feature, calculate the change in the model output caused by the addition or absence of each feature under different feature combinations, and calculate the weighted expectation of the contributions of all combinations to obtain the contribution value of each factor .
Citation Information
Patent Citations
Integrated landslide susceptibility evaluation method based on fusion model
CN120804908A
River and lake water bloom prediction method and system based on integrated diffusion learning model
CN121303474A