Landslide dam stability prediction method and system based on adaptive knowledge structure fusion
Patent Information
- Application Number
- CN202610949324.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]1.小样本与缺失并存时,模型性能对样本划分与数据质量高度敏感,容易出现看似准确但不稳定的结果;
[0037]1.本发明通过将IVAM经验公式以硬嵌入方式固定为网络基准分支,使模型从构建之初即具备明确的工程机理基准,神经网络仅需学习有限幅度的非线性偏差校正。在小样本训练条件下,本发明模型的回归精度和分类准确率高,且预测结果的方差小、训练稳定性高,有效抑制了传统机器学习在小样本场景下容易出现的过拟合与结果振荡问题。
Smart Images

Figure CN122797296A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of landslide dam stability prediction technology, and more specifically, to a method and system for landslide dam stability prediction based on adaptive knowledge structure fusion. Background Technology
[0002] Naturally formed landslide dams are characterized by their sudden onset, loose materials, heterogeneous structure, and rapid evolution. Once they overtower erosion or breach, they often trigger high-sediment-laden floods and catastrophic chain effects, posing a significant threat to downstream towns, transportation arteries, and critical infrastructure. Therefore, achieving rapid stability assessment and risk classification, as well as identifying high-risk dam types, within the short time window after the formation of a landslide dam, has direct engineering significance for guiding emergency response and engineering decision-making.
[0003] In recent years, machine learning methods have been increasingly used for rapid prediction of the stability of landslide dams due to their nonlinear fitting capabilities and fast inference advantages, achieving good results under certain data conditions. However, in the context of landslide dams, machine learning models still face the following prominent challenges:
[0004] 1. When small sample size and missing data coexist, model performance is highly sensitive to sample splitting and data quality, which can easily lead to seemingly accurate but unstable results.
[0005] 2. Most machine learning models are still black boxes, making it difficult to provide an auditable and transferable explanation of physical consistency, thus hindering their engineering application.
[0006] 3. The integration of physical constraints and data-driven approaches is still insufficient. In particular, how to use prior knowledge to stabilize training and suppress extrapolation bias when samples are insufficient remains a key issue.
[0007] Therefore, there is an urgent need for a method for predicting the stability of landslide dams that can deeply integrate empirical knowledge with data-driven approaches, and has good generalization ability and interpretability under small sample conditions. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for predicting the stability of landslide dams based on adaptive knowledge structure fusion. By embedding the IVAM empirical formula as an untrainable knowledge benchmark branch into the network output layer, and constructing an experience and data adaptive fusion module, a feature weighting mechanism guided by prior knowledge, and a multi-layer nonlinear correction unit, the invention achieves deep fusion of domain knowledge and measured data, significantly improving the adaptability, reliability, and interpretability of the model under small sample conditions.
[0009] To achieve the above objectives, the present invention adopts the following technical solution:
[0010] A method for predicting the stability of landslide dams based on adaptive knowledge structure fusion includes the following steps:
[0011] Step S1: Establish a landslide dam disaster database, select multiple samples from the database, and preprocess the parameters of each sample;
[0012] Step S2: Construct knowledge-driven features based on the IVAM empirical formula and build an untrainable IVAM baseline branch to output the empirical stability index for each landslide dam sample.
[0013] Step S3: Construct an adaptive knowledge structure fusion network, dynamically fusing the empirical stability index output by the IVAM baseline branch with the output of the data-driven residual correction branch to obtain the stability index of the model prediction.
[0014] Step S4: Design a multi-dimensional composite constraint training strategy, adopt a total loss function that includes knowledge constraint terms and data fitting terms, and train the parameters of the adaptive knowledge structure fusion network by adopting a phased collaborative optimization method.
[0015] Step S5: Use the trained adaptive knowledge structure fusion network to predict the stability indicators of the landslide dam to be evaluated, output the indicator values, and determine the stability of the landslide dam.
[0016] Furthermore, step S1 includes the following steps:
[0017] Step S11: Establish a database of landslide dam disasters;
[0018] Step S12: Select multiple samples from the database that have complete parameters for dam height, dam width, dam volume, catchment area, and material coefficient.
[0019] Step S13: Perform logarithmic transformation on the four variables of dam height, dam width, dam volume and drainage area, and verify the normality of the transformed data.
[0020] Step S14: Pearson correlation analysis was used to confirm that there was no serious multicollinearity among the variables.
[0021] Furthermore, in step S2, the constructed knowledge-driven features include logarithmic dam height, logarithmic dam width, logarithmic dam volume, logarithmic catchment area, and material coefficient.
[0022] Furthermore, in step S3, dynamic fusion includes: weighting and summing the empirical stability index output by the IVAM baseline branch and the correction amount output by the residual correction branch, with the weighting coefficients automatically determined by the current landslide dam sample and the critical threshold for stability discrimination.
[0023] Furthermore, for samples far from the critical threshold for discrimination but within the applicable range of the empirical formula, the weighting coefficients tend to be smaller, thus reducing the residual correction amount; for samples close to the critical threshold for discrimination or exceeding the applicable range of the empirical formula, the weighting coefficients automatically increase, thus increasing the residual correction amount; thereby achieving adaptive fusion of the empirical model and the data-driven model.
[0024] Furthermore, in step S4, the total loss function includes: a data fitting loss term based on labeled samples, an empirical value consistency loss term based on unlabeled knowledge sampling points, a gradient consistency loss term, a trend ranking loss term, and a regularization loss term for the correction magnitude of the constraint residual correction branch. Each loss term is weighted and summed using preset weight coefficients.
[0025] Furthermore, in step S4, the phased collaborative optimization includes: adopting a preheating phased growth strategy, assigning a higher weight to the data fitting loss term in the early stage of training, and then gradually increasing the weights of the empirical value consistency loss term, gradient consistency loss term, and trend ranking loss term; using the AdamW optimizer for training, and synchronously extracting knowledge sampling points in each training batch to participate in loss calculation.
[0026] This invention provides a stability prediction system for landslide dams based on adaptive knowledge structure fusion, the system comprising:
[0027] The data acquisition and preprocessing module is used to establish a database of landslide dam disasters, screen samples, and preprocess the parameters of the samples.
[0028] The IVAM baseline module, connected to the data acquisition and preprocessing module, is used to calculate and output an untrainable empirical stability index as a baseline branch based on the parameters of the samples.
[0029] The adaptive residual module is connected to the data acquisition and preprocessing module and the IVAM baseline module respectively. It is used to generate sample-level gating coefficients and residual corrections. It dynamically fuses the empirical stability index output by the IVAM baseline branch with the residual corrections to output the stability index predicted by the model.
[0030] The multi-source training module is connected to the data acquisition and preprocessing module, the IVAM benchmark module, and the adaptive residual module, respectively. It is used to construct a total loss function that includes data fitting loss, empirical consistency loss, gradient consistency loss, trend ranking loss, and regularization loss. A phased collaborative optimization strategy is used to train the adaptive knowledge structure fusion network.
[0031] The discrimination output module, connected to the adaptive residual module, is used to compare the stability index value output by the model with the zero threshold and output the judgment result of the landslide dam.
[0032] Furthermore, in the IVAM benchmark module, the input dam height and dam width are used to calculate the height-to-width ratio, and then linearly combined with the dam volume, catchment area and material coefficient according to the fixed coefficients of the IVAM empirical formula to output the empirical stability index.
[0033] Furthermore, the multi-source training module includes:
[0034] The preheating scheduling unit is used to assign high weights to the data fitting loss term in the early stage of training, and then gradually increase the weights of the empirical value consistency loss term, gradient consistency loss term, and trend ranking loss term.
[0035] The regularization unit is used to calculate the norm of the residual correction as a regularization loss term to limit the correction magnitude.
[0036] The beneficial effects of this invention are:
[0037] 1. This invention fixes the IVAM empirical formula as the network baseline branch through hard embedding, enabling the model to have a clear engineering mechanism baseline from the beginning of its construction. The neural network only needs to learn a limited amplitude of nonlinear bias correction. Under small sample training conditions, the model of this invention has high regression accuracy and classification accuracy, and the variance of the prediction results is small and the training stability is high, effectively suppressing the overfitting and result oscillation problems that are prone to occur in traditional machine learning in small sample scenarios.
[0038] 2. This invention constructs a feature importance weighting module based on mechanism and an adaptive fusion correction module combining experience and data. On the one hand, prior feature weights are constructed based on the analytical coefficients of the IVAM formula, guiding the model to prioritize variables with a more significant impact on stability. On the other hand, sample-level gating coefficients automatically determine the distance between the current landslide dam and the critical stability threshold, as well as the material type, dynamically adjusting the fusion weights of the empirical benchmark and data correction. For samples far from the discrimination threshold and where the empirical formula is reliable, the residual correction magnitude is automatically reduced to avoid overfitting; for samples close to the critical state or where the empirical formula has poor applicability, the correction magnitude is automatically increased, enabling the model to adapt to different engineering scenarios and achieving organic synergy between mechanism knowledge and data-driven approaches.
[0039] 3. This invention simultaneously introduces numerical consistency supervision, gradient consistency supervision, trend ranking supervision, and correction magnitude regularization during the training process, constructing a multi-dimensional composite constraint system. Gradient consistency constraints ensure that the model output's sensitivity to input variables remains consistent with the IVAM analytical gradient; trend ranking supervision constrains the model output's response to variable changes to conform to engineering principles. Compared to soft constraint schemes that only superimpose empirical formula residuals into the loss function, this invention guarantees the physical consistency of the model from three levels: function value, local sensitivity, and response direction. This gives the prediction results auditable and transferable engineering interpretability, overcoming the shortcomings of traditional black-box models that are difficult to meet the needs of engineering promotion. Attached Figure Description
[0040] Figure 1 This is a flowchart of a landslide dam stability prediction method based on adaptive knowledge structure fusion in this embodiment;
[0041] Figure 2 This is a framework diagram of the landslide dam stability prediction system based on adaptive knowledge structure fusion in this embodiment;
[0042] Figure 3a This is a scatter plot of the true and predicted values for a pure data-driven model with a training set of 40 data points in this embodiment.
[0043] Figure 3b This is a scatter plot of the true and predicted values for a pure data-driven model with a training set of 60 data points in this embodiment.
[0044] Figure 3c This is a scatter plot of the true and predicted values for a pure data-driven model with a training set of 80 data points in this embodiment.
[0045] Figure 3d This is a scatter plot of the true and predicted values of the model of the present invention with a training set of 40 data points and a knowledge sampling ratio r of 1 in this embodiment.
[0046] Figure 3e This is a scatter plot of the true and predicted values of the model of the present invention with a training set of 60 data points and a knowledge sampling ratio r of 1 in this embodiment.
[0047] Figure 3f This is a scatter plot of the true and predicted values of the model of the present invention with a training set of 80 data points and a knowledge sampling ratio r of 1 in this embodiment.
[0048] Figure 3g This is a scatter plot of the true and predicted values of the model of the present invention with a training set of 40 data points and a knowledge sampling ratio of r of 2 in this embodiment.
[0049] Figure 3hThis is a scatter plot of the true and predicted values of the model of the present invention with a training set of 60 data points and a knowledge sampling ratio of r of 2 in this embodiment.
[0050] Figure 3i This is a scatter plot of the true and predicted values of the model of the present invention with a training set of 80 data points and a knowledge sampling ratio of r of 2 in this embodiment.
[0051] Figure 3j This is a scatter plot of the true and predicted values of the model of the present invention with a training set of 40 data points and a knowledge sampling ratio of r of 3 in this embodiment.
[0052] Figure 3k This is a scatter plot of the true and predicted values of the model of the present invention with a training set of 60 data points and a knowledge sampling ratio of r of 3 in this embodiment.
[0053] Figure 3l This is a scatter plot of the true and predicted values of the model of the present invention with a training set of 80 data points and a knowledge sampling ratio r of 3 in this embodiment.
[0054] Figure labels: 1. Data acquisition and preprocessing module; 2. IVAM benchmark module; 3. Adaptive residual module; 4. Multi-source training module; 5. Discriminant output module. Detailed Implementation
[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0056] Example: A method for predicting the stability of landslide dams based on adaptive knowledge structure fusion. This method first constructs a landslide dam disaster database, performing statistical analysis and logarithmic transformation preprocessing on key features such as dam height, dam width, dam volume, catchment area, and material coefficients. Then, a pure data-driven baseline model is established. Next, an adaptive knowledge structure fusion network is proposed, embedding the IVAM empirical formula as an untrainable knowledge benchmark branch into the network output layer. The adaptive residual module implements functions such as mechanism-based feature importance weighting and adaptive fusion correction of experience and data. Simultaneously, a multi-dimensional composite constraint training strategy is constructed through a multi-source training module. Finally, the trained model is used to predict the stability of the landslide dam, outputting a stability index and determining whether it is stable or unstable.
[0057] Specifically, such as Figure 1 As shown, it includes the following steps:
[0058] Step S1: Establish a landslide dam disaster database, select multiple samples with the required parameters from the database, and preprocess the parameters of each sample.
[0059] This step aims to build a dataset for model training and validation, and to ensure data quality by eliminating variable scale differences through preprocessing.
[0060] Step S1 specifically includes the following steps:
[0061] Step S11: Collect historical landslide dam cases from around the world and establish a database containing 1,157 samples, covering information such as dam geometry parameters, hydrodynamic processes of the landslide dam lake, causes, material composition and stability status;
[0062] Step S12: Filter the dam height H, dam width W, dam volume V, and drainage area from the database. and material coefficient A small sample dataset consists of 99 landslide dam cases with all five parameters complete, including 48 stable landslide dams and 51 unstable landslide dams. Dam height H, dam width W, and dam volume V are used to characterize the morphological features of the landslide dams, and the drainage area is also included. Material coefficients are used to reflect the hydrodynamic conditions of landslide-dammed lakes. Used to describe the material composition characteristics of landslide dams;
[0063] Step S13 involves determining the dam height H, dam width W, dam volume V, and catchment area. Logarithmic transformation was performed on the four variables to reduce the numerical range and improve the skewed distribution. The normality of the logarithmically transformed data was verified by the Shapiro-Wilk test. The p-values of each variable after transformation were all greater than 0.05, which satisfies the normal distribution hypothesis.
[0064] Step S14: Use the Pearson product-moment correlation coefficient test to perform correlation analysis on each variable.
[0065] Among them, there is a strong positive correlation between the dam volume V and the dam width W, and between the dam volume V and the dam height H, reflecting the coupling characteristics between the volume and geometric scale of the landslide dam; at the same time, the dam width W and the drainage area... Dam volume V and drainage area There is a moderate positive correlation between dam height and drainage area. The correlation between them is relatively weak, indicating that the watershed area Its linear interpretability at the geometric scale is limited. In contrast, material coefficients... Related to dam height H, dam width W, dam volume V, and drainage area The correlations between the data are weak or very weak, indicating that material information and geometric scale information are highly complementary.
[0066] While a few combinations of variables showed strong correlations, these correlations were primarily concentrated between geometric scales and volumetric indices, failing to present a comprehensive high-correlation network structure, indicating a low risk of severe multicollinearity. Considering the significant randomness and nonlinearity of the formation and evolution of landslide dams, material coefficients... As a crucial factor reflecting the dam's structure and scour resistance, it can provide effective information for stability assessment that differs from that on a geometric scale. Therefore, in subsequent model construction, the dam height H, dam width W, dam volume V, and catchment area are retained. and material coefficient These five variables together serve as input features to improve the model's comprehensive ability to characterize the stability mechanism of landslide dams.
[0067] Step S2: Construct knowledge-driven features based on the IVAM empirical formula and build an untrainable IVAM baseline branch to output the empirical stability index for each landslide dam sample.
[0068] This step aims to embed the IVAM empirical formula into the network in a non-trainable form, providing a stable physical mechanistic benchmark for the model.
[0069] Step S2 specifically includes the following steps:
[0070] Step S21: Based on the IVAM empirical formula, select the dam height H, dam width W, dam volume V, and drainage area. and material coefficient As the original input variable;
[0071] Step S22, Construct knowledge-driven features:
[0072] First, construct the original input vector and the logarithmic knowledge feature vector, whose expressions are as follows:
[0073]
[0074]
[0075] In the formula, The height of the landslide dam is in meters. The width of the landslide dam is in meters. The volume of the landslide dam is... m 3 Enter the unit; The area of the landslide-dammed lake basin is expressed in square kilometers. is a material coefficient, dimensionless, with a value range of 0-1; It represents a logarithm with base 10.
[0076] Then, the above logarithmic knowledge features are standardized, and the calculation formula is as follows:
[0077]
[0078] In the formula, This is the mean vector of each feature; Let be the standard deviation vector of each feature.
[0079] Based on the material composition type of the landslide dam, the rock-dominated, rock-and-soil mixed, coarse-grained soil-dominated, and fine-grained soil-dominated types were mapped to value ranges of 0.70-1.00, 0.50-0.70, 0.25-0.50, and 0.00-0.25, respectively. Specific material coefficients were then determined using linear interpolation. The numerical value. The standardized features. As a knowledge-driven feature vector.
[0080] Step S23: Based on the IVAM method, construct a non-trainable IVAM baseline branch. This branch contains no trainable parameters, and its expression is:
[0081]
[0082] In the formula, To ensure consistency of units.
[0083] This branch contains no trainable parameters, ensuring that the model has a traceable empirical mechanistic baseline at any input. At the same time, the network no longer learns stability metrics from zero, but learns bias corrections with a limited magnitude on this baseline.
[0084] Step S24: Input the knowledge-driven feature vector into the IVAM baseline branch and output the empirical stability index for each landslide dam sample. This metric serves as the baseline solution for subsequent data-driven corrections.
[0085] Step S3: Construct an adaptive knowledge structure fusion network, dynamically fusing the empirical stability index output by the IVAM baseline branch with the output of the data-driven residual correction branch to obtain the stability index of the model prediction.
[0086] This step aims to learn local nonlinear deviations that cannot be covered by empirical formulas through data-driven residual correction branches, and to achieve dynamic adaptive fusion of empirical benchmarks and data corrections.
[0087] Step S3 specifically includes the following steps:
[0088] Step S31: To enable the model to accurately quantify the influence of different input factors on the stability index of the landslide dam, a mechanism-based feature importance prior vector is first constructed based on the analytical coefficients of the IVAM empirical formula. Then, this vector is dynamically fused with the feature weights obtained from the statistical analysis of measured data to jointly determine the final weight coefficients of each input factor. The expression for this final weight coefficient is as follows:
[0089]
[0090] In the formula, It is obtained by normalizing the absolute values of the IVAM analytical coefficients; Standardized knowledge-driven feature vectors; and These are the learnable parameters of the attention network; It uses a sigmoid activation function. Features after attention weighting. Inputting the subsequent network allows the model to prioritize variables that have a more significant impact on stability in its structure, while retaining the ability to adaptively adjust feature weights based on sample data.
[0091] Step S32: To improve the model's adaptability to different types and conditions of landslide dams, a dynamic correction mechanism is introduced, and the calculation results of each layer are updated according to the following formula:
[0092]
[0093]
[0094] In the formula, The channel gating coefficient is generated from the features of the current sample, and its value ranges from 0 to 1; For the first The hidden state of the layer; For the first The hidden state of the layer; , , and All are learnable parameters; For layer normalization operation; It is the hyperbolic tangent activation function; It uses an sigmoid activation function. This structure enables the network to use hidden channels and residual paths of different intensities for different landslide dam samples, which is equivalent to adaptively adjusting the network depth and feature channels according to the engineering conditions.
[0095] The final model output consists of the IVAM baseline value and the gated residual correction.
[0096]
[0097] In the formula, This is a stability metric for model predictions; Empirical stability metrics for the IVAM baseline branch output; For learnable but limited residual scaling parameters; The sample-level corrected gating weighting coefficient is defined as follows: ; Output for residual correction branch; These are the learnable parameters for the residual branch.
[0098] For samples far from the discrimination threshold but within the applicable range of the empirical formula, the weighting coefficients... The weighting coefficient tends to a smaller value, thus reducing the residual correction amount; for samples close to the critical threshold or exceeding the applicable range of the empirical formula, the weighting coefficient... The automatic increase in residual correction amount enables adaptive fusion of empirical and data-driven models.
[0099] Step S33: Dynamically add the empirical stability index output by the IVAM baseline branch to the gated residual correction result to obtain the final model-predicted stability index. .
[0100] Step S4: Design a multi-dimensional composite constraint training strategy, adopt a total loss function that includes knowledge constraint terms and data fitting terms, and train the parameters of the adaptive knowledge structure fusion network by adopting a phased collaborative optimization method.
[0101] This step aims to improve the generalization ability and physical interpretability under small sample conditions by training the network in a coordinated manner through multiple physical consistency constraints.
[0102] Step S4 specifically includes the following steps:
[0103] Step S41, construct the total loss function, which includes the following items:
[0104] The loss term, which is based on data fitting using labeled samples and supervised by mean squared error, is expressed as follows:
[0105]
[0106] In the formula, This represents the number of labeled samples. For the model to the first Predictive stability index for each sample; For the first The true stability label of each sample.
[0107] Based on the empirical value consistency loss term of unlabeled knowledge sampling points, the model output is constrained to not deviate from the IVAM benchmark overall, and its expression is:
[0108]
[0109] In the formula, This represents the number of unlabeled knowledge sampling points. For the first Input features of each sampling point For sampling point index; For the model to the first Predictive stability index for each sampling point; For the first IVAM baseline values for each sampling point.
[0110] The gradient consistency loss term constrains the local sensitivity of the model output to the input variables to maintain consistency with the analytical gradient of the IVAM formula. Its expression is:
[0111]
[0112] In the formula, To predict the output of the model based on standardized features The gradient; The output of the IVAM formula is as follows: The gradient; It is the weighted norm.
[0113] The trend ranking supervision term constructs trend sample pairs based on the direction of the coefficients in the IVAM formula, ensuring that they satisfy... > And it uses sorting constraints, the expression of which is:
[0114]
[0115] In the formula, The number of trend sample pairs; and For a pair of samples that satisfy the IVAM ordering relation; This is the preset interval parameter; and These are the prediction stability metrics of the model for positive and negative samples, respectively.
[0116] The supervision methods formed by the aforementioned loss terms do not directly require the model values to be exactly equal to empirical formulas. Instead, they constrain the model's response to changes in variable direction to conform to engineering principles, serving as a supplementary insight in addition to numerical residuals and gradient residuals. To prevent residual branches from excessively replacing empirical benchmarks, a correction magnitude regularization term is introduced, with the following expression:
[0117]
[0118] In the formula, For the first Individual sample-level corrected gated weighting coefficients; For the first Output of each residual correction branch.
[0119] The total loss function is the weighted sum of the above terms, and its expression is:
[0120]
[0121] In the formula, To follow the training rounds The changing weighting coefficients; , and Preset weights for each type of loss; is the regularization coefficient.
[0122] Step S42, adopt a preheating-style phased growth strategy, i.e., weighting coefficients. The values are small in the early stages of training, allowing the data fitting loss term to dominate, and then gradually increased. To gradually increase the weights of the empirical value consistency loss term, gradient consistency loss term, and trend ranking loss term, the model convergence is not unstable due to excessively strong mechanistic constraints in the early stages of training.
[0123] Step S43: Train the model using the AdamW optimizer, and simultaneously extract knowledge sampling points in each training batch to participate in the loss calculation. After training, the model can output a stability metric with physical consistency.
[0124] Through the above structural design, the model not only provides supervision using empirical formulas, but also explicitly inherits the IVAM mechanism in the network feedforward structure, and achieves local correction of complex engineering samples through adaptive gated residuals.
[0125] Step S5: Use the trained adaptive knowledge structure fusion network to predict the stability indicators of the landslide dam to be evaluated, output the indicator values, and determine the stability of the landslide dam.
[0126] This step aims to use a trained model to predict and classify the stability of the landslide dam to be evaluated.
[0127] Step S5 specifically includes the following steps:
[0128] Step S51: Calculate the dam height H, dam width W, dam volume V, and catchment area of the landslide dam to be evaluated. and material coefficient Input the trained adaptive knowledge structure fusion network to make the model output a stability index. ;
[0129] Step S52, make a judgment based on the preset discrimination threshold: when When the value is greater than 0, it is classified as a stable landslide dam. When the value is ≤0, it is determined to be an unstable landslide dam;
[0130] Step S53: Evaluate the model performance on the independent validation set, using absolute accuracy, conservative accuracy, and misclassification rate as evaluation metrics.
[0131] This embodiment also provides a landslide dam stability prediction system based on adaptive knowledge structure fusion, such as... Figure 2 As shown, the system includes a data acquisition and preprocessing module 1, an IVAM benchmark module 2, an adaptive residual module 3, a multi-source training module 4, and a discrimination output module 5.
[0132] The data acquisition and preprocessing module 1 is used to establish a landslide dam disaster database, screen samples, and perform logarithmic transformation on the parameters of the samples before standardization preprocessing.
[0133] Preferably, the data acquisition and preprocessing module 1 also includes a correlation analysis unit and a normality test unit.
[0134] The correlation analysis unit is used to perform Pearson correlation tests on the preprocessed variables to confirm that there is no serious multicollinearity.
[0135] The normality test unit is used to verify the normality of the logarithmically transformed data using the Shapiro-Wilk test.
[0136] IVAM baseline module 2 is connected to data acquisition and preprocessing module 1 and is used to directly calculate and output an untrainable empirical stability index as a baseline branch based on the parameters of the sample; this module does not contain any trainable parameters.
[0137] As a preferred option, in IVAM benchmark module 2, the input dam height and dam width are used to calculate the height-to-width ratio, and then linearly combined with the dam volume, catchment area and material coefficient according to the fixed coefficients of the IVAM empirical formula to output the empirical stability index.
[0138] The adaptive residual module 3 is connected to the data acquisition and preprocessing module 1 and the IVAM benchmark module 2 respectively. It is used for feature attention weighting and generates sample-level gating coefficients and residual corrections. It dynamically fuses the empirical stability index output by the IVAM benchmark branch with the residual corrections to output the stability index predicted by the model.
[0139] The multi-source training module 4 is connected to the data acquisition and preprocessing module 1, the IVAM benchmark module 2, and the adaptive residual module 3, respectively. It is used to construct a total loss function that includes data fitting loss, empirical consistency loss, gradient consistency loss, trend ranking loss, and regularization loss, and to train the adaptive knowledge structure fusion network using a phased collaborative optimization strategy.
[0140] As a preferred embodiment, the multi-source training module 4 includes a preheating scheduling unit and a regularization unit.
[0141] The preheating scheduling unit is used to assign a high weight to the data fitting loss term in the early stage of training, and then gradually increase the weight of the empirical value consistency loss term, gradient consistency loss term, and trend ranking loss term.
[0142] The regularization unit is used to calculate the norm of the residual correction as a regularization loss term to limit the correction magnitude.
[0143] The discrimination output module 5 is connected to the adaptive residual module 3 and is used to compare the stability index value output by the model with the zero threshold to output the judgment result of the landslide dam.
[0144] The method proposed in this embodiment will be verified below.
[0145] 1. Comparison of regression prediction performance:
[0146] From the established complete sample of 99 parameters, 40, 60, and 80 samples were selected as training sets, respectively, and the remainder were used as validation sets for comparative experiments between the pure data-driven model and the model of this invention. In the experiments, the knowledge sampling ratio r was set to 1, 2, and 3, respectively, and each parameter combination was trained 10 times. Statistics were then compiled. mean Standard deviation, mean RMSE, standard deviation RMSE, mean classification accuracy, and standard deviation of classification accuracy.
[0147] like Figures 3a-3l As shown, Figures 3a-3c The results are scatter plots of the true and predicted values for a pure data-driven model, with training sets containing 40, 60, and 80 data points respectively. Figures 3d-3f The above is a scatter plot of the true values versus predicted values of the model of this invention, with the knowledge sampling ratio r being 1 for all cases, and the training set data sizes being 40, 60, and 80 respectively. Figure 3g-Figure 3i The above is a scatter plot of the true values versus predicted values of the model of this invention, with the knowledge sampling ratio r being 2 for all cases, and the training set data sizes being 40, 60, and 80 respectively. Figures 3j-3l The above is a scatter plot of the true values versus predicted values of the model of this invention, with the knowledge sampling ratio r being 3 for all cases, and the training set data sizes being 40, 60, and 80 respectively.
[0148] Depend on Figures 3a-3l As shown, with the increase in sample size, the scatter points cluster more closely around the reference line, indicating that increasing the sample size can improve regression accuracy and reduce systematic error. Compared with a purely data-driven model, the model of this invention significantly reduces outliers in small sample scenarios. This is mainly because the IVAM baseline branch provides stable initial mechanistic values for the model, and the gated residual branch only learns a limited bias, avoiding unreasonable oscillations in ordinary neural networks when there are insufficient samples. As the knowledge sampling ratio r increases, the model gains more constraints in terms of numerical consistency, gradient consistency, and trend consistency, and the predicted point cloud converges further towards the reference line. When r continues to increase, the performance improvement tends to level off, indicating that there is a reasonable trade-off between the knowledge sampling ratio and computational cost.
[0149] Scatter plots are mainly used to display the distribution pattern of point clouds and the overall fitting trend, but it is difficult to quantitatively summarize the accuracy and stability of multiple settings. Therefore, under the same data partitioning and repeated training times, we further statistically analyzed the mean and standard deviation of the regression and classification indicators of each model under different sample sizes and knowledge sampling ratios, and summarized them into Table 1.
[0150] Table 1. Experimental Results of Model Performance Comparison
[0151]
[0152] Table 1 shows the comparison models as pure data-driven models. As can be seen from Table 1, under three different data scales, the model of this invention exhibits superior overall performance in both regression and classification compared to the pure data-driven model: R... 2 The overall mean is higher, the overall mean RMSE is lower, and the classification accuracy is significantly improved. This result indicates that by upgrading IVAM empirical knowledge from an "external loss constraint" to a "network structure baseline branch," and by adding adaptive gated residual correction and trend ranking supervision, the model can more effectively suppress overfitting, reduce unreasonable extrapolation, and improve discrimination stability under small sample conditions. Further comparison of different r values reveals that as the knowledge sampling ratio increases, the model performance generally shows a pattern of first improving and then stabilizing; when r is 2 or 3, the model achieves a relatively balanced performance in terms of accuracy, standard deviation, and computational cost, which can be considered a recommended setting for subsequent engineering applications.
[0153] 2. Comparison of classification prediction performance:
[0154] Based on the aforementioned database of 1157 landslide dam cases, further screening was conducted to identify domestic and international cases of collapsed and non-collapsed landslide dams with available measured data that had not previously been included in the modeling, thus constructing a model test set. The model of this invention was compared and validated with the BI method, DBI method, Is method, Ia method, Ls(AHWL) method, Ls(AHV) method, and the IVAM empirical formula. Absolute accuracy, conservative accuracy, and false positive rate were used for evaluation.
[0155] The absolute accuracy rate represents the proportion of samples whose calculated results perfectly match the actual stable state of the landslide dam; the conservative accuracy rate reflects the method's ability to make safer judgments, and its value can be characterized by the proportion of cases that are actually stable but are judged as unstable, together with the absolute accuracy rate; the misclassification rate refers to the proportion of landslide dams that are actually unstable but are incorrectly identified as stable. The calculation results for different methods are listed in Table 2.
[0156] Table 2 Comparison of Calculation Results of Stability Evaluation Methods for Landslide Dams
[0157]
[0158] Table 2 shows that different methods for evaluating the stability of landslide dams exhibit significant differences in their discrimination performance. Overall, the model of this invention performs best: only 2 cases were misclassified in 67 cases, with 65 accurate cases and 0 uncertain cases; its absolute accuracy and conservative accuracy both reached 97.01%, and its misclassification rate was the lowest at only 2.99%, demonstrating high discrimination reliability and stability. In contrast, traditional empirical discrimination indicators suffer from numerous uncertainties or low accuracy; while some methods have high conservative accuracy, they still exhibit a certain number of conservative judgments or misclassifications. In summary, this invention effectively reduces uncertainties and lowers the risk of misclassification through the collaborative work of IVAM hard-embedded baseline branch, adaptive gated residual correction, and multi-source knowledge supervision.
[0159] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for predicting the stability of landslide dams based on adaptive knowledge structure fusion, characterized in that, Includes the following steps: Step S1: Establish a landslide dam disaster database, select multiple samples from the database, and preprocess the parameters of each sample; Step S2: Construct knowledge-driven features based on the IVAM empirical formula and build an untrainable IVAM baseline branch to output the empirical stability index for each landslide dam sample. Step S3: Construct an adaptive knowledge structure fusion network, dynamically fusing the empirical stability index output by the IVAM baseline branch with the output of the data-driven residual correction branch to obtain the stability index of the model prediction. Step S4: Design a multi-dimensional composite constraint training strategy, adopt a total loss function that includes knowledge constraint terms and data fitting terms, and train the parameters of the adaptive knowledge structure fusion network by adopting a phased collaborative optimization method. Step S5: Use the trained adaptive knowledge structure fusion network to predict the stability indicators of the landslide dam to be evaluated, output the indicator values, and determine the stability of the landslide dam.
2. The method for predicting the stability of landslide dams based on adaptive knowledge structure fusion according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Establish a database of landslide dam disasters; Step S12: Select multiple samples from the database that have complete parameters for dam height, dam width, dam volume, catchment area, and material coefficient. Step S13: Perform logarithmic transformation on the four variables of dam height, dam width, dam volume and drainage area, and verify the normality of the transformed data. Step S14: Pearson correlation analysis was used to confirm that there was no serious multicollinearity among the variables.
3. The method for predicting the stability of landslide dams based on adaptive knowledge structure fusion according to claim 1, characterized in that, In step S2, the constructed knowledge-driven features include logarithmic dam height, logarithmic dam width, logarithmic dam volume, logarithmic catchment area, and material coefficient.
4. The method for predicting the stability of landslide dams based on adaptive knowledge structure fusion according to claim 1, characterized in that, In step S3, dynamic fusion includes: weighting and summing the empirical stability index output by the IVAM baseline branch and the correction amount output by the residual correction branch, with the weighting coefficients automatically determined by the current landslide dam sample and the critical threshold for stability discrimination.
5. The method for predicting the stability of landslide dams based on adaptive knowledge structure fusion according to claim 4, characterized in that, For samples that are far from the critical threshold for discrimination but within the applicable range of the empirical formula, the weighting coefficients tend to be smaller, thus reducing the residual correction amount; for samples that are close to the critical threshold for discrimination or exceed the applicable range of the empirical formula, the weighting coefficients automatically increase, thus increasing the residual correction amount. Achieve adaptive fusion of empirical models and data-driven models.
6. The method for predicting the stability of landslide dams based on adaptive knowledge structure fusion according to claim 1, characterized in that, In step S4, the total loss function includes: a data fitting loss term based on labeled samples, an empirical value consistency loss term based on unlabeled knowledge sampling points, a gradient consistency loss term, a trend ranking loss term, and a regularization loss term for the correction magnitude of the constraint residual correction branch. Each loss term is weighted and summed using preset weight coefficients.
7. The method for predicting the stability of landslide dams based on adaptive knowledge structure fusion according to claim 6, characterized in that, In step S4, the phased collaborative optimization includes: adopting a preheating phased growth strategy, assigning a higher weight to the data fitting loss term in the early stage of training, and then gradually increasing the weights of the empirical value consistency loss term, gradient consistency loss term, and trend ranking loss term; using the AdamW optimizer for training, and synchronously extracting knowledge sampling points in each training batch to participate in loss calculation.
8. A landslide dam stability prediction system based on adaptive knowledge structure fusion for implementing the method of claim 1, characterized in that, The system includes: The data acquisition and preprocessing module (1) is used to establish a landslide dam disaster database, screen samples, and preprocess the parameters of the samples; IVAM benchmark module (2), connected to data acquisition and preprocessing module (1), is used to calculate and output an untrainable empirical stability index as a benchmark branch based on the parameters of the sample. The adaptive residual module (3) is connected to the data acquisition and preprocessing module (1) and the IVAM benchmark module (2) respectively. It is used to generate sample-level gating coefficients and residual corrections, dynamically fuse the empirical stability index output by the IVAM benchmark branch with the residual corrections, and output the stability index predicted by the model. The multi-source training module (4) is connected to the data acquisition and preprocessing module (1), the IVAM benchmark module (2) and the adaptive residual module (3) respectively. It is used to construct a total loss function that includes data fitting loss term, empirical consistency loss term, gradient consistency loss term, trend ranking loss term and regularization loss term, and to train the adaptive knowledge structure fusion network using a phased collaborative optimization strategy. The discrimination output module (5) is connected to the adaptive residual module (3) and is used to compare the stability index value output by the model with the zero threshold to output the judgment result of the landslide dam.
9. The landslide dam stability prediction system based on adaptive knowledge structure fusion according to claim 8, characterized in that, In the IVAM benchmark module (2), the input dam height and dam width are used to calculate the height-to-width ratio, and then linearly combined with the dam volume, drainage area and material coefficient according to the fixed coefficients of the IVAM empirical formula to output the empirical stability index.
10. The landslide dam stability prediction system based on adaptive knowledge structure fusion according to claim 8, characterized in that, The multi-source training module (4) includes: The preheating scheduling unit is used to assign high weights to the data fitting loss term in the early stage of training, and then gradually increase the weights of the empirical value consistency loss term, gradient consistency loss term, and trend ranking loss term. The regularization unit is used to calculate the norm of the residual correction as a regularization loss term to limit the correction magnitude.