Stability judgment method and device of barrier dam, computer equipment and medium

By improving the Light GBM algorithm and combining Bayesian optimization algorithm, combining parameters are constructed to generate a richer feature set, which solves the problem of inaccurate stability prediction resulting in the missing data of the dam, and improves the accuracy of prediction.

CN120145210APending Publication Date: 2025-06-13ZHENGZHOU UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204184.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When the prior art deals with the evaluation of stability of the dam, the lack of data leads to inaccurate judgments and it is difficult to meet the data integrity requirements.

Method used

The Light GBM algorithm is improved by introducing the focus loss function FL, combined with the Bayesian optimization algorithm, and constructing combined parameters to generate a richer feature set and improving the prediction ability of the model.

Benefits of technology

In the absence of data, the accuracy of the stability prediction of the dam is improved, and a stability evaluation model of the dam under the absence of data is established, solving the problem of inaccurate stability prediction caused by data loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145210A_ABST
    Figure CN120145210A_ABST
Patent Text Reader

Abstract

The invention provides a stability judgment method and device for a barrier dam, computer equipment and a medium, and belongs to the technical field of barrier dam emergency processing, and the method comprises the steps: introducing a focus loss function (FL) to improve a Light GBM algorithm, and obtaining an FL-Light GBM algorithm; based on the obtained initial parameters related to the stability, combined parameters are constructed through linear combination, and training samples are constructed according to the initial parameters and the combined parameters; using the training sample to train an FL-Light GBM algorithm through a Bayesian optimization algorithm to obtain a stability prediction model; and obtaining initial parameters and combined parameters of the target weir dam, and inputting the initial parameters and the combined parameters into the stability prediction model to obtain an output stability judgment result of the target weir dam. Therefore, the accuracy of algorithm prediction can be improved through parameter combination on the basis of an improved algorithm under the condition of data missing, a barrier dam stability evaluation model under the condition of data missing is established, and the problem of inaccurate stability prediction caused by barrier dam data missing is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of emergency treatment of barrier dams, and particularly relates to a method, device, computer equipment and medium for judging the stability of barrier dams. Background Art

[0002] As a natural dam body, the formation cause of a barrier dam is complex, and the topographic and geomorphic features and the interaction between water and soil in the formation area are different from each other, resulting in a complex action mechanism of the influencing factors of the stability of the barrier dam. The stability of the barrier dam is not only affected by the geomorphic and geological conditions of the area where it is located, but also related to the dam body materials, structure and hydrodynamic conditions of the dam body. At the same time, due to the characteristics of easy breach of the barrier dam itself, it is difficult to collect the breach data of the dam, resulting in the lack of some characteristics in the barrier dam data set.

[0003] However, traditional evaluation methods cannot accurately judge the stability when key features are missing, and the research on the stability evaluation of barrier dams has higher requirements for data integrity. Therefore, in dealing with the problem of data loss commonly existing in the stability evaluation of barrier dams, there is an urgent need for a method that can accurately evaluate the stability of barrier dams under the condition of data loss. Summary of the Invention

[0004] In order to solve the problem that inaccurate judgment of the stability of the barrier dam is caused by data loss, the present invention provides a method, device, computer equipment and medium for judging the stability of the barrier dam.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] First, a method for judging the stability of a barrier dam is provided. The method includes:

[0007] Introduce the focal loss function FL to improve the Light GBM algorithm to obtain the FL-Light GBM algorithm;

[0008] Obtain the initial parameters of the barrier dam including dam body parameters and environmental parameters, and construct combined parameters directly related to the stability of the barrier dam by means of linear combination. According to the initial parameters, combined parameters and the corresponding stability states of the barrier dam, construct training samples;

[0009] Use the training samples to train the FL-Light GBM algorithm through the Bayesian optimization algorithm to obtain a stability prediction model;

[0010] Input the initial parameters and combined parameters of the target barrier dam obtained into the stability prediction model, and output the stability judgment result of the target barrier dam.

[0011] Optionally, the formula for defining the focal loss function FL is:

[0012]

[0013] Among them, p t has the following expression: p represents the probability value that the classification label y equals 1 during the model training process, α is the weight factor, and (1 - p t ) γ is the modulation factor for simple negative samples, and γ > 0 is an adjustable parameter.

[0014] Optionally, the initial parameters include the dam height, dam width, dam length, dam volume, reservoir capacity, basin area, material composition, and inducing factors; the combined parameters directly related to the stability of the barrier lake constructed by using the initial parameters in a linear combination manner include:

[0015] Width-to-height ratio, and its specific expression is: X 9 = W / H;

[0016] Dam shape coefficient, and its specific expression is:

[0017] Overtopping risk coefficient, and its specific expression is:

[0018] Lake shape coefficient, and its specific expression is:

[0019] Among them, W is the dam width, H is the dam height, V d is the dam volume, A is the basin area, and V l is the reservoir capacity.

[0020] Optionally, constructing training samples according to the initial parameters, combined parameters, and their corresponding barrier lake stability states includes:

[0021] Adding the combined parameters to the dataset of the initial parameters to obtain a new dataset, and randomly dividing the new dataset into a training set and a test set according to a preset ratio.

[0022] Optionally, the training of the FL-Light GBM algorithm by the Bayesian optimization algorithm includes:

[0023] Based on the Gaussian process of the Bayesian optimization algorithm, constructing a surrogate model of the objective function for any parameter combination of the FL-Light GBM algorithm by using the training samples;

[0024] Predicting multiple new parameter combinations and their corresponding objective functions and uncertainties through the surrogate model;

[0025] Construct a sampling function according to the prediction result. The sampling function is screened and optimized by the objective function and uncertainty of the new parameter combination to obtain the screened and optimized candidate parameter combinations.

[0026] Iteratively optimize multiple candidate parameter combinations to complete the optimization of the FL-Light GBM algorithm.

[0027] Secondly, a device for judging the stability of a barrier dam is provided. The device includes:

[0028] A construction module for improving the Ligh GBM algorithm by introducing a focal loss function FL to obtain the FL-Light GBM algorithm.

[0029] An acquisition module for acquiring the initial parameters of the barrier dam including dam body parameters and environmental parameters, constructing combined parameters directly related to the stability of the barrier dam by means of linear combination, and constructing training samples according to the initial parameters, combined parameters and their corresponding stability states of the barrier dam.

[0030] A training module for training the FL-Light GBM algorithm by means of the Bayesian optimization algorithm using the training samples to obtain a stability prediction model.

[0031] A judgment module for inputting the initial parameters and combined parameters of the target barrier dam obtained into the stability prediction model and outputting the stability judgment result of the target barrier dam.

[0032] In addition, a computer-readable storage medium is provided. The storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method for judging the stability of a barrier dam is implemented.

[0033] Finally, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned method for judging the stability of a barrier dam is implemented.

[0034] The method for judging the stability of a barrier dam provided by the present invention has the following beneficial effects:

[0035] First, the LightGBM algorithm is improved by introducing the focal loss function FL. By reducing the weights of easily classified samples, FL makes the model pay more attention to difficult-to-classify samples, improves the performance of the model when dealing with imbalanced datasets, and enhances the prediction accuracy of the model for the unstable state of the minority-class landslide dams. The combined parameters are constructed in a linear combination manner. On the one hand, it enriches the parameter data, and on the other hand, it explores the potential relationships between the initial parameters, strengthens the correlation with the stability of the landslide dams, generates a richer feature set in this way, helps the model capture more factors affecting the stability of the landslide dams, and improves the prediction ability of the model. Finally, it is optimized by the Bayesian algorithm, which can automatically search for the optimal combination of hyperparameters, thereby improving the prediction performance and generalization ability of the model. Through the above methods, in the case of missing data, based on the improved algorithm, the prediction accuracy of the algorithm can be improved through parameter combination, and thus a stability evaluation model of landslide dams under the condition of missing data is established, solving the problem of inaccurate stability prediction caused by missing landslide dam data. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the embodiments of the present invention and their design schemes, the accompanying drawings required for this embodiment will be briefly introduced below. The accompanying drawings in the following description are only partial embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0037] Figure 1 FIG. is a schematic diagram of the overall framework idea of a model provided by the present invention according to an exemplary embodiment.

[0038] Figure 2 FIG. is a schematic flowchart of a method for judging the stability of a landslide dam provided by the present invention according to an exemplary embodiment.

[0039] Figure 3 FIG. is a schematic flowchart of the LightGBM algorithm for processing missing values provided by the present invention according to an exemplary embodiment.

[0040] Figure 4 FIG. is a bar chart showing the relationship between the aspect ratio and stability provided by the present invention according to an exemplary embodiment.

[0041] Figure 5 FIG. is a bar chart showing the relationship between the dam shape coefficient and stability provided by the present invention according to an exemplary embodiment.

[0042] Figure 6 FIG. is a bar chart showing the relationship between the overtopping risk coefficient and stability provided by the present invention according to an exemplary embodiment.

[0043] Figure 7A bar graph showing the relationship between a lake shape coefficient and stability provided by the present invention according to an exemplary embodiment.

[0044] Figure 8 The figure is a schematic diagram of a Bayesian optimization algorithm flow chart provided by the present invention according to an exemplary embodiment.

[0045] Figure 9 A ROC curve diagram of a model test set provided by the present invention according to an exemplary embodiment.

[0046] Figure 10 A schematic diagram of a Youden point provided according to an exemplary embodiment of the present invention.

[0047] Figure 11 The present invention is a block diagram of a device for determining stability of a barrier dam according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to enable those skilled in the art to better understand the technical solution of the present invention and implement it, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the scope of protection of the present invention.

[0049] The present invention solves the problem that the existing research on the stability evaluation of landslide dams has higher requirements on the integrity of the characteristic data of landslide dams and insufficient evaluation under the condition of missing data. To achieve the above purpose, the present invention proposes a rapid evaluation model for the stability of landslide dams under the condition of missing data based on FL-Light GBM with feature combination and Bayesian optimization. Figure 1 As shown in the figure, firstly, a basic data database of landslide dam cases is constructed with reference to existing landslide dam research cases; at the same time, the influencing factors of landslide dam stability are analyzed, and 8 features including dam body parameters are determined, and the overall missing situation of the 8 features in the constructed database is analyzed; then, 4 new features are constructed by using the feature combination method of linear combination, the purpose of which is to further mine the information contained in the data and thus improve the model performance; secondly, the focal loss function (FL) is introduced to improve the Light GBM algorithm to obtain the FL-Light GBM algorithm, and the Bayesian optimization algorithm is used to find the optimal parameter combination; finally, the applicability of the model is analyzed using 4 actual cases of landslide dams.

[0050] The technical solutions provided by various embodiments of the present invention are described in detail below in conjunction with the accompanying drawings.

[0051] First, the present invention provides a method for determining the stability of a dam. Figure 2 As shown, the following steps are included:

[0052] S201. Introduce the focal loss function FL to improve the Light GBM algorithm to obtain the FL-Light GBM algorithm.

[0053] In this step, first, conduct a deconstruction analysis of the Light GBM algorithm. GBDT (Gradient Boosting Decision Tree) is an ensemble learning algorithm based on the Boosting idea. The basic idea of the algorithm is to use regression trees to iteratively update the weights of the training set, with the advantages of good training effect and not being prone to overfitting. LightGBM is an ensemble learning algorithm improved on the basis of traditional GBDT, which introduces two new technologies: Gradient-based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB).

[0054] (1) Gradient-based One-Side Sampling GOSS

[0055] The gradient-based one-side sampling algorithm starts from the perspective of reducing samples. When calculating the information gain, most samples with small gradients are excluded, and large-gradient samples that have a greater impact on the information gain are retained. When the model selects the splitting node of the decision tree, it selects the feature with the largest information gain for splitting. The variance after splitting is often used to measure the information gain after splitting of the gradient boosting tree.

[0056] Suppose O is the training set on the fixed node of the decision tree. The variance gain of the splitting feature j at the point d of this node is defined as the following formula (1):

[0057]

[0058] In the formula: n O = ΣI[x i ∈O], and n is the dimension of the data, and x i is the feature vector corresponding to sample i.

[0059] For feature j, in the GOSS algorithm, first, sort the training samples in descending order according to the absolute value of the gradient of the training samples; then retain the data subset A of the top-a×100% of the samples with larger gradients; for the samples A c with smaller gradients in the remaining set j(1 - a)×100%, further randomly sample to obtain a subset B with a size of b×|A c |; finally, divide the data according to the estimated variance gain of the subset A∪B, as shown in formula (2):

[0060]

[0061] In the formula: A l ={x i ∈A: x ij ≤d}, A r ={x i ∈A: x ij >d}, B l ={x i ∈B: x ij ≤d}, B r ={x i ∈B: x ij >d}, the coefficient is used to normalize the sum of gradients on the subset B to the size of A c .

[0062] Therefore, the GOSS algorithm uses the estimate of a smaller sample subset to replace the exact V on all data sets j (d) to determine the split point. While reducing the computational cost, the following formulas can be used to reduce the loss of computational accuracy of the GOSS algorithm and due to the results obtained by random sampling.

[0063] The approximation error is defined as Equation (3):

[0064]

[0065] There is a probability of at least 1 - δ to obtain Equation (4):

[0066]

[0067] Where

[0068] According to the above analysis, it can be obtained that

[0069] 1) The asymptotic approximation rate of the GOSS algorithm is If the split is relatively balanced, then At the same time The approximation error is controlled by the second term in Equation (4), and As n → ∞, the error will tend to 0.

[0070] 2) When a = 0, it is random sampling, which is a special case of the GOSS algorithm. In general, GOSS is better than random sampling.

[0071] Next, discuss the generalization performance of GOSS. Define the generalization error of GOSS to represent the difference between the variance gain obtained from the sampled training data and the potential true variance gain, and it can be as Equation (5):

[0072]

[0073] On the one hand, if the GOSS approximation is accurate, its generalization error will be very close to the error calculated using all the data. On the other hand, sampling diversifies the base learners, which helps improve the generalization performance of the model.

[0074] (2) Exclusive Feature Bundling EFB

[0075] The EFB algorithm is another important feature of Light GBM. For high-dimensional datasets with sparse data, EBF reduces the number of features by binding and fusing some features. It mainly solves two problems: (1) which features to bundle together; (2) how to bundle the features together.

[0076] (3) Principle of Light GBM Algorithm for Handling Missing Values

[0077] The Light GBM algorithm can automatically identify and process missing values in the dataset through the built-in sparse matrix perception algorithm. Its processing method can be Figure 3 illustrated as follows. For feature F0, during the training process, if there are missing values in feature F0, the processing steps are as follows:

[0078] 1) First, for the non-missing data of F0, calculate the split gain and compare their magnitudes, select the largest gain, and determine it as the split node (i.e., select a certain threshold of a certain feature);

[0079] 2) Then, for the missing data of F0, divide the missing values into the left subtree and the right subtree respectively, calculate the split gains of the left subtree and the right subtree respectively, select the direction with a larger gain, and use this direction as the split direction of the missing values.

[0080] During the prediction stage, if there are missing values in feature F0, it can be divided into the following two cases:

[0081] 1) If there were missing values in F0 during the training process, divide it according to the division direction of the missing values during the training process.

[0082] 2) If there were no missing values in F0 during the training process, divide the missing values into the default direction (left subtree).

[0083] (4) Improvement and Optimization of Light GBM Algorithm

[0084] To solve the problems of unbalanced feature distribution of barrier dams and missing data in the dataset, the Focal Loss function (FL) is introduced to improve the Light GBM algorithm to obtain the FL-Light GBM algorithm.

[0085] 1) Improvement of Focal Loss in Light GBM

[0086] The Focal Loss function is a loss function introduced in the cv field to deal with imbalanced datasets

[12] , and it can be regarded as a generalized form of the binary cross-entropy function. The loss function used by Light GBM to handle binary classification tasks is the binary cross-entropy function, and its calculation formula is shown in Equation (6):

[0087]

[0088] p represents the probability value that the classification label y equals 1 during the model training process. For the convenience of representation, let pt be defined as Equation (7):

[0089]

[0090] Then the cross-entropy can be expressed as Equation (8):

[0091] CE(p,y) = CE(p t ) = -log(p t ). (8)

[0092] Adding a weight factor α ∈ [0,1] before the cross-entropy loss function, the weight factor is used to change the proportion of stable and unstable barrier lakes, which can solve part of the sample imbalance problem. In addition, in imbalanced datasets, there are many easily classifiable simple negative samples, and the loss of these samples will dominate the gradient descent direction during training, submerging the minority class samples. Therefore, it is necessary to reduce the impact of simple negative samples. Adding a modulation factor (1 - p t ) γ , where γ > 0 is an adjustable parameter. The Focal Loss function is defined as Equation (9):

[0093]

[0094] During the model iteration process, when a sample is misclassified, the value of p for this sample t is very small, and the modulation factor (1 - p t ) is approximately equal to 1, and the loss value is hardly affected. When the sample is easy to classify, the value of p for this sample t is very large, and the modulation factor (1 - p t)Approximately equal to 0, so the weight of easy-to-separate samples decreases. The role of the adjustment factor can be enhanced by the focusing parameter γ. A large loss value indicates a relatively large training error for the sample. A small loss value means a relatively small training error for the sample. The goal of the model during training is to reduce the total loss value. The focal loss function self-adjusts the loss value based on the probability obtained by classifying the sample, and finally achieves the purpose of adjusting the sample weight. Ultimately, the algorithm pays more attention to difficult-to-classify samples and reduces the influence of easy-to-classify samples. The final form of the Focal Loss function is shown in Equation (10):

[0095] FL(p,y) = FL(p t ) = -α(1 - p t ) γ log(p t ). (10)

[0096] After introducing the Focal Loss function, not only can the problem of unbalanced sample class distribution be solved, but also for the abnormal distribution of barrier lake samples (such as samples at the classification boundary), the improved classification algorithm can effectively classify them. Combining with the data of the present invention, the parameter α is taken as 0.64 and γ is taken as 2.

[0097] S202. Based on the obtained initial parameters, construct combined parameters through linear combination, and construct training samples according to the initial parameters and combined parameters.

[0098] In this step, first, data screening is carried out, and barrier lake cases used for model establishment are screened out according to certain principles to reduce the uncertainty during model establishment. The screened data is used to construct new data through a feature combination method. The new data is added to the original data set to form a new data set, and the new data set is randomly divided into a training set and a test set according to a ratio of 8:2.

[0099] Specifically, obtain the initial parameters of the barrier lake including dam body parameters and environmental parameters, and through linear combination, use the initial parameters to construct combined parameters directly related to the stability of the barrier lake. Construct training samples according to the initial parameters, combined parameters, and the corresponding stability state of the barrier lake.

[0100] For example, through the collection and statistics of a series of databases containing different characteristic parameters established when studying the stability and breach parameters of barrier lakes in existing solutions, a database containing 2,179 barrier lake cases is established, including basic information such as whether the barrier lake is formed and whether it is stable, the location, name, occurrence time of the barrier lake, and whether the barrier lake is formed and whether it is stable. For geomorphological characteristics, the main collected data includes the dam length, dam width, dam height, and dam body volume of the barrier lake. For hydrodynamic characteristics, the main collected data includes the storage capacity and catchment area of the upstream barrier lake.

[0101] Based on past experience and considering the easy accessibility of data, hydrogeological conditions, and characteristic parameters of the barrier lake dam itself, etc., eight characteristics generated from relevant data are selected to consider the influencing factors of the stability of the barrier lake dam, as shown in Table 1.

[0102] Table 1 Statistical table of characteristics related to the stability of the barrier lake dam

[0103]

[0104]

[0105] As shown in the above table, the initial parameters include dam height, dam width, dam length, dam body volume, reservoir capacity, basin area, material composition, and inducing factors. Then, the overall data missing situation of the eight relevant characteristics of the stability of the above-mentioned barrier lake dam in the database is analyzed, and its missing overview is shown in Table 2.

[0106] Table 2 List of data missing situations

[0107]

[0108] To eliminate the influence caused by missing data, in this step, the method of feature combination is adopted to construct new data through the existing data. Feature combination, also known as feature cross, feature construction, etc., refers to constructing new features through the mutual combination of features, and its purpose is to further mine the information contained in the data, thereby improving the performance of the model. Considering the influence of the mutual combination of each feature of the barrier lake dam on the stability of the barrier lake dam, four new data are constructed in this step by means of linear combination.

[0109] (1) Aspect ratio, which refers to the ratio of the height of the barrier lake dam to the dam body, and the specific expression is:

[0110] X 9 = W / H.

[0111] According to the statistical analysis of barrier lake dam cases, the quantity and proportion of two types of barrier lake dam cases in each interval of the aspect ratio of the barrier lake dam are as Figure 4 shown. It can be seen that as the aspect ratio increases, the proportion of stable barrier lake dams gradually increases, indicating that the stability of the barrier lake dam is gradually increasing.

[0112] (2) Dam shape coefficient, which refers to the dimensionless ratio of the dam body volume to the dam height, and the specific expression is:

[0113]

[0114] Figure 5For the distribution of stable and unstable barrier lakes in each interval of the characteristic values in the statistical cases of barrier lakes in the database. As this value increases, the stability of the barrier lakes has been improved to a certain extent.

[0115] (3) The overtopping risk coefficient is the dimensionless ratio of the basin area to the dam height, and the specific expression is:

[0116]

[0117] Figure 6 For the distribution of stable and unstable cases in each interval of this characteristic value. As this value increases, the stability of the barrier lakes has been improved to a certain extent.

[0118] (4) The lake shape coefficient is the dimensionless ratio of the reservoir capacity and the dam height, and the specific expression is:

[0119]

[0120] where W is the dam width, H is the dam height, V d is the dam body volume, A is the basin area, V l is the reservoir capacity.

[0121] Figure 7 The distribution of two types of barrier lakes in each interval of the characteristic value in the database cases is counted. As this value increases, the stability of the barrier lakes has been reduced to a certain extent.

[0122] After the combined parameters are constructed, the combined parameters are added to the dataset of the initial parameters to obtain a new dataset. The new dataset is randomly divided into a training set and a test set according to a preset ratio. In the present invention, the new dataset is randomly divided into a training set and a test set in a ratio of 8:2.

[0123] S203. Using this training sample, train the FL-Light GBM algorithm through the Bayesian optimization algorithm to obtain a stability prediction model.

[0124] The Bayesian parameter optimization uses a Gaussian process, which can consider the parameter information searched before, has fewer iteration times, a fast convergence speed, and is robust for non-convex optimization problems, and can relatively easily find the local optimal solution. Therefore, the present invention uses a Bayesian optimization method based on a Gaussian process combined with cross-validation to search the hyperparameter space and find the optimal parameter combination.

[0125] ① Gaussian process

[0126] The Gaussian Process (GP) is a stochastic process in probability theory. It is a concept in probability theory that refers to the linear combination within a set of random variables that follow a normal distribution. Any linear combination of random variables in this process follows a normal distribution. When optimizing the parameters of Light GBM, the Gaussian process can be regarded as a combination of hyperparameters of the LightGBM model, which can be expressed as Equation (11):

[0127] f(x) ∼ GP(μ(x), k(x, x')); (11)

[0128] Among them, f(x) is the objective function; μ(x) is the mean function, μ(x) = E(f(x)); k(x, x') is the covariance function. Assume that the known parameter set is where f t = f(x t ), the next sampling point x t+1 . Assume that the mean of the prior distribution is 0, then the joint distribution of f and f t+1 is expressed as Equation (12):

[0129]

[0130] Among them, K is the matrix composed of k(x, x'); K t+1 is the matrix composed of k(x, x′ t+1 ); K t+1,t+1 is the matrix composed of k(x t+1 , x′ t+1 ); The posterior distribution P(f t+1 |D t+1 , x t , x t+1 ) is expressed as Equation (13):

[0131]

[0132] Among them, the mean of the posterior distribution of f t+1 variance

[0133] ② Acquisition function

[0134] After establishing the GP Gaussian process, an acquisition function is needed to determine the next sampling point x t+1 ​, common acquisition functions include probability of improvement (PI), Excepted Improvement (EI), and upper confidence bound (UCB). The present invention selects the simple and easy-to-use PI function. The next sampling point based on the PI strategy is shown in Equation (14):

[0135]

[0136] where Φ(·) represents the standard normal cumulative distribution function; f t (x + ) is the maximum value of the current objective function; ε is the trade-off coefficient.

[0137] The main process of the Bayesian optimization algorithm is as Figure 8 shown.

[0138] S204. Obtain the initial parameters and combined parameters of the target barrier lake dam, and input them into the stability prediction model to obtain the output stability judgment result of the target barrier lake dam.

[0139] In this step, it is necessary to apply the stability prediction model to actual use. Obtain the initial parameters of the target barrier lake dam to be predicted, construct the combined parameters according to the initial parameters in the above steps, and then input both the initial parameters and the combined parameters into the stability prediction model to obtain the stability prediction result of the target barrier lake dam.

[0140] To further illustrate the present invention, the following specific embodiments are also provided in this step:

[0141] 1. Data selection

[0142] Before model construction, it is necessary to select data from the database. If there are too many missing characteristic data in the samples, it will bring great uncertainty to the evaluation results. Therefore, it is necessary to screen the samples. The selection principle is: cases of barrier lake dams with no more than 2 missing characteristics and at least the characteristics of reservoir capacity or basin area. A total of 350 cases of barrier lake dams are selected to form the research data set, including 127 stable barrier lake dams and 223 unstable barrier lake dams. Some cases of barrier lake dams with missing characteristic data and the overall descriptive statistics are shown in Tables 3 and 4.

[0143] Table 3 Situations of some barrier lake dams with missing data

[0144]

[0145]

[0146] Table 5 Statistical table of 350 cases of barrier lake dams

[0147] Feature Name Number of Cases Missing Rate Minimum Value Maximum Value Average Value Standard Deviation Dam Height (m) 350 0 2.0 1825.0 58.24 139.74 Dam Width (m) 340 2.9% 20 5000 614.01 647.48 Dam Length (m) 342 2.3% 5.0 3750.0 341.13 351.29 <![CDATA[Watershed area (km 2 )]]> 263 24.9% 0.19 173484.00 2175.30 15971.56 <![CDATA[Volume of the dam body (10 6 m 3 )]]> 334 4.6% 0.0045 2200.00 28.3 187.26 <![CDATA[Storage capacity (10 6 m 3 )]]> 242 30.9% 0.00014 17000.00 184.52 1499.42 Inducing Factor 264 24.6% / / / / Material Composition 314 10.3% / / / /

[0148] 2. Validation of the Effectiveness of the Light GBM Algorithm

[0149] Before building the model, it is necessary to verify the effectiveness of the Light GBM algorithm and explore whether Light GBM can achieve a high accuracy rate under the condition of missing data. Compared with other algorithms, it is compared with the results of not imputing missing values and evaluating the stability after imputing specific values.

[0150] Experiments are carried out on the barrier lake dataset containing missing data. The dataset used in the experiment is the original dataset composed of 350 barrier lake cases with missing data.

[0151] First, the original dataset is divided into a training set and a test set in a ratio of 8:2. Then, the training set is imported into the random forest, XGBoost, and Light GBM algorithms to train the baseline model, and the performance of the baseline model is tested on the test set. In the data preprocessing stage, since all three algorithms are ensemble tree models, there is no need to normalize the features. Except for the Light GBM algorithm, the random forest and XGBoost use one-hot encoding for categorical features.

[0152] Secondly, the missing value imputation algorithm is used to complete the missing data to form a complete dataset, and then the model is trained and tested in the same way on the complete dataset. Through the experiments on the previous simulated dataset, the parameter K of the KNN imputation method takes the value of 7, the number of iterations of the MissForest imputation method is 10, and the number of imputations of the MICE multiple imputation method is 5 times, and the number of iterations is 10.

[0153] To ensure fairness, the training set and test set used by each algorithm before and after missing value imputation remain the same, and no operations other than necessary data encoding are performed. Each experiment is repeated 10 times, and the final evaluation index is the average AUC value of the ten experiments. The experimental results are shown in Table 5.

[0154] Table 5 Average AUC under Different Methods and Conditions

[0155] Missing Data Handling Method Random Forest XGBoost LightGBM KNN Imputation Method 0.755 0.756 0.748 MissForest Imputation Method 0.745 0.742 0.738 MICE Multiple Imputation Method 0.754 0.755 0.760 No Imputation 0.768 0.783 0.792

[0156] According to Table 5, the AUC value of the baseline model established by the Light GBM algorithm under the default parameters has reached a good level close to 0.8. It can be seen that the Light GBM algorithm has stronger applicability to the missing data of the barrier lake of the present invention.

[0157] 3. Establish an Evaluation Model

[0158] First, according to the feature combination method in the above steps, construct 4 new features to obtain a new dataset for establishing an evaluation model. Table 6 is a statistical table of the characteristics of 12 barrier lakes in the new dataset.

[0159] Table 6 Statistical Table of the Characteristics of Barrier Lakes in the New Dataset

[0160]

[0161]

[0162] Since Light GBM supports both missing values and categorical features, data preprocessing is not required before model establishment, and the model can be directly trained and tested on the dataset. The main steps for establishing the model are as follows:

[0163] Step1: Divide the new dataset into a training set and a test set according to a ratio of 8:2.

[0164] Step2: On the training set, use Bayesian optimization and 5-fold cross-validation to search for the optimal parameter combination, and use the average AUC value of cross-validation as the optimization target.

[0165] Step3: Set the optimal parameter combination to retrain the model on the training set, and then verify the generalization performance of the model on the test set. In addition, use the Light GBM algorithm and the FL-Light GBM algorithm to establish models on the original dataset without new features as a control.

[0166] To establish an evaluation model with good performance, it is generally necessary to adjust the parameters of the algorithm. Therefore, use Bayesian optimization to optimize and adjust the parameters of the FL-Light GBM algorithm. Select 6 parameters that affect the FL-Light GBM model for optimization. These parameters affect aspects such as the accuracy, training speed, and fitting degree of the model. Table 7 shows the definitions and value ranges of the hyperparameters.

[0167] Table 7 Optimized Parameters and Value Ranges

[0168] Number Parameter Value Range Definition 1 learning_rate (0.001,0.3) Learning Rate 2 max_depth (2,10) Maximum Depth of Tree 3 min_child_samples (2,100) Minimum Data Volume of Leaf Node 4 num_leaves (0,150) Number of Leaves per Tree 5 reg_lambda (0,6) Regularization Coefficient 6 subsample (0.7,1.0) Sampling Ratio per Iteration

[0169] When using the Bayesian optimization algorithm to search for optimal parameters, the selection of the target will affect the output optimal parameters. The AUC value, which has the strongest comprehensiveness among model evaluation indicators, is selected as the optimization target. 5-fold cross-validation is set, and the number of iterations is 30 times. Table 8 shows some results of the Bayesian optimization process. It can be seen that when the Bayesian optimization iterates 24 times, the target is the highest, and the optimal parameter combination is learning_rate: 0.1442, max_deepth: 7.26, min_child_samples: 56.88, num_leaves: 38.35, reg_lambda: 9.161, subsample: 0.7454.

[0170] Table 8 Some Results of Bayesian Optimization

[0171] Iter Target Parameter 1 Parameter 2 Parameter 3 Parameter 4 Parameter 5 Parameter 6 1 0.8435 0.2228 2.211 124.7 112.6 4.666 0.8829 2 0.8695 0.05286 2.592 172.0 114.8 7.972 0.601 3 0.8615 0.1442 7.26 56..88 38.35 9.161 0.7454 4 0.8913 0.1748 7.691 55.9 37.07 9.217 0.9811 5 0.8806 0.1856 2.007 2.116 147.4 7.088 0.9894 6 0.8182 0.001 2.0 2.0 10.0 10.0 0.6 7 0.7348 0.001 10.0 102.6 150.0 0.0 0.6 8 0.8802 0.3 10.0 40.53 92.41 10.0 1.0 9 0.8967 0.3 2.0 97.24 10.0 10.0 1.0 10 0.8034 0.001 2.0 2.0 111.9 0.0 1.0 11 0.8058 3.0 2.0 14.66 81.24 10.0 0.6 12 0.8315 1.389 9.959 6.06 84.17 9.31 0.9317 13 0.8897 0.1656 9.376 34.09 149.3 8.984 0.9742 14 0.8878 0.3 2.0 102.4 74.17 10.0 0.6 15 0.8145 0.3 10.0 24.56 122.8 10.0 1.0 … 24 0.8955 0.1442 7.26 56.88 38.35 9.161 0.7454 25 0.8468 0.3 10.0 125.5 10.0 0.0 0.6 …

[0172] Set the algorithm parameters to the above optimal parameter combination, retrain the optimal model on the training set, and then test the generalization ability of the model on the test set. When outputting the results, the ROC curve is plotted with the output probability values. The output results are as Figure 9 shown.

[0173] According Figure 9 to which, it can be seen that the AUC value of the model on the test set is 0.904, and the model also has good generalization ability on unknown data, reaching an excellent level.

[0174] Table 9 summarizes the results of the three models established by the present invention on the test set. All three models have been optimized by Bayesian parameters.

[0175] Table 9 Comparison of AUC of Three Models

[0176]

[0177] As can be seen from Table 9, compared with the unimproved Light GBM model, the AUC value of the improved FL-Light GBM model by the algorithm increases by 0.022, indicating the effectiveness of the proposed FL-Litgh GBM improvement algorithm of the present invention; the AUC of the model trained after adding new features is 0.904, and its performance is the best among the three models, which also proves that the idea of constructing new features by considering the interaction between landslide dam features is reasonable and effective.

[0178] It should be noted that in machine learning classification algorithms, the value directly output by the model is not a classification label, but the output probability value. By comparing it with the default threshold of 0.5 set by the model, if the model output probability value of the sample is greater than the threshold, it will be marked as 1, and if it is less than the threshold, it will be marked as 0, and then used as the final output result. Therefore, the choice of threshold will greatly affect the final output confusion matrix, and it is necessary to reasonably determine the threshold, that is, to determine the optimal threshold.

[0179] The present invention determines the optimal threshold by using the Youden index and the ROC curve.

[0180] The Youden index (Youden index, Y) is a statistic for evaluating the performance of a classification model. On the ROC curve, the point where Y takes the maximum value is the Youden point, and the corresponding threshold is the Youden threshold. The calculation formula is as follows:

[0181] Y = TPR - FPR (15)

[0182] In the formula, Y is the Youden index, TPR is the probability of correctly evaluating the positive samples among the positive classes in the original samples, and FPR is the probability of wrongly evaluating the positive samples among the negative classes in the original samples.

[0183] Figure 10 That is the Youden point determined according to the Youden index and the ROC curve, and the corresponding threshold is 0.535.

[0184] The evaluation results of the stability of 70 test set landslide dams before and after adjusting the output model threshold are shown in Table 10. Under the default threshold, among 45 cases of unstable landslide dams, 5 cases were misreported as stable, and the misjudgment rate was 7.1%. When the threshold was set to the optimal threshold of 0.535 obtained by using the ROC curve, 3 cases of unstable landslide dams were correctly classified. The misjudgment rate of the model under the optimal threshold was 2.9%, which was reduced by 4.2% compared with the misjudgment rate under the default threshold. The absolute accuracy rate of the model increased from 85.7% to 90.0%, maximizing the performance of the model.

[0185] Table 10 Comparison of test set evaluation results before and after threshold adjustment

[0186]

[0187] 4. Case application

[0188] The applicability of the model of the present invention is analyzed by using 4 actual cases of landslide dams, namely the Mellonggou landslide dam in Danba County, Sichuan Province, the Jiala landslide dam on the Yarlung Zangbo River in Tibet, the Nanba landslide dam formed by the Wenchuan earthquake, and the Usio landslide dam in Tajikistan.

[0189] Since some characteristic parameters of the above four barrier lakes are not fixed values but a range of values, and barrier lakes generally burst at their weakest positions, therefore, the minimum value of each characteristic value range is taken in the present invention for calculation. Table 11 shows the characteristic parameters taken for the calculation of each barrier lake and its stability.

[0190] Table 11 Specific information of four barrier lakes

[0191]

[0192] Table 12 shows the calculation results of the four barrier lakes by different evaluation methods and the corresponding stability evaluation results. It can be seen from the table that the stability evaluation results of the model established in the present invention for Meilonggou barrier lake, "10.29" Jiala barrier lake and Nanba barrier lake are unstable, and the stability evaluation result for Usoi barrier lake is stable, which is consistent with the actually observed stability.

[0193] Table 12 Stability evaluation results of each barrier lake

[0194]

[0195]

[0196] Using the above method, first, the Light GBM algorithm is improved by introducing the focal loss function FL. FL makes the model pay more attention to difficult-to-classify samples by reducing the weights of easy-to-classify samples, improves the performance of the model when dealing with imbalanced data sets, and improves the prediction accuracy of the model for the unstable state of minority-class barrier lakes. The combined parameters are constructed in a linear combination manner. On the one hand, it enriches the parameter data, and on the other hand, it mines the potential relationship between the initial parameters, so as to generate a richer feature set, which helps the model to capture more factors affecting the stability of the barrier lake and improves the prediction ability of the model. Finally, it is optimized by the Bayesian algorithm, which can automatically search for the optimal combination of hyperparameters, thereby improving the prediction performance and generalization ability of the model. Through the above method, it is possible to improve the prediction accuracy of the algorithm through parameter combination based on the improved algorithm in the case of missing data, thus establishing a stability evaluation model of the barrier lake under the condition of missing data and solving the problem of inaccurate stability prediction caused by missing data of the barrier lake.

[0197] Secondly, the present invention also provides a device for judging the stability of a barrier lake, as Figure 11 shown, including:

[0198] A construction module 1101, configured to improve the Light GBM algorithm by introducing the focal loss function FL to obtain the FL-LightGBM algorithm.

[0199] An acquisition module 1102, configured to acquire initial parameters of a barrier lake, including dam body parameters and environmental parameters, construct combined parameters directly related to the stability of the barrier lake in a linear combination manner by using the initial parameters, and construct training samples according to the initial parameters, the combined parameters, and their corresponding barrier lake stability states.

[0200] A training module 1103, configured to use the training samples to train the FL-Light GBM algorithm through a Bayesian optimization algorithm to obtain a stability prediction model.

[0201] A judgment module 1104, configured to input the initial parameters and combined parameters of the obtained target barrier lake into the stability prediction model and output the stability judgment result of the target barrier lake.

[0202] By adopting the above device, first, the Ligh GBM algorithm is improved by introducing a focal loss function FL. FL makes the model pay more attention to difficult-to-classify samples by reducing the weights of easy-to-classify samples, improves the performance of the model when dealing with unbalanced data sets, and improves the prediction accuracy of the model for the unstable state of minority-class barrier lakes. The combined parameters are constructed in a linear combination manner. On the one hand, it enriches the parameter data, and on the other hand, it excavates the potential relationship between the initial parameters, strengthens the correlation with the stability of the barrier lake, generates a richer feature set in this way, helps the model capture more factors affecting the stability of the barrier lake, and improves the prediction ability of the model. Finally, optimization is carried out through the Bayesian algorithm, and the optimal combination of hyperparameters can be automatically searched, thereby improving the prediction performance and generalization ability of the model. Through the above method, under the condition of data loss, based on the improved algorithm, the prediction accuracy of the algorithm can be improved through parameter combination, thereby establishing a stability evaluation model of the barrier lake under the condition of data loss and solving the problem of inaccurate stability prediction caused by data loss of the barrier lake.

[0203] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the steps of the stability judgment method of the barrier lake provided above. Figure 1 The steps of the stability judgment method of the barrier lake provided above.

[0204] The present invention also provides a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the steps of the stability judgment method of the barrier lake provided above. Figure 1 The steps of the stability judgment method of the barrier lake provided above.

[0205] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0206] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems) and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0207] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0208] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0209] It should be noted that the above specific implementation method can enable those skilled in the art to understand the invention more comprehensively, but does not limit the invention in any way. Therefore, although the present invention has been described in detail in this specification, those skilled in the art should understand that the invention can still be modified or replaced by equivalents; and all technical solutions and improvements that do not deviate from the spirit and scope of the invention are included in the protection scope of the patent for the invention. Any figure mark in the claims should not be regarded as limiting the claims involved.

Claims

1. A method for determining the stability of a landslide dam, characterized in that: The method comprises: The focus loss function FL is introduced to improve the Light GBM algorithm to obtain the FL-Light GBM algorithm; Acquire initial parameters of the barrier dam including dam body parameters and environmental parameters, construct combined parameters directly related to the stability of the barrier dam by using the initial parameters in a linear combination manner, and construct training samples according to the initial parameters, the combined parameters and the corresponding stability state of the barrier dam; Using the training samples, the FL-Light GBM algorithm is trained by a Bayesian optimization algorithm to obtain a stability prediction model; The obtained initial parameters and combined parameters of the target landslide dam are input into the stability prediction model, and the stability judgment result of the target landslide dam is output.

2. A method for determining the stability of a landslide dam according to claim 1, characterized in that: The formula for defining the focusing loss function FL is: Among them, p t The expression is: p represents the probability value of the classification label y equal to 1 during the model training process, α is the weight factor, (1-p t ) γ is the modulation factor of the simple negative sample, and γ>0 is an adjustable parameter.

3. A method for determining the stability of a landslide dam according to claim 1, characterized in that: The initial parameters include dam height, dam width, dam length, dam volume, reservoir capacity, basin area, material composition and inducing factors; the combination parameters directly related to the stability of the weir dam constructed by using the initial parameters in a linear combination manner include: Aspect ratio, the specific expression is: X9 = W / H; The dam body shape coefficient is expressed as follows: Overflow risk coefficient, the specific expression is: Lake shape coefficient, the specific expression is: Where W is the dam width, H is the dam height, V d is the volume of the dam, A is the basin area, V l For storage capacity.

4. A method for determining the stability of a landslide dam according to claim 1, characterized in that: Constructing training samples according to the initial parameters and combined parameters and their corresponding stability states of the barrier dam includes: The combined parameters are added to the data set of the initial parameters to obtain a new data set, and the new data set is randomly divided into a training set and a test set according to a preset ratio.

5. A method for determining the stability of a landslide dam according to claim 1, characterized in that: The training of the FL-Light GBM algorithm by the Bayesian optimization algorithm includes: Based on the Gaussian process of the Bayesian optimization algorithm, the training samples are used to construct a proxy model of the objective function of any parameter combination of the FL-Light GBM algorithm; Predict multiple new parameter combinations and their corresponding objective functions and uncertainties through surrogate models; Constructing an acquisition function according to the prediction result, wherein the acquisition function screens and optimizes the new parameter combination through its objective function and uncertainty to obtain a candidate parameter combination after screening and optimization; Iteratively optimize multiple candidate parameter combinations to complete the optimization of the FL-Light GBM algorithm.

6. A device for determining the stability of a landslide dam, characterized in that: The device comprises: A construction module is used to introduce the focus loss function FL to improve the Light GBM algorithm to obtain the FL-Light GBM algorithm; An acquisition module is used to acquire initial parameters of the barrier dam including dam body parameters and environmental parameters, construct a combination parameter directly related to the stability of the barrier dam by using the initial parameters in a linear combination manner, and construct a training sample according to the initial parameters and the combination parameters and their corresponding stability states of the barrier dam; A training module, used to train the FL-Light GBM algorithm using the training samples through a Bayesian optimization algorithm to obtain a stability prediction model; The judgment module is used to input the acquired initial parameters and combined parameters of the target dam into the stability prediction model, and output the stability judgment result of the target dam.

7. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

8. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 5 when executing the program.