Earth and rockfill dam break prediction method and related equipment

By combining the K-nearest neighbor algorithm and SMOTE oversampling with the LightGBM algorithm, the problem of incomplete data in the existing earth-rock dam break prediction method is solved, and a more accurate earth-rock dam break risk assessment is achieved with greater adaptability.

CN120670783APending Publication Date: 2025-09-19XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510818485.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing earth-rock dam failure prediction methods only consider a few failure factors, making it difficult to obtain a comprehensive and accurate assessment of earth-rock dam failure risk by taking into account multiple failure characteristic variables.

Method used

The K-nearest neighbor algorithm is used to fill the missing data of dam break characteristic variables, the SMOTE oversampling method is combined to perform data balancing, and the LightGBM algorithm is used to build a dam break prediction model.

Benefits of technology

It has achieved the establishment of complex mapping relationships based on a small amount of basic data, improved the scientificity and accuracy of dam break prediction, made it more adaptable, and was able to quickly and effectively assess the risk of earth-rock dam failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670783A_ABST
    Figure CN120670783A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of earth and rockfill dam break prediction methods, and discloses an earth and rockfill dam break prediction method, which takes dam break characteristic variables influencing a dam as input factors, introduces a K-nearest neighbor algorithm, carries out data filling, and then carries out equalization processing on the data. While filling of a large amount of data is ensured, the number of data iterations is greatly reduced, and the data is fully and uniformly distributed in a test specified range. The dam break prediction model is constructed based on the LightGBM algorithm, so that the dam break prediction of the model is more scientific and accurate, and the adaptability is higher. According to the method, a complex mapping relation can be established through a small amount of basic data, and the outburst prediction problem is accurately and efficiently solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of earth-rock dam failure prediction methods, and in particular relates to an earth-rock dam failure prediction method and related equipment. Background Art

[0002] Earth-rock dams, due to their low construction cost and strong adaptability to terrain and geology, have become the most widely used and fastest-growing dam type in dam construction. However, a dam failure can severely damage the downstream ecological environment and endanger people's lives and property. Therefore, there is an urgent need for objective, accurate, and rapid risk prediction methods for earth-rock dams to mitigate the risk of failure. Empirical methods for predicting earth-rock dam failures are subject to significant subjective influences and fail to account for factors affecting the failure, such as dam height, topographical and geological conditions, and geographic location.

[0003] On this basis, most existing methods for predicting earth-rock dam failures only consider a limited number of failure factors, making it difficult to comprehensively and accurately assess the failure risk of earth-rock dams by taking into account multiple failure characteristic variables. Therefore, it is necessary to develop a rapid and effective safety assessment and risk warning model design and optimization process for earth-rock dams to meet the actual needs of earth-rock dam failure prediction engineering construction. Summary of the Invention

[0004] The present invention provides a method for predicting earth-rock dam failure and related equipment, which solves the problem that most methods only consider a few failure factors and find it difficult to obtain a comprehensive and accurate assessment of the earth-rock dam failure risk while taking into account multiple failure characteristic variables.

[0005] To achieve the above object, the present invention provides the following technical solutions: A method for predicting earth-rock dam failure, comprising: Obtain the dam-break characteristic variables that affect the dam, and fill in the missing data of the dam-break characteristic variables that affect the dam using the K-nearest neighbor algorithm to obtain initial data; Perform equalization processing on the initial data to obtain a data set; Input the data set into the trained dam break prediction model and output the prediction results; Comparative analysis of the prediction model results; Among them, the dam break prediction model is built based on the LightGBM algorithm.

[0006] Preferably, the dam-break characteristic variables affecting the dam are obtained, and the missing data of the dam-break characteristic variables affecting the dam are filled using the K-nearest neighbor algorithm. The steps of obtaining the initial data are specifically as follows: Multiple dam-break characteristic variables were selected from the earth-rock dam case database, including control variables and indicator variables, with the ratio of control variables to indicator variables being 1:2. Use the K-nearest neighbor algorithm to fill in missing data, specifically: Calculate the Euclidean distance between the sample to be filled and other samples, and select the K most similar neighbors; For continuous variables, the weighted average of K neighbor eigenvalues ​​is taken as the filling value; For biclass variables, count the mode of the category labels in the K neighbors as the filling value.

[0007] Preferably, the specific formula for calculating the Euclidean distance between the sample to be filled and other samples is:

[0008] Where: D mn Indicates space m and n distance, S is the total number of input variables, xi m Indicates the i ( i ∈[1,..., S ]) variable m The value to be predicted, xi n Indicates the i The variable n A value to be predicted.

[0009] Preferably, the steps of performing equalization processing on the initial data to obtain the data set are specifically as follows: Use SMOTE oversampling method to deal with data imbalance problem: For minority class samples, calculate the distances to their K nearest neighbor samples. For each minority class sample, randomly select one of its K nearest neighbors and randomly select a point on the line connecting the original sample and the neighbor sample to generate a new sample. Synthetic samples are generated between minority class samples and randomly selected neighbor samples until the minority class and majority class samples are balanced.

[0010] Preferably, the training of the dam break prediction model is as follows: Obtain the dam-break characteristic variables that affect the dam, and fill in the missing data of the dam-break characteristic variables that affect the dam using the K-nearest neighbor algorithm to obtain initial data; Perform equalization processing on the initial data to obtain a data set; The dataset is divided into 70% training set and 30% test set; The LightGBM algorithm was used to train the model, and the parameters were set including the learning rate, the number of iterations (500), the maximum depth of the tree, and the number of leaf nodes to obtain the dam break prediction model; Among them, the Leaf-wise leaf growing strategy with depth restriction is used to replace the traditional layer-growing decision tree strategy.

[0011] Preferably, the steps of performing comparative analysis on the prediction model results are specifically as follows: The confusion matrix is ​​used to calculate the accuracy, precision, recall and F1 value of the dam break prediction model; The true positive rate and false positive rate corresponding to the dam break prediction model were calculated by changing the decision boundary of the classifier at different thresholds. The ROC curve was drawn and the AUC value was calculated to evaluate the model performance.

[0012] Preferably, the method for calculating the AUC value is:

[0013] Among them, M and N are the number of positive and negative samples respectively. is the ranking of positive samples.

[0014] A dam failure prediction system for earth-rock dams, comprising: Data filling module: used to obtain the dam-break characteristic variables that affect the dam, and fill the missing data of the dam-break characteristic variables that affect the dam using the K-nearest neighbor algorithm to obtain initial data; Data processing module: used to perform equalization processing on the initial data to obtain a data set; Prediction module: used to input the data set into the trained dam break prediction model and output the prediction results; Analysis module: used for comparative analysis of the output results of the prediction model; Among them, the dam break prediction model is built based on the LightGBM algorithm.

[0015] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of a method for predicting earth-rock dam failure are implemented.

[0016] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for predicting earth-rock dam failure.

[0017] Compared with the existing technology, the present invention has the following beneficial effects: The present invention provides a method for predicting earth-rock dam failure, which takes the characteristic variables affecting the dam failure as input factors, introduces the K-nearest neighbor algorithm, performs data filling, and then performs data balancing. While ensuring the filling of a large amount of data, the number of data iterations is greatly reduced, and the data is fully and evenly distributed within the specified range of the experiment. The dam failure prediction model is constructed based on the LightGBM algorithm, making the model dam failure prediction more scientific and accurate, and more adaptable. The method of the present invention can establish complex mapping relationships through a small amount of basic data, and accurately and efficiently solve the problem of failure prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of a method for predicting earth-rock dam failure according to the present invention; Figure 2 This is a flow chart of a method for predicting earth-rock dam failure according to an embodiment of the present invention; Figure 3 Statistics of case database of the embodiment of the present invention; Figure 4 A KNN-based calculation principle diagram of an embodiment of the present invention; Figure 5 This is a graph comparing the target variable training set ratios before and after SMOTE oversampling in an embodiment of the present invention; Figure 6 : This is a confusion matrix analysis diagram of each model in the embodiment of the present invention; Figure 7 This is a comparison chart of the accuracy of various combination models according to an embodiment of the present invention; Figure 8 This is the ROC curve of the LightGBM model in the embodiment of the present invention; Figure 9 This is a block diagram of an earth-rock dam failure prediction system according to the present invention. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0020] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.

[0021] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not require further definition or explanation in subsequent drawings.

[0022] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0023] like Figure 1 As shown, the present invention provides a method for predicting earth-rock dam failure, comprising: S1 obtains the dam-break characteristic variables affecting the dam, and fills the missing data of the dam-break characteristic variables affecting the dam using the K-nearest neighbor algorithm to obtain initial data; S2 performs equalization processing on the initial data to obtain a data set; S3 inputs the data set into the trained dam break prediction model and outputs the prediction results; S4 conducts comparative analysis of the prediction model results; Among them, the dam break prediction model is built based on the LightGBM algorithm.

[0024] Obtain the dam failure characteristic variables that affect the dam, and fill the missing data of the dam failure characteristic variables that affect the dam using the K-nearest neighbor algorithm. The steps for obtaining the initial data are as follows: Multiple dam-break characteristic variables were selected from the earth-rock dam case database, including control variables and indicator variables, with the ratio of control variables to indicator variables being 1:2. Use the K-nearest neighbor algorithm to fill in missing data, specifically: Calculate the Euclidean distance between the sample to be filled and other samples, and select the K most similar neighbors; For continuous variables, the weighted average of K neighbor eigenvalues ​​is taken as the filling value; For biclass variables, count the mode of the category labels in the K neighbors as the filling value.

[0025] The steps for equalizing the initial data and obtaining the data set are as follows: Use SMOTE oversampling method to deal with data imbalance problem: For minority class samples, calculate the distances to their K nearest neighbor samples. For each minority class sample, randomly select one of its K nearest neighbors and randomly select a point on the line connecting the original sample and the neighbor sample to generate a new sample. Synthetic samples are generated between minority class samples and randomly selected neighbor samples until the minority class and majority class samples are balanced.

[0026] The training of the dam break prediction model is as follows: Obtain the dam-break characteristic variables that affect the dam, and fill in the missing data of the dam-break characteristic variables that affect the dam using the K-nearest neighbor algorithm to obtain initial data; Perform equalization processing on the initial data to obtain a data set; The dataset is divided into 70% training set and 30% test set; The LightGBM algorithm was used to train the model, and the parameters were set including the learning rate, the number of iterations (500), the maximum depth of the tree, and the number of leaf nodes to obtain the dam break prediction model; Among them, the Leaf-wise leaf growing strategy with depth restriction is used to replace the traditional layer-growing decision tree strategy.

[0027] like Figure 2 As shown, another embodiment of the present invention provides a method for predicting earth-rock dam failure, comprising: Step 1: Determine the characteristic variables that affect dam failure: dam location N1, dam construction time N2, dam operation time X3, etc., and fill in the missing data using the K-Nearest Neighbors (KNN) algorithm. Step 2: Establish an earth-rock dam failure database based on overtopping failure, seepage failure, and instability failure of earth-rock dam cases, and use the SMOTE oversampling method to address the data imbalance problem in the database; Step 3: Build a dam break prediction model based on LightGBM and introduce an evaluation method for earth-rock dam break prediction model to evaluate the model prediction performance; Step 4: Establish the confusion matrix between different prediction models and implement comparative analysis of dam break prediction model evaluation results based on KNN-SMOTE-LightGBM.

[0028] Step 1 is as follows: Step 1.1, from the database of 1437 earth-rock dam cases, the statistical data is shown in Figure 3 After uniformly processing the data, 30 important variables were selected from the characteristic variables of these dams as model input features. The data were coded, among which there were 10 control variables, namely, the location of the dam, N 1. Dam construction time N 2. Dam operation time N 3. Dam type N 4. Dam height N 5. Total storage capacity of earth-rock dam N 6. Designed discharge flow of the dam N 26 , average rainfall N 27, the distance between normal reservoir water level and dam top N 28 , dam crest width N 29 , dam crest length N 30 . And 20 indicator variables, namely, the presence of flood phenomena exceeding the design flood control standard in the reservoir area N 7. The gate fails to rise or fall due to a malfunction N 8. Insufficient spillway discharge capacity N 9. Damage to flood discharge structures N 10 , Improper human operation N 11 , dam cracks N 12 , dam slope instability N 13 2. The dam construction quality is poor N 14 , The filter layer is clogged or has poor durability N 15 , improper seepage control N 16 , animal nests or plant roots N 17 , cracks appear on the dam foundation and pavement N 18 , dam settlement N 19 , slope instability N 20 , untimely maintenance N 21 , improper management N 22 Floods overflowed the roof N 23 , penetration damage N 24 , dam design discharge flow N 25, The parameter values ​​of specific variables are shown in Tables 1 and 2.

[0029] Table 1 Characteristic variable data values

[0030] Table 2 Description of characteristic variables

[0031] Step 1.2: Since there is a large amount of missing data in the database, which affects the accuracy of model training, the K-nearest neighbor algorithm is used based on the dam break characteristic variables determined in step 1.1. The advantage of this algorithm is that it calculates the distance between the sample to be predicted and all samples in the training set, finds the K nearest samples, and then predicts the result of the sample to be predicted based on the label (classification) or value (regression) of these K samples. Therefore, by calculating the Euclidean distance between the characteristic variable of the dam and other characteristic variables, the K most similar samples are selected as neighbors of the missing value based on the calculated distance. The weighted average of the characteristic values ​​of these K neighbors is used as the estimated value of the missing variable, and finally the missing value is replaced with the interpolated estimated value. The specific formula is as follows:

[0032] Where: D mn Indicates space m and n distance, S is the total number of input variables, xi m Indicates the i ( i ∈[1,..., S ]) variable m The value to be predicted, xi n Indicates the i The variable n A value to be predicted.

[0033] Step 1.3: For missing data in categorical variables, use the K-nearest neighbor algorithm. First, determine the K most similar neighbors. These neighbors are the samples closest to the missing value. For these K neighbors, obtain their classification labels. Count the number of occurrences of each category in these K neighbors. Based on the statistical results, select the category with the most occurrences (i.e., vote) as the result of filling the missing value of the dam variable. The specific calculation principle is shown in the figure. Figure 4 .

[0034] Step 1.4: Repeat the above steps for multiple iterations until all missing values ​​are filled. In step 2, specifically: Step 2.1: According to the SMOTE oversampling method, the algorithm synthesizes new samples between minority class samples to increase the number of minority classes. Specifically, for each minority class sample, SMOTE randomly selects a sample from its nearest minority class neighbor, and then generates a new synthetic sample between these two samples to deal with the data imbalance problem. For the minority class sample in the earth-rock dam database, the distance to its K nearest neighbor samples is calculated. The comparison of the target variable training set ratio before and after SMOTE oversampling is shown in Figure 5 .

[0035] Step 2.2: Select one of the nearest neighbor samples and randomly select a point on the line segment between the sample and the original sample as a new synthetic sample.

[0036] Step 2.3: Repeat the above steps until a sufficient number of synthetic samples are generated.

[0037] In step 3, the sample data is divided into training set and test set. 70% of the data is used to train the model, and 30% of the data is used to test the model performance. The LightGBM algorithm is used for prediction. The principle of the LightGBM algorithm is: given the sample data X ={( x i ,y i )} N i=1 , where: x represents the sample data, y represents the category label, f ( x ) represents the estimation function, the purpose of LightGBM is to f ( x ) find an approximate function f ( x ), so that the loss function L ( y,f ( x )) has the least expectations.

[0038] (2) Among them, T regression trees are used T t ( x )( t =1,..., T ) to obtain a final approximate model: (3) The decision tree can be represented as w q(x) , q∈{1,2,...,J} ,in J Indicates the number of leaves, q represents the decision rule of the tree, w Is a vector representing the sample weight of the leaf node. Therefore, LightGBM will be trained in a deterministic form at the tth iteration, as shown in formula (4): (4) In LightGBM, the objective function is approximated using Newton's fast approximation. After removing the constant term in equation (4), the simplified formula is shown in equation (5): (5) In formula (5), g i and h i Represents the first-order and second-order gradient statistics of the loss function. I j Represents the sample data set of leaves, and formula (4) can be converted into (6) For a certain tree structure q ( x ), the optimal weight of each leaf node w*j Score and F The extreme values ​​of are shown in Equations (7) and (8).

[0039] (7) (8) F*t It can be seen as a measure of the tree structure q Finally, the objective function after segmentation is added, that is, (9) Where: I L and I R The sample sets for the left and right branches are respectively. The LightGBM algorithm builds on previous decision-making algorithms by optimizing the computational efficiency and memory usage of traditional gradient boosted tree (GBDT) algorithms for large-scale data processing. Therefore, the data was imported into the LightGBM library. Model parameters such as the learning rate, number of trees, and tree depth were set. A leaf-wise growth strategy with depth constraints was used instead of the traditional layer-by-layer decision tree growth strategy. The training set data was used to train the dam break prediction model.

[0040] The specific algorithm procedure is as follows: data = pd.read_csv('D:\database\knn filled dataset.csv') X = data.iloc[:, 0:30] Y = data.iloc[:, 30] # Use LabelEncoder to convert the target variable to integer encoding le = LabelEncoder() Y_encoded = le.fit_transform(Y) # Load the dataset X_train, X_test, y_train, y_test = train_test_split(X, Y_encoded,test_size=0.3, random_state=100) # Create and train the model model=LGBMClassifier(learning_rate=0.01,n_estimators=500,max_depth=5,num_leaves=20) model.fit(X_train, y_train) # Make predictions on the test set y_pred = model.predict(X_test) y_pred_proba = model.predict_proba(X_test) # Convert the predicted probability into a binary coded matrix lb = LabelBinarizer() y_test_binarized = lb.fit_transform(y_test) y_pred_binarized = lb.transform(y_pred) # Calculate classification accuracy accuracy = accuracy_score(y_test, y_pred) print("Test accuracy:", accuracy) from sklearn.metrics import confusion_matrix import matplotlib.pyplot as plt import seaborn as sns The specific parameters are as follows: Optimal parameter values ​​for the LightGBM model

[0041] In step 4, specifically: Step 4.1: In order to objectively and accurately evaluate the prediction ability of the model, the confusion matrix and ROC curve are introduced to evaluate the dam break prediction model.

[0042] Step 4.2: Select the confusion matrix evaluation indicators model accuracy, precision, recall and F1 value, with the calculation formula as follows:

[0043] Where: True Positive (TP) represents the number of samples correctly classified as category 1; False Positive (FP) represents the number of samples that misclassify category 2 as category 1; False Negative (FN) represents the number of samples that misclassify category 1 as category 2; True Negative (TN) represents the number of samples that are correctly classified as category 2. For specific confusion matrix evaluation indicators of each model, see Figure 6 .

[0044] Step 4.3: By changing the decision boundary of the classifier at different thresholds, we calculate the corresponding TPR and FPR, thus obtaining points on the curve. Finally, we connect all the points to obtain the complete ROC curve. AUC is the area under the ROC curve, which can be solved using numerical calculation methods. The calculation formula is as follows:

[0045] Step 4.4: To demonstrate the effectiveness of the LightGBM model, we introduce the support vector machine model, random forest model, and XGBoost model to compare their prediction performance with the LightGBM model in the earth-rock dam database. Figure 7 Based on the comparative analysis of the evaluation results of the dam-break prediction model of each model, the accuracy of the earth-rock dam break prediction results of different models is obtained. Figure 8 .

[0046] This embodiment fully considers the detailed information of 1,437 dams and establishes a database. It takes important breach parameters as input factors and introduces the KNN data filling algorithm and the SMOTE oversampling algorithm to process the data. While ensuring the filling of a large amount of data, the number of data iterations is greatly reduced, and the data is fully and evenly distributed within the specified range of the experiment. The KNN model is used to fill in missing data values. On this basis, SMOTE synthesizes new minority class samples to balance the difference in the number of samples between the majority class and the minority class. The integration of LightGBM makes the model dam breach prediction more scientific and accurate, and more adaptable. The method of the present invention can establish complex mapping relationships through a small amount of basic data, and solve them accurately and efficiently.

[0047] like Figure 9As shown, the present invention provides an earth-rock dam failure prediction system, comprising: Data filling module: used to obtain the dam-break characteristic variables that affect the dam, and fill the missing data of the dam-break characteristic variables that affect the dam using the K-nearest neighbor algorithm to obtain initial data; Data processing module: used to perform equalization processing on the initial data to obtain a data set; Prediction module: used to input the data set into the trained dam break prediction model and output the prediction results; Analysis module: used for comparative analysis of the output results of the prediction model; Among them, the dam break prediction model is built based on the LightGBM algorithm.

[0048] An embodiment of the present invention provides a terminal device. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of each of the aforementioned method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in each of the aforementioned device embodiments are implemented.

[0049] The computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to accomplish the present invention.

[0050] The terminal device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0051] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0052] The memory may be used to store the computer programs and / or modules, and the processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory.

[0053] If the module / unit integrated in the terminal device is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media does not include electric carrier signal and telecommunication signal.

[0054] Although the embodiments of the present invention have been described above with reference to the accompanying drawings, the present invention is not limited to the above-mentioned specific embodiments and application fields. The above-mentioned specific embodiments are merely illustrative and instructive, and are not restrictive. A person skilled in the art, guided by the description, may devise various forms without departing from the scope of protection of the claims of the present invention, all of which fall within the scope of protection of the present invention.

Claims

1. A method for predicting earth-rock dam failure, characterized in that: include: Obtain the dam-break characteristic variables that affect the dam, and fill in the missing data of the dam-break characteristic variables that affect the dam using the K-nearest neighbor algorithm to obtain initial data; Perform equalization processing on the initial data to obtain a data set; Input the data set into the trained dam break prediction model and output the prediction results; Comparative analysis of the prediction model results; Among them, the dam break prediction model is built based on the LightGBM algorithm.

2. The earth-rock dam failure prediction method according to claim 1, characterized in that: Obtain the dam failure characteristic variables that affect the dam, and fill the missing data of the dam failure characteristic variables that affect the dam using the K-nearest neighbor algorithm. The steps for obtaining the initial data are as follows: Multiple dam-break characteristic variables were selected from the earth-rock dam case database, including control variables and indicator variables, with the ratio of control variables to indicator variables being 1:

2. Use the K-nearest neighbor algorithm to fill in missing data, specifically: Calculate the Euclidean distance between the sample to be filled and other samples, and select the K most similar neighbors; For continuous variables, the weighted average of K neighbor eigenvalues ​​is taken as the filling value; For biclass variables, count the mode of the category labels in the K neighbors as the filling value.

3. The earth-rock dam failure prediction method according to claim 1, characterized in that: The specific formula for calculating the Euclidean distance between the sample to be filled and other samples is: Where: D mn Indicates space m and n distance, S is the total number of input variables, xi m Indicates the i ( i ∈[1,..., S ]) variable m The value to be predicted, xi n Indicates the i The variable n A value to be predicted.

4. The earth-rock dam failure prediction method according to claim 1, characterized in that: The steps for equalizing the initial data and obtaining the data set are as follows: Use SMOTE oversampling method to deal with data imbalance problem: For minority class samples, calculate the distances to their K nearest neighbor samples. For each minority class sample, randomly select one of its K nearest neighbors and randomly select a point on the line connecting the original sample and the neighbor sample to generate a new sample. Synthetic samples are generated between minority class samples and randomly selected neighbor samples until the minority class and majority class samples are balanced.

5. The earth-rock dam failure prediction method according to claim 1, characterized in that: The training of the dam break prediction model is as follows: Obtain the dam-break characteristic variables that affect the dam, and fill in the missing data of the dam-break characteristic variables that affect the dam using the K-nearest neighbor algorithm to obtain initial data; Perform equalization processing on the initial data to obtain a data set; The dataset is divided into 70% training set and 30% test set; The LightGBM algorithm was used to train the model, and the parameters were set including the learning rate, the number of iterations (500), the maximum depth of the tree, and the number of leaf nodes to obtain the dam break prediction model; Among them, the Leaf-wise leaf growing strategy with depth restriction is used to replace the traditional layer-growing decision tree strategy.

6. The earth-rock dam failure prediction method according to claim 1, characterized in that: The specific steps for comparative analysis of the prediction model results are as follows: The confusion matrix is ​​used to calculate the accuracy, precision, recall and F1 value of the dam break prediction model; The true positive rate and false positive rate corresponding to the dam break prediction model were calculated by changing the decision boundary of the classifier at different thresholds. The ROC curve was drawn and the AUC value was calculated to evaluate the model performance.

7. The earth-rock dam failure prediction method according to claim 1, characterized in that: The method for calculating the AUC value is: Among them, M and N are the number of positive and negative samples respectively. is the ranking of positive samples.

8. A dam failure prediction system for earth-rock dams, characterized in that: include: Data filling module: used to obtain the dam-break characteristic variables that affect the dam, and fill the missing data of the dam-break characteristic variables that affect the dam using the K-nearest neighbor algorithm to obtain initial data; Data processing module: used to perform equalization processing on the initial data to obtain a data set; Prediction module: used to input the data set into the trained dam break prediction model and output the prediction results; Analysis module: used for comparative analysis of the output results of the prediction model; Among them, the dam break prediction model is built based on the LightGBM algorithm.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the earth-rock dam failure prediction method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the earth-rock dam failure prediction method according to any one of claims 1 to 7 are implemented.