Beach bank collapse prediction method and system based on dynamic mechanism and machine learning

CN120336998AActive Publication Date: 2025-07-18WUHAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510357746.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-18
Estimated Expiration
2045-03-25

Smart Images

  • Figure CN120336998A_ABST
    Figure CN120336998A_ABST
Patent Text Reader

Abstract

The invention provides a shoal collapse prediction method based on a dynamic mechanism and machine learning. The method comprises the following steps: step 1, collecting historical basic data of a river reach where a prediction section is located; 2, screening the basic data, and calculating a characteristic variable, a beach bank collapse probability target variable and a beach bank collapse width target variable based on the screened basic data; step 3, constructing a beach bank collapse probability prediction model and a beach bank collapse width prediction model; 4, evaluating the precision of the beach bank collapse probability prediction model and the beach bank collapse width prediction model; and 5, based on the beach bank collapse probability prediction model and the beach bank collapse width prediction model, respectively predicting the beach bank collapse probability and the beach bank collapse width in the future time period. According to the method, the rule can be automatically learned from multi-source historical data, the influence of multiple factors is considered, the nonlinear mapping relation between the input feature and the output target is established, quantitative prediction of beach bank collapse is achieved, and the method has great significance in reducing disaster risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of river regime prediction, and particularly relates to a beach bank retreat prediction method and system based on dynamic mechanism and machine learning. Background Technique

[0002] Beach bank retreat is an instability failure phenomenon that occurs on the slopes of water areas such as rivers, lakes, and seas under the combined action of natural and human factors, and it is one of the major problems that need to be solved urgently in the field of water conservancy projects. Beach bank retreat is characterized by strong suddenness, great destructive power, and wide influence range, seriously threatening the lives and property safety of people along the coast and restricting the economic development along the coast.

[0003] Traditional beach bank retreat prediction methods are mainly based on empirical formulas and mechanical models. Empirical formulas are usually established based on observational data under specific regions and conditions. However, due to the large differences in geological conditions, hydrological conditions, and human activities in different regions, the universality of empirical formulas is poor and it is difficult to be popularized and applied to other regions or conditions. In addition, empirical formulas usually only consider a few influencing factors and it is difficult to comprehensively reflect the complex mechanism of beach bank retreat. Mechanical models are based on theories such as soil mechanics and river dynamics, considering various factors to establish a slope stability analysis model, and judging the retreat risk by calculating the slope safety factor. However, mechanical models require input of a large number of physical and mechanical property parameters of soil, such as the cohesion, internal friction angle, permeability coefficient, etc. of the soil, and these parameters are often difficult to obtain accurately, especially for complex stratum structures and heterogeneous soils, the difficulty of parameter acquisition is greater, which greatly affects the prediction accuracy of mechanical models.

[0004] In recent years, machine learning technology has shown strong advantages in data mining and pattern recognition, providing new ideas for solving the above problems. Machine learning algorithms can automatically learn laws from existing data, establish a non-linear mapping relationship between input features and output targets, so as to achieve more accurate prediction. However, the current quantitative prediction research on beach bank retreat width is still relatively limited. Some models fail to fully incorporate the influence of multiple factors such as water and sediment, riverbed boundary, and previous riverbed deformation, and there is a lack of association with the model in feature selection. Summary of the Invention

[0005] The present invention combines multi-source historical data, constructs a river beach bank retreat prediction model embedded with a feature selection module based on dynamic mechanism and machine learning algorithm, and can realize the intelligent prediction of the retreat probability and retreat width of a specific beach bank under certain future conditions.

[0006] To solve the above technical problems, the present invention adopts the following technical solutions:

[0007] A beach bank retreat prediction method based on dynamic mechanism and machine learning, comprising the following steps:

[0008] Step 1. Collect the historical basic data of the river reach where the prediction section is located;

[0009] Step 2. Screen the basic data and calculate the characteristic variables, the target variable of the bank collapse probability, and the target variable of the bank collapse width based on the screened basic data;

[0010] Step 3. Preprocess the calculated characteristic variables, construct the prediction models of the bank collapse probability and the bank collapse width, and train the bank collapse probability prediction model based on the preprocessed characteristic variables and the target variable of the bank collapse probability, and train the bank collapse width prediction model based on the preprocessed characteristic variables and the target variable of the bank collapse width;

[0011] Step 4. Evaluate the accuracy of the bank collapse probability prediction model and the bank collapse width prediction model;

[0012] Step 5. Predict the bank collapse probability and the bank collapse width in the future period based on the bank collapse probability prediction model and the bank collapse width prediction model respectively.

[0013] Further, the historical basic data in Step 1 includes the data of the flow rate, water level, sediment concentration, and median grain size of bed sediment of the hydrological station in the river reach where the prediction section is located during the historical period, as well as the measured topographic data of the prediction section.

[0014] Further, the target variable of the bank collapse probability in Step 2 is: whether the bank deforms and collapses. Mark the bank collapse sample as 1 and the stable or siltation sample as 0.

[0015] Further, the target variable of the bank collapse width is: the bank deformation width. By overlaying the cross-sectional topographies of adjacent measurement times of the prediction section, determine the starting distances of the bank lips on the left and right sides of the bank, and calculate the difference in the starting distances of the bank lips between adjacent measurement times as the bank deformation width.

[0016] Further, the characteristic variables include: characteristic variables of water and sediment conditions, characteristic variables of river bed boundaries, and characteristic variables of previous river bed deformations.

[0017] Further, the characteristic variables of water and sediment conditions include: average flow rate, average sediment concentration, coefficient of variation of flow rate, average water flow scouring intensity in the future period and the previous 3 periods, cross-sectional average flow velocity, flow rate change rate, average water depth near the bank, and average flow velocity near the bank;

[0018] The characteristic variables of river bed boundaries include: floodplain area, floodplain river width, and river facies coefficient; bank height and slope; relative distance of thalweg from the bank, distance of toe from the bank, relative elevation difference between the lowest point near the bank and the thalweg, and relative elevation difference between the near-bank river beds on the same and opposite sides; retreat width of toe; median grain size of bed sediment; distance of regulation project from the bank;

[0019] The characteristic variables of the previous riverbed deformation include: the swing width of the thalweg, the scouring and silting thickness, the change in the beach slope, and the deformation width of the beach.

[0020] Furthermore, the preprocessing in step 3 includes:

[0021] Based on the IterativeImputer method, impute the missing or outlier feature values;

[0022] Examine the autocorrelation of the characteristic variables and eliminate the highly autocorrelated features;

[0023] Based on the Yeo-Johnson transformation, perform debiasing on the characteristic variables;

[0024] Convert the characteristic variable data into a distribution with a mean of 0 and a standard deviation of 1;

[0025] Randomly or in a certain order divide the preprocessed characteristic variables and the beach bank recession probability target variable into a training set, a validation set, and a test set, which are used as the input data for the beach bank recession probability prediction model.

[0026] In the same format, use the preprocessed characteristic variables and the beach bank recession width target variable as the input data for the beach bank recession width prediction model.

[0027] Furthermore, the construction of the beach bank recession probability prediction and beach bank recession width prediction models in step 3 is specifically as follows:

[0028] First, use Lasso regression to screen the candidate characteristic variables and eliminate the low-correlation features; then input the screened feature set into the KNN model for training, jointly tune the number of characteristic variables and other hyperparameters of the KNN algorithm, and determine the hyperparameter combination that makes the model achieve the optimal performance on both the training set and the validation set through grid search; based on the historical basic data, train the beach bank recession probability prediction model and the recession width prediction model respectively.

[0029] Furthermore, step 5 includes:

[0030] Based on the trained beach bank recession probability prediction model, predict the beach bank recession probability under specific conditions. If the probability prediction model predicts it as a recession, then further predict its recession width based on the beach bank recession width prediction model.

[0031] On the other hand, the present invention provides a beach bank recession prediction system based on a dynamic mechanism and machine learning, including:

[0032] A data collection module, which is used to collect the historical basic data of the river section where the prediction section is located;

[0033] A variable calculation module, which is used to screen the basic data and calculate characteristic variables, target variables of beach bank recession probability and target variables of beach bank recession width based on the screened basic data;

[0034] A model construction module, which is used to preprocess the calculated characteristic variables, construct a beach bank recession probability prediction model and a beach bank recession width prediction model, and train the beach bank recession probability prediction model based on the preprocessed characteristic variables and the target variables of beach bank recession probability, and train the beach bank recession width prediction model based on the preprocessed characteristic variables and the target variables of beach bank recession width;

[0035] An accuracy evaluation module, which is used to evaluate the accuracy of the beach bank recession probability prediction model and the beach bank recession width prediction model;

[0036] A prediction module, which is used to predict the beach bank recession probability and the beach bank recession width in a future period based on the beach bank recession probability prediction model and the beach bank recession width prediction model respectively.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] The method of the present invention can automatically learn the rules from multi-source historical data, consider the influence of multiple factors, establish a non-linear mapping relationship between input features and output targets, realize the quantitative prediction of beach bank recession, and is of great significance for reducing disaster risks. Description of the Drawings

[0039] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 It is a flow chart of the beach bank recession prediction method of the present invention;

[0041] Figure 2 It is a schematic diagram of the extraction of target variables and the determination of the toe of the slope and the lowest point near the shore of the present invention;

[0042] Figure 3 It is a data distribution diagram before and after the debiasing and standardization processing of the characteristic variables of the present invention;

[0043] Figure 4 It is a schematic diagram of the basis for adjusting the parameters of the beach bank recession probability prediction model of the present invention;

[0044] Figure 5 It is a schematic diagram of the basis for adjusting the parameters of the beach bank recession width prediction model of the present invention;

[0045] Figure 6 Schematic diagram of the prediction result of the beach bank recession probability prediction model of the present invention;

[0046] Figure 7 Schematic diagram of the prediction result of the beach bank recession width prediction model. Specific implementation manners

[0047] To make the above objects, features and advantages of the present application more obvious and understandable, the following will describe the specific implementation manners of the present application in detail with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0048] Embodiment 1

[0049] The overall calculation process of this embodiment is shown in Figure 1 . It should be noted that: ① The beach bank recession prediction is carried out for a single beach bank, and even the left and right beach banks of the same cross-section need to be predicted separately; ② The recession probability prediction and the recession width prediction are respectively realized based on independent models; ③ The beach bank recession width prediction only uses the recession samples.

[0050] Step 1. Collect the historical basic data of the river section where the prediction cross-section is located;

[0051] Specifically, it includes collecting the data of the flow rate, water level, sediment concentration and median grain size of bed sediment of the hydrological station in the river section where the prediction cross-section is located, as well as the measured topographic data of the prediction cross-section. If there is a hydrological station near the prediction cross-section, the data of this hydrological station is directly used. If the prediction cross-section is far from the hydrological station, the data at the prediction cross-section is obtained by linear interpolation according to the data of the upstream and downstream hydrological stations. To ensure sufficient topographic data, the prediction cross-section is generally selected as a fixed cross-section in the river. The increase in the amount of data usually helps to improve the prediction accuracy, so relevant data should be collected as comprehensively as possible.

[0052] Step 2. Screen the basic data and calculate the characteristic variables, the beach bank recession probability target variable and the beach bank recession width target variable based on the screened basic data;

[0053] The screening of the basic data includes:

[0054] Eliminate the noise samples; delete the noise samples that do not belong to the category of beach bank deformation, such as the samples of the change in the beach lip position caused by measurement errors or the formation of new channels by the water flow.

[0055] In this embodiment, the target variable of the beach bank recession probability is whether the beach bank deforms and recedes. The beach bank recession samples are marked as 1, and the stable or accretion samples are marked as 0. The target variable of the beach bank recession width is the deformation width of the beach bank. By comparing the cross-section topographies of adjacent surveys of the arbitrage prediction cross-section, the distances from the starting points of the beach lips on the left and right sides of the beach bank are determined, and the difference in the distances from the starting points of the beach lips between adjacent surveys is calculated as the deformation width of the beach bank.

[0056] By comparing the cross-section topographies of adjacent surveys of the arbitrage prediction cross-section, the distances from the starting points of the beach lips on the left and right sides of the beach bank are determined, and the difference in the distances from the starting points of the beach lips between adjacent surveys is calculated as the deformation width of the beach bank. It is set that when the beach bank recedes (accretes), the deformation width is negative (positive). As Figure 2 shown, within one flood season, the left and right sides of the cross-section beach bank receded by 226 m and 98 m respectively. The present invention includes a beach bank recession probability prediction model and a recession width prediction model. The beach bank recession probability prediction model is trained using all samples (including recession and non-recession samples), and the target variable is whether the beach bank deforms and recedes. The beach bank recession samples are marked as 1, and the stable or accretion samples are marked as 0. Due to the limitation of the river channel topography measurement accuracy, only the samples with the beach bank recession width between adjacent surveys greater than the actual measurement error are marked as recession (1). For the wandering section of the lower Yellow River, only the samples with the recession width between adjacent surveys exceeding 6 m are marked as recession (1). The beach bank recession width prediction model is trained only using the recession samples, and the target variable is the beach bank recession width.

[0057] The characteristic variables calculated in this embodiment include: water and sediment condition characteristic variables, riverbed boundary characteristic variables, and previous riverbed deformation characteristic variables.

[0058] Calculate the water and sediment condition characteristic variables. First, based on the typical water and sediment characteristic parameters in existing research, including variables such as average flow rate, average sediment concentration, flow variation coefficient, the average water flow scouring intensity in the future period and the previous 3 periods, and cross-section average flow velocity, etc., which characterize the cross-section average water and sediment conditions. At the same time, the present invention innovatively introduces the flow rate change rate to more accurately describe the influence of the dynamic change of the flow rate during flood and dry seasons on the stability of beach bank recession. In addition, variables characterizing the near-shore water flow conditions are calculated, including near-shore average water depth and near-shore average flow velocity, etc. If there is no hydrological cross-section near the prediction cross-section, its water and sediment process is obtained by linear interpolation according to the measured data of the upstream and downstream hydrological cross-sections according to the distance. The calculation methods of the flow variation coefficient, flow rate change rate, and near-shore water depth and flow velocity are mainly introduced below.

[0059] ① Flow variation coefficient (Q Cv )

[0060] The flow variation coefficient is used to characterize the dispersion degree of the flow rate series:

[0061]

[0062] Where: Q is the average flow rate (m 3 / s) of the predicted cross-section during the flood season (non-flood season); n is the number of days in the flood season (non-flood season); Q i is the average daily flow rate (m 3 / s) on the i-th day of the flood season (non-flood season).

[0063] ② Flow rate change rate (Q c )

[0064] The main flow path is closely related to the magnitude of the flow rate. If the flow rate changes significantly, the flow path is unstable, and drastic river regime changes such as main channel swing and beach bank erosion and recession are likely to occur. Define the flow rate change rate: When predicting beach bank changes during the flood season, it is the ratio of the average flow rate during this flood season to the average flow rate of the previous non-flood season; when predicting beach bank changes during the non-flood season, it is the ratio of the average flow rate of the previous flood season to the average flow rate of this non-flood season. When the value of this characteristic variable is relatively large, the probability of beach bank erosion and recession increases. Its calculation formula is as follows:

[0065] Q c = Q f / Q d (2)

[0066] Where: Q f is the average flow rate (m 3 / s) during the flood season; Q d is the average flow rate (m 3 / s) during the non-flood season.

[0067] ③ Average nearshore water depth (H b )

[0068] The average nearshore water depth during the flood season (non-flood season) is the difference between the average cross-section water level and the elevation of the beach bank toe:

[0069]

[0070] Where: is the average water level (m) of the calculated cross-section during the flood season (non-flood season); Ze is the elevation of the toe (m). The toe is defined as the location where the riverbed slope changes significantly below, such as where the riverbed becomes flat or the sign of the slope changes, as Figure 2 shown.

[0071] ④ Average nearshore flow velocity (U b )

[0072] Assume that the nearshore roughness coefficient and slope are the same as the cross-section average values. According to the Manning formula, the average nearshore flow velocity can be approximately calculated through the cross-section average flow velocity. Its calculation formula is:

[0073]

[0074] Where: is the average depth of the cross-section, which is the ratio of the cross-sectional area below Z to the river width, i.e.,

[0075] Calculate the characteristic variables of the riverbed boundary, including the main channel and bank slope morphology, the position of the main stream, the width of the toe recession, and the influence of regulation projects. Among them, the main channel morphology includes the floodplain area, the floodplain river width, and the river facies coefficient, etc.; the bank slope morphology includes the height and slope of the beach bank; the position of the main stream is characterized by the relative distance of the thalweg from the bank. To improve the description accuracy of the riverbed boundary characteristics, the present invention proposes variables such as the distance of the toe from the bank, the relative elevation difference between the lowest point near the bank and the thalweg, and the relative elevation difference between the near-bank riverbeds on the same side and the opposite side. The distance of the toe from the bank is the width of the beach that buffers the erosion of the bank slope in front of the beach bank. The larger this distance, the relatively safer the bank slope; the relative elevation difference between the lowest point near the bank and the thalweg can indirectly reflect the intensity of the near-bank water flow. The smaller the elevation difference, the relatively greater the erosion intensity of the near-bank water flow; the relative elevation difference between the near-bank riverbeds on the same side and the opposite side reflects the comparison of the water flow concentration degree on both sides of the river bank. The water flow tends to concentrate on the bank side with a lower elevation, and the stability of the beach bank on this side is relatively lower. By analyzing a large number of measured data, it is found that these variables have a significant impact on the probability and amplitude of beach bank collapse and recession, so they are introduced in this method. In addition, the width of the toe recession is calculated according to the average residual shear stress; the bed sediment particle size at the predicted cross-section is obtained by linear interpolation according to the measured value at the hydrological cross-section; the influence of the regulation project on the beach bank collapse and recession is quantified by the distance of the project from the bank.

[0076] ① Beach bank morphology

[0077] The position of the toe is as Figure 1 shown. The calculation formulas for the beach bank height (BH) and slope (BS) are:

[0078] BH = Z b - Z f (5)

[0079] BS = arctan[(Z b - Z f ) / |X f - X b |] (6)

[0080] In the formula: X f , Z f are the distance from the starting point of the toe and the elevation (m) respectively; Z b is the elevation of the beach lip (m).

[0081] ② Relative distance of the main stream from the bank

[0082] The distance of the toe from the bank (EB) can be combined with the relative distance of the thalweg from the bank (TB) to approximately characterize the degree of the main stream close to the bank. Their calculation methods are as follows:

[0083] EB = |X e - Xb | (7)

[0084] TB = |X t -X b | / B bf (8)

[0085] In the formula: X t , X e , X b are the starting distances (m) of the thalweg, the nearshore scouring point, and the left or right bank toe respectively; B bf is the bank-full channel width (m).

[0086] ③ Relative elevation difference (LT) between the lowest point on the nearshore and the thalweg

[0087] As Figure 2 shown, the lowest point on the nearshore is the point with the lowest elevation within the nearshore bed surface, one on each of the left and right banks. Sometimes the lowest point on the nearshore on one bank side coincides with the thalweg point of the cross-section. The relative elevation difference between the lowest point on the nearshore and the thalweg can approximately characterize the nearshore water flow scouring intensity, and its calculation formula is as follows:

[0088] LT = (Z l - Z t ) / H bf (9)

[0089] In the formula: Z l , Z t are the elevations (m) of the lowest point on the nearshore and the thalweg respectively; H bf is the bank-full water depth (m).

[0090] ④ Relative elevation difference (△Z) between the nearshore river beds on the left and right banks

[0091] The water flow is generally concentrated on the side of the nearshore river bed with a lower elevation. The relative elevation difference between the nearshore river beds on the left and right banks is calculated to characterize the comparison of the water flow concentration degree on the left and right banks. The smaller the value of this variable, the greater the probability of the bank slope instability and collapse. Its calculation formula is as follows:

[0092]

[0093] In the formula: are the average elevations (m) within the nearshore river beds on the left and right banks respectively.

[0094] ⑤ Retreating width of the toe of the slope (ΔW)

[0095] The scouring of the water flow on the toe of the slope makes the bank slope steeper, creating conditions for the occurrence of bank collapse. The following formula is used to calculate the retreating width of the toe of the slope scoured by the water flow:

[0096]

[0097] In the formula: kd is the scour coefficient, m 3 / (Ns), related to the characteristics of the soil itself and the incipient shear stress, k d = 2×10 -7 τ c -0.5 ; λ1 is the scour index, generally taken as 1.0; τ c is the incipient shear stress during soil scour, N / m 2 ; τ f is the shear stress of the nearshore current, N / m 2 , and it is assumed that it is proportional to the water depth, that is, τ f = γ w hJ, where γ w is the unit weight of water, N / m 3 ; h is the nearshore water depth, m; J is the longitudinal water surface slope.

[0098] In the wandering section of the lower Yellow River, the pores of the beach bank soil are relatively large and the beach bank soil is not yet fully compacted. The incipient shear stress of the soil is calculated by the following formula:

[0099] τ c = 6.68×10 2 ×d + 3.67×10 -6 / d (12)

[0100] In the formula: d is the soil particle size, m, and the median particle size of the bed sediment is approximately used to replace the median particle size of the soil at the toe of the slope.

[0101] ⑥ Influence of regulation works

[0102] The influence of regulation works is reflected by calculating the distance of the project from the shore (WB), which is the distance from the beach lip to the project, and the retreat width of the beach bank will not exceed this value.

[0103] WB = |X b - X w | (13)

[0104] In the formula: X w is the starting distance of the project (m).

[0105] Calculate the characteristic variables of the riverbed deformation in the early stage. Considering the lag of riverbed evolution, in addition to the current riverbed boundary conditions, this method supplements the calculation of the key variables of the previous riverbed evolution, including the swing width and scouring and silting thickness of the thalweg in the previous stage, the change of the beach bank slope, and the deformation width of the beach bank, etc. The calculation formulas are as follows:

[0106] ① Swing width of the thalweg in the previous stage (ΔB t )

[0107] ΔB t = r×|X t - X't | (14)

[0108] Wherein: When predicting the beach and bank changes during the flood season, X t , X' t are respectively the distances from the starting point of the thalweg before the flood of the current year and after the flood of the previous year (m); when predicting the beach and bank changes during the non-flood season, X t , X' t are respectively the distances from the starting point of the thalweg after the flood and before the flood of the current year (m); r is the control symbol parameter. If the thalweg swings towards the bank in the previous period, r = -1; if it swings towards the center of the river, r = 1.

[0109] ② The average scouring and silting thickness (ΔZ t )

[0110]

[0111]

[0112] Wherein: When predicting the beach and bank changes during the flood season, Z t , Z' t , Z” t are respectively the thalweg elevations before the flood of the current year, after the flood of the previous year, and before the flood of the previous year (m); when predicting the beach and bank changes during the non-flood season, Z t , Z' t , Z” t are respectively the thalweg elevations after the flood, before the flood of the current year, and after the flood of the previous year (m); r is the control symbol parameter. When the thalweg scours, r = -1; when it silts, r = 1.

[0113] ③ The change in beach and bank slope (△S) in the previous stage 1

[0114] ΔS = BS' - BS (17)

[0115] Wherein: When predicting the beach and bank changes during the flood season, BS and BS’ are respectively the beach and bank slopes before the flood of the current year and after the flood of the previous year; when predicting the beach and bank changes during the non-flood season, BS and BS’ are respectively the beach and bank slopes after the flood and before the flood of the current year. If the bank slope becomes steeper, △S < 0; otherwise, △S ≥ 0.

[0116] Step 3. Preprocess the calculated characteristic variables, construct a beach and bank recession probability prediction model and a beach and bank recession width prediction model, and train the beach and bank recession probability prediction model based on the preprocessed characteristic variables and the beach and bank recession probability target variable, and train the beach and bank recession width prediction model based on the preprocessed characteristic variables and the beach and bank recession width target variable;

[0117] The preprocessing of the calculated characteristic variables includes:

[0118] Based on the IterativeImputer method, missing or abnormal feature values are interpolated; limited by the accuracy of topographic measurement and water and sediment interpolation, some feature values are missing or have abnormal values. To deal with this situation, this method uses IterativeImputer for multiple interpolation.

[0119] The IterativeImputer method is different from simple mean and mode interpolation. It uses the regression model to predict missing values through iteration, making the interpolation results more accurate. For each missing value, IterativeImputer uses other features as input to train the regression model to predict the missing value. The initial interpolation usually adopts a simple strategy (such as mean interpolation). Then multiple iterations are performed. In each iteration, the interpolated features are used to predict the missing values of other features until convergence or the set number of iterations is reached. Its advantage is that it can use the relationship between all features to improve the accuracy of interpolation.

[0120] Test the autocorrelation of characteristic variables and remove highly autocorrelated features; high correlation between features may lead to multicollinearity, affecting the stability and accuracy of the model. Although the characteristic variables in step 3 have different physical meanings, they may still show high correlation in the data. For example, when the interannual fluctuations in water and sediment conditions are not large, the water and sediment characteristic parameters of the current year may be correlated with the average water and sediment characteristic parameters of the previous years; when the deep well is always close to a certain shore, the relative distance of the deep well from the shore may be correlated with the distance of the nearshore scouring point from the shore. Therefore, it is necessary to perform a feature correlation test and remove highly autocorrelated features. Specifically, first calculate the correlation coefficient matrix of all features; secondly, set a threshold (0.9), and the feature pairs above this threshold are considered to be highly correlated; finally, delete one of the features with a high correlation.

[0121] Based on the Yeo-Johnson transformation, the characteristic variables are debiased. In actual data, the characteristic variables may have a large skewness, and the data distribution may tilt to the left or right, and cannot present a symmetrical normal distribution, such as Figure 4 As shown in a. This skewness may have an adverse effect on the training of the model. Debiasing reduces skewness by transforming the data, making the data distribution closer to the normal distribution, reducing the impact of outliers, and helping the model to more accurately capture the relationship between features and target variables. Considering that the characteristic variables in the prediction of beach collapse are both positive and negative, this method uses Yeo-Johnson transformation to debias the features.

[0122] Convert the feature variable data into a distribution with a mean of 0 and a standard deviation of 1. The standardization process can balance the influence of each feature, enabling the model to utilize all feature information more fairly. If the features are not standardized, due to the large differences in their numerical ranges, it may affect the similarity measurement or distance calculation of the model, thereby reducing the performance of the model. Figure 3 Taking the feature variable of the river regime coefficient as an example, the data distributions before and after debiasing and standardization are shown. Before processing, the range of the river regime coefficient was 9.83 - 50.63, and the data was mainly concentrated in the range of 10 - 15, with an obvious left skew in the distribution; after processing, the variable range was -1.62 - 1.82, and the data distribution was more uniform.

[0123] Randomly or in a certain order, divide the preprocessed feature variables and the target variable of the bank collapse and recession probability into a training set, a validation set, and a test set, which serve as the input data for the bank collapse and recession probability prediction model.

[0124] According to the same format, take the preprocessed feature variables and the target variable of the bank collapse and recession width as the input data for the bank collapse and recession width prediction model.

[0125] In this embodiment, a bank collapse and recession prediction model based on the K-Nearest Neighbors (KNN) algorithm is constructed, and Lasso regression is innovatively embedded in the model for feature selection. Specifically, first, use Lasso regression to screen the candidate feature variables and eliminate the low-correlation features; then input the screened feature set into the KNN model for training to improve the prediction accuracy and generalization performance of the model. During the model optimization process, jointly tune the number of feature variables and other hyperparameters of the KNN algorithm, and determine the hyperparameter combination that enables the model to achieve the optimal performance on both the training set and the validation set through grid search. Based on historical data, train the bank collapse and recession probability prediction model and the bank collapse and recession width prediction model respectively.

[0126] The parameter optimization basis of the bank collapse and recession probability prediction model is as Figure 4 shown. By plotting the learning curves of the model accuracy and the number of neighboring points under different numbers of features, generate corresponding analysis diagrams for each number of features. The parameter optimization basis of the bank collapse and recession width prediction model is as Figure 5 shown. The model respectively lists the feature variables actually used by the model under different numbers of features in the training set and the validation set, as well as the goodness-of-fit indicators corresponding to different numbers of neighboring points. Based on the above results, the optimal number of features and the number of neighboring points can be selected to ensure that the model has both high prediction accuracy and generalization ability.

[0127] Step 4. Evaluate the accuracy of the bank collapse and recession probability prediction model and the bank collapse and recession width prediction model;

[0128] Evaluate the accuracy of the model on the test set. If the test set accuracy is also relatively high, it can be considered that the trained model has a certain generalization ability, and its prediction results have certain reference value, and the model can be applied for prediction (Step 5); if it does not meet the requirements, the model needs to be retrained (Step 4).

[0129] Step 5. Based on the beach bank recession probability prediction model and the beach bank recession width prediction model, respectively predict the beach bank recession probability and the beach bank recession width in the future period.

[0130] Input the sample data to be predicted, and based on the model trained in Step 5, predict the beach bank recession probability and the recession width under specific conditions. This method first applies the recession probability prediction model to output the recession probability (0-100%) and the prediction category (collapse or non-collapse) of each prediction sample, as Figure 6 shown. If the prediction result is collapse, then based on the recession width prediction model, output the recession width of each prediction sample, as Figure 7 shown.

[0131] Embodiment 2

[0132] This embodiment provides a beach bank recession prediction system based on a kinetic mechanism and machine learning, including:

[0133] A data collection module, which is used to collect the historical basic data of the river section where the prediction section is located;

[0134] A variable calculation module, which is used to screen the basic data and calculate the characteristic variables, the beach bank recession probability target variable and the beach bank recession width target variable based on the screened basic data;

[0135] A model construction module, which is used to preprocess the calculated characteristic variables, construct a beach bank recession probability prediction model and a beach bank recession width prediction model, and train the beach bank recession probability prediction model based on the preprocessed characteristic variables and the beach bank recession probability target variable, and train the beach bank recession width prediction model based on the preprocessed characteristic variables and the beach bank recession width target variable;

[0136] An accuracy evaluation module, which is used to evaluate the accuracy of the beach bank recession probability prediction model and the beach bank recession width prediction model;

[0137] A prediction module, which is used to predict the beach bank recession probability and the beach bank recession width in the future period based on the beach bank recession probability prediction model and the beach bank recession width prediction model respectively.

[0138] As described above, it is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application.

[0139] It should be understood that the parts not elaborated in detail in this specification all belong to the prior art.

[0140] It should be understood that the above description of the preferred embodiment is relatively detailed, and it should not be considered as a limitation to the protection scope of the present invention patent. Under the inspiration of the present invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the scope protected by the claims of the present invention, and all fall within the protection scope of the present invention. The scope of protection claimed in the present invention shall be subject to the appended claims.

Claims

1. A beach bank recession prediction method based on a kinetic mechanism and machine learning, characterized in that The steps include: Step 1. Collect basic historical data of the river section where the prediction section is located; Step 2. Screening the basic data and calculating the characteristic variables, the beach collapse probability target variable and the beach collapse width target variable based on the screened basic data; Step 3. Preprocess the calculated characteristic variables, construct a beach collapse probability prediction model and a beach collapse width prediction model, and train the beach collapse probability prediction model based on the preprocessed characteristic variables and the beach collapse probability target variable, and train the beach collapse width prediction model based on the preprocessed characteristic variables and the beach collapse width target variable; Step 4. Evaluate the accuracy of the beach collapse probability prediction model and the beach collapse width prediction model; Step 5. Based on the beach collapse probability prediction model and the beach collapse width prediction model, the beach collapse probability and the beach collapse width in the future period are predicted respectively.

2. The beach bank recession prediction method based on the kinetic mechanism and machine learning according to claim 1, characterized in that The historical basic data in step 1 include the flow, water level, sediment content and median particle size of bed sand of the hydrological station in the river section where the prediction section is located in the historical period, as well as the measured topographic data of the prediction section.

3. The beach bank recession prediction method based on the dynamic mechanism and machine learning according to claim 1, wherein The target variable of the probability of beach collapse in step 2 is whether the beach is deformed and collapsed. The beach collapse samples are marked as 1, and the stable or silted samples are marked as 0.

4. The beach bank recession prediction method based on the dynamic mechanism and machine learning according to claim 1, wherein The target variable of the beach collapse width is: beach deformation width. The cross-sectional topography of adjacent measurements of the predicted section is arbitrarily predicted to determine the starting distances of the beach lips on the left and right sides of the beach. The difference in the starting distances of the beach lips between adjacent measurements is calculated as the beach deformation width.

5. The beach bank recession prediction method based on the dynamic mechanism and machine learning according to claim 1, characterized in that, The characteristic variables include: water and sand condition characteristic variables, riverbed boundary characteristic variables, and early riverbed deformation characteristic variables.

6. The beach bank recession prediction method based on the kinetic mechanism and machine learning according to claim 4, characterized in that, The characteristic variables of water and sediment conditions include: average flow, average sediment content, flow variation coefficient, average water flow scouring intensity in the future period and the previous three periods, average flow velocity in the section, flow change rate, average water depth near the shore and average flow velocity near the shore; The riverbed boundary characteristic variables include: flat beach area, flat beach river width and river phase coefficient; beach bank height and slope; relative distance between the deep well and the shore, distance between the slope foot and the shore, relative elevation difference between the lowest point near the shore and the deep well, and relative elevation difference between the riverbed near the shore and the opposite side; width of the slope foot retreat; median particle size of the bed sand; distance from the shore to the regulation project; The variables of the early riverbed deformation characteristics include: the early deep channel swing width and scouring thickness, the change of beach slope and the width of beach deformation.

7. The beach bank recession prediction method based on the dynamic mechanism and machine learning according to claim 1, characterized in that The pretreatment in step 3 includes: Based on the IterativeImputer method, missing values or abnormal values of features are interpolated; Test the autocorrelation of characteristic variables and remove highly autocorrelated features; Based on the Yeo-Johnson transformation, the characteristic variables are debiased; Convert the feature variable data into a distribution with a mean of 0 and a standard deviation of 1; The pre-processed characteristic variables and the target variable of beach collapse probability are randomly or in a certain order divided into a training set, a validation set and a test set as input data of the beach collapse probability prediction model; The preprocessed characteristic variables and the target variable of beach collapse width are used as the input data of the beach collapse width prediction model in the same format.

8. The beach bank recession prediction method based on the kinetic mechanism and machine learning according to claim 1, characterized in that, In step 3, the specific construction of the beach bank recession probability prediction model and the beach bank recession width prediction model is as follows: First, use Lasso regression to screen candidate feature variables and eliminate low-correlation features. Subsequently, input the screened feature set into the KNN model for training, jointly tune the number of feature variables and other hyperparameters of the KNN algorithm, and determine the hyperparameter combination that enables the model to achieve optimal performance on both the training set and the validation set through grid search. Train the beach bank recession probability prediction model and the recession width prediction model respectively based on the historical basic data.

9. The beach bank recession prediction method based on the kinetic mechanism and machine learning according to claim 1, characterized in that Step 5 includes: Based on the trained beach bank recession probability prediction model, predict the beach bank recession probability under specific conditions. If the probability prediction model predicts it as a recession, then further predict its recession width based on the beach bank recession width prediction model.

10. A beach bank recession prediction system based on a kinetic mechanism and machine learning, characterized in that, It includes: A data collection module, which is used to collect the historical basic data of the river section where the prediction section is located. A variable calculation module, which is used to screen the basic data and calculate the feature variables, the beach bank recession probability target variable, and the beach bank recession width target variable based on the screened basic data. A model construction module, which is used to preprocess the calculated feature variables, construct the beach bank recession probability prediction model and the beach bank recession width prediction model, and train the beach bank recession probability prediction model based on the preprocessed feature variables and the beach bank recession probability target variable, and train the beach bank recession width prediction model based on the preprocessed feature variables and the beach bank recession width target variable. An accuracy evaluation module, which is used to evaluate the accuracy of the beach bank recession probability prediction model and the beach bank recession width prediction model. A prediction module, which is used to predict the beach bank recession probability and the beach bank recession width in the future period respectively based on the beach bank recession probability prediction model and the beach bank recession width prediction model. The beach bank recession prediction system based on the dynamic mechanism and machine learning is used to execute the steps in the beach bank recession prediction method based on the dynamic mechanism and machine learning according to any one of claims 1-9.

Citation Information

Patent Citations

  • River bank collapse early warning method and device based on multi-source data fusion

    CN115293241A

  • Bank collapse monitoring and early warning evaluation method based on stable slope ratio calculation

    CN118965688A

  • Mountainous area slope displacement prediction method based on mi-GRA and improved PSO-lstm

    WO2024001942A1

  • Method and apparatus for simulating rapid transport of sediment during reservoir discharging

    WO2024098951A1