River short-term water flow prediction method based on Bayesian optimization CNN-LSTM-Att
By optimizing the CNN-LSTM-Att model using Bayesian methods and combining CNN, LSTM, and self-attention mechanisms, the problems of insufficient hydrological data and hyperparameter configuration in traditional models are solved. This achieves high efficiency and accuracy in short-term river flow prediction, reduces errors, and improves model adaptability.
Patent Information
- Application Number
- CN202510979300.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional short-term river flow prediction models rely on specific hydrological data and simplified assumptions. When accurate data is lacking, the prediction accuracy is limited, and the hyperparameter configuration of hybrid models is difficult to optimize, resulting in insufficient generalization ability and poor local feature learning ability.
A Bayesian-optimized CNN-LSTM-Att model was constructed by combining a convolutional neural network (CNN), a long short-term memory model (LSTM), and a self-attention mechanism. The hyperparameters were tuned using the Bayesian optimization algorithm, and the model was trained using self-collected river flow data.
It significantly improves the accuracy of short-term water flow forecasting, reduces information complexity, enhances the ability to capture nonlinear trends, improves forecasting performance, reduces mean absolute error and root mean square error compared to other models, and increases the coefficient of determination.
Smart Images

Figure CN120952113A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of river flow prediction and deep learning, and specifically relates to a method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-Att. Background Technology
[0002] Global warming is exacerbating the frequency of extreme hydrological events, leading to an increasingly severe socioeconomic burden. The number of people threatened by floods each year is also steadily rising, and this trend is expected to intensify. Predicting short-term river flow allows for early warnings, providing downstream residents, cities, and farmland with time for evacuation and protection. Accurate river flow forecasting serves as a bridge between natural science and societal needs; its significance extends beyond its tangible value in disaster prevention and mitigation, permeating the overall planning for sustainable water resource utilization, ecological protection, and socioeconomic development.
[0003] Traditional short-term river flow forecasting relies primarily on physical models, requiring specific assumptions and extensive hydrological data for calibration. However, in practical applications, accurate hydrological data, such as rainfall, evaporation, and watershed characteristics, is often lacking, and the model is limited by reliance on simplifying assumptions and complex structures. With the development of artificial intelligence, data-driven models, due to their relatively low data requirements and strong versatility, can quickly simulate linear and nonlinear relationships between flow data, providing solutions for flow forecasting lacking calibration data. In some cases, their predictive performance even surpasses that of physical models, leading to their increasing widespread application in short-term water flow forecasting.
[0004] Studies have shown that deep learning models such as convolutional neural networks (CNN) and long short-term memory models (LSTM) are superior to machine learning (Malihe et al. Water Resources Management, 2024:1-20). However, the single LSTM model has limited prediction accuracy due to insufficient generalization ability and lack of local feature learning ability (Hu Leyi et al. Lake Science, 2024, 36(04):1241-1252). By integrating different models, multi-faceted feature extraction of time-series data is achieved, effectively improving the model's adaptability to complex dynamic systems and enhancing its ability to capture different time scales: the combination of LSTM and spatial recursion (SR) improves the accuracy of flow prediction in large watersheds (Yu et al. Hydrology and Earth System Sciences, 2024, 28(9): 2107-2122); the LSTM-based encoder-decoder structure (En-De-LSTM) decomposes multi-frequency components in hydrological sequences and is effectively applied to flow prediction at the Hankou hydrological station on the Yangtze River; the TimeGAN-LSTM model realizes monthly runoff prediction in undefined river basins (Miao Lei et al. Hydropower Energy Science, 2024, 42(11): 12-15). However, the hyperparameter configuration of mixed complex models plays a crucial role in model performance, and traditional trial-and-error methods are difficult to meet optimization requirements and cannot unleash the model's potential. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention proposes a short-term river flow prediction method based on Bayesian optimization CNN-LSTM-Att. This method integrates CNN, LSTM, and self-attention mechanisms, while introducing a Bayesian optimization algorithm to fine-tune hyperparameters, constructing a coupled BO-CNN-LSTM-Att model. This model is trained using field-collected data, efficiently completing the prediction task.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0007] A method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-ATT includes the following steps:
[0008] Step 1: Obtain the original water flow dataset for the target river channel during historical time periods, and preprocess the obtained original water flow dataset, dividing the preprocessed water flow dataset into training set and test set.
[0009] Step 2: Construct the CNN-LSTM-ATt model structure, including a self-attention mechanism layer and a fully connected layer, as well as a feature extraction layer, a self-attention mechanism layer, and a fully connected layer composed of CNN layers and LSTM layers connected in series. The CNN layer extracts features from the preprocessed water flow data and inputs the features extracted by the CNN layer into the LSTM layer, which captures the temporal features of the input features. The self-attention mechanism layer assigns weights to the temporal features output by the LSTM layer, and finally, the predicted short-term water flow data is output through the fully connected layer.
[0010] Step 3: Use the Bayesian algorithm to optimize the relevant hyperparameters of the feature extraction layer, and apply the optimized hyperparameter combination to the CNN-LSTM-Att model structure to obtain the Bayesian optimized CNN-LSTM-Att model structure, namely the BO-CNN-LSTM-Att model structure.
[0011] Step 4: Based on the training set, using the preprocessed historical water flow data of the target river as input and the water flow data of the target river within a preset time period after the input time period as output, train the constructed BO-CNN-LSTM-Att model structure to obtain the trained BO-CNN-LSTM-Att model; based on the test set, apply the BO-CNN-LSTM-Att model to obtain the predicted water flow data of the target river within a preset time period in the future to evaluate the model.
[0012] Step 5: Using the water flow data of the target river channel during the target time period as input, apply the BO-CNN-LSTM-ATt model to obtain the predicted water flow data of the target river channel during the future preset time period.
[0013] Furthermore, the process of obtaining the raw water flow dataset for historical time periods of the target river channel in step 1 includes:
[0014] S1. A self-made, remotely controlled, transceiver-integrated acoustic tomography system is deployed in staggered pairs on both banks of the target river channel. The acoustic tomography system includes a control system deployed at a site on the shore and underwater transducers fixed at a preset depth underwater. The control system and the corresponding underwater transducers are connected by a waterproof cable, and the horizontal baseline between the sites of the paired acoustic tomography systems on both banks forms a 45° angle with the direction of water flow. The reciprocity time of the acoustic waves is obtained by the underwater transducers corresponding to the paired acoustic tomography systems on both banks emitting sound waves to each other. Combined with X-ray acoustic correction and Doppler effect calibration, the river cross-sectional velocity is obtained. The water level is measured in real time using a synchronous water level gauge. The river cross-sectional area is calculated based on the water level. The water flow rate is obtained by multiplying the river cross-sectional area by the river cross-sectional velocity.
[0015] S2. Collect water flow data according to step S1 within a preset time interval during the historical period to form the original water flow dataset.
[0016] Furthermore, the raw water flow dataset obtained in step 1 is preprocessed, including normalization and tensor transformation, specifically as follows:
[0017] S1. The original water flow data is mapped to the interval [0, 1] by normalization. The normalization formula is:
[0018]
[0019] In the formula: l is the index of the original water flow data, x l This represents the l-th raw water flow data value. x represents the normalized value of the l-th raw water flow data. max x represents the maximum value in the original water flow dataset; min This represents the minimum value in the original water flow dataset;
[0020] S2. Perform tensor transformation on the normalized water flow data to meet the data input requirements of the CNN layer;
[0021] S3. Divide the preprocessed water flow dataset into training and test sets according to time order, with a ratio of 80% and 20%.
[0022] Furthermore, in step 2, the CNN layer consists of convolutional layers and pooling layers. The convolutional layer extracts local features through one-dimensional convolution, and the convolution formula is as follows:
[0023]
[0024] Where j is the data index of the output sequence, y j The j-th data point in the output sequence; k is the index of the weight parameters inside the convolution kernel, w k x is the k-th weight parameter of the convolution kernel, where K is the kernel size; j+k-1 For y j The corresponding convolution kernel window covers the k-th input data; after the convolution operation, the convolution result is processed by the non-linear activation function ReLU, and then a max pooling layer is used to achieve dimensionality reduction;
[0025] The pooled features are input into an LSTM layer. The neurons in the LSTM layer contain a forget gate, an input gate, an output gate, and a memory unit. The current input and the hidden state from the previous time step are input in parallel into the forget gate, input gate, output gate, and memory unit. The outputs of the forget gate and the input gate are proportionally summed to obtain the current memory unit state. The output gate controls the output ratio of the current memory unit state to obtain the final hidden state output. Specifically:
[0026] The forgetting gate determines the ratio of information retained to forgotten in memory units from the previous time step, and its formula is:
[0027] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0028] t is the current time step, f t For the output of the forget gate, W f Let b be the weight matrix of the forget gate. f For the bias of the forget gate, h t-1 The hidden state of the previous time step, x t Given the current input, σ() is the Sigmoid activation function;
[0029] The input gate determines the proportion of input information written to the memory unit at the current time step, and its formula is:
[0030] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0031]
[0032] i t The output of the input gate, For candidate memory values, W i and W c Here, are the weight matrices for the input gate and candidate memory, respectively; tanh is the hyperbolic tangent activation function; b i and b c These are the biases for the input gate and candidate memory, respectively;
[0033] The state of the current memory cell is updated based on the forget gate and the input gate, and its expression is:
[0034]
[0035] The output f of the forget gate t Controlling old memories C t-1 The retention of the input gate's output i t Controlling candidate memories The addition;
[0036] The output gate controls the output ratio of the current memory cell state, thereby obtaining the final output of the hidden state. The expressions are as follows:
[0037] o t=σ(W o ·[h t-1 ,x t ]+b o )
[0038] h t =o t ×tanh(C t )
[0039] o t W is the output of the output gate. o Let b be the weight matrix of the output gate. o h is the bias of the output gate. t This represents the hidden layer state at the current time step, i.e., the output of the LSTM layer.
[0040] The output of the LSTM layer is input into the self-attention mechanism layer, and the self-attention mechanism assigns weights to the temporal features output by the LSTM layer.
[0041] Finally, the prediction result is output through a fully connected layer.
[0042] Furthermore, the specific steps for the self-attention mechanism layer to assign weights to the hidden layer state features output by the LSTM layer in step 2 are as follows:
[0043] S1. Perform a linear transformation on the input hidden layer feature data, mapping it to a query, a key, and a value. The transformation formula is as follows:
[0044]
[0045] Where Q is the query matrix, K is the key matrix, V is the value matrix, and W is the value matrix. q W k W v These are the weight matrices for the query matrix, key matrix, and value matrix, respectively, and X is the input vector;
[0046] S2. Calculate the dot product of the query and each key to obtain the attention score, using the formula:
[0047]
[0048] d is the scaling factor, in the denominator Used to scale dot products and stabilize gradient flow;
[0049] S3. Normalize the attention score using the Softmax function to obtain the attention weights, the expression of which is:
[0050] weights = Softmax(scores)
[0051] S4. Calculate the weighted sum of V using the calculated attention weights to obtain the final attention output, expressed as:
[0052] output = weights·V
[0053] The output of the last time step of the attention mechanism calculation is taken and then connected to a fully connected layer to output the prediction result.
[0054] Furthermore, the specific steps for optimizing the relevant hyperparameters of the feature extraction layer using the Bayesian algorithm in step 3 are as follows:
[0055] S1. Select a preset combination of hyperparameters, including CNN kernel size, number of neurons in LSTM hidden layer, number of LSTM hidden layer layers, and learning rate, and define a search space for each hyperparameter.
[0056] S2. Define the objective function, including initializing the CNN-LSTM-ATt model structure according to preset hyperparameters, training the entire model structure using the training set for a preset number of iterations, calculating the loss for each iteration, iteratively optimizing the model parameters based on backpropagation, and finally calculating the average loss on the test set as the objective function value.
[0057] S3. Optimization, the specific process is as follows:
[0058] S3.1. Based on the preset random search points for each hyperparameter, randomly select multiple hyperparameter combinations, train the model structure, obtain the objective function values corresponding to each hyperparameter combination, and form an observation dataset.
[0059] S3.2. Based on the observed dataset, a Gaussian regression process is used as a surrogate model to fit the objective function, generating a probability distribution for predicting the value of the objective function; the Gaussian regression prediction formula is:
[0060]
[0061] a * Let μ(a) be the combination of hyperparameters to be predicted. * ) represents the hyperparameter combination a * The predicted value of the objective function is: cov(a) * ) represents the hyperparameter combination a * Covariance under the following conditions For the current hyperparameter combination a * C is a vector formed by the covariances of the various hyperparameter combinations in the observed dataset. ov To form the matrix of covariances among various hyperparameter combinations in the observation dataset, c ov For hyperparameter combination a * Its own covariance, Let be the variance term of the noise, n be the number of samples in the observation dataset, I be the identity matrix, and b be the vector of objective function values of each hyperparameter combination in the observation dataset.
[0062] S3.3. Based on the Gaussian regression process, the hyperparameter sample points are sampled by improving the acquisition function to select the next hyperparameter combination; the formula for the expected improvement is:
[0063] EI(x)=E[max(f(x)-f best ,0)]
[0064] EI(x) represents the desired improvement in the acquisition function, and f(x) is the predicted value of the objective function, given by a Gaussian process. best E is the currently known optimal value, and E is the expected value.
[0065] S3.4 Update the surrogate model. Repeat steps S3.2 and S3.3 within a preset number of iterations until the combination of hyperparameters corresponding to the lowest test set loss value is found. This combination is then used as the optimal hyperparameter combination and output.
[0066] Furthermore, the optimized hyperparameter combination of the Bayesian algorithm is as follows: the size of the CNN layer convolution kernel is 2, the number of neurons in the LSTM hidden layer is 200, the number of LSTM hidden layers is 10, and the learning rate is 0.008.
[0067] Furthermore, the hyperparameter combination optimized by the Bayesian algorithm is applied to the CNN-LSTM-Att model structure to obtain the BO-CNN-LSTM-Att model structure; based on the training set in step 1, the BO-CNN-LSTM-Att model structure is trained, and Adam is used as the training optimizer.
[0068] Furthermore, the mean absolute error (MAE), root mean square error (RMSE), and coefficient of determination (R²) were used. 2 The prediction performance of the BO-CNN-LSTM-ATt model is evaluated using the following formulas:
[0069]
[0070]
[0071] Where m is the number of samples in the original water flow dataset, and l∈[1,m], x l For the l-th raw water flow data value, This is the predicted water flow rate. The mean is the sample value; MAE reflects the model's average flow prediction error; RMSE is more sensitive to some extreme errors, and the combination of the two errors better reflects the accuracy of the model's water flow prediction; R 2 This reflects the model's fit; the closer the value is to 1, the better the model's predictive performance.
[0072] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0073] This invention integrates CNN, self-attention mechanism, and Bayesian optimization algorithm on the basis of LSTM model. It trains the model using self-collected river flow dataset to achieve accurate prediction of short-term water flow. It uses CNN convolutional neural network to extract local features of time series data, reducing information complexity and preventing overfitting. The integration of self-attention mechanism provides the model with stronger feature extraction capabilities and effectively captures nonlinear changing trends. For hyperparameter optimization, Bayesian optimization algorithm is added to improve the efficiency of hyperparameter combination adjustment and significantly improve prediction performance.
[0074] A comparison of the prediction performance of the BO-CNN-LSTM-Att model of this invention with LSTM, CNN-LSTM, and CNN-LSTM-Att models reveals that the BO-CNN-LSTM-Att model of this invention has the smallest mean absolute error (MAE) of 35.74, root mean square error (RMSE) of 44.96, and a high coefficient of determination (CDO) of 0.9728. Compared to LSTM, RMSE and MAE are reduced by 51.55% and 53.56%, respectively; compared to CNN-LSTM, RMSE and MAE are reduced by 33.35% and 35.13%, respectively; and compared to CNN-LSTM-Att, RMSE and MAE are reduced by 24.53% and 24.92%, respectively. Attached Figure Description
[0075] Figure 1 This is a flowchart of the BO-CNN-LSTM-ATt joint model algorithm of the present invention;
[0076] Figure 2 This is a schematic diagram of the acoustic tomography system layout of the present invention;
[0077] Figure 3 This is the original river flow data for this invention;
[0078] Figure 4 This is a schematic diagram of a classic CNN network structure;
[0079] Figure 5 This is a structural diagram of the basic LSTM model of the present invention;
[0080] Figure 6 This is a flowchart of the self-attention mechanism algorithm of the present invention;
[0081] Figure 7 This is the hyperparameter iterative optimization process of the present invention;
[0082] Figure 8 This is a comparison chart of the true and predicted values of the BO-CNN-LSTM-Att model of the present invention with those of LSTM, CNN-LSTM, and CNN-LSTM-Att.
[0083] Figure 9 This is a comparative analysis of the RMSE, MAE, and R2 of the BO-CNN-LSTM-Att model of the present invention with LSTM, CNN-LSTM, and CNN-LSTM-Att. Detailed Implementation
[0084] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0085] like Figure 1 As shown, the present invention provides a method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-ATt, comprising the following steps:
[0086] Step 1: Obtain the original water flow dataset for the target river channel during historical time periods, and preprocess the obtained original water flow dataset, dividing the preprocessed water flow dataset into training set and test set.
[0087] Step 2: Construct the CNN-LSTM-ATt model structure, including a self-attention mechanism layer and a fully connected layer, as well as a feature extraction layer composed of CNN layers and LSTM layers connected in series. The CNN layer extracts features from the preprocessed water flow data and inputs the features extracted by the CNN layer into the LSTM layer, which captures the temporal features of the input features. The self-attention mechanism layer assigns weights to the temporal features output by the LSTM layer, and finally outputs the predicted short-term water flow data through the fully connected layer.
[0088] Step 3: Use the Bayesian algorithm to optimize the relevant hyperparameters of the feature extraction layer, obtain the best hyperparameter combination, and obtain the Bayesian-optimized CNN-LSTM-Att model structure, namely the BO-CNN-LSTM-Att coupled model structure.
[0089] Step 4: Based on the training set, using the preprocessed historical water flow data of the target river as input and the water flow data of the target river within a preset time period after the input time period as output, train the constructed BO-CNN-LSTM-Att model structure to obtain the trained BO-CNN-LSTM-Att model; based on the test set, apply the BO-CNN-LSTM-Att model to obtain the predicted water flow data of the target river within a preset time period in the future to evaluate the model.
[0090] Step 5: Using the water flow data of the target river channel during the target time period as input, apply the BO-CNN-LSTM-ATt model to obtain the predicted water flow data of the target river channel during the future preset time period.
[0091] Step 1 involves obtaining the raw water flow dataset for historical time periods of the target river channel, which includes:
[0092] S1, such as Figure 2 As shown, a self-made, remotely controlled, transceiver-integrated acoustic tomography system is deployed in staggered pairs on both banks of the target river channel. The acoustic tomography system includes a control system deployed at onshore stations and underwater transducers fixed at a predetermined depth underwater. The control system and the corresponding underwater transducers are connected via waterproof cables, and the horizontal baseline between the paired acoustic tomography systems on both banks forms a 45° angle with the direction of water flow. The reciprocity time of sound waves is obtained by the underwater transducers of the paired acoustic tomography systems on both banks emitting sound waves to each other. Combined with X-ray acoustic correction and Doppler effect calibration, the river cross-sectional velocity is obtained. A synchronous water level gauge is used to measure the water level in real time. The river cross-sectional area is calculated based on the water level, and the flow rate is obtained by multiplying the river cross-sectional area by the river cross-sectional velocity.
[0093] S2. Collect water flow data according to step S1 within a preset time interval during the historical period to form the original water flow dataset.
[0094] The raw water flow dataset obtained in step 1 is preprocessed, including normalization and tensor transformation, specifically as follows:
[0095] S1. Normalize the traffic data to the interval [0, 1]. The normalization formula is:
[0096]
[0097] In the formula: l is the index of the original water flow data, x l This represents the l-th raw water flow data value. x represents the normalized value of the l-th raw water flow data. max x represents the maximum value in the original water flow dataset; minThis represents the minimum value in the original water flow dataset;
[0098] S2. Perform tensor transformation on the normalized water flow data to meet the data input requirements of the CNN layer;
[0099] S3, such as Figure 3 As shown, the preprocessed water flow dataset is divided into training and test sets in chronological order at a ratio of 80% and 20%.
[0100] like Figure 4-6 As shown, the CNN layer in step 2 consists of a convolutional layer and a pooling layer. The convolutional layer extracts local features through one-dimensional convolution, and the convolution formula is:
[0101]
[0102] Where j is the data index of the output sequence, y j The j-th data point in the output sequence; k is the index of the weight parameters inside the convolution kernel, w k x is the k-th weight parameter of the convolution kernel, where K is the kernel size; j+k-1 For y j The corresponding convolution kernel window covers the k-th input data; after the convolution operation, the convolution result is processed by the non-linear activation function ReLU, and then a max pooling layer is used to achieve dimensionality reduction;
[0103] After pooling, the features are input into the LSTM layer. The neurons in the LSTM layer contain a forget gate, an input gate, an output gate, and a memory unit. The current input and the hidden state from the previous time step are input to the forget gate, input gate, output gate, and memory unit in parallel. The outputs of the forget gate and the input gate are proportionally summed to obtain the current memory unit state. The output gate obtains the final hidden state output by controlling the output ratio of the current memory unit state. Specifically:
[0104] The forgetting gate determines the ratio of information retained to forgotten in memory units from the previous time step, and its formula is:
[0105] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0106] t is the current time step, f t For the output of the forget gate, W f Let b be the weight matrix of the forget gate. f For the bias of the forget gate, h t-1 The hidden state of the previous time step, x t Given the current input, σ() is the Sigmoid activation function;
[0107] The input gate determines the proportion of input information written to the memory unit at the current time step, and its formula is:
[0108] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0109]
[0110] i t The output of the input gate, For candidate memory values, W i and W c Here, are the weight matrices for the input gate and candidate memory, respectively; tanh is the hyperbolic tangent activation function; b i and b c These are the biases for the input gate and candidate memory, respectively;
[0111] The state of the memory unit is updated according to the forget gate and the input gate, and its expression is:
[0112]
[0113] The output f of the forget gate t Controlling old memories C t-1 The retention of the input gate's output i t Controlling candidate memories The addition;
[0114] The output gate controls the output ratio of the current memory cell state, thereby obtaining the final output of the hidden state. The expressions are as follows:
[0115] o t =σ(W o ·[h t-1 ,x t ]+b o )
[0116] h t =o t ×tanh(C t )
[0117] o t W is the output of the output gate. o Let b be the weight matrix of the output gate. o h is the bias of the output gate. t This is the hidden layer state;
[0118] The output of the LSTM layer is input into the self-attention mechanism layer, which then assigns weights to the hidden layer state features output by the LSTM layer. The specific steps are as follows:
[0119] S1. Perform a linear transformation on the input hidden layer feature data, mapping it to a query, a key, and a value. The transformation formula is as follows:
[0120]
[0121] Where Q is the query matrix, K is the key matrix, V is the value matrix, and W is the value matrix. q W k W v These are the weight matrices for the query matrix, key matrix, and value matrix, respectively, and X is the input vector;
[0122] S2. Calculate the dot product of the query and each key to obtain the attention score, using the formula:
[0123]
[0124] d is the scaling factor, in the denominator Used to scale dot products and stabilize gradient flow;
[0125] S3. Normalize the attention score using the Softmax function to obtain the attention weights, the expression of which is:
[0126] weights = Softmax(scores)
[0127] S4. Calculate the weighted sum of V using the calculated attention weights to obtain the final attention output, expressed as:
[0128] output = weights·V
[0129] The output of the last time step of the attention mechanism calculation is taken and then connected to a fully connected layer to output the prediction result.
[0130] The specific steps for optimizing the hyperparameters of the feature extraction layer using the Bayesian algorithm in step 3 are as follows:
[0131] S1. Select a preset combination of hyperparameters, including CNN kernel size, number of neurons in LSTM hidden layer, number of LSTM hidden layer layers, and learning rate, and define a search space for each hyperparameter.
[0132] S2. Define the objective function, including initializing the CNN-LSTM-ATt model structure according to preset hyperparameters, training the entire model structure using the training set for a preset number of iterations, calculating the loss for each iteration, iteratively optimizing the model parameters based on backpropagation, and finally calculating the average loss on the test set as the objective function value.
[0133] S3. Optimization, the specific process is as follows:
[0134] S3.1. Based on the preset random search points for each hyperparameter, randomly select multiple hyperparameter combinations, train the model structure, obtain the objective function values corresponding to each hyperparameter combination, and form an observation dataset.
[0135] S3.2. Based on the observed dataset, a Gaussian regression process is used as a surrogate model to fit the objective function, generating a probability distribution for predicting the value of the objective function; the Gaussian regression prediction formula is:
[0136]
[0137] a * Let μ(a) be the combination of hyperparameters to be predicted. * ) represents the hyperparameter combination a * The predicted value of the objective function under the given conditions, cov(a) * ) represents the hyperparameter combination a * Covariance under the following conditions For the current hyperparameter combination a * C is a vector formed by the covariances of the various hyperparameter combinations in the observed dataset. ov To form the matrix of covariances among various hyperparameter combinations in the observation dataset, c ov For hyperparameter combination a * Its own covariance, Let be the variance term of the noise, n be the number of samples in the observation dataset, I be the identity matrix, and b be the vector of objective function values of each hyperparameter combination in the observation dataset.
[0138] S3.3. Based on the Gaussian regression process, the hyperparameter sample points are sampled by improving the acquisition function to select the next hyperparameter combination; the formula for the expected improvement is:
[0139] EI(x)=E[max(f(x)-f best ,0)]
[0140] EI(x) represents the desired improvement in the acquisition function, and f(x) is the predicted value of the objective function, given by a Gaussian process. best E is the currently known optimal value, and E is the expected value.
[0141] S3.4 Update the surrogate model. Repeat steps S3.2 and S3.3 within a preset number of iterations until the combination of hyperparameters corresponding to the lowest test set loss value is found. This combination is then used as the optimal hyperparameter combination and output.
[0142] The hyperparameter combination optimized by the Bayesian algorithm is applied to the CNN-LSTM-Att model structure to obtain the BO-CNN-LSTM-Att model structure; the BO-CNN-LSTM-Att model structure is trained based on the training set in step 1, and Adam is used as the training optimizer.
[0143] Finally, the mean absolute error (MAE), root mean square error (RMSE), and coefficient of determination (R²) are used. 2 The predictive performance of the BO-CNN-LSTM-ATt model is evaluated.
[0144] Example
[0145] like Figure 5 As shown, the water flow data of the Jingjiang section of the Yangtze River was calculated and acquired in real time using a self-developed integrated acoustic tomography flow measurement instrument and a synchronous water level gauge. The dataset includes the water flow of the Jingjiang section of the Yangtze River from 15:00 on September 22, 2024 to 17:30 on September 29, 2024, collected every 5 minutes, resulting in 2017 consecutive data points, which constitute the original water flow dataset. After preprocessing, 80% of the data was selected as the training set and 20% as the test set in chronological order to test the accuracy and reliability of the model.
[0146] As shown in Table 1 and Figure 7 As shown in (a), (b), (c), and (d), the key hyperparameters are optimized using a Bayesian optimization algorithm: the learning rate, kernel size of the CNN layer, hidden size of the LSTM layer, and number of hidden layers (num_layers) are used as hyperparameter combinations to set the hyperparameter search space, initial random search points, and number of optimization iterations; the optimal hyperparameter configuration combination is determined through a Gaussian regression process: the kernel size of the CNN layer is 2, the number of hidden layers of the LSTM is 10, the number of hidden layer neurons is 200, and the learning rate is 0.008; the final hyperparameter combination is used in the CNN-LSTM-Att model structure to obtain the BO-CNN-LSTM-Att model structure; based on the training set, the constructed BO-CNN-LSTM-Att model structure is trained to obtain the BO-CNN-LSTM-Att model.
[0147] Table 1 Optimization parameter settings for model hyperparameters
[0148]
[0149] The target river flow was predicted using the BO-CNN-LSTM-Att model of this invention, the traditional LSTM, the CNN-LSTM, and the CNN-LSTM-Att model, respectively. Figure 8 As shown, the prediction curve of the model BO-CNN-LSTM-Att in this invention is closer to the true value, and it is more sensitive to the trend of water flow change; as Figure 9 As shown, the BO-CNN-LSTM-Att model of this invention has the smallest mean absolute error (MAE) and root mean square error (RMSE). Compared with the traditional LSTM model, RMSE and MAE are reduced by 51.55% and 53.56%, respectively; compared with the CNN-LSTM model, they are reduced by 33.35% and 35.13%; and compared with the CNN-LSTM-Att model, they are reduced by 24.53% and 24.92%. In addition, the BO-CNN-LSTM-Att model of this invention has the highest coefficient of determination, reflecting a high degree of adaptability to complex water flow dynamics. The decreasing error trend from left to right and the gradually increasing coefficient of determination line reflect the necessity and superiority of model fusion.
[0150] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made by those skilled in the art within the scope of the technology disclosed in this invention, based on the technical solution and inventive concept of the present invention, should be covered within the protection scope of this invention. Therefore, the protection scope of this invention should be determined by the scope of the claims.
Claims
1. A method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-ATt, characterized in that, Includes the following steps: Step 1: Obtain the original water flow dataset for the target river channel during historical time periods, and preprocess the obtained original water flow dataset, dividing the preprocessed water flow dataset into training set and test set. Step 2: Construct the CNN-LSTM-ATt model structure, including a self-attention mechanism layer and a fully connected layer, as well as a feature extraction layer composed of CNN layers and LSTM layers connected in series; the CNN layer extracts features from the preprocessed water flow data, and the features extracted by the CNN layer are input into the LSTM layer, which captures the temporal features of the input features; The self-attention mechanism layer assigns weights to the temporal features output by the LSTM layer, and finally outputs the predicted short-term water flow data through the fully connected layer. Step 3: Use the Bayesian algorithm to optimize the relevant hyperparameters of the feature extraction layer, and apply the optimized hyperparameter combination to the CNN-LSTM-Att model structure to obtain the Bayesian optimized CNN-LSTM-Att model structure, namely the BO-CNN-LSTM-Att model structure. Step 4: Based on the training set, using the preprocessed historical water flow data of the target river as input and the water flow data of the target river within a preset time period after the input time period as output, train the constructed BO-CNN-LSTM-Att model structure to obtain the trained BO-CNN-LSTM-Att model; based on the test set, apply the BO-CNN-LSTM-Att model to obtain the predicted water flow data of the target river within a preset time period in the future to evaluate the model. Step 5: Using the water flow data of the target river channel during the target time period as input, apply the BO-CNN-LSTM-ATt model to obtain the predicted water flow data of the target river channel during the future preset time period.
2. The method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-ATt as described in claim 1, characterized in that, Step 1 involves obtaining the raw water flow dataset for historical time periods of the target river channel, which includes: S1. A self-made, remotely controlled, transceiver-integrated acoustic tomography system is deployed in staggered pairs on both banks of the target river channel. The acoustic tomography system includes a control system deployed at a site on the shore and underwater transducers fixed at a preset depth underwater. The control system and the corresponding underwater transducers are connected by a waterproof cable, and the horizontal baseline between the sites of the paired acoustic tomography systems on both banks forms a 45° angle with the direction of water flow. The reciprocity time of the acoustic waves is obtained by the underwater transducers corresponding to the paired acoustic tomography systems on both banks emitting sound waves to each other. Combined with X-ray acoustic correction and Doppler effect calibration, the river cross-sectional velocity is obtained. The water level is measured in real time using a synchronous water level gauge. The river cross-sectional area is calculated based on the water level. The water flow rate is obtained by multiplying the river cross-sectional area by the river cross-sectional velocity. S2. Collect water flow data according to step S1 within a preset time interval during the historical period to form the original water flow dataset.
3. The method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-ATt as described in claim 2, characterized in that, The raw water flow dataset obtained in step 1 is preprocessed, including normalization and tensor transformation, specifically as follows: S1. Map the water flow data to the interval [0, 1] using a normalization formula. The normalization formula is: In the formula: l is the index of the original water flow data, x l This represents the l-th raw water flow data value. x represents the normalized value of the l-th raw water flow data. max x represents the maximum value in the original water flow dataset; min This represents the minimum value in the original water flow dataset; S2. Perform tensor transformation on the normalized water flow data to meet the data input requirements of the CNN layer; S3. Divide the preprocessed dataset into training and test sets in chronological order at a ratio of 80% and 20%.
4. The method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-ATt as described in claim 1, characterized in that, In step 2, the CNN layer consists of convolutional layers and pooling layers. The convolutional layer extracts local features through one-dimensional convolution. The convolution formula is: Where j is the data index of the output sequence, y j The j-th data point in the output sequence; k is the index of the weight parameters inside the convolution kernel, w k x is the k-th weight parameter of the convolution kernel, where K is the kernel size; j+k-1 For y j The corresponding convolution kernel window covers the k-th input data; after the convolution operation, the convolution result is processed by the non-linear activation function ReLU, and then a max pooling layer is used to achieve dimensionality reduction; The pooled features are input into an LSTM layer. The neurons in the LSTM layer contain a forget gate, an input gate, an output gate, and a memory unit. The current input and the hidden state from the previous time step are input in parallel into the forget gate, input gate, output gate, and memory unit. The outputs of the forget gate and the input gate are proportionally summed to obtain the current memory unit state. The output gate controls the output ratio of the current memory unit state to obtain the final hidden state output. Specifically: The forgetting gate determines the ratio of information retained to forgotten in memory units from the previous time step, and its formula is: f t =σ(W f ·[h t-1 ,x t ]+b f ) t is the current time step, f t For the output of the forget gate, W f Let b be the weight matrix of the forget gate. f For the bias of the forget gate, h t-1 The hidden state of the previous time step, x t Given the current input, σ() is the Sigmoid activation function; The input gate determines the proportion of input information written to the memory unit at the current time step, and its formula is: i t =σ(W i ·[h t-1 ,x t ]+b i ) i t The output of the input gate, For candidate memory values, W i and W c Here, are the weight matrices for the input gate and candidate memory, respectively; tanh is the hyperbolic tangent activation function; b i and b c These are the biases for the input gate and candidate memory, respectively; The state of the memory unit is updated according to the forget gate and the input gate, and its expression is: The output f of the forget gate t Controlling old memories C t-1 The retention of the input gate's output i t Controlling candidate memories The addition; The output gate controls the output ratio of the current memory cell state, thereby obtaining the final output of the hidden state. The expressions are as follows: the t =σ(W o ·[h t-1 ,x t ]+b o ) h t =o t ×tanh(C t ) o t For the output of the output gate, W o Let b be the weight matrix of the output gate. o h is the bias of the output gate. t This represents the hidden layer state at the current time step. The output of the LSTM layer is input into the self-attention mechanism layer, and the self-attention mechanism assigns weights to the hidden layer state features output by the LSTM layer. Finally, the prediction result is output through a fully connected layer.
5. The method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-ATt according to claim 4, characterized in that, The specific steps of weight allocation for the hidden layer state features output by the self-attention mechanism in step 2 are as follows: S1. Perform a linear transformation on the input hidden layer feature data, mapping it to a query, a key, and a value. The transformation formula is as follows: Where Q is the query matrix, K is the key matrix, V is the value matrix, and W is the value matrix. q W k W v These are the weight matrices for the query matrix, key matrix, and value matrix, respectively, and X is the input vector; S2. Calculate the dot product of the query and each key to obtain the attention score, using the formula: d is the scaling factor, in the denominator Used to scale dot products and stabilize gradient flow; S3. Normalize the attention score using the Softmax function to obtain the attention weights, the expression of which is: weights = Softmax(score) S4. Calculate the weighted sum of V using the calculated attention weights to obtain the final attention output, expressed as: output = weights·V The output of the last time step of the attention mechanism calculation is taken and then connected to a fully connected layer to output the prediction result.
6. The method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-ATt as described in claim 5, characterized in that, The specific steps for optimizing the hyperparameters of the feature extraction layer using the Bayesian algorithm in step 3 are as follows: S1. Select a preset combination of hyperparameters, including CNN kernel size, number of neurons in LSTM hidden layer, number of LSTM hidden layer layers, and learning rate, and define a search space for each hyperparameter. S2. Define the objective function, including initializing the CNN-LSTM-ATt model structure according to preset hyperparameters, training the entire model structure using the training set for a preset number of iterations, calculating the loss for each iteration, iteratively optimizing the model parameters based on backpropagation, and finally calculating the average loss on the test set as the objective function value. S3. Optimization, the specific process is as follows: S3.
1. Based on the preset random search points for each hyperparameter, randomly select multiple hyperparameter combinations, train the model structure, obtain the objective function values corresponding to each hyperparameter combination, and form an observation dataset. S3.
2. Based on the observed dataset, a Gaussian regression process is used as a surrogate model to fit the objective function, generating a probability distribution for predicting the value of the objective function; the Gaussian regression prediction formula is: a * Let μ(a) be the combination of hyperparameters to be predicted. * ) represents the hyperparameter combination a * The predicted value of the objective function under the given conditions, cov(a) * ) represents the hyperparameter combination a * Covariance under the following conditions For the current hyperparameter combination a * C is a vector formed by the covariances of the various hyperparameter combinations in the observed dataset. ov To form the matrix of covariances among various hyperparameter combinations in the observation dataset, c ov For hyperparameter combination a * Its own covariance, Let be the variance term of the noise, n be the number of samples in the observation dataset, I be the identity matrix, and b be the vector of objective function values of each hyperparameter combination in the observation dataset. S3.
3. Based on the Gaussian regression process, the hyperparameter sample points are sampled by improving the acquisition function to select the next hyperparameter combination; the formula for the expected improvement is: EI(x)=E[max(f(x)-f best ,0)] EI(x) represents the desired improvement in the acquisition function, and f(x) is the predicted value of the objective function, given by a Gaussian process. best E is the currently known optimal value, and E is the expected value. S3.4 Update the surrogate model. Repeat S3.2 and S3.3 within the preset number of iterations until the hyperparameter combination corresponding to the lowest test set loss value is found. This hyperparameter combination is then used as the final hyperparameter combination and output.
7. The method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-ATt as described in claim 6, characterized in that, The optimized hyperparameter combination of the Bayesian algorithm is as follows: the kernel size of the CNN layer is 2, the number of neurons in the LSTM hidden layer is 200, the number of LSTM hidden layers is 10, and the learning rate is 0.
008.
8. The method for predicting short-term river flow based on Bayesian optimized CNN-LSTM-ATt as described in claim 1, characterized in that, Using mean absolute error (MAE), root mean square error (RMSE), and coefficient of determination (R²) 2 The predictive performance of the model is evaluated using the following formulas: Where m is the number of samples in the original water flow dataset, and l∈[1,m], x l For the l-th raw water flow data value, This is the predicted water flow rate. The mean is the sample mean; MAE reflects the model's average force prediction error; RMSE is more sensitive to some extreme errors, and the combination of the two errors better reflects the accuracy of the model's water flow prediction; R 2 This reflects the model's fit; the closer the value is to 1, the better the model's predictive performance.