Daily runoff prediction method based on long-term and short-term component neural network (LSTCNet)
By introducing long-term short-term component neural network (LSTCNet) into the runoff prediction network, processing historical data and short-term data respectively, combining multiple neural network structures and AR models, the problem of ignoring the difference between long-term and short-term data in the existing technology is solved, and a higher precision daily runoff prediction is achieved.
Patent Information
- Application Number
- CN202311498823.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-11-13
AI Technical Summary
The existing runoff prediction network ignores the differences between long-term and short-term data, making it difficult to extract highly nonlinear and complex hydrological features, which in turn affects the accuracy of daily runoff prediction.
The daily runoff prediction method based on long-term short-term component neural network (LSTCNet) is adopted to process historical data and short-term data through long-term components and short-term components respectively. Combining the convolutional layer, GRU layer, attention layer and residual structure, long-term and short-term features are extracted and learned, and the linear dependence between variables is captured through the AR model.
The prediction accuracy of the Japanese runoff prediction network model is improved, and the complex characteristics of the hydrological system can be captured more accurately, which significantly improves the prediction ability of Japanese runoff.
Smart Images

Figure SMS_11 
Figure SMS_12 
Figure SMS_13
Abstract
Description
Technical Field
[0001] The present invention relates to the problem of runoff prediction in the field of hydrology, and in particular to a daily runoff prediction method based on a long-term short-term component neural network (LSTCNet). Background Art
[0002] Accurately predicting daily runoff is conducive to the effective allocation of water resources and provides strong guidance for flood control. Due to many factors such as meteorological conditions, surface mantle of the basin, and human activities, the runoff process is characterized by randomness, chaos, and fuzziness. Machine learning is an end-to-end learning method that can accurately predict the daily runoff of a basin without relying on experience or considering physical processes. Since it was introduced into the field of hydrology, it has flourished and shown great potential.
[0003] It is difficult for simple data-driven models to understand the complex characteristics of hydrological systems, so some researchers have proposed recurrent neural networks (RNNs) to learn dependencies between time series. Although RNNs have great advantages in processing time series tasks, they suffer from gradient explosion and gradient vanishing problems after multiple time steps of backpropagation. To address this bottleneck, researchers have proposed two different types of RNNs - long short-term memory neural networks (LSTM) and gated recurrent networks (GRU). They both summarize the cell states across all time steps, overcome the shortcomings of traditional RNNs, and learn long-term dependencies. In addition, with the widespread use of attention mechanisms, researchers have begun to apply AM to runoff prediction, further improving the prediction accuracy.
[0004] However, existing studies do not pay enough attention to the differences between long-term and short-term data. Although researchers have improved LSTM and GRU networks and proposed various network structures, these networks use the same structure to extract and learn long-term and short-term features, while ignoring the differences between them. Therefore, these networks are not sufficient to extract highly nonlinear and complex hydrological features, and accurate and reliable daily runoff prediction remains a challenging task for these networks. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a daily runoff prediction method based on a long-term short-term component neural network (LSTCNet) to improve the prediction accuracy of a daily runoff prediction network model.
[0006] The following technical measures are adopted to solve the above technical problems: a daily runoff prediction method based on long-term short-term component neural network (LSTCNet) mainly includes the following steps:
[0007] (1) Take a set of long-term time series historical data {Li} = L1, L2, ..., Ln and a short-term data S as input. The long-term data is fed to the long-term component, and the short-term data is fed to the nonlinear short-term component and the linear component.
[0008] (2) Long-term component processes long-term data: The long-term component consists of convolutional layers, GRU layers, and attention layers. Long-term data is input into the convolutional layer to extract the time distribution features and local dependencies between variables. The recurrent component is a recurrent layer with a GRU. The GRU layer automatically learns to extract the feature values of the features with the support of the attention mechanism and assigns different weights according to the importance of the features. The specific implementation details of this module are as follows:
[0009] The input matrix is L∈R T×D , the convolutional layer consists of k kernels of size ω1×D, and the convolutional layer output matrix vector is h L , which can be regarded as a length of T c A k-dimensional vector sequence, where T c =T-ω1+1, T represents the time step. L is input to two GRU layers n times, each GRU layer contains g units, and the output of the recurrent component is h t,n ∈R g , the output of the entire long-term component process is o L ∈R n×g .
[0010] (3) Short-term component processes short-term data: The short-term component uses the residual structure composed of convolutional layers, namely the RC layer, to make full use of multi-level feature information. The residual structure layer can fuse multi-level feature information, thereby improving the comprehensiveness of feature information.
[0011] The structure of the RC layer is as follows Figure 2 As shown in Figure 1, an RC layer contains two ReLU functions and two convolutional layers. Each convolutional layer contains j kernels with a kernel size of ω2×ω2, where the first ω2 is the time dimension and the second ω2 is the variable dimension. The input S of the RC layer (i) After the ReLU activation function, the convolutional layer is input to obtain the shallow feature matrix Then Input the next ReLU activation function and convolution layer to get a deep feature matrix Finally, by inputting S (i) and matrix Integrate the residual structure and get the output of RC as:
[0012] S (i+1) =ReLU(W s2 *ReLU(W s1 *S(i) +b s1 )+b s2 )+S (i) ,i=1,…,I (1)
[0013] Where I is the number of RC layers, W s1 and b s1 is the learnable parameter of the previous convolutional layer, W s2 and b s2 is the learnable parameter of the next convolutional layer, S (i) is the input of the RC layer, S (i+1) is the output; the filling is the same, the output is S (i+1) It will be an S (i+1) ∈R T×D Vector.
[0014] Then, the output S of RCs is transformed into (l+1) Input another convolutional layer, which consists of k kernels of size ω2×ω2, and the output is h S , which is connected through a fully connected layer (FCL) to obtain o S ∈R g , which is the prediction result of short-term data.
[0015] (4) The linear part processes short-term data: The linear part uses the AR model to enhance the ability to capture the linear dependence between variables and correct the predicted value of the nonlinear part. Assuming the input matrix is S, the AR model is expressed as follows:
[0016]
[0017] in, represents the output of the AR model, and b ar is a learnable parameter, s t-p is the input of the layer at time t, and q is the input window size, which represents the amount of past information required by the model.
[0018] (5) The final prediction of the model comes from the sum of the outputs of the long-term component, the short-term component, and the AR component. FCL is used to combine the long-term component output and the short-term component output of the nonlinear part. The output of FCL is:
[0019]
[0020] Where W nl and b nl is the learnable parameter in FCL, relatively, o L and S is the output of long-term and short-term historical data forecasts, Represents the nonlinear output of the model. Finally, by integrating the output of the nonlinear part and the linear part, the final prediction of the model is obtained:
[0021]
[0022] (6) The method of the present invention trains and predicts y by minimizing the mean absolute error between the predicted value and the observed value. t The Adam optimizer is used in the model training process, and its objective function is:
[0023]
[0024] Among them, N is the number of training samples and D is the dimension of the target data. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a structural block diagram of the long-term short-term component neural network (LSTCNet) proposed in the present invention.
[0026] Figure 2 Details of the residual structure (RC layer) composed of convolutional layers.
[0027] Figure 3 The data source for evaluating the model performance is the Qingxi River Basin map.
[0028] Figure 4 Comparison of the visualization results of LSTCNet's predicted values and observed values for the Qingxi dataset.
[0029] Figure 5 A detailed view comparing the LSTCNet predictions and observed values. DETAILED DESCRIPTION
[0030] The present invention is further described below in conjunction with the accompanying drawings. It is necessary to point out here that the following embodiments are only used to further illustrate the present invention and cannot be understood as limiting the scope of protection of the present invention. Persons skilled in the art in this field may make some non-essential improvements and adjustments to the present invention based on the above content of the present invention, which still fall within the scope of protection of the present invention.
[0031] (1) The method proposed in the present invention is developed on the Tensorflow deep learning framework. In order to learn the dependency more effectively, the parameter values shown in Table 1 are selected to train the model.
[0032] (2) To verify the prediction performance of the proposed method, historical runoff data of the Qingxi River Basin from the Sichuan Provincial Hydrology and Water Conservancy Bureau were selected. The locations of the basic hydrological stations and three rain gauge stations are shown in the figure below. Figure 3The site information, sequence length of the collected data and Pearson correlation coefficients (PCCs) are listed in Table 2. At the same time, a total of 20 years of data (1986-2005) were used for model calibration, and the statistical information of the training (1986-2003) and testing (2004-2005) data, such as the minimum, maximum and average values, are shown in Table 3.
[0033] (3) Five commonly used statistical indicators, namely mean absolute error (MAE), root mean square error (RMSE), Nash-Sutcliffe efficiency (NSE), correlation coefficient (CC) and Willmott index (WI), were used to quantitatively evaluate all experimental models. MAE and RMSE were used to evaluate the prediction accuracy and robustness of the proposed method; NSE was used to measure the ability to predict variables different from the mean and give the proportion of the initial variance accounted for by the model; CC represents the strength and direction of the linear relationship between the observed and predicted values.
[0034] (4) The runoff prediction network compared by the method of the present invention is:
[0035] Method 1: The method proposed by Yin Z, Feng Q, Wen X et al., reference “Design and evaluation of SVR, MARS and M5Tree models for 1, 2 and 3-day lead time forecasting of riverflow data in a semiarid mountainous catchment [C] / / Stoch Environ Res Risk Assess. 2018: 2457–2476”;
[0036] Method 2: The method proposed by Qi Y, Zhou Z, Yang L et al., reference “A Decomposition-Ensemble Learning Model Based on LSTM Neural Network for Daily ReservoirInflow Forecasting[C] / / Water Resour Manage.2019:4123–4139”;
[0037] Method 3: The method proposed by Y.wu, Z.Liu, W.Xu et al., reference “Context-Aware Attention LSTM Network for Flood Prediction[C] / / International Conference on Pattern Recognition(ICPR).2018:1301-1306”;
[0038] Method 4: The method proposed by Bi XY, Li B, Lu WL et al., reference “Daily runoff forecasting based on data-augmented neural network model[J] / / Journal of Hydroinformatics.2020:900–915”.
[0039] (5) As can be seen from Table 4, compared with MARS, LSTM, AMLSTM, and CAGANet runoff prediction networks, the method proposed in this invention achieves the best performance. Figure 4 The thick solid line in the upper middle part is the real surface rainfall time series, and the solid line and dotted line below represent the real value and predicted value of surface runoff, respectively. Figure 4 Detailed view of Figure 5 ,It can be observed that the overall prediction performance of LSTCNet is better, the peak prediction performance is close to the observed values at most points, and LSTCNet predicts small-scale floods better than large-scale floods.
[0040] (6) To demonstrate the effectiveness of the framework of this method, some ablation experiments were conducted on the original dataset. The experimental results are shown in Tables 5, 6, and 7. Table 5 shows the indispensable role of all components of LSTCNet, Table 6 shows the effectiveness of the Long+Short (i.e., using long-term components to process long-term data and short-term components to process short-term data) combination scheme, and Table 7 shows that the multi-level information generated by stacked RCs can improve accuracy, and the optimal number of RC layers is 3.
[0041] Table 1 Parameter selection
[0042] Parameters Value Number of the long - term historical data series n 6 Length of time step T 6 Dimension of variables D 6 Number of CNN layers k (the nonlinear part process long - term data) 16 <![CDATA[Size of kernelω1×D(the nonlinear part process long-term data)]]> [3,6] Number of GRU neurons g [16,16] Number of CNN layers k (the nonlinear part process short - term data) 16 <![CDATA[Size of kernelω2×ω2(the nonlinear part process short-term data)]]> [3,3] Number of CNN layers j (the structure of RC) 64 <![CDATA[Size of kernelω2×ω2(the structure of RC)]]> [3,3] Number of RC layers I 3 Size of the input window q 6 Size of batch 60 Number of epochs 300 Rate of learning 0.0001 Rate of dropout 0.2
[0043] Table 2 Qingxi hydrological station information
[0044]
[0045] Table 3 Qingxi dataset information
[0046]
[0047] Table 4 Comparison results of different methods on the Qingxi dataset
[0048] Models MAE RMSE NSE CC WI MARS 5.19 21.69 0.550 0.761 0.800 LSTM 4.36 17.28 0.753 0.873 0.889 AM - LSTM 3.93 16.22 0.816 0.907 0.934 CAGANet 3.11 12.23 0.854 0.918 0.918 Ours 2.90 11.92 0.855 0.947 0.961
[0049] Table 5 Comparison results after removing different components
[0050] Models MAE RMSE NSE CC WI Ours 2.90 11.92 0.855 0.947 0.961 Ours w / o long - term component 5.56 21.79 0.546 0.823 0.771 Ours w / o short - term component 4.63 23.57 0.469 0.795 0.707 Ours w / o AR component 2.99 13.13 0.835 0.916 0.940
[0051] Table 6 Comparison results of different combinations of components for processing long-term data and short-term data
[0052] Scheme MAE RMSE NSE CC WI Long - term component+Long - term component 3.20 13.55 0.825 0.910 0.927 Long - term component+Short - term component 2.90 11.92 0.855 0.947 0.961 Short - term component+Long - term component 5.01 22.31 0.524 0.808 0.756 Short - term component+Short - term component 3.18 12.98 0.839 0.924 0.943
[0053] Table 7 Prediction comparison results using different numbers of RCs
[0054] Nunber of RCs MAE RMSE NSE CC WI 1 3.45 15.69 0.768 0.891 0.918 2 3.34 14.13 0.809 0.908 0.938 3 2.90 11.92 0.855 0.947 0.961 4 3.32 13.55 0.825 0.914 0.944 5 3.31 14.55 0.789 0.895 0.937
Claims
1. A daily runoff prediction method based on long-term short-term component neural network (LSTCNet), characterized by The following steps are involved: (1) Taking a set of long-term time series historical data {Li} = L1, L2, ..., Ln and a short-term data S as input, the long-term data is fed to the long-term component, and the short-term data is fed to the nonlinear short-term component and the linear component; (2) Long-term component processes long-term data: The long-term component consists of convolutional layers, GRU layers, and attention layers. The convolutional layer is used to extract time-distributed features and local dependencies between variables. The recurrent component is a recurrent layer with GRU. The GRU layer automatically learns to extract feature values with the support of the attention mechanism and assigns different weights according to the importance of the features. The specific implementation details of this module are as follows: The input matrix is L∈R T×D , the convolutional layer consists of k kernels of size ω1×D, and the convolutional layer output matrix vector is h L , which can be regarded as a length of T c A k-dimensional vector sequence, where T c =T-ω1+1, T represents the time step, h L is input to two GRU layers n times, each GRU layer contains g units, and the output of the recurrent component is h t,n ∈R g , the output of the entire long-term component process is o L ∈R n×g ; (3) Short-term component processes short-term data: The short-term component uses a residual structure composed of convolutional layers (i.e., RC layer) to fully utilize multi-level feature information, thereby improving the comprehensiveness of feature information. The specific implementation details of this part are as follows: An RC layer contains two ReLU functions and two convolutional layers. Each convolutional layer contains j kernels with a kernel size of ω2×ω2, where the first ω2 is the time dimension and the second ω2 is the variable dimension. The input S of the RC layer is (i) After the ReLU activation function, the convolutional layer is input to obtain the shallow feature matrix Then Input the next ReLU activation function and convolution layer to get a deep feature matrix Finally, by inputting S (i) and matrix Integrate the residual structure and get the output of RC as: S (i+1) =ReLU(W s2 *ReLU(W s1 *S (i) +b s1 )+b s2 )+S (i) ,I=1,…,I (1) Where I is the number of RC layers, W s1 and b s1 is the learnable parameter of the previous convolutional layer, W s2 and b s2 is the learnable parameter of the next convolutional layer, and then the output S of RCs is transformed into (I+1) Input another convolutional layer, which consists of k kernels of size ω2×ω2, and the output is h S , which is connected through a fully connected layer (FCL) to obtain o S ∈R g , which is the prediction result of short-term data; (4) The linear part processes short-term data: The linear part uses the AR model to enhance the ability to capture the linear dependence between variables and correct the predicted value of the nonlinear part. Assuming the input matrix is S, the AR model can be expressed as: in, represents the output of the AR model, and b ar is a learnable parameter, s t- p is the input of the layer at time t, q is the input window size, which represents the amount of past information required by the model; (5) The final prediction of the model comes from the sum of the outputs of the long-term component, the short-term component and the AR component. FCL is used to combine the long-term component output and the short-term component output of the nonlinear part: Where W nl and b nl is the learnable parameter in FCL, relatively, o L and S is the output of long-term and short-term historical data forecasts, represents the nonlinear output of the model, and finally the final prediction of the model is obtained by integrating the output of the nonlinear part and the linear part: (6) The method of the present invention trains and predicts y by minimizing the mean absolute error between the predicted value and the observed value. t , the Adam optimizer is used in the model training process, and its objective function is: Among them, N is the number of training samples and D is the dimension of the target data.
2. A daily runoff prediction method based on long-term short-term component neural network (LSTCNet) as claimed in claim 1, characterized in that Different network structures are used to process long-term and short-term data.
3. A daily runoff prediction method based on long-term short-term component neural network (LSTCNet) as claimed in claim 1, characterized in that An attention mechanism (AM) is introduced to automatically select relative times among all time steps of long-term data.
4. A daily runoff prediction method based on long-term short-term component neural network (LSTCNet) as claimed in claim 1, characterized in that Utilize multi-layer residual structure to fuse multi-level features in short-term data.
5. A daily runoff prediction method based on long-term short-term component neural network (LSTCNet) as claimed in claim 1, characterized in that The traditional AR model is used as a linear neural network part to enhance the learning of linear dependencies between variables and modify the output prediction value.
6. A daily runoff prediction method based on long-term short-term component neural network (LSTCNet) for executing requirements 1 to 5.
Citation Information
Patent Citations
Long-term and short-term traffic flow prediction model construction method based on deep learning
CN110610232A
Short-term air traffic flow prediction method based on combined model
CN115293417A
Short-term power load prediction method suitable for power distribution area
CN115688993A