Intelligent daily runoff forecasting method and system based on mixed feature optimization and variation prediction, storage medium and electronic equipment

Through the Gray Wolf Optimization Algorithm, the variational modal decomposition parameters were optimized, combined with the Mamba2 and Transformer models, multi-scale features were extracted and error correction was performed, which solved the problem of insufficient adaptability of the existing runoff prediction model on nonlinear and non-stationary data, and achieved higher precision runoff prediction.

CN120561555APending Publication Date: 2025-08-29NORTH CHINA UNIV OF WATER RESOURCES & ELECTRIC POWER
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510720822.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

When the existing runoff prediction model deals with complex nonlinear and non-stationary data characteristics, it lacks adaptability and cannot effectively reflect the impact of climate change and extreme climate events, resulting in low prediction accuracy.

Method used

The Gray Wolf Optimization Algorithm is used to optimize the variational modal decomposition parameters, combined with the Mamba2 model, Transformer model and one-dimensional convolutional neural network, through mixed feature optimization and variational prediction, multi-scale features are extracted and error correction is performed to improve the adaptability and accuracy of the prediction model.

Benefits of technology

The adaptability and prediction accuracy of the runoff prediction model to non-stationary and nonlinear data is improved, providing a more reliable basis for flood prevention and disaster reduction and water resource scheduling decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561555A_ABST
    Figure CN120561555A_ABST
Patent Text Reader

Abstract

The invention discloses a daily runoff prediction method and system based on GWO-VMD feature reconstruction and a mixed deep learning model, a storage medium and electronic equipment. VMD parameters are optimized through a GWO algorithm, an original runoff sequence is decomposed, and multi-scale features are fused; high-precision prediction is realized by combining the local feature extraction capability of a Mamba2 model and the global time sequence modeling advantage of Transform; and the prediction error is further corrected by using the CNN. The method solves the problem that a traditional model is poor in adaptability to non-stable and non-linear data, and has remarkable application value in flood control and disaster reduction and water resource scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hydrological prediction, and in particular to a daily runoff intelligent forecasting method, system, storage medium and electronic equipment based on hybrid feature optimization and variational prediction. Background Art

[0002] Currently, water scarcity has become a key factor restricting socioeconomic development and ecological and environmental protection worldwide. Accurate daily runoff forecasts are crucial for water resource management and flood prevention and disaster reduction. However, existing runoff prediction models have limitations in handling complex nonlinear and nonstationary data characteristics and in terms of prediction accuracy. Traditional process-based physical models and data-driven statistical models lack adaptability in the face of climate change and environmental changes, and are unable to effectively reflect new hydrological characteristics or the impact of extreme climate events. In recent years, with the development of deep learning technology, although some results have been achieved in runoff prediction, the inherent nonstationarity and nonlinearity of hydrological time series still pose significant challenges to these models. Summary of the Invention

[0003] The purpose of the present invention is to provide a daily runoff intelligent forecasting method, system, storage medium and electronic equipment based on hybrid feature optimization and variational prediction, which can improve the adaptability of the runoff prediction model to non-stationary and nonlinear data.

[0004] The technical solution adopted in the present invention is:

[0005] A daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction includes the following steps:

[0006] Step 1: Use the Grey Wolf Optimization (GWO) algorithm to optimize the variational mode decomposition (VMD) parameters, then decompose the original runoff sequence to obtain multiple intrinsic mode functions (IMFs). Extract the statistical characteristics and time-frequency domain features of each IMF, aggregate them into a feature set, and fuse the original runoff sequence with the feature set to form a joint feature.

[0007] Step 2: Divide the joint features into training set, validation set and test set, and use minimum and maximum normalization to normalize the data to the [0,1] interval;

[0008] Step 3: Divide the normalized time series data into fixed-length time windows and generate tensor format to input into the Mamba2 model;

[0009] Step 4: Generate high-level feature representations of the time series through the Mamba2 model and input the high-level feature representations into the encoder-decoder structure of the Transformer model to mine deep temporal dependencies;

[0010] Step 5: Map the Transformer output to the prediction result space through a linear layer, and perform denormalization on the prediction result to restore it to the original numerical range;

[0011] Step 6: Use the one-dimensional convolutional neural network (CNN) to correct the error of the prediction results and output the final optimized daily runoff prediction value.

[0012] In the step 1, the Grey Wolf Optimization Algorithm (GWO) is used to optimize the variational mode decomposition (VMD) parameters, which specifically includes the following steps: setting the number of VMD modes, the balanced reconstruction error, the parameters of the modal bandwidth, the center frequency update step, and the population size and maximum number of iterations of the GWO;

[0013] The energy entropy of each IMF is used as the fitness function, and the VMD parameters are iteratively updated through the individual positions of the gray wolves until the optimal parameter combination is obtained.

[0014] In the step 2, the data in the joint special type is divided into a training set of 70%, a validation set of 20%, and a test set of 10%. Normalization is achieved by subtracting the minimum value of the feature and dividing it by the difference between the maximum and minimum values ​​of the feature.

[0015] The length of the time window in step 3 is dynamically adjusted according to the historical runoff data period, and the data loader converts the window data into batch tensors and inputs them into the Mamba2 model.

[0016] The step 4 specifically includes the following steps:

[0017] 4.1, Mamba2 model performs preliminary processing on the input time series data:

[0018] First, the convolutional layer performs a convolution operation on the input time series data. In the time dimension, the convolution kernel slides and interacts locally with the data, integrating the features of adjacent time steps and extracting local features. Next, the parameter generation network uses a small neural network to dynamically generate the parameters of the state-space model (SSM) based on the characteristics of the input sequence, allowing the model behavior to adapt to the input data in real time. After that, the block recursive algorithm layer divides the long sequence into fixed-size blocks. Each block is calculated independently, and the output of the previous block is used as the initial state for the calculation of the next block. This recursive processing of long sequences effectively reduces computational complexity and improves memory usage efficiency. Finally, the multi-head mechanism layer introduces multiple heads to process the input data in parallel. Different heads focus on different aspects of the sequence, comprehensively capturing the various patterns and dependencies in the sequence, and enhancing the model's ability to express time series data.

[0019] 4.2. Preprocess the initially processed features: Specifically, add position codes. Position codes are generated in the form of sine and cosine functions. Specifically, for each time step t and each feature dimension i, the value of the i-th dimension of the position code can be expressed as:

[0020]

[0021] Where H represents the feature dimension of the high-level representation, t represents the time step, and i represents the index of the feature dimension. The position encoding is added to the high-level representation element by element to obtain the feature representation with added position information.

[0022] 4.3. The Transformer model processes the feature representation with positional encoding through its encoder part; by stacking multiple layers of encoders, the Transformer model gradually extracts deeper feature information and relationship patterns in the input sequence; each layer of encoder will further explore the complex structure and dependency relationships in the sequence based on its input, and ultimately generate a high-level feature representation that contains rich semantic information and dynamic laws.

[0023] The step five specifically includes the following steps:

[0024] 5.1. First, the decoder of the Transformer model processes the feature representation of the encoder output; this ensures that the decoder can only rely on previously generated information when generating the output of each time step;

[0025] 5.2, then, the output of the masked multi-head self-attention mechanism sublayer is input into the encoder-decoder multi-head self-attention mechanism sublayer to interact the decoder output with the encoder output, so that the decoder can use the global feature information extracted by the encoder to generate more accurate prediction results;

[0026] 5.3. Next, the output of the encoder-decoder multi-head self-attention mechanism sublayer is input into the feedforward neural network sublayer, where the same nonlinear transformation operation as the encoder is performed to further process and enhance the feature representation.

[0027] 5.4. After processing by multiple layers of decoders, the final decoder output is obtained, which contains the prediction information for the time series data. A linear layer is then used to map the feature representation of the decoder output to the target prediction result space.

[0028] 5.5, finally, an inverse normalization operation is applied to the prediction results to restore them from the normalized range to the actual numerical range of the original data.

[0029] The denormalization operation in step five is specifically achieved by multiplying the normalized prediction result by the original data range and adding the minimum value.

[0030] An intelligent daily runoff forecasting system based on hybrid feature optimization and variational prediction includes:

[0031] The data preprocessing module is used to clean and correct raw runoff data, handle missing values ​​and outliers, and ensure data quality. It also extracts a variety of statistical and time series features to reflect the basic laws and changing trends of the data. The module also smooths, detrends, and normalizes the data to reduce noise and scale differences, improving model training results. In addition, it uses derived features to enhance the original information, providing a more comprehensive and accurate input basis for subsequent modeling and prediction.

[0032] The hybrid deep learning module combines the multi-scale self-attention advantages of Mamba2, the global dependency modeling capabilities of Transformer, and the local feature extraction capabilities of CNN to comprehensively characterize hydrological runoff time series data. Specifically, Mamba2 excels in capturing complex nonlinear temporal relationships, Transformer improves the capture of long-range dependencies, and CNN enhances sensitivity to local temporal features. The integration of these three enables the model to more accurately reflect the multi-level and multi-scale characteristics of runoff series, significantly improving the robustness and accuracy of predictions.

[0033] The Mamba2 model contains convolutional layers and recurrent neural network layers to extract local temporal features;

[0034] The Transformer model uses a multi-head self-attention mechanism to capture global temporal dependencies by calculating the scaled dot product attention weights of the query, key, and value matrices;

[0035] The CNN model consists of a one-dimensional convolutional layer, a pooling layer, and a fully connected layer. The convolution kernel is used to slide and extract the prediction error features. The pooling layer compresses the feature dimensions. The fully connected layer outputs the corrected prediction results for executing steps 4 to 6.

[0036] The result output module is used to generate and visualize the final prediction results.

[0037] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the device where the computer-readable storage medium is located executes the daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction.

[0038] An electronic device includes: a memory and a processor, wherein the memory stores a program that can be run on the processor, and when the processor executes the program, the daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction is implemented.

[0039] This solution optimizes the VMD parameters through the GWO algorithm to achieve effective decomposition of the original runoff series and integrate multi-scale features. Specifically, VMD itself, as an adaptive signal decomposition method, can to some extent solve the modal aliasing problem of empirical mode decomposition (EMD), but its performance is greatly affected by parameters such as the penalty factor α and the modal decomposition number k. Improper parameter selection will lead to poor decomposition results. The GWO algorithm has global optimization capabilities. By using the VMD penalty factor α and the number of modes k as optimization variables and evaluating the fitness function using indicators such as the root mean square error (RMSE) during the iterative process, it automatically searches for the optimal VMD parameter combination, improving the accuracy and effectiveness of VMD for decomposing non-stationary and nonlinear runoff series, and enabling more accurate extraction of multi-scale features from complex data.

[0040] The solution then combines the Mamba2 model's powerful local feature extraction capabilities with the Transformer's global time series modeling advantages to achieve high-precision predictions. The Mamba2 model, based on a structured state-space model and boasting linear time complexity, is highly efficient in capturing local feature information when processing long sequences of data. The Transformer excels at modeling long-range dependencies within sequences and, through its self-attention mechanism, can learn the temporal characteristics of data from a global perspective. The combination of these two approaches fully exploits data features at both local and global scales, effectively addressing the complex feature patterns of non-stationary and nonlinear data and improving prediction accuracy.

[0041] Furthermore, the solution uses a CNN model to correct prediction errors. CNN models have powerful feature extraction and pattern recognition capabilities, and can learn the characteristics and patterns in prediction errors. By analyzing and processing these error characteristics, they can reversely correct the prediction results, further improving overall prediction accuracy and providing a more reliable basis for decision-making in flood prevention and water resource scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 is a flow chart of the present invention;

[0044] Figure 2 The structure diagram of the VMD decomposition method of the present invention;

[0045] Figure 3The transformer model structure diagram of the present invention;

[0046] Figure 4 The structural diagram of the Mamba2 model of the present invention;

[0047] Figure 5 CNN structure diagram of the present invention. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.

[0049] like Figure 1 、 2 As shown in 3, the present invention includes the following steps:

[0050] Step 1: Use the Grey Wolf Optimization Algorithm (GWO) to optimize the variational mode decomposition (VMD) parameters, then decompose the original runoff sequence to obtain multiple intrinsic mode functions (IMFs); extract the statistical characteristics and time-frequency domain characteristics of each IMF, summarize them to form a feature set, and fuse the original runoff sequence and the feature set into a joint feature; the new joint feature is used as the input of the subsequent model. Specifically, it refers to using the Grey Wolf Optimization Algorithm (GWO) to optimize the variational mode decomposition (VMD) parameters, decompose the original runoff sequence, and better extract the characteristic information contained therein, and fuse it with the original sequence to provide better input data for subsequent modeling. The specific process is as follows:

[0051] First, determine the parameters for GWO-VMD. For the VMD portion, you need to set the number of modes (K), which determines the number of intrinsic mode functions (IMFs) into which the original sequence will be decomposed. Also, set the parameter (α) that balances the reconstruction error and modal bandwidth, as well as the parameter (τ) that controls the update step size of the modal center frequency. For the GWO portion, you need to determine parameters such as the population size (i.e., the number of gray wolves) and the maximum number of iterations.

[0052] The GWO algorithm is then run to optimize the VMD parameters. During each iteration, the wolf group updates its position based on the quality of the VMD decomposition results, adjusting the VMD parameters. Specifically, the group evaluates the decomposition performance of the current parameter combination by calculating metrics such as the energy entropy of each IMF after VMD decomposition. Based on these evaluation results, individual wolves update their positions according to the GWO algorithm's rules, continuously approaching the optimal VMD parameter combination.

[0053] When the GWO algorithm converges, the optimal VMD parameters are obtained. These optimal parameters are then used to perform VMD decomposition on the original runoff series, yielding a series of intrinsic mode functions (IMFs). These IMFs can reflect the characteristics of the original runoff series at different frequency scales. For example, low-frequency IMFs may reflect long-term runoff trends, while high-frequency IMFs may capture information such as short-term runoff fluctuations.

[0054] Next, we extract the features of these IMFs. Various feature extraction methods can be used, such as calculating statistical features such as the energy, mean, variance, and kurtosis of each IMF, or calculating its time-frequency domain characteristics. These features are aggregated to form a feature set.

[0055] Finally, the original runoff sequence is fused with the extracted feature set. This fusion method can be simple concatenation, where the values ​​of the original sequence are combined with the individual feature values ​​in the feature set in a certain order to form a new joint feature vector. More complex fusion strategies can also be employed, such as weighted summation, where different weights are assigned to each component based on its importance, resulting in a new joint feature. This new joint feature serves as input to subsequent models, more comprehensively reflecting the characteristics of the original runoff sequence and providing richer information for subsequent modeling and prediction.

[0056] Step 2: Divide the joint features into training set, validation set and test set, and use minimum and maximum normalization to normalize the data to the interval [0,1]. In the step 2, the data division ratio of the joint features is 70% for the training set, 20% for the validation set and 10% for the test set. Normalization is achieved by subtracting the minimum value of the feature and dividing it by the difference between the maximum and minimum values ​​of the feature.

[0057] In actual use, the method used in this application is minimum and maximum normalization, scaling the data to the range of [0,1] "means to reasonably divide the joint feature data obtained after step 1 and normalize it using the minimum and maximum normalization method to unify the data value range to the interval of [0,1], thereby facilitating subsequent model training and prediction. The specific process is as follows:

[0058] First, determine the dataset partitioning ratio. Typically, joint feature data is divided into a training set, a validation set, and a test set. For example, a 7:2:1 ratio might be used: 70% of the data is used as the training set for model training, 20% of the data is used as the validation set for evaluating and adjusting model performance during training, and 10% of the data is used as the test set for final evaluation of the model's generalization and predictive performance. The specific partitioning ratio can be adjusted based on the amount of data and actual needs.

[0059] Then, according to the determined division ratio, the joint feature data is randomly divided into training set, validation set, and test set. During the division process, it is necessary to ensure that the data distribution in each data set can well reflect the characteristics of the original data to avoid large deviations between data sets due to improper division.

[0060] Next, the divided data set is normalized by minimum and maximum normalization. For each feature, its normalized value (x') can be calculated by the following formula:

[0061]

[0062] Here, x represents a feature value in the original data, min(x) represents the minimum value of this feature among all values ​​in the entire dataset, and max(x) represents the maximum value of this feature among all values ​​in the entire dataset. This normalization method scales the value range of each feature to the interval [0, 1]. This eliminates the impact that differences in the dimensions and value ranges of different features may have on model training, allowing the model to treat each feature more fairly, thereby improving model training and prediction performance.

[0063] Step 3: Divide the normalized time series data into fixed-length time windows, each containing a certain number of time steps and corresponding observations. This is generated in tensor format and fed into the Mamba2 model. The length of the time window in step 3 is dynamically adjusted based on the period of historical runoff data. The data loader converts the window data into batch tensors before feeding them into the Mamba2 model. The data loader allows batch reading of data from the dataset and converting it into the tensor format required by the model.

[0064] In actual use, the joint feature time series data processed in step 2 is windowed to convert the time series data into a form suitable for model processing. The data is then loaded into the Mamba2 model in batches through the data loader to prepare for model training and prediction. The specific process is as follows:

[0065] First, determine the length of the time window (L). The time window length can be selected based on the actual problem requirements and the characteristics of the data. For example, in a runoff prediction problem, if you need to consider the impact of data from a longer time range on the current runoff, you can choose a longer time window length; if you are concerned about short-term dynamic changes, you can choose a shorter time window length. The choice of time window length affects the model's ability to capture the temporal dependence and dynamic characteristics of time series data.

[0066] Then, according to the determined time window length, the preprocessed joint feature time series data is sequentially slid into fixed-length time windows. Each time window contains L consecutive time steps and corresponding observations.

[0067] Next, convert the divided time window data into the tensor format required by the model. Based on the input requirements of the Mamba2 model, organize each time window data into a tensor. Assuming each time window has L time steps and each time step has D features (i.e., the dimensions of the joint features), each time window data can be represented as a tensor with a shape of (L, D).

[0068] Finally, the data loader is used to batch read these tensor-formatted time window data from the dataset. The data loader reads B time window data at a time from the dataset, based on a set batch size B, and combines them into a batch tensor of shape (B, L, D), which is then input into the Mamba2 model. The batch reading functionality of the data loader improves the efficiency of model training and prediction, enabling the model to more efficiently process large-scale time series data.

[0069] Step 4: During the modeling process, a high-level feature representation of the time series is generated through the Mamba2 model, which contains rich feature information and dynamic change patterns in the sequence. The high-level feature representation is input into the encoder-decoder structure of the Transformer model to mine deep temporal dependencies. The Mamba2 model in Step 4 includes convolutional layers and recurrent neural network layers for extracting local temporal features. The Transformer model adopts a multi-head self-attention mechanism to capture global temporal dependencies by calculating the scaled dot product attention weights of the query, key, and value matrices.

[0070] The step 4 specifically includes the following steps:

[0071] 4.1. The initial processing flow of the Mamba2 model for input time series data is complex and sophisticated. First, the convolutional layer performs a convolution operation on the input time series data. In the temporal dimension, the convolution kernel slides and interacts locally with the data, integrating features from adjacent time steps and extracting local features. Next, the parameter generation network dynamically generates parameters for state-space models (SSMs) using a small neural network based on the characteristics of the input sequence, allowing the model's behavior to adapt to the input data in real time. Next, the block-wise recursive algorithm layer divides the long sequence into fixed-size blocks. Each block is independently computed, and the output of the previous block serves as the initial state for the calculation of the next block. This recursive processing of long sequences effectively reduces computational complexity and improves memory efficiency. Finally, the multi-head mechanism layer introduces multiple heads to process the input data in parallel. Each head focuses on different aspects of the sequence, comprehensively capturing the various patterns and dependencies in the sequence, and enhancing the model's ability to express time series data.

[0072] 4.2. Preprocess the initially processed features: Specifically, add positional encoding; the positional encoding is generated in the form of sine and cosine functions. Next, the Mamba2 model may further process these initially extracted features to generate a more advanced representation. The role of positional encoding is to embed the temporal position information in the time series data into the feature representation, so that the Transformer model can distinguish data at different time steps. The positional encoding can be generated in the form of sine and cosine functions. For each time step t and each feature dimension i, the value of the i-th dimension of the positional encoding can be expressed as:

[0073]

[0074] Where H represents the feature dimension of the high-level representation, t represents the time step, and i represents the index of the feature dimension. The positional encoding is added to the high-level representation element by element to obtain a feature representation with added position information.

[0075] These high-level representations are then fed into the Transformer model. The core of the Transformer model is the self-attention mechanism, which can process the input sequence data in parallel and capture the relationship between any two positions in the sequence.

[0076] 4.3, the Transformer model processes the feature representation with positional encoding through its encoder part; by stacking multiple layers of encoders, the Transformer model gradually extracts deeper feature information and relationship patterns in the input sequence; each layer of encoder will further explore the complex structure and dependency relationships in the sequence based on its input, and finally generate a high-level feature representation that contains rich semantic information and dynamic rules. In actual use, the encoder is composed of multiple identical layers stacked together, and each layer contains two main sublayers: a multi-head self-attention mechanism sublayer and a feedforward neural network sublayer. In the multi-head self-attention mechanism sublayer, the input feature representation is first divided into multiple heads, and each head independently calculates the self-attention weight. For each head, the query (Q), key (K), and value (V) matrices are calculated, and then the self-attention weight is calculated by the following formula:

[0077]

[0078] Among them, d k The dimension of the key vector is represented by the softmax function, which normalizes the self-attention weights into a probability distribution. The self-attention calculation results of multiple heads are concatenated and passed through a linear transformation layer to obtain the output of the multi-head self-attention mechanism sub-layer. This output is then residually connected to the input and subjected to layer normalization to obtain the final output of the multi-head self-attention mechanism sub-layer.

[0079] Next, the output of the multi-head self-attention mechanism sublayer is input into the feedforward neural network sublayer. The feedforward neural network sublayer performs the same linear transformation and nonlinear activation operations on the feature representation of each position. Specifically, for the feature representation x of each position, the output of the feedforward neural network sublayer can be expressed as:

[0080] FFN(x)=max(0,xW1+b1)W2+b2

[0081] Where W1 and W2 are the weight matrices of the linear transformation, and b1 and b2 are the bias vectors. Through processing in the feedforward neural network sublayer, the feature representation can be further nonlinearly transformed, enhancing the model's expressive power. Similarly, the output of the feedforward neural network sublayer is residually connected to the input and layer normalized to obtain the final output of the encoder layer.

[0082] By stacking multiple layers of encoders, the Transformer model is able to gradually extract deeper feature information and relational patterns from the input sequence. Each encoder layer further explores the complex structure and dependencies within the sequence based on its input, ultimately generating a high-level feature representation that contains rich semantic information and dynamic patterns, providing a strong feature foundation for subsequent prediction tasks.

[0083] Step 5: After being processed by the encoder and decoder, the Transformer output is mapped to the prediction result space through a linear layer to complete the transformation from features to results. The prediction results are then denormalized and restored to the original numerical range.

[0084] In actual use, after the encoder and decoder of the Transformer model complete the feature extraction and modeling of the time series data, a linear layer is used to convert the feature representation of the model output into the target prediction result. Then, an inverse normalization operation is performed to restore the prediction result from the normalized range to the actual numerical range of the original data, thereby obtaining the final prediction result. The specific process is as follows:

[0085] 5.1. First, the decoder of the Transformer model processes the feature representations output by the encoder, ensuring that the decoder relies solely on previously generated information when generating its output at each time step. The decoder is also composed of multiple identical layers, each of which contains three main sublayers: a masked multi-head self-attention mechanism sublayer, an encoder-decoder multi-head self-attention mechanism sublayer, and a feedforward neural network sublayer. The masked multi-head self-attention mechanism sublayer is similar to the encoder's multi-head self-attention mechanism, but masks future information in the decoder input sequence to prevent information leakage during decoding. Specifically, a mask matrix is ​​added to the self-attention calculation, ensuring that each time step only focuses on information from the current and previous time steps. The mask matrix has the same shape as the self-attention weight matrix: elements below the diagonal are 0, indicating that the current time step cannot access information from future time steps, and elements above the diagonal are 1, indicating that the current time step can access information from previous time steps. This ensures that the decoder relies solely on previously generated information when generating its output at each time step.

[0086] 5.2, the output of the masked multi-head self-attention mechanism sublayer is then input into the encoder-decoder multi-head self-attention mechanism sublayer to interact the decoder output with the encoder output, allowing the decoder to utilize the global feature information extracted by the encoder to generate more accurate prediction results. In the encoder-decoder multi-head self-attention mechanism sublayer, the query (Q), key (K), and value (V) matrices are also calculated, but the query matrix comes from the decoder output, and the key matrix and value matrix come from the encoder output.

[0087] 5.3. Next, the output of the encoder-decoder multi-head self-attention mechanism sublayer is input into the feedforward neural network sublayer, where the same nonlinear transformation operation as the encoder is performed to further process and enhance the feature representation.

[0088] 5.4, ​​after processing by multiple layers of decoders, the final decoder output is obtained, which contains the prediction information for the time series data. A linear layer is then used to map the feature representation of the decoder output to the target prediction result space. The role of the linear layer is to convert the high-dimensional feature representation of the decoder output into a vector with the same dimension as the target prediction result. Assuming that the feature representation of the decoder output is a tensor with a shape of (L', D'), and the dimension of the target prediction result is (M), then the linear layer can be converted using the following formula:

[0089] y=xW+b

[0090] Among them, x represents the feature representation of the decoder output, W is the weight matrix of the linear layer, whose shape is (D',M), b is the bias vector, whose shape is M, and y represents the output of the linear layer, that is, the target prediction result.

[0091] 5.5. Finally, the prediction results are denormalized to restore them from the normalized range to the actual numerical range of the original data. Since the data was normalized by minimum and maximum normalization in step 2, the denormalization operation can be achieved by multiplying the normalized prediction results by the original data range and adding the minimum value, that is, the following formula:

[0092] yoriginal=ynormalized×(max(yoriginal)-min(yoriginal))+min(yoriginal)(0.1)

[0093] Among them, y normalized represents the normalized prediction result, max(yoriginal) and min(yoriginal) represent the maximum and minimum values ​​in the original data, respectively, and yoriginal represents the restored prediction result. By performing the inverse normalization operation, the prediction result can be converted to the same numerical range as the original data, thus obtaining a practical prediction value.

[0094] Step 6: Use a one-dimensional convolutional neural network (CNN) to correct the prediction error and output the final optimized daily runoff forecast. The CNN model in step 6 includes a one-dimensional convolutional layer, a pooling layer, and a fully connected layer. The convolution kernel is used to extract prediction error features, the pooling layer compresses the feature dimensions, and the fully connected layer outputs the corrected prediction result.

[0095] In actual use, after obtaining the preliminary prediction results of the Transformer model, the convolutional neural network (CNN) model is further used to adjust and optimize the prediction results, improve the accuracy of the prediction results through error correction, and enhance the stability of the model. The specific process is as follows:

[0096] First, the Transformer model's initial predictions are fed into the CNN model. The CNN model can use a one-dimensional convolutional architecture, with the time step as the convolution dimension. Assume the initial predictions are a tensor of shape (L, M), where (L) represents the number of time steps and (M) represents the dimension of the target prediction.

[0097] The CNN model then convolves the initial predictions through its convolutional layers. These layers extract local features from the initial predictions and smooth and adjust them. For example, a one-dimensional convolutional layer with a kernel size of 3 can be used. This layer slides the kernel along the timestep dimension, performing weighted summation of the predictions within each local region to generate a new feature representation.

[0098] Next, the output of the convolutional layer can be further processed, such as by downsampling through a pooling layer to reduce the time step dimension of the feature tensor while retaining important feature information. Pooling can be performed using methods such as maximum pooling or average pooling.

[0099] The pooled feature tensor is then fed into the fully connected layer for final prediction result adjustment and error correction. The fully connected layer integrates and transforms the feature information in the feature tensor to obtain the final prediction result. Assuming the dimension of the target prediction result is M, the fully connected layer can be converted using the following formula:

[0100] yfinal=x pooled W fc +b fc

[0101] Among them, xpooled represents the feature tensor after the pooling operation, Wfc is the weight matrix of the fully connected layer, and its shape is (C×L″,M), b fc is a bias vector with a shape of M, and yfinal represents the final prediction result after adjustment by the CNN model.

[0102] A daily runoff prediction system, comprising:

[0103] The data preprocessing module cleans and corrects raw runoff data, addressing missing values ​​and outliers to ensure data quality. It also extracts various statistical and time series features to reflect the underlying patterns and changing trends of the data. The module also smooths, detrends, and normalizes the data to reduce noise and scale variations, improving model training effectiveness. Furthermore, derived features are used to enhance the raw information, providing a more comprehensive and accurate input foundation for subsequent modeling and forecasting.

[0104] The hybrid deep learning module combines the multi-scale self-attention advantages of Mamba2, the global dependency modeling capabilities of Transformer, and the local feature extraction capabilities of CNN to comprehensively characterize hydrological runoff time series data. Specifically, Mamba2 excels in capturing complex nonlinear temporal relationships, Transformer improves the capture of long-range dependencies, and CNN enhances sensitivity to local temporal features. The integration of these three allows the model to more accurately reflect the multi-level and multi-scale characteristics of runoff series, significantly improving the robustness and accuracy of predictions.

[0105] The Mamba2 model contains convolutional layers and recurrent neural network layers to extract local temporal features;

[0106] The Transformer model adopts a multi-head self-attention mechanism to capture global temporal dependencies by calculating the scaled dot product attention weights of the query, key, and value matrices. In the Transformer model, both the encoder and decoder contain 6 layers, each layer contains 8 attention heads, and the hidden layer dimension is 512.

[0107] The CNN model consists of a one-dimensional convolutional layer, a pooling layer, and a fully connected layer. The convolution kernel is used to slide over the prediction error features, the pooling layer compresses the feature dimensions, and the fully connected layer outputs the corrected prediction results. In this solution, the CNN model is used to correct prediction errors and improve prediction accuracy. The one-dimensional convolutional layer within the model slides over the prediction error data sequence with a convolution kernel of size 3, taking a weighted sum of the error data for three consecutive time steps and adding a bias to extract local features. Next, a pooling layer with a window size of 2 compresses the output features of the convolution layer, using methods such as maximum or average pooling to reduce the feature dimensions, reduce the amount of computation, and prevent overfitting. Finally, the fully connected layer integrates and maps the feature vectors output by the pooling layer and outputs the corrected prediction results, providing an accurate decision-making basis for flood prevention and water resource scheduling.

[0108] The result output module is used to generate and visualize the final prediction results.

[0109] A computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the device containing the computer-readable storage medium executes the daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction as described above. The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory, and other memories.

[0110] An electronic device includes: a memory and a processor, wherein the memory stores a program that can be run on the processor, and when the processor executes the program, it implements the daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction as described above.

[0111] If the modules / units integrated in the electronic device described in this application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the processes in the above-mentioned embodiment methods by instructing the relevant hardware devices to complete them through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments.

[0112] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.

[0113] Computer-readable instructions are stored in the computer-readable storage medium, and the computer-readable instructions are executed by a processor in the electronic device to implement the daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction described in any of the above embodiments.

[0114] The following specific example is used to further explain the above. An important tributary of the Yellow River in China is taken as the research object. The tributary basin has complex terrain, significant seasonal differences in precipitation, and drastic changes in runoff, which poses challenges to local water resources management and flood control. The present invention provides a runoff intelligent prediction method based on GWO-VMD, Mamba2 model, Transformer model and CNN model, such as Figure 1 As shown, it mainly includes the following steps:

[0115] Step 1: Identify the study area and collect data

[0116] A specific river section of an important tributary of the Yellow River was assumed as the research basin, and daily hydrological and meteorological data of the basin from 2015 to 2025 were collected, including precipitation, runoff, water level, temperature, humidity, etc.

[0117] Step 2: Data preprocessing and exception handling

[0118] During the data preprocessing stage, the collected data were quality checked, and more than 2% of abnormal precipitation data (such as data that clearly exceeded the historical extreme range) were discovered and eliminated. Missing runoff data were filled using linear interpolation to ensure data integrity and accuracy.

[0119] Step 3: GWO-VMD modal decomposition and feature extraction

[0120] The original runoff series was decomposed using VMD technology optimized by GWO. The VMD mode number K was set to 5, the balance parameter α was set to 2000, and the number of iterations was set to 50. The population size of the GWO algorithm was set to 30, and the maximum number of iterations was set to 100. After VMD decomposition optimized by GWO, five intrinsic mode functions (IMFs) were obtained. Among them, the first IMF mainly reflects the long-term trend of runoff, the second IMF is closely related to seasonal changes, the third IMF is related to short-term fluctuations, and the remaining IMFs contain more complex noise information. The energy, variance and other characteristics of these IMFs are extracted and fused with the original runoff series to form a new joint feature as the model input.

[0121] Step 4: Data normalization and partitioning

[0122] The minimum and maximum normalization method is used to normalize the joint feature data to the range of [0,1]. The specific formula is:

[0123] x′=max(x)-min(x)x-min(x)

[0124] Here, x is the eigenvalue in the original data. Subsequently, the normalized data is randomly divided into training set, validation set, and test set in a ratio of 7:2:1.

[0125] Step 5: Time window division and data loading

[0126] The time window length is set to 30 days. A sliding window is applied based on the time step, dividing the time series data into fixed-length time windows. Each window contains 30 days of runoff data and its associated features. Using the data loader, the divided time window data is batch read and converted into the tensor format required by the model for input into the Mamba2 model.

[0127] Step 6: Mamba2 model processing and Transformer modeling

[0128] The Mamba2 model performs preliminary processing on input time series data, extracting local features and temporal dependencies to generate high-level representations. These high-level representations are then fed into the Transformer model, where the encoder and decoder further explore deep relationships and patterns. In the Transformer model, both the encoder and decoder contain six layers, each with eight attention heads, and a hidden layer dimension of 512.

[0129] Step 7: Generate prediction results and denormalize

[0130] The output of the Transformer model is mapped to the target prediction result space through the linear layer to obtain the preliminary prediction result. The inverse normalization formula is used to restore the prediction result from the range [0,1] to the actual numerical range of the original data:

[0131] yoriginal=ynormalized×(max(yoriginal)-min(yoriginal))+min(yoriginal)

[0132] Among them, ynormalized is the normalized prediction result.

[0133] Step 8: CNN model adjustment and error correction

[0134] A CNN model consisting of a one-dimensional convolutional layer, a pooling layer, and a fully connected layer was constructed to adjust the preliminary prediction results. The convolution kernel size of the convolutional layer was 3, and the window size of the pooling layer was 2. The prediction results were error corrected using the CNN model, and the optimized model parameters included the convolutional layer weight matrix and the fully connected layer weight matrix. After the model training was completed, the prediction results were evaluated, and the results showed that for a runoff peak event in 2025, the method of the present invention accurately predicted the runoff peak 7 days in advance, with a peak flow of 1500m 3 / s, and the error with the actual observed value is within 3%.

[0135] Model training and prediction

[0136] The model was trained using data from 2015 to 2024. The model parameters included the optimization parameters of GWO-VMD, the structural parameters of the Mamba2 model, the hyperparameters of the Transformer model, and the convolution kernel size and pooling window size of the CNN model. After training, the model was used to predict runoff on data from 2025. The model prediction results showed that for a runoff peak event in 2025, the proposed method accurately predicted the runoff peak seven days in advance, with a peak flow of 1500m 3 / s, and the error with the actual observed value is within 3%.

[0137] This example demonstrates the application of the method in predicting runoff for a tributary of the Yellow River. Using hypothetical data and parameters, it demonstrates that the method can effectively improve the accuracy and interpretability of runoff predictions, providing strong technical support for water resources management and flood prevention and disaster reduction in the basin.

[0138] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and other division methods may be used in actual implementation.

[0139] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0140] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

[0141] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products or apparatuses.

[0142] Note that the above are only preferred embodiments of the present invention and the principles of the technology used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention is described in detail through the above embodiments, the present invention is not limited to the specific embodiments described herein. Without departing from the concept of the present invention, it may also include many other effective embodiments, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction, characterized by: The following steps are involved: Step 1: Use the Grey Wolf Optimization (GWO) algorithm to optimize the variational mode decomposition (VMD) parameters, then decompose the original runoff sequence to obtain multiple intrinsic mode functions (IMFs). Extract the statistical characteristics and time-frequency domain features of each IMF, aggregate them into a feature set, and fuse the original runoff sequence with the feature set to form a joint feature. Step 2: Divide the joint features into training set, validation set and test set, and use minimum and maximum normalization to normalize the data to the [0,1] interval; Step 3: Divide the normalized time series data into fixed-length time windows and generate tensor format to input into the Mamba2 model; Step 4: Generate high-level feature representations of the time series through the Mamba2 model and input the high-level feature representations into the encoder-decoder structure of the Transformer model to mine deep temporal dependencies; Step 5: Map the Transformer output to the prediction result space through a linear layer, and perform denormalization on the prediction result to restore it to the original numerical range; Step 6: Use the one-dimensional convolutional neural network (CNN) to correct the error of the prediction results and output the final optimized daily runoff prediction value.

2. The daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction according to claim 1 is characterized by: In the step 1, optimizing the variational mode decomposition (VMD) parameters using the Grey Wolf Optimization (GWO) algorithm specifically includes the following steps: Set the number of modes, balanced reconstruction error, modal bandwidth parameters, center frequency update step size of VMD, and the population size and maximum number of iterations of GWO; The energy entropy of each IMF is used as the fitness function, and the VMD parameters are iteratively updated through the individual positions of the gray wolves until the optimal parameter combination is obtained.

3. The daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction according to claim 1 is characterized by: In the step 2, the data in the joint special type is divided into a training set of 70%, a validation set of 20%, and a test set of 10%. Normalization is achieved by subtracting the minimum value of the feature and dividing it by the difference between the maximum and minimum values ​​of the feature.

4. The daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction according to claim 1 is characterized by: The length of the time window in step 3 is dynamically adjusted according to the historical runoff data period, and the data loader converts the window data into batch tensors and inputs them into the Mamba2 model.

5. The daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction according to claim 1 is characterized by: The step 4 specifically includes the following steps: 4.1, Mamba2 model performs preliminary processing on the input time series data: First, the convolutional layer performs a convolution operation on the input time series data. In the time dimension, the convolution kernel slides and interacts locally with the data, integrating the features of adjacent time steps and extracting local features. Next, the parameter generation network uses a small neural network to dynamically generate the parameters of the state-space model (SSM) based on the characteristics of the input sequence, allowing the model behavior to adapt to the input data in real time. After that, the block recursive algorithm layer divides the long sequence into fixed-size blocks. Each block is calculated independently, and the output of the previous block is used as the initial state for the calculation of the next block. This recursive processing of long sequences effectively reduces computational complexity and improves memory usage efficiency. Finally, the multi-head mechanism layer introduces multiple heads to process the input data in parallel. Different heads focus on different aspects of the sequence, comprehensively capturing the various patterns and dependencies in the sequence, and enhancing the model's ability to express time series data. 4.

2. Preprocess the preliminarily processed features: specifically, add position codes; The position encoding is generated in the form of sine and cosine functions. Specifically, for each time step t and each feature dimension i, the value of the i-th dimension of the position encoding can be expressed as: Where H represents the feature dimension of the high-level representation, t represents the time step, and i represents the index of the feature dimension. The position encoding is added to the high-level representation element by element to obtain the feature representation with added position information. 4.

3. The Transformer model processes the feature representation with positional encoding through its encoder part; by stacking multiple layers of encoders, the Transformer model gradually extracts deeper feature information and relationship patterns in the input sequence; each layer of encoder will further explore the complex structure and dependency relationships in the sequence based on its input, and ultimately generate a high-level feature representation that contains rich semantic information and dynamic laws.

6. The daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction according to claim 1 is characterized by: The step five specifically includes the following steps: 5.

1. First, the decoder of the Transformer model processes the feature representation of the encoder output; this ensures that the decoder can only rely on previously generated information when generating the output of each time step; 5.2, then, the output of the masked multi-head self-attention mechanism sublayer is input into the encoder-decoder multi-head self-attention mechanism sublayer to interact the decoder output with the encoder output, so that the decoder can use the global feature information extracted by the encoder to generate more accurate prediction results; 5.

3. Next, the output of the encoder-decoder multi-head self-attention mechanism sublayer is input into the feedforward neural network sublayer, where the same nonlinear transformation operation as the encoder is performed to further process and enhance the feature representation. 5.

4. After processing by multiple layers of decoders, the final decoder output is obtained, which contains the prediction information for the time series data. A linear layer is then used to map the feature representation of the decoder output to the target prediction result space. 5.5, finally, an inverse normalization operation is applied to the prediction results to restore them from the normalized range to the actual numerical range of the original data.

7. The daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction according to claim 6 is characterized by: The denormalization operation in step five is specifically achieved by multiplying the normalized prediction result by the original data range and adding the minimum value.

8. A daily runoff intelligent forecasting system based on hybrid feature optimization and variational prediction, characterized by: include: The data preprocessing module is used to clean and correct raw runoff data, handle missing values ​​and outliers, and ensure data quality. It also extracts multiple statistical and time series features to reflect the basic laws and changing trends of the data. The module also smooths, detrends, and normalizes the data to reduce noise and scale differences, thereby improving model training results. In addition, derived features are used to enhance original information, providing a more comprehensive and accurate input basis for subsequent modeling and prediction; The hybrid deep learning module combines the multi-scale self-attention advantages of Mamba2, the global dependency modeling capabilities of Transformer, and the local feature extraction capabilities of CNN to comprehensively characterize hydrological runoff time series data. Specifically, Mamba2 excels in capturing complex nonlinear temporal relationships, Transformer improves the capture of long-range dependencies, and CNN enhances sensitivity to local temporal features. The integration of these three enables the model to more accurately reflect the multi-level and multi-scale characteristics of runoff series, significantly improving the robustness and accuracy of predictions. The Mamba2 model contains convolutional layers and recurrent neural network layers to extract local temporal features; The Transformer model uses a multi-head self-attention mechanism to capture global temporal dependencies by calculating the scaled dot product attention weights of the query, key, and value matrices; The CNN model consists of a one-dimensional convolutional layer, a pooling layer, and a fully connected layer. The convolution kernel is used to slide and extract the prediction error features. The pooling layer compresses the feature dimensions. The fully connected layer outputs the corrected prediction results for executing steps 4 to 6. The result output module is used to generate and visualize the final prediction results.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the device where the computer-readable storage medium is located executes the daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction as described in any one of claims 1 to 7.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a program that can be run on the processor, and when the processor executes the program, it implements the daily runoff intelligent forecasting method based on hybrid feature optimization and variational prediction as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Flood forecasting method and device based on intelligent optimization neural network

    CN121303477A

  • Flood forecasting method and device based on intelligent optimization neural network

    CN121303477B