Runoff forecasting method and system based on improved generative adversarial network
By improving the combination of generative adversarial network and variational autoencoder, the data complexity and randomness of existing runoff forecast models are solved, and the accuracy of runoff forecast and the stability of the model are improved.
Patent Information
- Application Number
- CN202510145627.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The hydrological data based on the existing runoff forecast model is complex and has great randomness during training, which makes it difficult to achieve accurate runoff forecasting.
Using the runoff forecasting method based on the improved generative adversarial network, by establishing the initial sample set, training the variant autoencoder model to generate reconstruction samples, combining the original data to form the reconstruction sample set, and using this sample set to train the generation adversarial network forecasting model.
It enhances the model's understanding of complex data structures, deeply explores the potential distribution rules of the initial sample data, improves the prediction accuracy of the prediction model after training, and reduces the training difficulty.
Smart Images

Figure CN120180858A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to runoff forecasting, and more specifically, relates to a runoff forecasting method and system based on an improved generative adversarial network. Background Art
[0002] In the field of runoff forecasting, the accuracy of runoff forecasting is crucial for optimizing reservoir operation strategies. However, the randomness and unpredictability of reservoir runoff pose great challenges to forecasting models. Traditional statistical and physical models often struggle to capture complex non-linear relationships and long-term dependence characteristics. To address these challenges, machine learning and deep learning technologies have been widely applied, which are mainly based on the relationship between input data and the observed object without including any artificially set physical parameters. In recent years, with the continuous development of machine learning technologies, learning-based models have become increasingly popular and their performance has been getting better and better.
[0003] However, learning-based data-driven models usually require a large amount of complete data to support model training, and the forecasting accuracy of the forecasting model is highly correlated with the specific data in the training dataset. Therefore, how to use limited hydrological data for runoff prediction has important research significance. The hydrological data usually has a complex data composition and large randomness when training existing runoff forecasting models, which brings difficulties to the training of forecasting models and makes it difficult to achieve accurate runoff forecasting. Summary of the Invention
[0004] In view of the above defects or improvement requirements of the prior art, the present invention provides a runoff forecasting method and system based on an improved generative adversarial network, which is used to solve the problem that the hydrological data usually has a complex data composition and large randomness when training existing runoff forecasting models, bringing difficulties to the training of forecasting models and making it difficult to achieve accurate runoff forecasting.
[0005] To achieve the above object, according to one aspect of the present invention, a runoff forecasting method based on an improved generative adversarial network is provided, including:
[0006] Establish an initial sample set according to historical runoff data;
[0007] Use the initial sample set to train a variational autoencoder model. The trained variational autoencoder model is used to generate reconstructed samples according to the initial samples in the initial sample set, and combine the initial samples and the reconstructed samples to form a reconstructed sample set;
[0008] Construct a forecasting model based on a generative adversarial network, and use the reconstructed sample set to train the forecasting model to obtain the trained forecasting model;
[0009] Use the trained prediction model to perform runoff prediction.
[0010] According to the runoff prediction method based on the improved generative adversarial network provided by the present invention, the specific steps of training the variational autoencoder model using the initial sample set include:
[0011] Obtain the KL divergence during the training process to measure the difference between the latent space distribution of the encoder output of the variational autoencoder model and the standard normal distribution;
[0012] Obtain the reconstruction loss during the training process to measure the difference between the reconstructed samples and the initial samples;
[0013] According to the KL divergence and the reconstruction loss, obtain the total loss, and use the total loss to train the variational autoencoder model.
[0014] According to the runoff prediction method based on the improved generative adversarial network provided by the present invention, the KL divergence The calculation formula is:
[0015]
[0016] In the formula, μ i represents the mean of the i-th latent variable; represents the variance of the i-th latent variable; D represents the dimension of the latent space; log(σ 2 ) represents the logarithmic variance;
[0017] The reconstruction loss The calculation formula is:
[0018]
[0019] In the formula, M is the number of initial samples; N is the number of features in the reconstructed samples; x ij represents the feature value of the j-th compressed representation of the i-th initial sample; represents the j-th reconstructed feature value in the reconstructed sample corresponding to the i-th initial sample;
[0020] The total loss The calculation formula is:
[0021]
[0022] According to the runoff prediction method based on the improved generative adversarial network provided by the present invention, the prediction model includes a generator and a discriminator. The generator is constructed by a gated recurrent unit, a bidirectional long short-term memory network and an attention mechanism, and the discriminator is constructed by a convolutional neural network.
[0023] According to the runoff forecasting method based on an improved generative adversarial network provided by the present invention, a Dropout layer is provided after the gated recurrent unit and the bidirectional long short-term memory network in the generator, and three fully connected layers are provided after the attention mechanism; the generator first processes the input data using the gated recurrent unit, and the output of the gated recurrent unit is passed to the bidirectional long short-term memory network. The output of the bidirectional long short-term memory network integrates the attention mechanism, and the features processed by the attention mechanism are gradually mapped and dimension-reduced through three fully connected layers, and finally mapped to an output dimension to generate a prediction result.
[0024] The discriminator includes four one-dimensional convolutional layers, a flattening layer, and three fully connected layers. The output of the convolutional layer is flattened through the flattening layer into a one-dimensional tensor suitable for processing by the fully connected layer, and then the discriminant result is output through three fully connected layers.
[0025] According to the runoff forecasting method based on an improved generative adversarial network provided by the present invention, training the forecasting model using the reconstructed sample set specifically includes:
[0026] Training the generator using the loss function of the generator, where the calculation formula of the loss function of the generator is as follows:
[0027]
[0028] In the formula, represents the generator loss; P g represents the sample distribution generated by the generator; D(x) represents the score of the discriminator for the input sample x.
[0029] According to the runoff forecasting method based on an improved generative adversarial network provided by the present invention, training the forecasting model using the reconstructed sample set specifically includes:
[0030] Training the discriminator using the loss function of the discriminator, where the calculation formula of the loss function of the discriminator is as follows:
[0031]
[0032] In the formula, represents the discriminator loss; P r represents the distribution of real data; D(x) represents the score of the discriminator for the input sample x; λ represents the coefficient of the gradient penalty term; represents the gradient penalty term; represents maximizing the score of the discriminator for real data; the second part represents minimizing the score of the discriminator for the generated data.
[0033] According to the runoff forecasting method based on an improved generative adversarial network provided by the present invention, the gradient penalty term is used to constrain the gradient norm of the discriminator, and its formula is as follows:
[0034]
[0035] In the formula, is the interpolation between the generated sample and the real sample, representing the random interpolation between the generator and the real data; represents the gradient norm of the discriminator on the interpolated sample; represents the distribution of the interpolated sample.
[0036] According to the runoff forecasting method based on an improved generative adversarial network provided by the present invention, the initial sample includes historical runoff data and corresponding feature data, and the feature data includes multiple data among average temperature, maximum temperature, minimum temperature, dew point temperature, rainfall, and historical daily-scale runoff of the previous 1 day, 2 days, and 3 days, the maximum and minimum daily-scale runoff of the current day.
[0037] In another aspect of the present invention, a runoff forecasting system based on an improved generative adversarial network is provided. The system includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it executes the runoff forecasting method based on an improved generative adversarial network described in any one of the above.
[0038] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the runoff forecasting method and system based on an improved generative adversarial network provided by the present invention:
[0039] 1. By constructing a variational autoencoder to reconstruct data and merging the original data into the subsequent forecasting model, the model's ability to understand complex data structures is enhanced, the potential distribution law of the initial sample data can be deeply mined, the potential correlation between different types of data in the initial sample data can be known, and reconstructed sample data with the same distribution law can be generated. The reconstructed sample data can be used to enhance the initial sample data. Therefore, training the forecasting model based on the reconstructed sample set is beneficial to reducing the training difficulty and improving the forecasting accuracy of the trained forecasting model;
[0040] 2. The forecasting model uses a gated recurrent unit and a bidirectional long short-term memory network to process the long-term and short-term dependencies in the runoff data, enhancing the dynamic modeling ability for time series data. The attention mechanism optimizes the model's ability to focus on important features by weighting key time steps in the input sequence, effectively improving the model's recognition of key features when processing complex time series data. By the discriminator's discrimination of the generated data, the generative adversarial training process is optimized, and the stability of the generative model is improved, thereby enhancing the forecasting performance of the runoff forecasting model;
[0041] 3. The application of the present invention to runoff prediction can improve the accuracy of runoff prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the structure diagram of the generative adversarial network provided by the present invention;
[0043] Figure 2 is the structure diagram of the variational autoencoder provided by the present invention;
[0044] Figure 3 is the flow chart of the runoff prediction method based on the improved generative adversarial network provided by the present invention;
[0045] Figure 4 is the structure diagram of the generator of the runoff prediction model based on the improved generative adversarial network provided by the present invention;
[0046] Figure 5 is the structure diagram of the discriminator of the runoff prediction model based on the improved generative adversarial network provided by the present invention;
[0047] Figure 6 is the comparison chart of the calculation results between the runoff prediction model based on the improved generative adversarial network provided by the present invention and other prediction models. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0049] Please refer to Figure 1 and Figure 2 , this embodiment provides a runoff prediction method based on an improved generative adversarial network, and the runoff prediction method based on the improved generative adversarial network includes:
[0050] Establish an initial sample set according to the historical runoff data;
[0051] Use the initial sample set to train the variational autoencoder model. The trained variational autoencoder model is used to generate reconstructed samples according to the initial samples in the initial sample set, and combine the initial samples and the reconstructed samples to form a reconstructed sample set;
[0052] Construct a prediction model based on the generative adversarial network, and use the reconstructed sample set to train the prediction model to obtain the trained prediction model;
[0053] Use the trained prediction model to perform runoff prediction.
[0054] In the prior art, generative adversarial networks are mainly applied to research fields such as image generation, data augmentation, natural language processing, image processing, and speech recognition. Currently, there are few runoff prediction methods based on improved generative adversarial networks. In this embodiment, a prediction model is constructed based on a generative adversarial network. During the training of the generative adversarial network, through the adversarial training of the generator and the discriminator, the accuracy of the runoff prediction data generated by the generator can be guaranteed; and this embodiment proposes to construct a variational autoencoder to reconstruct and enhance the initial samples. Through the variational autoencoder, the potential distribution law of the initial sample data can be deeply mined, the association between different types of data in the initial sample data can be better understood, and reconstructed sample data with the same distribution law can be generated. The initial sample data can be enhanced using the reconstructed sample data, and the sample data volume can be expanded. Therefore, training the prediction model based on the reconstructed sample set is beneficial to reducing the training difficulty and improving the prediction accuracy of the trained prediction model.
[0055] In some specific embodiments, the purpose of this embodiment is to propose a runoff prediction method based on an improved generative adversarial network. By introducing a variational autoencoder to learn the latent features of the data, reconstruct the data and merge the original data into the input of the prediction model, weight the key time steps in the input sequence by fusing the attention mechanism, couple the bidirectional long short-term memory network and the gated recurrent unit as the generator of the generative adversarial network, and use four one-dimensional convolutional layers and three fully connected layers as the discriminator to predict the runoff. Finally, compare the runoff output value calculated by the simulation with the actual runoff to evaluate the advantages and disadvantages of the runoff prediction method. The specific steps are as follows:
[0056] Step1: Data loading and preprocessing. First, load the initial sample set data, and divide the initial sample set data into a training set and a test set according to the ratio of 80% and 20%. To improve the generalization ability and stability of the model, normalize the sequence data, that is, the data in each initial sample;
[0057] Step2: Construction and training of the variational autoencoder model (VAE) and feature enhancement. A variational autoencoder is introduced for the feature enhancement of the prediction factors. The variational autoencoder model consists of an encoder and a decoder. The encoder maps the input feature data, that is, the initial sample data, to the latent space. The latent space represents the compressed representation of the input data. Each dimension of the latent space corresponds to a latent variable, and the encoder outputs the mean and log variance of these latent variables. The decoder then reconstructs the feature input data according to the latent variables. During the training process, the variational autoencoder optimizes the model parameters by minimizing the total loss. The total loss consists of two parts: the KL divergence loss and the reconstruction loss. During training, the Adam optimizer is used for gradient update.
[0058] Training the variational autoencoder model using the initial sample set specifically includes: obtaining the KL divergence loss during training to measure the difference between the latent space distribution output by the encoder and the standard normal distribution; obtaining the reconstruction loss during training to measure the difference between the reconstructed samples reconstructed by the decoder and the initial samples. According to the KL divergence and the reconstruction loss, obtain the total loss, and use the total loss to train the variational autoencoder model.
[0059] Use the Adam optimizer for training, calculate the KL divergence loss and the reconstruction loss, obtain the total loss, perform backpropagation to calculate the gradients and update the model parameters, and determine whether the iteration termination condition is reached (set the maximum number of iterations or set the total loss threshold). If the termination condition is not met, restart the next iteration process. If the termination condition is met, input the reconstructed feature data, i.e., the reconstructed samples, and the original feature data, i.e., the initial samples, into the prediction model together;
[0060] The KL divergence loss and the reconstruction loss are shown in equations (1) - (3); the KL divergence is used to measure the difference between the latent space distribution output by the encoder and the standard normal distribution. For each sample data, the calculation formula for the KL divergence is:
[0061]
[0062] In the formula, μ i represents the mean of the i-th latent variable; represents the variance of the i-th latent variable; D represents the dimension of the latent space, i.e., the total number of latent variables; log(σ 2 ) represents the log variance.
[0063] In the variational autoencoder, each input initial sample is a vector x i = [x i1 , x i2 , …, x iL , where L is the number of features in the initial sample. Each sample is reconstructed by the decoder to obtain a vector composed of N features The VAE uses binary cross-entropy to calculate the reconstruction loss. The reconstruction loss formula is:
[0064]
[0065] In the formula, M is the number of initial samples; N is the number of features in the reconstructed samples, and the number of reconstructed features is the same as the number of latent variables; x ij represents the feature value of the j-th compressed representation of the i-th initial sample, i.e., the value of the j-th latent variable generated by the encoder for the i-th sample; Denote the j-th reconstructed eigenvalue in the reconstructed sample corresponding to the i-th initial sample.
[0066] The total loss function of VAE consists of two parts: reconstruction loss and KL divergence. The formula is as follows:
[0067]
[0068] Step3: Sliding window data preparation. For example, the window size is set to 3. Each input reconstructed sample includes data from the first 3 time steps, and the output is the target value at the 4th time step.
[0069] Step4: Design and training of the improved generative adversarial network model. The prediction model includes a generator and a discriminator. The generator is constructed using a gated recurrent unit (GRU) and a bidirectional long short-term memory network (BiLSTM) combined with an attention mechanism; the GRU layer is used to process the input data, and the output of the GRU layer is then passed to the BiLSTM integrated attention mechanism layer to generate data.
[0070] Specifically, referring to Figure 4 , the generator first uses the GRU layer to process the input data. The GRU has a strong ability to capture short-term dependencies and can effectively extract local features and temporal dependencies in sequence data. The output of the GRU layer is then passed to the BiLSTM layer. The BiLSTM captures long-term dependencies in the sequence by processing data in both the forward and backward directions simultaneously, improving the model's ability to model sequence data.
[0071] To further improve the performance of the model, the present invention integrates an attention mechanism based on the output of the BiLSTM layer. The working process of this attention mechanism includes a linear transformation of the BiLSTM output to calculate the attention scores for each time step. Subsequently, these scores are normalized to obtain the attention weights. These weights are used to perform a weighted sum of the BiLSTM output to capture the feature information of important time steps.
[0072] First, calculate the attention weights. The specific formula is shown in equations (4) and (5):
[0073] a t = tanh(W a h t + b a ) (4)
[0074] In the formula, h is the hidden state, h t is the hidden state of the BiLSTM network at the t-th time step, with dimension d; W a is a learnable weight matrix with dimension d×d; b a is a bias term with dimension d; at is the calculated attention score, with dimension d.
[0075] The scores at each time step are mapped to scalar values through another linear transformation, and then the normalized exponential function (softmax function) is applied to obtain the weights. The attention weights are the contribution ratios of each time step, so the sum of the weights of all time steps is 1.
[0076]
[0077] In the formula, is a learnable weight matrix with dimension d×1, used to map the attention scores to a scalar. t′ represents all time steps, and α t is the normalized attention weight, representing the contribution ratio of each time step to the final output. ∑ t′ exp(W c a t′ ) is the weighted sum of all time steps, used to normalize all weights.
[0078] After that, the obtained attention weights are used to perform a weighted sum of the input hidden states to obtain the final weighted output h att , and the specific formula is shown in Equation (6):
[0079]
[0080] In the formula, h att is the output after the weighted sum of h t , and it will be passed as the output after the attention mechanism to the subsequent layers of the model. α t is the attention weight at the t-th time step, representing the contribution degree of the t-th time step to the final output.
[0081] In the specific implementation, both the GRU layer and the bidirectional LSTM layer have 512 hidden units, and both are respectively equipped with a Dropout layer to prevent overfitting and enhance the robustness of the model, that is, a Dropout layer is provided after the gated recurrent unit and the bidirectional long short-term memory network in the generator. The design of the attention mechanism includes a linear layer and a context vector calculation layer, responsible for optimizing feature weighting. Finally, three fully connected layers are provided after the attention mechanism, and the features processed by the attention mechanism are gradually mapped and dimensionality-reduced through the three fully connected layers, with the dimensions decreasing from 512 to 128, then to 64, and finally mapped to an output dimension to generate the prediction result.
[0082] Reference Figure 5, and then a discriminator is constructed using a Convolutional Neural Network (CNN), the losses of the generator and the discriminator are calculated, and the Adam optimizer is used for optimization during the training process, and the parameters of the generator and the discriminator are updated; it is judged whether the iteration termination condition is reached. If the termination condition is not satisfied, the next iteration process is started, otherwise the optimal prediction value is output. Specifically, the architecture of the discriminator includes four one-dimensional convolutional layers, which have 32, 64, 128, and 256 convolutional kernels respectively; then it includes a flattening layer and three fully connected layers. Subsequently, the output of the convolutional layer undergoes a flattening operation through the flattening layer to be flattened into a one-dimensional tensor suitable for processing by the fully connected layer, and then the discriminant result is output after passing through three fully connected layers. The fully connected layer includes three levels, which have 220, 220, and 1 neuron respectively.
[0083] Lipschitz continuity is a way to describe the rate of change of a function. For a function f: R n →R m (from an n-dimensional real space to an m-dimensional real space), if there exists a constant L such that for all x, y ∈ R n , the following formula (7) is satisfied:
[0084] ∥f(x) - f(y)∥ ≤ L∥x - y∥ (7)
[0085] Then this function is called Lipschitz continuous, where L is the Lipschitz constant, which measures the rate of change of the function. When L = 1, this function is said to satisfy the 1-Lipschitz condition. That is, for any two inputs x and y, the following formula (8) is satisfied:
[0086] ∥f(x) - f(y)∥ ≤ ∥x - y∥ (8)
[0087] The 1-Lipschitz condition requires that the output change of the function cannot exceed the amplitude of the input change. This condition has an important impact on the smoothness of the function. It ensures that the function does not have violent fluctuations and is usually used to ensure the stability of the optimization problem.
[0088] In WGAN-GP, the goals of the generator and the discriminator are to minimize the Wasserstein distance, which measures the difference between two distributions, and the formula is shown in (9):
[0089]
[0090] In the formula, W(P r , P g ) represents the Wasserstein distance, and P r represents the distribution of the real data. P grepresents the distribution of the generated data. f represents the output of the discriminator. Specifically, the value output by f(x) represents the degree to which the sample x is judged to be real data. ∥f∥ L ≤1 means that the function f must satisfy the Lipschitz condition, that is represents taking the supremum. That is, among all functions f that satisfy the Lipschitz condition, the function that maximizes the expected difference is selected. That is, for the real data distribution P r the expected value of the function f(x) is calculated for all samples under it, that is, the weighted average of the scores of the real data. x ∼ P r represents sampling from the real data distribution P r among them. represents calculating the expected value of the function f(x) for all samples under the generated data distribution P g that is, the weighted average of the scores of the generated data. x ∼ P r represents sampling from the real data distribution P g among them.
[0091] The training of the generator and the discriminator includes two main loss functions: the generator loss and the discriminator loss. Specifically, training the prediction model using the reconstructed sample set includes: training the generator using the loss function of the generator, and training the discriminator using the loss function of the discriminator.
[0092] The goal of the discriminator is to maximize the discrimination ability between real samples and generated samples. At the same time, to ensure the accuracy of the Wasserstein distance, the discriminator needs to satisfy the 1-Lipschitz condition. This condition ensures that the gradient of the discriminator will not be too large, thus avoiding instability in the optimization process.
[0093] The loss function of the discriminator includes the expected value of the real data, the expected value of the generated data, and the gradient penalty term. The specific formula of the loss function of the discriminator is shown in (10)(11).
[0094]
[0095] In the formula, represents the discriminator loss. P r represents the distribution of the real data. D(x) represents the score of the discriminator for the sample x. λ represents the coefficient of the gradient penalty term. represents the gradient penalty term. represents maximizing the score of the discriminator for the real data. The second part represents minimizing the score of the discriminator for the generated data. Among them, Denotes the expected value of the real data, that is, the average value of the function D(x) calculated for all samples from the real data distribution P r below. Denotes the expected value of the generated data, that is, the average value of the function D(x) calculated for all samples from the generated data distribution P g below.
[0096] Gradient penalty term Used to constrain the gradient norm of the discriminator to satisfy the 1-Lipschitz condition, and the formula is as follows:
[0097]
[0098] In the formula, is the interpolation between the generated sample and the real sample, representing the random interpolation between the generator and the real data. Denotes the gradient norm of the discriminator on the interpolated samples. Denotes the distribution of the interpolated samples, usually the interpolation distribution between the generated data and the real data.
[0099] The goal of the generator is to generate data as real as possible so that the discriminator considers these generated data to be real. Therefore, the loss function of the generator aims to maximize the score of the discriminator for the generated samples, and the specific formula of the loss function of the generator is shown in (12):
[0100]
[0101] In the formula, Denotes the generator loss. P g Denotes the sample distribution generated by the generator. D(x) denotes the score of the discriminator for the input sample x.
[0102] Step5: The model predicts and evaluates the results. Use the trained generative adversarial network to predict the runoff data, denormalize the output prediction values, compare the results with the original data, and conduct an evaluation.
[0103] To comprehensively evaluate the prediction performance and computational efficiency of the experimental design model, this embodiment selects five prediction evaluation indicators: Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), Coefficient of Determination (R2), Pearson correlation coefficient (r), and the formulas are as follows:
[0104]
[0105]
[0106]
[0107]
[0108]
[0109] In the formula, y i , denote observed values and predicted values respectively; are the mean of observed values and the mean of predicted values respectively; n is the length of the test set.
[0110] This specific example uses the hydrological element data of Yichang Station in the upper reaches of the Yangtze River as an example to verify the effect of the solution of the present invention. Figure 3 The overall process of the runoff forecasting method based on the improved generative adversarial network is shown to perform runoff prediction to reflect the effect achieved by the present invention.
[0111] The initial sample includes historical runoff data and corresponding feature data, and the feature data includes average temperature, maximum temperature, minimum temperature, dew point temperature, rainfall, and multiple data of the historical daily runoff of the previous 1 day, 2 days, and 3 days, and the maximum and minimum daily runoff of the current day. In this embodiment, the daily runoff data of Yichang Station from January 1, 2007 to December 31, 2021 are obtained to construct the initial sample set, and the feature data includes average temperature, maximum temperature, minimum temperature, dew point temperature, rainfall, and the historical daily runoff of the previous 1 day, 2 days, and 3 days, and the maximum and minimum daily runoff of the current day.
[0112] Step 1: Collect daily runoff data and characteristic data from Yichang Station from January 1, 2007 to December 31, 2021, remove missing values or outliers to ensure the integrity and accuracy of the data. Normalize all data to be in the interval [0, 1] to eliminate the influence of different dimensions. Divide the data set into a training set and a test set, with the training set accounting for 80% of the data and the remaining 20% as a test set.
[0113] Step 2: Construct a variational autoencoder model. The encoder maps the input to a 5-dimensional latent space through a multi-layer fully connected network, with each dimension having a latent variable. The decoder then maps the representation in the latent space back to the input space. The model uses the reparameterization trick to sample from the latent space with an approximate Gaussian distribution. The sum of the cross-entropy loss and the KL divergence loss is used as the training objective. The KL divergence term ensures that the latent distribution output by the encoder is as close as possible to the standard normal distribution. The Adam optimizer is used for training, and the learning rate is set to 1×10 -5 , the batch size is 128, and the number of training iterations is 300. The reconstruction loss and the KL divergence of the model are calculated in each round of training, and the parameters are updated through backpropagation. The trained VAE model is used to enhance the features, generating 5 new feature factors, which are then input into the prediction model together with the 10 feature factors of the original dataset.
[0114] Step 3: Use the sliding window method to convert the time series data into a format suitable for training the prediction model. The sliding window size is set to 3, and each input sample includes the data of the first 3 time steps, with the output being the target value at the 4th time step.
[0115] Step 4: Refer to Figure 4 and Figure 5 , the generator uses GRU and BiLSTM structures. The input is the enhanced feature data, and the output is the predicted target variable. The attention mechanism is used to perform weighted combination on the time series data to improve the prediction accuracy. The output dimension of the GRU layer is 512, and the output dimension of the bidirectional LSTM is 256. After passing through the attention layer, finally, 1 target value is output through the Linear layer. The discriminator adopts a convolutional neural network structure, using 4 convolutional layers to extract local features in the time series data. Finally, the true / false discrimination result is output through the fully connected layer. The number of output channels of the convolutional layers are 32, 64, 128, and 256 respectively, and finally the true / false determination result is output through 3 fully connected layers. The generator and the discriminator are updated alternately. The generator is updated once after training the discriminator 5 times. The Adam optimizer is used for optimization, and the learning rate is 1×10 -4 , the batch size is 128, and the number of training iterations is 100.
[0116] Step 5: Use the trained generator to predict the test set data and generate runoff prediction values. Denormalize the predicted latent features and convert them back to the scale of the original data. To evaluate the effectiveness of the forecasting model proposed in the present invention, in this embodiment, the long short-term memory network (LSTM); bidirectional long short-term memory network (BILSTM); gated recurrent unit (GRU); Wasserstein generative adversarial network with gradient penalty (WGAN-GP); generative adversarial network integrating gated recurrent unit, bidirectional long short-term memory network and gradient penalty (GBWGAN-GP); generative adversarial network integrating gated recurrent unit, bidirectional long short-term memory network, convolutional neural network and gradient penalty (CGBWGAN-GP); generative adversarial network integrating gated recurrent unit, bidirectional long short-term memory network, convolutional neural network, attention mechanism and gradient penalty (CGBWGAN-GP-ATT) are selected as experimental comparison models, and CGBWGAN-GP-ATT-VAE is the model of the method proposed in the present invention, which integrates gated recurrent unit, bidirectional long short-term memory network, convolutional neural network, variational autoencoder and generative adversarial network with gradient penalty. The prediction results are as Figure 6 shown. The prediction performance of the model is evaluated using indicators such as root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE), coefficient of determination (R2), and Pearson correlation coefficient (r). The calculation results table of the runoff forecasting model based on the improved generative adversarial network and other prediction models is shown in Table 1. Compared with the CGBWGANGP-ATT model, the coefficient of determination of the model proposed in the present invention increased by about 0.633%, and the MAE value decreased by about 225.854; entering the test period, compared with the CGBWGANGP-ATT model, the coefficient of determination of the model proposed in the present invention decreased by about 1.07%, and the MAE value decreased by about 296.719. This further confirms the advantage of introducing feature enhancement. Generally speaking, the runoff forecasting method based on the improved generative adversarial network has good performance in the runoff forecasting of Yichang Station.
[0117] Table 1 Calculation results table of the runoff forecasting model based on the improved generative adversarial network and other models
[0118]
[0119]
[0120] This embodiment of the present invention also provides a runoff forecasting system based on an improved generative adversarial network. The system includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it executes the runoff forecasting method based on the improved generative adversarial network described in any one of the above.
[0121] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A runoff forecasting method based on an improved generative adversarial network, characterized in that: include: Establish an initial sample set based on historical runoff data; The variational autoencoder model is trained using the initial sample set, the trained variational autoencoder model is used to generate reconstructed samples according to the initial samples in the initial sample set, and the initial samples and the reconstructed samples are combined to form a reconstructed sample set; Constructing a prediction model based on a generative adversarial network, training the prediction model using the reconstructed sample set, and obtaining the trained prediction model; The trained forecast model is used to forecast runoff.
2. The runoff forecasting method based on the improved generative adversarial network according to claim 1, characterized in that: Using the initial sample set to train the variational autoencoder model specifically includes: Obtaining the KL divergence during the training process is used to measure the difference between the potential space distribution of the encoder output of the variational autoencoder model and the standard normal distribution; Obtaining the reconstruction loss during training is used to measure the difference between the reconstructed sample and the initial sample; According to the KL divergence and the reconstruction loss, the total loss is obtained, and the variational autoencoder model is trained using the total loss.
3. The runoff forecasting method based on improved generative adversarial network according to claim 2, characterized in that: KL divergence The calculation formula is: In the formula, μ i represents the mean of the i-th latent variable; represents the variance of the ith latent variable; D represents the dimension of the latent space; log(σ 2 ) represents the logarithmic variance; Reconstruction loss The calculation formula is: Where M is the number of initial samples; N is the number of features in the reconstructed samples; x ij represents the eigenvalue of the jth compressed representation of the i-th initial sample; represents the jth reconstructed eigenvalue in the reconstructed sample corresponding to the i-th initial sample; Total loss The calculation formula is:
4. The runoff forecasting method based on an improved generative adversarial network according to any one of claims 1 to 3, characterized in that: The prediction model includes a generator and a discriminator, wherein the generator is constructed by a gated recurrent unit and a bidirectional long short-term memory network combined with an attention mechanism, and the discriminator is constructed by a convolutional neural network.
5. The runoff forecasting method based on improved generative adversarial network according to claim 4, characterized in that: A Dropout layer is provided after the gated recurrent unit and the bidirectional long short-term memory network in the generator, and three fully connected layers are provided after the attention mechanism; the generator first uses the gated recurrent unit to process the input data, and the output of the gated recurrent unit is passed to the bidirectional long short-term memory network. The output of the bidirectional long short-term memory network integrates the attention mechanism, and the features processed by the attention mechanism are gradually mapped and reduced in dimension through three fully connected layers, and finally mapped to an output dimension to generate a prediction result; The discriminator includes four one-dimensional convolutional layers, a flattening layer and three fully connected layers. The output of the convolutional layer is flattened through the flattening layer to be flattened into a one-dimensional tensor suitable for processing by the fully connected layer, and then passes through three fully connected layers to output the discrimination result.
6. The runoff forecasting method based on improved generative adversarial network according to any one of claims 1 to 3, characterized in that: Using the reconstructed sample set to train the prediction model specifically includes: The generator is trained using the generator's loss function, where the calculation formula of the generator's loss function is as follows: In the formula, represents the generator loss; P g represents the sample distribution generated by the generator; D(x) represents the score of the discriminator for the input sample x.
7. The runoff forecasting method based on improved generative adversarial network according to claim 6, characterized in that: Using the reconstructed sample set to train the prediction model specifically includes: The discriminator is trained using the loss function of the discriminator, wherein the calculation formula of the loss function of the discriminator is as follows: In the formula, represents the discriminator loss; P r Represents the distribution of real data; D(x) represents the score of the discriminator for the input sample x; λ represents the coefficient of the gradient penalty term; represents the gradient penalty term; Represents maximizing the score of the discriminator on the real data; the second part represents minimizing the score of the discriminator on the generated data.
8. The runoff forecasting method based on improved generative adversarial network according to claim 7, characterized in that: Gradient Penalty The gradient norm used to constrain the discriminator is as follows: In the formula, is the interpolation of the generated samples and the real samples, indicating the random interpolation between the generator and the real data; Represents the gradient norm of the discriminator on the interpolated samples; Represents the distribution of interpolation samples.
9. The runoff forecasting method based on improved generative adversarial network according to any one of claims 1 to 3, characterized in that: The initial sample includes historical runoff data and corresponding characteristic data, wherein the characteristic data includes average temperature, maximum temperature, minimum temperature, dew point temperature, rainfall, and multiple data of historical daily runoff of the previous 1 day, 2 days, and 3 days, and the maximum and minimum daily runoff of the current day.
10. A runoff forecasting system based on an improved generative adversarial network, characterized in that: The system includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the runoff forecasting method based on the improved generative adversarial network described in any one of claims 1 to 9 is executed.
Citation Information
Patent Citations
DDoS attack distinguishing method and system based on CVAE-WGAN-GP
CN118631562A
Multi-element sea wave forecasting method based on signal decomposition, electronic equipment and medium
CN118673388A
Lymph node CT detection system employing recurrent spatio-temporal attention mechanism
WO2020258611A1
Cited By
Low-voltage active transformer area line loss rate prediction method and system based on improved GAN
CN120822672A