CNN-Btransform power load prediction method based on WGAN enhanced data

Through the CNN-BiTransformer-BiLSTM method after WGAN enhanced data, the problems of insufficient accuracy and noise sensitivity in power load prediction are solved, and the load prediction with higher accuracy and robustness is achieved, and intelligent scheduling of the power system is supported.

CN120341836APending Publication Date: 2025-07-18SHENYANG INST OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510425454.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When traditional power load prediction methods deal with nonlinearity, time-varying and external factors, there are problems such as insufficient accuracy, limited long-term dependency modeling capabilities, and noise sensitivity.

Method used

WGAN augmentation data is used, combined with CNN, BiTransformer and BiLSTM methods to predict power loads. Through data augmentation, feature extraction, time-dependent modeling and deep feature learning, the final load prediction results are generated.

Benefits of technology

It improves the accuracy and robustness of power load prediction, enhances the model's adaptability and noise resistance to complex load modes, and provides reliable technical support for intelligent scheduling of power systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120341836A_ABST
    Figure CN120341836A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of power system load prediction, and particularly relates to a CNN-Btransform power load prediction method based on WGAN (Wide Generation Area Network) enhanced data. Performing data enhancement on the power load time sequence by adopting an adversarial network; extracting a feature map of the power load time sequence after data enhancement; the forward time dependence and the backward time dependence of bidirectional Transform modeling are used for extracting long-term time features from the feature map; learning deep features of the long-term features by using a bidirectional long-short-term memory network; and returning to a regression layer to generate a final load prediction result. According to the method, the power load prediction precision and stability under the condition of relatively less data can be effectively improved, and meanwhile, the anti-noise capability of the model and the adaptability to a complex load mode are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power system load forecasting, and particularly relates to a power load forecasting method based on CNN-BiTransformer-BiLSTM after enhancing data with WGAN. Background Art

[0002] With the continuous development of the power system and the increasing complexity of power demand, power load forecasting plays a crucial role in the planning, scheduling, and operation of the power system. Accurate power load forecasting can not only improve the stability and economy of power grid operation, but also optimize energy scheduling and reduce the uncertainty of power supply. However, due to the significant nonlinearity, time-varying nature of power load data and its susceptibility to external factors, traditional statistical methods and prediction methods based on shallow machine learning often face problems such as insufficient accuracy, limited ability to model long-term dependencies, and sensitivity to noise when dealing with these complexities. Summary of the Invention

[0003] The present invention provides a power load forecasting method based on CNN-Bitransformer after enhancing data with WGAN, which can effectively address the complexity of power load time series data and data missing problems to improve the accuracy and robustness of power load forecasting.

[0004] The present invention is implemented as follows:

[0005] A power load forecasting method, the method comprising:

[0006] Performing data enhancement on the power load time series using an adversarial network;

[0007] Extracting the feature map of the power load time series after data enhancement;

[0008] Using a bidirectional Transformer to model the forward and backward time dependencies to extract long-term time features from the feature map;

[0009] Adopting a bidirectional long short-term memory network to learn the deep features of the long-term time features;

[0010] Returning to the regression layer to generate the final load forecasting result.

[0011] Further, performing data enhancement on the power load time series using an adversarial network includes:

[0012] Using a generator G(z; θ G ) to accept a random vector z from a noise distribution P z to generate synthetic data The specific formula is:

[0013]

[0014] Among them, z represents the input random noise vector, which usually follows a standard normal distribution W is the weight matrix of the generator; b g is the bias term of the generator; f(·) is the ReLU function, which is used to introduce non-linear mapping;

[0015] The discriminator D(x; θ D ) is used to judge whether the input data is real data, and its mathematical expression is:

[0016] D(x) = g(W d x + b d )

[0017] Among them, x represents the input power load data; W d is the weight matrix of the discriminator; b d is the bias of the discriminator; g(·) uses the LeakyReLU activation function.

[0018] Furthermore, by minimizing the Wasserstein distance, the data distribution P G generated by the generator is made as close as possible to the real data distribution P data , and the objective function is:

[0019]

[0020] Among them, is the expected score of the real data on the discriminator; is the expected score of the generated data on the discriminator;

[0021] A gradient penalty term is added to the objective function:

[0022]

[0023] Among them, λ is the gradient penalty coefficient; represents the interpolation sample randomly sampled between the real data and the generated data; is the L2 norm of the gradient of the discriminator with respect to .

[0024] Furthermore, the feature map of the power load time series after data augmentation is extracted, and the calculation formula is:

[0025]

[0026] Among them, x t-i,d represents the input data at time step t - i and channel d; w i,d,cis the weight corresponding to the time step i, channel d, and output channel c in the convolutional kernel; b c is the bias of the output channel c; k is the time step window size of the convolutional kernel; D is the number of channels of the input data; y t,c is the eigenvalue of the output at time step t and output channel c.

[0027] Further, a bidirectional Transformer is used to model the forward and backward temporal dependencies to extract long-term temporal features from the feature map, including:

[0028]

[0029] where and represent the hidden states of the forward and backward Transformer units at time step t, respectively.

[0030] Further, a bidirectional long short-term memory network is used to learn the deep features of the long-term temporal features, including:

[0031] i t = σ(W i x t + U i h t-1 + b i )

[0032] f t = σ(W f x t + U f h t-1 + b f )

[0033] o t = σ(W o x t + U o h t-1 + b o )

[0034]

[0035] h t = o t ⊙ tanh(c t )

[0036] where x t represents the Transformer output feature at time step t; h t-1 is the hidden state of the previous time step; h t is the hidden state of the current time step; W * and U * are the weight matrices of the input and recurrent connections, respectively, b* is the bias term; σ(·) is the sigmoid activation function, and tanh(·) is the hyperbolic tangent activation function; ⊙ represents element-wise multiplication.

[0037] Further, return to the regression layer to generate the final load prediction result, including:

[0038] Adopt a multi-layer neural network structure and train it in combination with the mean squared error loss function. The calculation formula is:

[0039]

[0040] where W fc is the weight matrix of the fully connected layer; b fc is the bias of the fully connected layer; f FC (·) is a linear transformation, is the output.

[0041] A CNN-BiTransformer-BiLSTM power load prediction method based on data enhanced by WGAN. This method includes:

[0042] S1 Use the Wasserstein Generative Adversarial Network (WGAN) to perform data enhancement on the original power load data to improve the distribution consistency and generalization ability of the data;

[0043] S2 Adopt a Convolutional Neural Network (CNN) to extract local patterns in the power load time series to capture short-term features;

[0044] S3 Use a Bidirectional Transformer (BiTransformer) to model the forward and backward time dependencies to enhance the ability to capture long-term time features;

[0045] S4 Adopt a Bidirectional Long Short-Term Memory Network (BiLSTM) to further learn the deep features of the time series to enhance the adaptability of the model to complex load patterns;

[0046] S5 Generate the final load prediction result through the regression layer.

[0047] Further, in S1, the specific operations of the WGAN data enhancement module can be specifically divided into the design of the generator and the discriminator, where:

[0048] The generator G(z; θ G ) accepts the random vector z from the noise distribution P z and generates synthetic data Its specific formula is:

[0049]

[0050] Among them, z represents the input random noise vector, which usually follows a standard normal distribution W g is the weight matrix of the generator; b g is the bias term of the generator; f(·) is the ReLU function, which is used to introduce a non-linear mapping.

[0051] The discriminator D(x; θ D ) is used to judge whether the input data is real data, and its mathematical expression is:

[0052] D(x) = g(W d x + b d ) (2)

[0053] Among them, x represents the input power load data (real data or generated data); W d is the weight matrix of the discriminator; b d is the bias term of the discriminator; g(·) adopts the LeakyReLU activation function to ensure that there is also a non-zero gradient in the negative half-axis.

[0054] Furthermore, the training objective of WGAN is to minimize the Wasserstein distance so that the data distribution P G generated by the generator is as close as possible to the real data distribution P data , and its objective function is:

[0055]

[0056] Among them, is the expected score of the real data on the discriminator; is the expected score of the generated data on the discriminator

[0057] In order to satisfy the 1-Lipschitz condition of the discriminator, WGAN-GP (WGAN with gradient penalty) is adopted, and then a gradient penalty term is added to the objective function:

[0058]

[0059] Among them, λ is the gradient penalty coefficient; represents the interpolation sample randomly sampled between the real data and the generated data; is the L2 norm of the gradient of the discriminator with respect to .

[0060] Furthermore, in S2, the implementation of the CNN feature extraction module is as follows:

[0061] Convolution operation, and its calculation formula is:

[0062]

[0063] Among them, x t-i,d represents the input data at time step t - i and channel d; w i,d,c is the weight corresponding to time step i, channel d, and output channel c in the convolutional kernel; b c is the bias of output channel c; k is the time step window size of the convolutional kernel; D is the number of channels of the input data; y t,c is the eigenvalue of the output at time step t and output channel c.

[0064] Furthermore, in S3, the implementation of the BiTransformer time series modeling module is as follows:

[0065]

[0066] Among them and respectively represent the hidden states of the forward and backward Transformer units at time step t.

[0067] The Transformer unit mainly captures global time series dependencies from local features using the multi - head attention mechanism. After the linear transformation and encoding pre - processing steps, its formula is as follows:

[0068] Q = HW q , K = HWk, V = HW v (5)

[0069]

[0070] H′ = FFN(Attentionf)+H (7)

[0071] Among them, H is the input sequence, W q , W k , W v are the query, key, and value matrices respectively; d k is the dimension of the key; the softmax function is used to calculate the attention weights; H′ is the output feature, for the forward Transformer, it is For the backward Transformer, it is

[0072] Furthermore, in S4, BiLSTM is used to further learn the deep features of the time series, and its formula is:

[0073] i t = σ(W i x t +U i h t-1 +bi ) (8)

[0074] f t = σ(W f x t + U f h t-1 + b f ) (9)

[0075] o t = σ(W o x t + U o h t-1 + b o ) (10)

[0076]

[0077] h t = o t ⊙ tanh(c t ) (13)

[0078] Among them, x t represents the output feature of the Transformer at time step t; h t-1 is the hidden state of the previous time step; h t is the hidden state of the current time step; W * and U * are the weight matrices of the input and recurrent connections respectively, b * is the bias term; σ(·) is the sigmoid activation function, tanh(·) is the hyperbolic tangent activation function; ⊙ represents element-wise multiplication.

[0079] Finally, the results of the forward and backward temporal LSTMs are fused through an addition layer to obtain the final BiLSTM output:

[0080]

[0081] is the hidden state of the forward LSTM at the current time step; is the hidden state of the backward LSTM at the current time step.

[0082] Furthermore, in S5, the fully connected regression layer adopts a multi-layer neural network structure and is trained in combination with the mean squared error loss function, and its calculation formula is:

[0083]

[0084] Among them, W fc is the weight matrix of the fully connected layer; b fc is the bias of the fully connected layer; f FC(·) is a linear transformation.

[0085] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects:

[0086] The method of the present invention combines WGAN for data augmentation to improve data quality and the generalization ability of the model; uses a convolutional neural network (CNN) to extract local and global features to more effectively capture the time patterns of electric loads; and utilizes the global self-attention mechanism of a bidirectional Transformer to capture long-distance time dependencies. Finally, a bidirectional LSTM (BiLSTM) is used to more fully mine the long-term dependencies of time series data and improve the prediction accuracy. Compared with traditional methods, the present invention can not only enhance the model's learning ability for complex load patterns, but also effectively reduce the impact of data noise on the prediction results, providing more reliable technical support for the intelligent scheduling and optimal operation of power systems. Description of the Drawings

[0087] Figure 1 is a flowchart of the CNN-Bitransformer power load prediction method based on WGAN-enhanced data provided by the embodiments of the present invention;

[0088] Figure 2 is the prediction result of the method provided by the embodiments of the present invention. Detailed Embodiments

[0089] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0090] See Figure 1 As shown, a CNN-Bitransformer power load prediction method based on WGAN-enhanced data includes the following steps:

[0091] S1 Use a Wasserstein generative adversarial network (WGAN) to perform data augmentation on the original power load data to improve the distribution consistency and generalization ability of the data;

[0092] S2 Use a convolutional neural network (CNN) to extract the feature map of the power load time series after data augmentation, which can extract local patterns in the power load time series to capture short-term features;

[0093] S3 Use a bidirectional Transformer (BiTransformer) to model the forward and backward time dependencies to extract long-term time features from the feature map; to enhance the ability to capture long-term time features;

[0094] S4 uses a Bidirectional Long Short-Term Memory Network (BiLSTM) to further learn the deep features of the time series, so as to enhance the adaptability of the model to complex load patterns;

[0095] S5 generates the final load prediction result through the regression layer.

[0096] In one embodiment, in S1, the specific operation of implementing the WGAN data augmentation module can be divided into the design of the generator and the discriminator, where:

[0097] The generator G(z; θ G ) accepts the random vector z from the noise distribution P z and generates synthetic data The specific formula is:

[0098]

[0099] where z represents the input random noise vector, usually following the standard normal distribution W g is the weight matrix of the generator; b g is the bias term of the generator; f(·) is the ReLU function, which is used to introduce non-linear mapping.

[0100] The discriminator D(x; θ D ) is used to judge whether the input data is real data, and its mathematical expression is:

[0101] D(x) = g(W d x + b d ) (2)

[0102] where x represents the input power load data (real data or generated data), W d is the weight matrix of the discriminator; b d is the bias term of the discriminator, and g(·) uses the LeakyReLU activation function to ensure non-zero gradients on the negative half-axis.

[0103] In one embodiment, the training objective of the WGAN is to minimize the Wasserstein distance so that the data distribution P G generated by the generator is as close as possible to the real data distribution P data , and its objective function is:

[0104]

[0105] where, is the expected score of the real data on the discriminator; is the expected score of the generated data on the discriminator

[0106] To meet the 1-Lipschitz condition of the discriminator, WGAN-GP (WGAN with gradient penalty) is adopted, and a gradient penalty term is added to the objective function:

[0107]

[0108] where λ is the gradient penalty coefficient; represents an interpolation sample randomly sampled between the real data and the generated data; is the L2 norm of the gradient of the discriminator with respect to In one embodiment, in S2, the CNN feature extraction module is implemented as follows:

[0109] The convolution operation, and its calculation formula is:

[0110]

[0111]

[0112] where x t-i,d represents the input data at time step t - i and channel d; w i,d,c is the weight corresponding to time step i, channel d, and output channel c in the convolution kernel; b c is the bias of output channel c; k is the time step window size of the convolution kernel; D is the number of channels of the input data; y t,c t,c is the eigenvalue of the output at time step t and output channel c.

[0113] In one embodiment, in S3, the BiTransformer time series modeling module is implemented as follows:

[0114]

[0115] where and respectively represent the hidden states of the forward and backward Transformer units at time step t.

[0116] The Transformer unit mainly captures global time series dependencies from local features using the multi-head attention mechanism. After the linear transformation and the encoding preprocessing step, its formula is as follows:

[0117] Q = HW q K = HW k V = HW v (5)

[0118]

[0119] H′ = FFN(Attentionf) + H (7)

[0120] Among them, H is the input sequence, and W q , W k , W v are the query, key, and value matrices respectively; d k is the dimension of the key; the softmax function is used to calculate the attention weights; H' is the output feature. For the forward Transformer, it is For the backward Transformer, it is

[0121] In one embodiment, in S4, BiLSTM is used to further learn the deep features of the time series, and its formula is:

[0122] i t = σ(W i x t + U i h t-1 + b i ) (8)

[0123] f t = σ(W f x t + U f h t-1 + b f ) (9)

[0124] o t = σ(W o x t + U o h t-1 + b o ) (10)

[0125]

[0126] h t = o t ⊙ tanh(c t ) (13)

[0127] Among them, x t represents the Transformer output feature at time step t; h t-1 is the hidden state of the previous time step; h t is the hidden state of the current time step; W * and U * are the weight matrices of the input and recurrent connections respectively, b * is the bias term; σ(·) is the sigmoid activation function, tanh(·) is the hyperbolic tangent activation function; ⊙ represents element-wise multiplication.

[0128] Finally, the results of the forward and reverse temporal LSTMs are fused through an addition layer to obtain the final output of the BiLSTM:

[0129]

[0130] is the hidden state of the forward LSTM at the current time step; is the hidden state of the backward LSTM at the current time step.

[0131] Furthermore, in S5, the fully connected regression layer adopts a multi-layer neural network structure and is trained in combination with the mean squared error loss function. Its calculation formula is:

[0132]

[0133] where, W fc is the weight matrix of the fully connected layer; b fc is the bias of the fully connected layer; f FC (·) is a linear transformation.

[0134] In this embodiment, the annual electricity load data of Spain in 2018 is used as an example. The data has a step size of 15 minutes and 96 time steps per day. The interval prediction of the electricity load is realized through the following steps.

[0135] 1. Data preprocessing

[0136] The annual electricity load data of Spain in 2018 is processed by normalization to ensure the convergence during model training and reduce the errors caused by data scale differences.

[0137] 2. WGAN data augmentation

[0138] After the data preprocessing is completed, WGAN is used to augment the data to expand the training data and improve the generalization ability of the model.

[0139] 3. CNN feature extraction

[0140] After the data augmentation, CNN is used to extract the time series features. CNN captures the changing trends of daily electricity loads by learning the local patterns in the electricity load data.

[0141] 4. BiTransformer captures temporal dependencies

[0142] The extracted features are input into BiTransformer to capture the forward and backward temporal dependency relationships of the electricity load data. Through the processing of BiTransformer, the temporality of the electricity load can be better modeled, and the global perception ability of the data can be improved.

[0143] 5. Capturing Temporal Features with BiLSTM

[0144] Based on the output of BiTransformer, BiLSTM is used to capture the importance of different historical features and emphasize the feature maps that are important for power load forecasting, thereby improving the response speed and prediction accuracy.

[0145] 6. Result Verification

[0146] The electricity load data for the whole year of 2018 in Spain is used for verification, and the WGAN generates the simulated load data for 2017 in Spain. The results are as follows:

[0147] (1) Performance on the training set:

[0148] MAPE: 0.13763;

[0149] RMSE: 6.8952;

[0150] (2) Performance on the test set:

[0151] MAPE: 0.10665;

[0152] RMSE: 6.9367;

[0153] (3) Result Analysis

[0154] See Figure 2 The experimental results in show that the method of the present invention has achieved better prediction accuracy in the power load forecasting task, and is relatively low in terms of the RMSE and MAPE indicators, proving the effectiveness of the method in power load forecasting.

[0155] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A CNN-Bitransformer power load forecasting method based on enhanced data by WGAN, characterized in that, The method includes: Performing data augmentation on the power load time series using an adversarial network; Extracting the feature map of the power load time series after data augmentation; Extracting long-term time features from the feature map by using a bidirectional Transformer to model the forward and backward time dependencies; Using a bidirectional long short-term memory network to learn the deep features of the long-term time features; Returning to the regression layer to generate the final load prediction result.

2. The CNN-Bitransformer power load prediction method based on WGAN-enhanced data according to claim 1, wherein Performing data augmentation on the power load time series using an adversarial network, including: The generator G(z; θ G ) accepts a random vector z from the noise distribution P z and generates synthetic data The specific formula is: Among them, z represents the input random noise vector, which usually follows the standard normal distribution W is the weight matrix of the generator; b g is the bias term of the generator; f(·) is the ReLU function, which is used to introduce non-linear mapping; Use the discriminator D(x; θ D ) to determine whether the input data is real data, and its mathematical expression is: D(x) = g(W d x + b d ) Among them, x represents the input power load data; W d is the weight matrix of the discriminator; b d is the bias of the discriminator; g(·) uses the LeakyReLU activation function.

3. The CNN-Bitransformer power load prediction method based on WGAN-enhanced data according to claim 2, wherein By minimizing the Wasserstein distance, the data distribution P generated by the generator G is as close as possible to the true data distribution P data and the objective function is: Among them, is the scoring expectation of the real data on the discriminator; is the scoring expectation of the generated data on the discriminator; Adding a gradient penalty term to the objective function: Among them, λ is the gradient penalty coefficient; denotes the interpolation sample randomly sampled between the real data and the generated data; is the L2 norm of the gradient of the discriminator with respect to .

4. The CNN-Bitransformer power load prediction method based on WGAN-enhanced data according to claim 2, wherein Extracting the feature map of the power load time series after data augmentation, and the calculation formula is: where x t-i,d represents the input data at time step t-i and channel d; w i,d,c is the weight corresponding to time step i, channel d, and output channel c in the convolutional kernel; b c is the bias of output channel c; k is the time step window size of the convolutional kernel; D is the number of channels of the input data; y t,c is the eigenvalue of the output at time step t and output channel c.

5. The CNN-Bitransformer power load prediction method based on WGAN-enhanced data according to claim 1, wherein Extracting long-term time features from the feature map by using a bidirectional Transformer to model the forward and backward time dependencies, including: where and represent the hidden states of the forward and backward Transformer units at time step t, respectively.

6. The CNN-Bitransformer power load prediction method based on WGAN-enhanced data according to claim 1, wherein Using a bidirectional long short-term memory network to learn the deep features of the long-term time features, including: i t = σ(W i x t + U i h t-1 + b i ) f t = σ(W f x t + U f h t-1 + b f ) o t = σ(W o x t + U o h t-1 + b o ) h t = o t ⊙tanh(c t ) where x t represents the Transformer output feature at time step t; h t-1 is the hidden state of the previous time step; h t is the hidden state of the current time step; W * and U * are the weight matrices of the input and the recurrent connection respectively, b * is the bias term; σ(·) is the sigmoid activation function, tanh(·) is the hyperbolic tangent activation function; ⊙ represents element-wise multiplication.

7. The CNN-Bitransformer power load forecasting method based on enhanced data by WGAN according to claim 1, characterized in that, Returning to the regression layer to generate the final load prediction result, including: Adopting a multi-layer neural network structure and training in combination with a mean square error loss function, and the calculation formula is: Among them, W fc is the weight matrix of the fully connected layer; b fc is the bias of the fully connected layer; f FC (·) is a linear transformation, is the output.