Long-Term Electric Load Forecasting Method Based on Hierarchical Residual Self-Attention Neural Network
Through the method based on the hierarchical residual self-attention neural network, the power load data characteristics are extracted and encoded and decoded, the error accumulation and feature mining problems in long-term power load prediction are solved, and the long-term power load prediction with higher accuracy is achieved, and the stable operation of the power system is supported.
Patent Information
- Application Number
- CN202210048738.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-01-17
AI Technical Summary
The existing power load prediction models have problems with error accumulation and insufficient feature mining capabilities in long-term prediction, especially in high-dimensional data scenarios, which are difficult to effectively capture long-distance historical features. The existing tool models are only suitable for short-term predictions, making it difficult to meet the stability and real-time requirements of the power system.
The method based on hierarchical residual self-attention neural network is adopted to adaptively extract the trend terms, period terms, holiday terms and weather terms characteristics in the power load data, and the hierarchical residual self-attention network blocks are used for encoding and generative decoding, reconstructing future power load fluctuations, and improving the accuracy of the model's long-sequence prediction.
The feature mining capability and model generalization capability of power load prediction are improved, medium- and long-term prediction errors are reduced, better feedback guidance is provided, and effective support is provided for the stable operation and allocation of power systems.
Smart Images

Figure CN114529051B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of load forecasting in power energy systems, and particularly to a long-term power load forecasting method based on a hierarchical residual self-attention neural network. Background Art
[0002] Power load forecasting technology is an indispensable part of the intelligent power grid system and is actively applied to many scenarios. How to effectively control the power load to achieve supply-demand balance has become an important research direction in the operation and management of modern power systems. The power system aims to provide economic, reliable, and high-quality electric energy for various users and should always meet the load requirements of customers. At the same time, in the future, new energy will be the goal of the new power industry reform, which makes it a major challenge for future load forecasting to effectively ensure the stable operation of the power grid and reasonably formulate power generation plans. The core issue of load forecasting is how to obtain the historical change law of the forecasting object and its relationship with certain influencing factors. The forecasting model is actually a mathematical function expressing this change law, and the challenge of load forecasting lies in that it will be affected by many external factors, including power trading market factors, national policy factors, weather factors, residential electricity consumption habits factors, etc., which are the problems to be solved.
[0003] The models for load forecasting can essentially be classified as mathematical models for time series forecasting. The commonly used methods can be divided into: traditional statistical methods, machine learning-based methods, deep learning-based methods, and third-party tool prediction methods. (1) Traditional statistical methods, commonly including time series models such as Auto Regression (AR) and Auto Regression Moving Average (ARMA). These models have simple principles and are suitable for analyzing stationary sequences and simple non-stationary sequences at a small order of magnitude. However, they are not suitable for solving non-linear prediction scenarios. (2) Machine learning-based methods. Machine learning is a broad category, and there are many models suitable for solving non-linear predictions, commonly including Support Vector Machines (SVM), decision tree models, K-Nearest Neighbor models, etc., and even ensemble learning models with better prediction capabilities such as XGBoost and LightGBM. Machine learning models have well solved the non-linear problem, but in the prediction scenario of large-scale and high-dimensional data, they are restricted by their feature mining ability. Building a machine learning prediction model often requires manual processing of data features. (3) Deep learning-based methods. Deep learning models, due to their powerful fitting ability, can adaptively mine and learn data features and are very suitable for solving non-linear prediction problems. Commonly used methods include Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Gate Recurrent Unit (GRU), etc. Among them, recurrent neural networks represented by LSTM and GRU are widely used in sequence modeling and have good sequence capabilities. However, due to serial learning, recurrent neural networks will gradually lose the ability to learn long-distance historical features during training and there is an error accumulation phenomenon. Therefore, they are often used in combination with other deep learning models. (4) Third-party tool prediction method. In recent years, some large domestic and foreign companies have also open-sourced their self-developed time series prediction methods. For example, Facebook launched the Prophet model in 2017. This model comprehensively considers the trend term, periodic term, and holiday term of the time series. The model is easy to use and has stable prediction ability. Subsequently, Amazon launched the DeepAR model in 2018. This model uses a probability-based autoregressive inference method to reduce the uncertainty during prediction. The prediction accuracy of these tools is significant, but they can only perform short-term predictions and are not suitable for energy load scenarios with high real-time requirements and strong stability. Summary of the Invention
[0004] The objective of the present invention is to combine existing technologies and improve upon them, optimizing the modeling effect of the load forecasting model in the power load forecasting scenario. Specifically, the present invention uses a neural network approach for modeling, proposing a network structure based on a hierarchical residual self-attention mechanism for long-sequence forecasting of stable and highly periodic power load data.
[0005] To achieve the above-mentioned objective of the invention, the technical solution adopted by the present invention is as follows:
[0006] A long-term power load forecasting method based on a hierarchical residual self-attention neural network, comprising the following steps:
[0007] Step 1. Obtain the source data of the unit load sequence and the weather data monitored by sensors from the time series database;
[0008] Step 2. Clean the source data, extract features from the cleaned historical load data and weather data, respectively extracting four major features: the trend term, the periodic term, the holiday term, and the weather term of the load fluctuation. Perform data fusion on the historical load sequence data and the feature data to obtain a fusion vector for input to the next neural network modeling.
[0009] Step 3. Use the hierarchical residual self-attention neural network proposed by the present invention to encode the input sequence, extract and mine the important features therein, and perform model training;
[0010] Step 4. Perform generative encoding on the features extracted from the source historical load data to be predicted, and predict the load sequence within the next time step range.
[0011] The beneficial effects of the present invention: The model proposed by the present invention is based on the Transformer neural network, which uses a self-attention mechanism internally. Compared with traditional recurrent neural networks, it has a stronger ability to capture global features. On this basis, the present invention method improves the Transformer and uses it as a sub-network unit for feature extraction. At the same time, more residual connections are added, and the external unit is designed as a hierarchical structure. Each layer decomposes the time series, and the decomposed time series passes through the residual self-attention unit in turn. Finally, the time series is reconstructed through the convolutional network at the bottom layer and decoded generatively to predict a whole segment of future load data. This method is more excellent and flexible than traditional methods in terms of feature mining ability and model generalization ability. Implementing load forecasting through this method can well reduce the prediction error of medium- and long-term load forecasting, provide feedback guidance for the operation and allocation of power units, and ensure the stable operation of the power system. Description of the Drawings
[0012] Figure 1It is a flow chart of a long-term power load forecasting method based on a hierarchical residual self-attention neural network in an embodiment of the present invention;
[0013] Figure 2 It is a schematic diagram of the overall framework of a hierarchical residual self-attention neural network prediction model in an embodiment of the present invention;
[0014] Figure 3 It is a schematic diagram of the framework of the Transformer neural network model;
[0015] Figure 4 It is a schematic diagram of the framework of the residual neural network model;
[0016] Figure 5 is a schematic diagram of the framework of each layer of improved residual self-attention blocks in an embodiment of the present invention;
[0017] Figure 6 It is a schematic diagram of a framework for prediction using generative decoding in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The present invention will be further described below in conjunction with the accompanying drawings. Figure 1 As shown:
[0019] Step 1. Determine the start time T start and end time T end , use middleware services or data analysis software to read the load data within the specified time range from the database storing the unit load sequence X raw Similarly, read the weather data X collected by the sensor weather , skip to step 2.
[0020] Step 2: Extract sequence feature data from historical load data and weather data and perform data fusion with source historical load data, including the following sub-steps:
[0021] Step 2-1. Extraction of weather data features.
[0022] Encode the weather data collected by the sensor. The collected data includes at least temperature data, weather status data, timestamp data, etc. Analyze the data, remove the abnormal data with large deviation, and encode the temperature data X. weather (T) is normalized to the maximum and minimum values, where the normalization function is expressed as:
[0023]
[0024] Similarly, for other numerical weather-related data, the above method can be used for feature normalization, which effectively helps the subsequent feature fusion.weather (S) is often a categorical label, such as [sunny, cloudy, light rain, heavy rain, light snow, …]. For such data, the one-hot encoding method is used to convert it into numerical data. Specifically, each label will be encoded into a unique numerical value. The representation of one-hot encoding is as follows:
[0025] Status Sunny Cloudy Light rain Heavy rain Light snow …… Coding value 0 1 2 3 4 ……
[0026] Through the above method, the feature processing of weather data can be achieved.
[0027] Step 2-2. Extraction of the trend term feature and the periodic term feature of the historical load sequence.
[0028] The historical load sequence feature is the main factor affecting the future sequence trend. For the non-linear and time-varying features of the time series, it is decomposed through a shallow neural network. Before decomposition, data cleaning is required for the data, and the data is analyzed and the values with too large offsets are removed. Specifically, a mixed sequence decomposition layer neural network is defined, and the original input is X input According to the following process, the trend term feature and the periodic term feature can be generated:
[0029] X trend = MovingAvg(X input )
[0030] X period = X input - X trend
[0031] where MovingAvg is the moving average function, which is obtained by using an average pooling operation of a one-dimensional convolution. Through this operation, the overall trend term of the sequence fluctuation can be obtained, and then the periodic term can be obtained by subtracting the trend term from the original sequence
[0032] Step 2-3. Extraction of the holiday term feature of the historical load sequence.
[0033] In load forecasting, the existence of important holidays also affects the load trend to a certain extent. Specifically, for the timestamp X of the original load data extracted in Step 1 timestamp , the pandas and numpy libraries in the Python language are used to analyze the data, and the extended features of the date where each timestamp is located are calculated, including the month X month , the day number X day , the hour X hour , the minute X minute , the week X weekday , whether it is a working day X iswork , whether it is a holiday Xisholiday , whether it is a weekend or not, X isweekend For more fine-grained features, the pandas DataFrame library is used to parse and analyze the time during this process, as shown below:
[0034] X month , X day , X hour , X minute ,... = Extend(X timestamp )
[0035] X timestamp = Linear(Extend(X timestamp ))
[0036] Among them, Extend is a feature extension function that converts the extended multi-dimensional features into a data form with the same dimension as the source sequence through a non-linear conversion layer for subsequent feature fusion.
[0037] Step 2-4. Feature embedding and fusion.
[0038] Through the first three steps, steps 2-1, 2-2, and 2-3, the existing feature dataset X can be obtained weather , X trend , X period , X trend , X timestamp , Next, these features are fused. Here, an additive model is used for fusion and input to the subsequent hierarchical residual neural network, which is expressed as:
[0039]
[0040] Among them, DropOut is a common neuron inactivation rate function in neural network modeling, aiming to prevent overfitting. RELU is a common activation function. Finally, through an additive model, the fused features can be obtained.
[0041] Step 3. Use the hierarchical residual self-attention neural network proposed in the present invention to encode the input sequence, extract and mine important features therein. The overall schematic diagram of this model is as shown in Figure 2 In this embodiment, step 3 specifically includes the following sub-steps:
[0042] Step 3-1. Sequence feature decomposition.
[0043] A major innovation of the present invention is to use a hierarchical decomposition sequence modeling process to replace the traditional linear modeling process. It recursively decomposes the feature sequence according to the number of layers, and then uses the residual self-attention network to model the decomposition features of each layer, and finally can train a better feature expression at a deeper level. Specifically, the decomposition algorithm provided by the present invention includes odd-even decomposition and binary decomposition, wherein the pseudo code of the algorithm is expressed as follows:
[0044]
[0045] in, is the mixed feature sequence of the source input, Level is the number of preset layers, SplitSeries is the sequence decomposition function, and the default algorithm proposed in this invention is to use binary decomposition to decompose the two feature components X left ,X right They are input into the residual block for updating respectively, and we get Then continue to use Algorithm 1 for recursive decomposition until the layer limit is reached, and finally use the Merge function to return the merged sequence.
[0046] Step 3-2. Use a hierarchical residual self-attention neural network to extract information about feature components.
[0047] The prototype of the hierarchical residual self-attention neural network proposed in this invention is the Transformer network, and its architecture is shown in the figure Figure 3 As shown, specifically, the present invention uses a self-attention mechanism, which has more potential to mine dependencies between time series than LSTM and GRU. The self-attention mechanism emphasizes focusing on the overall situation and better preventing information loss. The present invention has modified the original Transformer in terms of training time and prediction accuracy. Specifically, the feedforward neural network in the original Transformer encoder is replaced with a convolutional network with a smaller number of parameters. At the same time, for the hierarchical structure proposed in this design, more cross-layer residual connections are added to stabilize the gradient changes during model training. The basic structure of the residual network is as follows Figure 4 Finally, this design simplifies the Transformer decoder layer and replaces it with a combination of a fully connected layer and a Gaussian error function. The overall modified architecture is as follows Figure 5 shown.
[0048] At each layer, the feature component X input Input into the model to obtain the timing feature information X with timing dependence dep , expressed as:
[0049] X dep =ResidualAttentionBlock(Xinput )
[0050] Step 3-2 specifically includes the following steps:
[0051] Step 3-2-1: Input each single time feature component X after division of each layer input into the multi-head residual self-attention block to obtain the encoded feature X emded . The multi-head residual self-attention mechanism is expressed as:
[0052] ResidualMultiHead(H) = Concat(head1, head2,... head n )W o
[0053] where ResidualMultiHead represents the multi-head residual self-attention layer, H represents the number of attention heads, and W o represents the weight vector, that is, performs a non-linear transformation on the feature vector after fusion of multiple heads, so as to map it to a specified length. head1, head2,... head n represents the output of the self-attention layer of each head. Concat is a tensor concatenation function. The calculation of each head is as follows:
[0054]
[0055] where Q i , K i , V i are obtained through non-linear transformation after encoding the input data in each head. Prev i is the probability matrix calculated by the previous layer of the multi-head self-attention layer, and its result is passed to the next layer. Stable and excellent performance can still be obtained in a very deep network structure. By using multiple heads, the final fused feature X attn is obtained. The representations of these variables are as follows:
[0056]
[0057] Step 3-2-2: Input the output feature of the multi-head self-attention layer into the first layer of regularization layer NormalizationLayer1 to generate the feature vector X norm1 , and generate its copy X norm2 . Input X norm1 into the second layer of one-dimensional convolutional network to obtain the encoded vector X conv . Input X conv and X norm2Make a connection. Through the second normalization layer, i.e., Normalization Layer2, generate the encoded temporal feature component Z that is then passed to the next self-attention layer. Meanwhile, the probability matrix Prev calculated in step 3-2-1 is also passed to the next layer, and the relevant expression is as follows: i It is also passed to the next layer, and the relevant expression is as follows:
[0058] X norm1 = NormalizationLayer1(X attn )
[0059] X norm2 = X norm1
[0060] X conv = Dropout(Relu(Conv1d(X norm1 )))
[0061] Z = NormalizationLayer2(X conv + X norm2 )
[0062] Step 3-2-3. Repeat steps 3-2-1 and 3-2-2, and use the same operations in each stacked residual attention unit in the encoder part of the hierarchical residual block.
[0063] Step 3-2-4. Input the finally encoded vector Z of the encoder into the decoder for decoding. The decoder is improved from the traditional Transformer structure and has been appropriately simplified, and the expression is as follows:
[0064] Z = Gelu(Linear(Dropout(Z)))
[0065] Among them, Dropout is a hyperparameter representing the neuron inactivation rate in the neural network, which plays a role in preventing overfitting. Linear is a simple non-linear transformation function, and GELU is the Gaussian error linear unit that performs better in sequence modeling and has the best comprehensive performance in multiple scenarios. Its expression is as follows:
[0066]
[0067] The temporal component feature decoded by the decoder will have very good context expression ability. Pass the temporal component Z to the next residual self-attention block.
[0068] Step 3-2-5. Loop through steps 3-2-1, 3-2-2, 3-2-3, 3-2-4 until the sequence cannot be divided (reaching the layer requirement).
[0069] Step 3-2-6. Time series reconstruction. Through the steps between 3-2-1 and 3-2-5, the original time component features have been segmented into many time component features of the same length and restored in the order of the relative positions of the original features. The following are the segmentation and reconstruction algorithm processes using the odd-even segmentation strategy and the binary segmentation strategy respectively:
[0070]
[0071]
[0072] Compress the reconstructed sequence in the above way. Use the mean square error between the compressed sequence values and the true sequence values as the loss function to update the parameters of the neural network, thereby training the network. Set the compression length to embed_len, and use X to represent the compressed vector embed to express:
[0073] X embed = Embed(X T , embed_len)
[0074] Finally, use the mean square error (MSE) as the loss function to update the model parameters:
[0075]
[0076] where is the predicted value, represented by X embed in the training stage, and Y T is the true value, represented by X true in the training stage.
[0077] Step 4 Set the prediction step size and perform generative decoding to predict the load sequence in the next time range. Specifically, assume that the feature data of the reconstructed sequence to be predicted has been obtained at this step. Similarly, it is necessary to set the compression length embed_len. Here, the compression length should be less than the length of the reconstructed sequence sequence_len. Here, the length of the reconstructed sequence is default set to 96 and the compression length is 48. Through this operation, the reconstructed sequence will be compressed to a part at the end of the specified length, as Figure 6 shown in the following is the expression of the whole process:
[0078] X embed = Embed(X T , embed_len)
[0079] After obtaining the compressed sequence, the present invention performs long sequence prediction by proposing a generative decoding method. By setting the prediction length predict_len, a zero tensor X of the same dimension as the prediction length is initialized. zero , and X embed is horizontally concatenated with X zero , and then recompressed. The length of this compression is predict_len, generating the load prediction X pred for the historical sequence:
[0080] X pred = Embed(Concat(X embed , X zero ), predict_len)
[0081] The above is the preferred implementation process of the present invention. All changes made according to the technology of the present invention that do not exceed the scope of the technical solution of the present invention in terms of the functions and effects produced belong to the protection scope of the present invention.
Claims
1. A long-term electric load forecasting method based on a hierarchical residual self-attention neural network, characterized in that The method includes the following steps: Step 1. Obtain the source data of the unit load sequence and the weather data monitored by sensors from the time series database; Step 2. Perform data cleaning on the source data, extract features from the cleaned historical load data and weather data, and respectively extract four major features: the trend term, the periodic term, the holiday term, and the weather term of the load fluctuation; fuse the historical load sequence data and the weather feature data to obtain a fusion vector for the input of the next neural network modeling; specifically include: use a convolutional neural network to extract the overall trend term and periodic term features in the original load sequence; In view of the nonlinear and time-varying characteristics of time series, the feature decomposition is performed through a shallow neural network. Before decomposition, the data is cleaned, analyzed and the values with excessive offset are removed. A mixed sequence decomposition layer neural network is defined, and the original input is X input Generate the trend term X by following the following process: trend Characteristic and periodic terms X period feature: X trend = MovingAvg(X input ) X period = X input - X trend where MovingAvg is a moving average function obtained by using an average pooling operation of a one-dimensional convolution; Use one-hot encoding to extract features from the holiday term and the weather term. Finally, use the additive idea to horizontally concatenate the source load sequence and all the extracted feature data, and perform transformation through a fully connected layer to obtain the fused time series feature vector; Step 3. Use a hierarchical residual self-attention neural network to encode the input sequence, extract and mine the important features therein, and perform model training; specifically include: adopt the idea of recursion, hierarchically perform feature downsampling decomposition on the time series feature vector, use the residual self-attention network to mine the features of each decomposed time series component, on the basis of reaching the decomposition depth, recombine the mined features according to the original relative positions, and convert them into prediction results through a one-dimensional convolutional layer. Iterate in this way continuously, use the Adam algorithm as the optimization algorithm, and use the mean square error between the predicted value and the true value as the loss function to perform model training; Step 4. Perform generative encoding on the features extracted from the source historical load data to be predicted, and predict the load sequence within the next time step range; specifically include: perform feature transformation on the source load data to be predicted through Steps 2 and 3, concatenate the transformed features with a all-zero vector initialized to the prediction length, pass the concatenated vector through the model trained in Step 3, perform generative encoding, and predict the load sequence fluctuation for the entire future period.
2. The long-term electric load forecasting method based on a hierarchical residual self-attention neural network according to claim 1, wherein: The prototype of the hierarchical residual self-attention neural network is the Transformer network. The feed-forward neural network in the encoder of the original Transformer network is replaced with a convolutional network, and at the same time, more cross-layer residual connections are added to smooth the gradient change during model training. The decoder layer in the Transformer network is replaced with a combination of a fully connected layer and a Gaussian error function.
3. The long-term electric load forecasting method based on a hierarchical residual self-attention neural network according to claim 1, wherein: The concatenation described in Step 4 uses horizontal concatenation and is compressed.
4. The long-term electric load forecasting method based on the hierarchical residual self-attention neural network according to claim 3, wherein: The compression length is the prediction length.
Citation Information
Patent Citations
GRU-NN power load level prediction method based on EMD-SVR-MLR and attention mechanism
CN112766078A
Method and device for predicting and evaluating operation state of electric power information acquisition system
CN112884008A