Deep LSTM network short-term load prediction method based on stacked auto-encoder

By adopting a deep LSTM network method based on stacked autoencoder in short-term power load prediction, the problems of insufficient cutting-edge fitting capabilities and high dependence on input data in the prior art are solved, and a more accurate and robust short-term load prediction effect is achieved.

CN120222335APending Publication Date: 2025-06-27GREATER BAY AREA INST FOR INNOVATION HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510285370.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient cutting-edge fitting capabilities and high dependence on input data in short-term power load prediction, resulting in low fitting between the prediction results and the actual load data, especially the prediction performance within 1 hour and less.

Method used

The short-term load prediction method of deep LSTM network based on stacked autoencoder is adopted. Through systematic collection and preprocessing of historical power load data, meteorological data and calendar data, load, meteorological factors and time-related characteristics are extracted, and the fitting ability and robustness of the model are enhanced by stacked autoencoder and deep LSTM network.

Benefits of technology

Improves the fitting ability of the load curve tip, enhances the linking ability to old time step values, reduces dependence on input data types, and achieves more accurate and robust short-term power load predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120222335A_ABST
    Figure CN120222335A_ABST
Patent Text Reader

Abstract

The invention discloses a short-term load prediction method for a deep LSTM network based on a stacked auto-encoder. The short-term load prediction method comprises the following specific steps: step 1, systematically collecting historical data; 2, original data extraction, cleaning and standardized preparation and preprocessing operation are carried out; 3, extracting data set load, meteorological external factors and time related characteristics; 4, carrying out self-encoder model supervision training; 5, finely adjusting the weight of the hidden state; and step 6, outputting a power load predicted value and evaluating the precision of the power load predicted value. Compared with the prior art, the method has the advantages that the fitting capability of the method at the tip of a load curve is improved by using the stacked auto-encoder, and input features with relatively large influence are endowed with relatively high weights so as to achieve relatively good prediction precision; the depth of the LSTM network is generally increased, and the capability of linking an old time step value to a current time step is enhanced; and an auto-encoder is constructed, so that the dependence of the method on input data types is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of short-term power load forecasting, and in particular to a short-term load forecasting method based on a deep LSTM network of stacked autoencoders. Background Art

[0002] With the introduction of the carbon peak and carbon neutrality goals, the proportion of renewable energy in primary energy consumption continues to increase. Although this brings more flexible, more open and highly intelligent development opportunities to the power grid, it also poses new challenges to the stability and management of the power grid due to the inherent uncertainty of clean energy. Short-term power load forecasting uses statistics, machine learning and other methods to mine the key factors affecting the load from historical data such as meteorology and date. The results are of great significance to the safe operation and flexible scheduling of the power grid, which can reduce unnecessary power generation, reduce resource waste, and have huge economic benefits.

[0003] Electric load forecasting requires the establishment of a reliable load forecasting method, which inputs historical load data and other reference data into the model to obtain the load forecast value for a specific period in the future. However, due to the emerging active distribution network, and the subsequent opening of energy markets and various ancillary service markets, the operation and control of the power system have brought many changes, mainly facing the problems of load forecasting deviation caused by the uncertainty and randomness of users' own behavior, the great influence of multi-dimensional external factors on electricity consumption behavior, and poor noise reduction of historical data. The traditional Long Short-Term Memory Network (LSTM) network prediction method has limitations in dealing with these problems, especially at the tip of the load curve, the prediction results have a low fit with the actual load data, resulting in poor prediction performance within 1 hour or less. In addition, for long time series input data, the traditional LSTM network prediction method has a weak ability to link the old time step value to the current time step, and is highly dependent on the input data type. Therefore, although a wide variety of forecasting methods with their own advantages have been proposed for power load forecasting, the accuracy and robustness of these methods still need to be improved, and they are not adaptable enough to the new conditions and data brought about by emerging active distribution networks, which limits the accuracy and flexibility of power dispatching. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a short-term load forecasting method based on a deep LSTM network of a stacked autoencoder, which solves the problems of insufficient cutting-edge fitting ability and high dependence on input data in the prior art and improves the accuracy and robustness of short-term load forecasting.

[0005] To solve the above technical problems, the technical solution provided by the present invention is: a short-term load forecasting method based on a stacked autoencoder deep LSTM network, and its specific steps are as follows:

[0006] Step 1: Systematically collect historical power load data, meteorological data, and calendar data;

[0007] Step 2: Perform operations such as extraction, cleaning, standardization preparation, and preprocessing on the original data, and construct new load, meteorological, and calendar data sets;

[0008] Step 3: Extract load, meteorological external factors, and time-related features from the data set, and use the input attention mechanism to highlight high-impact load sequences;

[0009] Step 4: Conduct supervised and unsupervised greedy hierarchical pre-training on the deep LSTM autoencoder model based on the stacked autoencoder;

[0010] Step 5: Fine-tune the weights of the hidden states extracted from each layer of the deep LSTM network, and correctly learn the time dependence related to the ultra-long sequence input data through the time series stacked autoencoder;

[0011] Step 6: Output the power load prediction value and evaluate the accuracy.

[0012] Furthermore, in Steps 1 and 2, through visualization and statistical analysis, the integrity and accuracy of the data set are confirmed. When there are missing values in the data set, the missing value processing is first performed. When the number of missing values in the daily data is small, the data before and after the missing part are averaged to make up for the missing data;

[0013] When the number of missing values is large, the data at the same time period of the previous and the next day are selected, and the missing data are filled in proportion considering the change trend.

[0014] Furthermore, in Step 3, when extracting load, meteorological, and time features, the stacked autoencoder is used to add a weight bias β to the tip load to highlight the high-impact load in the data set.

[0015] Furthermore, the basic structural idea of the deep LSTM network is to connect two LSTM networks with opposite directions. The forward LSTM obtains the past data information of the input sequence, and the backward LSTM obtains the future data information of the input sequence. The hidden state of the deep LSTM network can be described as:

[0016]

[0017] Furthermore, the training process of the stacked autoencoder is executed sequentially. During the encoding, decoding, and learning phases, the reconstruction process of the input sequence is completed under the mean squared error cost function. The trained encoder layer is stored and placed as the input for the next autoencoder. This process is repeated for all layers. The specific number of stacks of the proposed stacked autoencoder is determined through trial and error.

[0018] Furthermore, a supervised and unsupervised greedy hierarchical pre-training structure based on the RNN method is adopted to solve the problem of randomly setting the initial value of the hidden state of the deep LSTM network.

[0019] Furthermore, all frozen pre-trained LSTM network layers are transferred to the fine-tuning stage, allowing these layers to be adaptively adjusted according to the specific task, intelligently screening and focusing on key time series, and finally training the overall model.

[0020] After adopting the above structure, the present invention has the following advantages: The present invention uses the stacked autoencoder to improve the fitting ability at the tip of the load curve, assigns higher weights to the input features with greater influence, so as to achieve better prediction accuracy; generally increases the depth of the LSTM network, extracts more input features, and enhances the ability to link the old time step values to the current time step; constructs an autoencoder, and hybridly applies supervised and unsupervised learning methods to reduce the dependence of the method on the input data type, so as to achieve more accurate and robust short-term load prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flowchart of the short-term load prediction method for a deep LSTM network based on a stacked autoencoder;

[0022] Figure 2 is the content structure of the LSTM network;

[0023] Figure 3 is a schematic diagram of the structure of the deep LSTM network;

[0024] Figure 4 is a schematic diagram of the autoencoder architecture of the deep LSTM network;

[0025] Figure 5 is the load prediction curve of the prediction method for the LD_BS dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] For the convenience of those skilled in the art, the present invention will be further described below in conjunction with the embodiments and the drawings. The content mentioned in the embodiments is not a limitation to the present invention. It should be clear that the historical power load data targeted in the embodiments of the present invention is only one of the situations that can be accurately fitted by this load prediction method, and this method is generally applicable to various power load prediction scenarios.

[0027] As Figures 1-5 , the present invention provides a short-term load forecasting method based on a deep LSTM network of stacked autoencoders. The specific implementation steps are as shown in the appendix Figure 1 and are specifically elaborated as follows:

[0028] Step 1: Systematically collect historical power load data, meteorological data, and calendar data. Since power data involves the national economic operation status and is not easy to obtain, generally, publicly available datasets on the Internet are collected for experiments. This is also a conventional method adopted in similar research. In the example of this application, the original dataset comes from a certain city in a certain province of China from 2022 to 2023, with a total of 510 days of load curves, meteorological data, and calendar data, denoted as LD_BS. The measurement period of the load is 1 hour, so the dimension of the load curve is 24. The meteorological data includes the maximum and minimum wind speeds, the highest and lowest temperatures, etc. The calendar data uses the One-Hot technology to encode weekdays, weekends, and holidays.

[0029] Step 2: Perform operations such as extraction, cleaning, standardization preparation, and preprocessing on the original data, and construct new load, meteorological, and calendar datasets. Divide the integrated dataset into a training set, a validation set, and a test set in chronological order to ensure the consistency of load, meteorological, and calendar information for each day and each hour, and to ensure the generalization ability of the forecasting method on different time windows. Through visualization and statistical analysis, confirm the integrity and accuracy of the dataset. When there are missing values in the dataset, missing value processing can be performed first. When the number of missing values in daily data is small, the data before and after the missing part should be averaged to make up for the missing data; when the number of missing values is large, the data at the same time period of the previous / next day should be selected and filled with missing data at an equal ratio considering the change trend. Apply the min-max standardization technique to transform the data to the same scale and reduce the influence of dimensions. The specific calculation method is as follows:

[0030]

[0031] where Xn, X, Xmax, and Xmin are the processed data, the original data, and the maximum and minimum values of the samples, respectively.

[0032] Step 3: Extract load, meteorological external factors, and time-related features from the dataset, and use the input attention mechanism to highlight high-impact load sequences. Extract high-dimensional features from the input historical power load data by building a convolutional neural network architecture. The extracted feature vectors are transformed into a time series form and input into the gated recurrent unit, and the input attention mechanism is used to highlight high-impact load sequences. The historical load X is used as the input, and the matrix A is the output of the network. The description of the convolutional neural network is as follows:

[0033]

[0034] Among them, the outputs of convolutional layer 1, convolutional layer 2, pooling layer 1, and pooling layer 2 are C1, C2 respectively, P1, P2, W1, W2 are the weights of each layer, B1, B2, B3, B4 are the biases, β1, β2 are the biases highlighted by the input attention mechanism, and Sigmoid is the activation function. The input attention mechanism strengthens important information by assigning different weights to the input features of the method to avoid the problem of information loss due to too long sequences, making it easier for the method to capture long-range relationships in the sequence. The description of the weight bias β1 is as follows:

[0035] β1 = utanh(WA1 + B) (3)

[0036] Among them, u is the weight and B is the bias.

[0037] Step 4: Conduct supervised and unsupervised greedy hierarchical pre-training on the deep LSTM autoencoder model based on the stacked autoencoder. Each hidden layer neuron of the LSTM network is replaced by a unit called a memory block, and there is a constant error carousel (CEC) at the center of each memory block. As shown in the appendix Figure 2 , each memory block has one or more memory units and three input, output, and forget gates. The calculation process of the LSTM network is described as follows:

[0038]

[0039] Among them, At is the input at time t, ht is the hidden state at time t, W is the weight matrix, B is the bias of the LSTM, δ(x) is the activation function, and the subscripts i, f, o represent input, forget, and output respectively. In time processing, traditional LSTM usually ignores future information and only processes data in one direction. While the deep LSTM network can enhance the ability to link old time step values to the current time step, and its structural schematic diagram is shown in the appendix Figure 3 . The basic structural idea of the deep LSTM network is to connect two LSTM networks with opposite directions. The forward LSTM obtains the past data information of the input sequence, and the backward LSTM obtains the future data information of the input sequence. Therefore, the hidden state of the deep LSTM network can be described as:

[0040]

[0041] The schematic diagram of the deep LSTM network autoencoder architecture is shown in the appendix Figure 4, during the encoding, decoding, and learning phases, the reconstruction process of the input sequence is completed under the mean squared error cost function. The stacking process of the LSTM model is inspired by the hierarchical supervised and unsupervised greedy hierarchical pre-training RNN method, and a pre-training structure is adopted to solve the problem of randomly setting the initial values of the deep LSTM hidden states. The hyperparameter settings of the proposed deep LSTM network short-term load forecasting method based on stacked autoencoders are shown in Table I:

[0042] Table I Hyperparameter Settings of the Forecasting Method

[0043]

[0044] Step 5: Fine-tune the weights of the hidden states extracted from each layer of the deep LSTM network, and correctly learn the time dependence related to the ultra-long sequence input data through the time series stacked autoencoder. Transfer all the frozen pre-trained LSTM network layers to the fine-tuning stage, and finally train the overall model while addressing the challenge of randomly assigning initial network weights that becomes increasingly prominent as the data sequence length increases. The trained encoder layer is placed as the input layer of the second autoencoder. The training process of the autoencoder is executed sequentially, and the trained encoder layer is stored and placed as the input of the next autoencoder. This process will be repeated for all layers. Note that the specific depth of the proposed stacked autoencoder is determined through trial and error.

[0045] Step 6: Output of the power load prediction value and accuracy evaluation. In the power load prediction evaluation, the root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE) are selected to evaluate the accuracy of the prediction results.

[0046] RMSE is used to measure the deviation between the predicted value and the true value, and is defined as:

[0047]

[0048] MAE is used to measure the absolute error between the predicted value and the true value, and is defined as:

[0049]

[0050] MAPE is used to measure the relative error between the predicted value and the true value, and is defined as:

[0051]

[0052] The comparison of the prediction accuracies of the short-term load prediction method of the deep LSTM network with stacked autoencoders for different months of the LD_BS dataset is shown in Table II. The average MAPE is only 2.85%, indicating that the prediction method has little difference from the actual value and the prediction method has high accuracy. In addition, the MAPEs of multiple different months are all below 3.5%, indicating that the prediction method has strong adaptability to different months and high robustness of the method.

[0053] Table II Comparison of Prediction Accuracies for Different Months

[0054]

[0055] The comparison between the prediction curve of the short-term load prediction method of the deep LSTM network with stacked autoencoders for the LD_BS dataset on July 11 and the actual load curve is as attached Figure 5 shown. It can be seen that the fitting effect between the prediction curve and the historical load curve is good, and the trends of the curves are basically the same.

[0056] The above embodiments are preferred implementation schemes of the present invention. In addition, the present invention can also be implemented in other ways. Any obvious replacement without departing from the concept of the technical solution of the present invention is within the protection scope of the present invention.

[0057] After adopting the above method, compared with the traditional short-term load prediction method, the prediction method proposed by the present invention, based on the increase in the depth of the stacked autoencoder and the LSTM network, improves the fitting ability of the method at the tip of the load curve and the correlation strength of the forward and backward time series. By using the constructed stacked autoencoder and mixing the supervised and unsupervised learning pre-training methods, the weights of the hidden states of each layer of the pre-trained LSTM network are fine-tuned, achieving a more accurate and robust short-term load prediction effect.

[0058] In order to make it more convenient for those of ordinary skill in the art to understand the improvements of the present invention over the prior art, some of the drawings and descriptions of the present invention have been simplified, and for the sake of clarity, some other elements have also been omitted from this application document. Those of ordinary skill in the art should be aware that these omitted elements may also constitute the content of the present invention.

Claims

1. A short-term load forecasting method based on a deep LSTM network of stacked autoencoders, characterized in that: Its specific steps are as follows: Step 1: Systematic collection of historical power load data, meteorological data and calendar data; Step 2: Extraction, cleaning and standardization of raw data, preparation and preprocessing operations, and construction of new load, meteorological and calendar datasets; Step 3: Extract the load, meteorological external factors and time-related features of the dataset, and use the input attention mechanism to highlight the high-impact load sequences; Step 4: Conduct supervised and unsupervised greedy layered pre-training of the deep LSTM autoencoder model based on stacked autoencoders; Step 5: Fine-tune the weights of the hidden states extracted from each layer of the deep LSTM network to correctly learn the temporal dependencies associated with the very long sequence input data through the time series stacked autoencoder; Step 6: Output and accuracy evaluation of power load forecast value.

2. According to claim 1, a deep LSTM network short-term load forecasting method based on stacked autoencoders is characterized in that: Steps 1 and 2 above confirm the completeness and accuracy of the data set through visualization and statistical analysis. When the data set is missing, the missing values ​​are first processed. When the missing values ​​of daily data are small, the data before and after the missing part are averaged to make up for the missing data. When there are many missing values, select the data of the same period of the previous and next day, and fill the missing data in equal proportions while considering the changing trend.

3. The method for short-term load forecasting based on a deep LSTM network of stacked autoencoders according to claim 1, characterized in that: The step 3 adds a weight bias β to the tip load by using a stacked autoencoder when extracting load, meteorological and time features, so as to highlight the high impact load in the data set.

4. The method for short-term load forecasting based on a deep LSTM network of stacked autoencoders according to claim 1, characterized in that: The basic structural idea of ​​the deep LSTM network is to connect two LSTM networks in opposite directions. The forward LSTM obtains the past data information of the input sequence, and the backward LSTM obtains the future data information of the input sequence. The hidden state of the deep LSTM network can be described as:

5. The method for short-term load forecasting based on a deep LSTM network of stacked autoencoders according to claim 1, characterized in that: The training process of the stacked autoencoder is performed sequentially. During the encoding, decoding, and learning stages, the reconstruction process of the input sequence is completed under the mean squared error cost function. The trained encoder layer is stored and placed as the next autoencoder input. This process will be repeated for all layers. The specific number of stacking times of the proposed stacked autoencoder is determined by trial and error.

6. The method for short-term load forecasting based on a deep LSTM network of stacked autoencoders according to claim 1, characterized in that: The supervised and unsupervised greedy hierarchical pre-training structure based on the RNN method is used to solve the problem of random setting of the initial value of the hidden state of the deep LSTM network.

7. The method for short-term load forecasting based on a deep LSTM network of stacked autoencoders according to claim 1, characterized in that: Transferring all frozen pre-trained LSTM network layers to the fine-tuning stage allows these layers to be adaptively adjusted according to specific tasks, intelligently screening and focusing on key time series, and finally training the overall model.