A Holiday Load Forecasting Method and System Based on Deep Transfer Learning
By building a domain adversarial transfer learning network and Adapter method that improves Transformer, the problem of insufficient data in holiday load prediction is solved, and high-precision load prediction effect is achieved.
Patent Information
- Application Number
- CN202410791742.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-06-19
AI Technical Summary
The existing load prediction methods are difficult to directly apply to holiday load prediction. Traditional deep learning methods need to be retrained when the feature space changes, and lack clear basis for shared parameters, resulting in insufficient applicability of the model.
A domain-adversarial transfer learning network for improved Transformer is built, using conventional load samples as the source domain and holiday load samples as the target domain, and shared parameters are obtained through domain-adversarial transfer learning network training, and the pre-trained model is fine-tuned by the Adapter method to improve the accuracy of holiday load prediction.
By maximizing the similarity between holidays and non-holidays, the targetedness and accuracy of holiday load forecasting are improved, and the problem of insufficient data in holiday load forecasting is effectively solved.
Smart Images

Figure CN118676910B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a holiday load forecasting method and system based on deep transfer learning, belonging to the technical field of power load forecasting. Background Art
[0002] Accurate short-term load forecasting is the basic data support required for the daily unit commitment of the power system and is one of the keys to ensuring the balance between power grid supply and demand. With the improvement of science and technology and economic level, the load composition in China has become more and more complex, and its volatility and randomness have increased significantly.
[0003] Conventional load forecasting methods are usually based on a large number of historical samples, which can train the model sufficiently without considering the underfitting problem caused by insufficient samples. The holiday load has certain similarities with the conventional load fluctuation law. However, since the load is closely related to social activities, the holiday load and the conventional load fluctuation laws also show significant differences, and this difference is necessarily reflected in the historical samples. Therefore, the holiday load historical samples and the conventional load historical samples do not meet the data consistency requirements for modeling, and it is difficult to directly apply the conventional load forecasting model to holiday load forecasting.
[0004] Traditional deep learning methods are trained with a large number of specific samples to solve specific problems. Once the feature space changes, the model is no longer applicable and needs to be retrained. Transfer learning is a method to improve the model performance by applying the knowledge learned from the source domain to the target domain model. It focuses on the transfer and application of common knowledge in solving different problems. When the test sample distribution is different from the training sample, the transfer learning method does not need to train the model from scratch. It can retain the common knowledge obtained from the original model training and adjust the model with new data. Therefore, transfer learning has unique advantages in solving the problem of limited samples and has become a current research hotspot. However, currently, the shared parameters of the pre-trained model that need to be fixed in transfer learning are all obtained by experience without a clear basis. For this reason, the present invention is proposed. First, a domain adversarial transfer learning network model based on an improved Transformer is established to maximize the similarity between holidays and non-holidays, and a pre-trained model is jointly built to achieve the effect of reasonably optimizing the shared parameters. Then, the improved Transformer model trained by the domain transfer learning model is used as the pre-trained model, and the Adapter transfer learning method is adopted to fine-tune the parameters of the pre-trained model with holiday samples to improve the holiday load forecasting accuracy. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a holiday load forecasting method and system based on deep transfer learning. First, an improved Transformer domain adversarial transfer learning network is constructed. The Decoder structure of the Transformer model is discarded, and the Encoder part of the Transformer model is used as the classification predictor for domain transfer learning. The conventional load samples are used as the source domain, and the holiday load samples are used as the target domain to pre-train the model, so as to maximize the extraction of similar information and optimal shareable model parameters between the conventional load and the holiday load. Finally, the improved Transformer model obtained by training the domain adversarial learning network is used as the pre-trained model, and the Adapter method is used to fine-tune the pre-trained model parameters with the holiday load sample data as the target domain, so as to improve the pertinence and accuracy of holiday load forecasting.
[0006] Term Explanation:
[0007] Domain adversarial transfer learning network: Yaroslav Ganin et al. proposed the domain adversarial transfer learning network (Domain-Adversarial Training of Neural Networks, DANN) for characterizing domain adaptive learning and solving the problem of inconsistent data distributions between the source domain and the target domain. Its essence is a fusion-based transfer learning method based on instance samples and models. Its core goal is to construct a mapping between the source domain and the target domain, that is, the prediction model learned from the source domain can also be used for the target domain. In the traditional model-based transfer learning process, domain adaptability is usually represented by fixed features, that is, the training of the source domain and the target domain is divided into two stages. First, the model is pre-trained using the source domain to learn similar features, and then the model is fine-tuned using the target domain. DANN combines deep feature learning with domain adaptation and integrates them into the same training process, that is, embeds domain adaptation into the process of feature learning, so that the final decision result has both discriminative and similar features to the changes in the domain, and the features extracted from the source domain and the target domain have a similar distribution. Therefore, in the present invention, the domain adversarial transfer learning network is used to optimally extract the shareable parameters between the non-holiday model and the holiday model.
[0008] The architecture of DANN is as Figure 5 shown, including three key modules: a feature extractor, a class predictor, and a domain classifier. The feature extractor is used to map the data to a specific feature space to make it difficult for the domain classifier to distinguish the source of the data; the class predictor is used to classify the source domain data to assign the correct label as accurately as possible; the domain classifier is used to classify the data in the feature space to distinguish as accurately as possible whether the data comes from the source domain or the target domain.
[0009] Source domain and the target domain The samples together constitute the sample space When the sample x i After being input into the feature extractor, it is mapped into a feature vector F. Then the feature vector FS from the source domain is input into the class predictor to obtain the prediction result, and the feature vectors FS and FT from the source domain and the target domain are both input into the domain classifier to obtain the domain classification result. Then the loss function of the class prediction network is defined as
[0010]
[0011] where G f is the feature extractor network, and θ f are the parameters of the feature extractor network, G y is the class predictor network, and θ y are the parameters of the class predictor network
[0012] Define the loss function of the domain classifier network as
[0013]
[0014] where G d is the domain classification network, represents the source domain origin and is a vector of all zeros, is the target domain origin and is a vector of all ones.
[0015] GANN has to achieve two goals when training the model: one is to enable the domain classification network to correctly distinguish the feature origin, that is, to maximize the domain classification error; the other is to enable the prediction network to best match the source domain labels and minimize the prediction error. During the training process of the GANN model, the two goals compete with each other and finally achieve the Nash equilibrium. Then the total loss function of DANN can be defined as
[0016] L = L y + λL d
[0017] Therefore, the target task of DANN can be expressed as
[0018]
[0019] Transformer model: The Transformer model is a neural network with an attention mechanism proposed by Vaswanid et al. for processing time series data. It adopts a multi-layer encoder (Encoder) and decoder (Decoder) architecture. Each layer consists of a feed-forward neural network and multiple attention mechanism modules. The Encoder is used to encode the input sequence into a feature vector representation, and the Decoder is used to decode this vector representation into the target sequence. The structure of the Transformer model is as Figure 6as shown
[0020] The upper part represents a block structure of the Encoder of the Transformer model. Each block structure includes a multi-head attention mechanism and a fully connected feed-forward layer, and a normalization layer (Add&Norm) is added respectively. Usually, the Encoder is composed of multiple stacked block structures. The lower part represents a block structure of the Decoder of the Transformer model. Each block structure includes a masked multi-head attention mechanism, a multi-head attention mechanism and a fully connected feed-forward layer, and a normalization layer is added respectively. Similar to the Encoder, the Decoder is composed of multiple stacked block structures.
[0021] (1) Multi-head attention mechanism
[0022] The attention mechanism allows the training model to dynamically allocate different weights according to the correlation between different elements in the input data and the label, so as to mine strongly relevant information in the input data. It includes three key vectors: the query vector (Query), which represents the target to be focused on or retrieved; the key vector (Key), which represents the source to be matched or compared with the query vector; the value vector (Value), which represents the information to be weighted and summed according to the matching degree between the query vector and the key vector. Its structural principle is as Figure 7 as shown
[0023] Define the input sequence as X = {x1, x2, …, x t} where x t is the input feature corresponding to the t-th moment. The embedding network extracts features from the input vector to obtain the feature sequence F = {f1, f2, …, f t}. Define the query matrix, key matrix, and value matrix as the multiplication of the feature sequence, that is
[0024]
[0025] where Q = {q1, q2, …, q t} is the query sequence, K = {k1, k2, …, k t} is the key sequence, and V = {v1, v2, …, v t} is the value sequence. W Q , W K , W V are the query parameter matrix, key parameter matrix, and value parameter matrix respectively.
[0026] Then, calculate the weight of the corresponding value as
[0027]
[0028] Then The corresponding output is
[0029]
[0030] The multi - head attention mechanism is developed from the attention mechanism. The main difference is that in the multi - head attention mechanism, the query vector qi, the key vector k i , and the value vector v i are regarded as one "head". For multiple heads, for the input data x i it is necessary to construct multiple W Q , W K , W V query parameter matrices, key parameter matrices, and value parameter matrices, and multiply them with the feature sequence f i formed by them to obtain multiple groups of {q i , k i , v i}, as Figure 8 shown.
[0031] Transfer learning:
[0032] The three basic concepts of transfer learning are the source domain representing existing knowledge, the target domain that needs to be learned, and the task composed of the objective function and the learning result. Transfer learning is to transfer the knowledge and features learned in the source domain to the target domain, so that the model can better adapt to the data distribution characteristics in the target domain and improve the task effect.
[0033] In transfer learning, given a labeled source domain and a target domain with only a small number of samples The data distributions between the source domain and the target domain are similar but there are certain differences, that is, P(X s ) ≠ P(X t ). The goal of transfer learning is to learn the knowledge features in the target domain by means of the information of the source domain .
[0034] The technical solution of the present invention is as follows:
[0035] A holiday load forecasting method based on deep transfer learning, the steps are as follows:
[0036] (1) Divide the data set into a source domain (Source domain) and a target domain (Target domain);
[0037] (2) Construct the feature extractor and classification predictor of the domain adversarial transfer learning network, mine the similarity between the source domain and the target domain, optimize the model by minimizing the source domain prediction error, and obtain shared network parameters;
[0038] (3) Use the feature extractor and classification predictor trained by the domain adversarial transfer learning network as a pre-trained model (i.e., the improved Transformer model), fix the network parameters of the feature extractor, and use the Adapter Tuning transfer learning model to fine-tune the parameters of the classification predictor to improve the pertinence and accuracy of holiday load prediction;
[0039] (4) Use the classification predictor to perform load prediction, compare the difference between the prediction result and the real load data, evaluate the accuracy and performance of the model. At the same time, visualize and statistically analyze the prediction results to further illustrate the advantages and effects of the model.
[0040] Preferably according to the present invention, in step (1), the source domain is non-holiday data samples, and the target domain is holiday data samples.
[0041] Preferably according to the present invention, in step (1), data augmentation and data preprocessing are performed before dataset partitioning;
[0042] Data augmentation: Use TimeGAN to generate highly realistic synthetic holiday load datasets to expand the original holiday load sample set, thereby improving the balance of the data;
[0043] Data preprocessing: Perform preprocessing on the augmented holiday load data and the original data, including data cleaning, normalization, standardization, etc. operations to ensure the quality and consistency of the data.
[0044] Preferably according to the present invention, in step (2), the feature extractor of the domain adversarial transfer learning network is the embedding network on the Encoder side of the Transformer model, which extracts features from the input data, and the classification predictor uses the Encoder part of the Transformer model.
[0045] Preferably according to the present invention, in step (2), use the sample set composed of the source domain and the target domain to train the domain adversarial transfer learning network, indirectly achieving the pre-training of the pre-trained model and obtaining shareable network parameters.
[0046] Preferably according to the present invention, in step (3), the Adapter Tuning transfer learning model includes an input layer, a Transformer-Adapter layer, and a prediction layer;
[0047] Input layer: Preprocess the holiday-related samples and input them into the pre-trained improved Transformer encoder model;
[0048] Transformer-Adapter layer: For the non-linear and time-varying characteristics of the load, by embedding the Adapter module into the Encoder part of the Transformer model, fixing the Encoder attention mechanism part of the pre-trained Transformer model, training with holiday load samples, and fine-tuning the network parameters of the Adapter module and the fully connected layer to improve the prediction pertinence of the network;
[0049] Neil Houlsby et al. proposed in 2019 to introduce the Adapter fine-tuning technology into transfer learning. The Adapter freezes the main part of the pre-trained model, adds an Adapter module to the specific task layer, and fine-tunes the newly added module and some task layers. The Adapter technology is applied to the Encoder part of the Transformer model, as Figure 3 shown. The Adapter module receives the input from the previous layer, that is, the output and hidden representation of the Encoder part of the Transformer model. The input features are mapped to a lower-dimensional representation space through the downscaling layer to reduce the number of parameters; then the low-dimensional features are subjected to feature transformation through the non-linear layer to enhance the Adapter's non-linear feature mining ability; further through the upscaling layer, the low-dimensional features are remapped back to the original high-dimensional feature space, and the high-dimensional features are multiplied by a learnable parameter matrix to adjust the feature weights according to the feature importance; finally, after passing through the residual fully connected layer, the adjusted high-dimensional features are added to the original input to ensure that even if the parameters are close to zero, the Adapter module can still approximate the identity mapping and strengthen the information transmission within the module.
[0050] Prediction layer: The output of the fully connected layer is the predicted value. After obtaining the feature information through the Encoder part of the Transformer model, the load prediction is realized using the prediction layer.
[0051] A holiday load prediction system based on deep transfer learning, including:
[0052] Data input module, used to divide the data set into source domain and target domain;
[0053] Recognition module: Used to construct the feature extractor and classification predictor of the domain adversarial transfer learning network, mine the similarity between the source domain and the target domain, and obtain shared network parameters;
[0054] Use the feature extractor and classification predictor obtained by training the domain adversarial transfer learning network as a pre-trained model, fix the network parameters of the feature extractor, and use the Adapter Tuning transfer learning model to fine-tune the parameters of the classification predictor;
[0055] Prediction module: Use the classification predictor to perform load prediction.
[0056] The improved Transformer domain adversarial transfer learning network constructed in the present invention includes the following three cores:
[0057] (1) Design the domain adversarial transfer learning network framework
[0058] The key to transfer learning is to transfer the effective features in the source domain to the target domain. The transfer learning method based on the model usually adopts the fine-tuning method. This fine-tuning method means fixing some parameters learned by the pre-trained model in the source domain, and then fine-tuning the remaining parameters in the target domain. There is no clear basis for how to select the shared parameters. Usually, the feature extraction layer is fixed and the fully connected layer is fine-tuned. The present invention constructs a domain adversarial transfer learning network framework, optimizes the model by calculating the similarity between the source domain and the target domain and minimizing the source domain prediction error, and trains a feature extractor with shareable parameters.
[0059] (2) Establish the domain adversarial network of the improved Transformer model
[0060] The present invention optimizes and reconstructs the Transformer model, only retains the Encoder structure, abandons the Decoder structure to simplify the network complexity, and constructs a Transformer encoder in combination with the CNN-LSTM network to extract load features. The improved Transformer model, as the feature extractor and classification predictor that constitute the domain adversarial transfer learning, is the core module for realizing parameter sharing between the source domain and the target domain.
[0061] (3) Construct the model fine-tuning transfer learning model of Aadpter Tuning
[0062] The present invention uses a large number of source domain samples and limited target domain samples to train the domain adversarial network of the improved Transformer, tries to find the similarity between the source domain and the target domain as much as possible, and realizes the optimization of the domain adversarial network model. Then, use the improved Transformer model of the domain adversarial network as a pre-trained model, use the labeled holiday load data as input samples, fix the shared parameters of the feature extractor, and use the Adapter transfer learning method to fine-tune some parameters of the classification predictor to mine the specific laws of holiday loads and improve the pertinence and effectiveness of the model for holiday load prediction.
[0063] The beneficial effects of the present invention are as follows:
[0064] First, the present invention constructs a domain adversarial transfer learning network that improves the Transformer, discards the Decoder structure of the Transformer model, uses the Encoder part of the Transformer model as a classification predictor for domain transfer learning, and uses conventional load samples as the source domain and holiday load samples as the target domain to pre-train the model, maximizing the extraction of similar information and optimal shareable model parameters between conventional loads and holiday loads. Finally, the improved Transformer model obtained by training the domain adversarial learning network is used as a pre-trained model, and the Adapter method is used to fine-tune the pre-trained model parameters with holiday load sample data as the target domain, enhancing the pertinence and accuracy of holiday load prediction. Description of the Drawings
[0065] Figure 1 It is a framework diagram of the prediction model of the present invention;
[0066] Figure 2 It is a flow chart of the prediction model of the present invention;
[0067] Figure 3 It is a structural diagram of the Adapter Tuning transfer learning model of the present invention;
[0068] Figure 4 It is a structural diagram of the domain adversarial transfer learning network of the present invention;
[0069] Figure 5 It is a diagram of the existing DANN architecture;
[0070] Figure 6 It is a diagram of the existing Transformer model structure;
[0071] Figure 7 It is a schematic diagram of the existing multi-head attention mechanism;
[0072] Figure 8 It is a structural diagram of the existing multi-head attention mechanism;
[0073] Figure 9 It is a statistical chart of RMSE and MAE errors of the three prediction models in the present invention during the Spring Festival;
[0074] Figure 10 It is a prediction result diagram during the Spring Festival in the embodiment of the present invention;
[0075] Figure 11 It is a statistical chart of RMSE and MAE errors of the three prediction models in the present invention during the National Day;
[0076] Figure 12 Prediction result graph during the National Day holiday for the embodiment of the present invention;
[0077] Figure 13 Prediction result graph during the International Labor Day and Mid-Autumn Festival for the embodiment of the present invention;
[0078] Figure 14 RMSE and MAE error statistical graphs of three prediction models for the embodiment of the present invention during the International Labor Day;
[0079] Figure 15 RMSE and MAE error statistical graphs of three prediction models for the embodiment of the present invention during the Mid-Autumn Festival. Detailed implementation manners
[0080] The present invention will be further described below by way of embodiments in conjunction with the accompanying drawings, but not limited thereto.
[0081] Embodiment 1:
[0082] A holiday load prediction method based on deep transfer learning, the steps are as follows:
[0083] (1) Data augmentation: Use TimeGAN to generate a highly realistic synthetic holiday load data set to augment the original holiday load sample set, thereby improving the balance of the data;
[0084] Data preprocessing: Preprocess the augmented holiday load data and the original data, including operations such as data cleaning, normalization, and standardization, to ensure the quality and consistency of the data;
[0085] Divide the data set into a source domain and a target domain. The source domain is non-holiday data samples, and the target domain is holiday data samples;
[0086] (2) Construct the feature extractor and classification predictor of the domain adversarial transfer learning network, mine the similarity between the source domain and the target domain, optimize the model by minimizing the source domain prediction error, and obtain shared network parameters;
[0087] The feature extractor of the domain adversarial transfer learning network is the embedding network on the Encoder side of the Transformer model, which extracts features from the input data. The classification predictor uses the Encoder part of the Transformer model;
[0088] Use the sample set composed of the source domain and the target domain to train the domain adversarial transfer learning network, indirectly achieve the pre-training of the pre-trained model, and obtain shareable network parameters;
[0089] (3) Use the feature extractor and classification predictor obtained by training the domain adversarial transfer learning network as a pre-trained model (i.e., the improved Transformer model). Fix the network parameters of the feature extractor, and use the Adapter Tuning transfer learning model to fine-tune the parameters of the classification predictor to improve the pertinence and accuracy of holiday load forecasting;
[0090] The Adapter Tuning transfer learning model includes an input layer, a Transformer-Adapter layer, and a prediction layer;
[0091] Input layer: Preprocess the holiday-related samples and input them into the pre-trained improved Transformer encoder model;
[0092] Transformer-Adapter layer: For the non-linear and time-varying characteristics of the load, by embedding the Adapter module into the Encoder part of the Transformer model, fixing the Encoder attention mechanism part of the pre-trained Transformer model, training with holiday load samples, and fine-tuning the network parameters of the Adapter module and the fully connected layer to improve the prediction pertinence of the network;
[0093] The parameter fine-tuning process is as follows: The Adapter module receives the input from the previous layer, that is, the output and hidden representation of the Encoder part of the Transformer model. The input features are mapped to a lower-dimensional representation space through a downscaling layer to reduce the number of parameters; then the low-dimensional features are subjected to feature transformation through a non-linear layer to improve the Adapter's non-linear feature mining ability; further through an upscaling layer, the low-dimensional features are remapped back to the original high-dimensional feature space, and the high-dimensional features are multiplied by a learnable parameter matrix to adjust the feature weights according to the feature importance; finally, through a residual fully connected layer, the adjusted high-dimensional features are added to the original input to ensure that even if the parameters are close to zero, the Adapter module can still approximate the identity mapping and strengthen the information transmission within the module;
[0094] Prediction layer: The output of the fully connected layer is the predicted value. After obtaining the feature information through the Encoder part of the Transformer model, the load prediction is realized using the prediction layer.
[0095] (4) Use the classification predictor to perform load forecasting, compare the difference between the prediction result and the actual load data, evaluate the accuracy and performance of the model. At the same time, visualize and statistically analyze the prediction results to further illustrate the advantages and effects of the model.
[0096] Case study:
[0097] In this embodiment, the measured load data of a certain area in Shandong Province is used as test data for simulation experiments to verify the effectiveness of the proposed model. Using historical load data, meteorological data, and holiday feature data as inputs, the data from 2019 to 2020 is selected to train the model, and the holiday data in 2021 is used for test verification. The sampling interval of the load data is 15 minutes. The input data of the holiday load prediction model includes a multi-dimensional time series of the load values at the same time on the previous 5 days before the day to be predicted, meteorological features, week features, and holiday features, and the output data is the load to be predicted 24 hours in advance.
[0098] Verification of prediction effectiveness:
[0099] To verify the superiority of the method proposed in this embodiment, the Extreme Gradient Boosting (XGBT, replaced by model M1 in the following text), the improved Transformer model (without fine-tuning for direct prediction, M2), and the proposed method (M3) are respectively constructed for comparative verification to verify the superiority of the method proposed in this article. In this embodiment, the load prediction results of models M1, M2, and M3 during the four festivals of Spring Festival, Mid-Autumn Festival, May 1st International Labor Day, and National Day in 2021 are compared.
[0100] Table 1 respectively counts the prediction accuracies of the three models during the Spring Festival. From the comparison of the models, the prediction effects of models M1 and M2 have a large gap with the prediction errors of the M3 model proposed in this article. The RMSE indicators of model M3 are respectively reduced by 4.97% and 2.91% compared with model M2 and model M1, and the MAE indicators are on average reduced by 4.77% and 2.71%. It can be seen that the model in this embodiment can effectively improve the prediction accuracy of the model during the Spring Festival by migrating the similarity of non-holiday load characteristics.
[0101] To visually show the comparison of the prediction accuracies of the models, Figure 9 The prediction performances of the three prediction models during the Spring Festival are shown. Analyzing from the time, the RMSE and MAE indicators of the M3 model are relatively high on the 1st and 2nd days during the Spring Festival, indicating that the prediction effect is relatively poor. The RMSE and MAE indicators from the 3rd day to the 5th day are both less than 5.0%, and the prediction effect is better. The RMSE and MAE increase on the 6th and 7th days. It can be seen that the fluctuation law suddenly changes on the 1st and 2nd days of the holiday, resulting in a weaker temporal correlation between loads and a larger prediction error. As the holiday progresses, the temporal dependence of the samples during the holiday can be mined from the 3rd day to the 5th day, so the prediction effect is better. On the 7th day, the load rebounds, and the temporal law of the previous few days cannot well reflect its change, so the error increases slightly. Figure 10 Intuitively shows the prediction result curves of the three models during the Spring Festival.
[0102] Table 1: MAE and RMSE Values of Each Model during the Spring Festival
[0103]
[0104] Table 2 separately counts the prediction accuracies of the three models during the National Day. From the comparison of the models, the prediction effects of models M1 and M2 have a large gap with the prediction errors of the M3 model proposed in this embodiment. On the first day, the RMSE indicators of model M3 are 0.05% and 3.32% lower than those of models M1 and M2 respectively, and the MAE indicators are on average 0.27% and 3.7% lower. From the second day to the seventh day, the RMSE indicators of model M3 are on average 5.0% lower than those of models M2 and M1, and the MAE indicators are on average 4.7% lower. Generally speaking, the RMSE indicators of model M3 are on average 4.6% lower than those of models M2 and M1, and the MAE indicators are on average 4.2% lower. Thus, it can be seen that the model proposed in this embodiment can effectively improve the prediction accuracy of the model during the Spring Festival by migrating the similarity of the load characteristics of non-holiday periods. To visually show the comparison of the prediction accuracies of the models, Figure 11 shows the prediction performances of the three prediction models during the National Day.
[0105] Table 2: MAE and RMSE Values of Each Model during the National Day
[0106]
[0107] Analyzing from the time aspect, the RMSE and MAE indicators on the first day during the National Day are relatively high, indicating that the prediction effect is relatively poor. Analyzing from the perspective of the model M3 proposed in this article, the RMSE indicators from the second day to the fifth day are less than 5%, and the prediction effect is better. It can be seen that the sudden change in the fluctuation law on the first day of the holiday leads to a weakening of the temporal correlation between loads and a larger prediction error. As the holiday time goes by, the temporal dependence of the samples during the holiday can be mined from the second day to the fifth day, so the prediction effect is better. The load rises relatively fast at noon on the sixth and seventh days, and the prediction error slightly increases, which is slightly worse than the prediction results from the second day to the fifth day. Compared with the Spring Festival, the prediction accuracy is relatively better on the second day during the National Day. From the prediction curves of the Spring Festival and the National Day Figure 10 and Figure 12 it can be seen that the daily fluctuations of the load curves during the Spring Festival are quite different, while the daily fluctuations of the load curves during the National Day are relatively small. Therefore, the prediction accuracy during the National Day is more stable.
[0108] Table 3 and Table 4 respectively show the prediction situations of models M1, M2, and M3 during the International Workers' Day and the Mid-Autumn Festival, Figure 13It shows the predicted curve graphs during the May 1st International Labor Day and the Mid-Autumn Festival. Looking at the time, the load level during the Mid-Autumn Festival holiday decreased day by day, and on the last day, which was the Mid-Autumn Festival day, the load level dropped to the lowest; the first day of the May 1st International Labor Day was May 1st, the holiday day, so the overall load curve was relatively low. After that, the power load level directly rebounded and reached the highest on the last day of the holiday. From the model comparison, during the May 1st Labor Day, the RMSE index of model M3 was on average 2.49% lower than that of model M2 and model M1, and the MAE index was on average 2.18% lower; during the Mid-Autumn Festival, the RMSE index of model M3 was on average 1.27% lower than that of model M2 and model M1, and the MAE index was on average 1.67% lower. The comparison result of the load prediction error is consistent with that of the Spring Festival and the National Day, further verifying that the proposed model can effectively improve the prediction accuracy of the model during the Spring Festival by migrating the similarity of the load characteristics on non-holiday days. To visually show the comparison of the prediction accuracy of the models, Figure 14 and Figure 15 shows the prediction error performance of the three prediction models during the International Labor Day and the Mid-Autumn Festival.
[0109] Table 3: MAE and RMSE values of each model during the International Labor Day
[0110]
[0111] Table 4: MAE and RMSE values of each model during the Mid-Autumn Festival
[0112]
[0113] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A holiday load forecasting method based on deep transfer learning, characterized in that The steps are as follows: (1) Divide the dataset into a source domain and a target domain; (2) Construct a feature extractor and a classification predictor for the domain adversarial transfer learning network, mine the similarity between the source domain and the target domain, and obtain shared network parameters. The feature extractor of the domain adversarial transfer learning network is the embedding network on the Encoder side of the Transformer model, which extracts features from the input data, and the classification predictor uses the Encoder part of the Transformer model; (3) Use the feature extractor and classification predictor obtained by training the domain adversarial transfer learning network as a pre-trained model, fix the network parameters of the feature extractor, and fine-tune the parameters of the classification predictor using the Adapter Tuning transfer learning model. The Adapter Tuning transfer learning model includes an input layer, a Transformer-Adapter layer, and a prediction layer; Input layer: Preprocess the holiday-related samples and input them into the pre-trained improved Transformer encoder model; Transformer-Adapter layer: For the non-linear and time-varying characteristics of the load, by embedding the Adapter module into the Encoder part of the Transformer model, fixing the Encoder attention mechanism part of the pre-trained Transformer model, training with holiday load samples, and fine-tuning the network parameters of the Adapter module and the fully connected layer; Prediction layer: The output of the fully connected layer is the predicted value. After obtaining the feature information through the Encoder part of the Transformer model, use the prediction layer to achieve load prediction; (4) Use the classification predictor to perform load prediction.
2. The holiday load forecasting method based on deep transfer learning according to claim 1, characterized in that, In step (1), the source domain is the non-holiday data samples, and the target domain is the holiday data samples.
3. The holiday load forecasting method based on deep transfer learning according to claim 2, characterized in that, In step (1), data augmentation and data preprocessing are performed before dividing the dataset.
4. The holiday load forecasting method based on deep transfer learning according to claim 3, wherein Data augmentation: Use TimeGAN to generate a synthetic holiday load dataset to augment the original holiday load sample set.
5. The holiday load forecasting method based on deep transfer learning according to claim 4, wherein Data preprocessing: Preprocess the augmented holiday load data and the original data, including data cleaning, normalization, and standardization operations.
6. The holiday load forecasting method based on deep transfer learning according to claim 5, characterized in that In step (2), use the sample set composed of the source domain and the target domain to train the domain adversarial transfer learning network, indirectly pre-train the pre-trained model, and obtain shareable network parameters.
7. A holiday load forecasting system based on deep transfer learning, which is applied to the holiday load forecasting method based on deep transfer learning described in claim 1, and is characterized in that, Including: A data input module for dividing the dataset into a source domain and a target domain; An identification module for constructing a feature extractor and a classification predictor for the domain adversarial transfer learning network, mining the similarity between the source domain and the target domain, and obtaining shared network parameters; Use the feature extractor and classification predictor obtained by training the domain adversarial transfer learning network as a pre-trained model, fix the network parameters of the feature extractor, and fine-tune the parameters of the classification predictor using the Adapter Tuning transfer learning model; A prediction module for performing load prediction using the classification predictor.
Citation Information
Patent Citations
Pumping unit working condition diagnosis method based on transfer learning and ViT network
CN116467624A
Image processing method and apparatus, storage medium, and electronic device
WO2023098912A1