Electric Load Prediction Method and System Based on GLA-UNet Model Structure

Through the improved GLA-UNet model structure, combined with the global-local multi-scale perceived attention mechanism and multi-level memory bank, the accuracy of periodic and seasonal fluctuations in daily power load prediction is solved, and higher prediction accuracy and robustness are achieved.

CN119944671BActive Publication Date: 2025-07-11SHANGHAI UNIVERSITY OF ELECTRIC POWER
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510417343.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

Existing power load prediction models are difficult to accurately capture periodic and seasonal fluctuations on datasets with day as the minimum sampling unit, resulting in insufficient prediction accuracy and robustness, especially during periods of emergencies and seasonal changes.

Method used

The improved GLA-UNet model structure is adopted, combined with the global-local multi-scale perceived attention mechanism and multi-level memory library, and multi-level downsampling and upsampling are used to capture the multi-level features of time series data, and a dynamic sparse attention mechanism is introduced to optimize computing resource allocation.

Benefits of technology

The balance of accuracy, robustness and practicality in sparse daily power load prediction is achieved, especially the effective capture of periodic and seasonal patterns, reducing the prediction error of emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119944671B_ABST
    Figure CN119944671B_ABST
Patent Text Reader

Abstract

The present invention discloses an electric load forecasting method and system based on the GLA-UNet model structure, which relates to the technical field of electric load forecasting. The key points of the technical solution are: projecting single variable or multivariate data to a high-dimensional feature space through a linear mapping layer to generate an embedded vector sequence; performing multi-level downsampling on the embedded vector, each level including an average pooling layer and a GLA attention unit, outputting coding features of different time granularities; upsampling the deepest coding features step by step, each level fuses the coding features of the corresponding level and optimizes through the GLA attention unit, and outputs a reconstructed sequence with the same resolution as the input; adjusting the time length and feature dimension of the reconstructed sequence respectively through two linear layers to generate a power load forecast value. The present invention can effectively capture the multi-level features in time series data, and achieves a balance between accuracy, robustness and practicality in sparse daily-level power load forecasting, especially periodic and seasonal patterns.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric load forecasting, and more specifically, to an electric load forecasting method and system based on the GLA-UNet model structure. Background Art

[0002] In the new power system era led by smart grid technology, accurate forecasting of electricity demand has become the cornerstone of modern power system management. In building energy consumption management, precise electric load forecasting is not only a necessary condition for improving energy efficiency and reducing operating costs, but also the key to the deep integration of smart grid and renewable energy. At the same time, building load forecasting plays a core role in the local energy market and is an important basis for point-to-point energy trading decisions. With the rapid development of deep learning, deep learning methods have been widely applied in the field of electric load forecasting. Researchers have improved the electric load forecasting model continuously, as well as new methods and ideas provided by the time series forecasting field, all of which have improved the accuracy of electric load forecasting in order to optimize energy utilization efficiency and enhance management effectiveness.

[0003] Currently, most studies on electric load forecasting are aimed at datasets with hourly or minute-level minimum sampling units for forecasting. The data volume is relatively large, so good results are shown for both short-term and long-term forecasting. However, in reality, many buildings are limited by factors such as technical barriers, cost considerations, and privacy protection policies, and can only obtain load data with daily-level minimum sampling units. This actual situation has not been fully considered and explored in existing research. The data collected with daily-level minimum sampling units is not only sparse due to limited sample size, but also has more significant periodic and seasonal fluctuation characteristics. For example, in the case of a sudden power outage, the electric load value will form an obvious cliff-like drop; once the power supply is restored, the electricity demand will soar rapidly. In this case, although the daily-level data masks the specific moments, such extreme events not only disrupt the periodic pattern of the data, but also may cause the model to be difficult to accurately capture the normal daily and weekly fluctuations. In addition, for the alternation of seasons, the data on certain specific days may show significant increases or decreases, which further increases the complexity of forecasting. For example, the sudden increase in winter heating demand or the surge in summer air conditioner usage may generate abnormal peaks or valleys in the data of a certain day, thus affecting the overall seasonal trend analysis. The current research ignores the characteristics and challenges of the daily-level data described above, which may lead to a significant reduction in the performance of existing models in practical applications, and further affect the comprehensiveness and accuracy of energy management decisions.

[0004] Therefore, how to research and design an electric load forecasting technology that can overcome the above defects is an urgent problem for us to solve currently. Summary of the Invention

[0005] To address the deficiencies in the prior art, the objective of the present invention is to provide an electric load forecasting method and system based on the GLA-UNet model structure. By improving the traditional UNet model, a GLA-UNet model structure is established, making it more suitable for the field of electric load forecasting. It can effectively capture multi-level features in time series data and achieve a balance of accuracy, robustness, and practicality in sparse daily electric load forecasting, especially for periodic and seasonal patterns.

[0006] The above technical objective of the present invention is achieved through the following technical solutions:

[0007] In the first aspect, an electric load forecasting method based on the GLA-UNet model structure is provided, including the following steps:

[0008] Receive the original daily electric load data and covariate data, perform anomaly detection and / or filling operations, and output standardized time series data;

[0009] Project the preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate an embedded vector sequence;

[0010] Perform multi-level downsampling on the embedded vectors. Each level includes an average pooling layer and a GLA attention unit, and output encoded features with different time granularities;

[0011] Upsample the encoded features of the deepest layer step by step. Each level fuses the encoded features of the corresponding level and optimizes them through a GLA attention unit, and output a reconstruction sequence with the same resolution as the input;

[0012] Adjust the time length and feature dimension of the reconstruction sequence through two linear layers respectively to generate the electric load prediction value for the future period.

[0013] Furthermore, the execution of anomaly detection and / or filling operations includes any one or more of the following:

[0014] Perform deletion operations on data segments where the proportion of missing values or outliers is less than 5%;

[0015] Use a sliding window average to fill continuous abnormal regions, with a window size of 3 time points before and after;

[0016] Normalize the temperature and wind speed covariates, and perform one-hot encoding on the holiday variable.

[0017] Furthermore, the GLA attention unit performs the following operations:

[0018] Use multiple dilated convolutional layers with different dilation rates to extract multi-scale local features in parallel;

[0019] Capturing Long-Term Periodic Global Features through Learnable Global Token Vectors

[0020] Automatically generate gating coefficients based on input features, and weighted fuse multi-scale local features and long-term periodic global features according to the gating coefficients.

[0021] Furthermore, the generation expression of the gating coefficient is as follows:

[0022] ;

[0023] where represents the gating coefficient; represents the Sigmoid activation function; is the weight matrix; and are the outputs of the global attention layer and the local attention layer respectively.

[0024] Furthermore, the GLA attention unit also performs the following operations by introducing dynamic sparse attention:

[0025] Perform periodic intensity analysis on the input sequence to divide the high-fluctuation area through the local attention mechanism and the stable area through the global attention mechanism;

[0026] And, introduce deformable convolution in the decoding stage to dynamically adjust the convolution kernel sampling position according to the attention heat map.

[0027] Furthermore, the multi-level downsampling adopts three-level downsampling, and the sampling ratios are 1 / 2, 1 / 2, 1 / 2 respectively, and the sequence length of the output encoded features is 1 / 8 of the input length.

[0028] Furthermore, the upsampling operation adopts the bilinear interpolation algorithm, and the number of channels is adjusted through a 1×1 convolutional layer after interpolation.

[0029] Furthermore, adjusting the time length and feature dimension of the reconstructed sequence through two linear layers respectively includes:

[0030] Expand the sequence length to the target prediction length through the first linear layer;

[0031] Compress the feature dimension to the number of output variables through the second linear layer.

[0032] Furthermore, the method also includes:

[0033] Configure a multi-level memory bank composed of a short-term memory layer, a seasonal memory layer, and an event memory layer at the end of the encoder in the UNet structure; where

[0034] The short-term memory layer is used to store the load patterns of consecutive days and is updated using LSTM;

[0035] A seasonal memory layer for storing historical data of the same period by month and retrieving it through cosine similarity;

[0036] An event memory layer for recording the load response patterns of extreme events.

[0037] In a second aspect, an electric load prediction system based on the GLA-UNet model structure is provided, including:

[0038] A data preprocessing module for receiving original daily electric load data and covariate data, performing anomaly detection and / or filling operations, and outputting standardized time series data;

[0039] A feature embedding module for projecting preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate an embedded vector sequence;

[0040] A multi-scale encoding module for performing multi-level downsampling on the embedded vectors, with each level including an average pooling layer and a GLA attention unit, and outputting encoded features with different time granularities;

[0041] A context decoding module for performing upsampling on the deepest encoded features level by level, fusing the encoded features of the corresponding levels at each level and optimizing them through a GLA attention unit, and outputting a reconstructed sequence with the same resolution as the input;

[0042] A prediction output module for adjusting the time length and feature dimension of the reconstructed sequence through two linear layers respectively to generate the predicted value of the electric load for future periods.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] 1. The electric load prediction method based on the GLA-UNet model structure provided by the present invention. The UNet model is widely used in research fields such as images. The present invention improves the traditional UNet model to establish a GLA-UNet model structure, making it more suitable for the field of electric load prediction, capable of effectively capturing multi-level features in time series data, and achieving a balance of accuracy, robustness, and practicality in sparse daily electric load prediction, especially for periodic and seasonal patterns;

[0045] 2. The present invention proposes a global-local multi-scale perception attention mechanism module and embeds it into the UNet network to enhance the model's capture of the overall changes and local variations of sequence information;

[0046] 3. In the global-local multi-scale perception attention mechanism module of the present invention, a cross-scale correlation attention mechanism is proposed as the local attention mechanism. Trend features of different scales are extracted through multiple convolutional layers, and these features are fused through the multi-head attention mechanism, thereby realizing the model's ability to capture key trend features;

[0047] 4. In the global-local multi-scale perception attention mechanism module of the present invention, in order to enhance the module's ability to model global information, a multi-head global query attention mechanism is proposed. By adding global query tokens to each head of the attention, the model can focus on the context information of the overall sequence and better understand the overall structure of the sequence;

[0048] 5. The present invention also introduces dynamic sparse attention in the GLA module, which constitutes a hybrid routing mechanism with local-global attention. It can break through the full-sequence calculation mode of the traditional attention mechanism. By dynamically dividing regions, the computing resources are concentrated on key periods (such as the power outage recovery period), and the computational complexity is reduced;

[0049] 6. The present invention also configures a multi-level memory bank composed of a short-term memory layer, a seasonal memory layer, and an event memory layer at the end of the encoder in the UNet structure. By explicitly modeling the multi-level periodic characteristics of the power load, the prediction error for emergencies can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:

[0051] Figure 1 is the flowchart in Embodiment 1 of the present invention;

[0052] Figure 2 is the box plot before outlier processing in Embodiment 2 of the present invention;

[0053] Figure 3 is the box plot after outlier processing in Embodiment 2 of the present invention;

[0054] Figure 4 is the electricity consumption change diagram in Embodiment 2 of the present invention;

[0055] Figure 5 is the maximum temperature change diagram in Embodiment 2 of the present invention;

[0056] Figure 6 is the minimum temperature change diagram in Embodiment 2 of the present invention;

[0057] Figure 7 is the wind speed change diagram in Embodiment 2 of the present invention;

[0058] Figure 8 It is the graph of whether it is a holiday in Embodiment 2 of the present invention;

[0059] Figure 9 It is the comparison graph of the MSE of each model for predicting the next day in Embodiment 2 of the present invention;

[0060] Figure 10 It is the comparison graph of the MSE of each model for predicting the next two days in Embodiment 2 of the present invention;

[0061] Figure 11 It is the comparison graph of the MSE of each model for predicting the next three days in Embodiment 2 of the present invention;

[0062] Figure 12 It is the comparison graph of the MAE of each model for predicting the next day in Embodiment 2 of the present invention;

[0063] Figure 13 It is the comparison graph of the MAE of each model for predicting the next two days in Embodiment 2 of the present invention;

[0064] Figure 14 It is the comparison graph of the MAE of each model for predicting the next three days in Embodiment 2 of the present invention;

[0065] Figure 15 It is the comparison curve graph of the predicted value and the true value of the single-variable data set for predicting the next day in Embodiment 2 of the present invention;

[0066] Figure 16 It is the comparison curve graph of the predicted value and the true value of the single-variable data set for predicting the next two days in Embodiment 2 of the present invention;

[0067] Figure 17 It is the comparison curve graph of the predicted value and the true value of the single-variable data set for predicting the next three days in Embodiment 2 of the present invention;

[0068] Figure 18 It is the comparison curve graph of the predicted value and the true value of the multi-variable data set for predicting the next day in Embodiment 2 of the present invention;

[0069] Figure 19 It is the comparison curve graph of the predicted value and the true value of the multi-variable data set for predicting the next two days in Embodiment 2 of the present invention;

[0070] Figure 20 It is the comparison curve graph of the predicted value and the true value of the multi-variable data set for predicting the next three days in Embodiment 2 of the present invention;

[0071] Figure 21 It is the schematic diagram of the GLA-UNet model structure in Embodiment 3 of the present invention. Detailed implementation manners

[0072] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments and drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and are not intended to limit the present invention.

[0073] Embodiment 1: An electric load forecasting method based on the GLA-UNet model structure, as Figure 1 shown, includes the following steps:

[0074] S1: Receive the original daily power load data and covariate data, perform anomaly detection and / or filling operations, and output standardized time series data;

[0075] S2: Project the preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate an embedded vector sequence;

[0076] S3: Perform multi-level downsampling on the embedded vectors. Each level includes an average pooling layer and a GLA attention unit, and output encoded features with different time granularities;

[0077] S4: Upsample the encoded features of the deepest layer level by level. Each level fuses the encoded features of the corresponding level and optimizes them through a GLA attention unit, and output a reconstruction sequence with the same resolution as the input;

[0078] S5: Adjust the time length and feature dimension of the reconstruction sequence through two linear layers respectively to generate the predicted value of the future power load.

[0079] In some of the existing technologies, the UNet+BiLSTM architecture mainly processes spatio-temporal features serially, and cannot solve the contradiction between local mutations and global trends. Moreover, due to the sequential calculation mode of BiLSTM, the time complexity is relatively high, making it difficult to process long-cycle data.

[0080] In order to effectively capture multi-level features in time series data and achieve a balance of accuracy, robustness, and practicality in sparse daily power load forecasting, the present invention designs the GLA-UNet model structure. In the present invention, GLA refers to the global-local attention mechanism, and UNet is an encoder-decoder structure commonly used in image segmentation.

[0081] In step S1, the original daily power load data is the power consumption data sampled in days, and the covariate data includes but is not limited to temperature, wind speed, etc.

[0082] In some examples, the missing values and outliers of the data are first checked. For data with fewer missing and abnormal data, the deletion method is adopted to ensure data quality; in some examples, for data with more missing values and outliers, the average value of the data at three points before and after the outlier or missing value point is used for filling to ensure the stability of the sequence data to the greatest extent.

[0083] In some examples, only the original daily power load data can be input into the model for training to obtain univariate data; the number of characteristic variables of the univariate data is 1, and at this time, the size of the input data of the model is , where T is the length of the input sequence. In some examples, covariate data can also be added to the dataset to implement covariate data training. The covariates introduced by the multivariate data are the highest temperature, the lowest temperature, the wind speed, and whether it is a holiday. At this time, the size of the input data of the model is .

[0084] In step S2, for the input data , first, it passes through an embedding layer to map the data. The embedding layer consists of a linear layer, which maps the low-dimensional data to high-dimensional, so as to effectively capture the interaction between different features and provide more abundant information support for subsequent predictions.

[0085] In step S3, at this time, the size of the mapped input data is . Then, a method of extracting features in multiple layers is adopted. The stage is set to 3, and the average pooling layer is used to extract trend features. After each feature extraction, it passes through multiple GLA attention units (hereinafter referred to as GLA modules) to capture the global and local dependence relationships. The GLA module adopts the global-local attention mechanism.

[0086] After each layer of average pooling layer, the sequence length changes as follows:

[0087] ;

[0088] Among them, represents the feature output by the th layer; represents the feature output by the th layer; is the sequence length of the feature output by the th layer, is the padding size, is the convolution kernel size, is the stride length.

[0089] In some examples, in order to quickly capture the local trend information in the daily electricity load data and enable the model to effectively extract multi-scale features at different time scales, so as to capture the impact of extreme events, the present invention proposes a cross-scale correlation attention mechanism as a local attention mechanism, which is used as the local attention mechanism in the GLA module.

[0090] The cross-scale correlation attention mechanism proposed by the present invention has multiple convolutional layers with different kernel sizes and dilation rates for extracting trend features at different time scales. By applying these convolutions to the query (Q) and key (K), multiple representations sensitive to different time scales are generated, which are then concatenated and processed through the multi-head attention mechanism. The convolutional layers are designed to capture trends of different durations by adjusting their receptive fields. Larger kernel sizes and higher dilation rates enable the model to consider longer time spans, while smaller kernels and lower dilation rates focus on shorter periods. This multi-scale approach ensures that the attention mechanism can utilize the understanding of electricity load dynamics to achieve more accurate predictions.

[0091] Specifically, for a set of input sequences , where B represents the batch size, T represents the sequence length, and D represents the dimension after embedding. For the Q query matrix and the K key matrix, a set of convolutional kernels and dilation rates are used to extract trend information at different scales. The specific expression is:

[0092] ;

[0093] Among them, represents the multi-scale local features extracted by the convolutional kernel and the dilation rate ; represents the key matrix corresponding to , characterizing the correlation at different time scales; represents the number of convolutional layers; represents the value matrix, which mainly maps the multi-scale features extracted by the convolution to a dimension suitable for attention calculation through the linear transformation layer as the value matrix in the attention mechanism; represents the convolution operation; represents the padding operation. To maintain the sequence length unchanged after passing through the convolutional layer, is padded, and the padding size is .

[0094] For the and of the multiple time-scale trends extractedFusion is performed along the sequence length dimension, and the expression is:

[0095] ;

[0096] Among them, represents the fusion result; represents the fusion result; represents tensor concatenation; represents the sequence length dimension.

[0097] After merging and the size of is . Next, perform multi-head splitting:

[0098] ;

[0099] Among them, represents the query matrix of the th attention head after splitting; represents the key matrix of the th attention head after splitting; represents the value matrix of the th attention head after splitting; represents the number of subspaces into which the feature is decomposed; represents dimension reshaping; is the new key vector length, and , is the number of heads for multi-head partitioning.

[0100] Calculate the operation result of the attention mechanism through scaled dot product, that is:

[0101] ;

[0102] Among them, represents the core unit of the th attention head, which is used to capture the multi-scale temporal correlation of power load data through parallel computing; represents the normalization activation function.

[0103] Finally, merge the results of all heads and use a linear layer to output the final attention calculation result. The specific expression is:

[0104] ;

[0105] Among them, represents the final attention calculation result.

[0106] In some examples, in order to enhance the model's ability to capture global features and help the model better understand the long-term trends and periodic patterns of the data, the present invention proposes a multi-head global query attention mechanism as the global attention mechanism in the GLA module. The present invention introduces a global query token on the basis of the traditional multi-head attention mechanism to act as an "information convergence point", and similarly for a set of input sequences , where B represents the batch size, T represents the sequence length, and D represents the embedding dimension.

[0107] First, define a multi-head global query vector , where H is the number of attention heads, and initialize the multi-head query vector :

[0108] ;

[0109] Among them, is the weight initialization.

[0110] Reshape the input sequence into , and repeat the global query vector to match the number of batches and heads:

[0111] ;

[0112] Among them, represents the extended global query vector; represents the tensor replication operation.

[0113] At this time, , splice the global query vector into the input to get : .

[0114] And , then perform the calculation of the scaled dot attention mechanism:

[0115] ;

[0116] Among them, . Remove the global query vector from the attention mechanism calculation and reshape it back to the original shape to obtain the final output result:

[0117] ;

[0118] Among them, represents the multi-dimensional tensor after global attention calculation; represents restoring the original dimension of the input sequence.

[0119] In some examples, the GLA module also introduces a fusion mechanism for dynamically combining the outputs of the global and local attention layers. This mechanism is implemented through a gating unit, which is a fully connected layer based on the Sigmoid activation function to determine the weights given to the global and local contexts. Specifically, the output of the fusion mechanism can be expressed by the following formula:

[0120] ;

[0121] where, is the Sigmoid activation function, is the weight matrix, and and are the outputs of the global attention layer and the local attention layer respectively; is the output of the fusion mechanism. When the value of G is close to 1, the model relies more on the global context; conversely, when the value of G is close to 0, the model relies more on the local context. This dynamic weighting mechanism enables the GLA module to flexibly adjust the importance of global and local information according to the data characteristics.

[0122] In step S4, the task of the decoder part is to gradually restore to the length of the original input sequence based on the feature sequence extracted by the encoder. The present invention adopts an upsampling mechanism based on a fully connected layer and a skip connection technique to finally predict the future power load value.

[0123] First, starting from the deepest layer of the encoder and after passing through L layers of the GLA module for upsampling, after passing through L layers of the GLA module, the result at this time is denoted as and , and similarly, after three stages, the length of is restored.

[0124] The length restored in each stage is the same as the sequence length of the encoder part . Using the skip connection technique allows the decoder part to directly access the feature maps of the encoder part, thus retaining important local change information during the prediction process. Finally, the obtained .

[0125] In step S5, to obtain the final prediction result, two linear layers are used to separately change the two dimensions of the sequence length and the number of channels in the prediction result:

[0126] ;

[0127] where, is the first linear layer, which is used to adjust the sequence length from T to the required target prediction length ; is the output of the first linear layer;

[0128] ;

[0129] wherein, is the second linear layer for adjusting the number of channels to the required prediction length ; is the output of the second linear layer. Finally, .

[0130] In some examples, the present invention also introduces dynamic sparse attention in the GLA module, which forms a hybrid routing mechanism with local-global attention to achieve: performing periodic intensity analysis on the input sequence, automatically dividing the high-fluctuation area (dominated by local attention) and the stable area (dominated by global attention); introducing deformable convolution in the decoding stage to dynamically adjust the sampling position of the convolution kernel according to the attention heat map. The above method can break through the full-sequence calculation mode of the traditional attention mechanism, concentrate computing resources on critical periods (such as the power outage recovery period) through dynamic area division, and reduce the computing complexity.

[0131] In some examples, the present invention also configures a multi-level memory bank composed of a short-term memory layer, a seasonal memory layer, and an event memory layer at the end of the encoder in the UNet structure.

[0132] Among them, the short-term memory layer is used to store the load patterns of consecutive days and is updated using LSTM (Long Short-Term Memory Network); the seasonal memory layer is used to store historical data of the same period by month and is retrieved through cosine similarity; the event memory layer is used to record the load response patterns of extreme events. By explicitly modeling the multi-level periodic characteristics of the power load, the present invention can reduce the prediction error of unexpected events.

[0133] Example 2: Verification is carried out using the power load dataset of an office building

[0134] The present invention verifies the effectiveness of the proposed model by comparing the mean square error (MSE) and mean absolute error (MAE) of the results of the model on univariate and multivariate datasets, and verifies whether the introduced covariates will affect the accuracy of the power load.

[0135] Step 1: In this invention, a power load dataset of an office building in a certain city is used. The dataset takes a day as the minimum collection unit and covers the daily power load data for two consecutive years from January 1, 2022 to December 31, 2023, that is, there are 730 time nodes of power consumption data. The additional covariates introduced are the daily maximum temperature, minimum temperature, wind speed, and whether the day is a holiday. These added covariates are used to verify the effectiveness of the model in a multivariate dataset and whether it will affect the final prediction result. The specific description of the dataset is shown in Table 1.

[0136] Table 1 The constructed multivariate dataset

[0137]

[0138] Step 2: Screen the dataset for missing values and outliers. In this example, there are no missing values in the dataset. For outliers, box plots are drawn to determine whether there are outliers, and values higher than 4 times the standard deviation of the original data are defined as outliers. The comparison of box plots before and after removing outliers is as Figure 2 、 Figure 3 shown. As Figure 2 can be seen, there are obvious outliers in the data, and the maximum value is higher than 10000, but the number of outliers is small. It can be seen from the figure that there are only two outlier cases. Therefore, the deletion method is used to directly remove the outliers. The box plot of the dataset after removal is as Figure 3 shown. Removing outliers can ensure that the model training will not be affected by outliers, thus ensuring the training effect of the model.

[0139] Step 3: Process the processed dataset. First, the changes in power consumption, maximum temperature, minimum temperature, wind speed, and whether it is a holiday are plotted, respectively Figures 4 - 8 shown. As Figure 4 shown. The daily change in power load consumption shows certain internal rules, but at the same time, it is accompanied by significant fluctuations. This indicates that the power load data with a day as the minimum sampling unit not only reflects obvious periodic and seasonal characteristics but also reflects the impact of unforeseen events in daily life on power demand. The regularity of the change trends of the maximum temperature and minimum temperature is relatively obvious, as Figure 5 、 Figure 6As shown, it reached its peak around July and its trough around January in both years. The wind speed change did not show obvious regular characteristics. The monthly consumption change of the power load demonstrated the annual periodicity of the power load change, reflecting the change characteristics of the power consumption within 12 months of the whole year. This periodicity was mainly affected by seasonal factors, and different seasons brought different demand patterns for production activities and residents' lives. Taking Shanghai as an example, as a representative city in East China, the annual periodicity of its power load was particularly significant. According to on-site research, Shanghai divided a year into three main stages based on local climate characteristics: the heating season (from November to March of the following year), the cooling season (from May to September), and two shorter transitional seasons (April and October).

[0140] Step Four: To verify the effectiveness of the introduced multi-variable factors, first, calculate the Pearson correlation coefficient to analyze the correlation between the multi-variable factors and the power load. The Pearson correlation coefficient can be used to measure the degree of linear correlation between variables. Therefore, the Pearson correlation coefficient can calculate the correlation between the power load and multi-factor variables.

[0141] As shown in Table 2, there was a positive correlation between the maximum temperature, minimum temperature and the power load, that is, as the temperature increased or decreased, the power demand also increased accordingly. At the same time, the wind speed and whether it was a holiday showed a negative correlation with the power load, meaning that higher wind speed and non-working days tended to reduce power consumption. However, except for the factor of holidays, the relationship between other variables and the power load was not a simple linear correlation, indicating that relying solely on linear assumptions might not accurately verify the impact of multi-variable factors on the power load.

[0142] Table 2 Calculation Results of Pearson Correlation Coefficient

[0143]

[0144] Step Five: To verify the impact of introducing multi-variable factors on the prediction effect, by comparing the performance of multiple models of single variable (only considering historical power load data) and multi-variable (combining temperature, wind speed, holidays, etc.), verify the effectiveness of the model proposed in the present invention, and at the same time explore whether multi-variable factors have an impact on the model prediction performance for verification. The specific data set segmentation and hyperparameter selection are as follows.

[0145] The single-variable experiment only used the power load data of the past 16 days to predict the power load of the next 1, 2, and 3 days. The multi-variable experiment used the power load data of the past 16 days and supplementary covariates to predict the power load of the next 1, 2, and 3 days. The time span of the data set was from January 1, 2022 to December 31, 2023, totaling 730 days. The specific data segmentation was as follows.

[0146] Training set: It accounts for 70% of the total data, that is, the data from January 1, 2022 to May 26, 2023 is used for the training of the model. This time period covers multiple seasonal changes and can fully reflect the annual periodic characteristics of the electricity load.

[0147] Validation set: It accounts for 10% of the total data, that is, the data from May 27, 2023 to July 31, 2023 is used for the validation of the model.

[0148] Test set: It accounts for 20% of the total data, that is, the data from August 1, 2023 to December 31, 2023 is used for the final evaluation of the model performance.

[0149] Step Six: The proposed model GLA-UNet of the present invention was compared with several models such as Transformer, Informer, Reformer, and ITransformer. The hyperparameters of the comparison models are shown in Table 3.

[0150] Table 3 Training Hyperparameters of Each Model

[0151]

[0152] Step Seven: In order to produce reliable results, each model was trained 10 times separately, and the losses were averaged. When comparing the performances of different prediction time periods (predicting one day, two days, and three days), the differences shown by each model in terms of the mean squared error (MSE) and mean absolute error (MAE) metrics can be observed.

[0153] As can be seen from Tables 4 and 5, the GLA-UNet model shows the lowest errors in terms of MSE and MAE metrics for all prediction periods in both univariate and multivariate scenarios. For the univariate dataset, the prediction results are 0.141 and 0.318 (predicting one day), 0.308 and 0.397 (predicting two days), and 0.425 and 0.457 (predicting three days). After introducing multivariate factors, they are 0.102 and 0.251 (predicting one day), 0.194 and 0.318 (predicting two days), and 0.216 and 0.368 (predicting three days). This indicates that GLA-UNet has significant advantages in the power load forecasting task and can effectively capture the trend of power load changes over time. In contrast, the Reformer model uses Locality-Sensitive Hashing Attention (LSH Attention) to replace the traditional multi-head attention mechanism. However, power load forecasting requires precise capture of long-term dependencies, and the approximate nature of LSH may lead to information loss and affect prediction accuracy. Therefore, it shows higher MSE and MAE in both univariate and multivariate scenarios for each prediction period. Especially in the three-day prediction, the univariate prediction MSE reaches 0.982 and the MAE reaches 0.799, while the multivariate prediction MSE reaches 0.788 and the MAE reaches 0.700. This indicates that the model may face higher uncertainty when dealing with predictions over longer time periods. Compared with the Reformer model, the prediction performance of the Transformer, Informer, and ITransformer models has been greatly improved, but overall, they are still inferior to the GLA-UNet model.

[0154] Table 4 Results of each model on the test set of the univariate dataset

[0155]

[0156] Table 5 Results of each model on the test set of the multivariate dataset

[0157]

[0158] Step 8: Plot the change curves of MSE and MAE for each model under univariate and multivariate conditions. It can be observed from Figures 9 - 14 that Figures 9 - 14In it, "Univariate" means single variable and "Multivariate" means multiple variables. When predicting the results for the next day, compared with the model after introducing multiple variable factors, the mean squared error (MSE) of the multivariate prediction decreased by 5.1%, and the mean absolute error (MAE) decreased by 4.1%. When the prediction range is extended to the next two days, the MSE and MAE of the multivariate prediction have more obvious improvements, with the decline reaching 12.8% and 9.2% respectively. In the case of predicting the next three days, this difference further expands, and the MSE and MAE of the multivariate prediction decrease by 24.4% and 15.5% respectively compared with the multivariate model. It can be concluded that although introducing multiple variable factors can indeed improve the prediction performance of the model, in the short-term prediction (such as one to two days), its improvement effect on the overall prediction accuracy is relatively limited. As the prediction time span increases, the importance of multiple variable factors gradually emerges, playing a more crucial role in improving prediction accuracy. This indicates that in short-term prediction, the power load data occupies a relatively dominant position, and with the time span, the influence of multiple variable-related factors gradually appears. This also provides ideas for actual power load prediction. For short-term prediction tasks, it can be considered to sacrifice a certain prediction accuracy to effectively save the cost of data collection and processing without affecting the decision-making quality. When the prediction range is extended to the medium and long term, more potential influencing factors need to be considered more comprehensively. At this time, introducing multiple variable factors not only helps to improve the prediction accuracy but also enhances the model's adaptability to future uncertainties. Therefore, when conducting medium and long-term power load prediction, although obtaining and integrating more dimensional data may bring additional costs, in the long run, this can provide a more accurate and reliable basis for power system planning and operation, ensuring the effective allocation of resources and the stable operation of the system.

[0159] Step Nine: To more intuitively display the prediction results, select the comparison results of the predicted values and the true values of each model in the test set. The single-variable and multi-variable prediction results are as Figures 15 - 20 shown.

[0160] As can be seen from the figure, the proposed model GLA-UNet shows a better fitting effect compared to other models in univariate and multivariate prediction. In the overall range, it can be seen that the GLA-UNet model has a stronger ability to capture periodicity and shows the best fitting effect for the peaks in the global range. At the fluctuation points, the proposed model GLA-UNet has the smallest volatility in univariate and multivariate prediction compared to other models, and can relatively accurately reflect the power load trend during this period. After a short period of stability, the electricity load increases sharply, and it can be clearly seen that the GLA-UNet model has the strongest ability to capture this sudden situation. It not only responds quickly to the load increase but also maintains a certain degree of accuracy during the prediction process. For the same number of prediction days, the multivariate prediction results of each model are significantly better than the univariate prediction results. Especially in the capture of the overall periodicity, the univariate prediction has a weak ability to capture the peaks in the period. As the prediction length increases, the prediction results of each model can only capture the overall periodic changes, but the gap between the predicted value and the true value gradually increases. At the fluctuation points, the volatility of the univariate prediction results of each model is higher than that of the multivariate prediction results. Therefore, the univariate prediction has a relatively poor ability to capture the local characteristics of non-periodic fluctuations.

[0161] Embodiment 3: An electric load prediction system based on the GLA-UNet model structure, which is used to implement the electric load prediction method based on the GLA-UNet model structure described in Embodiment 1, as Figure 21 shown, including a data preprocessing module, a feature embedding module, a multi-scale encoding module, a context decoding module, and a prediction output module.

[0162] Among them, the data preprocessing module is used to receive the original daily power load data and covariate data, perform anomaly detection and / or filling operations, and output standardized time series data; the feature embedding module is used to project the preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate an embedded vector sequence; the multi-scale encoding module is used to perform multi-level downsampling on the embedded vector, each level including an average pooling layer and a GLA attention unit, and output encoded features with different time granularities; the context decoding module is used to upsample the deepest encoded features level by level, fuse the encoded features of the corresponding level at each level and optimize them through the GLA attention unit, and output a reconstruction sequence with the same resolution as the input; the prediction output module is used to adjust the time length and feature dimension of the reconstruction sequence through two linear layers respectively to generate the predicted value of the power load in the future period.

[0163] Working principle: By improving the traditional UNet model, the present invention establishes a GLA-UNet model structure, making it more suitable for the field of power load forecasting. It can effectively capture multi-level features in time series data and achieve a balance of accuracy, robustness, and practicality in sparse daily power load forecasting, especially for periodic and seasonal patterns.

[0164] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0165] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure One one process or multiple processes and / or blocks Figure One one block or multiple blocks.

[0166] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure One one process or multiple processes and / or blocks Figure One one block or multiple blocks.

[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure One one process or multiple processes and / or blocks Figure One one block or multiple blocks.

[0168] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An electric load forecasting method based on the GLA-UNet model structure, characterized in that, It includes the following steps: Receiving the original daily power load data and covariate data, performing anomaly detection and / or filling operations, and outputting standardized time series data; Projecting the preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate a sequence of embedding vectors; Performing multi-level downsampling on the embedding vectors, with each level containing an average pooling layer and a GLA attention unit, and outputting encoded features with different time granularities; Performing upsampling on the encoded features of the deepest layer level by level, fusing the encoded features of the corresponding levels at each level and optimizing through the GLA attention unit, and outputting a reconstruction sequence with the same resolution as the input; Adjusting the time length and feature dimension of the reconstruction sequence through two linear layers respectively to generate the power load prediction value for the future period; The GLA attention unit performs the following operations: Using multiple dilated convolutional layers with different dilation rates to extract multi-scale local features in parallel; Capturing long-term periodic global features through a learnable global token vector; Automatically generating a gating coefficient according to the input features, and weighted fusing the multi-scale local features and the long-term periodic global features according to the gating coefficient; The generation expression of the gating coefficient is: ; Among them, represents the gating coefficient; represents the Sigmoid activation function; is the weight matrix; and are the outputs of the global attention layer and the local attention layer respectively; The method further includes: Configuring a multi-level memory bank composed of a short-term memory layer, a seasonal memory layer, and an event memory layer at the end of the encoder in the UNet structure; wherein, The short-term memory layer is used to store the load patterns of consecutive days and is updated using LSTM; The seasonal memory layer is used to store historical data for the same period by month and is retrieved through cosine similarity; The event memory layer is used to record the load response patterns of extreme events.

2. The electric load prediction method based on the GLA-UNet model structure according to claim 1, wherein The performing of the anomaly detection and / or filling operations includes any one or more of the following: Performing a deletion operation on a data segment where the proportion of missing values or outliers is less than 5%; Using a sliding window average to fill the continuous anomaly region, with the window size being 3 time points before and after; Normalizing the temperature and wind speed covariates and performing one-hot encoding on the holiday variable.

3. The electric load prediction method based on the GLA-UNet model structure according to claim 1, characterized in that, The GLA attention unit also performs the following operations by introducing dynamic sparse attention: Performing periodic intensity analysis on the input sequence to divide the high-fluctuation region through the local attention mechanism and the stable region through the global attention mechanism; And, introducing deformable convolution in the decoding stage to dynamically adjust the sampling position of the convolution kernel according to the attention heat map.

4. The electric load prediction method based on the GLA-UNet model structure according to claim 1, wherein The multi-level downsampling adopts three-level downsampling, with the sampling ratios being 1 / 2, 1 / 2, and 1 / 2 respectively, and the sequence length of the output encoded features is 1 / 8 of the input length.

5. The electric load prediction method based on the GLA-UNet model structure according to claim 1, wherein The upsampling operation adopts a bilinear interpolation algorithm, and the number of channels is adjusted through a 1×1 convolutional layer after interpolation.

6. The electric load prediction method based on the GLA-UNet model structure according to claim 1, wherein The adjusting of the time length and feature dimension of the reconstruction sequence through two linear layers respectively includes: Expanding the sequence length to the target prediction length through the first linear layer; Compressing the feature dimension to the number of output variables through the second linear layer.

7. An electric load prediction system based on the GLA-UNet model structure, characterized in that, This system is used to implement the electric load prediction method based on the GLA-UNet model structure as described in any one of claims 1-6, including: A data preprocessing module, which is used to receive the original daily power load data and covariate data, perform anomaly detection and / or filling operations, and output standardized time series data; A feature embedding module, which is used to project the preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate a sequence of embedding vectors; A multi-scale encoding module, which is used to perform multi-level downsampling on the embedding vectors. Each level includes an average pooling layer and a GLA attention unit, and outputs encoded features with different time granularities; A context decoding module, which is used to perform upsampling on the deepest encoded features level by level. Each level fuses the encoded features of the corresponding level and optimizes them through a GLA attention unit, and outputs a reconstructed sequence with the same resolution as the input; A prediction output module, which is used to adjust the time length and feature dimension of the reconstructed sequence through two linear layers respectively to generate the predicted value of the power load in the future period.

Citation Information

Patent Citations

  • Building extraction model training method, extraction method, equipment and storage medium

    CN115439756A

  • Power load prediction method based on multivariable time sequence information interaction

    CN118333232A