Electric load prediction method and system based on GLA-UNet model structure

By improving the UNet model, introducing global-local multi-scale perceived attention mechanism and dynamic sparse attention, the problem that existing models are difficult to capture periodic and seasonal patterns when processing sparse daily power load data is solved, achieving higher prediction accuracy and robustness.

CN119944671AActive Publication Date: 2025-05-06SHANGHAI UNIVERSITY OF ELECTRIC POWER

Patent Information

Application Number
CN202510417343.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

When existing power load prediction models process sparse data with the minimum sampling unit of day, it is difficult to accurately capture periodic and seasonal patterns, resulting in a decrease in prediction accuracy and affecting the comprehensiveness and accuracy of energy management decisions.

Method used

The electric load prediction method based on the GLA-UNet model structure is adopted, and the global-local multi-scale perceived attention mechanism, multi-head global query attention mechanism and dynamic sparse attention are introduced to enhance the model's multi-level feature capture ability of time series data, especially when processing sparse daily power load data.

Benefits of technology

In the prediction of sparse daily power load, a balance of accuracy, robustness and practicality is achieved, especially in capturing periodic and seasonal patterns, which significantly improves the accuracy and reliability of the prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119944671A_ABST
    Figure CN119944671A_ABST
Patent Text Reader

Abstract

The invention discloses an electrical load prediction method and system based on a GLA-UNet model structure, and relates to the technical field of electrical load prediction, and the method is characterized in that univariate or multivariate data is projected to a high-dimensional feature space through a linear mapping layer, and an embedded vector sequence is generated; multi-level down-sampling is carried out on the embedded vector, each level comprises an average pooling layer and a GLA attention unit, and coding features of different time granularities are output; carrying out step-by-step up-sampling on the deepest coding features, fusing the coding features of the corresponding layer in each level, optimizing through a GLA attention unit, and outputting and inputting a reconstruction sequence with the same resolution; and respectively adjusting the time length and the feature dimension of the reconstructed sequence through the two linear layers, and generating a power load prediction value. According to the method, the multi-level features in the time sequence data can be effectively captured, and the balance of precision, robustness and practicability, especially periodic and seasonal modes, is realized in sparse day-level power load prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric load forecasting, and more specifically, to an electric load forecasting method and system based on a GLA-UNet model structure. Background Art

[0002] In the era of new power systems led by smart grid technology, accurate prediction of power demand has become the cornerstone of modern power system management. In building energy consumption management, precise power load forecasting is not only a necessary condition for improving energy efficiency and reducing operating costs, but also the key to the deep integration of smart grids and renewable energy. At the same time, building load forecasting plays a core role in the local energy market and is an important basis for point-to-point energy trading decisions. With the rapid development of deep learning, deep learning methods have been widely used in the field of power load forecasting. Researchers have continuously improved the power load forecasting model and the new methods and ideas provided in the field of time series forecasting to improve the accuracy of power load forecasting, in order to optimize energy utilization efficiency and improve management effectiveness.

[0003] At present, most studies on power load forecasting are aimed at predicting data sets with hours or minutes as the minimum sampling unit. The data volume is relatively large, which shows good results for short-term or long-term forecasts. However, in reality, many buildings are limited by technical barriers, cost considerations, and privacy protection policies, and can only obtain load data with days as the minimum sampling unit. This actual situation has not been fully considered and discussed in existing studies. The data collected with days as the minimum sampling unit is not only sparse due to the limited sample size, but also contains more significant periodic and seasonal fluctuation characteristics. For example, in the event of a sudden power outage, the power load value will form an obvious cliff-like drop; once the power supply is restored, the power demand will soar rapidly. In this case, although the daily data conceals the specific moment, this extreme event not only destroys the periodic pattern of the data, but also may make it difficult for the model to accurately capture normal daytime and weekly fluctuations. In addition, with the alternation of seasons, the data on certain specific days may increase or decrease significantly, which further increases the complexity of prediction. For example, a sudden increase in heating demand in winter or a surge in air conditioning use in summer may produce abnormal peaks or valleys in the data on a particular day, thus affecting the overall seasonal trend analysis. Current research ignores the characteristics and challenges of daily data described above, which may lead to a significant reduction in the performance of existing models in practical applications, thus affecting the comprehensiveness and accuracy of energy management decisions.

[0004] Therefore, how to research and design an electric load forecasting technology that can overcome the above-mentioned defects is an issue that we urgently need to solve. Summary of the invention

[0005] In order to address the deficiencies in the prior art, the purpose of the present invention is to provide an electric load forecasting method and system based on the GLA-UNet model structure. By improving the traditional UNet model, a GLA-UNet model structure is established to make it more suitable for the field of electric load forecasting, which can effectively capture the multi-level features in time series data and achieve a balance between accuracy, robustness and practicality in sparse daily electric load forecasting, especially periodic and seasonal patterns.

[0006] The above technical objectives of the present invention are achieved through the following technical solutions: In the first aspect, a method for predicting electric load based on the GLA-UNet model structure is provided, comprising the following steps: Receive original daily power load data and covariate data, perform anomaly detection and / or filling operations, and output standardized time series data; Project the preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate an embedding vector sequence; The embedding vector is downsampled at multiple levels, each level contains an average pooling layer and a GLA attention unit, and outputs encoding features of different time granularities; The deepest encoding features are upsampled level by level, and the encoding features of the corresponding level are fused at each level and optimized through the GLA attention unit to output a reconstructed sequence with the same resolution as the input; The time length and feature dimension of the reconstructed sequence are adjusted respectively through two linear layers to generate the power load forecast value for the future period.

[0007] Further, the performing of anomaly detection and / or filling operations includes any one or more of the following: Delete data segments where the number of missing values ​​or outliers is less than 5%; For continuous abnormal areas, sliding window average filling is used, and the window size is 3 time points before and after; The temperature and wind speed covariates are normalized, and the holiday variables are one-hot encoded.

[0008] Furthermore, the GLA attention unit performs the following operations: Use multiple atrous convolutional layers with different dilation rates to extract multi-scale local features in parallel; Capturing long-term periodic global features via learnable global token vectors The gating coefficients are automatically generated according to the input features, and the multi-scale local features and long-term periodic global features are weighted and fused according to the gating coefficients.

[0009] Furthermore, the generation expression of the gating coefficient is: ; in, represents the gating coefficient; Represents the Sigmoid activation function; is the weight matrix; and are the outputs of the global attention layer and the local attention layer respectively.

[0010] Furthermore, the GLA attention unit also performs the following operations by introducing dynamic sparse attention: Perform periodic strength analysis on the input sequence to separate high-volatility regions through local attention mechanisms and stable regions through global attention mechanisms; Also, deformable convolution is introduced in the decoding stage to dynamically adjust the convolution kernel sampling position according to the attention heat map.

[0011] Furthermore, the multi-level downsampling adopts three-level downsampling, and the sampling ratios are 1 / 2, 1 / 2, and 1 / 2 respectively, and the sequence length of the output coding feature is 1 / 8 of the input length.

[0012] Furthermore, the upsampling operation adopts a bilinear interpolation algorithm, and after interpolation, the number of channels is adjusted through a 1×1 convolution layer.

[0013] Furthermore, the step of adjusting the time length and feature dimension of the reconstructed sequence respectively through two linear layers includes: The sequence length is expanded to the target prediction length through the first linear layer; The feature dimensions are compressed to the number of output variables through the second linear layer.

[0014] Furthermore, the method also includes: The encoder end in the UNet structure is configured with a multi-level memory bank consisting of a short-term memory layer, a seasonal memory layer, and an event memory layer; The short-term memory layer is used to store the load patterns for consecutive days and is updated using LSTM; Seasonal memory layer, used to store historical data by month and retrieve it by cosine similarity; Event memory layer, used to record load response patterns to extreme events.

[0015] Secondly, an electric load forecasting system based on the GLA-UNet model structure is provided, including: A data preprocessing module, used to receive original daily power load data and covariate data, perform anomaly detection and / or filling operations, and output standardized time series data; Feature embedding module, which is used to project the preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate an embedding vector sequence; The multi-scale encoding module is used to perform multi-level downsampling of the embedding vector. Each level contains an average pooling layer and a GLA attention unit to output encoding features of different time granularities. The context decoding module is used to upsample the deepest encoding features step by step, fuse the encoding features of the corresponding layer at each level and optimize them through the GLA attention unit, and output a reconstructed sequence with the same resolution as the input; The prediction output module is used to adjust the time length and feature dimension of the reconstruction sequence through two linear layers to generate the power load prediction value for the future period.

[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. The electric load forecasting method based on the GLA-UNet model structure provided by the present invention. The UNet model is widely used in research fields such as images. The present invention improves the traditional UNet model to establish a GLA-UNet model structure, making it more suitable for the field of electric load forecasting, and can effectively capture the multi-level features in time series data. In the sparse daily electric load forecasting, the balance between accuracy, robustness and practicality is achieved, especially the periodicity and seasonal patterns; 2. The present invention proposes a global-local multi-scale perception attention mechanism module and embeds it into the UNet network to enhance the model's ability to capture overall changes and local changes in sequence information; 3. In the global-local multi-scale perception attention mechanism module, the present invention proposes a cross-scale correlation attention mechanism as a local attention mechanism, extracts trend features of different scales through multiple convolutional layers, and fuses these features through a multi-head attention mechanism, thereby realizing the model's ability to capture key trend features; 4. In the global-local multi-scale perception attention mechanism module, the present invention proposes a multi-head global query attention mechanism to enhance the module's ability to model global information. By adding a global query token to each head of attention, the model can focus on the contextual information of the entire sequence and better understand the overall structure of the sequence. 5. The present invention also introduces dynamic sparse attention in the GLA module, which forms a hybrid routing mechanism with local-global attention, which can break through the full sequence calculation mode of the traditional attention mechanism, and concentrate the computing resources on the key period (such as the power outage recovery period) through dynamic area division, thereby reducing the computational complexity; 6. The present invention also configures a multi-level memory library consisting of a short-term memory layer, a seasonal memory layer and an event memory layer at the encoder end in the UNet structure, which can reduce the prediction error of emergencies by explicitly modeling the multi-level periodic characteristics of power loads. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings: Figure 1 is a flow chart of Embodiment 1 of the present invention; Figure 2 is a box plot before outlier processing in Example 2 of the present invention; Figure 3 is a box plot after outlier processing in Example 2 of the present invention; Figure 4 is a graph showing changes in power consumption in Example 2 of the present invention; Figure 5 is a maximum temperature variation diagram in Example 2 of the present invention; Figure 6 is the minimum temperature variation diagram in Example 2 of the present invention; Figure 7 is a wind speed variation diagram in Example 2 of the present invention; Figure 8 It is the holiday change diagram in Embodiment 2 of the present invention; Fig. 9 This is a comparison chart of the MSE of various models predicted for the next day in Example 2 of the present invention; Fig.10 This is a comparison chart of the MSE of various models predicted for the next two days in Example 2 of the present invention; Fig.11 This is a comparison chart of the MSE of various models predicted for the next three days in Example 2 of the present invention; Fig.12 This is a comparison chart of the MAE of various models predicted for the next day in Example 2 of the present invention; Fig.13 This is a comparison chart of the MAE of various models predicted for the next two days in Example 2 of the present invention; Fig.14 This is a comparison chart of the MAE of various models predicted for the next three days in Example 2 of the present invention; Fig.15 It is a comparison curve chart of the predicted value and the actual value of the single variable data set predicted for the next day in Example 2 of the present invention; Fig.16 It is a curve chart comparing the predicted value and the actual value of the single variable data set predicted for the next two days in Example 2 of the present invention; Fig.17 It is a comparison curve chart of the predicted value and the actual value of the single variable data set predicted for the next three days in Example 2 of the present invention; Fig.18 is a comparison curve chart of the predicted value and the actual value of the multivariate data set predicted for the next day in Example 2 of the present invention; Fig.19is a comparison curve chart of the predicted values ​​and the actual values ​​of the multivariate data set predicted for the next two days in Example 2 of the present invention; Fig. 20 is a comparison curve chart of the predicted value and the actual value of the multivariate data set predicted for the next three days in Example 2 of the present invention; Fig.21 Schematic diagram of the GLA-UNet model structure in Example 3 of the present invention. DETAILED DESCRIPTION

[0018] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.

[0019] Example 1: Electric load prediction method based on GLA-UNet model structure, such as Figure 1 As shown, the following steps are included: S1: Receives original daily power load data and covariate data, performs anomaly detection and / or filling operations, and outputs standardized time series data; S2: Project the preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate an embedding vector sequence; S3: Perform multi-level downsampling on the embedding vector, each level contains an average pooling layer and a GLA attention unit, and outputs encoding features of different time granularities; S4: Upsample the deepest encoding features step by step, fuse the encoding features of the corresponding layer at each level and optimize them through the GLA attention unit, and output the reconstructed sequence with the same resolution as the input; S5: The time length and feature dimension of the reconstructed sequence are adjusted through two linear layers to generate the power load forecast value for the future period.

[0020] The UNet+BiLSTM architecture recorded in some existing technologies mainly processes spatiotemporal features in serial mode, which cannot resolve the contradiction between local mutations and global trends. In addition, due to the high time complexity caused by the sequential calculation mode of BiLSTM, it is difficult to process long-period data.

[0021] In order to effectively capture the multi-level features in time series data and achieve a balance between accuracy, robustness and practicality in sparse daily power load forecasting, the present invention designs a GLA-UNet model structure. In the present invention, GLA refers to the global-local attention mechanism, and UNet is an encoder-decoder structure commonly used in image segmentation.

[0022] In step S1 , the original daily power load data is the power consumption data with the day as the sampling unit, and the covariate data includes but is not limited to temperature, wind speed, etc.

[0023] In some examples, the missing values ​​and outliers of the data are checked first, and the deletion method is adopted to ensure the data quality for the data with fewer missing and outlier values. In some examples, for the data with more missing and outlier values, the average value of the data of the three points before and after the outlier or missing value is used to fill the data to ensure the stability of the sequence data to the greatest extent.

[0024] In some examples, only the original daily power load data can be input into the model for training to obtain univariate data; the number of feature variables of the univariate data is 1, and the size of the input data of the model is , where T is the input sequence length. In some examples, covariate data can also be added to the data set to implement covariate data training. The covariates introduced by multivariate data are the maximum temperature, the minimum temperature, the wind speed, and whether it is a holiday. In this case, the size of the input data of the model is .

[0025] In step S2, for the input data First, the data is mapped through an embedding layer, which consists of a linear layer, mapping low-dimensional data to high-dimensional data, thereby effectively capturing the interaction between different features and providing richer information support for subsequent predictions.

[0026] In step S3, the size of the input data after mapping is Then, we use a multi-layer feature extraction method, set the stage to 3, use the average pooling layer to extract trend features, and pass through a multi-layer GLA attention unit (hereinafter referred to as the GLA module) after each feature extraction to capture the overall and local dependencies. The GLA module uses a global-local attention mechanism.

[0027] After each average pooling layer, the sequence length changes to: ; in, Indicates Features of layer outputs; Indicates Features of layer outputs; For the The sequence length of the layer output features, is the filling size, is the convolution kernel size, is the stride length.

[0028] In some examples, in order to quickly capture local trend information in the original daily power load data and enable the model to effectively extract multi-scale features at different time scales to capture the impact of extreme events, the present invention proposes a cross-scale correlation attention mechanism as a local attention mechanism, as a local attention mechanism in the GLA module.

[0029] The proposed cross-scale associative attention mechanism has multiple convolutional layers with different kernel sizes and dilation rates to extract trend features at different time scales. By applying these convolutions to the query (Q) and key (K), multiple representations sensitive to different time scales are generated, which are subsequently concatenated and processed by a multi-head attention mechanism. The convolutional layers are designed to capture trends of different durations by adjusting their receptive fields. Larger kernel sizes and higher dilation rates enable the model to consider longer time spans, while smaller kernels and lower dilation rates focus on shorter periods. This multi-scale approach ensures that the attention mechanism can leverage the understanding of the power load dynamics, resulting in more accurate predictions.

[0030] Specifically, for a set of input sequences , where B represents the batch size, T represents the sequence length, and D represents the dimension after embedding. For the Q query matrix and the K key matrix, a set of convolution kernels of different sizes are used. and expansion rate , to extract trend information of different scales. The specific expression is: ; in, Represents the convolution kernel and expansion rate Extracted multi-scale local features; Representation and The corresponding bond matrix characterizes the correlations at different time scales; Indicates the number of convolutional layers; represents the value matrix, Mainly through the linear transformation layer Map the multi-scale features extracted by convolution to dimensions suitable for attention calculation as the value matrix in the attention mechanism; Represents the convolution operation; Represents the padding operation. In order to keep the sequence length unchanged after the convolution layer, Fill with a size of .

[0031] For the extracted trends at multiple time scales as well as Fusion is performed along the sequence length dimension, and the expression is: ; in, express The fusion result of express The fusion result of Represents tensor concatenation; Represents the sequence length dimension.

[0032] After the merger and The size is . Next, we will perform a multi-head split: ; in, Indicates the split The query matrix of the attention heads; Indicates the split The key matrix of the attention heads; Indicates the split The value matrix of the attention heads; Indicates the number of subspaces into which the feature is decomposed; Represents dimension reshaping; is the length of the new key vector, and , The number of lots divided for long positions.

[0033] The attention mechanism operation result is calculated by scaling the dot product, that is: ; in, Indicates The core unit of the attention head is used to capture the multi-scale temporal correlation of power load data through parallel computing; Represents the normalized activation function.

[0034] Finally, the results of all heads are combined and the linear layer is used to output the final attention calculation result. The specific expression is: ; in, Represents the final attention calculation result.

[0035] In some examples, in order to increase the model's ability to capture global features and help the model better understand the long-term trends and periodic patterns of the data, the present invention proposes a multi-head global query attention mechanism as the global attention mechanism in the GLA module. Based on the traditional multi-head attention mechanism, the present invention introduces a global query token to act as an "information convergence point". Similarly, for a set of input sequences , where B represents the batch size, T represents the sequence length, and D represents the embedding dimension.

[0036] First, define a multi-head global query vector , where H is the number of attention heads, and initializes the multi-head query vector : ; in, Initialize the weights.

[0037] The input sequence Reshape , and the global query vector Repeat to match the number of batches and headers: ; in, represents the expanded global query vector; Represents a tensor copy operation.

[0038] at this time, , concatenate the global query vector to the input Get : .

[0039] and , and then perform the calculation of the zoom click attention mechanism: ; in, The final output is obtained by removing the global query vector from the attention mechanism calculation and reshaping it back to its original shape: ; in, Represents the multidimensional tensor after global attention calculation; Represents restoring the original dimensions of the input sequence.

[0040] In some examples, the GLA module also introduces a fusion mechanism to dynamically combine the outputs of the global and local attention layers. This mechanism is implemented through a gating unit, which determines the weight given to the global and local context based on the fully connected layer of the Sigmoid activation function. Specifically, the output of the fusion mechanism can be expressed by the following formula: ; in, is the Sigmoid activation function, is the weight matrix, and and They are the outputs of the global attention layer and the local attention layer respectively; is the output of the fusion mechanism. When the G value is close to 1, the model relies more on the global context; conversely, when the G value is close to 0, the model relies more on the local context. This dynamic weighting mechanism enables the GLA module to flexibly adjust the importance of global and local information according to data characteristics.

[0041] In step S4, the task of the decoder part is to gradually restore the length of the original input sequence based on the feature sequence extracted by the encoder. The present invention adopts an upsampling mechanism based on a fully connected layer and a skip connection technology to ultimately predict the future power load value.

[0042] First, from the deepest layer of the encoder start, After passing through the L-layer GLA module, upsampling is performed. After passing through the L-layer GLA module, the result at this time is recorded as and , also after 3 stages Restore length.

[0043] The length of each stage of recovery With the encoder part The sequence lengths are the same, and the use of skip connection technology allows the decoder to directly access the feature maps of the encoder, thereby retaining important local change information during the prediction process. .

[0044] In step S5, in order to obtain the final prediction result, two linear layers are used to change the sequence length and the number of channels in the prediction result respectively: ; in, is the first linear layer, which is used to adjust the sequence length from T to the desired target prediction length ; is the output of the first linear layer; ; in, is the second linear layer, which is used to convert the number of channels Adjust to desired prediction length ; is the output of the second linear layer. Finally, .

[0045] In some examples, the present invention also introduces dynamic sparse attention in the GLA module, which forms a hybrid routing mechanism with local-global attention to achieve: periodic intensity analysis of the input sequence, automatic division of high volatility areas (dominated by local attention) and stable areas (dominated by global attention); introduce deformable convolution in the decoding stage, and dynamically adjust the convolution kernel sampling position according to the attention heat map. The above method can break through the full sequence calculation mode of the traditional attention mechanism, concentrate computing resources on key periods (such as power outage recovery period) through dynamic area division, and reduce the computational complexity.

[0046] In some examples, the present invention also configures a multi-level memory bank consisting of a short-term memory layer, a seasonal memory layer, and an event memory layer at the end of the encoder in the UNet structure.

[0047] Among them, the short-term memory layer is used to store the load pattern of multiple consecutive days, which is updated using LSTM (Long Short-Term Memory Network); the seasonal memory layer is used to store historical data of the same period by month, which is retrieved by cosine similarity; the event memory layer is used to record the load response pattern of extreme events. The present invention can reduce the prediction error of emergencies by explicitly modeling the multi-level periodic characteristics of power load.

[0048] Example 2: Verification using an office building power load dataset as an example The present invention verifies the effectiveness of the model proposed in the present invention by comparing the mean square error (MSE) and the mean absolute error (MAE) of the model results on univariate and multivariate data sets, and verifies whether the introduced covariates will affect the accuracy of the power load.

[0049] Step 1: In this invention, a power load dataset of an office building in a certain city is used. The dataset takes day as the minimum collection unit and covers daily power load data for two consecutive years from January 1, 2022 to December 31, 2023, that is, there are 730 time nodes of electricity consumption data. The additional covariates introduced are the daily maximum temperature, minimum temperature, wind speed, and whether the day is a holiday. These added covariates are used to verify the effectiveness of the model under a multivariate dataset and whether it will affect the final prediction results. The specific description of the dataset is shown in Table 1.

[0050] Table 1 Constructed multivariate dataset

[0051] Step 2: Screen the data set for missing values ​​and outliers. In this example, there are no missing values ​​in the data set. For outliers, draw a box plot to determine whether there are outliers. Define an outlier as one that is 4 times the standard deviation of the original data. Compare the box plots before and after removing the outliers. Figure 2 , Figure 3 As shown by Figure 2 It can be seen that there are obvious anomalies in the data, with the maximum value being higher than 10,000, but the number of outliers is small. From the figure, it can be seen that there are only two outliers. Therefore, the deletion method is used to directly remove the outliers. The box plot of the data set after removal is as follows: Figure 3 As shown in the figure. Deleting outliers can ensure that model training will not be affected by outliers, thereby ensuring the training effect of the model.

[0052] Step 3: Process the processed data set. First, draw the changes in power consumption, maximum temperature, minimum temperature, wind speed, and whether it is a holiday. Figure 4-Figure 8 As shown. Figure 4 The daily changes in electricity load consumption show certain internal rules, but are also accompanied by significant fluctuations, which shows that the electricity load data with daily as the minimum sampling unit not only reflects obvious periodic and seasonal characteristics, but also reflects the impact of unforeseen events in daily life on electricity demand. The regularity of the maximum and minimum temperature change trends is relatively obvious, as shown in Figure 2. Figure 5 , Figure 6 As shown in the figure, in both years, the peak value was reached around July and the valley value was reached around January. The change in wind speed did not show a more obvious regularity. The monthly consumption change of power load shows the annual periodicity of power load change, reflecting the changing characteristics of power consumption within 12 months of the year. This periodicity is mainly affected by seasonal factors. Different seasons bring different demand patterns for production activities and residents' lives. Taking Shanghai as an example, as a representative city in East China, its annual periodicity of power load is particularly significant. According to field research, Shanghai divides the year into three main stages according to local climate characteristics: heating season (November to March of the following year), cooling season (May to September) and two shorter transition seasons (April and October).

[0053] Step 4: To verify the effectiveness of the introduced multivariate factors, the Pearson correlation coefficient is first calculated to analyze the correlation between the multivariate factors and the power load. The Pearson correlation coefficient can be used to measure the degree of linear correlation between variables. Therefore, the Pearson correlation coefficient can calculate the correlation between the power load and the multivariate variables.

[0054] As shown in Table 2, there is a positive correlation between the maximum temperature, the minimum temperature and the power load, that is, as the temperature increases or decreases, the power demand also increases accordingly. At the same time, wind speed and whether it is a holiday show a negative correlation with the power load, which means that higher wind speeds and non-working days tend to reduce power consumption. However, except for the factor of holidays, the relationship between other variables and power load is not a simple linear association, which shows that relying solely on linear assumptions may not accurately verify the impact of multivariate factors on power load.

[0055] Table 2 Pearson correlation coefficient calculation results

[0056] Step 5: In order to verify the impact of introducing multivariate factors on the prediction effect, the effectiveness of the model proposed in this invention is verified by comparing the performance of multiple models with single variables (considering only historical power load data) and multivariate (combining temperature, wind speed, holidays, etc.), and at the same time, whether multivariate factors have an impact on the prediction performance of the model is verified. The specific data set segmentation and hyperparameter selection are shown below.

[0057] The univariate experiment only uses the power load data of the past 16 days to predict the power load of the next 1, 2, and 3 days. The multivariate experiment uses the power load data of the past 16 days and supplementary covariates to predict the power load of the next 1, 2, and 3 days. The time span of the data set is from January 1, 2022 to December 31, 2023, a total of 730 days. The specific data segmentation is as follows: Training set: 70% of the total data, that is, data from January 1, 2022 to May 26, 2023 are used for model training. This period covers multiple seasonal changes and can fully reflect the annual periodic characteristics of power load.

[0058] Validation set: 10% of the total data, that is, data from May 27, 2023 to July 31, 2023 is used to validate the model.

[0059] Test set: 20% of the total data, that is, data from August 1, 2023 to December 31, 2023 is used for the final model performance evaluation.

[0060] Step 6: The model GLA-UNet proposed in the present invention is compared with several models including Transformer, Informer, Reformer and ITransformer. The hyperparameters of the compared models are shown in Table 3.

[0061] Table 3. Hyperparameters of each model training

[0062] Step 7: In order to produce credible results, each model was trained 10 times and the losses were averaged. When comparing the performance of different prediction time periods (predicting one, two, and three days), the differences in the mean square error (MSE) and mean absolute error (MAE) indicators of each model can be observed.

[0063] As can be seen from Tables 4 and 5, the GLA-UNet model shows the lowest error in the MSE and MAE indicators for all forecast periods in both the univariate and multivariate factors. The prediction results of the univariate data set are 0.141 and 0.318 (forecast for one day), 0.308 and 0.397 (forecast for two days), and 0.425 and 0.457 (forecast for three days), respectively. After the introduction of multivariate factors, they are 0.102 and 0.251 (forecast for one day), 0.194 and 0.318 (forecast for two days), and 0.216 and 0.368 (forecast for three days). This shows that GLA-UNet has significant advantages in the power load forecasting task and can effectively capture the trend of power load changes over time. In contrast, the Reformer model uses LSH Attention to replace the traditional multi-head attention mechanism, but power load forecasting requires accurate capture of long-term dependencies. The approximate nature of LSH may lead to information loss and affect forecast accuracy. Therefore, both in single variables and in each forecast cycle after the introduction of multivariate factors, it shows higher MSE and MAE, especially in the three-day forecast, the single variable forecast MSE reaches 0.982, MAE reaches 0.799, and the multivariate forecast MSE reaches 0.788, MAE reaches 0.700. This shows that the model may face higher uncertainty when dealing with forecasts over longer time periods. Compared with the Reformer model, the Transformer, Informer, and ITransformer models have greatly improved their forecasting performance, but are still inferior to the GLA-UNet model overall.

[0064] Table 4 Test results of each model in the univariate dataset

[0065] Table 5 Test set results of each model in multivariate dataset

[0066] Step 8: Draw the change curves of MSE and MAE of each model under single variable and multivariate conditions. Figure 9-Figure 14 It can be observed that Figure 9-Figure 14Univariate in the model is a single variable, and Multivariate is a multivariate. When predicting the results of the next day, the mean square error (MSE) of the multivariate prediction is reduced by 5.1%, and the mean absolute error (MAE) is reduced by 4.1% compared with the model after introducing multivariate factors. When the prediction range is extended to the next two days, the MSE and MAE of the multivariate prediction are more significantly improved, with a decrease of 12.8% and 9.2%, respectively. When predicting the next three days, this difference is further expanded, and the MSE and MAE of the multivariate prediction are reduced by 24.4% and 15.5%, respectively, compared with the multivariate model. It can be concluded that although the introduction of multivariate factors can indeed improve the prediction performance of the model, its improvement effect on the overall prediction accuracy is relatively limited in the short-term prediction (such as one to two days). With the increase of the prediction time span, the importance of multivariate factors gradually becomes prominent, and it plays a more critical role in improving the prediction accuracy. It shows that in short-term prediction, power load data occupies a relatively dominant position, and the influence of multivariate related factors gradually emerges with the time span. This also provides ideas for actual power load forecasting. For short-term forecasting tasks, it is possible to consider sacrificing a certain degree of forecast accuracy without affecting the quality of decision-making, thereby effectively saving the cost of data collection and processing. When the forecast range is extended to the medium and long term, it is necessary to consider various potential influencing factors more comprehensively. At this time, the introduction of multivariate factors not only helps to improve the forecast accuracy, but also enhances the model's ability to adapt to future uncertainties. Therefore, when conducting medium- and long-term power load forecasting, although obtaining and integrating more dimensional data may incur additional costs, in the long run, this can provide a more accurate and reliable basis for power system planning and operation, ensuring the effective allocation of resources and the stable operation of the system.

[0067] Step 9: In order to more intuitively display the prediction results, select the predicted values ​​of each model in the test set and compare them with the true value curve. The single variable and multivariate prediction results are as follows: Figure 15-Figure 20 shown.

[0068] It can be seen from the figure that the model GLA-UNet proposed in the present invention shows a better fitting effect than other models in single variable and multivariate prediction. In the overall range, it can be seen that the GLA-UNet model has a stronger ability to capture periodicity, and shows the best fitting effect for the peak value of the global range. At the fluctuation point, the proposed model GLA-UNet has the smallest volatility compared with other models in single variable and multivariate prediction, and can relatively accurately reflect the power load trend during this period. After a short period of stability, the power load increased sharply again. It can be clearly seen that the GLA-UNet model has the strongest ability to capture this emergency. It not only responds quickly to the growth of load, but also maintains a certain degree of accuracy during the prediction process. In the same number of prediction days, the multivariate prediction results of each model are significantly better than the single variable prediction results, especially in the overall periodic capture. The single variable prediction has a weaker ability to capture the peak value in the cycle. As the prediction length increases, the prediction results of each model can only capture the overall periodic changes, but the gap between the predicted value and the true value gradually increases. At the fluctuation point, the volatility of the univariate prediction results of each model is higher than that of the multivariate prediction results. Therefore, the univariate prediction has relatively poor ability to capture the local characteristics of non-periodic fluctuations.

[0069] Embodiment 3: An electric load prediction system based on a GLA-UNet model structure, the system is used to implement the electric load prediction method based on a GLA-UNet model structure as described in Embodiment 1, Fig.21 As shown, it includes a data preprocessing module, a feature embedding module, a multi-scale encoding module, a context decoding module and a prediction output module.

[0070] Among them, the data preprocessing module is used to receive the original daily power load data and covariate data, perform anomaly detection and / or filling operations, and output standardized time series data; the feature embedding module is used to project the preprocessed univariate or multivariate data to the high-dimensional feature space through a linear mapping layer to generate an embedded vector sequence; the multi-scale encoding module is used to perform multi-level downsampling of the embedded vector, each level includes an average pooling layer and a GLA attention unit, and outputs encoding features of different time granularities; the context decoding module is used to upsample the deepest encoding features step by step, each level fuses the encoding features of the corresponding level and optimizes them through the GLA attention unit, and outputs a reconstructed sequence with the same resolution as the input; the prediction output module is used to adjust the time length and feature dimension of the reconstructed sequence through two linear layers respectively, and generate the power load forecast value for the future period.

[0071] Working principle: This paper establishes a GLA-UNet model structure by improving the traditional UNet model, making it more suitable for the field of power load forecasting. It can effectively capture the multi-level features in time series data and achieve a balance between accuracy, robustness and practicality in sparse daily power load forecasting, especially periodic and seasonal patterns.

[0072] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0073] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0074] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0075] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0076] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. The electric load forecasting method based on the GLA-UNet model structure is characterized by: The following steps are involved: Receive original daily power load data and covariate data, perform anomaly detection and / or filling operations, and output standardized time series data; Project the preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate an embedding vector sequence; The embedding vector is downsampled at multiple levels, each level contains an average pooling layer and a GLA attention unit, and outputs encoding features of different time granularities; The deepest encoding features are upsampled level by level, and the encoding features of the corresponding level are fused at each level and optimized through the GLA attention unit to output a reconstructed sequence with the same resolution as the input; The time length and feature dimension of the reconstructed sequence are adjusted respectively through two linear layers to generate the power load forecast value for the future period.

2. The electric load forecasting method based on the GLA-UNet model structure according to claim 1 is characterized in that: The performing of anomaly detection and / or filling operations includes any one or more of the following: Delete data segments where the number of missing values ​​or outliers is less than 5%; For continuous abnormal areas, sliding window average filling is used, and the window size is 3 time points before and after; The temperature and wind speed covariates are normalized, and the holiday variables are one-hot encoded.

3. The electric load forecasting method based on the GLA-UNet model structure according to claim 1 is characterized in that: The GLA attention unit performs the following operations: Use multiple atrous convolutional layers with different dilation rates to extract multi-scale local features in parallel; Capturing long-term periodic global features via learnable global token vectors The gating coefficients are automatically generated according to the input features, and the multi-scale local features and long-term periodic global features are weighted and fused according to the gating coefficients.

4. The electric load forecasting method based on the GLA-UNet model structure according to claim 3 is characterized in that: The generation expression of the gating coefficient is: ; in, represents the gating coefficient; Represents the Sigmoid activation function; is the weight matrix; and are the outputs of the global attention layer and the local attention layer respectively.

5. The electric load forecasting method based on the GLA-UNet model structure according to claim 3 is characterized in that: The GLA attention unit also performs the following operations by introducing dynamic sparse attention: Perform periodic strength analysis on the input sequence to separate high-volatility regions through local attention mechanisms and stable regions through global attention mechanisms; Also, deformable convolution is introduced in the decoding stage to dynamically adjust the convolution kernel sampling position according to the attention heat map.

6. The electric load forecasting method based on the GLA-UNet model structure according to claim 1 is characterized in that: The multi-level downsampling adopts three-level downsampling, and the sampling ratios are 1 / 2, 1 / 2, and 1 / 2 respectively, and the sequence length of the output coding feature is 1 / 8 of the input length.

7. The electric load forecasting method based on the GLA-UNet model structure according to claim 1 is characterized in that: The upsampling operation adopts a bilinear interpolation algorithm, and after interpolation, the number of channels is adjusted through a 1×1 convolution layer.

8. The electric load forecasting method based on the GLA-UNet model structure according to claim 1 is characterized in that: The two linear layers are used to adjust the time length and feature dimension of the reconstructed sequence respectively, including: The sequence length is expanded to the target prediction length through the first linear layer; The feature dimensions are compressed to the number of output variables through the second linear layer.

9. The electric load forecasting method based on the GLA-UNet model structure according to claim 1 is characterized in that: The method further includes: The encoder end in the UNet structure is configured with a multi-level memory bank consisting of a short-term memory layer, a seasonal memory layer, and an event memory layer; The short-term memory layer is used to store the load patterns for consecutive days and is updated using LSTM; Seasonal memory layer, used to store historical data by month and retrieve it by cosine similarity; Event memory layer, used to record load response patterns to extreme events.

10. The electric load forecasting system based on the GLA-UNet model structure is characterized by: include: A data preprocessing module, used to receive original daily power load data and covariate data, perform anomaly detection and / or filling operations, and output standardized time series data; Feature embedding module, which is used to project the preprocessed univariate or multivariate data into a high-dimensional feature space through a linear mapping layer to generate an embedding vector sequence; The multi-scale encoding module is used to perform multi-level downsampling of the embedding vector. Each level contains an average pooling layer and a GLA attention unit to output encoding features of different time granularities. The context decoding module is used to upsample the deepest encoding features step by step, fuse the encoding features of the corresponding layer at each level and optimize them through the GLA attention unit, and output a reconstructed sequence with the same resolution as the input; The prediction output module is used to adjust the time length and feature dimension of the reconstruction sequence through two linear layers to generate the power load prediction value for the future period.

Citation Information

Patent Citations

  • Building extraction model training method, extraction method, equipment and storage medium

    CN115439756A

  • Power load prediction method based on multivariable time sequence information interaction

    CN118333232A

  • Non-intrusive load monitoring method based on improved UNet architecture

    CN118364426A

  • Power load prediction method and system

    CN119026763A

  • CNN-BiGRU-Attention-based power load interval prediction method

    CN119558442A

Cited By

  • Sewage plant water quality prediction and management operation method, computer program product and electronic product

    CN120877962A