Power load data prediction method and device, electronic equipment and storage medium
Through the power load data prediction method based on the Transformer model, the target model constructed by the window encoder solves the problems of low calculation efficiency of power load prediction and insufficient generalization capability in the existing technology, and achieves more accurate and efficient power load prediction.
Patent Information
- Application Number
- CN202411936981.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
AI Technical Summary
When the existing power load prediction methods process large-scale and multi-dimensional data, they have low computing efficiency and insufficient model generalization capabilities, making it difficult to accurately predict short-term power loads, especially in areas with large coverage of distributed photovoltaics.
The power load data prediction method based on the Transformer model is adopted, including obtaining distributed photovoltaic power load data, preprocessing data, extracting load sequences, and predicting through the target model constructed by the window encoder.
Through the improved Transformer model and window encoder, the prediction accuracy and computing efficiency of power load data can be significantly improved, and is suitable for areas with large distributed photovoltaic coverage.
Smart Images

Figure CN120067926A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of electric load forecasting, and particularly to a method, device, electronic device and storage medium for predicting electric load data. Background Art
[0002] The electric power industry is an important basic industry in the country's development, which is related to the country's development in various fields. Therefore, whether short-term electric load forecasting can be accurately carried out is the key to the reliable operation of the power system. Summary of the Invention
[0003] In view of this, the purpose of this application is to propose a method, device, electronic device and storage medium for predicting electric load data.
[0004] Based on the above purpose, in the first aspect of this application, a method for predicting electric load data is provided, including:
[0005] Obtain distributed photovoltaic power load data;
[0006] Preprocess the distributed photovoltaic power load data to obtain target power load data;
[0007] Obtain a load sequence according to the target power load data;
[0008] Predict electric load data based on the load sequence using a target model, where the target model is constructed based on the Transformer model and includes a window encoder.
[0009] In the second aspect of this application, a device for predicting electric load is provided, including:
[0010] An obtaining module configured to obtain distributed photovoltaic power load data;
[0011] A preprocessing module configured to preprocess the distributed photovoltaic power load data to obtain target power load data;
[0012] A time series module configured to obtain a load sequence according to the target power load data;
[0013] A prediction module configured to predict electric load data based on the load sequence using a target model, where the target model is constructed based on the Transformer model and includes a window encoder.
[0014] In the third aspect of this application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method described in the first aspect.
[0015] In the fourth aspect of the present application, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method described in the first aspect.
[0016] As can be seen from the above, the present application provides a method, apparatus, electronic device, and storage medium for predicting power load data. The method includes obtaining distributed photovoltaic power load data, preprocessing the distributed photovoltaic power load data to obtain target power load data, obtaining a load sequence based on the target power load data, and predicting power load data based on the load sequence using a target model. The target model is constructed based on the Transformer model and includes a window encoder. By using the target model based on the window encoder, a more accurate prediction result of power load data can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present application or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 FIG. shows a schematic diagram of an exemplary unsupervised anomaly data monitoring method according to an embodiment of the present application.
[0019] Figure 2 FIG. shows a schematic diagram of the structure of an exemplary model according to an embodiment of the present application.
[0020] Figure 3 FIG. shows a schematic flowchart of an exemplary power load data prediction method according to an embodiment of the present application.
[0021] Figure 4 FIG. shows a schematic diagram of an exemplary power load data prediction apparatus according to an embodiment of the present application.
[0022] Figure 5 FIG. shows a schematic diagram of an exemplary electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] To make the objectives, technical solutions, and advantages of the present application more clear and understandable, the following further describes the present application in detail with reference to specific embodiments and the accompanying drawings.
[0024] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of this application should have the ordinary meanings understood by those with ordinary skills in the field to which this application belongs. The "first", "second" and similar terms used in the embodiments of this application do not indicate any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0025] The power industry is an important basic industry in the country's development, which is related to the country's development in various fields. Therefore, whether short-term power load forecasting can be accurately carried out is the key to the reliable operation of the power system.
[0026] Traditional power load forecasting methods mainly include time series forecasting models, regression analysis models, and exponential analysis models, etc. Among them, the time series analysis model requires less power load data volume, but the weights of external weather factors such as temperature and humidity are considered relatively low in the model, and the forecasting accuracy for power load data with high complexity and strong volatility is limited; although the regression analysis method can be applied to multi-dimensional variable models, its model is relatively simple and assumes that there is a certain linear correlation in the data, which does not conform to the characteristics of power data. Considering that there are many factors affecting power load, among which weather factors such as temperature, rainfall, and air humidity have a great impact on power load, and in recent years, as the coverage of distributed photovoltaic has become deeper and deeper, especially the proportion of low-voltage distributed photovoltaic on the distribution network side is relatively high, resulting in a greater impact of light and other factors on the regional power load, etc. Therefore, whether these factors can be accurately taken into account is the key to improving the forecasting accuracy. Although the exponential analysis model has certain advantages in dealing with non-linear relationships, it has a strong dependence on historical data, and when facing emergencies, the adaptability and flexibility of the model are insufficient.
[0027] With the development of artificial intelligence technology, machine learning and deep learning methods have begun to be applied in the field of power load forecasting. These methods can automatically identify and extract key factors affecting power load by learning a large amount of historical data, thereby improving the forecasting accuracy. For example, the neural network model can capture the complex relationship between power load and various influencing factors through its non-linear mapping ability, and then make a more accurate prediction of power load.
[0028] Existing power load forecasting methods still face problems such as low computational efficiency and insufficient model generalization ability when dealing with large-scale and multi-dimensional data. Therefore, researching a new power load forecasting method that can effectively integrate multiple forecasting models, improve forecasting accuracy and computational efficiency simultaneously, and consider the distributed photovoltaic coverage factor is of great significance for the stable operation of the power system.
[0029] Existing short-term load forecasting methods are sensitive to parameter settings and initial solutions. In some cases, different parameter settings and initial solutions may lead to significantly different results, resulting in unsatisfactory forecasting effects. Moreover, the time complexity is high, the operation efficiency is low, and the forecasting effect for medium- and long-term power loads is poor.
[0030] In addition, short-term power load forecasting may require a large amount of training data and parameter tuning work. Facing massive data, the extracted features cannot accurately represent the typical electricity consumption patterns of different categories, and a large amount of computing resources are required. At the same time, in the face of a large training dataset, how to efficiently extract features has become a core issue. Current feature extraction techniques may not be able to accurately capture the typical features of different categories of electricity consumption patterns, which may lead to deviations in forecasting results.
[0031] To at least solve the above problems, the present application provides a power load data forecasting method, device, electronic device, and storage medium. The method includes obtaining distributed photovoltaic power load data, preprocessing the distributed photovoltaic power load data to obtain target power load data, obtaining a load sequence according to the target power load data, and predicting power load data based on a target model according to the load sequence. The target model is constructed based on the Transformer model and includes a window encoder. Through the target model based on the window encoder, a more accurate forecasting result of power load data can be obtained.
[0032] Power system load forecasting is a key link to ensure the safe and economic operation of the power system, and it has a crucial impact on the reasonable arrangement of various aspects such as power generation, power transmission, and power distribution. With the introduction of the power market competition mechanism, higher requirements are put forward for the accuracy and speed of short-term load forecasting to meet the needs of market competition. The research on load forecasting has a history of decades, and there are many theories and methods, covering various different technologies and algorithms. With the continuous progress of science and technology, the research on new load forecasting methods is also constantly deepening.
[0033] As a data mining technique, the time series prediction method has shown extraordinary potential in dealing with non-linear, high-dimensional problems and generalization ability, and is widely used in many fields such as pattern recognition and regression problem processing. Compared with traditional neural network and machine learning methods, the time series prediction method has significant advantages in some aspects. For example, it can better capture the complex patterns and dynamic changes in the data, thus providing more accurate prediction results.
[0034] First, preprocess the collected power load data. Due to reasons such as collection device failures and negligence of collection personnel during work, the power load data collected in real life often has problems such as data missing, data anomalies, and data duplication. And the quality of the data directly affects the effect of subsequent data analysis. Therefore, it is necessary to improve the data quality through methods such as patching or removal. The embodiments of this application perform some preprocessing work on the data, process the problems such as data missing, anomalies, and duplications existing in the original data according to certain rules, and perform normalization processing on the data to obtain the target power load data, providing a more reliable data set for subsequent analysis and research.
[0035] In some embodiments, a data decomposition algorithm can be used to decompose the collected original data. Empirical Mode Decomposition (EMD) decomposes the data according to its own time characteristics. The Intrinsic Mode Function (IMF) components decomposed contain different characteristics of different frequency patterns of the original sequence data. The role of EMD is to obtain stationary components after decomposing non-stationary data.
[0036] In this process, the phenomenon of mode mixing will occur. In view of this, in some embodiments, the Complete Ensemble Empirical Mode Decompostion with Adaptive Noise (CEEMDAN) algorithm can be used for data decomposition. The CEEMDAN algorithm is developed from the EMD method and effectively avoids the phenomenon of mode mixing. The decomposition steps of the CEEMDAN algorithm are divided into three parts:
[0037] The first part is to add N times of adaptive Gaussian white noise to the original load data y to obtain the load data y with added noise, perform EMD decomposition on the new load data (for example, the first power load data), and perform ensemble averaging on the obtained N first mode components N , to obtain the first intrinsic mode component IMF 1, then calculate the residual after subtracting the first eigenmode component. The first eigenmode component and the first residual r 1 are respectively shown in the following formulas:
[0038]
[0039] r 1 = y - IMF 1
[0040] Second part, after adding N times of adaptive Gaussian white noise to the first residual r 1 , the target residual is obtained, and the target residual is decomposed into N second mode components Continue to solve to obtain the second eigenmode component IMF 2 and the second residual r 2 are respectively shown in the following formulas:
[0041]
[0042] r 2 = r 1 - IMF 2
[0043] Third part, repeat the above steps to obtain k eigenmode components IMF i and the final third residual r k , the preprocessed original load data y (for example, the target power load data) can be expressed as:
[0044]
[0045] Compared with the EMD algorithm, the CEEMDN algorithm can effectively reduce the reconstruction error, has a faster calculation speed, and can reduce the number of low-frequency IMF components with very small amplitudes. Therefore, the embodiments of this application select this algorithm as the method for decomposing non-stationary sequences.
[0046] In order to further improve the overall prediction accuracy of the model, in the distributed photovoltaic power load prediction scenario of the embodiments of this application, an outlier detection method is also introduced. Through this iterative optimization process, the model can continuously adjust and optimize the prediction results, thereby significantly improving the accuracy and reliability of short-term power load prediction.
[0047] In the process of distributed photovoltaic power load forecasting, the training of the forecasting model is based on historical data. If the load data is directly input into the forecasting model, it is easy to cause deviation in the prediction of future data by the forecasting model, thereby reducing the accuracy. At the same time, it will also bring the consequences of disturbing the forecasting model and affecting the generalization ability. To improve the accuracy of the subsequent forecasting model, it is necessary to detect and process the abnormal load data to reduce the impact of abnormal data. According to whether the data is accompanied by abnormal labels, anomaly detection techniques can be subdivided into supervised, semi-supervised, and unsupervised anomaly detection. In the field of supervised anomaly detection, this technique is similar to supervised machine learning classification or regression algorithms in the context of data imbalance. For data anomalies, the error value between the actual output and the actual anomaly label is used as the loss function for backpropagation, and after multiple iterations, an optimal anomaly detection model is obtained. Unsupervised detection mainly focuses on revealing the similarities and differences between samples in the dataset, identifying abnormal data different from most data, without the need to know the data anomaly situation in advance. The unsupervised model directly analyzes and processes the data to learn the internal characteristics of the data, and optimizes the model through the training of unlabeled data. Supervised learning and unsupervised learning are not completely separated. Semi-supervised anomaly detection methods are in between, suitable for the situation of a small amount of labeled abnormal data and a large amount of unlabeled normal data. The model is trained by combining the labeled abnormal data and the unlabeled normal data to learn the pattern of abnormal data.
[0048] Due to its many advantages, unsupervised anomaly detection technology has wide applicability in a variety of tasks and data types, especially in the case where abnormal data is difficult to obtain and label, its advantages are more prominent. For example, in the fields of financial fraud detection, network security monitoring, medical diagnosis, etc., unsupervised detection technology can effectively identify abnormal behaviors or abnormal states, and can also perform effective anomaly detection even in the absence of sufficient labeled data. In addition, unsupervised detection technology also shows its high efficiency in processing large-scale datasets because they do not require cumbersome preprocessing or labeling of data.
[0049] Figure 1 The figure shows a schematic diagram of an exemplary unsupervised abnormal data monitoring method according to an embodiment of the present application.
[0050] As Figure 1As shown, common time-series unsupervised anomaly detection algorithms mainly include statistical outlier detection (e.g., Z-score box plot, Grubbs test, and non-parametric regression), density-based outlier detection (e.g., local outlier factor and Gaussian mixture model), clustering-based outlier detection (e.g., K-means algorithm and DBSCAN algorithm), distance-based outlier detection (e.g., K-nearest neighbor algorithm), and other methods, such as machine learning-based outlier detection (e.g., isolation forest, support vector machine, and autoencoder), etc. These techniques each have their own characteristics and applicable scenarios. Statistical-based techniques rely on the statistical characteristics of data, such as mean and standard deviation, to identify outliers; density-based techniques focus on the local density of data points, and outlier points are usually located in low-density regions; clustering-based techniques identify outlier points that do not belong to any cluster by grouping data; while machine learning-based techniques use complex algorithm models to learn the normal behavior patterns of data, thereby identifying abnormal behaviors that deviate from these patterns.
[0051] In some embodiments, the distributed photovoltaic power load data can be subjected to anomaly detection to obtain target data, an anomaly data label is acquired, and it is determined whether the target data is anomaly data according to the error value between the target data and the anomaly data label. In response to the error value being greater than or equal to a threshold, the detected target data is determined to be anomaly data. The anomaly data is processed (e.g., removed or modified, etc.) to obtain target power load data.
[0052] After preprocessing the data, time series prediction is performed on the processed data. A time series records a certain phenomenon or indicator at different time points at the time level. Time series analysis is a method widely used in different fields. According to the number of variables of the research object, it is divided into univariate time series and multivariate time series. The expression of the multivariate time series X is as follows:
[0053] X = {x 1 , x 2 , …, x n}
[0054] where n represents the length of the sequence, representing an M-dimensional variable at this time point. The variables in the multivariate time series x 1 , x 2 , …, x n contain more information and are correlated, which can reflect the relationships and interactions between different variables. Similarly, a more complex model often requires more advanced statistical methods.
[0055] It can be divided into discrete type and continuous type according to the data type. The former has data values only at finite time points, while the latter has data values between any two time points. According to the data variability, it can be divided into stable type and non-stable type. For a stable time series, statistical indicators such as the mean, variance, and autocorrelation coefficient should be similar in different time periods. The non-stable type often shows obvious trends and periodic patterns. In actual production and life, the data generally has weak stability. In some embodiments, the mean and covariance can be calculated by the following formulas:
[0056]
[0057] γ(s,t) = cov(x s ,x t ) = E[(x s -μ s )[(x t -μ t )]]
[0058] Among them, μ and E(x t ) represent the mean, f t (t) represents the time series data function, γ(s,t) represents the covariance between the time series x s and the time series x t , E[(x s -μ s )[(x t -μ t )]] represents the covariance between the time series x s and the time series x t , μ s and μ t respectively represent the means of the time series x s and the time series x t .
[0059] The components of a time series refer to the factors or components that affect the changes in the time series, including trends, seasonal variations, cyclic periods, and random fluctuations. Trends represent the overall change level presented by the time series in the long term; seasonal variations represent the periodic changes presented in the short term; cyclic periods represent the periodic changes presented by the time series in the long term, usually longer than seasonal variations; random fluctuations are used to represent the impact of unpredictable external factors on the time series. Dividing the time series Y(t) according to the components, the expression of the model is as follows:
[0060] Y(t) = T(t) + S(t) + C(t) + I(t)
[0061] Among them, T(t) represents the trend, S(t) represents the seasonal variation, C(t) represents the cycle period, and I(t) represents the random fluctuation. To further improve the accuracy, the time series can be divided into linear and non-linear parts, and the corresponding expressions are as follows:
[0062] Y(t) = l(t) + N(t)
[0063] Among them, L(t) represents the linear element, and N(t) represents the non-linear element.
[0064] In some embodiments, descriptive statistical methods can be used to summarize the data characteristics in the time series. For the basic statistical description of the time series, describe the relationship and correlation degree between variables, such as mean, variance, correlation coefficient, etc. Common correlation coefficients are as shown in the following formula:
[0065]
[0066] Among them, γ Pearson represents the Pearson correlation coefficient, ρ Spearman represents the Spearman correlation coefficient, Γ Kendall represents the Kendall correlation coefficient, n represents the number of sequence elements, X i and Y i respectively represent the data points of the two variables X and Y at the i-th observation value, and respectively represent the average values of the two variables X and Y, and a and b respectively represent the number of consistent and inconsistent element pairs.
[0067] The statistical quantities of non-stationary time series change with time, which is not conducive to establishing a reliable time series model. In view of this, in some embodiments, the data in the time series can then be smoothed. Common methods include the difference method, logarithmic transformation, exponential smoothing method, etc. The formula for single exponential smoothing is as follows:
[0068]
[0069] Among them, represents the predicted value at time t + 1, x t is the actual value at time t, s t (1) 、s t-1 (1) are the smoothed values at times t and t - 1, and α is the smoothing coefficient, representing the influence degree of historical data.
[0070] Time series model parameter estimation refers to the process of determining the values of various parameters in the time series model by analyzing and fitting historical data during the time series modeling process. Attention should be paid to the problems of overfitting and underfitting of the model to improve the prediction ability and stability of the model. Common methods for estimating time series model parameters include maximum likelihood estimation, least squares estimation, and Bayesian estimation.
[0071] In the embodiments of this application, an improved Transformer model is used for prediction. At the same time, the model does not directly use the complete load sequence as the training input, but inputs the time slices before the current time stamp of the sequence. Considering the daily cycle characteristics of the load sequence, the sequence is converted into a sliding window, and when the window length is insufficient, it is filled by replication. At the same time, an efficient adversarial training method is added to amplify the deviation, and the model consists of two encoders and two decoders.
[0072] Figure 2 The structural schematic diagram of an exemplary model according to the embodiments of this application is shown.
[0073] The improved Transformer structure in the embodiments of this application is as Figure 2 shown. The model may include an encoder, a window encoder, and two decoders (for example, a first decoder and a second decoder). The encoder may further include a first multi-head attention mechanism, a first layer of normalization, a target feed-forward layer, and a second layer of normalization. The window encoder may further include a masked multi-head attention mechanism, a third layer of normalization, a second multi-head attention mechanism, and a fourth layer of normalization. The first decoder may further include a first feed-forward layer and a first activation layer, and the second decoder may further include a second feed-forward layer and a second activation layer. Among them, C represents the complete load sequence, W represents the load sliding window sequence, O 1 and O 2 respectively represent the reconstruction outputs of the two decoders (for example, the first reconstruction output and the second reconstruction output).
[0074] The encoder acts on the complete load sequence, and the window encoder acts on the sliding window sequence with context information. At the same time, mask processing is performed on the multi-head attention mechanism in the window encoder to shield the data at subsequent positions to prevent the decoder from viewing future data during training. In some embodiments, the load sequence can be input into the first multi-head attention mechanism to calculate the correlation between different positions in the load sequence to obtain a first target load sequence, where the first target load sequence is an attention-weighted load sequence. The first target load sequence is input into the first layer of normalization for normalization processing to obtain a second target load sequence, where the second target load sequence is a normalized attention-weighted load sequence. The second target load sequence is input into the target feed-forward layer for non-linear transformation to obtain a first vector, and the first vector is input into the second layer of normalization for normalization processing to obtain a second vector. The power load data is predicted based on the second vector.
[0075] The output of the encoder is used as V and K by the window encoder for the attention operation using the encoded input window as Q, to better utilize the context information of the load sequence. Among them, V represents the value vector, which is used to store information related to the input sequence. Each input position has a corresponding value vector, and these value vectors are weighted and summed in the attention mechanism to generate the output; K represents the key vector, which is used to store information related to the input sequence. The key vector is used to calculate the similarity with the query vector to determine the contribution degree of each input element to the output; Q represents the query vector, which is used to extract information related to the input sequence. The query vector is compared with all key vectors to obtain the relevant value vectors. The query vector, key vector, and value vector are obtained through linear transformation from the original input vector (usually the word embedding representation).
[0076] In some embodiments, according to the load sequence, a load sliding window sequence can be obtained, and the load sliding window sequence is input into the masked multi-head attention mechanism to obtain a first target load sliding window sequence, where the first target load sliding window sequence is an attention-weighted load sliding window sequence. The first target load sliding window sequence is input into the third layer of normalization for normalization processing to obtain a second target load sliding window sequence. The second target load sliding window sequence and the second vector are input into the second multi-head attention mechanism to obtain a third vector, where the second vector includes the value vector and the key vector, and the third vector is the attention-weighted second vector. The third vector is input into the fourth layer of normalization to obtain a fourth vector, and the power load data is predicted based on the fourth vector.
[0077] The above operation uses the input time series window and the complete sequence to generate attention weights to capture the time trend within the input sequence, enabling the model to perform parallel operations on multiple batches of time series windows, thus significantly improving the training time of the proposed method. The decoder is the same, and its role is to output the reconstruction of the load. The model quickly infers the broader time trend in the data, mainly divided into the following two stages.
[0078] The first stage is input reconstruction, generating an approximate reconstruction of the input window. The encoder generates attention weights to extract features from the input and capture the time trend in the sequence. The masked multi-head attention layer replaces the sum and subsequent parts with larger negative numbers, and then the weights become 0, ensuring that the prediction at the current position only depends on the known outputs at previous positions. The encoder output is window-encoded and used as V and K for the attention operation using the encoded input window as Q. The reconstruction output of the input is completed through the decoder. The goal of the first stage is to minimize the reconstruction error. The loss functions of the two decoders are shown as follows:
[0079] L 1 = ‖O 1 - W‖ 2
[0080] L 2 = ‖O 2 - W‖ 2
[0081] Among them, O 1 and O 2 respectively represent the reconstruction outputs of the two decoders, and W represents the load window sequence.
[0082] The second stage utilizes the deviation generated in the first stage, that is, the difference between the predicted value and the true value. This deviation serves as a prior for modifying the attention weights in the second stage, helping the attention layer inside the encoder to extract the time trend, focusing on the subsequences with higher deviations, assigning higher activation degrees to the inputs with larger deviations, amplifying the deviations, and obtaining the final predicted output. The second stage adopts an output adversarial loss. The first encoder perfectly reconstructs the input load window sequence, and the second encoder maximizes the reconstruction difference. The corresponding loss function in this stage is shown as follows:
[0083]
[0084] Among them, is the output reconstruction of the second decoder in the second stage. The adversarial training method in the two stages can amplify the deviation, and the reconstruction error plays an activating role in the attention mechanism part of the encoder. At the same time, it can capture the short-term time trend of the window encoder. represents the minimum value of the reconstruction difference, represents the maximum value of the reconstruction difference.
[0085] In some embodiments, the fourth vector can be input into the first feed-forward layer and the second feed-forward layer respectively for non-linear transformation to obtain a fifth vector and a sixth vector respectively. The fifth vector and the sixth vector are input into the first activation layer and the second activation layer respectively to obtain a first reconstruction output and a second reconstruction output respectively. The predicted power load data is obtained based on the first reconstruction output and the second reconstruction output.
[0086] In the embodiments of the present application, the empirical mode decomposition (EMD) algorithm and the unsupervised anomaly data detection method are first used to preprocess the massive load data, enabling the data to be visually processed, improving the discrimination between normal data and abnormal data, effectively avoiding the disturbance of abnormal data to the prediction model, thereby reducing the prediction accuracy, and ensuring the quality of the power load data. The improved Transformer model is adopted to solve the problem of long time series processing, better utilize the context relationship of the data, and an efficient two-stage adversarial training method is added. Two decoders are used to reconstruct the output. In the second stage, the error generated in the first stage is amplified, and a higher activation degree is given to the input with a larger deviation to amplify the deviation, obtaining a more accurate preliminary prediction result.
[0087] In some embodiments, it can be considered to use an improved long short-term memory (LSTM) algorithm for power load data prediction. This algorithm fine-tunes and optimizes the hyperparameters of the LSTM neural network by introducing the grey wolf optimization algorithm or the particle swarm optimization algorithm. This approach helps to avoid setting parameters solely based on experience, thereby reducing errors caused by subjective judgment. The LSTM neural network is particularly suitable for dealing with power load prediction problems because it can fully consider the existence of numerous influencing factors in the power system. These factors constitute high-dimensional data, and the LSTM network performs well in processing such data, effectively improving the prediction accuracy and performance.
[0088] Here, in combination with specific embodiments, a power load prediction method based on an improved Transformer model and considering the distributed photovoltaic coverage is described for the present invention.
[0089] The embodiments of the present application select the distributed photovoltaic power load data as the experimental object. Given that the distributed photovoltaic power load data exhibits significant time-varying, uncertain, and diverse characteristics, and is significantly affected by factors such as weather conditions, seasonal changes, temperature fluctuations, and user behavior patterns, the data presents non-regular characteristics. Therefore, the embodiments of the present application use a method combining anomaly data detection and complete empirical mode decomposition (CEEMD) to preprocess the data.
[0090] First, the embodiments of the present application apply unsupervised anomaly detection technology to preliminarily process power load data to reduce the potential interference of data outliers on the data set. Subsequently, the complete empirical mode decomposition algorithm is used to decompose the sequence load data processed by anomaly detection, and it is split into several stable intrinsic mode components and residual function components.
[0091] An improved Transformer model is used to predict each preprocessed component. The specific parameters of the Transformer are as follows: sequence input length 672; sequence output length 168; window size 24; number of training times 128; regularization coefficient 0.1; number of iterations 500; data sampling interval 1 hour, time span from February 1, 2010 to February 28, 2010. The training set and the test set are divided according to a ratio of 3:1, and the loss function used is:
[0092]
[0093] When conducting data analysis, we first need to carefully compare and analyze the test results with the actual observed values. Through this comparison, we can calculate the prediction accuracy. To more comprehensively evaluate the prediction effect, we usually compare the adopted prediction model with several commonly used prediction models. In the comparison process, we adopt a series of standardized model evaluation indicators, including root mean square error (RMSE), mean absolute percentage error (MAPE), and mean absolute error (MAE). Through the comprehensive consideration of these indicators, we can more accurately evaluate the prediction performance of the model.
[0094] Figure 3 shows a schematic flowchart of an exemplary power load data prediction method according to an embodiment of the present application. As Figure 3 shown, the method may include the following steps.
[0095] In step 302, distributed photovoltaic power load data is obtained.
[0096] In step 304, the distributed photovoltaic power load data is preprocessed to obtain target power load data.
[0097] In some embodiments, the preprocessing of the distributed photovoltaic power load data to obtain target power load data further includes: adding adaptive Gaussian white noise to the distributed photovoltaic power load data to obtain first power load data; decomposing the first power load data to obtain multiple first modal components; performing ensemble averaging on the multiple first modal components to obtain the first intrinsic mode component; and obtaining the target power load data according to the distributed photovoltaic power load data and the first intrinsic mode component.
[0098] In some embodiments, obtaining the target power load data according to the distributed photovoltaic power load data and the first intrinsic mode component further includes: obtaining a residual according to the difference between the distributed photovoltaic power load data and the first intrinsic mode component; adding adaptive Gaussian white noise to the residual to obtain a target residual; decomposing the target residual to obtain a plurality of second mode components; performing ensemble averaging on the plurality of second mode components to obtain a second intrinsic mode component; and obtaining the target power load data according to the residual and the second intrinsic mode component.
[0099] In some embodiments, preprocessing the distributed photovoltaic power load data to obtain target power load data further includes: performing anomaly detection on the distributed photovoltaic power load data to obtain target data; obtaining an anomaly data label, and determining whether the target data is anomaly data according to an error value between the target data and the anomaly data label; in response to the error value being greater than or equal to a threshold, determining that the target data is anomaly data; and processing the anomaly data to obtain the target power load data.
[0100] In step 306, a load sequence is obtained according to the target power load data.
[0101] In step 308, based on the load sequence, power load data is predicted according to a target model, the target model is constructed based on a Transformer model, and the target model includes a window encoder.
[0102] In some embodiments, the target model includes an encoder, and the encoder includes a first multi-head attention mechanism, a first layer of normalization, a target feed-forward layer, and a second layer of normalization. Predicting the power load data based on the target model according to the load sequence further includes: inputting the load sequence into the first multi-head attention mechanism to calculate the correlation between different positions in the load sequence to obtain a first target load sequence, where the first target load sequence is an attention-weighted load sequence; inputting the first target load sequence into the first layer of normalization for normalization processing to obtain a second target load sequence, where the second target load sequence is the normalized attention-weighted load sequence; inputting the second target load sequence into the target feed-forward layer for non-linear transformation to obtain a first vector; inputting the first vector into the second layer of normalization for normalization processing to obtain a second vector; and predicting the power load data according to the second vector.
[0103] In some embodiments, the target model further includes a window encoder, and the window encoder includes a masked multi-head attention mechanism, a third layer normalization, a second multi-head attention mechanism, and a fourth layer normalization. The predicting of the power load data according to the second value vector further includes: obtaining a load sliding window sequence according to the load sequence; inputting the load sliding window sequence into the masked multi-head attention mechanism to obtain a first target load sliding window sequence, where the first target load sliding window sequence is an attention-weighted load sliding window sequence; inputting the first target load sliding window sequence into the third layer normalization for normalization processing to obtain a second target load sliding window sequence; inputting the second target load sliding window sequence and the second vector into the second multi-head attention mechanism to obtain a third vector, where the second vector includes a value vector and a key vector, and the third vector is the attention-weighted second vector; inputting the third vector into the fourth layer normalization to obtain a fourth vector; predicting the power load data according to the fourth vector.
[0104] In some embodiments, the target model further includes a first decoder and a second decoder. The first decoder includes a first feed-forward layer and a first activation layer, and the second decoder includes a second feed-forward layer and a second activation layer. The predicting of the power load data according to the fourth vector further includes: respectively inputting the fourth vector into the first feed-forward layer and the second feed-forward layer for non-linear transformation to respectively obtain a fifth vector and a sixth vector; respectively inputting the fifth vector and the sixth vector into the first activation layer and the second activation layer to respectively obtain a first reconstructed output and a second reconstructed output; obtaining the predicted power load data according to the first reconstructed output and the second reconstructed output.
[0105] The method proposed in this application is a short-term power load prediction model that considers the impact of distributed photovoltaic grid connection based on the combination of Transformer and temporal convolutional network. This model first uses the Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN) algorithm to decompose the original time series load data into multiple stable Intrinsic Mode Function (IMF) components and a residual (Res). Then, a combined model of Convolutional Neural Network (CNN) and bidirectional Transformer is used to predict each component one by one. In this way, the model can make full use of the advantages of CNN in feature extraction and the powerful ability of Transformer in processing time series data.
[0106] It should be noted that the method of the embodiment of the present application can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiment of the present application, and these multiple devices will interact with each other to complete the described method.
[0107] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0108] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides a power load prediction device.
[0109] Referring to Figure 4 , the power load prediction device includes:
[0110] An acquisition module 401, configured to acquire distributed photovoltaic power load data.
[0111] A preprocessing module 402, configured to preprocess the distributed photovoltaic power load data to obtain target power load data.
[0112] The preprocessing module 402 is further configured to add adaptive Gaussian white noise to the distributed photovoltaic power load data to obtain first power load data; decompose the first power load data to obtain a plurality of first modal components; perform ensemble averaging on the plurality of first modal components to obtain a first intrinsic mode component; and obtain the target power load data according to the distributed photovoltaic power load data and the first intrinsic mode component.
[0113] The preprocessing module 402 is further configured to obtain a residual according to the difference between the distributed photovoltaic power load data and the first intrinsic mode component; add adaptive Gaussian white noise to the residual to obtain a target residual; decompose the target residual to obtain a plurality of second modal components; perform ensemble averaging on the plurality of second modal components to obtain a second intrinsic mode component; and obtain the target power load data according to the residual and the second intrinsic mode component.
[0114] The preprocessing module 402 is further configured to perform anomaly detection on the distributed photovoltaic power load data to obtain target data; obtain anomaly data labels, and determine whether the target data is abnormal data according to the error value between the target data and the anomaly data labels; in response to the error value being greater than or equal to a threshold, determine that the target data is abnormal data; process the abnormal data to obtain the target power load data.
[0115] The time series module 403 is configured to obtain a load sequence according to the target power load data.
[0116] The prediction module 404 is configured to predict power load data based on the target model according to the load sequence, the target model is constructed based on the Transformer model, and the target model includes a window encoder.
[0117] Wherein, the target model includes an encoder, the encoder includes a first multi-head attention mechanism, a first layer of normalization, a target feed-forward layer, and a second layer of normalization. The prediction module 404 is further configured to input the load sequence into the first multi-head attention mechanism, calculate the correlation between different positions in the load sequence to obtain a first target load sequence, and the first target load sequence is an attention-weighted load sequence; input the first target load sequence into the first layer of normalization for normalization processing to obtain a second target load sequence, and the second target load sequence is the normalized attention-weighted load sequence; input the second target load sequence into the target feed-forward layer for non-linear transformation to obtain a first vector; input the first vector into the second layer of normalization for normalization processing to obtain a second vector; predict power load data according to the second vector.
[0118] Wherein, the target model further includes a window encoder, the window encoder includes a masked multi-head attention mechanism, a third layer of normalization, a second multi-head attention mechanism, and a fourth layer of normalization. The prediction module 404 is further configured to obtain a load sliding window sequence according to the load sequence; input the load sliding window sequence into the masked multi-head attention mechanism to obtain a first target load sliding window sequence, and the first target load sliding window sequence is an attention-weighted load sliding window sequence; input the first target load sliding window sequence into the third layer of normalization for normalization processing to obtain a second target load sliding window sequence; input the second target load sliding window sequence and the second vector into the second multi-head attention mechanism to obtain a third vector, the second vector includes a value vector and a key vector, and the third vector is an attention-weighted second vector; input the third vector into the fourth layer of normalization to obtain a fourth vector; predict power load data according to the fourth vector.
[0119] Among them, the target model further includes a first decoder and a second decoder. The first decoder includes a first feed-forward layer and a first activation layer, and the second decoder includes a second feed-forward layer and a second activation layer. The prediction module 404 is further configured to respectively input the fourth vector into the first feed-forward layer and the second feed-forward layer for non-linear transformation to respectively obtain a fifth vector and a sixth vector; input the fifth vector and the sixth vector into the first activation layer and the second activation layer respectively to respectively obtain a first reconstruction output and a second reconstruction output; and obtain the predicted power load data according to the first reconstruction output and the second reconstruction output.
[0120] For convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present application, the functions of each module can be implemented in one or more software and / or hardware.
[0121] The device in the above embodiment is used to implement the corresponding power load prediction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.
[0122] Based on the same technical concept, corresponding to the method in any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the power load prediction method described in any of the above embodiments.
[0123] Figure 5 FIG. shows a schematic diagram of an exemplary electronic device according to an embodiment of the present application. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0124] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0125] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and called and executed by the processor 1010.
[0126] The input / output interface 1030 is used to connect to an input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0127] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).
[0128] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0129] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary for implementing the solutions of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0130] The electronic device in the above embodiment is used to implement the corresponding power load prediction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0131] Based on the same technical concept, corresponding to the method in any of the above embodiments, the present application also provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the power load prediction method as described in any of the foregoing embodiments.
[0132] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0133] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the power load forecasting method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0134] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of brevity.
[0135] In addition, for simplicity of description and discussion, and so as not to make the embodiments of the present application difficult to understand, the known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0136] Although the present application has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0137] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present application.
Claims
1. A method for predicting power load data, comprising: Obtain distributed photovoltaic power load data; Preprocessing the distributed photovoltaic power load data to obtain target power load data; Obtaining a load sequence according to the target power load data; According to the load sequence, power load data is predicted based on a target model, wherein the target model is constructed based on a Transformer model, and the target model includes a window encoder.
2. The method of claim 1, wherein: The preprocessing of the distributed photovoltaic power load data to obtain target power load data further includes: Adding adaptive Gaussian white noise to the distributed photovoltaic power load data to obtain first power load data; Decomposing the first power load data to obtain a plurality of first modal components; Performing lumped averaging on the multiple first modal components to obtain a first eigenmodal component; The target power load data is obtained according to the distributed photovoltaic power load data and the first intrinsic mode component.
3. The method of claim 2, wherein: The step of obtaining the target power load data according to the distributed photovoltaic power load data and the first intrinsic mode component further comprises: Obtaining a residual according to a difference between the distributed photovoltaic power load data and the first intrinsic mode component; Adding adaptive Gaussian white noise to the residual to obtain a target residual; Decomposing the target residual to obtain a plurality of second modal components; Performing lumped averaging on the plurality of second modal components to obtain a second eigenmodal component; The target power load data is obtained according to the residual and the second eigenmode component.
4. The method of claim 1, wherein: The preprocessing of the distributed photovoltaic power load data to obtain target power load data further includes: Performing anomaly detection on the distributed photovoltaic power load data to obtain target data; Acquire an abnormal data label, and determine whether the target data is abnormal data according to an error value between the target data and the abnormal data label; In response to the error value being greater than or equal to a threshold, determining that the target data is abnormal data; The abnormal data is processed to obtain the target power load data.
5. The method of claim 1, wherein: The target model includes an encoder, the encoder includes a first multi-head attention mechanism, a first layer of standardization, a target feedforward layer, and a second layer of standardization, and the predicting of power load data based on the target model according to the load sequence further includes: Inputting the load sequence into the first multi-head attention mechanism, calculating the correlation between different positions in the load sequence to obtain a first target load sequence, where the first target load sequence is an attention-weighted load sequence; Inputting the first target load sequence into the first layer for normalization processing to obtain a second target load sequence, wherein the second target load sequence is the standardized attention-weighted load sequence; Inputting the second target load sequence into the target feed-forward layer for nonlinear transformation to obtain a first vector; Input the first vector into the second layer for normalization to obtain a second vector; Power load data is predicted based on the second vector.
6. The method of claim 5, wherein: The target model further includes a window encoder, the window encoder includes a masked multi-head attention mechanism, a third layer of normalization, a second multi-head attention mechanism, and a fourth layer of normalization, and the predicting of power load data according to the second value vector further includes: According to the load sequence, a load sliding window sequence is obtained; Inputting the load sliding window sequence into the masked multi-head attention mechanism to obtain a first target load sliding window sequence, wherein the first target load sliding window sequence is an attention-weighted load sliding window sequence; Inputting the first target load sliding window sequence into the third layer normalization for normalization processing to obtain a second target load sliding window sequence; Input the second target load sliding window sequence and the second vector into the second multi-head attention mechanism to obtain a third vector, wherein the second vector includes a value vector and a key vector, and the third vector is the attention-weighted second vector; Input the third vector into the fourth layer for standardization to obtain a fourth vector; Power load data is predicted based on the fourth vector.
7. The method of claim 6, wherein: The target model further includes a first decoder and a second decoder, the first decoder includes a first feedforward layer and a first activation layer, the second decoder includes a second feedforward layer and a second activation layer, and the predicting of power load data according to the fourth vector further includes: Inputting the fourth vector into the first feed-forward layer and the second feed-forward layer for nonlinear transformation respectively, so as to obtain a fifth vector and a sixth vector respectively; Inputting the fifth vector and the sixth vector into the first activation layer and the second activation layer respectively to obtain a first reconstruction output and a second reconstruction output respectively; Predicted power load data is obtained according to the first reconstruction output and the second reconstruction output.
8. A power load forecasting device, comprising: An acquisition module is configured to acquire distributed photovoltaic power load data; A preprocessing module is configured to preprocess the distributed photovoltaic power load data to obtain target power load data; A time series module is configured to obtain a load sequence according to the target power load data; The prediction module is configured to predict the power load data based on the target model according to the load sequence, wherein the target model is constructed based on the Transformer model, and the target model includes a window encoder.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.