Multi-user short-term power load prediction method of improved Transform model
By improving the internal solution of the Transformer model, Nystrom self-attention and spatial attention module, combined with the Monte Carlo stochastic inactivation method, the problems of spatial correlation and multi-user probability prediction in short-term power load prediction are solved, and efficient multi-user load prediction is achieved.
Patent Information
- Application Number
- CN202510421549.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
AI Technical Summary
The existing Transformer model cannot effectively capture the spatial correlation between sequences in short-term power load prediction, there is a memory bottleneck, and traditional methods cannot provide multiple users' probability prediction and cannot meet the load prediction needs of massive users.
The internal solution module, Nystrom self-attention mechanism, spatial attention module and Monte Carlo stochastic inactivation probability prediction method are adopted to decompose the input data through sliding average and Fourier transform, and the Nystrom method is used to approximate the self-attention mechanism, conduct deep feature mining, and probability prediction is performed through Monte Carlo stochastic inactivation method.
It improves the accuracy and robustness of short-term multi-user load prediction, can output point prediction and probability prediction results simultaneously, reduces time complexity, and improves prediction accuracy and efficiency.
Smart Images

Figure CN120341834A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of new power systems, and specifically to a multi-user short-term power load forecasting method for improving the Transformer model. Background Art
[0002] In recent years, the power big data technology has made great progress. The power system collects, stores, processes, and analyzes various types of massive power information data, and gradually realizes the logical summary of the data in the industry. Power load forecasting is an important link to improve the grid dispatching level, plays a crucial role in the safe and stable operation of the entire power system, and has important significance and practical value for accurately and efficiently forecasting short-term power loads. However, short-term power load forecasting is affected by multiple uncertain factors such as weather, temperature, and holidays, resulting in low forecasting accuracy. At the same time, due to the volatility and uncertainty of the energy consumption pattern and the load forecasting requirements of a large number of users, the traditional forecasting methods targeting only a single object cannot meet the new situation of a large number of users. Reducing the time complexity and improving the forecasting accuracy on the basis of fully considering the external factors affecting power load forecasting and targeting multiple users are the key issues that the current short-term power load forecasting models need to solve.
[0003] The common methods for power load forecasting mainly include statistical methods and machine learning methods. However, different technical defects have emerged for the above methods. For example, although the statistical method has a simple structure and a fast training speed, it cannot reflect the non-linear characteristics of the power load sequence. When the machine learning method faces a complex power system, the mining of data features is slightly insufficient.
[0004] Deep learning has the ability to efficiently represent data and model complex non-linear relationships in power load forecasting. Among them, the Transformer model, as a mainstream model in deep learning models, is widely used in the field of power load forecasting due to its advantages such as supporting parallelism and fast training speed by abandoning network structures such as recurrent neural network (RNN) and convolutional neural network (CNN).
[0005] However, it is found in practical applications that the standard Transformer model cannot capture the spatial correlation between sequences and cannot process highly volatile sequences. In the traditional Transformer model, both the decoder and the encoder adopt networks based on the self-attention mechanism, and its quadratic time complexity and high memory utilization increase quadratically with the input length L, resulting in a memory bottleneck when the traditional Transformer is applied to short-term power load forecasting.
[0006] At the same time, in the existing technology for multi-user load forecasting, it should be noted that the existing spatio-temporal methods can only provide deterministic forecasts and cannot provide probabilistic forecasts. Moreover, with the increasing uncertainty brought about by the high proportion of new energy penetration into the power grid in recent years, the role of probabilistic forecasting has been significantly improved in market transactions, operation scheduling and other aspects. Most of the existing load forecasting models emphasize the time characteristics of individual users while ignoring the spatial correlation of multiple users, which brings unnecessary troubles to power load forecasting.
[0007] Moreover, the existing technology also suffers from the lack of user-level data. The system-level load sequence is stable and has strong periodic laws, resulting in the research of load forecasting mainly focusing on the system level. However, due to the subjectivity and randomness of user electricity consumption behavior, which is easily affected by factors such as meteorological conditions and market electricity prices, the user load is often irregular and volatile. Therefore, accurate user and load forecasting still cannot be achieved. Summary of the Invention
[0008] The purpose of the present invention is to provide a multi-user short-term power load forecasting method that improves the Transformer model to solve the problems in the current market proposed in the above background technology.
[0009] To achieve the above purpose, the present invention provides the following technical solution: A multi-user short-term power load forecasting method that improves the Transformer model, including an S1 internal decomposition module, an S2 Nystrom self-attention mechanism, an S3 spatial attention module, and an S4 probability forecasting based on Monte Carlo random inactivation; specifically as follows:
[0010] S1 internal decomposition module: Set an internal decomposition block based on the idea of moving average and Fourier transform, and integrate it into the framework to decompose the input data sequence into trend, season, holiday and residual variables;
[0011] S2 Nystrom self-attention mechanism: Use the Nystrom method to approximate the softmax matrix in the standard self-attention mechanism to cope with the challenges of quadratic time complexity and memory usage in the Transformer;
[0012] S3 spatial attention module: Use convolution operations to deeply mine the features of the seasonal part after time series decomposition;
[0013] S4 probability forecasting based on Monte Carlo random inactivation: Use the Monte Carlo random inactivation method to extend the model to multi-user load probability forecasting.
[0014] Preferably, the decomposition formula of the S1 internal decomposition module is as follows:
[0015] y(t) = g(t) + s(t) + h(t) + ε(t)
[0016] Among them, y(t) is a subsequence of the sequence varying with time t; g(t) is a subsequence of the trend varying with time t; s(t) is a subsequence of the periodic variation with time t; h(t) is a subsequence of the holiday varying with time t; ε(t) is a subsequence of the error term varying with time t, and the error term represents any special transformation for which the model is not suitable; the trend part is decomposed by the moving average method; the seasonal part is approximated by the Fourier series for any smooth seasonal effect; the holiday part is a custom list represented by the unique name of the event or holiday.
[0017] Preferably, the specific calculation formulas for g(t), s(t), z(t), and h(t) are as follows:
[0018]
[0019] Among them, AvgPool is the average filtering convolution operation; Padding is the zero-padding width size on both sides; x is the sequence value; i is a numerical value, generally taking i = 10; p is the annual data, generally taking p = 365.25; z(t) is the generated regression quantity matrix; D i (i = 1,..., n) is the set of past and future dates of this holiday; K is the parameter assigned to each holiday; a i , b i are coefficients.
[0020] Preferably, the calculation formula for the softmax matrix S of the S2Nystrom self-attention mechanism is as follows:
[0021]
[0022] Among them, softmax(·) is the column-wise normalization function; Q is the query vector; K is the key vector; V is the value vector; d k is the dimension;
[0023] S is approximated by the orthogonal technique in the Nystrom method, and the calculation formula is as follows:
[0024]
[0025] Among them, A S ∈R m×n , B S ∈R m×(b-m) , F S ∈R (b-m)×m , C S ∈R (b-m)×(b-m) , d q is the dimension model, m is the number of columns of the sampling matrix, and b is the number of rows of the matrix Q;
[0026] Through Nystrom decomposition, the calculation formula of S is as follows:
[0027]
[0028] Among them, is the Moore-Penrose generalized inverse of A S , A S , B S , F S are all sub-matrices obtained by sampling the matrix S, denoted as
[0029] The calculation formula of the matrix S is as follows:
[0030]
[0031] Preferably, in the S3 spatial attention module, the seasonal part S de decomposed from the second internal decomposition block of the decoder is subjected to a convolution operation f conv using a 3×3 convolution kernel to convert it into a feature map F. The 3×3 convolution kernel has a relatively low computational complexity, and the calculation formula is as follows:
[0032] F = M c * S de
[0033] Among them, * represents the convolution operation; M C is the c convolution filters used.
[0034] Preferably, the spatial information in the original features is transformed into another space while retaining the key information, and the calculation formula is as follows:
[0035]
[0036] Among them, σ is the sigmoid function; f 7×7 is the convolution operation with a filter size of 7×7; and are the channel descriptions obtained by performing average pooling AvgPool(F) and max pooling MaxPool(F) on one channel dimension of F respectively; M s (F) is the spatial attention map, which is used to describe the important position information of the feature map in the space.
[0037] Preferably, the probability prediction of the S4 based on Monte Carlo dropout specifically includes the following: Given a set of N observations X = {x1, x2,..., x N} and Y = {y1, y2,..., y N}, f W (·) is a neural network with parameters W. When the test sample x *When inputting the training model, the probability distribution p(y * |x * ) of the predicted value is obtained by calculating the marginalized posterior distribution as follows:
[0038] p(y * |x * ) = ∫ w p(y * |f w (x * ))p(W|X,Y)dW.
[0039] Preferably, the variance of the prediction distribution quantifies the prediction uncertainty, which can be further decomposed using the law of total variance as:
[0040] Var(y * |x * ) = Var[E(y * |W,x * )] + E[Var(y * |W,x * )] = Var(f W (x * )) + σ 2
[0041] Among them, the variance is decomposed into two terms: Var(f W (x * )) reflects the model uncertainty; σ 2 reflects the internal noise in the data generation process. The key to estimating the model uncertainty is the posterior distribution p(W|X,Y), also known as Bayesian inference. Here, MC dropout is used to approximate the model uncertainty. The specific process is as follows: Given a new input x * , calculate the output of the neural network with random inactivation for each layer. This random feedforward is repeated B times to obtain Then approximate the model uncertainty through the sample variance to obtain:
[0042]
[0043] Among them,
[0044] Preferably, according to Var(y * |x * ) = Var[E(y * |W,x * )] + E[Var(y * |W,x * )] = Var(f W (x * )) + σ2 Can be approximated as
[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0046] By improving the Transformer model, changing the convention of sequence decomposition preprocessing, embedding an internal decomposition module, a Nystrom self-attention mechanism module, and a spatial attention module, the present model has a deep decomposition architecture. The internal decomposition module can decompose more predictable components from complex time patterns, taking into account external factors affecting power load forecasting, such as meteorological factors and date factors. The spatio-temporal attention mechanism is improved to effectively extract the dynamic spatio-temporal dependence relationships among highly volatile residential users. And the Nystrom self-attention mechanism is introduced to approximate the standard self-attention mechanism in an approximate manner, effectively reducing the time complexity. Using the Monte Carlo random inactivation method, the model is enabled to output both point prediction and probability prediction results simultaneously. The accuracy and robustness of short-term multi-user load point prediction and probability prediction are improved.
[0047] The model of the present invention is mainly divided into an input layer, a time series decomposition layer, an encoder layer, a decoder layer, and an output layer. Historical data is used as input and encoded through positional encoding, and then decomposed through the time series decomposition layer. The historical load data is decomposed into trend, season, holiday, and random parts. The encoder layer gradually eliminates the trend term to obtain the season, holiday, and random parts. The decoder layer models the season, holiday, and random parts for in-depth feature mining. Spatial attention operations are performed on the mined features. The spatial attention operation can locate effective information in space and perform feature weighting to learn important spatial features. And spatio-temporal related prediction results are obtained through the Nystrom self-attention mechanism and an accumulative manner. Brief Description of the Drawings
[0048] The drawings, as a part of the present invention, are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, but do not constitute an improper limitation to the present invention. Obviously, the drawings in the following description are only some embodiments, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0049] Figure 1 It is a schematic diagram of the Transformer network structure;
[0050] Figure 2 It is a schematic diagram of the network structure of the improved Transformer model for the multi-user short-term power load forecasting method of the present invention. Detailed implementation mode
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] Please refer to Figure 1 and Figure 2 , a multi-user short-term power load forecasting method for improving the Transformer model, including an S1 internal decomposition module, an S2 Nystrom self-attention mechanism, an S3 spatial attention module, and an S4 probability prediction based on Monte Carlo random inactivation;
[0053] In specific implementation, the S1 internal decomposition module: Set an internal decomposition block based on the ideas of moving average and Fourier transform, and integrate it into the framework to decompose the input data sequence into trends, seasons, holidays, and residual variables.
[0054] As a further explanation, the above decomposition formula is specifically as follows:
[0055] y(t) = g(t) + s(t) + h(t) + ε(t)
[0056] Among them, y(t) is the subsequence of the sequence changing with time t; g(t) is the subsequence of the trend changing with time t; s(t) is the subsequence of the periodicity changing with time t; h(t) is the subsequence of the holiday changing with time t; ε(t) is the subsequence of the error term changing with time t, and the error term represents any special transformation that the model is not suitable for; the trend part is decomposed by the moving average method; the seasonal part uses Fourier series to approximate any smooth seasonal effect; the holiday part is a custom list, represented by the unique name of the event or holiday. The specific calculation formulas of g(t), s(t), z(t), and h(t) are as follows:
[0057]
[0058] Among them, AvgPool is the average filtering convolution operation; Padding is the zero-padding width size on both sides; x is the sequence value; i is a numerical value, generally taking i = 10; p is the annual data, generally taking p = 365.25; z(t) is the generated regression quantity matrix; D i (i = 1,..., n) is the set of past and future dates of this holiday; K is the parameter assigned to each holiday; a i , b i are coefficients.
[0059] S2Nystrom self-attention mechanism: Use the Nystrom method to approximate the softmax matrix in the standard self-attention mechanism to address the challenges of quadratic time complexity and memory usage in Transformer.
[0060] As a further explanation, the softmax matrix S is calculated as follows:
[0061]
[0062] Among them, softmax(·) is a column-by-column normalization function; Q is the query vector; K is the key vector; V is the value vector; d k For dimension.
[0063] S is approximated by the orthogonal technique in the Nystrom method, and the calculation formula is as follows:
[0064]
[0065] Among them, A S ∈R m×n , B S ∈R m×(b-m) , F S ∈R (b-m)×m , C S ∈R (b-m)×(b-m) , d q is the dimensional model, m is the number of columns of the sampling matrix, and b is the number of rows of the matrix Q.
[0066] Through Nystrom decomposition, the calculation formula of S is as follows:
[0067]
[0068] in, A S The Moore-Penrose generalized inverse, A S , B S 、F S They are all sub-matrices obtained by sampling the matrix S, denoted as
[0069] The calculation formula of matrix S is as follows:
[0070]
[0071] S3 spatial attention module: Use convolution operation to perform deep feature mining on the seasonal part after time series decomposition; decompose the seasonal part S from the second internal decomposition block of the decoder de , a convolution operation is performed through a 3×3 convolution kernel f convConverted into the feature map F, the 3×3 convolution kernel has a relatively low computational complexity.
[0072] As a further illustration, the calculation formula is as follows:
[0073] F = M c * S de
[0074] where * represents the convolution operation; M C is the c convolution filters used.
[0075] Transform the spatial information in the original features into another space and retain the key information. The calculation formula is as follows:
[0076]
[0077] where σ is the sigmoid function; f 7×7 is the convolution operation with a filter size of 7×7; and are the channel descriptions obtained from an average pooling AvgPool(F) and a max pooling MaxPool(F) of one channel dimension of F respectively; M s (F) is the spatial attention map, which is used to describe the important position information of the feature map in the space.
[0078] S4 is the probability prediction based on Monte Carlo dropout: Using the Monte Carlo dropout method to extend the model to multi-user load probability prediction.
[0079] As a further illustration, specifically as follows:
[0080] Given a set of N observations X = {x1, x2, …, x N} and Y = {y1, y2, …, y N}, f W (·) is a neural network with parameters W. When the test sample x * is input into the training model, the probability distribution p(y * |x * ) of the predicted value is obtained by calculating the marginalized posterior distribution as:
[0081] p(y * |x * ) = ∫ w p(y * |f w (x * ))p(W|X,Y)dW.
[0082] The variance of the prediction distribution quantifies the prediction uncertainty, which can be further decomposed using the law of total variance as:
[0083] Var(y * |x * ) = Var[E(y * |W,x * )] + E[Var(y * |W,x * )] = Var(f W (x * )) + σ 2
[0084] Among them, the variance is decomposed into two terms: Var(f W (x * )) reflects model uncertainty; σ 2 reflects the internal noise in the data generation process. The key to estimating model uncertainty is the posterior distribution p(W|X,Y), also known as Bayesian inference. Here, MC dropout is used to approximate model uncertainty. The specific process is as follows: Given a new input x * , calculate the output of the neural network with dropout for each layer. This random feedforward is repeated B times to obtain Then, approximate the model uncertainty through sample variance to obtain:
[0085]
[0086] Among them,
[0087] According to Var(y * |x * ) = Var[E(y * |W,x * )] + E[Var(y * |W,x * )] = Var(f W (x * )) + σ 2 It can be approximated as:
[0088]
[0089] It should be noted that:
[0090] (1) Influence of input features on the model:
[0091] To illustrate the influence of input features on the load forecasting model, different variable combinations are used as the input of the model respectively, and the above improved Transformer model-based multi-user short-term electric load forecasting method is used for forecasting. The load forecasting effects of different input features are compared through experiments, and the results are shown in Table 1.
[0092]
[0093] Table 1 Influence of characteristic parameters on load forecasting
[0094] As shown in Table 1 above, when only dry-bulb temperature and humidity are used as model inputs, the prediction accuracy is the lowest; when wet-bulb temperature is added on this basis, the e of the prediction model ME 、e MRE are increased by 1.94% and 11.54% respectively. When humidity is added on this basis, the e of the prediction model ME 、e MRE are increased by 3.96% and 26.93% respectively. Since the higher the wet-bulb temperature, the higher the load, and the lower the temperature, the lower the load, the influence of temperature on prediction is more obvious. When dew-point temperature, dry-bulb temperature, wet-bulb temperature, and humidity are input into the model, the e of the prediction model ME 、e MRE are increased by 11.13% and 31.94% respectively, which indicates that the combined influence of humidity and temperature affects the accuracy of load forecasting. Only by fully considering the influence of temperature and humidity can the prediction accuracy of the model be more accurate.
[0095] When date factors and meteorological factors are input into the model, the prediction model can fully learn the effective information and implicit laws in the data, and the prediction effect is the best.
[0096] (2) Influence of the attention mechanism on the prediction model:
[0097] To verify the effectiveness of the improved Transformer model, compared with the model based on the self-attention mechanism and the model without the self-attention mechanism Noformer, the performance indicators are shown in Table 2.
[0098]
[0099] Table 2 Evaluation indicators of models using different attention mechanisms
[0100] As shown in Table 2 above, compared with the Noformer model without the self-attention mechanism, the prediction errors of the model based on the improved Transformer model and the model based on the self-attention mechanism are smaller, and the e ME are reduced by 16.45% and 12.01% respectively, and the e MRE are reduced by 8.93% and 6.29% respectively. This proves that the model based on the improved Transformer model can significantly improve the prediction ability of the model;
[0101] Compared with the model based on the self-attention mechanism, the prediction error of the model based on the improved Transformer model is smaller and the training speed is faster, which indicates that the improved Transformer model we proposed can significantly improve the prediction effect of the model.
[0102] (3) Validation of the effectiveness of the internal decomposition module:
[0103] To verify the effectiveness of the internal decomposition module, the following operations are performed on the original load training data respectively: directly predict without decomposition; predict after decomposition using the decomposition module; predict after decomposition using the proposed internal decomposition module. For the decomposed subsequences, predictions are made, and the prediction results of each sequence are superimposed to obtain the final load prediction results of the three methods. The evaluation indexes of each method are shown in Table 3.
[0104]
[0105] Table 3 Evaluation indexes using different decomposition methods
[0106] As shown in Table 3 above, compared with directly predicting using the original data without decomposition, the prediction errors using the decomposition module and the internal decomposition module are smaller, ME decreasing by 9.17% and 15.92% respectively. MRE decreasing by 5.65% and 9.30% respectively. It can be found that the prediction effect can be improved by the method of decomposition and then prediction.
[0107] Compared with the model based on the decomposition module, the model based on the internal decomposition module has a smaller prediction error. It shows that predicting after passing through the internal decomposition module is more conducive to improving the prediction accuracy.
[0108] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A multi - user short - term electric load forecasting method for improving the Transformer model, characterized in that It includes an S1 internal decomposition module, an S2 Nystrom self-attention mechanism, an S3 spatial attention module, and an S4 probability prediction based on Monte Carlo dropout; specifically as follows: S1 internal decomposition module: Set an internal decomposition block based on the ideas of moving average and Fourier transform, and integrate it into the framework to decompose the input data sequence into trend, season, holiday, and residual variables; S2 Nystrom self-attention mechanism: Use the Nystrom method to approximate the softmax matrix in the standard self-attention mechanism to address the challenges of quadratic time complexity and memory usage in Transformer; S3 spatial attention module: Use convolutional operations to perform in-depth feature mining on the seasonal part after time series decomposition; S4 probability prediction based on Monte Carlo dropout: Use the Monte Carlo dropout method to extend the model to multi-user load probability prediction.
2. The multi - user short - term electric load forecasting method for improving the Transformer model according to claim 1, wherein: The decomposition formula of the S1 internal decomposition module is as follows: y(t) = g(t) + s(t) + h(t) + ε(t) where y(t) is the subsequence of the sequence changing with time t; g(t) is the subsequence of the trend changing with time t; s(t) is the subsequence of the periodicity changing with time t; h(t) is the subsequence of the holiday changing with time t; ε(t) is the subsequence of the error term changing with time t, and the error term represents any special transformation that the model is not suitable for; the trend part is decomposed using the moving average method; the seasonal part is approximated by Fourier series for any smooth seasonal effect; the holiday part is a custom list represented by the unique name of the event or holiday.
3. The multi-user short-term electric load forecasting method for improving the Transformer model according to claim 2, characterized in that: The specific calculation formulas of g(t), s(t), z(t), and h(t) are as follows: Among them, AvgPool is the average filtering convolution operation; Padding is the zero-padding width size on both sides; x is the sequence value; i is a numerical value, generally taking i = 10; p is the annual data, generally taking p = 365.25; z(t) is the generated regression quantity matrix; D i (i = 1,…, n) is the set of past and future dates of this holiday; K is the parameter assigned to each holiday; a i 、b i are coefficients.
4. The multi - user short - term electric load forecasting method for improving the Transformer model according to claim 1, characterized in that: The calculation formula of the softmax matrix S of the S2 Nystrom self-attention mechanism is as follows: Among them, softmax(·) is a column-wise normalization function; Q is the query vector; K is the key vector; V is the value vector; d k is the dimension; S is approximated through the orthogonal technique in the Nystrom method, and the calculation formula is as follows: Among them, A S ∈R m×n , B S ∈R m×(b-m) , F S ∈R (b-m)×m , C S ∈R (b-m)×(b-m) , d q is the dimensionality model, m is the number of columns of the sampling matrix, and b is the number of rows of the matrix Q; Through Nystrom decomposition, the calculation formula of S is as follows: Among them, is the Moore-Penrose generalized inverse of A, S A, S B, S F S are all submatrices obtained by sampling the matrix S, denoted as The calculation formula of matrix S is as follows:
5. The multi-user short-term electric load forecasting method for improving the Transformer model according to claim 1, characterized in that: In the S3 spatial attention module, the seasonal part S decomposed from the second internal decomposition block of the decoder de , is subjected to a convolution operation f through a 3×3 convolution kernel conv to be converted into a feature map F. The 3×3 convolution kernel has a relatively small computational complexity, and the calculation formula is as follows: F = M c *S de where, * represents the convolution operation; M C are c convolution filters used.
6. The multi - user short - term electric load forecasting method for improving the Transformer model according to claim 5, wherein: Transform the spatial information in the original features into another space and retain the key information. The calculation formula is as follows: where σ is the sigmoid function; f 7×7 is a convolution operation with a filter size of 7×7; and are the channel descriptions obtained by average pooling AvgPool(F) and max pooling MaxPool(F) in one channel dimension of F respectively; M s (F) is the spatial attention map, which is used to describe the important position information of the feature map in space.
7. The multi - user short - term electric load forecasting method for improving the Transformer model according to claim 1, characterized in that: The probability prediction based on Monte Carlo random inactivation in S4 specifically includes the following: Given a set of N observations X = {x1, x2, …, x N} and Y = {y1, y2, …, y N}, f W (·) is a neural network with parameters W. When the test sample x * is input into the training model, the probability distribution p(y * | x * ) of the predicted value is obtained by calculating the marginalized posterior distribution as follows: p(y * |x * ) = ∫ w p(y * |f w (x * ))p(W|X,Y)dW.
8. The multi-user short-term electric load forecasting method for improving the Transformer model according to claim 7, characterized in that: The variance of the prediction distribution quantifies the prediction uncertainty, which can be further decomposed using the law of total variance into: Var(y * |x * ) = Var[E(y * |W,x * )] + E[Var(y * |W,x * )] = Var(f W (x * )) + σ 2 Among them, the variance is decomposed into two terms: Var(f W (x * )) reflects model uncertainty; σ 2 reflects the internal noise in the data generation process. The key to estimating model uncertainty is the posterior distribution p(W|X,Y), also known as Bayesian inference. Here, MC dropout is used to approximate model uncertainty. The specific process is as follows: Given a new input x * , calculate the output of the neural network with dropout at each layer. This random feedforward is repeated B times to obtain Then, approximate the model uncertainty through sample variance to obtain: Among them, 9. The multi - user short - term electric load forecasting method for improving the Transformer model according to claim 8, characterized in that: According to Var(y * |x * ) = Var[E(y * |W,x * )] + E[Var(y * |W,x * )] = Var(f W (x * )) + σ 2 can be approximated as
Citation Information
Cited By
Multi-user unified load prediction method based on industry embedding and parameter fine tuning
CN121327493A
A multi-user unified load prediction method based on industry embedding and parameter fine-tuning
CN121327493B
Campus electrical load prediction method based on Transform model
CN121417169A