A fractional calculus energy reduction guiding method for enhancing deep transformer-attention integrated prediction
By combining a Transformer-Attention network and a fractional-order stochastic dynamic calculus controller, the energy reduction guidance signal is predicted and output, which solves the problem of energy imbalance in the integrated energy system and improves system stability and energy utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-15
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies fail to effectively consider the impact of energy consumption status, energy consumption coefficient, season, temperature and access rules on energy consumption in integrated energy systems, resulting in an imbalance between energy supply and consumption. Furthermore, existing methods are limited to a single object or fail to effectively describe the nonlinear relationship between energy consumption and time and the random fluctuations of noise signals.
A fractional-order calculus method with enhanced deep Transformer-Attention ensemble prediction is adopted, which combines Transformer-Attention network, temporal attention unit and fractional-order stochastic dynamic calculus controller. By predicting the baseline energy consumption and outputting energy reduction guidance signal, the energy consumption of the system is guided and the energy consumption is reduced.
It improved energy efficiency, enhanced system stability, promoted the integration of renewable energy, reduced energy consumption of the integrated energy system, and solved the problem of energy supply and consumption imbalance.
Smart Images

Figure CN116700011B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of energy control technology of power systems, artificial intelligence and the field of calculus application in mathematical application, and relates to a control method of artificial intelligence and comprehensive energy system, and is suitable for long-term energy reduction guidance of comprehensive energy system. BACKGROUND
[0002] The patent with the patent name "A long-term price guidance method for multi-group distributed flexible energy service providers" applied on August 25, 2020 and the application number 2020108660490 proposes a long-term dynamic game strategy under incomplete information, only considers the long-term dynamic game between the research objects, and does not add the influencing factors of the research objects to the control strategy. The patent with the patent name "A long-term price guidance method for dynamic differential control of flexible energy hybrid network" applied on January 18, 2022 and the application number 2021111974623 only considers the integer order differential of random dynamic differential, but the energy consumption and time are in a nonlinear relationship and have a non-Markov property, and the integer order random dynamic differential has limitations in describing the relationship of the non-Markov property. The patent with the patent name "A fractional order long-term price guidance method for enhancing deep attention bidirectional prediction" applied on October 18, 2022 and the application number 2022112767254 uses fractional order random dynamic differential for energy consumption guidance of electric vehicles, but this method is only limited to electric vehicles as a single object and does not consider the entire energy system. The influence of single electric vehicle energy consumption on the entire energy system is very small, and the energy consumption management of the entire energy system is lacking.
[0003] Therefore, a fractional calculus energy reduction guidance method for enhancing deep Transformer-Attention integrated prediction is proposed, which can consider the influence of energy consumption state, energy consumption coefficient, season, temperature and access rules on energy consumption; the method can play the regulation function of energy consumption in the comprehensive energy system and solve the problem of imbalance between energy supply and energy consumption in the comprehensive energy system; the method uses the Transformer-Attention network and the efficient time series prediction network combined with the time series attention unit to predict the benchmark energy consumption of users, which can solve the problem of predicting user energy consumption; the method introduces fractional order integration of Wiener process into the controller, which can better describe the random fluctuations of noise signals and random disturbances of the system; the method starts from the perspective of the comprehensive energy system, guides the system energy through the energy reduction guidance signal, reduces energy consumption while meeting the experience of energy consumption side; in the long run, the method reduces the energy consumption of the comprehensive energy system and improves the stability of the comprehensive energy system. SUMMARY
[0004] The application provides a fractional calculus energy reduction guiding method for enhancing deep Transformer-Attention integrated prediction, which combines a Transformer-Attention network, an efficient time series prediction network combined with a time series attention unit and a fractional random dynamic calculus controller, and is used for long-term energy reduction guiding of a comprehensive energy system, and has the functions of improving energy utilization, improving the stability of the comprehensive energy system, promoting renewable energy integration into the comprehensive energy system and reducing energy consumption.
[0005] Step (1): establishing an operation framework for energy reduction guiding of the comprehensive energy system; the comprehensive energy system obtains energy production from power plants, boilers and fossil fuels, collects expected energy consumption from the energy consumption side, compares the energy production and the expected energy consumption, finds the optimal energy reduction guiding signal through the energy reduction guiding method, and then sends the energy reduction guiding signal to the energy consumption side to guide industrial, commercial and residential energy consumption and reduce energy consumption.
[0006] Step (2): proposing fractional calculus control for enhancing deep Transformer-Attention integrated prediction, predicting the benchmark energy consumption of the comprehensive energy system through the Transformer-Attention network and the efficient time series prediction network combined with the time series attention unit.
[0007] Firstly, data preprocessing is performed, and the processed data is subjected to feature extraction; then, the feature data obtained after processing are predicted by the Transformer-Attention network method and the efficient time series prediction network combined with the time series attention unit, and the predicted benchmark energy consumption is obtained after selecting the prediction results.
[0008] The Transformer-Attention network is a deep network architecture with multi-head attention as a basic operation unit, which can balance between obtaining long-term dependence and low time complexity; the prediction accuracy is improved by proposing ProbSparse self-attention mechanism, self-attention distillation mechanism and generative decoder.
[0009] The Transformer-Attention network is composed of an input layer, an encoder, a decoder, a full connection layer and an output layer.
[0010] First, the historical consumption of electric energy, thermal energy and fuel are all input into the encoder to be encoded to obtain a mapping sequence, and then the target value to be predicted in the long sequence is filled in as zero, and the mapping sequence obtained by the encoder is input into the decoder together to directly generate the predicted output element to be obtained; the encoder will input the tth input sequence into the encoder to obtain a mapping sequence, and then the target value to be predicted in the long sequence is filled in as zero, and the mapping sequence obtained by the encoder is input into the decoder together to directly generate the predicted output element to be obtained; the encoder will input the tth input sequence
[0011] (1)
[0012] wherein, is an encoding matrix; t refers to the tth time; is a real matrix set with a size of ; is a real number; is the length of the sequence; d is the input dimension; The ProbSparse self-attention mechanism adopted by the Transformer-Attention network is different from the standard self-attention mechanism. The scaled dot product pair adopted by the standard self-attention mechanism is:
[0013]
[0014] (2)
[0015] wherein, is a standard self-attention mechanism value; Q is a query vector; K refers to a queried vector; V is a content vector; is the transpose of the queried vector; is a normalized exponential function; is the length of Q; is the length of K; is a real matrix set with a size of ; is a real matrix set with a size of ; is a real matrix set with a size of ; is a real matrix set with a size of
[0016] The ProbSparse self-attention mechanism defines the scaled dot product pair of the standard attention mechanism as a probability form kernel smoother: (3)
[0017] wherein, is a probability form self-attention mechanism value; i, j, l refer to the number of rows; is the ith row of Q; is the jth row of K; is the lth row of K; is the jth row of V; is a probability function; is a conditional probability distribution under . It refers to the sum of conditional probabilities;
[0018] Probability function in kernel smoother The asymmetric exponential kernel used is:
[0019] (4)
[0020] In the formula, yes Transpose of;
[0021] The probability distribution of the query vector satisfies a uniform distribution as follows:
[0022] (5)
[0023] In the formula, This refers to the uniform distribution of the query vector;
[0024] if Nearly uniform distribution If the self-attention becomes value V, it is redundant for predicting the output. Therefore, the similarity between distributions p and q is used to distinguish important parts of the sequence. The similarity measured by the Kullback-Leibler divergence is as follows:
[0025] (6)
[0026] In the formula, It refers to the Kullback-Leibler divergence between distributions p and q; Transpose of; It is a logarithmic function with base e;
[0027] Remove from Kullback-Leibler divergence This constant defines the sparsity metric for the i-th query vector as:
[0028] (7)
[0029] In the formula, It refers to the sparsity measure of the i-th query vector;
[0030] In standard self-attention mechanisms, the sparsity metric yields the attention value for the ProbSparse self-attention mechanism:
[0031] (8)
[0032] In the formula, This refers to the attention value of the ProbSparse self-attention mechanism; is a sparse matrix of the same size as Q, containing only the sparsity measure ;
[0033] The distillation process of attention values is to extract the attention values of the ProbSparse self-attention mechanism, give priority to the better features with dominant characteristics, and generate focused self-attention feature mapping in the next layer by calculating the ProbSparse self-attention value of each element in each encoding matrix; The distillation process from the nth layer to the (n+1) layer is:
[0034] (9)
[0035] wherein, is a sparse matrix of the same size as Q, containing only the sparsity measure is the self-attention feature mapping of the nth layer; is a sparse matrix of the same size as Q, containing only the sparsity measure is the self-attention feature mapping of the (n+1) layer; is a maximum pooling function; is an activation function; is a 1-dimensional convolution filter in the time dimension using the activation function; is a basic operation in the multi-head ProbSparse self-attention and attention block;
[0036] The generative decoder is stacked by 2 identical multi-head attention layers, and the generative prediction can alleviate the problem of speed decline in long-term prediction; the vector input to the decoder is:
[0037] (10)
[0038] wherein, is the vector input to the decoder; is the start token vector; is the placeholder vector of the target sequence, each element of which is 0; is the length of the start token vector; is the length of the placeholder vector of the target sequence; is a concatenation operation function;
[0039] After decoding operations on and , the vector is passed through a fully connected layer to obtain the predicted energy consumption:
[0040] (11)
[0041] wherein, is the energy consumption predicted by the Transformer-Attention network; is a set of real number matrices of size is a set of real number matrices; is the dimension of the output data; is the o-th vector of the target output;
[0042] The efficient time series prediction network with time series attention unit is not a recurrent neural network, but uses an attention mechanism to process time evolution in parallel; the efficient time series prediction network with time series attention unit decomposes the time series attention into two parts: static attention and dynamic attention; the static attention uses small core depth convolution and dilated convolution to achieve a large receptive field, thereby capturing the long-time dependence of the sequence; the dynamic attention learns the time weight by using the difference of the inter-time attention, thereby capturing the changing trend between sequences; the efficient time series prediction network with time series attention unit uses a difference dispersion regularization method to optimize the loss function of time series prediction learning; the difference dispersion regularization method converts the difference between the predicted value and the true value into a probability distribution, and calculates the Kullback-Leibler divergence between them, so that the efficient time series prediction network with time series attention unit learns the inherent change rule in time series; the historical consumption of electric energy, thermal energy and fuel is input into the input matrix of the efficient time series prediction network with time series attention unit:
[0043] (12)
[0044] wherein, T refers to the length of the input time series; is a set of real number matrices of size T;
[0045] The predicted value of the neural network mapping is:
[0046] (13)
[0047] wherein, is the predicted value of the neural network model mapping; is a neural network model;
[0048] The forward difference between the predicted value of the neural network model mapping and the true value is:
[0049] (14)
[0050] wherein, is the forward difference of the predicted value of the neural network mapping; is the i+1-th predicted value of the neural network model mapping; is the i-th predicted value of the neural network model mapping; is the forward difference of the true value; is referred to as a true value is referred to as the i+1th data is referred to as a true value is referred to as the i+1th data
[0051] The forward difference is converted into a probability by The function is:
[0052] (15)
[0053] In the formula, is referred to as a probability distribution function is referred to as dynamic attention is referred to as static attention is referred to as a temperature coefficient; exp() is an exponential function with e as the base
[0054] The differential entropy regularization function is obtained by calculating the Kullback-Leibler divergence between the probability distribution and
[0055] (16)
[0056] In the formula, is referred to as a differential entropy regularization function is the length of the time series to be predicted
[0057] The efficient time series prediction network combined with the time series attention unit is trained in an end-to-end manner in a completely unsupervised manner, and the evaluation difference loss function is composed of a mean square error loss and a constant weighted differential entropy regularization:
[0058] (17)
[0059] In the formula, is an evaluation difference loss function is a constant
[0060] The weight parameters of the efficient time series prediction network combined with the time series attention unit can be solved by the loss function:
[0061] (18)
[0062] In the formula, is the value of the solved weight parameters; argmin is the solution corresponding to the minimum value of the objective function is a weight parameter
[0063] The efficient time series prediction network combined with the time series attention unit is used to predict the subsequent , the energy consumption predicted by the efficient time series prediction network combined with the time series attention unit is
[0064] (19)
[0065] wherein, denotes the energy consumption predicted by the efficient time series prediction network combined with the time series attention unit; denotes a real matrix set with the size of
[0066] Step (3): The enhanced deep Transformer-Attention integrated prediction fractional calculus control is used for energy reduction guidance of the comprehensive energy system; the predicted benchmark energy consumption is input into the fractional random dynamic calculus controller, and the fractional random dynamic calculus controller outputs the energy reduction guidance signal;
[0067] The energy consumption of the comprehensive energy system includes power consumption, heat consumption and fuel consumption. The benchmark energy consumption obtained by the benchmark energy consumption prediction function through the energy consumption predicted by the Transformer-Attention network and the efficient time series prediction network combined with the time series attention unit is:
[0068] (20)
[0069] wherein, denotes the predicted benchmark energy consumption; denotes the benchmark energy consumption prediction function;
[0070] The energy consumption state differential output by the fractional random dynamic calculus controller is:
[0071] (21)
[0072] wherein, denotes the order of the fractional calculus; denotes the energy consumption state; denotes the fractional differential of the energy consumption state; denotes the energy obtained by the comprehensive energy system; denotes the energy consumption predicted by the fractional random dynamic calculus controller; denotes the fractional differential of time; denotes the noise intensity; denotes the fractional integral of the Wiener process; denotes the Wiener process;
[0073] An integrated energy system comprises an energy generation side and an energy consumption side. The energy consumption is affected by the change in energy consumption output from a fractional-order stochastic dynamic calculus controller, which is:
[0074] (twenty two)
[0075] In the formula, This refers to the change in energy consumption; () refers to a logical function; , , and These refer to logical functions. The coefficients of the energy consumption state function, energy consumption coefficient function, seasonal function, temperature function, and admission rule function are contained in parentheses; Sta() refers to the energy consumption state function. It is the parameter of the energy consumption state function Sta(); Coe() refers to the energy consumption coefficient function. This refers to the energy consumption coefficient; This refers to the parameters of the energy consumption coefficient function Coe(); Wea() refers to the seasonal function. This refers to seasonal conditions; Tem() refers to the temperature function. This refers to air temperature; Rul() refers to the admission rule function. This refers to the admission rules; This refers to the parameter of energy consumption;
[0076] The predicted electricity load is obtained using a fractional-order stochastic dynamic differential controller:
[0077] (twenty three)
[0078] In the formula, This refers to the proportion of renewable energy; () refers to a symbolic function;
[0079] The symbolic function S() is:
[0080] (twenty four)
[0081] Logical functions ()for:
[0082] (25)
[0083] In the formula, It is a logical function The parameters of ();
[0084] Step (4): The energy reduction guide signal is generated by the fractional order stochastic dynamic calculus controller considering the dynamic energy consumption changes caused by the energy consumption state, energy consumption coefficient, season, air temperature and access rules;
[0085] The energy consumption influencing factors, the predicted benchmark energy consumption and the energy amount provided by the integrated energy system are taken as the input variables of the fractional order stochastic dynamic calculus controller, and the energy reduction guide signal is taken as the output variable;
[0086] The energy consumption state function Sta() is:
[0087] (26)
[0088] In the formula, 、 、 and respectively refer to the coefficients of the skew degree, the constant term of the change amount, the quadratic term of the change amount and the sixth term of the change amount in the energy consumption state function Sta() controlling the energy consumption state;
[0089] The energy consumption coefficient function Coe() is:
[0090] (27)
[0091] In the formula, is the total number of splines; z refers to the zth spline; is the I spline function;
[0092] The season function Wea() is:
[0093] (28)
[0094] In the formula, is the sine function;
[0095] The air temperature function Tem() is:
[0096] (29)
[0097] The access rule function Rul() is:
[0098] (30)
[0099] The predicted benchmark energy consumption and the energy supply amount are taken as the input variables of the fractional order stochastic dynamic calculus controller, the predicted energy consumption is output through the fractional order stochastic dynamic differential equation, and the optimal energy reduction guide signal is solved by using the objective function; the function for solving the energy reduction guide signal by using the fractional order stochastic dynamic calculus controller is:
[0100] (31)
[0101] period is a prediction period; is a prediction energy consumption function taking energy consumption coefficient as a variable;
[0102] Step (5): the energy reduction guide signal is applied to the integrated energy system to guide the energy consumption side to use energy, improve energy utilization, strengthen the stability of the integrated energy system, promote the integration of renewable energy, and reduce the energy consumption of the integrated energy system.
[0103] The present application has the following advantages and effects compared with the prior art:
[0104] (1) The present application uses data preprocessing function, Transformer-Attention network and efficient time series prediction network combined with time series attention unit to predict the benchmark energy consumption of users, and introduces fractional order integral of Wiener process into fractional order stochastic dynamic calculus controller, which can improve the accuracy of energy reduction guide signal.
[0105] (2) The present application considers the influence of energy consumption state, energy consumption coefficient, season, temperature and access rule on energy consumption from the perspective of integrated energy system, and guides the energy consumption side to consume energy through energy reduction guide signal, thereby reducing the energy consumption of integrated energy system.
[0106] (3) Compared with the patent with the name of "a long-term price guide method for multi-group distributed flexible energy service providers" applied on August 25, 2020, with the application number of 2020108660490, the present application not only considers the dynamic game between energy production and energy consumption, but also considers the influencing factors of energy consumption in the energy reduction guide of integrated energy system.
[0107] (4) Compared with the patent with the name of "a long-term price guide method and dynamic differential method for dynamic differential control of flexible energy hybrid network" applied on January 18, 2022, with the application number of 2021111974623, the present application adds fractional order stochastic differential to the controller, which can better describe the nonlinear relationship between energy consumption and time and the non-Markov property.
[0108] (5) Compared with the patent with the name of "a fractional order long-term price guide method for enhancing deep attention bidirectional prediction" applied on October 18, 2022, with the application number of 2022112767254, the present application expands the research object from single electric vehicle to integrated energy system, and adds fractional order integral of Wiener process in the controller, which can better describe the random fluctuations of noise signal and random disturbances of system, and improve the accuracy of energy consumption guide signal. BRIEF DESCRIPTION OF DRAWINGS
[0109] Figure 1 is the fractional calculus energy reduction guiding method for enhanced deep Transformer-Attention integrated prediction of the method of the present application.
[0110] Figure 2 is the Transformer-Attention network of the method of the present application.
[0111] Figure 3 is the high-efficiency time series prediction network combined with the time series attention unit of the method of the present application. DETAILED DESCRIPTION
[0112] The present application proposes a fractional calculus energy reduction guiding method for enhanced deep Transformer-Attention integrated prediction, which is described in detail in combination with the drawings as follows:
[0113] Figure 1 is the fractional calculus energy reduction guiding method for enhanced deep Transformer-Attention integrated prediction of the method of the present application. First, raw energy consumption data is obtained from industry, commerce, and residence and data preprocessing is performed. Then, the preprocessed data is predicted by the Transformer-Attention network and the high-efficiency time series prediction network combined with the time series attention unit, respectively, to determine the predicted benchmark energy consumption. Finally, the predicted benchmark energy consumption is input into the fractional random dynamic calculus controller, combined with energy consumption influencing factors and energy production, to output the optimal energy reduction guiding signal to guide the consumption of energy by industry, commerce, and residence.
[0114] Figure 2 is the Transformer-Attention network of the method of the present application. First, the encoder receives a large number of long sequence inputs. The Transformer-Attention network uses ProbeSparse self-attention instead of standard self-attention. Then, multi-head attention is distilled, and the attention after the distillation operation is a kind of pyramid dependency relationship. Through layer-by-layer distillation, the dominant attention is extracted to the decoder, greatly reducing the network size. Finally, the decoder receives long sequence inputs, fills the target elements with zeros, calculates the weighted attention group of the feature map, and immediately predicts the output elements in a generative style.
[0115] Figure 3The high-efficiency timing prediction network combined with the timing attention unit is the method of the present application. The overall structure of the high-efficiency timing prediction network combined with the timing attention unit is data input, encoder, timing attention unit, decoder and data output. In the timing attention unit, first, the static attention generated by the actual value is transformed through small kernel depth convolution, dilatable depth convolution and 1*1 convolution to obtain a first output vector. Then, the dynamic attention mapped by the neural network is obtained through the average pooling layer and the full connection layer to obtain a second output vector. Finally, the two output vectors are combined and sent to the decoder, and the final data output is generated by the decoder.
[0116] The above only describes the preferred embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A fractional calculus energy reduction guiding method for enhancing deep Transformer-Attention integrated prediction, characterized in that, The method combines a Transformer-Attention network, an efficient time series prediction network combined with a time series attention unit, and a fractional order stochastic dynamic calculus controller to guide long-term energy reduction of a comprehensive energy system. Step (1): Establish an operation framework for energy reduction guidance of a comprehensive energy system; the comprehensive energy system obtains energy production from power plants, boilers, and fossil fuels, and collects expected energy consumption from the energy consumption side. The comprehensive energy system compares the energy production and the expected energy consumption, finds the optimal energy reduction guidance signal through the energy reduction guidance method, and then sends the energy reduction guidance signal to the energy consumption side to guide industrial, commercial, and residential energy consumption and reduce energy consumption. Step (2): Propose a fractional calculus control for enhancing deep Transformer-Attention integrated prediction, and predict the benchmark energy consumption of a comprehensive energy system through a Transformer-Attention network and an efficient time series prediction network combined with a time series attention unit. First, perform data preprocessing and feature extraction on the processed data; then, use the Transformer-Attention network method and the efficient time series prediction network combined with the time series attention unit to predict the benchmark energy consumption after processing the feature data. The Transformer-Attention network is a deep network architecture with multi-head attention as the basic operation unit, which can balance between long-term dependence and low time complexity. The prediction accuracy is improved by proposing ProbSparse self-attention mechanism, self-attention distillation mechanism, and generative decoder. The Transformer-Attention network consists of an input layer, an encoder, a decoder, a fully connected layer, and an output layer. First, the historical consumption of electric energy, thermal energy and fuel are all input into the encoder to be encoded to obtain a mapping sequence, and then the target value to be predicted in the long sequence is filled with zero, and the mapping sequence obtained by the encoder is input into the decoder together to directly generate the predicted output element to be obtained; the encoder encodes the tth input sequence into a matrix: (1) wherein denotes the encoding matrix; t denotes the time instant; denotes the set of real matrices of size ; and denotes a real number; denotes the length of the sequence; d denotes the input dimension; The ProbSparse self-attention mechanism used by the Transformer-Attention network is different from the standard self-attention mechanism. The scaled dot product pair used by the standard self-attention mechanism is: (2) wherein, denotes a standard self-attention mechanism value; Q is a query vector; K denotes a vector being queried; V is a content vector; denotes the transpose of the vector being queried; is a normalized exponential function; is the length of Q; is the length of K; denotes a set of real matrices of size ; denotes a set of real matrices of size ; The ProbSparse self-attention mechanism defines the scaled dot-product pair of the standard attention mechanism as a kernel smoother in the form of a probability: (3) In the formula, denotes a probability-form self-attention mechanism value; i, j, l refer to row numbers; refers to the i-th row of Q; refers to the j-th row of K; refers to the l-th row of K; refers to the j-th row of V; refers to the probability function; the conditional probability distribution of X under X; refers to the sum of conditional probabilities; Probability function in kernel smoother The asymmetric exponential kernel used is: (4) wherein is the transpose of The probability distribution of the query vector satisfies the uniform distribution: (5) wherein denotes a uniform distribution of the query vector; If near uniform distribution then the self-attention becomes the value V, which is redundant for predicting the output, and thus the similarity between the distributions p and q is employed to distinguish important parts in the sequence, measured by the Kullback-Leibler divergence: (6) wherein denotes the Kullback-Leibler divergence between the distributions p and q; the transpose of is the natural logarithm function with base e; In the Kullback-Leibler divergence This constant defines the sparsity measure of the ith query vector as (7) wherein denotes the sparsity measure of the i-th query vector; In the standard self-attention mechanism, the attention value of the ProbSparse self-attention mechanism is obtained by the sparsity measure: (8) wherein is the attention value of the ProbSparse self-attention mechanism; is a sparse matrix of the same size as Q, containing only the sparsity measure ; By calculating the ProbSparse self-attention value of each element in each encoding matrix, the distillation process of the attention value is to extract the attention value of the ProbSparse self-attention mechanism, give the better feature with the dominant feature the privilege of the better feature, and generate a focused self-attention feature mapping in the next layer; The distillation process from the n-th layer to the (n+1)-th layer is: (9) wherein denotes self-attention feature map of the n-th layer; denotes self-attention feature map of the (n+1)-th layer; denotes max-pooling function; denotes activation function; denotes performing 1-dimensional convolution filter in time dimension with activation function; denotes basic operations in multi-head ProbSparse self-attention and attention block; The generative decoder is stacked by two identical multi-head attention layers, and the generative prediction can alleviate the problem of speed decline in long-term prediction. The vector input to the decoder is: (10) wherein is a vector input to the decoder; is a start token vector; is a placeholder vector for the target sequence, each element of which is 0; is a length of the start token vector; is a length of the placeholder vector for the target sequence; is a concatenation operation function; By performing a decoding operation on the vector and the predicted energy consumption of the vector through a fully connected layer is: (11) In the formula, refers to the energy consumption predicted by the Transformer-Attention network; refers to the output value, and the size of the output value is a set of real number matrices; refers to the dimension of the output data; refers to the oth vector constituting the target output; The efficient time series prediction network combined with the time series attention unit does not use a recurrent neural network, but uses an attention mechanism to process time evolution in parallel. The efficient time series prediction network combined with the time series attention unit decomposes the time series attention into two parts: static attention and dynamic attention. Static attention uses small kernel depth convolution and dilated convolution to achieve a large receptive field, thereby capturing long-term dependencies of sequences. The dynamic attention learns the time sequence weight by using different time sequence attentions, thereby capturing the changing trend between sequences; The efficient time sequence prediction network combined with the time sequence attention unit uses a difference dispersion regularization method to optimize the loss function of time sequence prediction learning; the difference dispersion regularization method converts the difference between the predicted value and the true value into a probability distribution and calculates the Kullback-Leibler dispersion between them, so that the efficient time sequence prediction network combined with the time sequence attention unit learns the inherent change rule in the time sequence; the historical consumption of electric energy, thermal energy and fuel is input into the input matrix of the efficient time sequence prediction network combined with the time sequence attention unit, and the predicted value of the neural network mapping is: (12) In the formula, T refers to the length of the input time series; refers to a set of real number matrices with a size of T; The predicted value of the neural network mapping is: (13) In the formula, denotes a predicted value mapped by the neural network model; is a neural network model; The forward difference between the predicted value of the neural network model mapping and the true value is: (14) wherein is the forward difference of the predicted value of the neural network mapping; refers to the i+1th predicted value of the neural network model mapping; refers to the i+1th predicted value of the neural network model mapping; is the forward difference of the true value; refers to the true value the i+1th data; refers to the true value the i+1th data; Forward difference is transformed into probability by function is: (15) wherein denotes a probability distribution function; denotes dynamic attention; denotes static attention; denotes temperature coefficient; exp() is the exponential function with base e; By computing the Kullback-Leibler divergence between the probability distributions and The differential divergence regularization function is obtained by computing the Kullback-Leibler divergence between the probability distributions (16) In the formula, denotes a microdispersion regularization function; is the length of the time series to be predicted; The efficient temporal prediction network combined with the timing attention unit is trained in an end-to-end manner in a fully unsupervised manner by a mean square error loss and a constant The loss function for evaluating the difference in the weighted differential dispersion regularization is: (17) wherein is a loss function that evaluates the difference; is a constant; The weight parameters of the efficient time sequence prediction network combined with the time sequence attention unit are solved by the loss function: (18) In the formula, is the value of the solved weight parameter; argmin is the solution corresponding to the minimum value of the objective function; is a weight parameter; The efficient time series prediction network with the time series attention unit is to predict the subsequent from time t+1. The efficient time series prediction network with the time series attention unit can learn the mapping from The energy consumption obtained by the efficient time series prediction network with the time series attention unit is: (19) In the formula, refers to the energy consumption predicted by the efficient time sequence prediction network combined with the time sequence attention unit; refers to a set of real number matrices with a size of . Step (3): The fractional order calculus control integrated with the enhanced deep Transformer-Attention prediction is used for energy reduction guidance of the comprehensive energy system; the predicted benchmark energy consumption is input into the fractional order random dynamic calculus controller, and the fractional order random dynamic calculus controller outputs the energy reduction guidance signal; The energy consumption of the comprehensive energy system includes electric power consumption, thermal energy consumption and fuel consumption, and the benchmark energy consumption obtained by the benchmark energy consumption prediction function through the energy consumption predicted by the Transformer-Attention network and the efficient time sequence prediction network combined with the time sequence attention unit is: (20) In the formula, refers to the predicted reference energy consumption amount; refers to the reference energy consumption amount prediction function; The energy consumption state differential output by the fractional order random dynamic calculus controller is: (21) In the formula, denotes the order of fractional calculus; denotes the energy consumption state; denotes the fractional differential of the energy consumption state; denotes the energy amount obtained by the comprehensive energy system; denotes the energy consumption amount predicted by the fractional random dynamic calculus controller; denotes the fractional differential of time; denotes the noise intensity; denotes the fractional integral of the Wiener process; denotes the Wiener process; The comprehensive energy system includes energy generation side and energy consumption side, and the change amount of the energy consumption amount by using the energy consumption amount output by the fractional order random dynamic calculus controller is: (22) In the formula, is the change in the amount of energy consumption; () is a logical function; , , and are coefficients of the energy consumption state function, the energy consumption coefficient function, the season function, the temperature function and the access rule function in the logical function ; Sta() is the energy consumption state function; is a parameter of the energy consumption state function Sta(); Coe() is the energy consumption coefficient function; is the energy consumption coefficient; is a parameter of the energy consumption coefficient function Coe(); Wea() is the season function; is the season; Tem() is the temperature function; is the temperature; Rul() is the access rule function; is the access rule; is a parameter of the amount of energy consumption; The predicted electricity load is obtained by the fractional order random dynamic differential controller: (23) In the formula, denotes the proportion of the amount of renewable energy; () denotes the sign function; The sign function S() is: (24) Logic function () is: (25) wherein is a parameter of the logic function () Step (4): Considering the dynamic energy consumption change caused by the energy consumption state, the energy consumption coefficient, the season, the temperature and the access rule, the energy reduction guidance signal is generated by the fractional order random dynamic calculus controller; The energy consumption influencing factors, the predicted benchmark energy consumption and the energy amount provided by the comprehensive energy system are used as the input variables of the fractional order random dynamic calculus controller, and the energy reduction guidance signal is used as the output variable; The energy consumption state function Sta() is: (26) In the formula, , , and respectively refer to the coefficients of the skew degree, the constant term of the change amount, the quadratic term of the change amount and the sixth term of the change amount in the energy consumption state function Sta(). The energy consumption coefficient function Coe() is: (27) wherein denotes the total number of splines; z denotes the zth spline; denotes the Ith spline function; The season function Wea() is: (28) wherein denotes the sine function; The temperature function Tem() is: (29) The access rule function Rul() is: (30) The predicted benchmark energy consumption and the energy supply amount are used as the input variables of the fractional order random dynamic calculus controller, the predicted energy consumption is output by the fractional order random dynamic differential equation, and the optimal energy reduction guidance signal is solved by using the objective function; the function for solving the energy reduction guidance signal by using the fractional order random dynamic calculus controller is: (31) In the formula, period refers to a prediction period; refers to a prediction energy consumption function with energy consumption coefficient as a variable. Step (5): The energy reduction guidance signal is applied to the comprehensive energy system to guide the energy consumption side to use energy.
Citation Information
Patent Citations
Power transformer load prediction method based on Transform model
CN115622047A
Electric vehicle daily charging demand curve prediction method based on attention mechanism
CN115730710A