ShuffleTransform multi-head attention mechanism accurate load prediction power generation scheduling method

Through the ShuffleTransformer multi-head attention mechanism combined with parallel convolutional neural network and Transformer structure, the problem that existing power load prediction methods are difficult to accurately predict long-term and complex scenario loads is solved, and accurate prediction of power loads is achieved, which improves prediction accuracy and computing efficiency.

CN120146299APending Publication Date: 2025-06-13GUANGXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510271641.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-09
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing power load prediction methods are difficult to accurately predict long-term loads and complex scene loads, and large-scale convolutional neural networks are difficult to deploy, and a single neural network is prone to ignore feature correlation.

Method used

The ShuffleTransformer multi-head attention mechanism is adopted, combined with parallel convolutional neural network and Transformer structure, and the data correlation and long-term dependencies are captured through parallel operation of dual-channel output and self-attention mechanism to improve prediction accuracy.

Benefits of technology

The long and short-term sequence prediction accuracy of power load is improved, the calculation amount and running time are reduced, the feature correlation is fully utilized, and the accuracy of prediction results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146299A_ABST
    Figure CN120146299A_ABST
Patent Text Reader

Abstract

The invention provides a ShuffleTransform multi-head attention mechanism accurate load prediction power generation scheduling method, and the structure of the ShuffleTransform multi-head attention mechanism accurate load prediction power generation scheduling method is characterized in that two optimized lightweight Shuffle networks are operated in parallel, and are inputted to a Transform structure. A convolutional layer is added in an encoder module of a Transform, and the convolutional layer and Multi-headAttention interact with each other. According to the ShuffleTransform multi-head attention mechanism accurate load prediction power generation scheduling method, the problem that a power load prediction result is not accurate enough can be solved, long and short sequence prediction results are optimized, the output of a generator set is reasonably distributed, and the minimization of energy consumption is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of electric power energy, intelligent systems and automation control, artificial intelligence and machine learning, data science, optimization, demand response and scheduling, and predictive analysis. It relates to a predictive optimization method and is applicable to the control of power system scheduling. Background Art

[0002] Load data is greatly affected by various factors such as climate and economy, and its output exhibits characteristics of volatility, non-linearity, and non-stationarity, which increases the difficulty of predicting the load and has a greater impact on the stability and operation of the power system. Currently, prediction methods are divided into direct prediction methods and indirect prediction methods from the perspective of the characteristics of the prediction model. The indirect prediction method uses physical models to drive electrical, meteorological, and thermal processes, with high modeling difficulty and high requirements for data real-time performance. In the direct prediction method, neural networks are widely used in related data prediction due to their strong learning ability and adaptability, especially showing excellent performance in dealing with non-linearity and volatility.

[0003] However, there are problems with existing large convolutional neural networks that are difficult to deploy in practical applications due to their complex models; single neural networks are prone to ignoring the correlation and mutual influence between features when analyzing features; and power load prediction methods have problems of being unable to accurately predict in the long term and unable to achieve accurate prediction in complex scenarios.

[0004] Therefore, a precise load prediction power generation scheduling method based on the ShuffleTransformer multi-head attention mechanism is proposed. A lightweight Shufflenet is used to optimize the model, significantly reducing the computational amount and accelerating the running speed; parallel computing with dual-channel output is introduced, and after the parallel computing with dual-channel output, through the interaction between the convolutional layer and self-attention, the purpose of deeply exploring the correlation degree of the input data is achieved; two models are mixed to extract the advantages of the two models to capture richer and deeper features and relationships; the Transformer is used to extract the temporal features of long-term dependencies in the sequence, and the self-attention mechanism captures the short-term dependencies of the load to improve the prediction accuracy of long and short sequences. Summary of the Invention

[0005] The present invention proposes a precise load prediction power generation scheduling method based on the ShuffleTransformer multi-head attention mechanism, which combines two parallel convolutional neural networks and Transforms to deeply explore the correlation degree of the input data, has the ability to improve the training efficiency, and realizes the improvement of the prediction accuracy of long and short sequences of power generation. The steps in the use process are as follows:

[0006] Step (1): There is The total power load data of a dispatching area to be powered is divided into matrix data, including a training set and a test set:

[0007] (1)

[0008] Among them, is the total number of input matrices; n is the number of rows in the training set, is the number of rows in the test set, and y is the number of columns in the test set and the training set; is the input dimension of the training set; is the input dimension of the test set; is the output dimension of the training set; is the output dimension of the test set;

[0009] Step (2): Input the matrix into the parallel convolutional neural network of the ShuffleTransformer multi-head attention mechanism;

[0010] (2)

[0011] Among them, , are the two output quantities of the parallel convolutional neural network, is the input matrix, where D is the depth or number of channels, H is the height, W is the width, and T is the sequence length; is a network model;

[0012] The convolutional mathematical model of Shuffle is:

[0013] (3)

[0014] Among them is the element of the output feature map at the i-th layer, j-th column, k-th row, and m-th channel; is the element of the input feature map at the d-th channel, i-th layer, j-th column, and k-th row; is the weight of the convolutional kernel at the m-th output channel, d-th input channel, and position (x, y, z); represents the value at the position of the d-th channel, i + x-th row, j + y-th column, and k + z-th time step of the input data; , and respectively represent one less than the sizes of the convolutional kernel in the depth, height, and width directions;

[0015] The batch normalization and activation layer mathematical model of Shuffle is:

[0016] (4)

[0017] (5)

[0018] (6)

[0019] (7)

[0020] where denotes calculating the mean value; is the data point of the s-th sample; denotes the total number of data; denotes calculating the variance; Normalize using the mean value and variance: is a very small constant to avoid the formula from being invalid due to a zero denominator; Scale and translate the normalized data to restore the data; is the scaling factor used to adjust the scale of the normalized data, is the offset factor used to adjust the offset of the normalized data;

[0021] Grouping, shuffling, depthwise separable convolution, and residual connection operations of ShuffleNet:

[0022] (8)

[0023] (9)

[0024] (10)

[0025] (11)

[0026] where, , C is the number of channels. Grouped convolution divides the number of channels C into g groups, and each group has C / g channels; is the output of the g-th group; is the input of the g-th group; is the convolution kernel of the -th channel in the g-th group; is the convolution operation; is the shuffled feature map; is the index permutation after the shuffling operation; is the input of the g-th group; is the 1x1 convolution kernel of the denotes summing the convolution results of all grouped channels; It is a 3x3 convolutional kernel; F(X) is a feed-forward network applied to the input, including a non-linear activation function and a fully connected layer; is the output feature map; Y is the addition of the input and the output of the feed-forward network;

[0027] Step (3): Concatenate the two output data from step (2) with the matrix from step (1) to form an enhanced input matrix, aiming to enrich the data representation ability;

[0028] Step (4): Input the enhanced input matrix from step (3) into the Transformer structure of the ShuffleTransformer multi-head attention mechanism;

[0029] The convolutional batch normalization activation of the Transformer is the same as that of Shufflenet;

[0030] The input of the Transformer and the embedding part convert the input sequence into a vector representation, and at the same time introduce position encoding to represent the position information of each position in the sequence; The composition of the position encoding of the Transformer is:

[0031] (12)

[0032] where position represents the position in the input sequence, represents the even-dimensional index of the position encoding vector, represents the odd-dimensional index of the position encoding vector, represents the dimension of the input embedding vector; and are respectively the th and th dimension values of the position encoding vector; and are the sine and cosine functions respectively;

[0033] The composition of the query Q, key K, and value V matrices of the multi-head attention mechanism of the Transformer is:

[0034] (13)

[0035] where the input sequence X undergoes linear transformation through three different weight matrices 、 and to obtain the query Q, key K, and value V matrices;

[0036] The composition of the attention scores of the multi-head attention mechanism of the Transformer is:

[0037] (14)

[0038] (15)

[0039] Among them, is the dimension of the key vector, is the transpose of K; is the attention score matrix; AttentionWeights is the attention weight, normalized by the function;

[0040] The mathematical model of the weighted sum of the multi-head attention mechanism of Transformer is:

[0041] (16)

[0042] Among them, O is the output sequence, composed of all weighted input elements combined; N represents the number of attention heads; is the attention weight of the th input element; is the

[0043] The core composition of the multi-head attention mechanism of Transformer is:

[0044] (17)

[0045] Among them, h is that the query, key, and value vectors are divided into h heads, is the concatenation of the attention of each head, , and are the attention outputs of the 1st, 2nd, and hth heads respectively, is the multi-head attention mechanism;

[0046] The mathematical model of the residual connection and layer normalization of the multi-head attention mechanism of Transformer is:

[0047] (18)

[0048] (19)

[0049] Among them, is layer normalization, is the output after the residual connection, the output after layer normalization;

[0050] The feed-forward of Transformer consists of the following parts:

[0051] (20)

[0052] Among them, is the output of the fully connected layer, is the output of the self-attention layer, is the connection layer; the updated will be used as the final output and sent to the next encoder layer or decoder layer;

[0053] The decoder part of the Transformer consists of:

[0054] The task of the decoder is to convert the output of the encoder into the final output sequence; except for the masking process, the decoder process is the same as the encoder process; the decoder process is input embedding, positional encoding, masking, self-attention mechanism, multi-head attention, feed-forward network, output layer, and loop processing;

[0055] The mathematical model of masked self-attention is:

[0056] (21)

[0057] Among them, is the masking matrix, which is used to prevent the decoder from seeing future information when calculating the attention weights; is the calculation of the attention mechanism;

[0058] Step (5): According to the network structure trained from the training set, repeat the process and apply it to the test set, and output the load prediction result of the test set;

[0059] Step (6): Input the final predicted result into the power generation scheduling program, and use the power generation scheduling program to give instructions on the power generation of the generator in the next 24 hours, and minimize the energy consumption by optimizing the output allocation of the generator sets.

[0060] The present invention has the following advantages and effects compared with the prior art:

[0061] (1) The existing load data is affected by various factors, and its output shows characteristics of volatility, non-linearity, and non-stationarity, making it difficult to mine data correlation and increasing the difficulty of predicting the load. However, the present invention can capture data features and deeply explore data correlation through the interaction of the parallel operation dual-channel output and the multi-head attention mechanism of the Transformer structure.

[0062] (2) The existing Shufflenet method focuses on reducing the computational complexity through efficient channel shuffle operations and grouped convolutions, and performs well in local feature extraction. However, for long sequence data or data with complex global dependencies, its feature extraction ability is relatively limited. In contrast, the present invention combines with the self-attention mechanism of the Transformer method, which can avoid only focusing on local features, pays attention to the degree of association between each position in the computational sequence and other positions, and effectively captures long-range dependencies.

[0063] (3) The existing Transformer has high computational resources and time investment, and its computational complexity is quadratic with respect to the sequence length. Especially when dealing with long sequence data, the computational amount will increase sharply. The present invention reduces the computational amount through the grouped convolution and channel shuffle operations of the Shufflenet method, which can avoid the computational bottleneck of the Transformer method when dealing with long sequences, thereby accelerating the training and inference speed.

[0064] (4) The existing single-network method has the problem of easily ignoring the correlation between features, resulting in an inaccurate prediction result for a single feature analysis. The present invention considers a parallel hybrid model, and the combined prediction method can make full use of the advantages of each method to improve the accuracy of the prediction result. Brief Description of the Drawings

[0065] Figure 1 is the overall framework diagram of the method of the present invention.

[0066] Figure 2 is the Shufflenet process framework diagram of the method of the present invention.

[0067] Figure 3 is the Transformer process framework diagram of the method of the present invention. Detailed Embodiment

[0068] A ShuffleTransformer multi-head attention mechanism accurate load prediction power generation scheduling method proposed by the present invention is described in detail as follows in conjunction with the accompanying drawings:

[0069] Figure 1 is the overall framework diagram of the method of the present invention. The specific implementation steps of the present invention are as follows:

[0070] (1)Input the load data and divide it into a training set and a test set. (2)To improve the calculation efficiency and accuracy, preprocess the data and convert it into matrix form. (3)Input the matrix into a parallel convolutional neural network to extract spatial features and temporal features respectively, optimize the model rate, and improve the accuracy. After the two obtained prediction data are extracted, they are fused in the splicing layer, and the features of the second stage are input into the Transformer layer. (4)The Transformer layer will dynamically combine spatio-temporal features according to the dependency and correlation degree of different features and learn, so as to optimize the prediction accuracy of the model. (5)Output the data processed by the Transformer layer as the final prediction result. (6)Input the final prediction result into the power generation scheduling program, and use the power generation scheduling program to give instructions on the power generation of the generator in the next 24 hours, optimize the output allocation of the generator set, and minimize the energy consumption.

[0071] Figure 2 It is the Shufflenet process framework diagram of the method of the present invention. First is the specific structure of the Shuffle part. The input passes through convolution, batch normalization, and activation layers. Then it passes through two basic units, namely (a) the unit structure with a stride of 1 and (b) the unit structure with a stride of 2. The core functions of Shufflenet, grouped convolution and channel shuffle, exist in the unit structure. (a) shows the process of grouped convolution, and (b) shows the process of grouped convolution and channel shuffle. Finally, the prediction data is output.

[0072] Figure 3 It is the Transformer process framework diagram of the method of the present invention. First is the specific structure of the Transformer part. The input passes through convolution, batch normalization, and activation layers. Then it passes through the encoder structure. The specific process of the encoder structure is input embedding, position encoding, self-attention mechanism, multi-head attention, feed-forward network, and output layer. (a) shows the specific principle of the self-attention mechanism, and (b) shows the multi-head attention mechanism composed of the self-attention mechanism. The decoder structure is similar to the encoder structure. The specific process is input embedding, position encoding, masking, self-attention mechanism, multi-head attention, feed-forward network, and output layer; the essential difference is that the sequence input to the decoder is blocked. Finally, the prediction data is output.

[0073] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. A ShuffleTransformer multi-head attention mechanism accurate load forecasting power generation scheduling method, characterized in that: Combining two parallel convolutional neural networks with Transform is used to deeply explore the correlation of input data, which has the ability to improve training efficiency and improve the accuracy of long-term and short-term series prediction of power generation. The steps in the use process are: Step (1): Yes The total power load data of the power generation dispatching area is divided into a training set and a test set: (1) in, is the total number of input matrices; n is the number of rows in the training set, is the number of test set rows, and y is the number of columns in the test set and the training set; is the input dimension of the training set; is the input dimension of the test set; is the output dimension of the training set; is the output dimension of the test set; Step (2): Input the matrix into the parallel convolutional neural network of the ShuffleTransformer multi-head attention mechanism; (2) in, , are the two outputs of the parallel convolutional neural network, is the input matrix, where D is the depth or number of channels, H is the height, W is the width, and T is the sequence length; It is a network model; The convolution mathematical model of Shuffle is: (3) in is the element of the output feature map at layer i, column j, row k, and channel m; is the element of the input feature map at the dth channel, ith layer, jth column, and kth row; is the weight of the convolution kernel at the mth output channel, dth input channel, position (x, y, z); Represents the value at the d-th channel, i+x-th row, j+y-th column, and k+z-th time step of the input data; , and Respectively represent the size of the convolution kernel in depth, height and width minus one; The mathematical model of Shuffle's batch normalization and activation layer is: (4) (5) (6) (7) in Indicates the calculation of mean; It represents the data point of the sth sample; Indicates the total number of data; Indicates the calculation of variance; Normalize using mean and variance: is a very small constant, which prevents the formula from being invalid due to the denominator being zero; Scale and translate the normalized data to restore the data; is the scaling factor used to adjust the scale of the normalized data, is the offset factor used to adjust the offset of the normalized data; ShuffleNet's grouping, shuffling, depth-wise separable convolution and residual connection operations: (8) (9) (10) (11) in, , C is the number of channels, and grouped convolution divides the number of channels C into g groups, each group has C / g channels; is the output of the gth group; It is Group input; It is the g group The convolution kernel of channels; is a convolution operation; is the feature map after shuffling; is the index permutation after the shuffle operation; It is Group input; It is 1x1 convolution kernel with channels; Indicates the sum of the convolution results of all grouped channels; is a 3x3 convolution kernel; F(X) is a feed-forward network applied to the input, including non-linear activation functions and fully connected layers; is the output feature map; Y is the sum of the input and the output of the feedforward network; Step (3): concatenate the two output data of step (2) with the matrix of step (1) to form an enhanced input matrix to achieve the purpose of enriching the expressive power of the data; Step (4): Input the enhanced input matrix of step (3) into the Transformer structure of the ShuffleTransformer multi-head attention mechanism; The convolutional batch normalization excitation of Transformer is consistent with Shufflenet; The input and embedding part of the Transformer converts the input sequence into a vector representation, and introduces position encoding to represent the position information of each position in the sequence; the Transformer position encoding is composed of: (12) Among them, position represents the position in the input sequence, represents the even-numbered dimension index of the position encoding vector, represents the odd-numbered dimension index of the position encoding vector, represents the dimension of the input embedding vector; and are the position encoding vectors and The value of the dimension; and are the sine and cosine functions respectively; The query Q, key K, value V matrix of the Transformer multi-head attention mechanism is composed as follows: (13) Among them, the input sequence X is passed through three different weight matrices , and Perform linear transformation to obtain query Q, key K and value V matrices; The composition of the attention score of the Transformer's multi-head attention mechanism is: (14) (15) in, is the dimension of the key vector, is the transpose of K; is the attention score matrix; AttentionWeights is the attention weight, The function is normalized; The mathematical model of the weighted summation of the Transformer's multi-head attention mechanism is: (16) Where O is the output sequence, consisting of all weighted input elements Combination; N represents the number of attention heads; It is The attention weights of the input elements; It is input elements; The core structure of Transformer's multi-head attention mechanism is: (17) where h is the query, key and value vectors are divided into h heads, is to stitch together the attention of each head, , and are the attention outputs of the 1st, 2nd and hth heads respectively, It is a multi-head attention mechanism; The mathematical model of the residual connection and layer normalization of the Transformer's multi-head attention mechanism is: (18) (19) in, is the layer normalization, is the output after residual connection, The output of the layer after normalization; The Transformer feedforward consists of the following parts: (20) in, is the output of the fully connected layer, is the output of the self-attention layer, is the connection layer; updated It will be sent to the next encoder layer or decoder layer as the final output; The decoder part of Transformer consists of: The task of the decoder is to convert the output of the encoder into the final output sequence; the decoder process is the same as the encoder process except for the mask process; the decoder process is input embedding, position encoding, masking, self-attention mechanism, multi-head attention, feedforward network, output layer and loop processing; The mathematical model of masked self-attention is: (21) in, is a mask matrix used to prevent the decoder from seeing future information when calculating attention weights; is the calculation of the attention mechanism; Step (5): Based on the network structure trained by the training set, repeat the process and apply it to the test set, and output the load forecasting result of the test set; Step (6): Input the final prediction result into the power generation scheduling program, and use the power generation scheduling program to give instructions on the power generation of the generator in the next 24 hours, so as to minimize energy consumption by optimizing the output distribution of the generator sets.