A Power Load Forecasting Method Based on a Dual-Channel Cross-Attention Network

Through the power load prediction method based on the dual-channel cross attention network, the problems of non-stationarity, feature extraction limitations and nonlinear mapping capabilities of load time series data prediction in the prior art are solved, and higher prediction accuracy and model generalization capabilities are achieved.

CN119669732BActive Publication Date: 2025-06-20JIAHE CO CREATION (DALIAN) INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510180416.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-20
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

The existing load time series data prediction methods have non-stationarity problems, feature extraction limitations, and insufficient nonlinear mapping capabilities.

Method used

The power load prediction method based on the dual-channel cross attention network is adopted, and the parameters are optimized by variational modal decomposition and adaptive particle swarm algorithm to extract the key features of the power load data. Combined with a bidirectional long and short-term memory network and Transformer model, and feature enhancement is performed through improved cross-attention mechanism and multi-layer perceptron MLP-weighted fusion of Kolmogolov-Arnold network KAN.

Benefits of technology

Effectively extract key features in power load data, smoothing non-stationary components, improving the model's ability to handle complex timing tasks, significantly improving the accuracy of load time series prediction and model generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669732B_ABST
    Figure CN119669732B_ABST
Patent Text Reader

Abstract

A power load forecasting method based on a dual-channel cross-attention network belongs to the field of time series prediction technology in deep learning. This method includes the following steps: obtaining a power load data set and preprocessing it, applying VMD to decompose the power load data into multiple local components, and iteratively optimizing the VMD parameters through the APSO algorithm, screening the subsequence with the lowest complexity and combining it with the original features as the model input; then constructing a dual-channel cross-attention network model, using the parallel mechanism of the BiLSTM channel and the Transformer channel to process time series data, enhancing the feature expression ability through an improved cross-attention layer, and using an MLP fusion KAN network to perform power load forecasting. The present invention conducts time series prediction of power load by constructing a dual-channel cross-attention network, with high prediction accuracy and robustness, significantly improving the accuracy of load time series prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of time series prediction in deep learning, and particularly relates to a power load prediction method based on a dual-channel cross-attention network. Background Art

[0002] Time series data prediction has important application value in the field of power load, especially showing significant advantages in the prediction of power transformer temperature. By modeling and analyzing historical temperature, load, voltage and other related data, time series data prediction can accurately predict the future temperature change trend of the transformer, and identify potential overheating risks or abnormal states in advance. This not only helps in the health management and fault warning of equipment, but also can optimize the operation and maintenance strategies of the power system, extend the equipment life, reduce the downtime risk and operation and maintenance costs, thus ensuring the safe, stable and efficient operation of the power system.

[0003] There are various methods for predicting load time series data. According to different prediction principles, they can be divided into two categories: traditional prediction methods and deep learning-based prediction methods. Traditional prediction methods often have low prediction accuracy and high requirements for data, and are mainly applicable to stationary and smooth data. Recently developed deep learning-based methods have achieved great success in many application fields of time series prediction. Most of the model architectures used are GRU / LSTM models, but such models can only extract information between time dimensions and cannot obtain information of time series data in parallel from multiple dimensions. Moreover, in the face of complex, non-linear and non-stationary time series data, there are still problems such as non-stationarity, feature extraction limitations and insufficient non-linear mapping ability in existing research.

[0004] Therefore, there is an urgent need for a new technical solution in the prior art to solve this problem. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: to provide a power load prediction method based on a dual-channel cross-attention network to solve the technical problems such as non-stationarity, feature extraction limitations and insufficient non-linear mapping ability existing in the existing load time series data prediction methods.

[0006] A power load prediction method based on a dual-channel cross-attention network includes the following steps, and the following steps are carried out sequentially:

[0007] Step S1: Obtain the power load dataset, preprocess the power load data therein, then apply variational mode decomposition (VMD) to decompose the time series of power load data into different modal data components. Iteratively optimize the VMD parameters through the adaptive particle swarm optimization (APSO), obtain the optimized decomposition sequence and combine it with the original power load features as new features to form hybrid feature data. Process the hybrid feature data through the sliding window technique and Z-Score standardization in sequence, and divide it into a training set and a test set. Use the data in the training set and test set as the input vectors of the model, and the vectors are three-dimensional feature vectors containing a set number of time steps.

[0008] Step S2: Construct a deep learning network model based on a dual-channel cross-attention mechanism. Use the bidirectional long short-term memory network (BiLSTM) model and the Transformer model as the BiLSTM channel and the Transformer channel respectively for parallel computing to achieve feature extraction, and then combine the improved cross-attention mechanism and the multi-layer perceptron (MLP) to perform weighted fusion on the network model of the Kolmogorov-Arnold network (KAN) for feature enhancement.

[0009] Step S3: Use the data in the training set as the input to train the constructed dual-channel cross-attention network model, and continuously optimize the model until convergence meets the set requirements.

[0010] Step S4: Use the trained network model to predict the power load data.

[0011] The preprocessing in Step S1 includes filling the missing values with the mean value, replacing the outliers with the mean value, and then performing normalization.

[0012] In optimizing the VMD decomposition parameters in Step S1, select the sample entropy as the fitness function to quantitatively evaluate the quality of parameter combinations, continuously iterate with the goal of minimizing the sample entropy to find the global optimal parameters, then use the global optimal parameters to perform VMD decomposition on the load values, screen the subsequence with the lowest complexity, and then calculate the sample entropy of the subsequence and perform enhancement through the method of adaptive threshold screening to obtain the optimized decomposition sequence.

[0013] The calculation formula of the sample entropy is as follows:

[0014] (1);

[0015] In the formula, represents the total number of data points; represents the similarity threshold; and respectively represent that when the embedding dimension is and the distance between vectors is less than probability; representing the sample entropy of a time subsequence of length .

[0016] The formula for Z-Score normalization in step S1 is:

[0017] (2);

[0018] where, is the current th data point, is the average value of the entire data set, is the standard deviation of the entire data set, Z is the th standardized score of the data point.

[0019] In step S2, the BiLSTM channel uses a bidirectional long short-term memory network BiLSTM and residual connection technology to take the sequence data X t =(x1, x2,..., x T ) as the input data X forward , pass it to the forward long short-term memory network LSTM, and obtain a forward output result . At the same time, reverse the original sequence data X t to get X backward =(x T , x T-1 ,..., x1), and input it into the reverse LSTM to obtain a reverse output result . Subsequently, splice the forward and reverse output results to achieve feature integration, so as to fully extract the time dependence between the long term and the short term, and add the integrated result to the original sequence data X t to achieve residual connection to further optimize the feature expression, and finally obtain the feature extraction result vector of this channel through a fully connected layer mapping;

[0020] For the Transformer channel, first take the sequence data X t =(x1, x2,..., x T ) as the input and pass it to a 1D convolutional neural network CNN to obtain containing local semantic information, and then The Positional Encoding is input into the Transformer Encoder. The self-attention mechanism of the Transformer Encoder is used to calculate the attention scores, followed by residual connection and layer normalization. Then, it is further processed by a feed-forward neural network to obtain a time-dependent feature sequence that combines global and local temporal features. Finally, the result vector of the channel feature extraction is obtained through mapping by a fully connected layer. ;

[0021] The feature sequences obtained from the two channels and are subjected to feature fusion through an improved cross-attention mechanism. The attention weights between different positions are calculated crosswise and weighted summation is performed on the obtained different attention weights. The resulting sequence is the key feature extracted.

[0022] The method for enhancing features by the improved cross-attention mechanism and the multi-layer perceptron MLP weighted fusion of the Kolmogorov - Arnold network KAN in step S2 is as follows: First, the Transformer channel features are used as the query vector matrix , and the BiLSTM channel features are used as the key vector matrix and the value vector matrix respectively. The attention weights based on the Transformer channel features are calculated through the attention mechanism, and a weighted temporal feature representation is generated ; Subsequently, the BiLSTM channel features are used as the query vector matrix , and the Transformer channel features are used as the key vector matrix and the value vector matrix respectively. Again, the attention weights based on the BiLSTM channel features are calculated through the attention mechanism, and a weighted temporal feature representation is generated . Then, the two attention weights are fused in a context dynamic weight manner to obtain a fused cross-attention feature representation. Finally, it is input into the network model of the multi-layer perceptron MLP fused with the Kolmogorov - Arnold network KAN for prediction, and its calculation formula is as follows:

[0023] (3);

[0024] (4);

[0025] (5);

[0026] (6);

[0027] (7);

[0028] (8);

[0029] Wherein, is the feature vector extracted by the Transformer channel; is the feature vector extracted by the BiLSTM channel; , and are the learnable weight matrices for generating the query vector matrix , key vector matrix and value vector matrix in the Transformer channel, respectively; , and are the learnable weight matrices for generating the query vector matrix , key vector matrix and value vector matrix in the BiLSTM channel, respectively; is the dimension size of the key vector; is the normalization operation; is the context feature with a sequence length of ; FC represents the fully connected mapping; is the activation function; represents the context weight; crossAttention is the output of the improved cross-attention mechanism; the matrix represents the non-linear mapping. The matrix in the layer is composed of a linear transformation matrix and a set of trainable univariate spline functions . Among them, is the spline function corresponding to the element at the( , i , j ) position in the linear transformation matrix of the layer. Different layers are connected through the composite operation and represent the results predicted by KAN and MLP, respectively; in the MLP calculation, represents the weight matrix of the layer, and represent the bias terms of the first layer and the layer, respectively;

[0030] The convergence judgment of the model training in step S3 is based on the loss function setting. The loss function uses huber_loss to calculate the parameter gradient of the model and reversely updates the parameters of the model. The training process is expressed as:

[0031] (9);

[0032] In the formula, Indicates The true value of the data; Indicates The predicted value of data; Represents the threshold used to switch the range of square loss and absolute loss; and The difference between the two is less than , then the loss function uses square loss; and The difference between the two is greater than , linear loss is used to avoid outliers from having too much impact on the loss function.

[0033] Through the above design scheme, the present invention can bring the following beneficial effects:

[0034] The present invention adopts improved variational mode decomposition technology to decompose and reconstruct the power load time series, uses variational mode decomposition VMD optimized by adaptive particle swarm algorithm APSO to decompose the load data, and then calculates the sample entropy of different modes, adaptively sets the threshold to screen the decomposition sequence that meets the conditions, so as to generate an enhanced decomposition sequence RIMF ,This method effectively extracts the key features in the power load data, ,smoothes the non-stationary components, and improves the model’s ,ability to handle complex time series tasks.

[0035] The present invention designs an improved cross-attention mechanism for the dual-channel hybrid model to fuse the dual-channel features, thereby enhancing the feature fusion capability of the model and effectively improving the accuracy of load time series prediction.

[0036] The present invention realizes prediction through the weighted fusion design of the multi-layer perceptron MLP network and the Kolmogorov-Arnold network KAN, enhances the nonlinear feature mapping capability of the model, and effectively improves the accuracy of load time series prediction and the generalization ability of the model.

[0037] In summary, the present invention improves the accuracy of existing time series prediction and can be promoted in multiple time series prediction fields such as power load prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments:

[0039] Figure 1 It is a flowchart of a power load forecasting method based on a dual-channel cross-attention network of the present invention;

[0040] Figure 2 It is a decomposition diagram of data in a power load forecasting method based on a dual-channel cross-attention network of the present invention after original variational mode decomposition (VMD);

[0041] Figure 3 It is a decomposition diagram of data in a power load forecasting method based on a dual-channel cross-attention network of the present invention after variational mode decomposition (VMD) optimized by an adaptive particle swarm optimization (APSO) algorithm;

[0042] Figure 4 It is a schematic diagram of the model structure of a dual-channel cross-attention network model (VBTCKN) in a power load forecasting method based on a dual-channel cross-attention network of the present invention;

[0043] Figure 5 It is an application flowchart of a power load forecasting method based on a dual-channel cross-attention network of the present invention;

[0044] Figure 6 It is a network architecture diagram of a Kolmogorov-Arnold network (KAN) adopted in a power load forecasting method based on a dual-channel cross-attention network of the present invention;

[0045] Figure 7 It is a training loss diagram comparing the dual-channel cross-attention network model (VBTCKN) with other models in a power load forecasting method based on a dual-channel cross-attention network of the present invention;

[0046] Figure 8 It is a comparison diagram of the predicted value and the true value of the dual-channel cross-attention network model (VBTCKN) in a power load forecasting method based on a dual-channel cross-attention network of the present invention;

[0047] Figure 9 It is a diagram of the RMSE change results of the dual-channel cross-attention network model (VBTCKN) and other models at different prediction steps in the ETTh1 dataset in a power load forecasting method based on a dual-channel cross-attention network of the present invention;

[0048] Figure 10 It is a diagram of the RMSE change results of the dual-channel cross-attention network model (VBTCKN) and other models at different prediction steps in the ETTh2 dataset in a power load forecasting method based on a dual-channel cross-attention network of the present invention;

[0049] Figure 11 The figure showing the change results of RMSE of the dual-channel cross-attention network model VBTCKN and other models at different prediction steps in the ETTm1 dataset in a power load forecasting method based on a dual-channel cross-attention network according to the present invention;

[0050] Figure 12 The figure showing the change results of RMSE of the dual-channel cross-attention network model VBTCKN and other models at different prediction steps in the ETTh2 dataset in a power load forecasting method based on a dual-channel cross-attention network according to the present invention;

[0051] Figure 13 The ablation comparison result figure of the dual-channel cross-attention network model VBTCKN in a power load forecasting method based on a dual-channel cross-attention network according to the present invention;

[0052] Figure 14 The figure showing the loss comparison of the dual-channel cross-attention network model VBTCKN and other models at a fixed length in a power load forecasting method based on a dual-channel cross-attention network according to the present invention;

[0053] Figure 15 The figure showing the change results of RMSE of the dual-channel cross-attention network model VBTCKN and other models at different prediction steps in the GEFCom2014-E dataset in a power load forecasting method based on a dual-channel cross-attention network according to the present invention;

[0054] Figure 16 The figure showing the change results of RMSE of the dual-channel cross-attention network model VBTCKN and other models at different prediction steps in the Weather dataset in a power load forecasting method based on a dual-channel cross-attention network according to the present invention. Detailed implementation manners

[0055] To better understand the purpose, structure and function of the present invention, the following further describes in detail a power load forecasting method based on a dual-channel cross-attention network according to the present invention with reference to the accompanying drawings. Embodiment

[0056] Figure 1 A flow schematic diagram of a power load forecasting method based on a dual-channel cross-attention network according to an embodiment of the present invention is provided, and the specific steps of the method are as follows:

[0057] Step S1: Obtain a power load dataset. After preprocessing and data augmentation of the power load data in the power load dataset, divide and generate a training set and a test set. The specific steps are as follows:

[0058] Step S1-1: Preprocess the obtained power load data, including filling missing values with the mean, replacing outliers with the mean, and constructing a dataset suitable for power load forecasting.

[0059] Step S1-2: Perform variational mode decomposition (VMD) on the preprocessed dataset to decompose the power load time series into different modal data components. Use the adaptive particle swarm optimization (APSO) algorithm to optimize the VMD decomposition parameters. Select sample entropy as the fitness function to quantitatively evaluate the quality of parameter combinations. Continuously iterate with the goal of minimizing sample entropy to find the global optimal parameters. Then, use the global optimal parameters to perform VMD decomposition on the load values, screen the subsequence with the lowest complexity, calculate the sample entropy of the subsequence, and enhance it through the method of adaptive threshold screening to obtain the optimized decomposition sequence.

[0060] Among them, the sample entropy calculation formula is as follows: (1);

[0061] In the formula, represents the total number of data points; represents the similarity threshold; and respectively represent the probabilities that the distance between vectors is less than and when the embedding dimensions are ; represents the sample entropy of the time subsequence with length .

[0062] The search strategy of the adaptive particle swarm optimization (APSO) algorithm mimics the group behavior of bird flocks and fish schools, guiding the particles to search for the global optimal solution, dynamically adjusting the inertia weight and learning factors to balance global exploration and local exploitation, and adaptively adjusting the parameters according to the evolutionary state of the population to ensure that the population evolves in a better direction in each iteration. The update formulas for velocity and position of the adaptive particle swarm optimization (APSO) algorithm during iteration are as follows:

[0063] (10);

[0064] (11);

[0065] In the formula, represents the velocity of particle at the -th iteration, represents the inertia weight, which is used to control the influence degree of the current velocity of the particle on the new velocity, and are learning factors, and are random numbers between 0 and 1, is the particle The currently searched optimal position is the optimal position of the entire particle swarm is the particle The spatial position at the current moment is the particle At the The current position at the

[0066] For the VMD decomposition sequence optimized by the adaptive particle swarm optimization algorithm APSO, calculate its sample entropy respectively and enhance it by using the method of adaptive threshold screening to obtain the enhanced decomposition sequence RIMF , and the calculation formula is as follows

[0067] (12);

[0068] (13);

[0069] (14);

[0070] (15);

[0071] In the formula, the similarity threshold is an adaptive parameter / are the maximum and minimum threshold factors respectively, which are set to 0.1 and 0.4 respectively in the present invention is the base of the natural logarithm is the sample entropy of the current decomposition sequence IMF is IMF The standard deviation of is the total number of data is the current The is the average value of the entire data set is the adaptive threshold for screening the signal x. The original variational mode decomposition VMD decomposition sequence and the variational mode decomposition VMD decomposition sequence optimized by the adaptive particle swarm optimization algorithm APSO are respectively as Figure 2 and Figure 3 shown

[0072] Step S1-3: Obtain the decomposed sequence optimized by the adaptive particle swarm optimization algorithm APSO and combine it with the original power load characteristics as new features to form hybrid feature data. Process the hybrid feature dataset using the sliding window technique, set the window size to 60, and each batch contains a three-dimensional feature vector B*T*C with 10 time steps. B represents the number of batch samples, T represents the time step, and C represents the feature dimension to construct the input vector of the model.

[0073] Step S1-4: Standardize the input vectors at different times through the Z-Score function and divide the standardized power load data to obtain a training set and a test set; the Z-Score function is as follows:

[0074] (2);

[0075] In the formula, is the standard deviation of the entire dataset, is the standardized score of the

[0076] Step S2: Construct a deep learning network model based on the dual-channel cross-attention mechanism. The deep learning network model of the dual-channel cross-attention mechanism of the present invention is as Figure 4 shown, consisting of an architecture with a BiLSTM channel and a Transformer channel in parallel, combined with an improved cross-attention mechanism and a multi-layer perceptron MLP to fuse the Kolmogorov-Arnold network KAN. Among them, the BiLSTM channel consists of a bidirectional long short-term memory network BiLSTM layer, a random dropout Dropout layer, a residual connection, and a fully connected module. First, for the three-dimensional vector sequence data X t =(x1, x2,..., x T ) processed by the improved variational mode decomposition VMD and the sliding window technique, input it into the bidirectional long short-term memory network BiLSTM layer to extract features. The data will pass through a 64-unit bidirectional long short-term memory network BiLSTM layer, and the random dropout Dropout layer is applied to prevent overfitting. Then, the data will pass through a residual connection to fuse the extracted temporal features with the original temporal features. Finally, the data will pass through a fully connected layer with 64 neurons to map the feature dimension extracted by the bidirectional long short-term memory network BiLSTM layer to 32 to extract the key temporal information. In addition, for the three-dimensional vector sequence data X t =(x1, x2,..., x T) will also be input into the Transformer channel at the same time to achieve parallel processing. The Transformer channel consists of a 1D convolution layer, a Transformer encoder Transformer Encoder layer and a fully connected layer. First, the time series data will undergo an upsampling process to map the feature dimension to 32. Then, the data will enter a 1D convolution layer with a convolution kernel size of 1 and a hidden layer dimension of 64 to obtain local information of the time series data. Then the data will enter the Transformer Encoder layer with 2 heads, 2 layers and a hidden dimension of 64. The Transformer Encoder layer contains the self-attention mechanism Self-Attention, residual connection Add and hierarchical normalization Norm, as well as the feedforward neural network Feed Forward to extract rich global time series information. Then the data will pass through a fully connected layer with 64 neurons to map the feature dimension extracted by the channel to 32 to extract the key time series information. The features extracted by the two channels are then input into the improved cross attention mechanism. The Transformer channel features are first used as the query vector matrix , with BiLSTM channel features as key vector matrices Sum vector matrix , the attention weight based on the Transformer channel features is calculated through the attention mechanism, and the weighted temporal feature representation is generated ; The BiLSTM channel features are then used as the query vector matrix , with Transformer channel features as key vector matrices Sum vector matrix , the attention weight based on the BiLSTM channel features is calculated again through the attention mechanism, and the weighted temporal feature representation is generated , and then the two attention weights are weightedly fused using the contextual dynamic weight method to obtain a fused cross-attention feature representation. Finally, the output of the cross-attention mechanism will enter the multi-layer perceptron MLP and the Kolmogorov-Arnold network KAN respectively to obtain the predicted value and perform weighted fusion, thereby obtaining the output predicted value of the network model of the multi-layer perceptron MLP fused with the Kolmogorov-Arnold network KAN. The KAN network in the present invention has 2 layers, the multi-layer perceptron MLP has 1 layer, and β is 0.5, such as Figure 6 Shown is a diagram of the KAN network architecture used in the present invention.

[0077] The BiLSTM channel uses a bidirectional long short-term memory network BiLSTM and residual connection technology to convert the sequence data X t =(x1, x2,..., xT ) As the input data X forward , it is passed to the forward long short-term memory network LSTM to obtain a forward output result , and at the same time, the original sequence data X t is reversed to obtain X backward = (x T , x T-1 ,..., x1), and it is input into the reverse LSTM to obtain a reverse output result . Subsequently, the forward and reverse output results are concatenated by Concat to achieve feature integration, so as to fully extract the temporal dependence between the long term and the short term, and the integrated result is added to the original sequence data X t to implement residual connection to further optimize the feature expression. Finally, the feature extraction result vector of this channel is obtained through the fully connected layer mapping FC .

[0078] For the Transformer channel, first, the sequence data X t = (x1, x2,..., x T ) is passed as input to the 1D convolutional neural network CNN to obtain containing local semantic information. Then, is combined with the positional encoding PositionalEnconding and input into the Transformer encoder Transformer Encoder. The self-attention mechanism Self-Attention of the Transformer encoder is used to calculate the attention scores, and then through the residual connection Add and layer normalization Norm, and then further processed by the feed-forward neural network Feed Forward to obtain the time-dependent feature sequence combining global and local temporal features. Finally, the result vector of the feature extraction of this channel is obtained through the fully connected layer mapping FC .

[0079] The feature sequences and obtained from the two channels are fused by an improved cross-attention mechanism to extract temporal information, and then input into the network model of the multi-layer perceptron MLP fused with the Kolmogorov-Arnold network KAN for prediction. Its calculation formula is as follows:

[0080] (3);

[0081] (4);

[0082] (5);

[0083] (6);

[0084] (7);

[0085] (8);

[0086] Wherein, is the feature vector extracted by the Transformer channel; is the feature vector extracted by the BiLSTM channel; , and are the learnable weight matrices for generating the query vector matrix , the key vector matrix and the value vector matrix in the Transformer channel, respectively; , and are the learnable weight matrices for generating the query vector matrix , the key vector matrix and the value vector matrix in the BiLSTM channel, respectively; is the dimension size of the key vector; is the normalization operation; is the context feature with a sequence length of ; FC represents the fully connected mapping; is the activation function; represents the context weight; crossAttention is the output of the improved cross-attention mechanism; the matrix represents the non-linear mapping. The matrix in the layer is composed of a linear transformation matrix and a set of trainable univariate spline functions . Among them, is the spline function corresponding to the element at the ( , i , j ) position in the layer linear transformation matrix. Different layers are connected by the composition operation and represent the results predicted by KAN and MLP, respectively; in the MLP calculation, represents the weight matrix of the and represent the bias terms of the first layer and the is the weight term for prediction result fusion, is the final output time series prediction result.

[0087] Step S3: Use the generated training set and test set as model inputs, and train the constructed dual-channel cross-attention network model. Considering the different requirements for extracting long-term and short-term time series information and global time series information, input the power load data into the BiLSTM channel and the Transformer channel respectively for parallel computing for feature extraction, and then combine the improved cross-attention mechanism and the multi-layer perceptron MLP weighted fusion of the Kolmogorov-Arnold network KAN network model for feature enhancement. Continuously optimize the dual-channel cross-attention network model until convergence and performance stability are achieved. The performance stability means that the dual-channel cross-attention network model reaches the index of low error, that is, high accuracy;

[0088] In this embodiment, when training the constructed dual-channel cross-attention network model, set the number of iterations to 100, the ratio of the training set to the test set is 9:1, the learning rate of Adam is defaulted to 0.001, and the batch size is set to 128, adopt the sliding window mechanism, set the window size to 60, and adopt the huber_loss regression problem loss function. During the training process, the feature extraction part of the dual-channel cross-attention mechanism and the time series prediction part of the multi-layer perceptron MLP weighted fusion of the Kolmogorov-Arnold network KAN, its training process can be expressed as:

[0089] (9);

[0090] In the formula, represents the true value of the th data; represents the predicted value of the th data; represents the threshold value used to switch the range of the square loss and the absolute loss; The difference between and is less than , then the square loss is adopted for the loss function; The difference between and is greater than

[0091] , then the linear loss is adopted to avoid the excessive influence of outliers on the loss function. Figure 7 Figure 8 The training loss of the present invention and the training loss results of other models are as

[0092] Step S4: Use the trained deep learning network model based on the dual-channel cross-attention mechanism to predict the power load data. Figure 5 The application flow chart of the method used in the present invention is shown.

[0093] The present invention uses the following several common performance indicators to measure the prediction performance of the deep learning network model based on the dual-channel cross-attention mechanism:

[0094] Mean Absolute Error (MAE): It is a statistical indicator that measures the average difference between the predicted value and the actual value. It is the average of the absolute values of the deviations of all individual observations from the arithmetic mean, describing the average of the absolute values of the differences between the predicted value and the true value. That is, the smaller the MAE, the more accurate the prediction of the model. The calculation formula is as follows:

[0095] (16);

[0096] In the formula, represents the total amount of data.

[0097] Mean Squared Error (MSE): It is a measure that represents the degree of difference between the estimator and the estimated quantity, describing the degree of difference between the predicted value and the true value. Similar to MAE, the smaller the MSE, the more accurate the prediction of the model. The calculation formula is as follows:

[0098] (17).

[0099] Root Mean Square Error (RMSE): It is an index used to measure the prediction accuracy of the prediction model on continuous data, representing the average deviation degree between the predicted value and the true value. Similar to MAE, the smaller the RMSE, the more accurate the prediction of the model. The calculation formula is as follows:

[0100] (18).

[0101] Coefficient of Determination ( ) : It is an index that evaluates the consistency between the predicted value and the actual value, and its value range is from 0 to 1. The closer it is to 1, the better the model fits the data. The calculation formula is as follows:

[0102] (19);

[0103] In the formula, represents the average value of all true data.

[0104] Set up comparative experiments, ablation experiments, and predict evaluation metrics: Use the following several models, CNN[1], CNN-BiGRU[2], LSTM, BiLSTM, EMD-BiLSTM[3], VMD-BiLSTM[4], and BiLSTM-attention-LSTM[5], as comparative experiments. The results are as Figure 9 , Figure 10 , Figure 11 , Figure 12 shown. It can be seen that in the four sub-datasets of the power transformer dataset ETT under the dual-channel cross-attention network model VBTCKN proposed by the present invention, the RMSE of multi-step prediction is the lowest.

[0105] Use BiLSTM, BiLSTM-Transformer, BiLSTM-Transformer-CrossAttention, BiLSTM-Transformer-CrossAttention-KAN, VMD-BiLSTM, VMD-Transformer as ablation experiments. The results are as Figure 13 shown, verifying the rationality and effectiveness of the model design of the present invention.

[0106] To further verify the applicability of the dual-channel cross-attention network, the dual-channel cross-attention network model VBTCKN was also compared with other models in the GEFCom2014-E dataset under the conditions of prediction lengths of 10 and 30, including eight curves. The results are as Figure 14 shown. It can be seen that the dual-channel cross-attention network model VBTCKN proposed by the present invention is still superior to other models. The specific values are shown in Table 1.

[0107] Table 1 is a comparison table of this model with existing mainstream time series prediction models

[0108] The specific content of these comparative models is publicly available at:

[0109] [1] K. Wang et al., "Multiple convolutional neural networks formultivariate time series prediction," Neurocomputing, vol. 360, pp. 107-119,2019.

[0110] [2] He, J. Liu, F. Yang, X. Yang, S. Guo, and G. Sun, "Short-termPrediction of Distribution Voltage Based on CNN-BiGRU," 2023 4thInternational Conference on Advanced Electrical and Energy Systems (AEES),Shanghai, China, 2023, pp. 361-365.

[0111] [3] L. Thi Minh Lien, V. Quoc Anh, N. Duc Tuyen and G. Fujita, "Prediction of State-of-Health and Remaining-Useful-Life of Battery Based onHybrid Neural Network Model," in IEEE Access, vol. 12, pp. 129022-129039,2024.

[0112] [4] S. Ghimire, R. C. Deo, D. Casillas-Pérez, and S. Salcedo-Sanz, "Electricity demand error corrections with attention bi-directional neuralnetworks," Energy, vol. 291, p. 129938, 2024.

[0113] To verify the generalization performance of the model of the present invention, the present invention also made a comparison with other mainstream models in the GEFCom2014-E and Weather datasets. The results are as Figure 15 、 Figure 16 shown. It can be seen that in the multi-step prediction in different datasets, the RMSE of the model proposed by the present invention is the lowest.

[0114] In summary, it can be known that the prediction accuracy of the power load prediction method based on the dual-channel cross-attention network of the present invention is higher than that of the existing time series prediction methods.

[0115] It will be understood that the present invention is described by way of some embodiments, and those skilled in the art will be aware that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Additionally, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A power load forecasting method based on a dual-channel cross-attention network, characterized by: The method comprises the following steps, and the following steps are performed in sequence: Step S1: Obtain a power load data set, pre-process the power load data therein, then apply variational mode decomposition VMD to decompose the power load data time series into data components of different modes, iteratively optimize VMD parameters through adaptive particle swarm algorithm APSO, obtain the optimized decomposition sequence and combine it with the original power load feature as a new feature to form mixed feature data, and sequentially process the mixed feature data through sliding window technology and Z-Score standardization, and divide it into training set and test set, and use the data in the training set and test set as the input vector of the model, and the vector is a three-dimensional feature vector containing a set time step; Step S2: Construct a deep learning network model based on a dual-channel cross-attention mechanism, use a bidirectional long short-term memory network BiLSTM model and a Transformer model as BiLSTM channels and Transformer channels for parallel calculation to achieve feature extraction, and then combine the improved cross-attention mechanism and the multi-layer perceptron MLP weighted fusion Kolmogorov-Arnold network KAN network model for feature enhancement. Specifically, first use the Transformer channel feature as the query vector matrix , with BiLSTM channel features as key vector matrices Sum vector matrix , the attention weight based on the Transformer channel features is calculated through the attention mechanism, and the weighted temporal feature representation is generated ; The BiLSTM channel features are then used as the query vector matrix , with Transformer channel features as key vector matrices Sum vector matrix , the attention weight based on BiLSTM channel features is calculated again through the attention mechanism, and the weighted temporal feature representation is generated Then, the two attention weights are fused using the contextual dynamic weight method to obtain the fused cross-attention feature representation, which is finally input into the network model of multi-layer perceptron MLP fused with Kolmogorov-Arnold network KAN for prediction; Step S3: Use the data in the training set as input to train the constructed dual-channel cross-attention network model, and continue to optimize the model until convergence meets the set requirements; Step S4: Use the trained network model to predict power load data.

2. The method for power load forecasting based on a dual-channel cross attention network according to claim 1 is characterized in that: The preprocessing in step S1 includes filling missing values ​​with mean values, replacing abnormal values ​​with mean values, and then normalizing.

3. The power load forecasting method based on a dual-channel cross attention network according to claim 1 is characterized in that: In the optimization of VMD decomposition parameters in step S1, sample entropy is selected as a fitness function to quantify the quality of parameter combinations, and the global optimal parameters are found by iterating continuously with the goal of minimizing sample entropy. The load value is then subjected to VMD decomposition using the global optimal parameters, and a subsequence with the lowest complexity is screened. The sample entropy of the subsequence is then calculated and enhanced by adaptive threshold screening to obtain an optimized decomposition sequence. The sample entropy calculation formula is as follows: (1); In the formula, N Represents the total number of data points; represents the similarity threshold; and They are respectively represented in the embedding dimension and When the distance between vectors is less than probability; Indicates the length is The sample entropy of a time subseries.

4. The method for power load forecasting based on a dual-channel cross attention network according to claim 1 is characterized in that: The formula for Z-Score standardization in step S1 is: (2); In the formula, For the current data points, is the average value of the entire data set, is the standard deviation of the entire data set, Z For the The normalized scores of the data points.

5. The method for power load forecasting based on a dual-channel cross attention network according to claim 1 is characterized in that: In step S2, the BiLSTM channel uses a bidirectional long short-term memory network BiLSTM and residual connection technology to convert the sequence data X t =(x1, x2,..., x T ) as input data X forward , passed to the positive long short-term memory network LSTM, and a positive output result is obtained , and the original sequence data X t Reverse and get X backward =(x T , x T-1 ,...,x1), and input it into the reverse LSTM to get a reverse output result Then, the forward and reverse output results are spliced ​​to achieve feature integration, so as to fully extract the time dependency between long-term and short-term, and the integrated results are combined with the original sequence data X t Add them to realize residual connection to further optimize the feature expression, and finally obtain the feature extraction result vector of the channel through full connection layer mapping ; The Transformer channel first transforms the sequence data X t =(x1, x2,..., x T ) is passed as input to a 1D convolutional neural network CNN to obtain local semantic information , then Combined with the positional encoding, it is input into the Transformer encoder. The self-attention mechanism of the Transformer encoder is used to calculate the attention score. Then, it is further processed through residual connection and hierarchical normalization, and then through the feedforward neural network to obtain a time-dependent feature sequence combining global and local temporal features. Finally, the result vector of the channel feature extraction is obtained through the fully connected layer mapping. ; The feature sequence obtained by the dual channels and The improved cross-attention mechanism is used to perform feature fusion, cross-calculate and obtain the attention weights between different positions, and the different attention weights are weighted summed up. The obtained sequence is the key feature to be extracted.

6. The method for power load forecasting based on a dual-channel cross attention network according to claim 1 is characterized in that: In the method for feature enhancement using the improved cross attention mechanism and the network model of the multi-layer perceptron MLP weighted fusion Kolmogorov-Arnold network KAN in step S2, the specific calculation formula is as follows: (3); (4); (5); (6); (7); (8); In the formula, is the feature vector extracted by the Transformer channel; is the feature vector extracted by the BiLSTM channel; , and They are respectively used in the Transformer channel to generate the query vector matrix , key vector matrix and the value vector matrix The learnable weight matrix of , and They are used to generate query vector matrices in the BiLSTM channel. , the key vector matrix and the value vector matrix The learnable weight matrix of is the dimension size of the key vector; It is a normalization operation; The sequence length is contextual features; FC stands for fully connected mapping; is the activation function; represents the context weight; crossAttention is the output of the improved cross attention mechanism; matrix It represents a nonlinear mapping. Matrix of layers It is composed of a linear transformation matrix and a set of trainable univariate spline functions Composition, among which It is The first ( i , j ) elements corresponding to the spline function, different layers are operated by composite operations Connect to form a hierarchical mapping relationship; x represents the output feature of the cross attention mechanism; and Respectively represent the results predicted by KAN and MLP; in MLP calculation, represents the weight matrix of the Lth layer, and Represent the bias terms of the first layer and the Lth layer respectively; is the weight item for the fusion of prediction results, It is the final output time series prediction result.

7. The method for power load forecasting based on a dual-channel cross attention network according to claim 1 is characterized in that: The convergence judgment of the model training in step S3 is based on the loss function setting. The loss function uses huber_loss to calculate the parameter gradient of the model and reversely updates the parameters of the model. The training process is expressed as: (9); In the formula, Indicates The true value of the data; Indicates The predicted value of data; Represents the threshold used to switch the range of square loss and absolute loss; and The difference between the two is less than , then the loss function uses square loss; and The difference between the two is greater than , linear loss is used to avoid outliers from having too much impact on the loss function.

Citation Information

Patent Citations

  • Load prediction method based on fusion of attention mechanism and spatio-temporal characteristics

    CN118940919A

  • Photovoltaic power prediction method and system based on KAN

    CN119482456A