Multi-channel deep learning power load prediction method based on attention mechanism

By adopting a multi-channel deep learning method based on attention mechanism in power load prediction, combined with fluctuation-smooth feature fusion and improved fishing optimization algorithm, the problem of low prediction accuracy when load data is volatile in the prior art is solved, and more efficient and stable power load prediction is achieved.

CN120045869APending Publication Date: 2025-05-27CHINA THREE GORGES PROJECTS DEV CO LTD +2

Patent Information

Application Number
CN202510043247.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, when predicting power loads, especially when the load data sequence is volatile, the advantages of the attention mechanism are difficult to reflect, resulting in low prediction accuracy.

Method used

A multi-channel deep learning method based on attention mechanism is adopted, and the hyperparameters are optimized by setting fluctuation channels and smooth channels, combined with an improved leader-guided fishing optimization algorithm, to achieve fluctuation-smooth feature fusion and improve prediction accuracy.

Benefits of technology

It improves the accuracy and stability of power load prediction, enhances the system's ability to cope with power load uncertainty, and improves the overall efficiency of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045869A_ABST
    Figure CN120045869A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-channel deep learning power load prediction method based on an attention mechanism. The method comprises the following steps: carrying out data preprocessing operation of abnormal value repair on power load data by adopting IQR quartile distance and cubic spline interpolation; setting a fluctuation channel and a smooth channel according to fluctuation characteristics and smooth characteristics of the power load data, and constructing a deep learning load prediction model based on fluctuation-smooth characteristic fusion based on the fluctuation channel and the smooth channel; an improved leader-sleeve-guided fishing optimization algorithm is adopted to optimize hyper-parameters of the deep learning load prediction model based on fluctuation-smooth feature fusion; training a load prediction model of deep learning based on fluctuation-smooth feature fusion by adopting the power load data subjected to data preprocessing operation; and predicting the power load by adopting the trained load prediction model. The prediction stability is improved; the comprehensive efficiency of power load prediction is improved, and the coping capacity of the system to the uncertainty of the power load is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of load forecasting, and in particular to a power load forecasting method based on multi-channel deep learning with an attention mechanism. Background Art

[0002] With the large-scale grid connection of renewable energy power generation, its uncertainty and volatility have exacerbated the complexity and randomness of power loads, posing challenges to the stable operation of the energy system and also to accurately capturing the characteristics of load data by load forecasting models. Precise power load forecasting technology is an effective means to address this challenge. According to the forecasting principle, load forecasting can be generally classified into four categories: physical models, statistical models, artificial intelligence models, and hybrid models. Deep learning models can adaptively learn and process non-linear data using complex neural network architectures, and have the characteristics of strong generalization ability and high stability, and are widely used in the field of time series forecasting; applying the Recurrent neural network (RNN) that specifically processes sequence information to load forecasting, the RNN has problems of gradient disappearance or explosion; applying the Long short-term memory (LSTM) and Gate recurrent unit (GRU) which are variants of RNN to load forecasting, but the LSTM and GRU models are highly dependent on input data and lack the generalization ability for multivariate data; proposing to fuse different models to form a hybrid model to solve the problem of low forecasting accuracy; compared with a single model, the hybrid model has better forecasting performance. A new short-term power system load forecasting method based on Graph Convolutional Network and (Long Short-Term Memory, abbreviated as LSTM) is proposed; in order to mine key feature information in data, an attention mechanism is introduced into the forecasting model; however, the advantages of the attention mechanism are difficult to reflect when the load data sequence fluctuates greatly, and thus when the predicted load data sequence fluctuates greatly, the forecasting accuracy is not high. Summary of the Invention

[0003] Object of the Invention: In order to overcome the deficiencies existing in the prior art, the present invention provides a power load forecasting method based on multi-channel deep learning with an attention mechanism, which performs feature fusion on the fluctuation characteristics and smooth characteristics of power load data to forecast the load, and obtains more accurate and higher-stability load forecasting.

[0004] Technical solution: To achieve the above object, a power load forecasting method based on multi-channel deep learning with an attention mechanism of the present invention includes a data preprocessing operation of repairing outliers in power load data by using the IQR interquartile range and cubic spline interpolation; setting a fluctuation channel and a smoothing channel according to the fluctuation characteristics and smoothing characteristics of the power load data, and constructing a load forecasting model of deep learning based on the fusion of fluctuation-smoothing characteristics based on the fluctuation channel and the smoothing channel; using an improved leader-guided fishing optimization algorithm to optimize the hyperparameters of the load forecasting model of deep learning based on the fusion of fluctuation-smoothing characteristics; using the power load data after the data preprocessing operation to train the load forecasting model of deep learning based on the fusion of fluctuation-smoothing characteristics; using the trained load forecasting model of deep learning based on the fusion of fluctuation-smoothing characteristics to forecast the power load.

[0005] Furthermore, a data preprocessing operation of repairing outliers in power load data by using the IQR interquartile range and cubic spline interpolation; using the IQR interquartile range to detect outliers in the power load data, and the calculation process of the IQR interquartile range is as follows:

[0006] IQR = Q 3 - Q 1

[0007] In the formula, Q1 is the minimum value of the lower 25% of the data points in the power load dataset, called the first quartile; Q3 is the maximum value of the upper 25% of the data points in the power load dataset, called the third quartile;

[0008] By calculating the quartiles and interquartile range of the data, the data is divided into 3 different intervals, and the calculation process is as follows:

[0009] LOW = Q 1 - 1.5 * IQR

[0010] UP = Q 3 + 1.5 * IQR

[0011] When a sample value is less than LOW or greater than UP, then it is determined that the sample value is an outlier, and the sample value is deleted from the power load dataset.

[0012] Furthermore, cubic spline interpolation is used to fill in the deleted outliers in the power load dataset. Define the interpolation point set {x 0 , x 1 ,..., x n} and the corresponding function values {y 0 , y 1 ,..., y n}, x i represents the abscissa of the i-th known data point, and yi represents the corresponding to xi The corresponding ordinate; for each interval [x i , x i +1], define a cubic polynomial:

[0013] S i (x) = a i + b i (x - x i ) + c i (x - x i ) 2 + d i (x - x i ) 3

[0014] Wherein, a i , b i , c i , d i are all undetermined coefficients, and S i (x) is the ordinate corresponding to x i on the spline curve. The spline curve must pass through each data point (x i , y i ), that is, for each i, there is S i (x) = y i ;

[0015] And each x i must satisfy the following conditions:

[0016] S i-1 (x i ) = S i (x i )

[0017] S′ i-1 (x i ) = S′ i (x i )

[0018] S″ i-1 (x i ) = S″ i (x i )

[0019] When the above conditions are satisfied, the defined cubic polynomial is not only continuous in function value at each x i , but also its first derivative and second derivative are continuous.

[0020] Furthermore, use the boundary conditions to determine a i , b i , c i and d i, the boundary condition is a natural boundary condition, that is, the second derivative is 0 at the endpoints. The calculation process is as follows:

[0021] S′(x 0 )=0

[0022] S″(x n )=0

[0023] By using the boundary condition, a system of linear equations is constructed to solve for the coefficients a i , b i , c i and d i ;

[0024] The cubic spline interpolation function S(x) is a combination of defined cubic polynomials. In each interval [x i , x i +1], S(x) is defined by the corresponding polynomial S i (x). The calculation process is as follows:

[0025]

[0026] Through the abscissa x of the missing outlier, the corresponding S(x) is obtained, which is the substitute value of the missing outlier. The outliers in the power load data are deleted by using the IQR interquartile range and cubic spline interpolation, and the preprocessed power load data is obtained by filling in the calculated substitute value.

[0027] Furthermore, for the load prediction model based on deep learning with fluctuation-smoothing feature fusion, the channel composition of each moment's data is judged by the volatility n. One group is the fluctuation channel, and the other group is the smoothing channel; the calculation process of the volatility n defines the fluctuation area as the integral of the load curve over the current data interval, denoted as A truth :

[0028] A truth =∫x(t)dt

[0029] Define the test area A test as the comparison area for judging volatility. When there is no fluctuation in the data, the load curve is approximately a straight line. Select the first data x 0 in the i-th interval as the starting point and the last data x N as the ending point, and calculate the area of the triangle formed by this section of the load curve with respect to the X-axis as the test area. The calculation process is as follows:

[0030] A test =(x 0 +x N )×N / 2

[0031] The volatility η, an indicator representing the volatility of load data, is defined as follows:

[0032]

[0033] When the interval volatility is small, the value of η approaches 0; when the interval volatility is large, the absolute value of η tends to 1.

[0034] Furthermore, the fluctuation channel is used to process load data segments with large volatility, capture the fluctuation characteristics of the data, and through the processing of specific forget gates, input gates, and output gates, output the hidden layer state containing fluctuation information

[0035]

[0036] In the formula, F is the parameter of the fluctuation channel LSTM;

[0037] The smoothing channel is used to process relatively smooth load data segments, capture the continuous information and tiny change characteristics of the data, and through the processing of specific forget gates, input gates, and output gates, output the hidden layer state that preserves the smoothing information

[0038]

[0039] In the formula, S is the parameter of the smoothing channel LSTM.

[0040] Furthermore, the fishing optimization algorithm based on leader guidance improves the quality of the initial population through the Logistic-tent chaotic mapping method, enhancing the possibility of the population finding the global optimal solution; it introduces a leader-guided capture strategy, that is, taking the individual with the best fishing situation as a reference to guide the position update direction of other individuals;

[0041] Initialize the population using the Logistic-tent chaotic mapping to initialize the first population of fishermen Fisher1;

[0042]

[0043] In the formula, i is the population sequence number, and the value range is i ∈ {2, 3, 4... N}; r x is the control parameter, r x ∈ (0, 4).

[0044] Furthermore, the core of the leader-guided fishing optimization algorithm lies in improving the method of randomly selecting reference fishermen in the original algorithm to guiding the population position update through the three individuals with the best fishing efficiency in the fishermen population; by guiding through the individual with the best fishing efficiency, the convergence accuracy and convergence speed of the population are accelerated, and at the same time, the potential area of the optimal solution can be provided to improve the global convergence ability of the population; the mathematical model of the leader-guided capture strategy is expressed as follows:

[0045]

[0046]

[0047] In the formula, Leader T is the position of the leader in the T-th iteration; and are the top three individuals in terms of fishing efficiency; Exp is the error weight coefficient, and r s is the surrounding step size.

[0048] Beneficial effects: A power load forecasting method based on multi-channel deep learning with an attention mechanism of the present invention constructs a deep learning model based on a fluctuation channel and a smooth channel, which can autonomously distinguish the data fluctuation section and the smooth section to fully mine the fluctuation information in the fluctuation section and the minute change information in the smooth section. In addition, an improved fishing optimization algorithm is specifically designed to optimize the hyperparameters of the deep learning model to improve the prediction stability; enhance the comprehensive efficiency of power load forecasting and improve the system's ability to cope with the uncertainty of power load. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is the structure diagram of the load forecasting model for deep learning based on fluctuation-smooth feature fusion;

[0050] Figure 2 is the schematic diagram of the LCFOA algorithm mechanism;

[0051] Figure 3 is the schematic diagram of the predicted load results of the same prediction model using different optimization algorithms;

[0052] Figure 4 is the comparative analysis diagram of the key evaluation indicators of the predicted load results of the same prediction model with different optimization algorithms;

[0053] Figure 5 is the schematic diagram of the predicted load results of different prediction models with the same optimization algorithm;

[0054] Figure 6 is the schematic diagram of the difference between the predicted results of each model and the original data of different prediction models with the same optimization algorithm. DETAILED DESCRIPTION OF THE INVENTION

[0055] The present invention will be further described below in conjunction with the accompanying drawings.

[0056] As Figure 1 shown, a power load forecasting method based on multi-channel deep learning with an attention mechanism includes a data preprocessing operation of repairing outliers in power load data using the IQR interquartile range and cubic spline interpolation; setting a fluctuation channel and a smoothing channel according to the fluctuation characteristics and smoothing characteristics of the power load data, and constructing a load forecasting model of deep learning based on the fusion of fluctuation-smoothing features based on the fluctuation channel and the smoothing channel; using an improved leader-guided fishing optimization algorithm to optimize the hyperparameters of the load forecasting model of deep learning based on the fusion of fluctuation-smoothing features; training the load forecasting model of deep learning based on the fusion of fluctuation-smoothing features using the power load data after the data preprocessing operation; and predicting the power load using the trained load forecasting model of deep learning based on the fusion of fluctuation-smoothing features.

[0057] A data preprocessing operation of repairing outliers in power load data using the IQR interquartile range and cubic spline interpolation; detecting outliers in the power load data using the IQR interquartile range, and the calculation process of the IQR interquartile range is as follows:

[0058] IQR = Q 3 - Q 1

[0059] In the formula, Q1 is the minimum value of the lower 25% of the data points in the power load dataset, called the first quartile; Q3 is the maximum value of the upper 25% of the data points in the power load dataset, called the third quartile.

[0060] By calculating the quartiles and interquartile range of the data, the data is divided into 3 different intervals, and the 3 intervals are [-∞, LOW], [LOW, UP] and [UP, +∞], and the calculation process is as follows:

[0061] LOW = Q 1 - 1.5 * IQR

[0062] UP = Q 3 + 1.5 * IQR

[0063] When a sample value is less than LOW or greater than UP, it is determined that the sample value is an outlier, and the sample value is deleted from the power load dataset; that is, the sample values belonging to the intervals [-∞, LOW] and [UP, +∞] are determined as outliers and deleted, and the sample values belonging to the interval [LOW, UP] are normal values.

[0064] Use cubic spline interpolation to fill in the outliers deleted in the power load dataset. Define the set of interpolation points {x 0 , x 1 ,..., x n} and the corresponding function values {y 0 , y 1 ,..., y n}. Since a day contains 96 sample points, n = 96. ; x i represents the abscissa of the i-th known data point, and x i also represents any value in the set of interpolation points. y i represents the ordinate corresponding to x i , and y i also represents any value in the corresponding function values. For each interval [x i , x i + 1], define a cubic polynomial:

[0065] S i (x) = a i + b i (x - x i ) + c i (x - x i ) 2 + d i (x - x i ) 3

[0066] where a i , b i , c i , d i are all undetermined coefficients, and S i (x) is the ordinate corresponding to x i on the spline curve. The spline curve must pass through each data point (x i , y i ), that is, for each i, S i (x) = y i ;

[0067] And each x i must satisfy the following conditions:

[0068] S i-1 (x i ) = S i (x i )

[0069] S′ i-1 (x i ) = S′ i (x i )

[0070] S″ i-1 (x i ) = S″ i (x i )

[0071] When the above conditions are met, the defined cubic polynomial is not only continuous in function value at each x i , but also its first derivative and second derivative are continuous; ensuring that the spline curve is smooth throughout the data range.

[0072] Use boundary conditions to determine a i , b i , c i and d i , and the boundary conditions are natural boundary conditions, that is, the second derivative is 0 at the endpoints. The calculation process is as follows:

[0073] S′(x 0 ) = 0

[0074] S″(x n ) = 0

[0075] Through the boundary conditions, construct a system of linear equations to solve for the coefficients a i , b i , c i and d i ;

[0076] The cubic spline interpolation function S(x) is a combination of the defined cubic polynomials. In each interval [x i , x i + 1], S(x) is defined by the corresponding polynomial S i (x). The calculation process is as follows:

[0077]

[0078] Through the abscissa x of the missing outliers, obtain the corresponding S(x), which is the substitute value of the missing outliers. Delete the outliers in the power load data through the IQR interquartile range and cubic spline interpolation, and fill in the calculated substitute value to obtain the power load data after the preprocessing operation. Implementing outlier processing through IQR and Spline can provide high-quality and highly correlated power load data input for the power load prediction model.

[0079] The core of the Long Short-Term Memory (LSTM) network is the cell state, which is similar to a memory storage unit and can store both long-term and short-term memories simultaneously. In the cell state, long-term memories are updated slowly and are mainly used to preserve long-term dependency information, while short-term memories are updated rapidly to reflect recent changes. To effectively manage the cell state, three recurrent gate units are set in the hidden layer of the LSTM, namely the forget gate unit, the input gate unit, and the output gate unit. These gate units are responsible for controlling the preservation, writing, and reading of information respectively, ensuring that important information is retained while filtering out unimportant information. Although LSTM performs well in processing long sequences, it is difficult to accurately capture the dependency relationships of all key information when faced with extremely complex or high-noise data. To overcome this limitation, the attention mechanism is introduced, which not only enhances the model's sensitivity to important information but also improves the model's ability to process complex data.

[0080] Although the LSTM network with attention mechanism (LSTM-Attention) has significant advantages in processing long-term sequence data, it cannot take into account the fluctuation characteristics in the sequence. Especially when predicting complex data with large fluctuations, it faces great challenges. Usually, in relatively smooth data intervals, LSTM-Attention can accurately depict the load curve. However, in intervals with large fluctuations, the prediction accuracy is poor and it cannot provide accurate prediction results. Therefore, to solve the problem of low accuracy of the LSTM-Attention model in predicting the smooth and fluctuating segments of load data, a load prediction model based on the fusion of fluctuation-smoothing features in deep learning (FSD-LSTM-Attention) is proposed.

[0081] The LSTM layer of the load prediction model based on the fusion of fluctuation-smoothing features in deep learning consists of two LSTM channels. The channel composition for processing data at each moment is determined by the volatility n. One group is the fluctuation channel, and the other group is the smoothing channel. In the calculation process of the volatility n, the fluctuation area is defined as the integral of the load curve over the current data interval, denoted as A truth :

[0082] A truth = ∫x(t)dt

[0083] Define the test area A test as the comparison area for judging volatility. When there is no fluctuation in the data, the load curve is approximately a straight line. Select the first data x 0 in the i-th interval as the starting point, and the last data x NTaking the end point, calculate the area of the triangle formed by this section of the load curve with respect to the X-axis as the test area. The calculation process is as follows:

[0084] A test =(x 0 +x N )×N / 2

[0085] The volatility η, an index representing the volatility of load data, is defined as:

[0086]

[0087] When the volatility of the interval is small, the value of η approaches 0, and this load data is classified into the smooth channel group; when the volatility of the interval is large, the absolute value of η approaches 1, and this load data is classified into the fluctuating channel group.

[0088] The fluctuating channel is used to process load data segments with large volatility, capture the fluctuation characteristics of the data, and through the processing of specific forget gates, input gates, and output gates, output the hidden layer state containing fluctuation information

[0089]

[0090] In the formula, F is the parameter of the LSTM in the fluctuating channel; f t F is the output of the forget gate of the LSTM in the fluctuating channel, is the output of the input gate of the LSTM in the fluctuating channel, is the output of the output gate of the LSTM in the fluctuating channel, is the input time series at time t of the LSTM in the fluctuating channel, is the value of the memory cell at the previous time t - 1 of the LSTM in the fluctuating channel, is the value of the information storage unit at the current time t of the LSTM in the fluctuating channel.

[0091] The smooth channel is used to process load data segments with small volatility and relatively smooth characteristics, capture the continuous information and tiny change characteristics of the data, and through the processing of specific forget gates, input gates, and output gates, output the hidden layer state that preserves the smooth information

[0092]

[0093]

[0094] In the formula, S is the parameter of the LSTM in the smooth channel; f t S is the output of the forget gate of the LSTM in the smooth channel, is the output of the input gate of the smooth-channel LSTM, is the output of the output gate of the smooth-channel LSTM, is the input time series at time t of the smooth-channel LSTM, is the value of the memory cell at the previous time t-1 of the smooth-channel LSTM, is the value of the information storage unit at the current time t of the smooth-channel LSTM.

[0095] The fluctuation channel can improve the focusing ability of the LSTM network on data with significant fluctuation characteristics at the current moment, while the smooth channel improves the ability of the LSTM network to focus on the subtle change characteristics of the load sequence; the fluctuation characteristics extracted by the fluctuation channel and the smooth characteristics extracted by the smooth channel are fused by an intelligent fusion gate to obtain the fluctuation-smooth characteristics.

[0096] The power load data after preprocessing operations is used as the input data for training and input into the input layer of the load prediction model FSD-LSTM-Attention based on the fusion of fluctuation-smooth characteristics in deep learning; after passing through two LSTM layers, the volatility of the power load data is judged respectively. The load data with large volatility is used to extract fluctuation characteristics through the fluctuation channel, and the load data with small volatility is used to extract smooth characteristics through the smooth channel; the fluctuation-smooth characteristics obtained by fusing the fluctuation characteristics and the smooth characteristics by an intelligent fusion gate pass through the attention mechanism layer, and the weights of the attention mechanism are obtained and placed on the task of predicting the load data. Finally, it is output to the fully connected layer, and the data of the initial input layer is transmitted back to the fully connected layer in reverse to obtain the predicted result; the LCFOA algorithm is used to optimize the hyperparameters in the load prediction model FSD-LSTM-Attention based on the fusion of fluctuation-smooth characteristics in deep learning, and finally the final predicted load result is output by the output layer. The trained load prediction model FSD-LSTM-Attention based on the fusion of fluctuation-smooth characteristics in deep learning can predict the power load data of a certain day more accurately and efficiently.

[0097] The Fish-catching Optimization Algorithm (CFOA) is a new meta-heuristic optimization algorithm based on human behavior. It simulates the behavior of fishermen catching fish in a pond to achieve algorithm search and parameter optimization. CFOA mainly consists of two stages: the exploration stage and the exploitation stage. In the exploration stage, search is conducted based on individual catching behavior and collective catching behavior respectively. In the exploitation stage, it simulates fishermen surrounding the fish school and conducts exploitation based on the collective fishing strategy. The CFOA algorithm has poor search ability in the early exploration stage, is prone to falling into local optimal solutions, and cannot search for the position of the global optimal solution, resulting in a significant reduction in its exploitation ability. Therefore, a Leader-guidance-based CFOA (LCFOA) is proposed.

[0098] As Figure 2 shown, the Leader-guidance-based Fish-catching Optimization Algorithm improves the quality of the initial population through the Logistic-tent chaotic mapping method, enhancing the possibility of the population finding the global optimal solution; it introduces the leader-guidance capture strategy, that is, taking the individual with the best fishing situation as a reference to guide the position update direction of other individuals, thereby accelerating the search speed of the population.

[0099] Initialize the population using the Logistic-tent chaotic mapping to initialize the first population of fishermen, Fisher1.

[0100]

[0101] In the formula, i is the population sequence number, and the value range is i ∈ {2, 3, 4... N}; r x is the control parameter, r x ∈ (0, 4); mean represents the average value of Fisher i-1 .

[0102] The core of the Leader-guidance-based Fish-catching Optimization Algorithm is to improve the method of randomly selecting a reference fisherman in the original algorithm to guiding the population position update through the 3 individuals with the best fishing efficiency in the fisherman group; guiding through the individual with the best fishing efficiency can accelerate the convergence accuracy and speed of the population, and at the same time can provide a potential area for the optimal solution, improving the global convergence ability of the population. The mathematical model of the leader-guidance capture strategy is expressed as follows:

[0103]

[0104]

[0105] In the formula, Leader T is the position of the leader in the T-th iteration; and The top three individuals in terms of fishing efficiency; Exp is the error weight coefficient, r s is the surrounding step size, s is a random number in the range [0, 1], and R is a random number in the range [0, 1]. is the leader-guided fishing optimization algorithm, is the original fishing optimization algorithm. The leader-guided fishing optimization algorithm obtained by calculation is used to find the optimal hyperparameters in the load prediction model FSD-LSTM-Attention based on the fusion of fluctuation-smoothing features of deep learning.

[0106] The mean squared error MSE, coefficient of determination R 2 , mean absolute percentage error MAPE, and mean absolute error MAE are used to comprehensively evaluate the power load prediction results; the load prediction process based on the FSD-LSTM-Attention model is as follows:

[0107] Outlier data in the power load dataset is removed through the interquartile range (IQR), and the data is complemented based on cubic spline interpolation; the parameters of the FSD-LSTM-Attention and LCFOA models are initialized, and the initialized parameters are shown in Table 1; the hyperparameters of FSD-LSTM-Attention are optimized through the LCFOA algorithm to improve the prediction stability of the power load; on the premise of the optimal hyperparameters, the load value for the next day is predicted through FSD-LSTM-Attention; the prediction performance of the proposed model is comprehensively evaluated through MAE, MAPE, MSE, and R 2

[0108] Table 1 Parameter settings of FSD-LSTM-Attention and LCFOA models

[0109]

[0110]

[0111] ​In Table 1, Input Dimension is the input dimension, which refers to the number of features of the input data or the dimension of the input vector; Output Dimension is the output dimension, which refers to the number of features or the dimension of the vector output by the model; numEpochs is the number of training epochs, which refers to the number of times the entire training dataset is traversed; MinBatchSize is the minimum batch size, and the batch size refers to the number of samples used in one parameter update; Number Of LSTM Layers is the number of LSTM layers, which refers to the number of LSTM units stacked in the LSTM network; Optimizer is the optimizer, and Loss Function is the loss function.

[0112] Example 1

[0113] Under the condition that the prediction model remains unchanged, that is, all using the FSD-LSTM-Attention model of the present invention, the original CFOA algorithm, GWO, and GWCA are respectively used as the comparison algorithms of the LCFOA algorithm in this solution to search for the hyperparameters of the FSD-LSTM-Attention model to test the effectiveness and feasibility of the LCFOA algorithm in the field of power load forecasting.

[0114] As Figure 3 shown, compared with other comparison algorithms, Figure 3 (a) It can be seen that the model based on the LCFOA algorithm has significant advantages in the accuracy of load forecasting; specifically, compared with the FSD-LSTM-Attention model based on the CFOA algorithm, the FSD-LSTM-Attention model based on the GWO algorithm, and the FSD-LSTM-Attention model based on the GWCA algorithm; the load forecasting curve obtained by the FSD-LSTM-Attention model based on the LCFOA algorithm in this solution has a higher degree of fit with the actual load curve. Figure 3 (b) It can be seen that the average value of the prediction error of the model based on the LCFOA algorithm is closest to 0, and the violin plot is the flattest, indicating that the prediction results are relatively concentrated and there are fewer outliers; from Figure 3 (c) It can be seen that the prediction results of LSTM-Attention based on the LCFOA algorithm basically fall on the 1:1 line, indicating that the prediction results are closest to the true values, and the scatter plot has no discontinuity and the distribution is optimal.

[0115] As Figure 4As shown, a comparative analysis on four key evaluation metrics is presented; the results indicate that the FSD-LSTM-Attention model based on the LCFOA algorithm exhibits the lowest error value on the MSE metric, suggesting that its prediction accuracy is significantly superior to other models. Additionally, the FSD-LSTM-Attention model based on the LCFOA algorithm reaches the highest value on the R2 metric, further confirming its superiority in data fitting. Figure 4 (e), Figure 4 (f), Figure 4 (g) and Figure 4 (h) further confirm the competitiveness of the model based on the LCFOA algorithm in fitting accuracy by comparing the fitting results of the load prediction values of the four algorithms with the actual load values.

[0116] Table 2 Comparison of Evaluation Metrics of Load Prediction Models with Different Optimization Algorithms

[0117]

[0118]

[0119] As can be seen from Table 2, the R2 value of LCFOA is 99.90%, which is 1.9%, 1.1%, and 0.8% higher than those of other algorithms respectively, indicating a very high degree of fit between its prediction curve and the actual load curve. At the same time, the FSD-LSTM-Attention model based on the LCFOA algorithm reduces by approximately 50% on the MSE compared to other algorithms, with relatively high prediction accuracy.

[0120] Example 2

[0121] To verify the improvement effect of the proposed FSD-LSTM-Attention model, under the condition that the optimization algorithm is the LCFOA algorithm, a comparison is made with the load prediction results of CNN and BIGRU; the power load prediction results obtained by each model are as Figure 5 shown.

[0122] As Figure 5 shown, Figure 5 (a) Comparing the prediction results of each model, specifically, the FSD-LSTM-Attention model demonstrates excellent prediction ability. When there are significant fluctuations in the power load, the prediction curve of the FSD-LSTM-Attention model can still depict the actual load trajectory and accurately capture every key node of the load change. Figure 5(b) The rose diagram compares the evaluation indicators of the three models. The arc radius of the index R2 corresponding to the FSD-LSTM-Attention model is the longest, indicating the best fitting effect; the arcs of MAE and RMSE are the shortest, indicating that the proposed model has the smallest prediction error.

[0123] As Figure 6 shown, it focuses on showing the differences between the prediction results of each model and the original data. From Figure 6 (a), it can be seen that the LCFOA-FSD-LSTM-Attention model has more samples with the same value, its prediction result is closer to the original data, and the prediction accuracy is the highest. Figure 6 (b), Figure 6 (c) and Figure 6 (d) are box-violin diagrams. By comparing these three diagrams, it can be analyzed that the average predicted value of the LCFOA-FSD-LSTM-Attention model is 10758.58, which is closest to the average value of the original data, 10758.47. This shows that the LCFOA-FSD-LSTM-Attention model has a relatively high accuracy in the overall prediction trend.

[0124] Table 3 Comparison of prediction indicators of different models

[0125]

[0126] As can be seen from Table 3, the LCFOA-FSD-LSTM-Attention model shows the most competitive performance in terms of prediction error and direction accuracy. The LCFOA-FSD-LSTM-Attention model is superior to other models in all evaluation indicators, demonstrating strong competitive ability. Specifically, the R 2 of the LCFOA-FSD-LSTM-Attention model exceeds 99%, indicating that the model shows strong data fitting ability; at the same time, the MAPE is controlled below 0.3%, indicating that the model can accurately depict the complex power load prediction curve. Compared with LCFOA-CNN, the MSE and MAE of the LCFOA-FSD-LSTM-Attention model are reduced by 96.75% and 87.02% respectively; compared with LCFOA-BIGRU, the MSE and MAE of the LCFOA-FSD-LSTM-Attention model are reduced by 93.25% and 77.89% respectively.

[0127] In response to the challenges in electric load forecasting, especially the uncertainty and volatility problems brought about by the integration of renewable energy, this paper proposes a novel deep learning model based on dual-channel LSTM and attention mechanism. In Case 1, the average value and standard deviation of LCFOA are 3.38ⅹ10-223, and its convergence result is closest to the theoretical optimal value of 0. In Case 3, the R2 of the proposed model exceeds 99.9% and the MAPE is controlled below 0.3%, indicating the effectiveness and strong prediction performance of the proposed model in the field of load forecasting.

[0128] The above is only a description of the preferred embodiments of the present invention. Those of ordinary skill in the art can make several modifications and optimizations based on the above disclosure without departing from the basic principle content. These improvements and optimizations should be regarded as the protection scope understood by the present invention.

Claims

1. A power load forecasting method based on multi-channel deep learning of attention mechanism, characterized by: This includes data preprocessing operations to repair outliers in power load data using IQR interquartile range and cubic spline interpolation; setting fluctuation channels and smoothing channels based on fluctuation characteristics and smoothing characteristics of power load data, and building a load forecasting model based on deep learning based on fluctuation-smoothness feature fusion based on fluctuation channels and smoothing channels; An improved leader-guided fishing optimization algorithm is used to optimize the hyperparameters of the load forecasting model based on deep learning with fluctuation-smooth feature fusion; the load forecasting model based on deep learning with fluctuation-smooth feature fusion is trained using power load data after data preprocessing; The power load is predicted using a trained deep learning load forecasting model based on fluctuation-smooth feature fusion.

2. The power load forecasting method based on multi-channel deep learning of attention mechanism according to claim 1 is characterized in that: The IQR interquartile range and cubic spline interpolation are used to perform data preprocessing operations on the power load data to repair outliers; the IQR interquartile range is used to detect outliers in the power load data. The IQR interquartile range calculation process is as follows: IQR=Q3-Q1 Where Q1 is the minimum value of the lower 25% of the data points in the power load data set, called the first quartile; Q3 is the maximum value of the upper 25% of the data points in the power load data set, which is called the third quartile; By calculating the quartiles and interquartile ranges of the data, the data is divided into three different intervals. The calculation process is as follows: LOW=Q1-1.5*IQR UP=Q3+1.5*IQR When a sample value is less than LOW or greater than UP, the sample value is determined to be an abnormal value and is deleted from the power load data set.

3. The power load forecasting method based on multi-channel deep learning of attention mechanism according to claim 2 is characterized in that: Cubic spline interpolation is used to fill in the outliers deleted in the power load data set, and the interpolation point set {x0, x1, ..., x n ) and the corresponding function values ​​{y0, y1, ..., y n }, x i represents the horizontal coordinate of the i-th known data point, and yi represents the horizontal coordinate of x i The corresponding vertical coordinate; For each interval [x i , x i +1], define a cubic polynomial: S i (x)=a i + b i(x-x i )+c i (x-x i ) 2 +d i (x-x i ) 3 In the formula, a i , b i , c i , d i are all undetermined coefficients, S i (x) is the x on the spline curve i Corresponding to the ordinate, the spline must pass through each data point (x i ,y i ), that is, for each i, there is S i (x) = y i ; And every x i The following conditions are met: S i-1 (x i )=S i (x i ) S′ i-1 (x i )=S′ i (x i ) S″ i-1 (x i )=S″ i (x i ) When the above conditions are met, the defined cubic polynomial is i Not only is the function value continuous, but its first-order derivative and second-order derivative are also continuous.

4. The power load forecasting method based on multi-channel deep learning of attention mechanism according to claim 3 is characterized in that: Use boundary conditions to determine a i , b i , c i and d i , the boundary condition is a natural boundary condition, that is, the second-order derivative is 0 at the endpoint, and the calculation process is as follows: S′(x0)=0 S″(x n )=0 By using boundary conditions, a linear system of equations is constructed to solve the coefficient a i , b i , c i and d i ; The cubic spline interpolation function S(x) is a combination of cubic polynomials defined in each interval [x i , x i +1], S(x) is given by the corresponding polynomial S i (x) is defined and the calculation process is as follows: Through the horizontal coordinate x of the missing outlier value, we get its corresponding S(x), which is the substitute value of the missing outlier value. The outliers in the power load data are deleted through the IQR interquartile range and cubic spline interpolation, and the calculated substitute value is filled in to obtain the power load data after preprocessing.

5. The power load forecasting method based on multi-channel deep learning of attention mechanism according to claim 1 is characterized in that: The load forecasting model based on deep learning of fluctuation-smooth feature fusion determines the channel composition of the data at each moment through the fluctuation rate n, one group is the fluctuation channel, and the other group is the smooth channel; in the calculation process of the fluctuation rate n, the fluctuation area is defined as the integral of the load curve over the current data interval, denoted as A truth : A truth =∫x(t)dt Define the test area A test To determine the comparison area of ​​volatility, when there is no fluctuation in the data, the load curve is approximately a straight line. The first data x0 of the i-th interval is selected as the starting point and the last data x N As the end point, calculate the area of ​​the triangle enclosed by the load curve relative to the X-axis as the test area. The calculation process is as follows: A test =(x0+x N )×N / 2 The index volatility η that characterizes the volatility of load data is: When the interval volatility is small, the value of η approaches 0; when the interval volatility is large, the absolute value of η approaches 1.

6. The power load forecasting method based on multi-channel deep learning of attention mechanism according to claim 5 is characterized in that: The volatility channel is used to process load data segments with large volatility, capture the volatility characteristics of the data, and output the hidden layer state containing the volatility information through the processing of specific forget gates, input gates and output gates. Where, F is the parameter of the LSTM of the fluctuation channel; The smoothing channel is used to process relatively smooth load data segments, capture the continuous information and slight change characteristics of the data, and output the hidden layer state that preserves the smoothing information through the processing of specific forget gates, input gates, and output gates. Where S is the parameter of the smooth channel LSTM.

7. The power load forecasting method based on multi-channel deep learning of attention mechanism according to claim 1 is characterized in that: The leader-guided fishing optimization algorithm improves the quality of the initialization population through the Logistic-tent chaotic mapping method, and enhances the possibility of the population finding the global optimal solution; introduces the leader-guided capture strategy, that is, the individual with the best fishing situation is used as a reference to guide the position update direction of other individuals; The Logistic-tent chaotic mapping is used to initialize the population, and the first fisher population Fisher1 is initialized; Where i is the population sequence number, and its value range is i∈{2,3,4L N}; r x is the control parameter, r x ∈(0,4).

8. The power load forecasting method based on multi-channel deep learning of attention mechanism according to claim 7 is characterized in that: The core of the leader-guided fishing optimization algorithm is to improve the method of randomly selecting reference fishermen in the original algorithm to guide the population position update through the three individuals with the best fishing benefits in the fisher group; guiding through the individuals with the best fishing benefits can accelerate the convergence accuracy and convergence speed of the population, and at the same time can provide the potential area of ​​the optimal solution and improve the global convergence ability of the population; the mathematical model of the leader-guided capture strategy is expressed as follows: In the formula, Leader T is the position of the leader in the Tth iteration; and are the top three individuals in terms of fishing efficiency; Exp is the error weight coefficient, r s is the encirclement step length.

Citation Information

Patent Citations

  • Regional power distribution network short-term load prediction method

    CN112116144A

  • Capacity configuration method for photo-thermal power station of wind-solar new energy base

    CN114825381A

  • Power load prediction method based on VMD and improved grey wolf algorithm

    CN117394342A

  • Industrial short-term power load interval prediction method, device, equipment and medium

    CN118211719A

  • Power system vulnerability assessment method based on node transient voltage fluctuation index

    CN118693807A

Cited By

  • Heterogeneous computing power platform load prediction method

    CN120523608A