Power load prediction method based on deep learning and data driving
By adopting a DCNN and BiGRU structure with a feature weighting mechanism and a multi-layer attention mechanism in power load prediction, the problem of feature fusion and weight adjustment in the prior art is solved, and higher prediction accuracy and adaptability are achieved, providing technical support for the efficient operation of the power system.
Patent Information
- Application Number
- CN202510465571.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing power load prediction methods are difficult to efficiently integrate historical and future characteristics, dynamically adjust feature weights, and optimize long-term dependency modeling, resulting in insufficient prediction accuracy.
The expanded convolutional encoder (DCNN) with a feature weighting mechanism and a bidirectional gated cyclic unit (BiGRU) decoder structure with a multi-layer attention mechanism is adopted. The feature weights are adaptively allocated through the feature weighting mechanism, and the multi-layer attention mechanism is used to extract similar daily and time attention weights, and the information flow is dynamically adjusted.
It improves the adaptability of the model to complex load modes, improves prediction accuracy, and verifies its effectiveness through public data sets, providing technical support for the efficient operation of the power system.
Smart Images

Figure CN120016472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of short-term load forecasting for power systems, and in particular to a load forecasting method for a dilated convolutional encoder with a feature weighting mechanism and a bidirectional gated recurrent unit decoder with a multi-layer attention mechanism. Background Art
[0002] In recent years, with the rapid development of power system intelligence and data acquisition technology, power load forecasting methods have gradually become a research hotspot. Traditional load forecasting methods are difficult to fully explore the complex nonlinear relationship between high-dimensional historical data and external factors. Deep learning can effectively capture the time series characteristics and potential patterns in the data by constructing multi-layer neural networks, thereby improving prediction accuracy.
[0003] In existing research, convolutional neural networks are suitable for capturing local features, while gated recurrent units are suitable for processing long dependencies in time series. However, a single model has limitations when processing multi-time scale features or dynamically weighted input features. To this end, some scholars have proposed a solution that combines attention mechanisms. For example, models based on attention mechanisms can adaptively focus on important time steps or features, while dilated convolutional neural networks (DCNNs) can capture a wider range of temporal dependencies by expanding the receptive field.
[0004] At present, most prediction methods are implemented through feature engineering or simple network structures, but how to efficiently integrate historical and future features, dynamically adjust feature weights, and optimize long-term dependency modeling are still technical problems that need to be solved. Therefore, a structure that integrates feature weighting mechanism, DCNN encoder, and bidirectional gated recurrent unit (BiGRU) decoder with multi-layer attention mechanism is proposed, which can not only improve the adaptability of the model to complex load patterns, but also verify its prediction effect through public data sets, providing technical support for the efficient operation of power systems. Summary of the invention
[0005] A method for power load forecasting based on deep learning and data-driven, the method comprising the following operations:
[0006] First, the ISO-NE dataset is used to select hourly historical electricity load demand data.
[0007] Second, a feature weighting mechanism is applied to assign weights to the input features.
[0008] Again, develop a DCNN encoder.
[0009] Then, a multi-layer attention mechanism is used to select similar days and calculate the context vector to assist BiGRU in decoder prediction.
[0010] Finally, the entire DCNN encoder with a feature weighting mechanism and the BiGRU decoder structure with a multi-layer attention mechanism are constructed as the load forecasting model. The model is then verified using the data set, and the prediction effect of the model is judged based on the evaluation indicators.
[0011] The present invention is a power load forecasting method based on deep learning and data-driven, using the ISO-NE dataset. The input features include power load demand, dry-bulb temperature and dew point temperature features. is the historical time step, To predict future time steps, As feature dimension, separate historical features ,express The dimension is , separating out future features ,Right now The dimension is , set the historical target to ,Right now The dimension is , the historical characteristics and future features Concatenate to input matrix ,exist China Construction Indicates The characteristic dimension of the moment element is , the explanation of dimension symbols will not be repeated in the future.
[0012] The present invention applies a feature weighting mechanism to assign weights to input features, and the process specifically includes:
[0013] Step (1) For each time step ,Will The feature attention score vector is obtained by linear layer transformation: in, is the activation function, represents the linear layer weights, represents the linear layer bias.
[0014] Step (2) Refine the meaning of feature dimensions and get , indicating the Time step The attention score vector of features, ,right application Function, calculates feature attention weights: in For the The feature attention weights for each time step.
[0015] Step (3) assign feature attention weights and Multiply element by element to obtain the weighted total feature: in Represents element-by-element multiplication to separate weighted historical features and weighted future features .
[0016] The present invention develops a DCNN encoder, and the process specifically includes:
[0017] Step (1) Concatenate weighted historical features at each moment With historical goals Get Input , for each time have , after transposing, we get , used as the input of DCNN, if the input dimension Different from the initial number of channels , dimensionality reduction through a linear layer: in, Adapt the weights for the dimension reduction layer, is the dimension reduction layer bias, , the overall , after transposing, we get .
[0018] Step (2) Input the encoder Block, let the expansion factor be , , is the number of convolution blocks, is the number of expansion layers, in the block Neidi The layer output is also used as the The input of the layer is For Block The first layer input, so the The output of the layer is as follows: in, is the convolution kernel size, For Block exist The Layer output, For Block No. The convolution kernel of the layer weights, For Block No. The bias of the layer, is the activation function.
[0019] Step (3) Block The whole block output With Block The last layer output Make residual connections to get blocks The whole block output : ;
[0020] If the block Output Channel Number and Block Inconsistency, for the correction block The whole block output , through which Convolution performs projection: in, For Block Projected convolution weights, For Block Projection convolution bias, then With Block The last layer output Make residual connections to get blocks The whole block output : .
[0021] Step (4) After all convolution blocks, the total output is obtained , passing the output through the adaptation layer: in, is the adaptation layer weight, is the adaptation layer bias, thus obtaining the encoder hidden representation sequence As the DCNN output, Refers to one of them, is the number of DCNN output channels.
[0022] The present invention uses a multi-layer attention mechanism to select similar days and calculate context vectors to assist BiGRU in decoder prediction. The process specifically includes:
[0023] Step (1) weights historical features Divided by period, each period includes time steps, a total of cycles, there are , refactoring it into , while weighting the future features to reconstruct ,copy Get .
[0024] Step (2) for the cycle No. time steps, calculate the adjusted weighted historical features With weighted future features Similarity: in is the normalized periodic attention weight, which is calculated using the inverse of the similarity.
[0025] Step (3) for each prediction time step ,have , using BiGRU's previous prediction time step The forward hidden state of , backward hidden state , current weighted future features Constructing query vector , calculate the temporal attention score : in, is the temporal attention score weight, It is both the number of output channels of the aforementioned DCNN encoder and the dimension of the BiGRU hidden state. is the temporal attention score bias, The function is used to connect and normalize the current time attention distribution: in Also called temporal attention weight.
[0026] Step (4) for each cycle , using the temporal attention weights to the DCNN encoder output Calculate weighted encoder output : in To weight the information of each cycle, it is also called the context vector.
[0027] Step (5) converts the context vector With current future features Splicing to get the input of BiGRU , and update the status: in, for The forward hidden state at any moment, for The backward hidden state at the moment, The function is the operation process of the bidirectional gated recurrent unit, which is concatenated and passed through the fully connected layer to generate a prediction: in, is the activation function, is the weight of the fully connected layer, is the weight of the fully connected layer.
[0028] The present invention constructs the entire DCNN encoder with feature weighting mechanism and the BiGRU decoder structure with multi-layer attention mechanism as the load forecasting model, and obtains the final prediction sequence as .
[0029] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: (1) The present invention divides the input features into historical features and future features, adaptively allocates feature weights, and enhances the auxiliary role of features in prediction; (2) The present invention dynamically adjusts the relationship between input features and convolution channels by dimensionality reduction, and adjusts the relationship between output channels of different convolution blocks by projection; (3) The present invention uses a multi-layer attention mechanism to extract similar days, performs period and time weighting for historical information, updates the information flow of BiGRU, and improves prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a specific flow chart of a power load forecasting method based on deep learning and data-driven according to the present invention;
[0031] Figure 2 is a structural diagram of the attention mechanism in the example of the present invention to obtain the weighted encoder output;
[0032] Figure 3 It is a structural diagram of BiGRU dynamic update in an example of the present invention. DETAILED DESCRIPTION
[0033] The present invention is further described below in conjunction with the accompanying drawings. The embodiments are to more clearly illustrate the technical solution of the present invention, and the application scope of the present invention is not limited thereto and can be applied to many fields. Here is just one example.
[0034] like Figure 1 As shown, a power load forecasting method based on deep learning and data-driven is specifically divided into the following steps:
[0035] First, using the ISO-NE dataset, the input features include power load demand, dry bulb temperature and dew point temperature. is the historical time step, To predict future time steps, As feature dimension, separate historical features ,express The dimension is , separating out future features ,Right now The dimension is , set the historical target to ,Right now The dimension is , the historical characteristics and future features Concatenate to input matrix ,exist China Construction Indicates The characteristic dimension of the moment element is , the explanation of dimension symbols will not be repeated in the future.
[0036] Secondly, a feature weighting mechanism is applied to assign weights to input features. The process specifically includes:
[0037] Step (1) For each time step ,Will The feature attention score vector is obtained by linear layer transformation: in, is the activation function, represents the linear layer weights, represents the linear layer bias.
[0038] Step (2) Refine the meaning of feature dimensions and get , indicating the Time step The attention score vector of features, ,right application Function, calculates feature attention weights: in For the The feature attention weights for each time step.
[0039] Step (3) assign feature attention weights and Multiply element by element to obtain the weighted total feature: in Represents element-by-element multiplication to separate weighted historical features and weighted future features .
[0040] Next, develop the DCNN encoder, which includes:
[0041] Step (1) Concatenate weighted historical features at each moment With historical goals Get Input , for each time have , after transposing, we get , used as the input of DCNN, if the input dimension Different from the initial number of channels , dimensionality reduction through a linear layer: in, Adapt the weights for the dimension reduction layer, is the dimension reduction layer bias, , the overall , after transposing, we get .
[0042] Step (2) Input the encoder Block, let the expansion factor be , , is the number of convolution blocks, is the number of expansion layers, in the block Neidi The layer output is also used as the The input of the layer is For Block The first layer input, so the The output of the layer is as follows: in, is the convolution kernel size, For Block exist The Layer output, For Block No. The convolution kernel of the layer weights, For Block No. The bias of the layer, is the activation function.
[0043] Step (3) Block The whole block output With Block The last layer output Make residual connections to get blocks The whole block output : ;
[0044] If the block Output Channel Number and Block Inconsistency, for the correction block The whole block output , through which Convolution performs projection: in, For Block Projected convolution weights, For Block Projection convolution bias, then With Block The last layer output Make residual connections to get blocks The whole block output : .
[0045] Step (4) After all convolution blocks, the total output is obtained , passing the output through the adaptation layer: in, is the adaptation layer weight, is the adaptation layer bias, thus obtaining the encoder hidden representation sequence As the output of the DCNN encoder, Refers to one of them, is the number of DCNN output channels.
[0046] Then, a multi-layer attention mechanism is used to select similar days and calculate context vectors to assist BiGRU in decoder prediction. The process specifically includes:
[0047] Step (1) weights historical features Divided by period, each period includes time steps, a total of cycles, there are , refactoring it into , while weighting future features Refactored to ,copy Get .
[0048] Step (2) for the cycle No. time steps, calculate the adjusted weighted historical features With weighted future features Similarity: in is the normalized periodic attention weight, which is calculated using the inverse of the similarity.
[0049] Step (3) for each prediction time step ,have , using BiGRU's previous prediction time step The forward hidden state of , backward hidden state , current weighted future features Constructing query vector , calculate the temporal attention score : in, is the temporal attention weight, It is both the number of output channels of the aforementioned DCNN encoder and the dimension of the BiGRU hidden state. is the temporal attention bias, The function is used to connect and normalize the current time attention distribution: in It is also the temporal attention weight.
[0050] Step (4) is as follows Figure 2 As shown, for each cycle , using the temporal attention weights to the DCNN encoder output Calculate weighted encoder output : in To weight the information of each cycle, it is also called the context vector.
[0051] Step (5) is as follows Figure 3 As shown, the context vector With current future features Splicing to get the input of BiGRU , and update the status: in, for The forward hidden state at any moment, for The backward hidden state at the moment, The function is the operation process of the bidirectional gated recurrent unit, which is concatenated and passed through the fully connected layer to generate a prediction: in, is the activation function, is the weight of the fully connected layer, is the weight of the fully connected layer.
[0052] Finally, the entire DCNN encoder with feature weighting mechanism and the BiGRU decoder with multi-layer attention mechanism are constructed as the load forecasting model, and the final prediction sequence is obtained as follows: .
[0053] In the examples of the present invention, the mean absolute percentage error (MAPE [%]) and the mean absolute mean error (NRMSE [%]) are selected for verification, and the formulas for the two are: In the formula, yes The load forecast value at the time, yes The actual value of the load at the moment, is the minimum value of the true load value, is the maximum value of the true value of the load, To predict future time steps.
[0054] The prediction results are shown in Table 1. When other model conditions are the same, the model proposed in the present invention is compared with the CNN, BiGRU and CNN-BiGRU models to predict the average load in the next 24 hours, and the evaluation index is used to judge the effectiveness of the model.
[0055] Table 1 Model evaluation index results table: Model Category MAPE (%) NRMSE(%) CNN 4.644 6.733 BiGRU 4.317 6.327 CNN-BiGRU 5.156 7.108 Prediction model of the present invention 2.653 3.634 From the results in Table 1, it can be seen that the effect achieved by the prediction model of the present invention is optimal, which can prove the benefits of the present invention.
Claims
1. A power load forecasting method based on deep learning and data-driven, characterized in that: The method comprises the following operations: first, using the ISO-NE dataset, selecting its hourly historical power load demand data; second, applying a feature weighting mechanism to assign weights to input features; Secondly, develop a dilated convolutional encoder; then, use a multi-layer attention mechanism to select similar days and calculate context vectors to assist the bidirectional gated recurrent unit in decoder prediction; finally, construct the entire dilated convolutional encoder with a feature weighting mechanism and a bidirectional gated recurrent unit decoder with a multi-layer attention mechanism as a load forecasting model, and then use the data set to verify the model, and judge the prediction effect of the model based on the evaluation indicators. The specific process is as follows: First, using the ISO-NE dataset, the input features include power load demand, dry bulb temperature and dew point temperature. is the historical time step, To predict future time steps, As feature dimension, separate historical features ,express The dimension is , separating out future features ,Right now The dimension is , set the historical target to ,Right now The dimension is , the historical characteristics and future features Concatenate to input matrix ,exist China Construction Indicates The characteristic dimension of the moment element is , the explanation of dimension symbols will not be repeated in the future; Secondly, a feature weighting mechanism is applied to assign weights to input features. The process specifically includes: Step (1) For each time step ,Will The feature attention score vector is obtained by linear layer transformation: in, is the activation function, represents the linear layer weights, represents the linear layer bias; Step (2) Refine the meaning of feature dimensions and get , indicating the Time step The attention score vector of features, ,right application Function, calculates feature attention weights: in For the Feature attention weights for time steps; Step (3) assign feature attention weights With input Multiply element by element to get the weighted total feature: in Represents element-by-element multiplication to separate weighted historical features and weighted future features ; Next, we develop a dilated convolutional encoder, which includes: Step (1) Concatenate weighted historical features at each moment With historical goals Get Input , for each time have , after transposition, we get the dilated convolution encoder input ; Step (2) Input the encoder Block, let the expansion factor be , , is the number of convolution blocks, is the number of expansion layers, in the block Neidi The layer output is also used as the The input of the layer is For Block The first layer input, so the The output of the layer is as follows: in, is the convolution kernel size, For Block exist The Layer output, For Block No. The convolution kernel of the layer weights, For Block No. The bias of the layer, is the activation function; Step (3) Block The whole block output With Block The last layer output Make residual connections to get blocks The whole block output : Step (4) After all convolution blocks, the total output is obtained , passing the output through the adaptation layer: in, is the adaptation layer weight, is the adaptation layer bias, thus obtaining the encoder hidden representation sequence As the output of the dilated convolutional encoder, Refers to one of them, is the number of dilated convolution output channels; Then, a multi-layer attention mechanism is used to select similar days and calculate the context vector to assist the bidirectional gated recurrent unit in decoder prediction. The process specifically includes: Step (1) weights historical features Divided by period, each period includes time steps, a total of cycles, there are , refactoring it into , while weighting future features Refactored to ,copy Get ; Step (2) for the cycle No. time steps, calculate the adjusted weighted historical features With weighted future features Similarity: in is the normalized periodic attention weight, which is calculated using the inverse of the similarity; Step (3) for each prediction time step ,have , using the bidirectional gated recurrent unit to predict the previous time step The forward hidden state of , backward hidden state , current weighted future features Constructing query vector , calculate the temporal attention score : in, is the temporal attention score weight, It is both the number of output channels of the aforementioned diffusion convolutional encoder and the hidden state dimension of the bidirectional gated recurrent unit. is the temporal attention score bias, Function is used to connect and normalize the current time attention distribution : in Also called temporal attention weight; Step (4) for each cycle , using the temporal attention weights to the output of the dilated convolutional encoder Calculate weighted encoder output : in To weight the information of each cycle, it is also called context vector; Step (5) converts the context vector With the current weighted future features Splicing to obtain the input of the bidirectional gated recurrent unit , and update the status: in, for The forward hidden state at any moment, for The backward hidden state at the moment, The function is the operation process of the bidirectional gated recurrent unit, which is concatenated and passed through the fully connected layer to generate a prediction: in, is the activation function, is the weight of the fully connected layer, Bias for the fully connected layer; Finally, the entire dilated convolutional encoder with feature weighting mechanism and the bidirectional gated recurrent unit decoder with multi-layer attention mechanism are constructed as the load forecasting model, and the final prediction sequence is obtained as follows: .
2. The method for power load forecasting based on deep learning and data-driven according to claim 1, characterized in that: The developed dilated convolutional encoder stage concatenates weighted historical features at each moment With historical goals Get Input , for each time have , is the historical time step, if the input dimension Different from the initial number of channels , dimensionality reduction through a linear layer: in, Adapt the weights for the dimension reduction layer, is the dimension reduction layer bias, , the overall , after transposing, we get , used as input to the dilated convolutional encoder.
3. The method for power load forecasting based on deep learning and data-driven according to claim 1, characterized in that: The developed dilated convolutional encoder stage, if the block Output Channel Number and Block Inconsistency, for the correction block The whole block output , through which Convolution performs projection: in, For Block Projected convolution weights, For Block Projection convolution bias, then With Block The last layer output Make residual connections to get blocks The whole block output : , in order to calculate the total output of all convolutional blocks.
Citation Information
Patent Citations
Sequence-to-sequence power load prediction method based on multi-dimensional gating circulation unit
CN116384572A
Short-term load prediction method based on self-attention encoder and time convolutional network
CN118554422A