Short-term load prediction method based on multi-scale space-time DCNN feature extraction

By adopting a multi-scale spatiotemporal DCNN feature extraction method in short-term load prediction, combined with recursive feature pyramids and expanded convolutions, the problem of difficulty in capturing the long-term dependence and global features of load data is solved in the existing technology, and more efficient and accurate load prediction is achieved.

CN119994875AInactive Publication Date: 2025-05-13CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Patent Information

Application Number
CN202510068314.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the long-term dependence and global characteristics in load data in short-term load prediction, and has high computational complexity and memory footprint.

Method used

A feature extraction method based on multi-scale spatiotemporal DCNN is adopted to construct multi-level feature representations through recursive feature pyramids, and wider context information is captured using expanded convolution, while timing feature modeling is performed in combination with bidirectional gating recurrent units.

Benefits of technology

It improves the accuracy of short-term load prediction, can better capture complex patterns and long-term dependencies in load data, and reduces the computational complexity and memory usage of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119994875A_ABST
    Figure CN119994875A_ABST
Patent Text Reader

Abstract

The invention discloses a short-term load prediction method based on multi-scale space-time DCNN feature extraction. The short-term load prediction method specifically comprises the following steps: S1, acquiring and preprocessing historical load data and related influence factor data of a power system; s2, carrying out correlation analysis on the load data, determining the maximum correlation time lag of a load time sequence, and further determining the size of an optimal sliding time window; s3, constructing a backbone network based on a recursive feature pyramid, adopting expansion convolution as a feature extraction network, performing time sequence feature modeling in combination with a bidirectional gating cycle unit, performing back propagation by using calculation loss in a joint loss function, and updating model parameters; and S4, predicting load data in a future short period by using the trained model, and evaluating the prediction accuracy of the model through related evaluation indexes. According to the method, a time sequence is analyzed through ACF and PACF, and a sliding time window is scientifically set. And the RFP-DCNN-BiGRU model fully combines the capabilities of the DCNN and the BiGRU in feature extraction and time series data modeling, so that more accurate short-term load prediction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system prediction, and in particular to a short-term load prediction method based on multi-scale spatiotemporal DCNN feature extraction. Background Art

[0002] In recent years, with the construction of smart grids and the development of advanced measurement systems in my country, power grid companies have saved a large amount of original power load data with multi-source, multi-state and heterogeneous characteristics. Behind these massive power data, there is richer and deeper information. In the current context of smart grids and power big data, how to use the new generation of data mining technology and artificial intelligence technology in load characteristic analysis and load forecasting to improve the power grid companies' refined load management and formulate better differentiated power supply strategies is of great significance.

[0003] Short-term power load forecasting refers to the technology of predicting the changing trend of power load in the next few hours to days. It is the basis for the power system to achieve efficient scheduling and optimize resource allocation, and provide decision support for power companies, operators and government departments. Therefore, an efficient model enables it to perform prediction calculations faster while ensuring prediction accuracy, so that the prediction results can reflect real-time load changes in a timely manner. This is very important for the operation and scheduling of the power system, and can help power companies make reasonable decisions.

[0004] At present, there are many methods for short-term load forecasting, but the commonly used method is to first use CNN and TCN to mine the features between load data to form a new feature vector, and then input the extracted feature vector into the variant LSTM for prediction. However, for complex and massive data such as load data, CNN cannot effectively capture long-term dependencies when processing time series data due to the lack of sufficient receptive field, and it takes a lot of time to extract load features. TCN performs well in capturing local features of time series data, but when processing very long time series, its computational complexity and memory usage are high, and in some cases it may not be able to effectively capture global features. Based on this, a short-term load forecasting method based on multi-scale spatiotemporal DCNN feature extraction is proposed. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a short-term load forecasting method based on multi-scale spatiotemporal DCNN feature extraction. The method first uses a recursive feature pyramid to construct a multi-level feature representation, extracting and reconstructing subtle changes in the load data layer by layer, thereby enhancing the expressiveness of the features. Subsequently, dilated convolution is applied to process these features to expand the receptive field to capture a wider range of contextual information. In this way, the model can effectively retain local features while improving the understanding of global information, thereby better coping with complex patterns in load data and ultimately improving the accuracy of short-term load forecasting.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] A short-term load forecasting method based on multi-scale spatiotemporal DCNN feature extraction includes the following steps:

[0008] (1) Obtain historical load data of the power system and multi-dimensional factors affecting the load, such as meteorological data, and pre-process the data;

[0009] (2) Perform correlation analysis on the load data to determine the maximum correlation lag of the load time series. This lag point can be used as the optimal size of the sliding time window to reconstruct the high-dimensional load features into a new low-dimensional load feature set.

[0010] (3) Construct a backbone network based on recursive feature pyramid, use dilated convolution as the feature extraction network, and combine it with bidirectional gated recurrent units for temporal feature modeling. Use the loss calculated in the joint loss function for back propagation to update the model parameters.

[0011] (4) Use the trained model to predict future short-term load data and evaluate the accuracy of the model prediction using relevant evaluation indicators.

[0012] Furthermore, the data preprocessing in step S1 includes filling missing values, correcting outliers and normalizing the original load data. The normalization formula is as follows:

[0013]

[0014] In the formula, x* is the normalized value, x is the original data, and x min is the minimum value of the data, x max is the maximum value of the data.

[0015] Furthermore, in step S2, ACF and PACF are combined to analyze the correlation time lag of the load data, and the maximum correlation time lag of the load sequence is determined by this method as the size of the sliding time window. The ACF formula is as follows:

[0016]

[0017] In the formula, ρ r is the autocorrelation coefficient of lag period k, x t is the time series value, is the mean of the time series, and T is the length of the time series. The PACF formula is as follows:

[0018]

[0019] In the formula, φ kk is the partial autocorrelation coefficient of lag period k, r k is the autocorrelation coefficient.

[0020] Furthermore, the specific steps of step S3 are:

[0021] S301. The training set first passes through a DConv / s2d1 convolution layer with a step size of 2 to extract the basic low-frequency features of the load data, reduce the computational complexity, and reduce the spatial dimension. The output feature dimension is Where C1 is the number of channels.

[0022] S302. The data is divided into two paths, where the main path directly transmits the features after dimensionality reduction without deep processing; the residual path enters the residual module for deep feature extraction.

[0023] In the residual module, the data is passed through the first DConv / s1d1 convolutional layer to extract the short-term variation characteristics in the time series load, keeping the spatial size unchanged, and outputting the features Where C2 is the new number of channels. After the second DConv / s1d2 convolutional layer, a jump connection is performed to directly add the input features to the current output features, retaining the original features and enhancing the robustness of the model. The features after the jump connection pass through an additional DConv / s1d4 convolutional layer to generate updated features

[0024]

[0025] S304. Set the feature X of the main path shortcut The deep feature X3 generated by the residual path is concatenated through the concat operation, and the new feature dimension generated after concatenation is The fused features are dimensionally adjusted through a DConv / s1d1 convolution layer to output features.

[0026] S305.CSP module output X CSP It is sent to the subsequent layers for higher-level feature extraction. The features extracted from the high-level layers By returning to the initial stage through the feedback path, the feedback characteristic X RFP Through a DConv / s1d1 convolution to adjust the dimension, and the underlying feature X CSP Through recursive updating, the underlying features are continuously integrated with high-level semantic information to generate load features with more multi-scale and deep meanings.

[0027] S306. Input the feature set extracted by RFP-DCNN into the BiGRU model, and combine historical and future information to perform time series modeling through the GRU layers in the forward and backward directions. The forward GRU gradually processes the input features from time step t=1 to t=T / 2; the backward GRU gradually processes the input features from time step t=T / 2 to t=1; finally, the hidden states of the forward and backward GRUs are concatenated to obtain the output features of the BiGRU, which now contain complete time series information.

[0028] The output features of S307.BiGRU are fed into a fully connected layer for load forecasting, and the final output is the forecast sequence The predicted load value corresponding to each time step.

[0029] S308. Constructed joint loss function L all as follows:

[0030] L all =λ1L MSE +λ2L MAPE

[0031] Among them, λ1, λ2 are adjustable parameters used to balance the contribution of each loss term, L MSE is the mean square error loss function, L MAPE is the mean absolute percentage error loss function, and finally the model is jointly optimized by minimizing the sum of losses.

[0032] Furthermore, the specific steps of step S4 are:

[0033] S401. Divide the feature set reconstructed by the sliding time window into a training set, a validation set, and a test set in a ratio of 8:1:1;

[0034] S402. Build the RFP-DCNN-BiGRU model and input the training set into the RFP-DCNN-BiGRU model. The model first uses the RFP-DCNN neural network to extract load features from the input feature set at different scales, data dynamics, and different abstract levels, and then inputs the extracted features into the BiGRU for load prediction.

[0035] S403. Input the validation set data into the model obtained by the training set. The RFP-DCNN-BiGRU model will output the prediction results, and update the model parameters by calculating the loss and performing back propagation. Continuously iterate to minimize the value of the loss function and obtain the optimized RFP-DCNN-BiGRU model.

[0036] S404. Input the test set into the trained RFP-DCNN-BiGRU model to obtain the prediction results, and evaluate the accuracy of the model prediction through relevant evaluation indicators.

[0037] S405. The relevant evaluation indicators are the root mean square error (RMSE) and the coefficient of determination (R 2 ), the formula is as follows:

[0038]

[0039] In the formula, y t , and They represent the average of the actual load value, model prediction value and actual load value at time t. The smaller the RMSE, the better the R 2 The larger the value, the smaller the error between the predicted value and the true value, and the more accurate the load forecast.

[0040] The beneficial effects of the present invention are: 1) By combining ACF and PACF to analyze the load time series, its maximum correlation time lag can be determined, so as to scientifically set the size of the sliding time window. ACF provides the overall correlation range, and PACF determines the key lag point. The combination of the two ensures that the window can capture the periodicity and long-term dependence characteristics of the load, and reduce the interference of irrelevant information, providing a reliable input feature basis for load forecasting. 2) The backbone network is constructed using a recursive feature pyramid, and a single dilated convolution path is used as the backbone convolution network for feature extraction. The feature extraction network allows the network to effectively capture feature information of different scales through a multi-level recursive structure, and by expanding the receptive field of the convolution kernel, it can extract a larger range of contextual information while keeping the size of the feature map unchanged, so as to better capture the long-term dependency characteristics. 3) BiGRU is used to predict the load. Compared with the ordinary LSTM, it can consider both past and future information at the same time, and reduces the number of parameters of the LSTM network unit, shortens the training time of the model, and can more accurately capture the complex fluctuation characteristics of the power load, which helps to achieve faster and more accurate short-term load forecasting. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a logical architecture diagram;

[0042] Figure 2 It is the structural diagram of the proposed model;

[0043] Figure 3 It is the load forecasting flow chart; DETAILED DESCRIPTION

[0044] In order to make the purpose, technical scheme and advantages of the present invention more clear, the present invention is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For ordinary technicians in this technical field, as long as various changes are within the spirit and scope of the present invention defined and determined by the attached claims, all inventions and creations using the concept of the present invention are protected.

[0045] Refer to the attached Figure 1 , a short-term load forecasting method based on multi-scale spatiotemporal DCNN feature extraction, comprising the following steps:

[0046] S1. Obtain historical load data of the power system and multi-dimensional factor data affecting the load, such as meteorological data, and pre-process the data;

[0047] S2. Perform correlation analysis on load data to determine the maximum correlation time lag of load time series, and use this lag point as the optimal size of the sliding time window to reconstruct high-dimensional load features into a new low-dimensional load feature set;

[0048] S3. Build a backbone network based on recursive feature pyramid, use dilated convolution as the feature extraction network, and combine bidirectional gated recurrent units for temporal feature modeling. Use the loss calculated in the joint loss function for back propagation to update the model parameters.

[0049] S4. Use the trained model to predict the load data in the short term in the future, and evaluate the accuracy of the model prediction through relevant evaluation indicators.

[0050] Specifically, the meteorological data in step S1 includes wind speed, wind direction, temperature, weather, etc.

[0051] Specifically, the data preprocessing in step S1 includes filling missing values, correcting outliers and normalizing the original load data. The normalization formula is as follows:

[0052]

[0053] In the formula, x* is the normalized value, x is the original data, and x min is the minimum value of the data, x max is the maximum value of the data.

[0054] Specifically, in step S2, ACF and PACF are combined to analyze the correlation time lag of the load data, and the maximum correlation time lag of the load sequence is determined by this method as the size of the sliding time window. The ACF formula is as follows:

[0055]

[0056] In the formula, ρ r is the autocorrelation coefficient of lag period k, xt is the time series value, is the mean of the time series, and T is the length of the time series. The PACF formula is as follows:

[0057]

[0058] In the formula, φ kk is the partial autocorrelation coefficient of lag period k, r k is the autocorrelation coefficient.

[0059] Specifically, the specific steps of step S3 are:

[0060] S301. The training set first passes through a DConv / s2d1 convolution layer with a step size of 2 to extract the basic low-frequency features of the load data, reduce the computational complexity, and reduce the spatial dimension. The output feature dimension is Where C1 is the number of channels.

[0061] S302. The data is divided into two paths, where the main path directly transmits the features after dimensionality reduction without deep processing; the residual path enters the residual module for deep feature extraction.

[0062] In the residual module, the data is passed through the first DConv / s1d1 convolutional layer to extract the short-term variation characteristics in the time series load, keeping the spatial size unchanged, and outputting the features Where C2 is the new number of channels. After the second DConv / s1d2 convolutional layer, a jump connection is performed to directly add the input features to the current output features, retaining the original features and enhancing the robustness of the model. The features after the jump connection pass through an additional DConv / s1d4 convolutional layer to generate updated features

[0063]

[0064] S304. Set the feature X of the main path shortcut The deep feature X3 generated by the residual path is concatenated through the concat operation, and the new feature dimension generated after concatenation is The fused features are dimensionally adjusted through a DConv / s1d1 convolution layer to output features.

[0065] S305.CSP module output X CSP It is sent to the subsequent layers for higher-level feature extraction. The features extracted from the high-level layers By returning to the initial stage through the feedback path, the feedback characteristic X RFP Through a DConv / s1d1 convolution to adjust the dimension, and the underlying feature X CSP Through recursive updating, the underlying features are continuously integrated with high-level semantic information to generate load features with more multi-scale and deep meanings.

[0066] S306. Input the feature set extracted by RFP-DCNN into the BiGRU model, and combine historical and future information to perform time series modeling through the GRU layers in the forward and backward directions. The forward GRU gradually processes the input features from time step t=1 to t=T / 2; the backward GRU gradually processes the input features from time step t=T / 2 to t=1; finally, the hidden states of the forward and backward GRUs are concatenated to obtain the output features of the BiGRU, which now contain complete time series information.

[0067] The output features of S307.BiGRU are fed into a fully connected layer for load forecasting, and the final output is the forecast sequence The predicted load value corresponding to each time step.

[0068] S308. Constructed joint loss function L all as follows:

[0069] L all =λ1L MSE +λ2L MAPE

[0070] Among them, λ1, λ2 are adjustable parameters used to balance the contribution of each loss term, L MSE is the mean square error loss function, L MAPE is the mean absolute percentage error loss function, and finally the model is jointly optimized by minimizing the sum of losses.

[0071] Specifically, build the DCNN-RFP-BiGRU model according to the attached Figure 2, the DCNN includes an expanded convolution layer, a pooling layer and a flattening layer. The expanded convolution layer expands the receptive field by inserting intervals between the convolution kernel elements, so that the local pattern and global features of the input data can be extracted simultaneously. This method effectively retains multi-scale information and enhances the model's ability to capture complex load features; the pooling layer can retain important features and reduce the data dimension by performing dimensionality reduction operations on the features of each convolution channel, thereby reducing the number of model parameters, enhancing robustness, effectively preventing overfitting and accelerating the training convergence speed of the model; the flattening layer is used to map the pooled high-dimensional features into a one-dimensional vector, thereby providing a unified data input format for the subsequent BiGRU module, and realizing the effective connection between CNN and sequence models. The RFP realizes feature interaction and fusion of multi-layer features at multiple scales through feature recursion. Specifically, low-level convolutional features and high-level convolutional features share the same receptive field through recursive updates, thereby improving the network's ability to model global and local features. The BiGRU, or bidirectional gated recurrent unit, is an extended variant of the GRU that can capture contextual information in a time series bidirectionally through forward and backward hidden state propagation. Compared with unidirectional GRU or LSTM, BiGRU can not only learn the correlation features of the previous and next moments in the load time series more comprehensively, but also has higher computational efficiency and parameter optimization capabilities, thereby significantly improving the model's prediction performance for complex load data.

[0072] Furthermore, the DCNN-RFP is a multi-scale receptive field combined with DCNN. It uses RFP embedded in DCNN as the backbone network, replaces ordinary convolution with dilated convolution to expand the receptive field, and improves the multi-scale feature extraction capability. The CSP module separates and fuses features between different stages to optimize computational efficiency while maintaining the integrity of feature transfer. The recursive mechanism further strengthens the multiple updates and refinements of features, so that the output of each stage not only depends on the input of the current layer, but also combines the historical features of the previous layers. A weighted fusion strategy is used to fuse features to effectively suppress noise interference, and global average pooling is performed on the output feature map to reduce the dimension. Finally, the feature map is input into the bidirectional gated recurrent unit (BiGRU) to realize the temporal feature modeling of load data. The formula is as follows:

[0073]

[0074] In the above formula, f i is the output feature of the i-th layer, d represents the dilation factor, n represents the filter size, td·j represents the past direction, R i (f i ) is the output after feedback and then sent to the bottom-up backbone network.

[0075] Furthermore, the BiGRU is an improved model expanded on the basis of GRU. For each training sequence, two GRU models are established in the forward and reverse directions, and the hidden layer nodes of the two models are connected to the same output layer. This data processing method can provide complete historical and future information for each time point in the output layer input sequence. Therefore, the BiGRU network has the ability to learn the relationship between past and future load influencing factors and current load, which helps to extract the characteristics of load data. Compared with BiLSTM, BiGRU has a simpler structure, fewer parameters, higher computational efficiency, and very similar performance. The formula is as follows:

[0076]

[0077] In the above formula, h t for t The hidden layer state at the moment; is the output of the forward hidden layer at time t; is the output of the reverse hidden layer at time t; w t and v t They represent the forward hidden state corresponding to the bidirectional GRU at time t. and the reverse hidden state The corresponding weight; b t Represents the bias corresponding to the hidden layer state at time t.

[0078] See attached Figure 3 , in step S4, the specific steps of constructing the prediction model are as follows:

[0079] S401. Divide the feature set reconstructed by the sliding time window into a training set, a validation set, and a test set in a ratio of 8:1:1;

[0080] S402. Build the RFP-DCNN-BiGRU model and input the training set into the RFP-DCNN-BiGRU model. The model first uses the RFP-DCNN neural network to extract load features from the input feature set at different scales, data dynamics, and different abstract levels, and then inputs the extracted features into the BiGRU for load prediction.

[0081] S403. Input the validation set data into the model obtained by the training set. The RFP-DCNN-BiGRU model will output the prediction results, and update the model parameters by calculating the loss and performing back propagation. Continuously iterate to minimize the value of the loss function and obtain the optimized RFP-DCNN-BiGRU model.

[0082] S404. Input the test set into the trained RFP-DCNN-BiGRU model to obtain the prediction results, and evaluate the accuracy of the model prediction through relevant evaluation indicators.

[0083] S405. The relevant evaluation indicators are the root mean square error (RMSE) and the coefficient of determination (R 2 ), the formula is as follows:

[0084]

[0085] In the formula, y t , and They represent the average of the actual load value, model prediction value and actual load value at time t. The smaller the RMSE, the better the R 2 The larger the value, the smaller the error between the predicted value and the true value, and the more accurate the load forecast.

[0086] Although the specific implementation of the invention is described in detail in conjunction with the drawings, it should not be understood as limiting the scope of protection of this patent. Within the scope described in the claims, various modifications and variations that can be made by those skilled in the art without creative work still fall within the scope of protection of this patent.

Claims

1. A short-term load forecasting method based on multi-scale spatiotemporal DCNN feature extraction, characterized in that: Includes steps: (1) Obtain historical load data of the power system and multi-dimensional factors affecting the load, such as meteorological data, and pre-process the data; (2) Perform correlation analysis on the load data to determine the maximum correlation lag of the load time series. This lag point can be used as the optimal size of the sliding time window to reconstruct the high-dimensional load features into a new low-dimensional load feature set. (3) Construct a backbone network based on recursive feature pyramid, use dilated convolution as the feature extraction network, and combine it with bidirectional gated recurrent units for temporal feature modeling. Use the loss calculated in the joint loss function for back propagation to update the model parameters. (4) Use the trained model to predict future short-term load data and evaluate the accuracy of the model prediction using relevant evaluation indicators.

2. The short-term load forecasting method based on multi-scale spatiotemporal DCNN feature extraction according to claim 1 is characterized in that: In step S1, data preprocessing includes filling missing values, correcting outliers and normalizing the original load data. The normalization formula is as follows: In the formula, x* is the normalized value, x is the original data, and x min is the minimum value of the data, x max is the maximum value of the data.

3. The short-term load forecasting method based on multi-scale spatiotemporal DCNN feature extraction according to claim 1 is characterized in that: In step S2, ACF and PACF are combined to analyze the correlation time lag of the load data. This method is used to determine the maximum correlation time lag of the load sequence as the size of the sliding time window. The ACF formula is as follows: In the formula, ρ r is the autocorrelation coefficient of lag period k, x t is the time series value, is the mean of the time series, T is the length of the time series, and the PACF formula is as follows: In the formula, φ kk is the partial autocorrelation coefficient of lag period k, r k is the autocorrelation coefficient.

4. The short-term load forecasting method based on multi-scale spatiotemporal DCNN feature extraction according to claim 1 is characterized in that: In step S3, the defined network structure uses RFP embedded in DCNN as the backbone network, replaces ordinary convolution with dilated convolution to expand the receptive field and improve the multi-scale feature extraction capability; separates and fuses features between different stages through the CSP module, optimizes computational efficiency while maintaining the integrity of feature transfer; The recursive mechanism further strengthens the multiple updates and refinements of features, so that the output of each stage not only depends on the input of the current layer, but also combines the historical features of the previous layers; a weighted fusion strategy is used to fuse the features, effectively suppress noise interference, and perform global average pooling on the output feature map to reduce the dimension; finally, the feature map is input into BiGRU to realize the temporal feature modeling of load data, and the constructed joint loss function L all as follows: L all =λ1L MSE +λ2L MAPE Among them, λ1, λ2 are adjustable parameters used to balance the contribution of each loss term, L MSE is the mean square error loss function, L MAPE is the mean absolute percentage error loss function, and finally the model is jointly optimized by minimizing the sum of losses. The joint loss provides a flexible framework for deep learning models that can optimize multiple tasks or aspects at the same time, thereby guiding the model to learn useful feature representations more comprehensively and effectively during the training process.

5. The short-term load forecasting method based on multi-scale spatiotemporal DCNN feature extraction according to claim 1 is characterized in that: In step S4, the relevant evaluation indicators are the root mean square error (RMSE) and the coefficient of determination (R 2 ), the formula is as follows: In the formula, y t , and They represent the average values ​​of the actual load value, model predicted value and actual load value at time t respectively.

Citation Information

Patent Citations

  • Short-term power load and carbon emission prediction method and system based on DCNN-LSTM-AE-AM

    CN115936185A

  • Lightweight DeepLabV3 + image semantic segmentation method and device

    CN116704190A

  • Short-term load prediction method based on multiple views

    CN117313039A

  • Power transmission line inspection image detection method based on deep convolutional neural network

    CN117541535A

  • CNN-GRU and ARIMA model-based power load prediction method

    CN118316033A

Cited By

  • Method for predicting energy consumption of air conditioning system of railway station

    CN121542610A