IES multi-element load prediction method based on VMD and combined deep neural network

By applying a multi-load prediction method based on VMD and combined deep neural networks in an integrated energy system, the problem of insufficient multi-load prediction accuracy in the prior art is solved, and high-precision prediction of IES multi-load is achieved, especially when dealing with complex load characteristics and strong coupling.

CN120045910APending Publication Date: 2025-05-27CHINA THREE GORGES PROJECTS DEV CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510043244.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art lacks prediction accuracy in multi-load prediction of integrated energy systems, especially when dealing with complex IES load characteristics and strong coupling between multi-loads, traditional methods and shallow neural networks are difficult to meet the accuracy requirements.

Method used

The IES multivariate load prediction method based on VMD and combined deep neural network is adopted to decompose the multivariate load through VMD to obtain multiple subsequences of different frequency scales; then the time convolution network and graph convolution network are used to extract the timing characteristics and coupling characteristics of the multivariate data, and combine the long-term and short-term memory network to capture the long-term dependence relationship of the multivariate data to achieve the prediction of the IES multivariate load.

Benefits of technology

By reducing the non-stationarity of the original load data, the coupling characteristics between multiple data are fully explored, and the load prediction accuracy is improved, especially in the integrated energy system, which shows excellent multi-load prediction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045910A_ABST
    Figure CN120045910A_ABST
Patent Text Reader

Abstract

The invention discloses an IES multivariate load prediction method based on a VMD and a combined deep neural network, and the method comprises the steps: employing a VMD module to carry out the decomposition of IES multivariate complex load original data, and obtaining a plurality of subsequences of different frequency scales; a TCN time convolution module and a GCN graph convolution module are combined to obtain a plurality of spatial-temporal feature extraction modules, and a spatial-temporal feature extraction network formed by the plurality of spatial-temporal feature extraction modules is adopted to extract spatial-temporal features of a plurality of subsequences with different frequency scales; capturing a long-term dependency relationship of multivariate data in a plurality of subsequences with different frequency scales by adopting an LSTM network; and fusing the long-term dependency relationship of the multivariate data in the plurality of subsequences with different frequency scales and the spatial-temporal characteristics of the plurality of subsequences with different frequency scales through point multiplication to obtain a multivariate load prediction result. While the non-stationarity of the load data is reduced, the characteristics of the multivariate data on each layer are fully mined, and the method has excellent prediction precision in the multivariate load prediction of the integrated energy system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multivariate load forecasting for integrated energy systems, and in particular to an IES multivariate load forecasting method based on VMD and a combined deep neural network. Background Art

[0002] At present, the methods used for load forecasting are mainly divided into traditional methods and machine learning-based methods. However, as the IES load characteristics become increasingly complex, the prediction accuracy of traditional methods is limited and can no longer meet the accuracy requirements. Load forecasting based on machine learning is an effective method to improve prediction accuracy, which includes shallow neural networks and deep learning; however, due to its limited feature extraction capabilities, shallow neural networks usually require human experience to assist and are not suitable for complex systems such as IES.

[0003] It is worth noting that the general load forecasting transfers the traditional load forecasting ideas of the power system to the multi-load forecasting of the integrated energy system, and does not take into account the strong coupling between multiple loads such as cold, hot, and electric; some machine learning methods are used to extract the characteristics of a single load, build a multi-task learning framework, and use the sharing mechanism to learn the coupling information provided by different subtasks. The results show that its prediction accuracy is higher than that of the single-task prediction model; but the sharing mechanism of the multi-task framework can only learn limited coupling information and cannot fully extract the coupling information between multiple loads. The correlation coefficients between multiple loads are calculated using statistical methods, and the obtained coupling feature matrix is ​​used as part of the input to quantitatively describe the coupling between loads, but the construction of the coupling feature matrix increases the difficulty of model analysis; the non-stationarity and strong coupling of IES multi-load data are not considered. Summary of the invention

[0004] Purpose of the invention: In order to overcome the shortcomings of the prior art, the present invention provides an IES multivariate load forecasting method based on VMD and combined deep neural network, which uses VMD to decompose the multivariate load to obtain the multivariate load components, extracts the time series characteristics of the multivariate data through the time convolution network and the graph convolution network, and mines the coupling characteristics between the multivariate data; uses the long short-term memory network to capture the long-term dependency of the multivariate data, so as to realize the prediction of the IES multivariate load.

[0005] Technical solution: To achieve the above object, the IES multi-load prediction method based on VMD and combined deep neural network includes using a VMD variational mode decomposition module to decompose the original IES multi-complex load data, and decomposing it into multiple subsequences with different frequency scales; combining a TCN temporal convolutional module and a GCN graph convolutional module to obtain several spatio-temporal feature extraction modules, and using a spatio-temporal feature extraction network composed of several spatio-temporal feature extraction modules to extract the spatio-temporal features of multiple subsequences with different frequency scales; using an LSTM network to capture the long-term dependence relationship of the multi-source data in multiple subsequences with different frequency scales; fusing the long-term dependence relationship of the multi-source data in multiple subsequences with different frequency scales and the spatio-temporal features of multiple subsequences with different frequency scales through dot product to obtain the result of multi-load prediction.

[0006] Further, the VMD variational mode decomposition module decomposes the original IES multi-complex load data into K modal components v k (t) of multi-loads, that is, multiple subsequences with different frequency scales; using the obtained multiple subsequences with different frequency scales as the input data of the spatio-temporal feature extraction network and the input data of the LSTM network;

[0007] Obtain the analytical signal of v k (t) through Hilbert transform, and modulate the central band of v k (t) to the corresponding baseband, estimate the bandwidth using the Gaussian smoothness of the demodulated signal, and the constrained variational problem is expressed as:

[0008]

[0009] In the formula, {v k} = {v 1 ,.., v k} is the modal component after VMD decomposition, ω k = {ω 1 ,.., ω k} is the central frequency of each component, δ(t) is the impulse function, is the gradient calculation, s(t) is the original signal;

[0010] Introduce a quadratic penalty factor and a Lagrange multiplier to transform the constrained variational problem into an unconstrained variational problem, and the expression of the extended Lagrangian function is:

[0011]

[0012] Further, use the alternating direction multiplier method to iteratively update ω k n+1 、τ n+1, the saddle point of the Lagrangian function is obtained, that is, the optimal solution; the update expressions are as follows:

[0013]

[0014] In the formula, n is the number of iterations, ω is the frequency, and γ is the noise tolerance parameter; are the Fourier transforms of v k (t), τ(t), and s(t) respectively.

[0015] Furthermore, several spatio-temporal feature extraction modules all include a TCN-a temporal convolutional module, a TCN-b temporal convolutional module, and a GCN graph convolutional module; the input ends of both the TCN-a temporal convolutional module and the TCN-b temporal convolutional module are simultaneously used as the input end of the spatio-temporal feature extraction module;

[0016] Subsequences of multiple different frequency scales are used as the first input data and input to the input end of the first spatio-temporal feature extraction module among several spatio-temporal feature extraction modules after convolution operations; the outputs of the TCN-a temporal convolutional module and the TCN-b temporal convolutional module in the first spatio-temporal feature extraction module are fused through dot product operations and then input to the input end of the GCN graph convolutional module of the first spatio-temporal feature extraction module; the input at the input end of the first spatio-temporal feature extraction module and the output at the output end of the GCN graph convolutional module of the first spatio-temporal feature extraction module are connected by residual connection to obtain the second input data;

[0017] The second input data is input to the input end of the second spatio-temporal feature extraction module, and the outputs of the TCN-a temporal convolutional module and the TCN-b temporal convolutional module in the second spatio-temporal feature extraction module are fused through dot product operations and then input to the input end of the GCN graph convolutional module of the second spatio-temporal feature extraction module; the input at the input end of the second spatio-temporal feature extraction module and the output at the output end of the GCN graph convolutional module of the second spatio-temporal feature extraction module are connected by residual connection to obtain the third input data;

[0018] And so on, until the input at the input end of the Nth spatio-temporal feature extraction module and the output at the output end of the GCN graph convolutional module of the Nth spatio-temporal feature extraction module are connected by residual connection to obtain the final output data.

[0019] Furthermore, in several of the spatio-temporal feature extraction modules, the data obtained after the TCN-a temporal convolutional module of the Nth spatio-temporal feature extraction module performs temporal convolution processing on the input data and then passes through the tanh activation function is multiplied by the data obtained after the TCN-b temporal convolutional module of the Nth spatio-temporal feature extraction module performs temporal convolution processing on the input data and then passes through the sigmoid activation function to obtain the output H of the temporal convolutional module through dot product operation t ;

[0020]

[0021] where Θ 1 , Θ 2 , b, and c are all learnable parameters, is the dot product operation, * represents convolution, g(·) is the tanh activation function, σ(·) is the sigmoid activation function, and H t is the output of the temporal convolution module, and x is the input data.

[0022] Furthermore, the outputs H t of each temporal convolution module in several of the spatio-temporal feature extraction modules all perform skip connections, and the sums of each skip connection are calculated sequentially until the sum of the data obtained after the sum of the Nth skip connection is added to the final output data output by the Nth spatio-temporal feature extraction module to obtain the spatio-temporal features of multiple subsequences with different frequency scales.

[0023] Furthermore, the core of the TCN temporal convolution module is to use causal dilated convolution as the temporal convolution layer to extract the temporal features of the time series; causal convolution introduces a dilation factor to perform discontinuous convolution operations on the time series, applies the convolution kernel to a region larger than its length to obtain causal dilated convolution; the calculation process is as follows:

[0024]

[0025] where g(t) is the data output at time t, and f n (i) is the ith filter, and x t-di is the data input at time t - di, d is the dilation factor, and s is the size of the convolution kernel.

[0026] Furthermore, the GCN graph convolution module uses a graph learning layer to learn the adjacency matrix of the graph; the multivariate load data and external influencing factors are used as the inputs of the nodes of the graph learning layer, and the corresponding node embedding features are generated, and then the similarity matrix between the nodes is calculated using the node embedding features, and the similarity matrix is non-linearly transformed to obtain the adjacency matrix A; the calculation process is as follows: The calculation process is as follows:

[0027] A = ReLU(tanh(α(M 1 M 2 T - M 2 M 1 T )))

[0028] for i = 1, 2,.., N

[0029] M 1=tanh(αE 1 Z 1 )

[0030] M 2 =tanh(αE 2 Z 2 )

[0031] idx = argtopm(A[i, :])

[0032] A = [i, -idx] = 0

[0033] where E 1 and E 2 represent the randomly initialized source node embedding feature and target node embedding feature respectively, Z 1 and Z 2 are both trainable parameters of the model; tanh is the activation function, α represents the saturation rate of the activation function, argtopm(·) represents selecting the m nodes closest to the central node as its neighbor nodes, and A is the adjacency matrix;

[0034] During the inter - layer propagation process of the GCN graph convolution module, a part of the original state of the node is retained, and the outputs of each layer are summed after linear transformation, so that the model can capture the information of deeper - layer neighbor nodes while retaining the original local information; the calculation process is as follows:

[0035]

[0036] where is the normalized adjacency matrix after adding self - connection; D is the degree matrix of nodes, I is the identity matrix; k is the depth of the GCN graph convolution module; β is the retention rate, βH 0 is the original state of the node retained before each propagation; W (k) is the parameter matrix.

[0037] Beneficial effects: The IES multi-load prediction method based on VMD and combined deep neural network decomposes multi-loads by using VMD to reduce the non-stationarity of the original load data; the multi-load components obtained after VMD decomposition and external influencing factors are used as inputs, the time series features of multi-source data are extracted through a temporal convolutional network, and the coupling features between multi-source data are mined by using a graph convolutional network; the long short-term memory network is used to capture the long-term dependence relationship of multi-source data; and the IES multi-load is predicted. The graph adjacency matrix generated by the graph learning layer in the GCN graph convolutional module can characterize the coupling relationship strength between multi-loads, as well as between the load and influencing factors, and to a certain extent reflects the importance of neighbor nodes to the central node; it can reduce the number of neighbor nodes of the central node and reduce the complexity of the graph; the spatio-temporal feature extraction module can learn the strong coupling between IES multi-loads, obtain coupling features, and improve the load prediction accuracy; the multi-load prediction method can fully mine the features at all levels of multi-source data while reducing the non-stationarity of load data, and has excellent prediction accuracy in the multi-load prediction of integrated energy systems. Description of the Drawings

[0038] Figure 1 It is a schematic diagram of the prediction model framework based on VMD-TCN-GCN-LSTM;

[0039] Figure 2 It is a schematic diagram of causal dilated convolution;

[0040] Figure 3 It is a schematic diagram of the temporal convolutional network structure;

[0041] Figure 4 It is a schematic diagram of the LSTM network structure;

[0042] Figure 5 It is the historical data of multi-loads;

[0043] Figure 6 The multi-step prediction results of the electrical load at 24 moments of each model;

[0044] Figure 7 The multi-step prediction results of the cooling load at 24 moments of each model;

[0045] Figure 8 The multi-step prediction results of the heating load at 24 moments of each model. Detailed Embodiments

[0046] The present invention will be further described below with reference to the accompanying drawings.

[0047] As Figure 1As shown, the IES multi-load prediction method based on VMD and combined deep neural network includes using the VMD variational mode decomposition module to decompose the original data of IES multi-component complex load, and decomposing it into multiple subsequences with different frequency scales; combining the TCN temporal convolutional module and the GCN graph convolutional module to obtain several spatio-temporal feature extraction modules, and using the spatio-temporal feature extraction network composed of several spatio-temporal feature extraction modules to extract the spatio-temporal features of multiple subsequences with different frequency scales; using the LSTM network to capture the long-term dependence relationship of multi-component data in multiple subsequences with different frequency scales; multiplying the long-term dependence relationship of multi-component data in multiple subsequences with different frequency scales and the spatio-temporal features of multiple subsequences with different frequency scales through dot product to obtain the result of multi-load prediction.

[0048] As Figure 1 shown, it is the network structure of the VMD-TCN-GCN-LSTM prediction model. Considering the non-stationarity of multi-load, first perform VMD decomposition on multi-load to obtain relatively stable modal components of multiple electrical loads, cooling loads, and heating loads, and then use the obtained multi-load components and influencing factors such as temperature and wind speed as input data. After processing through 1×1 convolution to the specified dimension, input it to the spatio-temporal feature extraction module composed of the TCN temporal convolutional module and the GCN graph convolutional module intertwined to fully extract the temporal features and coupling features of multi-component data; the temporal features and coupling features are combined to obtain the spatio-temporal features of multiple subsequences with different frequency scales.

[0049] The VMD variational mode decomposition, Variational Mode Decomposition is an adaptive and completely non-recursive signal decomposition technology. It overcomes the endpoint effect and modal component aliasing problems existing in empirical mode decomposition, and can reduce the non-stationarity of time series with high complexity and strong non-linearity. The VMD variational mode decomposition module decomposes the original data of IES multi-component complex load into K modal components v k (t) of multi-load, that is, multiple subsequences with different frequency scales; use the obtained multiple subsequences with different frequency scales as the input data of the spatio-temporal feature extraction network and the input data of the LSTM network;

[0050] Obtain the analytical signal of v k (t) through Hilbert transform, and modulate the central band of v k (t) to the corresponding baseband, and estimate the bandwidth using the Gaussian smoothness of the demodulated signal. The constrained variational problem is expressed as:

[0051]

[0052] In the formula, {v k} = {v 1 ,..., vk} is the modal component after VMD decomposition, ω k = {ω 1 ,..., ω k} is the center frequency of each component, δ(t) is the impulse function, is the gradient calculation, s(t) is the original signal;

[0053] Introduce the quadratic penalty factor and Lagrange multiplier to transform the constrained variational problem into an unconstrained variational problem. The expression of the extended Lagrangian function is:

[0054]

[0055] Use the alternating direction multiplier method to iteratively update ω k n+1 , τ n+1 , and obtain the saddle point of the Lagrangian function, that is, the optimal solution; is the k-th modal component obtained after the (n + 1)-th iteration of the composite data, ω k n+1 is the center frequency of the k-th component obtained after the (n + 1)-th iteration, τ n+1 is the Lagrange multiplier after the (n + 1)-th iteration; the update expressions are as follows:

[0056]

[0057] In the formula, n is the number of iterations, ω is the frequency, and γ is the noise tolerance parameter; are the Fourier transforms of v k (t), τ(t), s(t) respectively, is the numerical value of the Fourier transform of the Lagrange multiplier after the n-th iteration.

[0058] Specifically, the alternating direction multiplier method transforms the original constrained optimization problem into a series of unconstrained sub-problems by introducing the Lagrange multiplier, and then approximates the optimal solution of the original problem by alternately optimizing these sub-problems, which enables the VMD algorithm to converge to a better solution more quickly and stably.

[0059] Such as Figure 1 shown, several spatio-temporal feature extraction modules all include the TCN-a temporal convolutional module, the TCN-b temporal convolutional module, and the GCN graph convolutional module; the input ends of the TCN-a temporal convolutional module and the TCN-b temporal convolutional module are both used as the input end of the spatio-temporal feature extraction module at the same time;

[0060] Sub - sequences of multiple different frequency scales are used as the first input data after convolution operations through 1×1Conv and then input to the input end of the first spatio - temporal feature extraction module among several spatio - temporal feature extraction modules; the outputs of the TCN - a temporal convolution module and the TCN - b temporal convolution module in the first spatio - temporal feature extraction module are fused through dot - product operation and then input to the input end of the GCN graph convolution module in the first spatio - temporal feature extraction module; the input at the input end of the first spatio - temporal feature extraction module and the output at the output end of the GCN graph convolution module in the first spatio - temporal feature extraction module are connected by residual connection to obtain the second input data; that is, the coupled feature output from the output end of the GCN graph convolution module in the first spatio - temporal feature extraction module.

[0061] The second input data is input to the input end of the second spatio - temporal feature extraction module. The outputs of the TCN - a temporal convolution module and the TCN - b temporal convolution module in the second spatio - temporal feature extraction module are fused through dot - product operation and then input to the input end of the GCN graph convolution module in the second spatio - temporal feature extraction module; the input at the input end of the second spatio - temporal feature extraction module and the output at the output end of the GCN graph convolution module in the second spatio - temporal feature extraction module are connected by residual connection to obtain the third input data; that is, the coupled feature output from the output end of the GCN graph convolution module in the second spatio - temporal feature extraction module.

[0062] By analogy, until the input at the input end of the Nth spatio - temporal feature extraction module and the output at the output end of the GCN graph convolution module in the Nth spatio - temporal feature extraction module are connected by residual connection to obtain the final output data.

[0063] As Figure 1 shown, among several said spatio - temporal feature extraction modules, the data obtained after the TCN - a temporal convolution module in the Nth spatio - temporal feature extraction module performs temporal convolution processing on the input data and then passes through the tanh activation function is multiplied by the data obtained after the TCN - b temporal convolution module in the Nth spatio - temporal feature extraction module performs temporal convolution processing on the input data and then passes through the sigmoid activation function, and the dot - product operation is used for fusion to obtain the output H of the temporal convolution module. t ; The outputs of two parallel TCN temporal convolution modules are multiplied to obtain the output of the temporal convolution module. For a given input x, the calculation formula of the temporal convolution module can be expressed as:

[0064]

[0065] In the formula, Θ 1 、Θ 2 、b, and c are all learnable parameters. is the dot - product operation, * represents convolution, g(·) is the tanh activation function, σ(·) is the sigmoid activation function, and Ht is the output of the temporal convolution module, and x is the input data.

[0066] After the TCN-a temporal convolution module in the first spatio-temporal feature extraction module performs temporal convolution processing on the input data, it is then processed by the tanh activation function to obtain the TCN-a data of the first spatio-temporal feature extraction module; the TCN-b temporal convolution module in the first spatio-temporal feature extraction module performs temporal convolution processing on the input data, and then passes through the sigmoid activation function to obtain the TCN-b data of the first spatio-temporal feature extraction module; the TCN-a data of the first spatio-temporal feature extraction module and the TCN-b data of the first spatio-temporal feature extraction module are fused by dot product operation to obtain the output H1 of the temporal convolution module.

[0067] After the TCN-a temporal convolution module in the second spatio-temporal feature extraction module performs temporal convolution processing on the input data, it is then processed by the tanh activation function to obtain the TCN-a data of the second spatio-temporal feature extraction module; the TCN-b temporal convolution module in the second spatio-temporal feature extraction module performs temporal convolution processing on the input data, and then passes through the sigmoid activation function to obtain the TCN-b data of the second spatio-temporal feature extraction module; the TCN-a data of the second spatio-temporal feature extraction module and the TCN-b data of the second spatio-temporal feature extraction module are fused by dot product operation to obtain the output H2 of the temporal convolution module; and so on, calculating the output H of the temporal convolution module in each spatio-temporal feature extraction module among several said spatio-temporal feature extraction modules. t .

[0068] The output H of each temporal convolution module in several said spatio-temporal feature extraction modules t All perform skip connections, and the sums of each skip connection are calculated in turn until the sum of the data obtained after the sum of the Nth skip connection ends is added to the final output data output by the Nth spatio-temporal feature extraction module to obtain the spatio-temporal features of multiple subsequences with different frequency scales.

[0069] In the first spatio-temporal feature extraction module, the outputs of two parallel TCN (Temporal Convolutional Network) temporal convolutional modules are multiplied to obtain the output H1 of the temporal convolutional module. Then, the output H1 and multiple subsequences of different frequency scales are used as input data for a summation operation to obtain the first summation result. In the second spatio-temporal feature extraction module, the outputs of two parallel TCN temporal convolutional modules are multiplied to obtain the output H2 of the temporal convolutional module, and then the output H2 and the first summation result are used for a summation operation to obtain the second summation result. In the third spatio-temporal feature extraction module, the outputs of two parallel TCN temporal convolutional modules are multiplied to obtain the output H3 of the temporal convolutional module, and then the output H3 and the second summation result are used for a summation operation to obtain the third summation result. And so on. In the Nth spatio-temporal feature extraction module, the outputs of two parallel TCN temporal convolutional modules are multiplied to obtain the output Ht of the temporal convolutional module, and then the output Ht and the (N - 1)th summation result are used for a summation operation to obtain the Nth summation result. The spatio-temporal features of multiple subsequences of different frequency scales are obtained by summing the Nth summation result and the final output data output by the Nth spatio-temporal feature extraction module.

[0070] As Figure 2 shown, the core of the TCN temporal convolutional module is to use causal dilated convolution as the temporal convolutional layer to extract the temporal features of the time series. In causal convolution, all data has a one-to-one causal relationship in the time dimension. This characteristic makes causal convolution have strict time constraints. If we want to improve its memory ability for time series, we can only increase the number of convolutional layers and the size of the convolutional kernel, but this often leads to an increase in the number of trainable parameters of the network model and increases the difficulty of model training. While dilated convolution can solve the problems existing in the above-mentioned causal convolution. When the number of convolutional layers is relatively small, causal convolution introduces a dilation factor to perform discontinuous convolution operations on the time series, applies the convolutional kernel to a region larger than its length, and obtains causal dilated convolution. As Figure 2 shown, d = 1 means sampling each point, d = 2 means sampling every two points. As the number of network layers increases, the receptive field of the convolutional kernel grows exponentially, so that global information can be learned with fewer layers. The calculation process is as follows:

[0071]

[0072] In the formula, g(t) is the data output at time t, f n (i) is the i-th filter, x t-di is the data input at time t - di, d is the dilation factor, and s is the size of the convolutional kernel.

[0073] As Figure 3As shown in the figure, the structure of the TCN time convolution module includes a causal dilated convolution unit, a weight normalization unit, an activation function unit, and a dropout unit. To solve the problem of gradient vanishing in deep neural networks, the TCN time convolution module introduces a residual connection, and sums the input data and the output of the deep neural network as the input for the next stage. IES multivariate data is a typical time series. TCN uses causal dilated convolution as the time convolution layer, enabling the convolution kernel to perform convolution operations on IES multivariate data at different time steps, thereby learning the variation laws and characteristics in the time dimension. Therefore, the TCN network can fully extract the temporal features of IES multivariate data.

[0074] The GCN graph convolution module is an extension of the convolutional neural network in non-Euclidean space. For data with a graph structure, it can be described as G=(V, E), where V represents the set of nodes in graph G, and E represents the set of edges in graph G. Since there are coupling relationships among IES multivariate loads, as well as between loads and influencing factors such as temperature and wind speed; this coupling relationship refers to the mutual association and influence among different types of loads, as well as between loads and influencing factors, and its essence is an implicit graph structure. Therefore, the GCN graph convolution module is used to learn this coupling relationship. When constructing the adjacency matrix in a conventional GCN network, if there is an edge connection between nodes, the corresponding position in the adjacency matrix is represented as 1, otherwise it is represented as 0. However, there is no actual edge connection relationship among multivariate loads, nor between loads and their influencing factors. Moreover, this method treats the influence of all neighbor nodes on the central node as the same, without considering that the influence degrees of different neighbor nodes on the central node are different. Paying too much attention to neighbor nodes with less influence on the central node will have a negative impact on the model performance.

[0075] Therefore, the GCN graph convolution module adopts a graph learning layer to learn the adjacency matrix of the graph. Since there is no actual edge connection relationship between influencing factors such as temperature and wind speed and multivariate loads, but an implicit graph structure, the multivariate load data and external influencing factors such as temperature and wind speed are used as the input of the nodes in the graph learning layer, and the corresponding node embedding features are generated. Then, the similarity between nodes is calculated using the node embedding features, that is, by calculating M 1 M 2 T -M 2 M 1 T ; A similarity matrix can be obtained. The elements in the similarity matrix represent the similarity between the input data of different nodes. The similarity matrix is non-linearly transformed to obtain the adjacency matrix A. Node pairs with high similarity will have larger values in the adjacency matrix, indicating a strong coupling relationship between the node pairs. The calculation process is as follows:

[0076] A = ReLU(tanh(α(M 1 M 2 T -M 2 M 1 T )))

[0077] for i = 1, 2,.., N

[0078] M 1 = tanh(αE 1 Z 1 )

[0079] M 2 = tanh(αE 2 Z 2 )

[0080] idx = argtopm(A[i, :])

[0081] A = [i, -idx] = 0

[0082] where E 1 , E 2 respectively represent the source node embedding features and target node embedding features randomly initialized, Z 1 , Z 2 are both trainable parameters of the model; tanh is an activation function, the hyperbolic tangent function, with an output range of (-1, 1), which is an activation function in neural networks to increase output non-linearity, expressed as (e x - e-x) / (e x + e-x); a represents the saturation rate of the activation function, argtopm(·) represents selecting the m nodes closest to the central node as its neighbor nodes to ensure the sparsity of the graph; A is the adjacency matrix; M 1 is the output matrix obtained after the aggregation transformation of the source node embedding feature matrix, M 2 is the output matrix obtained after the aggregation transformation of the target node embedding feature matrix; A[i, :] is all the data in the i-th row of the adjacency matrix, and A = [i, -idx] = 0 mainly makes the data of the selected nodes closest to the central node 0 to ensure the sparsity of the graph.

[0083] To prevent the node hidden state from converging to a single point during information propagation, resulting in information loss, during the inter-layer propagation of the GCN graph convolution module, a part of the original state of the node is retained, and the outputs of each layer are summed after linear transformation, so that the model can capture the information of deeper neighbor nodes while retaining the original local information; the calculation process is as follows:

[0084]

[0085] In the formula, is the normalized adjacency matrix after adding self-connection; D is the degree matrix of nodes, which is a diagonal matrix; I is the identity matrix, with all diagonal data being 1 and other data being 0; k is the depth of the GCN graph convolution module; β is the retention rate, and βH 0 is the original state of the nodes retained before each propagation; W (k) is the parameter matrix, and H (k) is the k-order depth information propagation result; Hout is the key information related to the prediction result obtained after filtering and screening the information propagation result, that is, the final screening result; the final screening result can be used as the coupled features mined by the GCN graph convolution module.

[0086] The graph learning layer in the GCN graph convolution module automatically generates the adjacency matrix A through the randomly initialized node embedding features of the multivariate load raw data, and after normalizing the adjacency matrix A, the normalized adjacency matrix after adding self-connection is obtained. It is input to the GCN graph convolution module, and at the same time, the output of the time convolution module is used as the input of the GCN graph convolution module. The GCN graph convolution module combines the normalized adjacency matrix after adding self-connection and the output Ht of the time convolution module to achieve the extraction of coupled features between multivariate data;

[0087] To solve the problems of gradient explosion and gradient disappearance during the training process of the model, the input of each TCN time convolution module and the output of the GCN graph convolution module are connected by residual connection and then input to the next spatio-temporal feature extraction module. At the same time, skip connections are made to the output of each TCN time convolution module. Its essence is a convolution of Li×1, where Li represents the length of the input sequence of the i-th skip connection. The sum of each skip connection is calculated in turn until the end of the last skip connection. The spatio-temporal features of multiple subsequences with different frequency scales extracted by the spatio-temporal feature extraction network formed by combining TCN and GCN are convolved through a 1×1Conv convolution operation, and the result output by the fully connected layer capturing the long-term dependence relationship of the LSTM module is fused by dot product. After the result fused by dot product is processed by the ReLU function, finally, the predicted result is obtained by reconstructing the model output.

[0088] As Figure 4 shown, the LSTM network is a variant of the recurrent neural network (RNN), which can avoid the problem of gradient disappearance encountered during the training of RNN and is good at capturing the long-term change trend of time series; as Figure 4 shown, the internal structure of the LSTM network consists of a forget gate, an input gate, and an output gate.

[0089] Forget gate f t :

[0090] f t = σ(W f · [h t-1 , x t + b f )

[0091] In the formula, W f represents the weight matrix, and b f represents the bias;

[0092] Input gate v t :

[0093] v t = σ(W v [h t-1 , x t + b v )

[0094] g t = tanh(W g [h t-1 , x t + b g )

[0095]

[0096] In the formula, W v and W g both represent the weight matrix, and b v , b g both represent the bias; represents the Hadamard product;

[0097] Output gate o t :

[0098] o t = σ(W o [h t-1 , x t + b o )

[0099]

[0100] In the formula, W o represents the weight matrix, and b o represents the bias; represents the Hadamard product;

[0101] Considering the long-term dependencies among historical multivariate data can enhance the model's fitting ability for load data; the LSTM network finely controls the information flow through its internal gating mechanism. The forget gate controls the information discarded in the cell unit, the input gate controls the information added to the cell unit, the output gate controls the output of the cell unit, and the cell unit information can be propagated across multiple time steps, enabling the LSTM to capture the long-term dependencies of multivariate data.

[0102] Embodiment

[0103] Given that the constructed prediction model needs to perform predictive analysis on multiple loads at the same time, this paper selects the root mean square error (RMSE), the mean absolute percentage error (MAPE), and the weighted mean absolute percentage error (WMAPE) as evaluation indicators. Among them, RMSE is used to measure the average deviation degree between the predicted values and the true values of various loads; MAPE is used to evaluate the relative error of the prediction model; WMAPE focuses on the weights of different load types and reflects the overall performance of the prediction model. The expressions of each evaluation indicator are as follows:

[0104]

[0105] where y i ′, y i are the predicted value and the actual value of the load at the i-th moment, respectively; N is the number of samples; α ele , α cool , α heat are the weights of the electric, cooling, and heating loads, which are taken as 0.5, 0.3, and 0.2 in this paper, respectively; are the mean absolute percentage errors of the electric, cooling, and heating loads, respectively.

[0106] The model in this paper is built based on the Python 3.8 environment and the Pytorch framework. Given that the hyperparameters of deep neural networks have a great impact on the model performance and there is no perfect theory to guide the selection of hyperparameters at present, after multiple experimental tests, the number of network layers is set to 5 in this paper. A dropout layer is added after each hidden layer to prevent overfitting. The value of dropout is set to 0.3, the training batch size is set to 64, the number of training epochs is set to 120, the initial learning rate is set to 0.0001, the saturation rate α of the activation function is set to 3, the retention rate β during the propagation process between GCN layers is set to 0.05, the depth k of GCN is set to 3, and the Adam algorithm is used as the optimization algorithm; when using VMD to decompose the multivariate load, grid search is used to determine the number of modes and the secondary balance parameter. Finally, the number of modes K is set to 6, and the secondary balance parameter is set to 2000.

[0107] Taking the multivariate load dataset of Arizona State University's Tempe campus from January 1, 2020 to January 28, 2021 as the sample data, with a sampling interval of 1 hour and a total of 9456 data, and using January 2021 as the test set, the remaining data is divided into a training set and a validation set according to an 8:2 ratio. The meteorological data comes from the Phoenix Sky Harbor International Airport weather station closest to the Tempe campus. The electrical load, cooling load, heating load, solar radiation, carbon emissions, wind speed, temperature, dew point, air pressure, precipitation, month, hour, and holiday situation are selected as the input features of the prediction model. Since there are abnormal data situations during the data recording process, which may have a great impact on the prediction performance of the model, it is necessary to preprocess the original data. The quartile method is used for outlier detection, and Akima cubic interpolation is used for correction and replacement. Finally, the measurement unit of the multivariate load is unified to MW to obtain the historical multivariate load curve, as Figure 5 shown.

[0108] Taking the historical data of the multivariate load as the original data, it is decomposed using the VMD variational mode decomposition module to obtain the decomposition results of the electrical, cooling, and heating loads; the subsequences of the electrical, cooling, and heating loads after VMD decomposition have better stationarity, and the central frequencies of the decomposed components are different, effectively avoiding mode mixing and fully retaining the characteristics of the original data, which is beneficial for the subsequent prediction model to mine. To make the input data have the same measurement scale, before importing the input data into the prediction model, this paper uses the maximum-minimum normalization method to process it. The specific formula is as follows:

[0109]

[0110] In the formula, x g represents the data after normalization, x represents the data to be normalized, and x max , x min represent the maximum and minimum values of the input data respectively.

[0111] Multi-step multi-variable load forecasting includes single-step multi-variable load forecasting and multi-step multi-variable load forecasting. Due to problems such as dynamic changes, data noise, and uncertainties in the load in actual engineering, single-step load forecasting methods often have difficulty accurately capturing its characteristics. Therefore, multi-step multi-variable load forecasting is commonly used in actual engineering. Thus, the VMD-TCN-GCN-LSTM forecasting model proposed in this paper is compared with four basic forecasting models, namely the CNN network (Convolutional Neural Network), the LSTM network (Long Short-Term Memory), the GRU network (Gated Recurrent Unit, a type of recurrent neural network RNN architecture for processing sequential data), and the CNN-LSTM network (a deep learning model that combines the convolutional neural network CNN and the long short-term memory network LSTM) for multi-step multi-variable load forecasting to verify the effectiveness of the proposed forecasting model in multi-step load forecasting. The input sequence step size is set to 168, and the output sequence step size is set to 24, that is, the load values for the next 24 time points are predicted at once using the historical data of the previous week. The predicted curves of electricity, cooling, and heating loads are as Figures 6 - 8 shown, and the prediction errors are shown in Table 1 below.

[0112] Table 1 Comparison of errors of each model in multi-step multi-variable load forecasting of the integrated energy system

[0113]

[0114] As Figures 6 - 8 shown, and as shown in Table 1, the comparison table of errors of each model in multi-step multi-variable load forecasting of the integrated energy system; since CNN is not good at extracting the temporal features of multi-variable data, its prediction effect is the worst when predicting the load values for the next 24 time points at once; due to the internal structure of LSTM and GRU having memory units, their ability to extract the features of multi-variable data is better than that of CNN, and their prediction accuracy is higher than that of CNN. However, due to multi-step forecasting, that is, outputting the load values for the next 24 time points at once, the prediction accuracy is limited; the CNN-LSTM combined model can capture the long-term dependence relationship of multi-variable data while extracting the local features of multi-variable data, and its prediction effect is better than that of the other three single models; since the model proposed in this paper fully considers the features at all levels of multi-variable data, its prediction accuracy is the highest when predicting the next 24 time points. From the prediction curves, compared with the other four comparison models, the prediction results of the proposed forecasting model in this paper are more in line with the actual load change curve, verifying that the proposed forecasting model in this paper has higher prediction accuracy in multi-step multi-variable load forecasting.

[0115] The above is only a description of the preferred embodiments of the present invention. Those of ordinary skill in the art can make several modifications and optimizations based on the above disclosure without departing from the basic principles. These improvements and optimizations should be regarded as the scope of protection understood by the present invention.

Claims

1. IES multivariate load forecasting method based on VMD and combined deep neural network, characterized by: The method includes using a VMD variational mode decomposition module to decompose the 1ES multivariate complex load original data to obtain multiple sub-sequences of different frequency scales; obtaining several spatiotemporal feature extraction modules by combining a TCN time convolution module and a GCN graph convolution module, and using a spatiotemporal feature extraction network composed of several spatiotemporal feature extraction modules to extract the spatiotemporal features of sub-sequences of multiple frequency scales; using an LSTM network to capture the long-term dependencies of multivariate data in sub-sequences of multiple frequency scales; fusing the long-term dependencies of multivariate data in sub-sequences of multiple frequency scales and the spatiotemporal features of sub-sequences of multiple frequency scales through point multiplication to obtain the result of multivariate load prediction.

2. The IES multivariate load forecasting method based on VMD and combined deep neural network according to claim 1 is characterized in that: The VMD variational modal decomposition module decomposes the IES multivariate complex load original data into K multivariate load modal components v k (t), that is, multiple subsequences of different frequency scales; the obtained multiple subsequences of different frequency scales are used as input data of the spatiotemporal feature extraction network and the input data of the LSTM network; By Hilbert transform, we can get v k (t) and v k The center band of (t) is modulated to the corresponding baseband, and the bandwidth is estimated using the Gaussian smoothness of the demodulated signal. The constrained variational problem is expressed as: In the formula, {v k }={v1,...,v k } is the modal component after VMD decomposition, ω k ={ω1, ..., ω k } is the center frequency of each component, δ(t) is the impulse function, is the gradient calculation, s(t) is the original signal; The quadratic penalty factor and Lagrange multiplier are introduced to transform the constrained variational problem into an unconstrained variational problem. The extended Lagrangian function expression is:

3. The IES multivariate load forecasting method based on VMD and combined deep neural network according to claim 2 is characterized in that: Iterative update using alternating direction multiplier method ω k n+1 , τ n+1 , find the Lagrangian function saddle point, which is the optimal solution; the update expression is as follows: Where n is the number of iterations, ω is the frequency, and γ is the noise tolerance parameter; They are v k The Fourier transform of τ(t), τ(t), and s(t).

4. The IES multivariate load forecasting method based on VMD and combined deep neural network according to claim 1 is characterized in that: Several spatiotemporal feature extraction modules include a TCN-a temporal convolution module, a TCN-b temporal convolution module and a GCN graph convolution module; the input ends of the TCN-a temporal convolution module and the TCN-b temporal convolution module are both used as the input ends of the spatiotemporal feature extraction module; After convolution operation, multiple subsequences of different frequency scales are input as first input data to the input end of the first spatiotemporal feature extraction module among the plurality of spatiotemporal feature extraction modules; the outputs of the TCN-a temporal convolution module and the TCN-b temporal convolution module in the first spatiotemporal feature extraction module are fused by point multiplication operation and input to the input end of the GCN graph convolution module of the first spatiotemporal feature extraction module; the input of the input end of the first spatiotemporal feature extraction module and the output of the GCN graph convolution module of the first spatiotemporal feature extraction module are residually connected to obtain the second input data; The second input data is input to the input end of the second spatiotemporal feature extraction module, and the outputs of the TCN-a time convolution module and the TCN-b time convolution module in the second spatiotemporal feature extraction module are fused by point multiplication operation and input to the input end of the GCN graph convolution module of the second spatiotemporal feature extraction module; the input of the input end of the second spatiotemporal feature extraction module and the output of the output end of the GCN graph convolution module of the second spatiotemporal feature extraction module are residually connected to obtain the third input data; And so on, until the input of the input end of the Nth spatiotemporal feature extraction module and the output of the output end of the Nth spatiotemporal feature extraction module GCN graph convolution module are residually connected to obtain the final output data.

5. The IES multivariate load forecasting method based on VMD and combined deep neural network according to claim 4 is characterized in that: In the plurality of spatiotemporal feature extraction modules, the data obtained by performing temporal convolution processing on the input data of the TCN-a temporal convolution module of the Nth spatiotemporal feature extraction module and then processing it with the tanh activation function are combined with the data obtained by performing temporal convolution processing on the input data of the TCN-b temporal convolution module of the Nth spatiotemporal feature extraction module and then processing it with the sigmoid activation function, and the output H of the temporal convolution module is obtained by performing a dot multiplication operation. t ; In the formula, Θ1, Θ2, b, and c are all learnable parameters. is a dot multiplication operation, * represents convolution, g(·) is the tanh activation function, σ(·) is the sigmoid activation function, H t is the output of the temporal convolution module, and x is the input data.

6. The 1ES multivariate load forecasting method based on VMD and combined deep neural network according to claim 5 is characterized in that: The output H of each time convolution module in the spatiotemporal feature extraction modules is t All jump connections are made, and each jump connection is summed in turn until the data obtained by the summation of the Nth jump connection is summed with the final output data output by the Nth spatiotemporal feature extraction module to obtain the spatiotemporal features of multiple subsequences of different frequency scales.

7. The 1ES multivariate load forecasting method based on VMD and combined deep neural network according to claim 1 is characterized in that: The core of the TCN temporal convolution module is to use causal dilated convolution as the temporal convolution layer to extract the temporal features of the time series; the causal convolution introduces a dilation factor to perform discontinuous convolution operations on the time series, and applies the convolution kernel to a region larger than its length to obtain a causal dilated convolution; the calculation process is as follows: In the formula, g(t) is the data output at time t, f n (i) is the i-th filter, x t-di is the data input at time t-di, d is the dilation factor, and s is the size of the convolution kernel.

8. The IES multivariate load forecasting method based on VMD and combined deep neural network according to claim 1 is characterized in that: The GCN graph convolution module uses a graph learning layer to learn the adjacency matrix of the graph; multivariate load data and external influencing factors are used as inputs of the graph learning layer nodes, and corresponding node embedding features are generated, and then the node embedding features are used to calculate the similarity between nodes to obtain a similarity matrix, and the similarity matrix is ​​subjected to a nonlinear transformation to obtain an adjacency matrix A; the calculation process is as follows: A=ReLU(tanh(α(M1M2 T -M2M1 T ))) for i = 1, 2, ..., N M1=tanh(αE1Z1) M2=tanh(αE2Z2) idx = argtopm(A[i,:]) A=[i,−idx]=0 Where E1 and E2 represent the randomly initialized source node embedding features and target node embedding features, respectively; Z1 and Z2 are both trainable parameters of the model; tanh is the activation function, a represents the saturation rate of the activation function, argtopm(·) represents selecting the m nodes closest to the central node as its neighbor nodes, and A is the adjacency matrix; During the inter-layer propagation of the GCN graph convolution module, part of the original state of the node is retained, and the output of each layer is summed after linear transformation, so that the model can capture the information of deeper neighbor nodes while retaining the original local information; the calculation process is as follows: In the formula, is the normalized adjacency matrix after adding self-connection; D is the degree matrix of the node, I is the identity matrix; k is the depth of the GCN graph convolution module; β is the retention rate, βH0 is the original state of the node retained before each propagation; W (k) is the parameter matrix.

Citation Information

Patent Citations

  • Comprehensive energy system multi-element load prediction method and system based on deep learning

    CN115936218A