Office building energy consumption prediction method and system
By decomposing office building energy consumption into modal components and using GCN and Transformer models, the problem of low energy consumption prediction accuracy in the prior art is solved, and more accurate energy consumption trend prediction is achieved.
Patent Information
- Application Number
- CN202510685367.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The prior art is difficult to accurately capture the nonlinear and dynamic characteristics of office building energy consumption, especially when dealing with the synergistic effects between different energy consumption loads and the influence of external factors, the prediction accuracy is low.
By decomposing the load into multiple modal components, using graph convolutional network (GCN) to find the coordinated energy consumption relationship of each group of loads, and combining the Transformer model to extract spatiotemporal features to build an energy consumption prediction model.
It improves the accuracy of energy consumption prediction, can more effectively capture the synergistic relationship between loads and the influence of external factors, and provides more accurate prediction of future energy consumption trends.
Smart Images

Figure CN120217609A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of energy consumption prediction for office buildings, and particularly relates to a method and system for predicting energy consumption of office buildings. Background Art
[0002] With the intensification of the energy crisis and the enhancement of environmental protection awareness, the prediction of energy consumption in office buildings has become an important means of energy conservation and emission reduction. Traditional energy consumption prediction methods mostly rely on the statistical analysis of historical data and are difficult to accurately capture the non-linear and dynamic characteristics of building energy consumption. In recent years, deep learning technology has made remarkable progress in the field of time series prediction, especially the graph convolutional neural network (GCN) and the Transformer model have shown excellent performance in processing complex sequence data.
[0003] Building energy consumption is usually composed of multiple loads (such as electricity, air conditioning, lighting, equipment, etc.), and there are complex collaborative energy consumption relationships among these loads. For example, there may be a collaborative relationship between the lighting load and the air conditioning load during office hours. That is, when the lighting load increases, the air conditioning load may also increase accordingly. Especially in summer or winter, the load changes of lighting and air conditioning are often interdependent. Traditional models such as ARIMA and linear regression usually model the energy consumption sequence as a whole, ignoring the interaction and dynamic coupling characteristics between these different loads. This processing method makes it difficult for these models to accurately capture the synergistic effect between loads, resulting in a significant deviation between the prediction results and the actual energy consumption situation. Moreover, building energy consumption is not only affected by historical energy consumption data but also closely related to external factors (such as weather data, equipment operation status, etc.). These external factors play an important role in the prediction of building energy consumption. Especially, the impact of weather changes on the air conditioning load is significant. High-temperature weather usually leads to a sharp increase in the air conditioning load, while low-temperature weather may cause changes in the heating load. However, existing models usually only rely on historical energy consumption data and lack the ability to effectively fuse multi-source data. This limitation of a single data source results in the prediction accuracy of traditional models being significantly limited when facing external environmental changes. The building energy consumption sequence usually contains two significant characteristics: on the one hand, there is a part with strong regularity, such as periodic electricity consumption (for example, the daily load fluctuation of the air conditioning system); on the other hand, there is a part with strong randomness, such as sudden electricity consumption (for example, the start of temporary equipment or energy consumption fluctuations caused by special activities). However, existing models such as long short-term memory networks, although able to handle some long-term and short-term dependencies, still face difficulties in dealing with sudden high fluctuations. Especially when sudden events cannot effectively learn these irregular changes from past sequences, their accuracy is often inferior to the prediction of the regular part. Building energy consumption data has significant multi-scale characteristics, covering changes at different time scales, such as hourly fluctuations, daily periodic fluctuations, weekly periodic fluctuations, and seasonal changes, etc. These multi-scale characteristics reflect the dynamic changes of building energy consumption within different time ranges and can have a significant impact on short-term fluctuations and long-term trends. Traditional models such as LSTM, although having advantages in capturing long-term dependencies, its structure is mainly modeled based on a single time scale and cannot effectively handle the combined effects of seasonal fluctuations and daily periodic changes. Similarly, although the GRU network is efficient, it also faces similar problems and cannot adapt to the differences in electricity consumption patterns between summer and winter and other periodic fluctuations at the same time. This makes it difficult for traditional models to provide accurate results when predicting building energy consumption, especially in scenarios with multi-scale characteristics. Summary of the Invention
[0004] In view of the deficiency that the existing technology cannot effectively model the synergy between different energy consumption loads, the present invention proposes an office building energy consumption prediction method and system, which improves the prediction accuracy by decomposing the load into multiple groups of modal components with typical spectral characteristics corresponding to different electricity consumption principles, and uses GCN to find the synergistic energy consumption relationship of each group of loads, thereby solving the problems existing in the existing technology.
[0005] A method for predicting energy consumption of an office building comprises the following steps: Collect the load sequence of different energy-using equipment in the current office building; Decompose the load sequence of different energy-using equipment into multiple modal components, and calculate the sample entropy of each modal component; reorganize the corresponding modal components according to the sample entropy of each modal component; use each reorganized modal component as a node, use dynamic time warping to calculate the temporal similarity between each modal component, and use the temporal similarity as the edge between nodes to construct a graph structure; The graph structure is input into the GCN-Transformer model, and the graph convolution operation of the graph convolution network GCN is used to extract the spatial features of the nodes in the graph structure. The spatial features of the nodes are input into the Transformer model, and the temporal features of the spatial features are extracted by introducing the self-attention mechanism to obtain the spatiotemporal features. By inputting the spatiotemporal features into the fully connected layer, the energy consumption prediction value for a period of time in the future is obtained.
[0006] Furthermore, dynamic time warping (DTW) is used to obtain the similarity between the modal components. , specifically including the following steps: Define X and Y as two modal component sequences with length T : Calculate each time step i , j The Euclidean distance between , and obtain a cost matrix of size T×T; where, x i is the sequence X in i The value of the time step; y j is the sequence Y in j The value of the time step; For each point ( i , j ), calculate its minimum cumulative distance: ; in, Indicates moving from above to ( i , j ), Indicates moving from the left to (i , j the cumulative distance of indicating the cumulative distance from moving diagonally to ( i , j ); calculating step by step the i , j of each point ( ) until ; then the DTW distance between sequences X and Y is: .
[0007] Furthermore, the graph convolution operation of the graph convolutional network GCN is used to extract the spatial features of the nodes in the graph structure, expressed as: ; where represents the node feature matrix of the l th layer; represents the weight matrix of the l th layer, used to perform a linear transformation on the nodes; is the degree matrix, a matrix composed of the diagonal elements of the adjacency matrix A; is the activation function; the adjacency matrix A is composed of , expressed as: ; where is the set threshold; is the Laplacian matrix, .
[0008] Furthermore, the spatial features of the nodes are input into the Transformer model, and the temporal features of the spatial features are extracted by introducing the self-attention mechanism, specifically including the following steps: Input the output into the Transformer to extract the time dependence. Since the Transformer requires position information, the input sequence of the Transformer is expressed as: , where , is the position encoding, expressed as: , t is the current time step, d is the total dimension of the input features; In the multi-head attention layer, the input feature generates query, key, and value matrices through linear transformation; Calculate the self-attention weights based on the query, key, and value matrices, and calculate the multi-head attention output in parallel according to the self-attention weights to obtain ; Perform a residual connection between the multi-head attention output and the input sequence to obtain features ; Perform calculations in the input feed-forward network to obtain features ; ; For the obtained features and perform a residual connection and then layer normalization to obtain the final temporal features extracted by the Transformer .
[0009] Furthermore, the variational mode decomposition method is used to decompose the load sequence into multiple mode components, which specifically includes the following steps: Construct the mathematical model of variational mode decomposition: ; where are the mode components, is the k th mode component; are the central frequencies of the mode components, is the central frequency corresponding to the k th mode; is the Dirac function; K is the decomposition layer number; j 1 is the imaginary unit; * is the convolution operator; is the original signal; k represents a positive integer; By introducing the penalty factor and the Lagrange multiplier operator , the mathematical model of variational mode decomposition is transformed into: ; In the formula is the Lagrange multiplier, is the penalty factor; Through iterative optimization, solve the variational mode decomposition model; in each iteration, update and the central frequency until convergence to obtain the final and the corresponding central frequency.
[0010] Furthermore, calculating the sample entropy for each mode component specifically includes the following steps: Form a vector subsequence of dimension m by arranging each mode component in sequence: , where ; each mode component is a time series composed of N data; Define the vector subsequence The distance between it and another vector subsequence is the absolute value of the maximum difference among the corresponding elements of the two; where denotes the starting position index of the subsequence w and s denotes the starting position index of another subsequence For a given , count the number of whose distance from is less than or equal to the similarity threshold r . Denote it as s , . , For , starting from the w -th subsequence, the number of other m -dimensional subsequences whose distance from it is less than or equal to the threshold r is defined as . Calculate the average value of all . This average value represents the matching probability of any two m -dimensional sequences under the threshold r . Increase the dimension to m +1, calculate the number of whose distance from is less than or equal to r . Denote it as , . Calculate the number of other w -dimensional subsequences whose distance from the m+ -th subsequence is less than or equal to the threshold r . Denote it as . Calculate the average value of all . This average value represents the matching probability of any two m +1-dimensional sequences under the threshold r . Then the sample entropy is expressed as: . In the formula: m is the sequence dimension; N is the number of samples.
[0011] Furthermore, reconstituting the corresponding modal component according to the sample entropy of each modal component specifically includes: The number of samples contained in the modal components that define each different frequency characteristic is N p , take out the first i Sample Components samp i , then we get N p The silhouette coefficient SC of the samples Np It is expressed as: ; in, i Represents a positive integer, 1≦ i ≦ N p ; a i Represents a specific sample samp i The average distance between other similar samples; b i For sample samp i The average distance between sample points of different categories excluding itself; according to N p The silhouette coefficient of each sample is calculated. All modal components of the same cluster are reorganized using the weighted average method. The energy of each modal component is defined as: ; For each cluster group , calculate the normalized weight based on the energy: ;in, j is a positive integer, It is k The energy of the modal components, It is i The energy of the modal components; Represents the modal component in the cluster group to which it belongs The weight within Then the reconstructed modal component is expressed as: .
[0012] The present invention also includes an office building energy consumption prediction system, comprising: The collection module is used to collect the load sequence of different energy-using equipment in the current office building; A graph structure construction module is used to decompose the load sequences of different energy - using devices into multiple modal components, calculate the sample entropy for each modal component; recombine the corresponding modal components according to the sample entropy of each modal component; use the recombined modal components as nodes, calculate the temporal similarity between each modal component using dynamic time warping, and use the temporal similarity as the edge between nodes to construct a graph structure; A prediction module is used to input the graph structure into the GCN - Transformer model, extract the spatial features of the nodes in the graph structure using the graph convolution operation of the graph convolutional network GCN, input the spatial features of the nodes into the Transformer model, extract the temporal features of the spatial features by introducing the self - attention mechanism to obtain spatio - temporal features; obtain the predicted energy consumption value for a future period of time by inputting the spatio - temporal features into a fully - connected layer.
[0013] The present invention provides an office building energy consumption prediction method, which has the following beneficial effects: In the present invention, the collaborative energy - using relationship between modal components is modeled by calculating the graph convolutional network (GCN). The modal components are recombined according to the sample entropy, and each recombined modal component is regarded as a node in the graph structure. The temporal similarity between modal components is calculated by dynamic time warping, ensuring that the model can automatically identify the implicit associations between modal components, thereby enhancing the prediction ability of future energy consumption trends; the spatial features between nodes are extracted by using graph convolution operations, effectively capturing the collaborative energy - using relationship between modal components. This spatial modeling ability enables the model to identify complex associations between different modal components. At the same time, the Transformer model enhances the time - dependence modeling ability through its self - attention mechanism, which can globally focus on the long - term trends and periodic changes in historical load data, so as to better capture the long - range dependence relationship in the energy consumption sequence; the combination of such spatial and temporal features enables the model to more comprehensively capture the multi - scale characteristics of the energy consumption sequence. Especially in the face of complex and changing energy consumption scenarios, the model can take into account both spatial collaboration and time dependence, thus providing more accurate prediction results. Description of the Drawings
[0014] Figure 1 It is the flowchart of the office building energy consumption prediction method in the embodiment of the present invention; Figure 2 It is the framework diagram of the Transformer model in the embodiment of the present invention. Detailed Embodiments
[0015] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0016] The present invention proposes an office building energy consumption prediction method based on sample entropy sequence component recombination GCN-Transformer. By analyzing the frequency division technology, the loads with different electricity consumption principles are decomposed into a group of modal components with typical spectral characteristics, and GCN is used to find the collaborative energy consumption relationship of each group of loads, so as to improve the prediction accuracy. As Figure 1 shown, the method specifically includes the following steps: S1. Data collection.
[0017] Data collection: Obtain energy consumption data from the energy consumption monitoring system, smart meters or relevant energy management platforms, that is, time series load data, simply referred to as load sequences. These data may include energy consumption values such as electricity, gas, and water at different time scales (such as hourly, daily, monthly). Ensure that the data covers a long enough time period to capture various change patterns of the energy consumption data.
[0018] S2. Data preprocessing: Data cleaning: Remove noise and missing values. For missing values, methods such as mean filling ( is the data point, n is the number of data points), or the linear interpolation formula ( t is the time point corresponding to the missing value, and are adjacent time points, and are the data values at the corresponding time points) can be used.
[0019] Data standardization: Min-Max standardization: , where x is the original data, and are the minimum and maximum values in the data set respectively.
[0020] Z-score standardization: , is the mean of the data set, is the standard deviation.
[0021] Data division: Divide the data set into a training set, a validation set and a test set according to the ratio of 70%:15%:15%, ensuring that the data distributions of each subset are similar.
[0022] S3. Sequence decomposition and recombination: (1)Load sequence decomposition: Use the variational mode decomposition (VMD) technique to decompose the original energy consumption sequence into multiple modal components with typical spectral characteristics. Variational mode decomposition is a signal processing method that can effectively improve the stationarity of time series with high complexity and strong nonlinearity, and obtain subsequences that contain multiple different frequency scales and are relatively stationary after decomposition.
[0023] Construct the mathematical model of variational mode decomposition: ; where are the modal components, is the k th modal component; are the central frequencies of the modal components, is the central frequency corresponding to the k th mode; is the Dirac function; K is the decomposition layer number; j 1 is the imaginary unit; * is the convolution operator; is the original signal; k represents a positive integer; in the specific solution process, the central frequency of each modal component is obtained through iterative calculation according to the principle of the minimum sum of modal bandwidths, so as to realize the separation of each modal component.
[0024] (2)By introducing the penalty factor and the Lagrange multiplier , the constrained problem is transformed into an unconstrained variational problem.
[0025] ; In the formula, is the Lagrange multiplier, is the penalty factor.
[0026] K is the key parameter of variational mode decomposition. If the value of K is too small, it will lead to under-decomposition and modal mixing phenomenon; on the contrary, if the value of K is too large, it will lead to over-decomposition, generate redundant modes, and increase the prediction calculation amount. In order to improve the accuracy of K parameter setting, the present invention proposes a method for determining the number of modes based on the central frequency, and evaluates the rationality of the value of K by observing the change of the central frequency of the newly generated components during the process of increasing the number of modes K.
[0027] (3)Through iterative optimization, solve the variational mode decomposition model; in each iteration, update and the central frequency until convergence, and obtain the final and the corresponding central frequency.
[0028] (4) Reorganization of sequence components based on sample entropy: Usually, the number of modal components obtained by decomposing the load sequence is large, and each modal component usually has a certain similarity. If each modal component is predicted separately, the time cost of model training will increase significantly and the prediction error will be amplified. Therefore, the present invention introduces sample entropy (SE) to quantify the complexity of each modal component obtained by decomposing the load sequence, and uses it as the basis for modal component reorganization. When the number of samples is limited, the sample entropy of each modal component is calculated, and the specific process is as follows: Each modal component is given by N The time series is composed of data, and each modal component is organized into a set of dimensions according to the sequence number. m A vector subsequence of: ,in, ; Define a vector subsequence with another vector subsequence The distance between is the absolute value of the maximum difference between the corresponding elements of the two: ; In the formula, w Represents a subsequence The starting position index of s Represents another subsequence The starting position index of w different); For a given ,statistics and The distance between them is less than or equal to the similarity threshold r of s ( , ) and record it as ;for , Defined as: ; Indicates that from w Starting from subsequences, how many other m The distance between the dimension subsequence and it is less than or equal to the threshold r; Defined as: ; Yes all The average value of any two m dimensional sequence at the threshold r The probability of matching; Increase the dimension tom +1, calculate and the number of those with a distance less than or equal to r is denoted as ; is defined as: ; represents starting from the w th subsequence, how many other m+ 1-dimensional subsequences have a distance less than or equal to the threshold r; is defined as: ; is the average value of all , representing the matching probability of any two m +1-dimensional sequences under the threshold r ; Then the sample entropy is defined as: ; In the formula: m is the sequence dimension; N is the number of samples.
[0029] According to the calculation result of the sample entropy, the present invention introduces K k-means clustering to recombine each modal component. Since the number of clusters J has an important impact on the clustering effect; the present invention uses the silhouette coefficient to evaluate the clustering effect of each component, and then selects the optimal number of clusters to ensure reducing the training time cost while improving the prediction accuracy. The range of the silhouette coefficient is [-1,1], and the closer it is to 1, the better the clustering effect.
[0030] Assume that the number of samples contained in each modal component is N p , take the i (1 ≤ i ≤ N p )th sample component samp i from multiple different modal components, then N p the silhouette coefficient SC Np of the ; In the formula: a i represents a specific sample sampi The average distance between samples of the same type; b i For sample samp i The average distance between sample points of different categories excluding itself.
[0031] according to N p The silhouette coefficient of each sample is calculated. All modal components of the same cluster are reorganized using the weighted average method. The energy of each modal component is defined as: .
[0032] For each cluster group , calculate the normalized weight based on the energy: ;in, It is k The energy of the modal components, It is i The energy of the modal components; Represents the modal component in its category The weight within.
[0033] Then the reconstructed modal component is expressed as: .
[0034] VMD can effectively separate different modal components by adaptively determining the number of modes, thereby capturing the multi-scale characteristics of the energy consumption series. Sample entropy is calculated for each modal component to evaluate its complexity and randomness. Sample entropy is an indicator to measure the complexity of time series, which can effectively distinguish the characteristics of different modal components. Usually, the number of components obtained by decomposing the load series is large, and each component usually has a certain similarity. If each component is predicted separately, the time cost of model training will be significantly increased, and the prediction error will be amplified. Therefore, sample entropy (SE) is introduced to quantify the complexity of different modal components obtained by decomposing the load series, and this is used as the basis for the reorganization of modal components. According to the sample entropy calculation results, K-means clustering is introduced to reorganize each modal component.
[0035] S4. GCN-Transformer model construction: Use a graph convolutional network (GCN) to extract the collaborative energy consumption features of the load. In the real world, there is a large amount of data that shows the characteristics of disordered arrangement and lack of fixed time or spatial order, which is called non-Euclidean structured data. In addition, the number of adjacent nodes of different nodes may be different, and each node may have a unique surrounding structure. Due to the unique structure of these data, traditional convolutional neural networks are no longer applicable to process them. Graph convolutional neurons interact and extract features of node information along the edges of the graph without destroying the graph structure of the data, which makes graph convolutional neurons have unique advantages in processing graph-structured data. Input the reorganized modal components into the GCN model for prediction. GCN is used to capture the collaborative energy consumption relationship between different modal components. By constructing a graph structure, the reorganized modal components are used as nodes and the load relationship is used as edges, and graph convolutional operations are used to extract the features between nodes.
[0036] Construct edges based on dynamic time warping (DTW). The core idea is to measure the temporal similarity between different modal components (IMFs). Calculate the DTW distance between any two modal components. If the DTW distance between two modal components is less than the set threshold θ, an edge is constructed. The specific steps are as follows: Let X and Y be two modal component sequences with length T : Calculate the Euclidean distance between each time step i , j to obtain a cost matrix of size T×T.
[0037] Recursively calculate the DTW distance matrix: ; Finally DTW the distance is: ; The calculation process of graph convolutional operations can be expressed as: ; In the formula: represents the node feature matrix of the l th layer, that is, the node representation of the network at the l th layer; represents the weight matrix of the l th layer, used to perform linear transformation on the nodes; is the adjacency matrix of the graph plus a self-connection, representing the connection relationship between nodes; is the degree matrix, which is a matrix composed of the diagonal elements of the adjacency matrix A; is the activation function.
[0038] Enhancing Temporal Dependence by Incorporating Transformer: The VMD-GCN load forecasting model mainly relies on GCN to learn the collaborative relationships between modal components (IMFs). However, it has certain limitations in time series modeling. GCN mainly propagates features based on a static adjacency matrix, insufficiently utilizes time information, and is difficult to capture the long-term change trends of the load. At the same time, it can only handle the collaborative relationships between local modal components and cannot dynamically adjust the correlations between modal components, resulting in the prediction results being sensitive to short-term fluctuations and having limited ability to identify sudden abnormal loads. To address these issues, after introducing Transformer, the Transformer model framework diagram is as shown in Figure 2 Figure [Figure number not provided in the original]. The model enhances the temporal dependence modeling ability through the self-attention mechanism, enabling it to globally focus on historical load data, breaking through the limitation of GCN that can only learn local load collaborative relationships, thereby more accurately predicting long-term trends and reducing errors caused by short-term fluctuations. In addition, Transformer can effectively capture long-range dependence relationships in the sequence, further improving the prediction accuracy and making the load prediction more stable and reliable.
[0039] Calculate the similarity between each modal component, and use dynamic time warping (DTW) to calculate the temporal distance between modal components: Let X and Y be two modal component sequences with length T: Calculate the Euclidean distance i 、 j between each time step , obtaining a cost matrix of size T×T; where, is the value of sequence X at the i -th time step; is the value of sequence Y at the j -th time step; For each point ([[]] i , j ), calculate its minimum cumulative distance: ; where, represents the cumulative distance moving from above to ([[]] i , j ), represents the cumulative distance moving from the left to ([[]] i , j ); represents the cumulative distance moving from the diagonal to ([[]] i , j ); Gradually calculate the i , j of each point ([[]] ), until (T,T).
[0040] Then the DTW distance between sequences X and Y is as follows: .
[0041] Construct the GCN spatial graph structure: Define the graph , the node set represents each modal component, and the edge set is constructed from the similarity between each modal component. If is less than the set threshold , then an edge is established; the adjacency matrix A is composed of , which is expressed as: ; According to the adjacency matrix, construct the normalized Laplacian matrix: ; where is the degree matrix, and its diagonal element D ii represents the degree of node i , which is expressed as .
[0042] The calculation formula of the GCN layer is: ; where represents the node feature matrix of the l -th layer, that is, the node representation of the network at the l -th layer; represents the weight matrix of the l -th layer, which is used to perform a linear transformation on the nodes; is the adjacency matrix of the graph structure, representing the connection relationship between nodes; is the degree matrix, which is a matrix composed of the diagonal elements of the adjacency matrix A ; is the activation function; is the feature matrix of the (l + 1)-th layer, which is the spatial feature extracted after passing through the GCN layer.
[0043] 4) Transformer extracts temporal features: After GCN processes the spatial relationship, the output is fed into the Transformer to extract the temporal dependence, and the input sequence is: ; where is the positional encoding.
[0044] ; Among them, t is the current time step (time index), d is the total dimension of the input features.
[0045] In the multi-head attention layer, the input features will generate query ( ), key ( ), and value matrices through linear transformation: .
[0046] Then calculate the self-attention weights: ; Among them: the query matrix , the key matrix , the value matrix , is the feature dimension, is the trainable parameter.
[0047] Using multi-head attention ( h heads), then calculate h groups of attention outputs in parallel: ; Among them h is the number of heads, is the transformation matrix.
[0048] Residual connection: ; Feed-forward network: ; Among them, W 1 is the weight matrix of the first layer, b 1 is the bias of the first layer, W 2 is the weight matrix of the second layer, b 2 is the bias of the second layer.
[0049] After passing through the feed-forward network, the obtained feature F is connected with through residual connection, and then through layer normalization to obtain: ; Output: Send the final feature extracted by the Transformer into the fully connected layer for prediction, and obtain the future energy consumption prediction value : ; Among them, is the MLP network.
[0050] S5. Model Training and Prediction: The mean squared error (MSE) loss is used as the loss function to measure the error between the model's predicted values and the true load values, which is defined as follows: ; where is the predicted value of the model, is the true load value, is the number of samples.
[0051] To prevent the model from overfitting, regularization is introduced to constrain the model, and the regularization term is defined as: ; where is the regularization coefficient.
[0052] The final total loss function consists of the mean squared error loss and the regularization term: ; where is the weight of the regularization term in the total loss.
[0053] The model is trained using the training set data. By iteratively adjusting the model parameters, the loss function is minimized. In each iteration, the gradient of the loss function with respect to the model parameters is calculated, and the parameters are updated using an optimizer.
[0054] Model Performance Evaluation: The trained model is applied to the test set, and each time series sample in the test set is predicted. The energy consumption prediction values for a future period (e.g., the next 24 hours or the next week) are obtained. Multiple performance metrics are calculated, including but not limited to: Mean Absolute Error (MAE), which is used to evaluate the average absolute deviation between the predicted values and the true values: .
[0055] where is the predicted value at the i th time point, is the true value at the i th time point.
[0056] Root Mean Square Error (RMSE), which punishes large errors more severely and can better reflect the overall level of prediction errors: .
[0057] Mean Absolute Percentage Error (MAPE), which represents the prediction error in the form of a percentage, making it easier to intuitively understand the degree of error: .
[0058] Compare the performance metrics of the model (such as mean squared error MSE, mean absolute error MAE, coefficient of determination R², etc.) under different building types (such as office buildings, shopping malls, residences, etc.) and different time scales (such as hours, days, months), and analyze its performance in different scenarios.
[0059] To visually display the prediction accuracy of the model, a comparison graph of the actual energy consumption and the predicted energy consumption can be drawn. Use a line graph to present the changing trends of both over time and observe the matching situation between the predicted values and the actual values. At the same time, use a scatter plot to show the relationship between the actual values and the predicted values, and add a reference line of y = x to evaluate the deviation degree of the predicted values. In addition, mark the areas with large errors in the graph and analyze the possible reasons, such as abnormal events or data noise. Furthermore, conduct a statistical analysis of the prediction errors to evaluate the reliability of the model. Calculate the 95% confidence interval to measure the fluctuation range of the prediction errors, and analyze the distribution of the errors (such as normal distribution, skewed distribution) to determine whether there is a systematic bias in the model (such as continuous overestimation or underestimation). In addition, calculate the statistical indicators of the errors (such as mean, standard deviation, skewness, kurtosis) to more precisely quantify the prediction performance of the model and provide a basis for optimizing the model.
[0060] The effects that the present invention can achieve: (1) Extract load collaborative relationship features: Model the collaborative energy consumption relationship between modal components through a graph convolutional network (GCN). Regard each modal component as a node in the graph structure, and calculate the temporal similarity between modal components through dynamic time warping (DTW) to construct an adjacency matrix, ensuring that the model can automatically identify the implicit associations between modal components. Thereby enhancing the prediction ability for future energy consumption trends. Traditional energy consumption prediction methods usually rely on time series modeling with a fixed window and are difficult to dynamically adjust the relationship between modal components. While GCN can adjust the load collaborative relationship under different time periods and different building usage scenarios through the adaptive learning of the graph structure, improving the generalization ability of the model and making the prediction more stable and accurate.
[0061] (2) Integrating spatial and temporal features: The combination of the Graph Convolutional Network (GCN) and the Transformer model enables the comprehensive modeling of the spatial collaborative relationships and temporal dependencies of modal components. By constructing a graph structure, GCN takes different modal components as nodes and temporal similarity as edges, and uses graph convolution operations to extract the spatial features between nodes, effectively capturing the collaborative energy consumption relationships between modal components. This spatial modeling ability allows the model to identify complex associations between different modal components, such as the mutual influence between air conditioners, lighting, and equipment power consumption. At the same time, the Transformer model enhances the temporal dependency modeling ability through its Self-Attention mechanism, which can globally focus on the long-term trends and periodic changes in historical load data. The Self-Attention mechanism enables the model to dynamically allocate weights and focus on the historical time points that are most important for the current prediction, thus better capturing the long-range dependencies in the energy consumption sequence. The combination of these spatial and temporal features enables the model to more comprehensively capture the multi-scale characteristics of the energy consumption sequence: This dual modeling ability significantly improves the stability and reliability of the prediction. Especially in the face of complex and changing energy consumption scenarios, the model can take into account both spatial collaboration and temporal dependence, thus providing more accurate prediction results.
[0062] (3) Reducing computational complexity: Traditional energy consumption prediction methods usually require separate modeling for each modal component. Especially after using techniques such as Variational Mode Decomposition (VMD) to decompose the original energy consumption sequence into multiple components, the model needs to process the features and prediction tasks of each component separately. This component-by-component processing method not only leads to a significant increase in computational volume but also significantly prolongs the model training time. Especially when facing large-scale energy consumption data, the issues of computational resource consumption and training efficiency are particularly prominent. To solve this problem, the present invention innovatively introduces sample entropy to quantify the complexity of different modal components obtained from the decomposition of the load sequence. Sample entropy can effectively measure the randomness and complexity of time series, thereby distinguishing the characteristics of different modal components. By calculating the sample entropy of each component, the model can identify modal components with similar complexity or characteristics, and then use the K-means clustering algorithm to reorganize these components. This clustering and reorganization method based on sample entropy reduces the number of modal components that need to be modeled separately, further optimizing the efficiency of the model and making the prediction process more efficient.
[0063] (4) Improving the generalization ability of the model: By introducing regularization terms and dynamic time warping technology, the model overfitting problem is effectively prevented, and the generalization ability of the model is improved from different angles. First, by adding the L2 regularization term to the loss function, the model can constrain the growth of parameters and avoid overfitting caused by excessive parameter values. Secondly, the dynamic time warping (DTW) technology is used to dynamically adjust the correlation between modal components, further enhancing the generalization ability of the model. DTW can flexibly capture the dynamic relationship between modal components by measuring the temporal similarity between different modal components, rather than relying solely on fixed time alignment. This method is particularly suitable for processing energy consumption data with nonlinear time offset, thereby reducing the impact of abnormal data on prediction results.
[0064] (5) Dynamically adjust the correlation of modal components: The present invention uses dynamic time warping (DTW) to construct a graph structure to dynamically measure the temporal similarity between different modal components. By setting the threshold θ, the model can automatically adjust the correlation between them according to the similarity of different modal components, making feature extraction more accurate and effectively avoiding the limitations brought by fixed correlation settings.
[0065] The advantage of the present invention is that it enhances the model's ability to identify sudden abnormal loads, especially in scenarios with large load fluctuations such as office buildings, and can effectively deal with sudden energy consumption changes caused by factors such as meetings, centralized startup of office equipment, or temporary increase or decrease of personnel. In addition, the dynamic adjustment mechanism enables the model to respond quickly to sudden load changes, avoiding the decline in prediction performance due to short-term drastic fluctuations, and improving the robustness and stability of the overall prediction.
[0066] Based on the same inventive concept, the present invention also proposes an office building energy consumption prediction system, comprising: The collection module is used to collect the load sequence of different energy-using equipment in the current office building.
[0067] The graph structure construction module is used to decompose the load sequence of different energy-using equipment into multiple modal components and calculate the sample entropy of each modal component; according to the sample entropy of each modal component, the corresponding modal components are reorganized; each reorganized modal component is used as a node, and dynamic time warping is used to calculate the timing similarity between each modal component, and the time similarity is used as the edge between the nodes to construct a graph structure.
[0068] A prediction module is used to input the graph structure into the GCN-Transformer model, extract the spatial features of nodes in the graph structure by using the graph convolution operation of the graph convolutional network (GCN), input the spatial features of the nodes into the Transformer model, extract the temporal features of the spatial features by introducing the self-attention mechanism, and obtain spatio-temporal features; by inputting the spatio-temporal features into a fully connected layer, an energy consumption prediction value for a future period of time is obtained.
[0069] As mentioned above, the above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. An energy consumption prediction method for office buildings, characterized in that, It includes the following steps: Collect the load sequences of different energy-using devices in the current office building; Decompose the load sequences of different energy-using devices into multiple modal components, calculate the sample entropy for each modal component; according to the sample entropy of each modal component, reorganize its corresponding modal component; use the reorganized modal components as nodes, and use dynamic time warping to calculate the temporal similarity between each modal component, and use the temporal similarity as the edge between nodes to construct a graph structure; Input the graph structure into the GCN-Transformer model, use the graph convolution operation of the graph convolutional network GCN to extract the spatial features of the nodes in the graph structure, input the spatial features of the nodes into the Transformer model, and extract the temporal features of the spatial features by introducing the self-attention mechanism to obtain spatio-temporal features; by inputting the spatio-temporal features into the fully connected layer, obtain the energy consumption prediction value for a future period of time.
2. The energy consumption prediction method for an office building according to claim 1, wherein Obtain the similarity between each modal component through Dynamic Time Warping (DTW). , which specifically includes the following steps: Define that X and Y are two sequences of modal components, with a length of T : Calculate for each time step i , j the Euclidean distance to obtain a cost matrix of size T×T; where x i is the value of sequence X at the i -th time step; y j is the value of sequence Y at the j -th time step; For each point ( i , j ), calculate its minimum cumulative distance: ; Among them, represents the cumulative distance from moving from above to ( i , j ); represents the cumulative distance from moving from the left to ( i , j ); represents the cumulative distance from moving diagonally to ( i , j ); Calculate step by step for each point ( i , j ) of , until ; Then the DTW distance between sequences X and Y is: .
3. The energy consumption prediction method for an office building according to claim 2, characterized in that, The extraction of the spatial features of the nodes in the graph structure by using the graph convolution operation of the graph convolutional network GCN is expressed as: ; Among them, represents the node feature matrix of the l th layer; represents the weight matrix of the l th layer, which is used to perform a linear transformation on the nodes; is the degree matrix, which is a matrix composed of the diagonal elements of the adjacency matrix A; is the activation function; the adjacency matrix A is composed of ; is expressed as: ; among them, is the set threshold; is the Laplacian matrix, .
4. A method for predicting the energy consumption of an office building according to claim 3, characterized in that, The input of the spatial features of the nodes into the Transformer model and the extraction of the temporal features of the spatial features by introducing the self-attention mechanism specifically include the following steps: The output Extract the time dependence in the input Transformer. Since the Transformer requires positional information, the input sequence of the Transformer is represented as: , where , is the positional encoding, represented as: , t is the current time step, d is the total dimension of the input features; In the multi-head attention layer, the input features generate query, key, and value matrices through linear transformation; Calculate self-attention weights based on the query, key, and value matrices, and calculate the multi-head attention output in parallel according to the self-attention weights to obtain ; Perform a residual connection between the multi-head attention output and the input sequence to obtain features ; Input into the feedforward network for calculation to obtain features ; The obtained features are subjected to residual connection with and then layer normalization to obtain the final temporal features extracted by the Transformer .
5. A method for predicting the energy consumption of an office building according to claim 1, characterized in that, Use the variational mode decomposition method to decompose the load sequence into multiple modal components, which specifically includes the following steps: Construct the mathematical model of variational mode decomposition: ; wherein are the modal components, is the k th modal component; is the center frequency of each modal component, is the k th center frequency corresponding to the mode; is the Dirac function; K is the decomposition level; j 1 is the imaginary unit; * is the convolution operator; is the original signal; k represents a positive integer; By introducing a penalty factor and the Lagrange multiplier , the mathematical model of variational mode decomposition is transformed into: ; where is the Lagrange multiplier, is the penalty factor; Solve the variational mode decomposition model through iterative optimization; in each iteration, update and the central frequency until convergence to obtain the final and the corresponding central frequency.
6. The energy consumption prediction method for an office building according to claim 5, characterized in that The calculation of the sample entropy for each modal component specifically includes the following steps: Group each modal component by sequence number to form a vector subsequence with a dimension of m : , where ; each modal component is a time series composed of N data. Define a vector subsequence and another vector subsequence The distance between is the absolute value of the maximum difference among the corresponding elements of the two; where w represents the starting position index of the subsequence , s represents the starting position index of another subsequence . For a given ,statistics and The distance between them is less than or equal to the similarity threshold r of s The number of , , ;for , from w subsequences, the others m The distance between the dimension subsequence and its subsequence is less than or equal to the threshold r The number is defined as ; Calculate all average value , this average value represents the probability of matching between any two m dimensional sequences at the threshold r ; Increase the dimension to m +1 and calculate the number of with a distance less than or equal to r denoted as , ; Calculate starting from the w th subsequence, the number of other m+ 1-dimensional subsequences with a distance less than or equal to the threshold r denoted as ; Calculate all the average value of , which represents the probability of matching between any two m +1-dimensional sequences at the threshold r ; The sample entropy is expressed as: ; Wherein: m is the sequence dimension; N is the number of samples.
7. A method for predicting the energy consumption of an office building according to claim 6, wherein, The reorganization of the corresponding modal component according to the sample entropy of each modal component specifically includes: Define the number of samples included in the modal components of each different frequency feature as N p , and take the i th sample component from multiple different subsequences samp i . Then, N p silhouette coefficients SC of the samples Np are expressed as: ; Among them, i represents a positive integer, 1 ≤ i ≤ N p ; a i represents the average distance between a specific sample samp i and other similar samples; b i is the average distance between a sample samp i and the sample points of different categories excluding itself; According to N p the silhouette coefficients of the samples, for all modal components in the same cluster, the weighted average method is used for recombination, and the energy of each modal component is defined as: ; For each clustering group , calculate the normalized weight based on energy: ; where j is a positive integer, is the energy of the k th modal component, is the energy of the i th modal component; represents the weight of the modal component within its affiliated clustering group . Then the reconstructed modal component representation is obtained as follows: .
8. An office building energy consumption prediction system, characterized in that, It includes: A collection module for collecting the load sequences of different energy-using devices in the current office building; A graph structure construction module for decomposing the load sequences of different energy-using devices into multiple modal components, calculating the sample entropy for each modal component; according to the sample entropy of each modal component, reorganizing its corresponding modal component; using the reorganized modal components as nodes, and using dynamic time warping to calculate the temporal similarity between each modal component, and using the temporal similarity as the edge between nodes to construct a graph structure; A prediction module for inputting the graph structure into the GCN-Transformer model, using the graph convolution operation of the graph convolutional network GCN to extract the spatial features of the nodes in the graph structure, inputting the spatial features of the nodes into the Transformer model, and extracting the temporal features of the spatial features by introducing the self-attention mechanism to obtain spatio-temporal features; by inputting the spatio-temporal features into the fully connected layer, obtain the energy consumption prediction value for a future period of time.
Citation Information
Patent Citations
Short-term power load prediction method combining VMD decomposition and time convolution network
CN114358389A
TGCN-GRU ultra-short-term load prediction method and device based on VMD and electronic equipment
CN114548532A
Short-term power load prediction method combining CDC-VMD and echo state network
CN116090627A
Space-time wind speed prediction method and device and medium
CN117828997A
Payment load prediction method based on spatial-temporal feature clustering and double-layer dynamic graph convolution
CN118535938A
Cited By
Three-dimensional human body posture estimation method and system
CN120977018A
A method and system for three-dimensional human pose estimation
CN120977018B
Apartment air conditioner energy consumption optimization regulation and control method and system based on electricity utilization data analysis
CN121112453A
Building cluster energy consumption dynamic optimization method and system based on big data and AI prediction
CN121724203A