Electrical load prediction method and device based on spatial-temporal correlation
By constructing a multi-layer spatiotemporal model and using interpretable spatiotemporal attention converter to extract features and combining STEF-DHNet model for prediction, the problem of low accuracy and reliability of power load prediction in the prior art is solved, and more efficient and interpretable prediction results are achieved.
Patent Information
- Application Number
- PCT/CN2024/129425
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-11-01
- Publication Date
- 2025-05-08
AI Technical Summary
When the existing electrical load prediction methods deal with spatiotemporal correlations, the model complexity increases, resulting in a decrease in interpretability, which in turn affects the accuracy and reliability of the prediction.
The power load prediction method based on spatiotemporal correlation is adopted to construct a multi-layer spatiotemporal model through spatiotemporal modeling, and the spatial and temporal characteristics of power loads are extracted using an interpretable spatiotemporal attention converter and fuse it into the STEF-DHNet model for prediction.
It improves the accuracy and model applicability of power load prediction, reduces the complexity of the model, and enhances the interpretability of the prediction results.
Smart Images

Figure CN2024129425_08052025_PF_FP_ABST
Abstract
Description
Power load forecasting method and device based on time-space correlation
[0001] This application claims priority to Chinese patent application No. CN202311452701.4, filed on November 3, 2023, entitled "Power Load Forecasting Method and Apparatus Based on Spatiotemporal Correlation." The disclosure of that prior application is incorporated herein by reference in its entirety. Technical Field
[0002] The present application relates to the technical field of load forecasting, and in particular to a method and device for predicting electricity load based on spatiotemporal correlation. Background Art
[0003] Electricity load forecasting is an essential component of power system planning and the foundation of its economic operation. It is crucial for both planning and operation. Electricity load is affected by both time and space. These two influences are not independent but rather intertwined, hence the term "electricity load has spatiotemporal dependence."
[0004] The current forecasting method for electricity load data with temporal and spatial dependencies is to construct temporal and spatial correlations separately and then combine them to obtain the spatiotemporal correlation. However, this method increases the complexity of the model, reduces the interpretability of the spatiotemporal correlations, and makes it difficult to verify the credibility of the model, resulting in low accuracy and reliability of electricity load forecasts. Technical issues
[0005] The embodiments of the present application provide a method and device for electricity load prediction based on spatiotemporal correlation to solve the problem of low accuracy and reliability of electricity load prediction. Technical Solutions
[0006] In a first aspect, an embodiment of the present application provides a method for predicting power load based on spatiotemporal correlation, comprising:
[0007] Based on the historical power load data and historical environmental data of different power consumption nodes, spatiotemporal modeling is performed to obtain a multi-layer spatiotemporal model. Each layer of the model is a directed graph, representing the spatial causal relationship at a fixed time, and the directed edges between each layer of the model represent the temporal causal relationship. The multi-layer spatiotemporal model is:
[0008] Among them, V t ={v 1,t ,…,v 4,t};
[0009] The causal structure of the multi-layer spatiotemporal model is defined based on Assumptions 1 and 2. Assumption 1 states that for s, s'∈S, t, t'∈T, the multi-layer network G = (v, ε, τ) satisfies the following conditions:
[0010] ①Assume t≤t', then
[0011] ②For s≠s′ and (v s,t ,v s′,t )∈ε, if and only if (v s,t ,v s′,t′ )∈ε; In addition, If and only if
[0012] In the above hypothesis 1, ① represents irreversibility, and ② represents spatial causality that is uniform over time;
[0013] The hypothesis 2 is: (v s,1 ,v s,1 )∈ε, if (s,s′)∈{(1,3),(1,4),(2,4),(4,3)}, otherwise,
[0014] The aforementioned hypothesis 2 represents the spatial causal relationship based on prior knowledge;
[0015] The power consumption nodes include B1, B2, B3, B4, and B5. B5 is the upstream node of B1, B2, B3, and B4, that is, the power consumption of B1, B2, B3, and B4 comes from B5. The environmental data includes three different temperatures P1 , P1 , P3, and power consumption control strategy D, C1 = {P1, P1, P3}, C2 = {B5}, C3 = {D} and C4 = {B1, B2, B3, B4}, node v s,t Represents cluster C s All available observations in , s∈S={1,2,3,4}, and time t∈T={1,...,τ};
[0016] Based on the interpretable spatiotemporal attention transformer, the multi-layer spatiotemporal model is subjected to feature extraction and prior knowledge fusion to obtain the spatiotemporal characteristics of power load. The interpretable spatiotemporal attention transformer includes a spatial causal attention network, a temporal attention network, and a spatial dependency comparison module.
[0017] The spatial causal attention network is:
[0018] in, yes The self-attention is from arrive The mapping has an attention weight matrix of d1×d2 The triplet of As a spatial feature matrix, it corresponds to X weighted by the causal constraint weights in Assumption 2 t ;
[0019] The temporal attention network is:
[0020] in, M T is the decoder mask; M T The upper diagonal elements in are -∞, so that (A (1) ) ij = 0, i < j, and the i-th row of the temporal attention network consists only of The j-th row vector of
[0021] The temporal attention network converts the following Merge as input value and use the following formula As query and key, the dimension of query and key is reduced through variable selection network;
[0022] Where vec(·) represents the flattening mapping,
[0023] When tB<t′≤t, each row yes Corresponding to The reduced vector of ;
[0024] The spatial dependency comparison module is:
[0025] in, The contrasting features of Given, for The reconstruction matrix, The (t′-t+B) row vector of and Both are obtained through spatial causal attention networks, The attention weights are used to explain and quantify spatial effects;
[0026] Scan comparison After that, the final output of the encoder of the interpretable spatiotemporal attention transformer is given by establish; Input to VSN2(·), VSN2(·) transmits the summary information to the decoder;
[0027] When tB<t′≤t, In the formula The connection matrix with the context vector, where:
[0028] Evaluate variable returns by using the variable selection weights of VSN2(·) the importance of
[0029] Based on the spatiotemporal characteristics of power load, the real-time power load data of the target power node and the real-time environmental data, the trained STEF-DHNet model is used for prediction to obtain the predicted power load data of the target power node at a specified future time.
[0030] In one possible implementation, spatiotemporal modeling is performed based on historical power load data and historical environmental data of different power consumption nodes. The resulting multi-layer spatiotemporal model includes:
[0031] The historical power load data and historical environmental data of different power consumption nodes are layered according to time to obtain a multi-layer network;
[0032] Based on the influence relationship between the historical power load data and historical environmental data of different power consumption nodes, similar data are divided into the same cluster to obtain multiple clusters;
[0033] Construct temporal causal hypothesis relationships and spatial causal hypothesis relationships based on the influence relationships between clusters;
[0034] Based on the self-attention module, the temporal causal hypothesis relationship and the spatial causal hypothesis relationship are screened to obtain the spatiotemporal causal relationship;
[0035] Adding spatiotemporal causal relationships to a multi-layer network results in a multi-layer spatiotemporal model.
[0036] In one possible implementation, feature extraction and prior knowledge fusion are performed on the multi-layer spatiotemporal model based on an interpretable spatiotemporal attention transformer to obtain the spatiotemporal characteristics of power load, including:
[0037] Based on the spatial causal attention network, the multi-layer spatiotemporal model is converted into multiple spatial feature matrices;
[0038] Based on the temporal attention network, each spatial feature matrix is compressed to obtain multiple spatial feature vectors, and each spatial feature vector is combined into a spatiotemporal feature matrix according to the temporal causal relationship;
[0039] Based on the spatial dependency comparison module, spatial causal prior knowledge is added to the spatiotemporal feature matrix to obtain the spatiotemporal characteristics of power load.
[0040] In one possible implementation, the STEF-DHNet model includes a first convolutional layer, a second convolutional layer, a flattening module, L fully connected layers, and an LSTM layer connected in sequence; where L is the number of input data.
[0041] In one possible implementation, based on the spatiotemporal characteristics of power load, the real-time power load data of the target power node, and real-time environmental data, a trained STEF-DHNet model is used for prediction. The predicted power load data of the target power node at a specified future time includes:
[0042] Based on the first convolution layer and the second convolution layer, feature extraction is performed on the real-time power load data of the target power node to obtain power load features;
[0043] The real-time environmental data of the target power consumption node is gridded and combined with the power load characteristics to input into the flattening module to obtain flattened data;
[0044] Extract predictive features from the flattened data based on L fully connected layers;
[0045] Based on the LSTM layer, the prediction features and spatiotemporal features of power load are predicted to obtain the predicted power load data of the target power node at a specified future time.
[0046] In one possible implementation, before using the trained STEF-DHNet model for prediction, the following is also included:
[0047] Construct training sets, validation sets, and test sets based on the historical power load data and historical environmental data of the target power consumption nodes;
[0048] Using the mean absolute error as the loss function, the STEF-DHNet model is adversarially trained based on the training set, validation set, and test set to obtain the trained STEF-DHNet model.
[0049] In one possible implementation, the STEF-DHNet model is adversarially trained based on the training set, validation set, and test set. The trained STEF-DHNet model includes:
[0050] Generate adversarial samples based on the training set and the policy network; the policy network consists of a sequentially connected spatiotemporal encoder, a spatial layer, a temporal layer, and a multi-head attention decoder;
[0051] The STEF-DHNet model is adversarially trained based on the reward function, training set, and adversarial samples, and verified based on the validation set and test set to obtain the trained STEF-DHNet model.
[0052] In one possible implementation, verification based on the validation set and the test set includes:
[0053] During the training process, the rolling error of the STEF-DHNet model is updated based on the test set and the given time length. If the rolling error is qualified, the training is completed, otherwise the training continues.
[0054] In one possible implementation, constructing a training set, a validation set, and a test set based on historical power load data and historical environmental data of a target power node includes:
[0055] Clean the historical power load data and historical environmental data of the target power consumption node to obtain clean data;
[0056] Smoothing the clean data to obtain smoothed data;
[0057] Add time information to the smoothed data, and use the historical power load data and historical environmental data at the same time as a sample data;
[0058] Extract features from each piece of sample data, add the extracted features to the corresponding sample data, obtain multiple pieces of feature sample data and form a data set;
[0059] The dataset is divided into training set, validation set and test set according to the preset ratio.
[0060] In a second aspect, an embodiment of the present application provides a device for predicting power load based on spatiotemporal correlation, comprising:
[0061] The spatiotemporal modeling module is used to perform spatiotemporal modeling based on historical power load data and historical environmental data of different power consumption nodes to obtain a multi-layer spatiotemporal model. Each layer of the model is a directed graph, representing the spatial causal relationship at a fixed time, and the directed edges between each layer of the model represent the temporal causal relationship. The multi-layer spatiotemporal model is:
[0062] Among them, V t ={v 1,t ,…,v 4,t};
[0063] The causal structure of the multi-layer spatiotemporal model is defined based on Assumptions 1 and 2. Assumption 1 states that for s, s'∈S, t, t'∈T, the multi-layer network G = (ν, ε, τ) satisfies the following conditions:
[0064] ①Assume t≤t', then
[0065] ②For s≠s′ and (v s,t ,v s′,t )∈ε, if and only if (vs,t′ ,v s′,t′ )∈ε; In addition, If and only if
[0066] In the above hypothesis 1, ① represents irreversibility, and ② represents spatial causality that is uniform over time;
[0067] The hypothesis 2 is: (v s,1 ,v s′,1 )∈ε, if (s,s′)∈{(1,3),(1,4),(2,4),(4,3)}, otherwise,
[0068] The aforementioned hypothesis 2 represents the spatial causal relationship based on prior knowledge;
[0069] The power consumption nodes include B1, B2, B3, B4, and B5. B5 is the upstream node of B1, B2, B3, and B4, that is, the power load of B1, B2, B3, and B4 comes from B5. The environmental data includes three different temperatures P1, P1, and P3, as well as the power consumption control strategy D. C1 = {P1, P1, P3}, C2 = {B5}, C3 = {D}, and C4 = {B1, B2, B3, B4}. Node v s,t Represents cluster C s All available observations in , s∈S={1,2,3,4}, and time t∈T={1,...,τ};
[0070] A feature extraction module is used to extract features and fuse prior knowledge from the multi-layer spatiotemporal model based on an interpretable spatiotemporal attention transformer to obtain the spatiotemporal characteristics of power load. The interpretable spatiotemporal attention transformer includes a spatial causal attention network, a temporal attention network, and a spatial dependency comparison module.
[0071] The spatial causal attention network is:
[0072] in, yes The self-attention is from arrive The mapping has an attention weight matrix of d1×d2 The triplet of As a spatial feature matrix, it corresponds to X weighted by the causal constraint weights in Assumption 2 t ;
[0073] The temporal attention network is:
[0074] in, M T is the decoder mask; M T The upper diagonal elements in are -∞, so that (A (1) ) ij = 0, i < j, and the i-th row of the temporal attention network consists only of The j-th row vector of
[0075] The temporal attention network converts the following Merge as input value and use the following formula As query and key, the dimension of query and key is reduced through variable selection network;
[0076] Where vec(·) represents the flattening mapping,
[0077] When tB<t′≤t, each row yes Corresponding to The reduced vector of ;
[0078] The spatial dependency comparison module is:
[0079] in, The contrasting features of Given, for The reconstruction matrix, The (t′-t+B) row vector of and Both are obtained through spatial causal attention networks, The attention weights are used to explain and quantify spatial effects;
[0080] Scan comparison After that, the final output of the encoder of the interpretable spatiotemporal attention transformer is given by establish; Input to VSN2(·), VSN2(·) transmits the summary information to the decoder;
[0081] When tB<t′≤t, In the formula The connection matrix with the context vector, where:
[0082] Evaluate variable returns by using the variable selection weights of VSN2(·) the importance of
[0083] The load forecasting module is used to predict the target power node's power load at a specified future time using the trained STEF-DHNet model based on the spatiotemporal characteristics of the power load, the real-time power load data of the target power node, and the real-time environmental data. Beneficial effects
[0084] The present application provides a method and device for electricity load forecasting based on spatiotemporal correlation. The present application first utilizes the hierarchical structure of electricity load data and environmental data, and constructs historical electricity load data and historical environmental data of different electricity nodes into a multi-layer spatiotemporal model according to spatial causality and temporal causality, and converts the spatiotemporal dependency of electricity load data and environmental data into spatiotemporal causality; then, based on an interpretable spatiotemporal attention converter, feature extraction is performed and prior knowledge is integrated, and complete spatiotemporal characteristics of electricity load can be obtained without the need to fuse spatial causality and temporal causality; finally, the spatiotemporal causality contained in the spatiotemporal characteristics of electricity load is integrated into the STEF-DHNet model to perform electricity load forecasting, thereby improving the forecast accuracy and model applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0086] FIG1 is a flowchart of an implementation method for power load forecasting based on spatiotemporal correlation according to an embodiment of the present application;
[0087] FIG2 is a flowchart of an implementation of feature extraction and prior knowledge fusion based on an interpretable spatiotemporal attention converter provided by an embodiment of the present application;
[0088] FIG3 is a flowchart of an implementation of prediction based on the STEF-DHNet model provided in one embodiment of the present application;
[0089] FIG4 is a flowchart of an implementation of adversarial training provided in an embodiment of the present application;
[0090] FIG5 is a flowchart of an implementation method for power load forecasting based on spatiotemporal correlation according to another embodiment of the present application;
[0091] FIG6 is a schematic structural diagram of an electricity load prediction device based on spatiotemporal correlation provided in an embodiment of the present application. Modes for Carrying Out the Invention
[0092] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0093] In recent years, due to the complexity of the spatiotemporal data involved, traditional machine learning (ML) models such as support vector machines (SVM), gradient boosting machines (GBM), and their modified extreme gradient boosting (XGBoost) have been used for spatiotemporal prediction. However, with the increasing use of deep learning techniques, attention has shifted to leveraging these models to improve accuracy and efficiency. Deep learning models are able to capture more complex and nonlinear patterns in spatiotemporal data, resulting in more accurate and reliable predictions.
[0094] Graph neural networks are neural models that reflect relationships between graphs through message passing between nodes in the graph. Another type of model that has been proposed is grid-based models, which have their own advantages. These models use a grid architecture to describe the relationship between spatiotemporal variables, typically using a convolutional neural network (CNN) to model the spatial dependencies between different time increments. Many grid-based models, such as the Spatial-Temporal Dynamic Network (STDN) and the Deep Multi-View Spatial-Temporal Network (DMVST-Net), have proven effective in capturing complex spatiotemporal patterns in data and have been applied in many studies.
[0095] This application introduces a grid-based deep learning model, which provides a more accurate representation of real-world scenarios by considering the actual spatiotemporal complexity of external factors, and introduces rolling error to evaluate the accuracy of the model in practical applications. It improves the adversarial robustness of spatiotemporal predictions through a reinforcement-based method, a spatiotemporal attention-based policy network, and a new self-knowledge distillation regularization module. It significantly improves the applicability of the deep learning model through a new converter neural network model based on causal relationships based on prior knowledge.
[0096] In order to make the purpose, technical solutions and advantages of this application clearer, specific embodiments will be described below with reference to the accompanying drawings.
[0097] Example 1:
[0098] Referring to FIG1 , which shows a flow chart of an implementation method of a power load forecasting method based on spatiotemporal correlation provided by an embodiment of the present application, the details are as follows:
[0099] Step 101: Perform spatiotemporal modeling based on historical power load data and historical environmental data of different power consumption nodes to obtain a multi-layer spatiotemporal model; wherein each layer of the model is a directed graph, representing a spatial causal relationship at a fixed time, and the directed edges between each layer of the model represent a temporal causal relationship.
[0100] In this embodiment, a directed graph D refers to an ordered triple (V(D), A(D), ψD), where V(D) is a set of vertices, A(D) is a set of directed edges, and ψD is an association function that maps each directed edge to a pair of vertices. The spatiotemporal forecasting of electricity load differs from other common forecasting problems. Electricity load data is time series data with a temporal relationship, and its values are affected by time-related factors. Compared to common forecasting problems, the characteristics and patterns of the time dimension need to be considered; electricity load data typically exhibits cyclical and seasonal patterns. For example, daily electricity load may have different peaks and valleys during the day and at night; electricity load is affected by a variety of factors, including but not limited to weather, weekends versus weekdays, and holidays. Compared to other common forecasting problems, more external factors need to be considered and incorporated into the forecasting model; the collection of electricity load data is relatively complex, requiring the installation of specialized monitoring equipment and real-time or periodic data collection. Sample data can be more easily obtained in common forecasting problems.
[0101] Collecting electricity load data:
[0102] Install power monitoring equipment: Install power monitoring equipment, such as smart meters and sensors, at locations where load forecasting is needed. Real-time data collection uses appropriate hardware and software tools to collect load data periodically or in real time. Remote data collection can be achieved through IoT technology.
[0103] Collect relevant environmental data: In addition to electricity load data, environmental data related to electricity consumption should also be collected, such as temperature, humidity, weather, electricity control strategies, etc.
[0104] Deep learning-based models have attracted much attention for their intuitive spatiotemporal modeling. There are two mainstream types of deep learning models. The first method combines a typical prediction model with a graph neural network (GNN). GNN-based models use a graph convolution network (GCN) to learn spatial dependencies and a recurrent neural network (RNN), temporal attention, or a temporal convolutional network (TCN) to learn temporal patterns. However, these models have limitations because the model structure is not flexible enough to include heterogeneous types of spatiotemporal predictors across sites. The second method is to restrict the model architecture to enforce spatiotemporal dependencies. These models are designed for specific domains or spatial structures, so their architecture needs to be modified according to the new spatial structure. To this end, in this embodiment, a multi-layer spatiotemporal model is used to characterize the spatiotemporal causal relationship between power load data and environmental data.
[0105] Step 102, based on the interpretable spatiotemporal attention converter, feature extraction and prior knowledge fusion are performed on the multi-layer spatiotemporal model to obtain the spatiotemporal characteristics of the power load; wherein the interpretable spatiotemporal attention converter includes a spatial causal attention network, a temporal attention network and a spatial dependency comparison module.
[0106] In this embodiment, the spatiotemporal attention weights in the Interpretable Spatiotemporal Attention Transformer describe spatiotemporal causal relationships through a multi-layer masked network. This model extends the existing transformer by alternating spatiotemporal masks, thereby incorporating prior knowledge into the model's feature learning. Compared to existing spatiotemporal prediction models, this model has two significant advantages: first, it allows for heterogeneous predictors at each site, allowing for flexible regression to fit the causal network; second, it is applicable to partially determined causal structures. By providing interpretable and diverse information that satisfies temporal causal relationships, this model significantly improves the applicability of deep learning models.
[0107] Step 103: Based on the spatiotemporal characteristics of power load, the real-time power load data of the target power node, and the real-time environmental data, a trained STEF-DHNet (Spatiotemporal External Factors Based Deep Hybrid Network) model is used to perform prediction to obtain the predicted power load data of the target power node at a specified future time.
[0108] In this embodiment, with the popularity of deep learning techniques, researchers are increasingly using these methods to address the challenges of spatiotemporal prediction. These deep learning methods can be divided into two main areas: grid-based and graph-based. Grid-based models have their own advantages. These models use a grid architecture to describe the relationship between spatiotemporal variables, typically using convolutional neural networks (CNNs) to model the spatial dependencies between different time increments. Many grid-based models have been shown to be effective in capturing complex spatiotemporal patterns in data and have been applied in many studies.
[0109] Research has shown that deep neural networks (DNNs) based on grid systems can achieve superior results compared to traditional machine learning techniques. Carefully designing appropriate DNN architectures is crucial for using them to analyze complex spatiotemporal data. One grid-based model utilizes latent representations and recurrent neural networks (RNNs) to capture spatiotemporal dynamics. However, this approach cannot account for periodic patterns in the data.
[0110] Some scholars have proposed the DMVST-Net (Deep Multi-View Spatial-Temporal Network) framework, which uses a long short-term memory network (LSTM) with a temporal view to capture the correlation between future power load and recent time points, uses a spatial view to understand local spatial correlations through a local CNN, and uses a semantic view to identify correlations between regions with similar temporal patterns. However, this approach has limitations because it uses a local CNN and can only work within a small area. It also does not incorporate external factors as spatiotemporal data into the model, which are key components for accurately predicting power load.
[0111] The model leverages the strengths of CNN and LSTM layers to make predictions. It is designed to capture the spatial and temporal dependencies of electricity load data and effectively incorporate external factors into their true complexity. By accounting for the true spatiotemporal complexity of external factors, the model provides a more accurate representation of real-world scenarios.
[0112] The embodiment of the present application first utilizes the hierarchical structure of power load data and environmental data, and constructs the historical power load data and historical environmental data of different power nodes into a multi-layer spatiotemporal model according to spatial causality and temporal causality, and converts the spatiotemporal dependencies of power load data and environmental data into spatiotemporal causality; then, based on an interpretable spatiotemporal attention converter, feature extraction is performed and prior knowledge is integrated, and the complete spatiotemporal characteristics of power load can be obtained without the need to fuse spatial causality and temporal causality; finally, the spatiotemporal causality contained in the spatiotemporal characteristics of power load is integrated into the STEF-DHNet model to perform power load forecasting, thereby improving the forecast accuracy and model applicability.
[0113] In one possible implementation, spatiotemporal modeling is performed based on historical power load data and historical environmental data of different power consumption nodes. The resulting multi-layer spatiotemporal model includes:
[0114] The historical power load data and historical environmental data of different power consumption nodes are layered according to time to obtain a multi-layer network;
[0115] Based on the influence relationship between the historical power load data and historical environmental data of different power consumption nodes, similar data are divided into the same cluster to obtain multiple clusters;
[0116] Construct temporal causal hypothesis relationships and spatial causal hypothesis relationships based on the influence relationships between clusters;
[0117] Based on the self-attention module, the temporal causal hypothesis relationship and the spatial causal hypothesis relationship are screened to obtain the spatiotemporal causal relationship;
[0118] Adding spatiotemporal causal relationships to a multi-layer network results in a multi-layer spatiotemporal model.
[0119] In this example, when the spatiotemporal causal structure is only partially determined, a multilayer network for spatiotemporal modeling is introduced to simultaneously build models based on temporal dependence, spatial causality, and variable correlations. This paves the way for the subsequent interpretable spatiotemporal attention transformer. This module enables representation learning of the multilayer network, determining both spatial and temporal causal relationships, and describing spatiotemporal causal relationships.
[0120] Multilayer networks are a useful tool for modeling patterns between variables with hierarchical structures, such as in biomedicine and community detection. Spatial causal relationships are modeled as time-fixed directed graphs on a single layer. Temporal dependencies between causal graphs are represented by directed edges.
[0121] Use a multi-layer network to learn the representation of spatiotemporal variables. For example, in a certain scenario, the power consumption nodes include B1, B2, B3, B4, and B5, where B5 is the upstream node of B1, B2, B3, and B4, that is, the power load of B1, B2, B3, and B4 comes from B5. The environmental data includes three different temperatures P1, P1, and P3, as well as the power consumption control strategy D. For this scenario, the above factors can be divided into four clusters: C1 = {P1, P1, P3}, C2 = {B5}, C3 = {D}, and C4 = {B1, B2, B3, B4}. Define a node v s,t Represents cluster C s All available observations in , s∈S={1,2,3,4}, and time t∈T={1,...,τ}, τ belongs to the maximum value of time. s,t The causal structure between is modeled by directed edges in the causal graph G, and the temporal dependencies between causal graphs are represented by directed edges. Let G be a multilayer network, a tuple defined by a node set v, an edge set ε, and a layer set, then:
[0122] Where V t ={v 1,t ,…,v 4,t The causal structure of the model is defined by Assumptions 1 and 2:
[0123] Assumption 1 (Temporal Causality): For s, s′∈S, t, t′∈T, the multilayer network G = (v, ε, τ) satisfies the following conditions:
[0124] ①Assume t≤t', then
[0125] ②For s≠s′ and (v s,t ,v s′,t )∈ε, if and only if (v s,t′ ,v s′,t′ )∈ε. In addition, If and only if
[0126] In Assumption 1, ① represents irreversibility and ② represents uniform spatial causality over time. Based on Assumption 2, spatiotemporal causality is established on a multi-layer network.
[0127] Hypothesis 2 (Spatial Causality): (v s,1 ,v s′,1 )∈ε, if (s,s′)∈{(1,3),(1,4),(2,4),(4,3)}. Otherwise,
[0128] Assumption 2 specifically defines spatial causal relationships based on prior knowledge. Directed edges require an estimate of a situation. However, in this study, they are assumed to be known because we focus on building an embedded feature space for spatiotemporal data. If the time indicator is omitted, the causal relationship is summarized as v1→v3, v1→v4, v2→v4, v3→v4, where the arrows are causal relationships. For example, v1→v3 means that v1 is the cause or parent of v3, and v1∈Pa(v3) represents the causal relationship. Use v i ∈Pa(v j ) to represent v i and v j The causal relationship between them.
[0129] Attention is a specific mapping from one sequence to another. Let V be a t×d matrix, each row vector of which represents an element in a sequence of length t, then attention returns a t×d' matrix V' of the input V. In the study, the attention function reduces the characteristics of spatiotemporal data in a multi-layer network. Attention uses two matrices Q∈R t×d and K∈R t× w is associated with the target and input sequence. The attention of V is defined as:
[0130] where softmax is the row-wise softmax function. In particular, the following is the attention weight:
[0131] Identify or control a feature of the training feature from the attention weight. Let a ij For (A) ij , v i and v′ i are the i-th row vectors of V and Attention(Q,K,V) respectively. This is easy to prove. Let M be a t×t matrix, and (M) ij Assigned to negative infinity, then a ij = 0, excluding the inclusion of feature v i ′ of v j Therefore, by assigning negative infinity to the elements of M, we can cut off the directed edges in the row vectors of V and construct the causal relationships of V.
[0132] When Q, K, and V represent the same sequence, the above formula Attention(Q, K, V) is called self-attention. In the self-attention module, these three matrices have different representations in the same sequence. Let X∈R t×d ′ is the input sequence, then the features in the self-attention are Q=XW Q , K=XW K , V=XW V Given, where W∈R d '×d . Since attention is paid to a t×d ′ to R t×d The mapping representation contains three weight parameter matrices. The self-attention of input X is represented by the masking matrix M×Z.
[0133] Where W=(W Q ,W K ,W V ), ⊙ is the element-wise product operator. In this paper, W is a tuple of trainable weight matrices, and M is a known masking matrix.
[0134] In one possible implementation, feature extraction and prior knowledge fusion are performed on the multi-layer spatiotemporal model based on an interpretable spatiotemporal attention transformer to obtain the spatiotemporal characteristics of power load, including:
[0135] Based on the spatial causal attention network, the multi-layer spatiotemporal model is converted into multiple spatial feature matrices;
[0136] Based on the temporal attention network, each spatial feature matrix is compressed to obtain multiple spatial feature vectors, and each spatial feature vector is combined into a spatiotemporal feature matrix according to the temporal causal relationship;
[0137] Based on the spatial dependency comparison module, spatial causal prior knowledge is added to the spatiotemporal feature matrix to obtain the spatiotemporal characteristics of power load.
[0138] In this embodiment, the multi-layer spatiotemporal model is input into the interpretable spatiotemporal attention transformer to extract the spatiotemporal characteristics of the power load. This allows the data to be adapted to the causal network, incorporates prior knowledge into feature learning, and improves the applicability of the deep learning model. Referring to Figure 2, the structure and function of each component of the interpretable spatiotemporal attention transformer are as follows:
[0139] (1) Spatial Causal Attention Network
[0140] The spatial causal attention network is a self-attention network that embeds a set of observation variables according to known spatial causal relationships. The embedded features represent the aggregated information of all sites in a fixed time, and their spatial causal relationships are represented by the spatial masking matrix M. S reflect.
[0141] All variables are initially embedded On. Set is x i,t In X t The embedding function in X t Represents the aggregated information of all sites within a fixed time. is u j,t in U t The embedded function in the following formula is: Represent the embedding functions of different matrices respectively. Let the spatial embedding vector be:
[0142] For 1≤i≤p, is a feature containing temporal information, which plays the role of dynamic and trainable position encoding in the transformer. is a p×d1 matrix, is the i-th row vector, which is the set of spatial embedding vectors for all positions. Then, the spatial causal attention network is defined as:
[0143] It is Self-attention, which is from arrive The mapping has an attention weight matrix of d1×d2 The triplet of M S represents the spatial masking matrix, As a spatial feature matrix, it corresponds to X weighted by the causal constraint weights in Assumption 2 t .
[0144] (2) Time Attention Network
[0145] The temporal attention network is a specially designed self-attention network that transforms a series of B past spatial features obtained in a spatial causal attention network. The temporal attention network is consistent with the similar concept of designing a transformer decoder, which satisfies the predictability of forward feed. The difference of the temporal attention network is that it uses reduced input to efficiently calculate the attention weights. The temporal attention network converts the following equation into Merge as input value and use the following formula As queries and keys, the dimensionality of the queries and keys is reduced through the variable selection network (VSN).
[0146] Where vec(·) represents the flattening mapping,
[0147] When tB<t′≤t, each row yes Corresponding to The VSN layer compresses the region information matrix by converting it into a single vector. The subscript of VSN1(·) is used to distinguish it from other functions of the same form. The temporal attention network is defined as:
[0148] in, MT is the decoder mask. M T The upper diagonal elements in are -∞, so that (A (1) ) ij = 0, i < j, and the i-th row of the temporal attention network consists only of When j≤i, the temporal attention network retains the irreversibility of temporal features. represents the attention weight matrix of Q, represents the attention weight matrix of K.
[0149] (3) Spatial dependence comparison module
[0150] Temporal Attention Network uses the spatial collapse feature of VSN, which leads to the spatial causal relationship caused by Hypothesis 2 in Add the comparison step of the construction space related features, for set up for The reconstruction matrix, The (t′-t+B) row vector of . The contrasting features of gives:
[0151] and Both are obtained through the spatial causal attention network, which fuses temporal information and spatial information together to make representation learning richer. The attention weights are used to explain and quantify the spatial effect.
[0152] Scan comparison After that, the final output of the encoder of the interpretable spatiotemporal attention transformer is given by Establish. is input to VSN2(·), which transmits the summary information to the decoder. In the formula, let When tB<t′≤t, The connection matrix with the context vector, where:
[0153] Evaluate variable returns by using the variable selection weights of VSN2(·) The importance of is the predicted value of the spatiotemporal feature.
[0154] (4) Decoder
[0155] A novel decoder architecture using global and local context vectors in a feed-forward network (FFN) layer is introduced. The decoder contains two VSN layers: a global VSN layer and a local VSN layer. Construct a global context vector:
[0156] Local VSN creates a local context vector by encapsulating the local embedded temporal features in the following formula.
[0157] in, and are the trainable weight vector and bias vector respectively. The two VSNs define a pooled context vector after time t:
[0158] in
[0159] Next, the encoder output and the pooled context vector are concatenated as Aggregate the forward and backward available features at time t.
[0160] In order to enrich the temporal features, the self-attention network is transformed into
[0161] in
[0162] is the last feature of the decoder. The model can evaluate temporal importance through the attention weights of the last self-attention layer. Using the attention weights, we can calculate the past time points that the model pays attention to and diagnose their consistency.
[0163] The quantile output layer is an FFN layer that returns k-step-ahead predictions at the quantile level q. as follows:
[0164] in is a trainable parameter, express The (B+k)th row of , k=1,...,τ.
[0165] Note that the decoder directly predicts Instead of recursion, a decoder designed for direct prediction can improve performance by avoiding error accumulation that leads to biased predictions due to its simple and effective model structure.
[0166] In the transformer-based prediction method, the encoder part is usually used for feature extraction, and the decoder is used to generate prediction results. In this embodiment, only the complete interpretable spatiotemporal attention transformer is trained, and then the spatiotemporal features of the power load are extracted using the above structures (1)-(3). Subsequent prediction is performed based on the spatiotemporal features of the power load output by the spatial dependency comparison module.
[0167] (5) Loss function
[0168] The composite quantile loss (CQL) is introduced to predict multiple quantiles. First, the quantile loss (QL) is defined as:
[0169] Among them I A (a) Returns 1 for a∈A, otherwise returns 0. The interpretable spatiotemporal attention transformer is trained by minimizing CQL, which is defined as follows:
[0170] Where W represents the entire weight and bias parameters, and T is the set of time points in the training dataset.
[0171] In one possible implementation, the STEF-DHNet model includes a first convolutional layer, a second convolutional layer, a flattening module, L fully connected layers, and an LSTM layer connected in sequence; where L is the number of input data.
[0172] In this example, data processed by the Interpretable Spatiotemporal Attention Transformer is fed into the STEF-DHNet model, further improving the accuracy of spatiotemporal forecasting. This model is designed to capture the spatial and temporal dependencies of electricity load data and effectively incorporate external factors into their true complexity. By accounting for the actual spatiotemporal complexity of external factors, the model provides a more accurate representation of real-world scenarios.
[0173] The STEF-DHNet model combines the strengths of CNN and LSTM layers for forecasting, while also considering the spatiotemporal nature of external factors by integrating them into the network architecture. The model uses a time lag of L = 4 hours to predict the electricity load for the next hour.
[0174] In one possible implementation, based on the spatiotemporal characteristics of power load, the real-time power load data of the target power node, and real-time environmental data, a trained STEF-DHNet model is used for prediction. The predicted power load data of the target power node at a specified future time includes:
[0175] Based on the first convolution layer and the second convolution layer, feature extraction is performed on the real-time power load data of the target power node to obtain power load features;
[0176] The real-time environmental data of the target power consumption node is gridded and combined with the power load characteristics to input into the flattening module to obtain flattened data;
[0177] Extract predictive features from the flattened data based on L fully connected layers;
[0178] Based on the LSTM layer, the prediction features and spatiotemporal features of power load are predicted to obtain the predicted power load data of the target power node at a specified future time.
[0179] In this embodiment, as shown in FIG3 , the specific process of prediction based on the STEF-DHNet model is as follows:
[0180] The first two layers of the network are convolutional layers. i (.|Θ i ) represents the operations performed by the i-th layer of the network (including batch normalization), and the parameter vector Θ i Expressed as: D t-1 =f2(f1(E t-1 |Θ1)|Θ2),
[0181] For each layer, K=32 kernels are used with a kernel size of 3×3. The output D of these two layers is t-1 is a tensor of size L×W×H×K. To include external factors, D is transformed along the last axis. t-1 With the external factor tensor F t-1 Connected together, a unified representation of the grid and time information and external factors is created as: C t-1 =CONCAT(D t-1 ,F t-1 ),
[0182] The obtained tensor C t-1 The size of is L×W×H×(K+M), where M is the number of external factors. Note that M external factors are used for each region. In the next step, only C t-1 The spatial dimensions of are flattened into a new shape of L×W×H(K+M), denoted as B t-1 . B t-1 =FLATTEN(C t-1 ),
[0183] Before passing the data to the LSTM layer, L fully connected layers (FC-D) are used, each with an output size of d, to generate dense outputs for each time lag. The resulting shape of size L×d is denoted as A t-1, the operation can be written as:
[0184] in is the parameter vector of the fully connected layer. The LSTM layer follows the fully connected layer and aims to capture the temporal patterns in the data. t =LSTM(A t-1 |Θ L ),
[0185] where Θ L is the parameter vector of the LSTM layer.
[0186] The prediction features and the spatiotemporal features of the power load are combined and input into the output of the LSTM layer using g t Represented, it is fed into another FC-D of size N to produce the output prediction result.
[0187] in represents the parameter vector of the second fully connected layer. The final fully connected layer is then reshaped into a tensor of dimension W×H, effectively generating the predicted electricity load Taking into account spatiotemporal causality and external factors at the same time.
[0188] The introduced STEF-DHNet model has the following parameter vector characteristics.
[0189] In one possible implementation, before using the trained STEF-DHNet model for prediction, the following is also included:
[0190] Construct training sets, validation sets, and test sets based on the historical power load data and historical environmental data of the target power consumption nodes;
[0191] Using the mean absolute error as the loss function, the STEF-DHNet model is adversarially trained based on the training set, validation set, and test set to obtain the trained STEF-DHNet model.
[0192] In this embodiment, to train STEF-DHNet, the mean absolute error (MAE) loss function is defined as:
[0193] Once the loss is calculated, the model is trained using backpropagation. During backpropagation, the gradient of the loss function with respect to the weights is calculated. The gradient is then used to update the weights using the ADAM (Adaptive Moment Estimation) optimizer. By iteratively minimizing the loss function, the model gradually learns to make more accurate predictions.
[0194] The mean absolute error (MAE), the root mean squared error (RMSE), and the mean absolute percentage error (MAPE) are used as evaluation metrics for performance comparison. The metrics are as follows:
[0195] The above metrics are evaluated on the three datasets at one-hour intervals.
[0196] In one possible implementation, the STEF-DHNet model is adversarially trained based on the training set, validation set, and test set. The trained STEF-DHNet model includes:
[0197] Generate adversarial samples based on the training set and the policy network; the policy network consists of a sequentially connected spatiotemporal encoder, a spatial layer, a temporal layer, and a multi-head attention decoder;
[0198] The STEF-DHNet model is adversarially trained based on the reward function, training set, and adversarial samples, and verified based on the validation set and test set to obtain the trained STEF-DHNet model.
[0199] In this embodiment, preprocessed data is input into adversarial training with the goal of enhancing the robustness of the model by augmenting the training data with adversarial examples, so that the model is robust. This involves modeling the node selection problem as a combinatorial optimization problem and using a reinforcement learning-based approach to learn the optimal node selection strategy. In addition, self-knowledge distillation is used as a new training technique to address the challenge of evolving adversarial nodes, thereby avoiding the "forgetting problem." As shown in Figure 4, the specific process of adversarial training is as follows:
[0200] 1. Adversarial training and formula
[0201] Adversarial training involves using adversarial examples generated by adversarial attacks during training to improve the robustness of the model. Adversarial training can be formulated as a min-max optimization problem.
[0202] Where θ represents the model parameters, x' is the adversarial sample, B(x,ε)={x+δ|||δ|| p ≤ε} denotes the set of adversarial examples with the maximum perturbation budget ε, where δ denotes the set of adversarial examples. θ (·) denotes the deep learning model, and y denotes the ground truth.
[0203] This application studies the application of traditional adversarial training methods in electricity load forecasting and introduces an adversarial training formula.
[0204] The following is an example of spatiotemporal adversarial training. It is based on the insight that the key to improving model robustness is to actively identify and focus on the most extreme cases of adversarial perturbations. Specifically, as follows, the worst-case scenario in the electricity load forecasting model involves both spatial and temporal aspects. From the temporal aspect, the attacker can inject adversarial perturbations in the feature space. In order to effectively defend against various types of attacks, it is crucial to deeply explore the worst-case scenario in the adversarial perturbation space, which is similar to the approach taken in the field of image recognition. From a spatial perspective, a dynamic node selection method is designed in each training epoch to maximize the internal loss and ensure that all nodes have a fair chance of being selected. To achieve this, a subset of nodes that exhibit spatiotemporal dependencies is dynamically selected from the full set of nodes in each training iteration.
[0205] First, define the perturbation space that allows hostility as follows: ψ(x' t )={x t +Δ t I t |||I t ||0≤η,||Δ t || p ≤ε},
[0206] Where x′ t is an example of spatiotemporal adversarial perturbation, Δt is the spatiotemporal adversarial perturbation, and the matrix I t ∈{0,1} n×n is the adversarial node indicator, which is a diagonal matrix with the jth diagonal element representing the node v j Whether it is selected as an adversarial node at time t. Specifically, if node v j If a node is selected as an adversarial node, the diagonal element jth of the matrix is equal to 1, otherwise it is 0. The parameter η is the budget for the number of nodes, and ε is the budget for the adversarial perturbation.
[0207] The adversarial training method for spatiotemporal prediction is described as follows:
[0208] x′ t-τLt ={x′ t-τ ,…,x′ t} is the hostile state from time period t-τ to t. train Represents the time step set of all training samples. L AT(·) represents the user-specified adversarial training loss function, which can include commonly used metrics such as Mean Squared Error (MSE) or others. The goal of the inner maximization is to find the optimal adversarial perturbation that maximizes the loss. In the outer minimization, the model parameters are updated to minimize the prediction loss.
[0209] 2. Strengthen optimal node subset learning
[0210] The problem of selecting the optimal node subset from a set of n spatiotemporal distributed data sources is formulated as a combinatorial optimization problem. The problem instance is denoted as s, which consists of n nodes represented by spatiotemporal features: t-τ:t From time slot t-τ to t. The goal is to select η nodes from the complete set of n nodes, consisting of the node subset Ω=(ω1,…,ω η ) represents, where ω k ∈{v1,…,v n},ω k ≠ω k′
[0211] Given a problem instance s, the goal is to learn a random policy by decomposing the probability of the solution using the chain rule Parameters The policy network uses this information to determine the optimal subset of nodes to select in order to explore the most extreme cases of adversarial perturbations at each training iteration.
[0212] The policy network consists of an encoder and a decoder. The encoder generates geographically distributed data embeddings, and the decoder generates Ω sequences.
[0213] 2.1 Strategic Network Design
[0214] The policy network is based on the spatiotemporal feature function x t-τ:t The solution Ω is obtained by taking Ω as input. It consists of a spatiotemporal encoder and a multi-head attention decoder. The encoder converts spatiotemporal features into embeddings, and the decoder constructs the solution in an autoregressive manner, selecting a node at a time and using the previous selection to select the next node until a complete solution is generated.
[0215] 1) Spatiotemporal Encoder. A spatiotemporal encoder similar to GraphWaveNet is used to convert the spatiotemporal data of electricity load into embeddings. The spatiotemporal encoder receives spatiotemporal data as input and produces node embeddings as output. The spatiotemporal encoder typically consists of multiple spatiotemporal layers and temporal layers.
[0216] 2) Spatial layer. Adaptive graph convolution is adopted as the spatial layer to capture spatial dependencies. The information aggregation method is based on the diffusion model, allowing traffic signals to diffuse for L steps. By aggregating the hidden states of adjacent nodes, the hidden layer embedding is updated through adaptive graph convolution.
[0217] Where Z′ l is the output of the l-th layer's implicit embedding, and W i is the model parameter at depth i, and A ada is the learnable adjacency matrix.
[0218] 3) Temporal layer. The model uses a gated temporal layer to process sequential data. It is defined as follows.
[0219] Where σ is the sigmoid function, and are the model parameters, ★ is the unfolded convolution operation, and ⊙ is the element-wise multiplication. E l is the input of the l-th block and also the output of the l - 1-th block. The following formula is used to add a residual link for each block. E l+1 =Z l +E l ,
[0220] The hidden states of different layers are concatenated and passed into two Multilayer Perceptrons (MLPs) to obtain the final node embeddings. F = MLP(∥ l=1 Z l ),
[0221] Where F is the set of node embeddings, and the average value of all node embeddings is represented as the graph embedding, which can be expressed as F i is the embedding of node v i .
[0222] 4) Multi-head attention decoder. The decoder generates a node sequence Ω by iteratively selecting a single node ω k at each step k, while using the encoder's embedding and the output ω k′ from the previous steps (for k' < k) as inputs.
[0223] Specifically, the decoder's input includes the graph embedding and the embedding of the last node, where the embedding of the first selected node is a learned embedding. The decoder calculates the probability that each node is selected as the adversarial node, while considering computational efficiency. During the decoding process, the context is represented by a special context node. For this purpose, a attention layer is calculated on top of the decoder in combination with an attention-based decoder, and the message is only sent to the context node. The context node embedding is defined as follows:
[0224] in, is the graph embedding and v is the learned embedding at the first iteration step. Embedding of the last selected node for the k-1 iterations.
[0225] In order to update the context node embedding of the message information, the multi-head attention method is used to calculate the new context node embedding: U' (c) =∥ j=1 MHA j (q (c) ,k j ,v j ),
[0226] in is self-focused, and
[0227] To calculate the probability of the next node, the key and value are taken from the initial node embedding. q = W Q U' (c) ,k i =W K H′ i ,
[0228] First, we calculate the pairwise number of a single attention head and query all nodes using the new context node.
[0229] Among them, C is a constant, and the selected node is r j =-∞ for shielding.
[0230] Then, the final probability of the node is calculated using the chain rule and the softmax function, and the probability of each node is calculated according to the softmax function.
[0231] where p i is the probability of node i, ω k is the current node. The node with the highest probability is selected from all nodes as the next sampling node.
[0232] 2.2 Balanced Reward Function Design
[0233] The main challenge in learning a policy network is evaluating the solution Ω generated by the policy network. One approach is to use an internal loss (calculated using the solution Ω) as a reward, with larger values indicating better solutions. However, as training progresses and the model becomes more robust, the internal loss is expected to decrease, which can lead to incorrect feedback and suboptimal solutions. To address this issue, a reward function balancing strategy is introduced. Instead of using the internal loss alone, the results generated by the policy network are compared with those generated by a baseline node selector, and the difference is used as a reward. This approach provides stable and effective feedback to the policy network, helping to alleviate the problem of reducing the internal loss during training.
[0234] Specifically, first solve Ω=(ω1,…,ω k ) Get the adversarial node index I t A set of ω k ∈{v1,...,v n}, use the following function:
[0235] Where I (i,i),t Indicates I t The i-th diagonal element at time step t.
[0236] In order to improve computational efficiency, we do not use gradient-based methods to calculate adversarial samples. Instead, we directly extract a random variable Δ from a probability distribution π(Δ) to calculate the adversarial sample for adversarial training. t =x t +Δ·I t ,
[0237] In the implementation, a uniform distribution in the range of (-ε,ε) is selected as the perturbation source Δ.
[0238] To evaluate the performance of the prediction model when using nodes in the solution as adversarial nodes, the cost function is calculated as follows:
[0239] is the MSE loss,
[0240] In order to ensure that the policy network receives stable and effective feedback, a balancing strategy is implemented for the reward. Specifically, a baseline node selector (such as Random selector, randomly select nodes, etc.) is used to select nodes as the solution Ω b The results generated by the policy network are then compared with the baseline results, and the difference is used as the reward. It is expressed as follows: r(Ω (p) )=L(Ω (p) )-L(Ω (b) ),
[0241] where Ω (p) is the solution generated by the policy network, Ω (b) is the solution generated by the baseline selector, and is aligned with the policy network selector and the baseline selector using superscripts (p) and (b), respectively. In this way, using the balanced reward function r(Ω (p) ) as a reward signal to guide the policy network to update the solution Ω. In practice, a heuristic method is adopted as a baseline selector to select nodes named TNDS.
[0242] 3.2.3 Strategy Network Training
[0243] The training of the policy network is completed by alternately training the policy network and the spatiotemporal prediction model in an adversarial manner. Specifically, the policy network is trained according to the input variables. Generate a solution sequence, denoted as Ω. Then calculate the equilibrium reward and use it to update the policy network. Then, use the final node selection metric to calculate the adversarial example, denoted as Optimization is performed using the Projected Gradient Descent (PGD) method, as follows:
[0244] in, Operator is used to convert variables The maximum perturbation of is limited to an estimated value. The adversarial example of the i-th iteration is expressed as: is the step length, I t Select metrics for the final nodes obtained from the policy network, is the mean square error loss function.
[0245] Then, the spatiotemporal prediction model is trained on the adversarial samples, and the prediction model loss is optimized as follows:
[0246] To train the policy network, the loss function is defined as follows:
[0247] Where C is a constant. The policy network is optimized using gradient descent and reinforcement learning, using the Adam optimizer.
[0248] 3. Regularized Adversarial Training
[0249] Another challenge in spatiotemporal prediction adversarial training is instability, which occurs when adversarial nodes constantly change during training. This can cause the model to be unable to effectively remember all historical adversarial node instances, resulting in a lack of robustness to stronger attacks, commonly known as the "forgetting problem." To address this issue, knowledge distillation (KD) is used to transfer knowledge from the teacher model to the student model. Previous research has shown that KD can improve the adversarial robustness of the model.
[0250] However, the traditional teacher model is static and cannot provide dynamic knowledge. To overcome this limitation, a new self-knowledge distillation regularization for adversarial training is introduced. Specifically, the model of the previous epoch is used as the teacher model, which means that the current spatiotemporal prediction model is trained using the knowledge extracted from the previous model. In this way, the current model can learn from the experience of adversarial attacks from the previous model. The knowledge distillation loss is defined as follows:
[0251] in is the knowledge distillation loss (e.g., MSE), is the teacher model, and the model trained last time is used. In summary, the final adversarial training loss is defined as follows:
[0252] where α is a parameter that controls the amount of knowledge transferred from the teacher model. Note that in the first training epoch, The function is directly used as the adversarial training loss.
[0253] Adversarial training of a spatiotemporal prediction model. The training process is divided into two phases. In the first phase, the policy network is trained using an algorithm. In the second phase, adversarial nodes are selected using the pre-trained policy network to improve computational efficiency. PGD is then used to compute adversarial examples. Finally, the prediction model parameters are updated using the Adam optimizer.
[0254] In one possible implementation, verification based on the validation set and the test set includes:
[0255] During the training process, the rolling error of the STEF-DHNet model is updated based on the test set and the given time length. If the rolling error is qualified, the training is completed, otherwise the training continues.
[0256] In this example, the STEF-DHNet model is evaluated based on rolling error, demonstrating the model’s effectiveness in long-term predictions. This metric takes the model’s previous outputs as input to generate subsequent outputs, making it an effective method for evaluating model accuracy over longer periods of time.
[0257] The available data samples were split into three non-overlapping parts: training, validation, and test sets. The training set was used to train and validate the models, the validation set was used to test the fit of each model on unknown test data, and the test set was used to roll the model over a given time. For each data split, the corresponding training, testing, and rolling errors (MAE, RMSE, and MAPE) were obtained for all methods and all datasets.
[0258] Unlike traditional performance metrics that rely on the model's predictions at a specific time, rolling error accounts for the accumulation of errors across multiple predictions over a longer period of time. The term "rolling" refers to the use of a rolling window approach to calculate the metric, where the prediction at time t is used as input to generate the next input for the model, which then generates the subsequent prediction, and so on. Rolling error metrics provide valuable insights into the accuracy of a model over an extended period of time and help select the best model that minimizes error over time without the need for frequent retraining.
[0259] In one possible implementation, constructing a training set, a validation set, and a test set based on historical power load data and historical environmental data of a target power node includes:
[0260] Clean the historical power load data and historical environmental data of the target power consumption node to obtain clean data;
[0261] Smoothing the clean data to obtain smoothed data;
[0262] Add time information to the smoothed data, and use the historical power load data and historical environmental data at the same time as a sample data;
[0263] Extract features from each piece of sample data, add the extracted features to the corresponding sample data, obtain multiple pieces of feature sample data and form a data set;
[0264] The dataset is divided into training set, validation set and test set according to the preset ratio.
[0265] In this embodiment, the collected power load data may contain noise or missing data. To reduce the impact of these interference factors on the model, the data can be cleaned and smoothed. Furthermore, this model performs spatiotemporal power load forecasting, requiring the filtering of data and modification of data types. The data can also be timestamped, feature extracted, and split.
[0266] (1) Data cleaning
[0267] Clean the collected power load and environmental data to remove outliers and noise. Statistical methods, interpolation methods, or machine learning methods can be used to fill in missing values and repair outliers.
[0268] (2) Data smoothing
[0269] Data smoothing: Smoothing of power load data with large fluctuations to reduce the interference of noise on the model.
[0270] Data smoothing is a statistical technique used to reduce noise and volatility in data. It processes raw data within a specific time window or spatial range to make it smoother and more continuous. Common data smoothing methods include moving average, weighted moving average, and exponential smoothing. These methods combine a certain number of adjacent data points into a single value to reduce noise and volatility. This data smoothing process produces a more stable and continuous sequence of data points, reducing the impact of noise and volatility on analysis and forecasting results. It is important to note that when applying data smoothing, there is a trade-off between the degree of smoothing and the risk of information loss. Excessive smoothing may result in loss of data detail, while insufficient smoothing may not effectively reduce noise and volatility. Therefore, when applying data smoothing methods, it is important to select appropriate parameters and smoothing degree based on the specific situation.
[0271] The moving average method is a commonly used data smoothing technique. It smoothes data by calculating the average of data points within a specific time window. The moving average method effectively reduces short-term noise and volatility while preserving long-term trends. Longer time windows better capture trends but may be less responsive to rapidly changing signals. Shorter time windows, on the other hand, are more sensitive to rapidly changing signals but may increase noise and volatility.
[0272] The steps of the moving average method are as follows:
[0273] ① Determine the length of the time window, that is, how many adjacent data points to consider. Longer time windows can reduce noise and volatility, but may result in larger delays.
[0274] ② Add up the data points in the time window and divide by the length of the time window to get the average value.
[0275] ③ Use the calculated average value as the new data point to replace all data points in the corresponding time window in the original data.
[0276] ④ Slide the time window, move it forward one unit (for example, slide it forward one time interval), and repeat steps ② and ③.
[0277] ⑤ Repeat the above steps until all data points are processed.
[0278] In addition to the simple moving average, there are other moving average variations, such as weighted moving average and exponentially smoothed moving average. These employ different weights or decay factors when calculating the average to better adapt to different data characteristics. Choosing the appropriate time window length and moving average method depends on the specific problem and data characteristics to reduce noise and volatility.
[0279] (3) Timestamp processing: If the collected data does not have timestamp information, it is necessary to add a timestamp to each data point. Timestamps can be generated based on the collection frequency and start time. Timestamps are converted into a time series format that can be analyzed and predicted, such as date-time or time interval.
[0280] (4) Feature extraction: Extract useful features from the raw data. In addition to electricity load data, environmental data related to electricity consumption can also be used to extract features. For example, daily, weekly, and monthly average temperatures can be extracted as features. Useful feature variables are extracted as needed and added to the dataset.
[0281] (5) Data splitting: The data set is divided into a training set, a validation set, and a test set according to a certain ratio (the test set is used to calculate the rolling error).
[0282] Example 2:
[0283] In Example 2, as shown in FIG5 , the steps for performing prediction using the power load prediction method based on spatiotemporal correlation provided by this application are as follows:
[0284] (1) First, the collected power load data and external environment data are preprocessed, including data cleaning, data splitting, feature extraction, timestamp processing and data smoothing;
[0285] (2) The STEF-DHNet model is trained adversarially based on the preprocessed data to improve the adversarial robustness of the prediction results;
[0286] (3) Perform spatiotemporal modeling based on the preprocessed data to obtain a multi-layer spatiotemporal model;
[0287] (4) Extracting spatiotemporal characteristics of power load from a multi-layer spatiotemporal model based on an interpretable spatiotemporal attention transformer;
[0288] (5) The spatiotemporal characteristics of the electricity load are input into the STEF-DHNet model trained in step (2) to perform electricity load forecasting, obtain a forecast result with improved accuracy, and evaluate the performance of the STEF-DHNet model through the rolling error.
[0289] As can be seen from the above, this application aims to improve the performance of spatiotemporal prediction of power load and enhance the robustness of spatiotemporal prediction. The key technologies of this application are as follows:
[0290] 1. We introduce a grid-based deep learning model, called STEF-DHNet, for spatiotemporal forecasting of electricity load. This model combines CNN and LSTM layers. CNN layers capture spatial dependencies and effectively model interactions between regions. Long short-term memory (LSTM) layers are deployed to capture temporal dependencies, including nonlinear relationships between current forecasts and past observations. Leveraging the strengths of CNN and LSTM layers, this model effectively incorporates the true complexity of external factors while maintaining computational efficiency.
[0291] 2. We introduce a performance metric called rolling error to evaluate the accuracy of our model in real-world applications. Unlike traditional performance metrics that rely on the model's predictions at a specific time, rolling error accounts for the accumulation of errors across multiple predictions over a longer period of time, taking the model's previous outputs as input to generate subsequent outputs, making it an effective method for evaluating model accuracy over longer periods of time. The rolling error metric provides valuable insights into the model's accuracy over a longer period of time and helps select the best model that minimizes error over time without the need for frequent retraining. Our model also outperforms state-of-the-art methods on this metric, demonstrating its ability to generate accurate predictions without the need for ongoing retraining.
[0292] 3. A new framework for improving the adversarial robustness of spatiotemporal predictions is introduced. This involves dynamically selecting a subset of nodes as adversarial examples, which not only reduces overfitting but also improves defense against dynamic adversarial attacks. To solve the task of selecting a subset from the total set of nodes, a reinforcement learning-based approach is introduced to learn the optimal node selection strategy. The node selection problem is modeled as a combinatorial optimization problem, and a policy-based network is used to learn a node selection strategy that maximizes the internal loss. A spatiotemporal attention-based policy network is designed to simulate spatiotemporal distributed data. In order to evaluate the results produced by the policy network, a balanced reward function strategy is introduced to provide stable and effective feedback to the policy network and alleviate the problem of internal loss reduction during training. To overcome the forgetting problem, a new self-knowledge distillation regularization module is introduced for adversarial training, in which the current model is trained using knowledge extracted from the adversarial attack experience of the previous model. In addition, self-knowledge distillation is used as a new training technique to address the challenge of evolving adversarial nodes, thereby avoiding the "forgetting problem".
[0293] 4. We introduce a multi-quantile prediction neural network model with spatiotemporal causal structure. This model extends the existing transformer by alternating spatiotemporal masks, thereby incorporating prior knowledge into the model's feature learning. In this proposed model, spatial and temporal causal relationships can be easily determined through the weights of the attention layer, and the importance of each variable can be identified through the VSN layer included in the model.
[0294] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0295] The following are device embodiments of the present application. For details not fully described therein, please refer to the corresponding method embodiments described above.
[0296] FIG6 shows a schematic diagram of the structure of a power load prediction device based on spatiotemporal correlation according to an embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown, which are described in detail as follows:
[0297] As shown in FIG6 , the power load prediction device 6 based on spatiotemporal correlation includes:
[0298] The spatiotemporal modeling module 61 is configured to perform spatiotemporal modeling based on historical power load data and historical environmental data of different power consumption nodes to obtain a multi-layer spatiotemporal model; wherein each layer of the model is a directed graph representing a spatial causal relationship at a fixed time, and the directed edges between each layer of the model represent a temporal causal relationship;
[0299] A feature extraction module 62 is configured to extract features and fuse prior knowledge from the multi-layer spatiotemporal model based on an interpretable spatiotemporal attention converter to obtain spatiotemporal characteristics of power load; wherein the interpretable spatiotemporal attention converter includes a spatial causal attention network, a temporal attention network, and a spatial dependency comparison module;
[0300] The load forecasting module 63 is used to predict the target power node's predicted power load data at a specified future time using the trained STEF-DHNet model based on the spatiotemporal characteristics of the power load, the real-time power load data of the target power node, and the real-time environmental data.
[0301] In a possible implementation, the spatiotemporal modeling module 61 is specifically configured to:
[0302] The historical power load data and historical environmental data of different power consumption nodes are layered according to time to obtain a multi-layer network;
[0303] Based on the influence relationship between the historical power load data and historical environmental data of different power consumption nodes, similar data are divided into the same cluster to obtain multiple clusters;
[0304] Construct temporal causal hypothesis relationships and spatial causal hypothesis relationships based on the influence relationships between clusters;
[0305] Based on the self-attention module, the temporal causal hypothesis relationship and the spatial causal hypothesis relationship are screened to obtain the spatiotemporal causal relationship;
[0306] Adding spatiotemporal causal relationships to a multi-layer network results in a multi-layer spatiotemporal model.
[0307] In a possible implementation, the feature extraction module 62 is specifically configured to:
[0308] Based on the spatial causal attention network, the multi-layer spatiotemporal model is converted into multiple spatial feature matrices;
[0309] Based on the temporal attention network, each spatial feature matrix is compressed to obtain multiple spatial feature vectors, and each spatial feature vector is combined into a spatiotemporal feature matrix according to the temporal causal relationship;
[0310] Based on the spatial dependency comparison module, spatial causal prior knowledge is added to the spatiotemporal feature matrix to obtain the spatiotemporal characteristics of power load.
[0311] In one possible implementation, the STEF-DHNet model includes a first convolutional layer, a second convolutional layer, a flattening module, L fully connected layers, and an LSTM layer connected in sequence; where L is the number of input data.
[0312] In a possible implementation, the load forecasting module 63 is specifically configured to:
[0313] Based on the first convolution layer and the second convolution layer, feature extraction is performed on the real-time power load data of the target power node to obtain power load features;
[0314] The real-time environmental data of the target power consumption node is gridded and combined with the power load characteristics to input into the flattening module to obtain flattened data;
[0315] Extract predictive features from the flattened data based on L fully connected layers;
[0316] Based on the LSTM layer, the prediction features and spatiotemporal features of power load are predicted to obtain the predicted power load data of the target power node at a specified future time.
[0317] In a possible implementation, the load forecasting module 63 is further configured to:
[0318] Before using the trained STEF-DHNet model for prediction, a training set, validation set, and test set are constructed based on the historical power load data and historical environmental data of the target power consumption node;
[0319] Using the mean absolute error as the loss function, the STEF-DHNet model is adversarially trained based on the training set, validation set, and test set to obtain the trained STEF-DHNet model.
[0320] In a possible implementation, the load forecasting module 63 is specifically configured to:
[0321] Generate adversarial samples based on the training set and the policy network; the policy network consists of a sequentially connected spatiotemporal encoder, a spatial layer, a temporal layer, and a multi-head attention decoder;
[0322] The STEF-DHNet model is adversarially trained based on the reward function, training set, and adversarial samples, and verified based on the validation set and test set to obtain the trained STEF-DHNet model.
[0323] In a possible implementation, the load forecasting module 63 is specifically configured to:
[0324] During the training process, the rolling error of the STEF-DHNet model is updated based on the test set and the given time length. If the rolling error is qualified, the training is completed, otherwise the training continues.
[0325] In a possible implementation, the load forecasting module 63 is specifically configured to:
[0326] Clean the historical power load data and historical environmental data of the target power consumption node to obtain clean data;
[0327] Smoothing the clean data to obtain smoothed data;
[0328] Add time information to the smoothed data, and use the historical power load data and historical environmental data at the same time as a sample data;
[0329] Extract features from each piece of sample data, add the extracted features to the corresponding sample data, obtain multiple pieces of feature sample data and form a data set;
[0330] The dataset is divided into training set, validation set and test set according to the preset ratio.
[0331] The embodiment of the present application first utilizes the hierarchical structure of power load data and environmental data, and constructs the historical power load data and historical environmental data of different power nodes into a multi-layer spatiotemporal model according to spatial causality and temporal causality, and converts the spatiotemporal dependencies of power load data and environmental data into spatiotemporal causality; then, based on an interpretable spatiotemporal attention converter, feature extraction is performed and prior knowledge is integrated, and the complete spatiotemporal characteristics of power load can be obtained without the need to fuse spatial causality and temporal causality; finally, the spatiotemporal causality contained in the spatiotemporal characteristics of power load is integrated into the STEF-DHNet model to perform power load forecasting, thereby improving the forecast accuracy and model applicability.
[0332] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0333] Those skilled in the art will appreciate that the templates, units, and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0334] If the module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned embodiments of the power load prediction method based on spatiotemporal correlation. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium, etc.
[0335] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for predicting power load based on spatiotemporal correlation, characterized in that: include: Based on the historical power load data and historical environmental data of different power consumption nodes, spatiotemporal modeling is performed to obtain a multi-layer spatiotemporal model; wherein each layer of the model is a directed graph, representing the spatial causal relationship of a fixed time, and the directed edges between each layer of the model represent the temporal causal relationship; the multi-layer spatiotemporal model is: Among them, V t =[v 1,t , ..., v 4,t }; The causal structure of the multi-layer spatiotemporal model is defined based on assumptions 1 and 2. Assumption 1 is: for s, s'∈S, t, t'∈T, the multi-layer network G = (v, ε, τ) satisfies the following conditions: ①Assume t≤t', then ②For s≠s′ and (v s,t ,v s′,t )∈ε, if and only if (v s,t′ ,v s′,t′ )∈ε; In addition, If and only if In the above hypothesis 1, ① represents irreversibility, and ② represents uniform spatial causality over time; The hypothesis 2 is: s,1 , v s′,1 )∈ε, if (s, s′)∈{(1, 3), (1, 4), (2, 4), (4, 3)}, otherwise, The hypothesis 2 represents the spatial causal relationship based on prior knowledge; The power consumption nodes include B1, B2, B3, B4, and B5. B5 is the upstream node of B1, B2, B3, and B4, that is, the power load of B1, B2, B3, and B4 comes from B5. The environmental data includes three different temperatures P1, P1, and P3, as well as the power consumption control strategy D. C1 = {P1, P1, P3}, C2 = {B5}, C3 = {D}, and C4 = {B1, B2, B3, B4}. The node v s,t Represents cluster C s All available observations in , s∈S={1,2,3,4}, and time t∈T={1,...,τ}; Based on the interpretable spatiotemporal attention converter, feature extraction and prior knowledge fusion are performed on the multi-layer spatiotemporal model to obtain the spatiotemporal characteristics of power load; wherein the interpretable spatiotemporal attention converter includes a spatial causal attention network, a temporal attention network and a spatial dependency comparison module; The spatial causal attention network is: in, yes The self-attention is from arrive The mapping has an attention weight matrix of d1×d2 The triplet of As a spatial feature matrix, it corresponds to X weighted by the causal constraint weights in Assumption 2 t ; The temporal attention network is: in, M T is the decoder mask; M T The upper diagonal elements in are -∞, so that (A (1) ) ij = 0, i < j, and the i-th row of the temporal attention network consists only of The j-th row vector of The temporal attention network converts the following Merge into input values and use the following formula As query and key, the dimension of query and key is reduced through variable selection network; Where vec(·) represents the flattening mapping, When tB<t′≤t, each row yes Corresponds to The reduced vector of ; The spatial dependency comparison module is: in, The contrasting features of Given, for The reconstruction matrix, The (t′-t+B) row vector of and Both are obtained through the spatial causal attention network, The attention weights are used to explain and quantify the spatial effect; Scan comparison After that, the final output of the encoder of the interpretable spatiotemporal attention transformer is given by to establish; to Input to VSN2(·), VSN2(·) transmits the summary information to the decoder; When tB<t′≤t, In the formula The connection matrix with the context vector, where: Evaluate variable returns by using the variable selection weights of VSN2(·) The importance of Based on the spatiotemporal characteristics of the power load, the real-time power load data of the target power node and the real-time environmental data, the trained STEF-DHNet model is used for prediction to obtain the predicted power load data of the target power node at a specified future time.
2. The method for predicting power load based on spatiotemporal correlation according to claim 1, characterized in that: Based on the historical power load data and historical environmental data of different power consumption nodes, spatiotemporal modeling is carried out to obtain a multi-layer spatiotemporal model including: The historical power load data and historical environmental data of different power consumption nodes are layered according to time to obtain a multi-layer network; Based on the influence relationship between the historical power load data and historical environmental data of different power consumption nodes, similar data are divided into the same cluster to obtain multiple clusters; Construct temporal causal hypothesis relationships and spatial causal hypothesis relationships based on the influence relationships between clusters; The temporal causal hypothesis relationship and the spatial causal hypothesis relationship are screened based on a self-attention module to obtain a spatiotemporal causal relationship; The spatiotemporal causal relationship is added to the multi-layer network to obtain a multi-layer spatiotemporal model.
3. The method for predicting power load based on spatiotemporal correlation according to claim 2, characterized in that: The feature extraction and prior knowledge fusion of the multi-layer spatiotemporal model based on the interpretable spatiotemporal attention converter are performed to obtain the spatiotemporal features of the power load, including: Converting the multi-layer spatiotemporal model into a plurality of spatial feature matrices based on the spatial causal attention network; Compressing each spatial feature matrix based on the temporal attention network to obtain a plurality of spatial feature vectors, and combining each spatial feature vector into a spatiotemporal feature matrix according to the temporal causal relationship; Based on the spatial dependency comparison module, spatial causal prior knowledge is added to the spatiotemporal feature matrix to obtain the spatiotemporal features of the power load.
4. The method for predicting power load based on spatiotemporal correlation according to claim 1, characterized in that: The STEF-DHNet model includes a first convolutional layer, a second convolutional layer, a flattening module, L fully connected layers and an LSTM layer connected in sequence; wherein L is the number of input data.
5. The method for predicting power load based on spatiotemporal correlation according to claim 4, characterized in that: The predicted power load data of the target power node at a specified future time is obtained by using the trained STEF-DHNet model for prediction based on the spatiotemporal characteristics of the power load, the real-time power load data of the target power node and the real-time environmental data, including: Based on the first convolution layer and the second convolution layer, feature extraction is performed on the real-time power load data of the target power node to obtain power load features; Gridding the real-time environmental data of the target power consumption node, and combining it with the power consumption load characteristics and inputting it into the flattening module to obtain flattened data; Extracting prediction features from the flattened data based on the L fully connected layers; The prediction features and the spatiotemporal features of the power load are predicted based on the LSTM layer to obtain the predicted power load data of the target power node at a specified future time.
6. The method for predicting power load based on spatiotemporal correlation according to claim 4, characterized in that: Before using the trained STEF-DHNet model for prediction, it also includes: Constructing a training set, a validation set, and a test set based on the historical power load data and historical environmental data of the target power node; Taking the mean absolute error as the loss function, adversarial training is performed on the STEF-DHNet model based on the training set, the validation set and the test set to obtain a trained STEF-DHNet model.
7. The method for predicting power load based on spatiotemporal correlation according to claim 6, characterized in that: The adversarial training of the STEF-DHNet model based on the training set, the validation set and the test set to obtain the trained STEF-DHNet model includes: Generate adversarial samples based on the training set and the strategy network; wherein the strategy network includes a spatiotemporal encoder, a spatial layer, a temporal layer, and a multi-head attention decoder connected in sequence; The STEF-DHNet model is adversarially trained based on the reward function, the training set and the adversarial sample, and is verified based on the verification set and the test set to obtain a trained STEF-DHNet model.
8. The method for predicting power load based on spatiotemporal correlation according to claim 7, characterized in that: The verifying based on the verification set and the test set includes: During the training process, the rolling error of the STEF-DHNet model is updated based on the test set and the given time length. If the rolling error is qualified, the training is completed, otherwise the training continues.
9. The method for predicting power load based on spatiotemporal correlation according to claim 6, characterized in that: The constructing of a training set, a validation set and a test set based on the historical power load data and historical environment data of the target power node comprises: Cleaning the historical power load data and historical environment data of the target power consumption node to obtain clean data; Performing smoothing processing on the cleaned data to obtain smoothed data; Adding time information to the smoothed data, and taking the historical power load data and historical environment data at the same time as a piece of sample data; Extract features from each piece of sample data, add the extracted features to the corresponding sample data, obtain multiple pieces of feature sample data and form a data set; The data set is divided into a training set, a validation set and a test set according to a preset ratio.
10. A device for predicting power load based on spatiotemporal correlation, characterized in that: include: The spatiotemporal modeling module is used to perform spatiotemporal modeling based on the historical power load data and historical environmental data of different power consumption nodes to obtain a multi-layer spatiotemporal model; wherein each layer of the model is a directed graph, representing the spatial causal relationship of a fixed time, and the directed edges between each layer of the model represent the temporal causal relationship; the multi-layer spatiotemporal model is: Among them, V t = {v 1,t ,...,v 4,t }; The causal structure of the multi-layer spatiotemporal model is defined based on assumptions 1 and 2. Assumption 1 is: for s, s'∈S, t, t'∈T, the multi-layer network G = (v, ε, τ) satisfies the following conditions: ①Assume t≤t', then ②For s≠s′ and (v s,t ,v s′,t )∈ε, if and only if (v s,t′ ,v s′,t′ )∈ε; In addition, If and only if In the above hypothesis 1, ① represents irreversibility, and ② represents uniform spatial causality over time; The hypothesis 2 is: s,1 ,v s′,1 )∈ε, if (s,s′)∈{(1,3),(1,4),(2,4),(4,3)}, otherwise, The hypothesis 2 represents the spatial causal relationship based on prior knowledge; The power consumption nodes include B1, B2, B3, B4, and B5. B5 is the upstream node of B1, B2, B3, and B4, that is, the power load of B1, B2, B3, and B4 comes from B5. The environmental data includes three different temperatures P1, P1, and P3, as well as the power consumption control strategy D. C1 = {P1, P1, P3}, C2 = {B5}, C3 = {D}, and C4 = {B1, B2, B3, B4}. The node v s,t Represents cluster C s All available observations in , s∈S={1,2,3,4}, and time t∈T={1,...,τ}; A feature extraction module, used to extract features and fuse prior knowledge on the multi-layer spatiotemporal model based on an interpretable spatiotemporal attention converter to obtain spatiotemporal characteristics of power load; wherein the interpretable spatiotemporal attention converter includes a spatial causal attention network, a temporal attention network and a spatial dependency comparison module; The spatial causal attention network is: in, yes The self-attention is from arrive The mapping has an attention weight matrix of d1×d2 The triplet of As a spatial feature matrix, it corresponds to X weighted by the causal constraint weights in Assumption 2 t ; The temporal attention network is: in, M T is the decoder mask; M T The upper diagonal elements in are -∞, so that (A (1) ) ij = 0, i < j, and the i-th row of the temporal attention network consists only of The j-th row vector of The temporal attention network converts the following Merge into input values and use the following formula As query and key, the dimension of query and key is reduced through variable selection network; Where vec(·) represents the flattening mapping, When tB<t′≤t, each row yes Corresponds to The reduced vector of ; The spatial dependency comparison module is: in, The contrasting features of Given, for The reconstruction matrix, The (t′-t+B) row vector of and Both are obtained through the spatial causal attention network, The attention weights are used to explain and quantify the spatial effect; Scan comparison After that, the final output of the encoder of the interpretable spatiotemporal attention transformer is given by to establish; to Input to VSN2(·), VSN2(·) transmits the summary information to the decoder; When tB<t′≤t, In the formula The connection matrix with the context vector, where: Evaluate variable returns by using the variable selection weights of VSN2(·) The importance of The load prediction module is used to predict the target power node at a specified future time by using the trained STEF-DHNet model based on the spatiotemporal characteristics of the power load, the real-time power load data of the target power node and the real-time environmental data.
Citation Information
Patent Citations
Power load prediction method and device based on space-time attention mechanism, and medium
CN113610277A
Electrical load prediction method based on time series data periodicity
CN114519471A
Evaluation method of urban spatio-temporal data prediction causal model
CN116227756A
Electrical load prediction method and device based on space-time correlation
CN117175588A
Method for predicting the destination location of a vehicle
US20230194288A1
Cited By
Power line health state evaluation and prediction method and system based on big data
CN120146319A
Strong convective weather forecasting method and device based on space-time cross convergence mapping
CN120233467A
Electric energy metering error correction method and system based on data fusion
CN120234767A
Power distribution network electric vehicle charging load prediction method and system
CN120235324A
Power distribution equipment cooperative control method and system based on edge calculation
CN120262704A