A power grid data missing completion method based on generative adversarial network and graph embedding
By employing an adversarial generative network with self-purification processing and a dual-task discriminator, false data injection is identified and isolated, and high-fidelity completed data is generated. This solves the problem of power grid data missingness and attack intertwining, and improves the robustness and availability of power grid data completion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 王欣宇
- Filing Date
- 2026-04-10
- Publication Date
- 2026-07-10
Smart Images

Figure CN122364667A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart grid data processing technology, and in particular to a method for missing data completion in power grids using adversarial generative networks and graph embedding. Background Technology
[0002] With the development of Advanced Metering Systems (AMIs) for smart grids, a large amount of grid measurement data provides crucial support for panoramic state perception and optimized scheduling of the power system. However, due to factors such as sensor failures, communication link interruptions, or extreme weather, the actual collected grid measurement data often exhibits significant spatiotemporal gaps, severely restricting the safe operation of downstream advanced applications. In recent years, data-driven deep learning methods have gradually become mainstream, among which spatiotemporal data completion models combining graph embedding and generative adversarial networks (GANs) have gained significant industry favor. These models first capture the spatial topological correlations of grid nodes through graph convolutional neural networks and extract temporal dependencies using temporal networks. Subsequently, through a zero-sum game process between the generator and discriminator, they implicitly learn the intrinsic distribution characteristics of multidimensional spatiotemporal data, thereby achieving the inference and reconstruction of missing grid measurement data.
[0003] While existing technologies have achieved some success in ideal data environments, they still have significant limitations. First, real-world power grids, facing random data gaps, are highly vulnerable to malicious spurious data injection (FDI) attacks. In such cases, traditional models, based on the strong assumption of absolutely pure known observation data, lack mechanisms for identifying and cleaning contaminated data. Once abnormally dirty data is injected into a local node, this data will amplify errors across the entire topology network through graph embedding aggregation mechanisms, causing the final completed value to deviate completely from the actual power grid operating conditions. Second, conventional generative adversarial networks focus solely on fitting purely numerical statistical distributions, ignoring the stringent physical operating laws and topological smoothness characteristics of the power grid's underlying structure. Furthermore, existing single-task discriminators can only unidirectionally determine the authenticity of data, failing to provide fine-grained quantification of cleanliness for outliers mixed in with known data. Therefore, a highly robust missing data completion mechanism with attack resistance and self-cleaning capabilities is urgently needed to address the technical bottleneck of sharply reduced completion accuracy and poor usability caused by the interplay of malicious attacks and data gaps in complex, untrusted network environments. Summary of the Invention
[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.
[0005] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides a method for power grid data missing completion that counteracts generative networks and graph embeddings, to address the problems mentioned in the background art.
[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a method for completing missing power grid data using adversarial generative networks and graph embedding, comprising:
[0007] The original power grid measurement data containing missing values is self-cleaned to identify suspected spoof data injection attack data points in the power grid measurement data and generate a cleanup mask.
[0008] Based on the topology of the power grid, the self-cleaned power grid measurement data, and the cleanup mask, the spatiotemporal embedding features of the power grid measurement data are extracted.
[0009] Data completion is performed using an adversarial generative network consisting of a generator and a dual-task discriminator. The generator generates completed data based on the self-cleaned power grid measurement data and the spatiotemporal embedding features. The dual-task discriminator is used to determine the authenticity of the input data and the cleanliness of the data points in the input data.
[0010] The completed data generated by the generator is filled into the missing positions of the power grid measurement data to obtain the completed power grid data.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0012] 1. This invention introduces a dual self-purification process that combines Kirchhoff's physical laws with spatial topological smoothness, identifying and isolating spoofed data injection attacks at the input source. Combined with a constructed dual mask, it prevents the cascading amplification of spatial errors from tampered dirty data during graph convolution aggregation, thus endowing the power grid state perception system with self-purification capabilities.
[0013] 2. This invention constructs a dual-task discriminator adversarial generative network (PGN), which not only retains the ability of conventional GPNs to determine the authenticity of the overall data distribution but also endows the discriminator with the ability to quantify the cleanliness of local nodes. Driven by the joint loss function, the generator is forced to output high-fidelity data that strictly conforms to the operating conditions of the power grid when filling data gaps caused by attacks. This washes away traces of topological damage left by malicious tampering and solves the problem that the intertwining of missing data and attacks may lead to model mode collapse.
[0014] 3. Furthermore, this invention employs a hard-addressing fusion mechanism based on a credibility matrix at the final output end. This ensures that the AI-generated supplementary data does not overwrite or contaminate real sensor observations during normal operation, acting only as a high-dimensional virtual sensor in natural communication blind spots or attack-isolated areas. This maximizes the absolute integrity and engineering usability of the output data, effectively connecting to and directly supporting advanced applications such as downstream power grid state estimation and optimal power flow. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0016] Figure 1 This is a flowchart illustrating the overall process of a power grid data missing completion method using adversarial generative networks and graph embedding, as described in one embodiment of the present invention. Detailed Implementation
[0017] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0018] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0019] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0020] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0021] Example 1
[0022] Reference Figure 1 This is the first embodiment of the present invention, which provides a method for power grid data missing completion that opposes generative networks and graph embeddings, including:
[0023] S1. Perform self-cleaning processing on the original power grid measurement data containing missing values, identify suspected spurious data injection attack data points in the power grid measurement data, and generate a cleanup mask.
[0024] It should be noted that in practical advanced measurement systems for smart grids, data is not only subject to random loss due to communication failures, but is also highly vulnerable to malicious data injection attacks. If a measurement matrix containing maliciously tampered data points is directly input into the subsequent graph embedding network, these errors will be amplified along the graph topology, leading to mode collapse of the entire generative adversarial network. Therefore, this invention employs a self-cleaning processing method that combines underlying physical laws with spatial topological constraints.
[0025] Furthermore, a physical consistency check based on the physical operating laws of the power grid is first performed. Since spoofed data injection attacks are often based on purely mathematical data manipulation, attackers find it difficult to simultaneously account for complex nonlinear power system equations. Therefore, this embodiment evaluates the physical consistency of the data by calculating the residuals between power grid measurement data and the predicted values of at least one power flow model.
[0026] Specifically, the original power grid measurement data matrix for the input time period is set as follows: ,in The total number of power grid nodes. Let be the time step. Represents a node At any moment The actual observed values (e.g., node voltage magnitude or phase angle) are used. A baseline estimate of the current grid state is obtained using an AC power flow model or a linearized DC power flow model, yielding model predictions. Based on this, define nodes. At any moment physical residuals It can be represented as:
[0027]
[0028] Furthermore, based on this, the physical residuals are mapped to physical confidence scores using a Gaussian kernel function. :
[0029]
[0030] in, This is the tolerance hyperparameter for physical residuals.
[0031] It needs to be explained that if a certain measurement value deviates significantly from the power flow prediction range constrained by Kirchhoff's current / voltage law (i.e., ... If the value is extremely large, then its physical confidence level is... The sharp decline to near zero indicates that the data is likely maliciously injected dirty data that defies the laws of physics.
[0032] Furthermore, due to the presence of naturally missing values (NaN) in the raw measurement data, conventional power flow solvers may not be able to directly calculate them. Therefore, before performing physical consistency and topology smoothness checks, the system must first temporarily fill in the missing points using grid state estimates or historical steady-state values to ensure the convergence of the baseline power flow prediction. Moreover, when calculating physical residuals and topology discrepancies, confidence assessment is only performed on data points that actually have physical observations (i.e., those without naturally missing values). For data points that are already in a naturally missing state, the confidence calculation process is skipped, and the missing value mask generated by the system later takes over.
[0033] Furthermore, following the physical consistency check, a topology smoothness check based on the power grid topology is performed. Besides physical constraints, adjacent nodes in the power grid space typically exhibit smooth evolution characteristics in electrical quantities (e.g., bus voltage). Malicious tampering with a single or local node will inevitably disrupt this local topology smoothness. Therefore, in this embodiment, we perform the topology smoothness check by quantifying the difference between the data of any node and the data of its first- or multi-order neighboring nodes in the power grid topology graph.
[0034] Specifically, define nodes At any moment Topological dissimilarity for:
[0035]
[0036] in, Represents a node The set of neighboring nodes, The number of neighboring nodes; This refers to the edge weights in the power grid topology adjacency matrix (which can be set based on the admittance between nodes or the physical line distance to reflect strong electrical correlation). Furthermore, it is important to emphasize that when calculating the dissimilarity, if neighboring nodes... If the data itself contains naturally missing values, these will be removed during summation and the denominator will be adjusted accordingly.
[0037] Similarly, topological dissimilarity is converted into topological confidence score. :
[0038]
[0039] in, This represents the tolerance for topological differences.
[0040] It should be noted that by performing a topology smoothness test based on the power grid topology, even if an attacker constructs an attack vector that satisfies the baseline power flow equation (blind zone attack), such a local spatial mutation will still be captured by the topology smoothness test because it is difficult for the attacker to simultaneously tamper with all the surrounding sensor nodes of the node.
[0041] Furthermore, to prevent misjudgments based on a single verification criterion (e.g., normal topology switching operations in the power grid may cause temporary topology non-smoothness), this embodiment also combines physical consistency verification results and topology smoothness verification results to calculate the nodes. At any moment Overall confidence level :
[0042]
[0043] in, This is the weighting balance coefficient, which can be adaptively configured according to the actual operating characteristics of the power grid (for example, in areas with severe load fluctuations, this weighting balance coefficient can be appropriately increased to increase the weight of physical constraints).
[0044] Furthermore, based on the comprehensive confidence level, a sanitization mask matrix with the same dimension as the measurement matrix is generated according to preset rules (e.g., a hard threshold truncation mechanism). For the elements in the matrix :
[0045]
[0046] in, This is a preset security confidence threshold. When This indicates that the data point has been identified as dirty data suspected of being injected by a fake data attack.
[0047] Subsequently, the generated purification mask is used to shield and isolate the original measurement data.
[0048] Specifically, will The original observations at the corresponding locations are forcibly reset to a missing state (i.e., assigned the NaN label).
[0049] It should be noted that through the above processing, the original measurement matrix not only includes naturally missing points caused by communication failures, but also adds actively isolated abnormal dirty data missing points. The advantage of doing so is that rather than allowing subsequent processing to be at risk of being contaminated with erroneous information, it is better to decisively remove these erroneous information first, and then use the characteristics of globally uncontaminated healthy nodes, in conjunction with subsequent graph embedding and adversarial generative networks, to re-infer and complete these positions, thereby giving the entire power grid system a self-purification capability against attacks.
[0050] S2. Based on the power grid topology, self-cleaned power grid measurement data, and cleanup mask, extract the spatiotemporal embedding features of the power grid measurement data.
[0051] It should be noted that the power grid measurement data in the power grid measurement matrix are not isolated pixels in Euclidean space, but rather non-Euclidean graph structure data that strongly depends on the physical wiring method of the power grid (spatial topology) and the dynamic evolution of load / generation (time series). Therefore, the present invention constructs a cascaded architecture combining a graph convolutional network (GCN) and a gated recurrent unit (GRU) for the power grid system. By mapping discrete and incomplete measurement signals into high-dimensional, continuous spatiotemporal embedding vectors, it provides the generator with prior contextual guidance.
[0052] Furthermore, in order to extract the spatiotemporal embedding features of power grid measurement data, a multi-dimensional graph node feature matrix needs to be constructed first. In order for the graph neural network to be able to perceive data missingness (i.e., to distinguish whether the data is naturally missing or actively cleaned up due to an attack), this embodiment has made special designs at the input end of the graph convolutional network.
[0053] Specifically, for each time step The power grid measurement data after self-purification The cleanup mask generated by S1 And a missing mask used to identify the locations of missing original power grid data. Feature dimensions are concatenated. Specifically, when original data is missing due to natural causes such as communication failures, the corresponding missing mask element is set to 1; if the data was collected normally and is not missing, the corresponding missing mask element is set to 0. This constructs the time-series data. The initial input node feature matrix of the graph convolutional network :
[0054]
[0055] in, This indicates a splicing operation along the feature dimension.
[0056] It should be noted that by introducing a dual mask (i.e., a cleanup mask and a missing mask) as auxiliary features, the neural network can learn autonomously during the aggregation process to assign extremely low aggregation weights to neighboring nodes that are missing or judged as dirty data, thereby avoiding the propagation of erroneous information across nodes.
[0057] Furthermore, after constructing the node feature matrix, spatial features are extracted using a graph convolutional network. The aforementioned multidimensional graph node feature matrix is input into a multi-layer graph convolutional network. In each layer, nodes interact with their neighboring nodes based on the physical topology of the power grid.
[0058] Specifically, let the adjacency matrix of the power grid topology be... (Where, each element in the matrix represents the admittance value or connection state of the transmission line), its degree matrix is... To prevent gradient vanishing caused by multi-layer aggregation, this embodiment uses a normalized Laplacian matrix for spatial information aggregation. For the... The update formula for layer graph convolution is expressed as:
[0059]
[0060] in, The adjacency matrix with added self-loops (meaning that nodes not only receive neighbor information but also retain their own information, which conforms to the node's own voltage inertia law). for The degree matrix; For the first Hidden features of the layer, and ; This is the learnable weight matrix for this layer, used to extract spatial electrical features of a specific frequency band; It is the ReLU nonlinear activation function.
[0061] Specifically, after After layer graph convolution, time steps are extracted. Spatial feature matrix containing global topological correlation of power grid .
[0062] Furthermore, while extracting spatial features as described above, a gated recurrent unit (GRU) is used to capture temporal dependencies, forming spatiotemporal embedded features. Since the operating state of the power grid (e.g., peak load variations) exhibits significant temporal continuity, we extract spatial feature sequences from each time step. Inputs are sequentially fed into a gated loop unit to capture their time-dependent dynamics.
[0063] Specifically, at any given moment The internal operation logic of the gated loop unit is set as follows:
[0064] Regarding the update gate :
[0065]
[0066] For the reset door :
[0067]
[0068] Candidate hidden state :
[0069]
[0070] The final hidden state at the current moment :
[0071]
[0072] in, This is the hidden state from the previous moment; and These are all parameter matrices and bias terms that the network can learn; is the Sigmoid activation function; * represents the Hadamard product.
[0073] It's important to explain that the reset gate determines the degree to which past spatial states influence the current state. For example, when a sudden topology switch occurs in the power grid (e.g., circuit breaker tripping), the system steady state is disrupted. In this case, the network will automatically learn to bring the reset gate close to 0 to cut off interference from historical normal states on the current fault state. The update gate, on the other hand, determines the proportion of historical information retained. For instance, during periods of stable power grid operation, when the load curve transitions smoothly, the update gate will maintain a higher value, thus relying on historical states to smooth the current spatial characteristics. Ultimately, the continuous hidden state sequence output by the GRU network, which serves as the extracted spatiotemporal embedding feature, is denoted as... .
[0074] S3. Data completion is performed using an adversarial generative network consisting of a generator and a dual-task discriminator. The generator generates completed data based on the self-cleaned power grid measurement data and spatiotemporal embedding features. The dual-task discriminator is used to determine the authenticity of the input data and the cleanliness of the data points in the input data.
[0075] Furthermore, conventional generative adversarial networks typically employ a single-task discriminator, which relies solely on the statistical distribution of numerical values to distinguish the authenticity of data matrices. This coarse-grained discrimination method cannot identify localized spoofed data injection attacks mixed in with the actual operation of the power grid. To overcome this technical problem, the present invention introduces a dual-task mechanism, enabling the discriminator to not only distinguish between genuine and spoofed data but also to determine whether the data has been maliciously contaminated.
[0076] Furthermore, firstly, we construct a generator, denoted as... The generator's task is to perform high-fidelity reconstruction of missing or isolated dirty data based on contextual clues. In this embodiment, the generator's input includes: raw power grid measurement data retaining only known observations. Missing mask indicating the location of missing data Cleanup mask And the spatiotemporal embedding features extracted by S2 The generator outputs the completed measurement data matrix. Its forward propagation process can be represented as:
[0077]
[0078] It is important to emphasize that, in terms of network structure, the generator can employ a decoder architecture that includes deconvolutional layers or multilayer perceptrons (MLPs) to generate data by mapping spatiotemporal embedded features back to the original node data space.
[0079] Specifically, in this embodiment, the generator adopts a structure that cascades feature concatenation with a multi-layer fully connected network. The generator first uses a concatenation layer to concatenate the original power grid measurement data (containing only known observations), the missing mask, and the cleaned-up mask along the feature dimension into a joint initial tensor. Then, the spatiotemporal embedding features are used as conditional constraint variables and fused with this joint initial tensor through element-wise addition or channel concatenation. Finally, the fused features are input into a decoding network consisting of at least two fully connected layers and a one-dimensional transposed convolutional layer, and the decoding output is a completed data matrix of the same dimension as the original data. Simultaneously, the shared feature extraction backbone of this dual-task discriminator uses a multi-layer one-dimensional convolutional neural network (1D-CNN) to extract local sequence features. The backbone network ends with two parallel branch networks: the first branch maps to scalar probabilities via a fully connected layer and a sigmoid activation function, denoted as . The second branch is mapped to a probability matrix with the same dimension as the input through a deconvolutional network symmetric to the backbone network, denoted as... .
[0080] Furthermore, after constructing the generator, a dual-task discriminator also needs to be built. In this embodiment, the discriminator adopts a network topology combining a shared feature extraction backbone and dual prediction heads. When an arbitrary measurement data matrix is input... When the data (which could be actual measurement data or supplementary data from the generator) is used, its output is divided into two parts:
[0081] The first part is the output of the authenticity judgment header, which is... This represents the probability that the input data as a whole belongs to the real power grid data.
[0082] The second part is the output of the cleanliness assessment head, which is... It represents the probability that each node in the input matrix and each data point at each time step is not attacked and is in a clean state (i.e., the ability to reconstruct the clean mask).
[0083] Furthermore, in the discriminator optimization phase of adversarial training, data containing real power grid measurements (i.e., uncontaminated historical health datasets) is utilized. and its corresponding real desiccant mask and fake data generated by the generator. As training samples. Since the discriminator aims to maximize its ability to distinguish between real and fake data while reconstructing a cleanliness mask, the discriminator's loss function is... It can be defined as a weighted sum of the authenticity classification loss and the cleanliness reconstruction loss:
[0084]
[0085] in, This represents the expectation operation. To balance the hyperparameters for the weights of the dual tasks, The cleanliness reconstruction loss is calculated using binary cross-entropy:
[0086]
[0087] It should be noted that by introducing cleanliness reconstruction loss, the discriminator can be forced to not only focus on the overall spatiotemporal distribution of the data when extracting features, but also to learn the local node consistency features that conform to the Kirchhoff law of the power grid, thus enabling it to quantify the cleanliness probability of each data point in the mixed data.
[0088] Furthermore, during the generator optimization phase of adversarial training, the present invention also fixes the parameters of the discriminator and trains the generator using a joint loss function. This joint loss function aims to achieve three objectives: fooling authenticity judgments, fooling cleanliness judgments, and anchoring to real and reliable data.
[0089] Specifically, the joint loss function The specific mathematical expression is:
[0090]
[0091] It should be explained that in this joint loss function, the first term is the realism adversarial loss: This is used to drive the overall data manifold generated by the generator to approximate the actual operating conditions of the power grid. The second item is cleanliness to counteract losses: This loss term is used during adversarial training to drive the cleanliness judgment result of the completed data generated by the generator to tend towards a preset target value representing the cleanliness state of the data (i.e., the matrix is all 1s). Its purpose is to force the generator to output clean values that conform to the physical laws of a healthy power grid when filling in the gaps in dirty data isolated due to attacks, thereby fundamentally washing away the topological damage left by the fake data injection attack. The third term is the data consistency loss: ,in, A trusted anchor mask matrix is defined if and only if the data was not originally missing. And the clean data that was determined to be attack-free ( )hour, Only the corresponding element is 1; all other data points that are missing or contaminated have a corresponding position of 0. Denotes the squared L2 norm. and This is the penalty coefficient.
[0092] It should be noted that the data consistency loss is introduced because in the implicit distributed reconstruction of generative adversarial networks, it is necessary to use known and absolutely reliable sensor observations in the power grid as physical anchors. This loss can drive the completed data to maintain a high degree of consistency with the original physical observations at reliable data points, which can not only accelerate model convergence, but also ensure the strict availability and fidelity of the output completed power grid matrix in power dispatching applications.
[0093] Furthermore, by employing the aforementioned alternating optimization strategy of the generator and dual-task discriminator under backpropagation, a converged adversarial generative network model can be obtained. At this point, the generator already possesses the ability to perform highly robust inference in complex and untrusted network environments.
[0094] Specifically, in this embodiment, the alternating optimization strategy is trained using Min-Max game theory and the Adam optimizer. The training process is as follows:
[0095] (i) Fix the generator parameters, extract a batch of real power grid measurement data samples from the training set, and have the generator generate a batch of fake samples. Calculate the gradient based on the discriminator's loss function and backpropagate to adjust the network parameters of the dual-task discriminator. Next iteration update;
[0096] (ii) Fix the parameters of the dual-task discriminator, use the outputs of the generator and discriminator to calculate the gradient based on the joint loss function, and perform an iterative update of the generator's network parameters;
[0097] (iii) According to the preset initial learning rate (e.g., set to 1e) -4 Repeat steps (i) and (ii) above alternately until the discriminator can no longer distinguish between real and fake data (i.e., the value output by the authenticity discriminator approaches 0.5) and the joint loss function converges and stabilizes, thus completing the training of the adversarial generative network model.
[0098] S4. Fill the missing positions of the power grid measurement data with the completed data generated by the generator to obtain the completed power grid data.
[0099] It should be noted that although the completed measurement data matrix reconstructed by the adversarial generative network in S3 closely approximates the actual power grid operating state in terms of statistical distribution and physical topology, any data-driven artificial intelligence model will inevitably have a certain inference variance. For high-security smart grid dispatching, real telemetry data that has not been attacked and has not experienced communication failures (e.g., data collected by normally operating PMUs or RTUs) represents the absolute physical reality of the system and must not be overwritten by the values generated by the neural network. Therefore, this embodiment adopts a hard addressing fusion mechanism based on a credibility matrix to ensure the absolute high fidelity of the final data.
[0100] Specifically, the process of filling the missing positions with the completed data generated by the generator is performed using the following mask fusion formula to obtain the final completed power grid data matrix. :
[0101]
[0102] in, This is the final output of the completed power grid measurement data matrix.
[0103] It needs to be explained that in the trusted anchor mask matrix, a node is considered trusted only if, at any given time, its data has not experienced natural missing data (i.e., missing mask). It also did not suffer from a fake data injection attack (i.e., sanitization mask). The element at that position is 1 only when the condition is met; otherwise, it is 0. Therefore, the first part of the formula... This essentially acts as a physical truth anchor, preserving the absolutely safe and reliable original observation data from the power grid intact, avoiding the introduction of any unnecessary algorithmic noise. The second part of the formula... It is equivalent to a high-dimensional virtual sensor. For natural data blind spots caused by communication interruption, or dirty data positions that have been forcibly cut off and isolated by the system in S1 due to malicious tampering (the value of the corresponding position is 1 at this time), the system will use the health value inferred by the generator under spatiotemporal physical constraints to seamlessly fill them.
[0104] It should be noted that, through the aforementioned hard addressing fusion mechanism, the final power grid data matrix not only inherits the observations from actual power grid sensors to the greatest extent possible, but also fills in data gaps and attack-contaminated blind spots with reasonable inferences that conform to Kirchhoff's laws and topological evolution principles. The completed power grid data matrix has extremely high integrity and physical consistency, and can be directly used as a reliable input source for downstream advanced power grid applications (such as power grid panoramic state estimation, optimal power flow calculation, and fault location). This solves the technical bottleneck of power grid dispatching system malfunctions caused by data loss and malicious attacks in complex and untrusted network environments, and improves the digital twin sensing capability and risk resistance robustness of the smart grid.
[0105] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for power grid data missing completion using adversarial generative networks and graph embeddings, characterized in that, include: The original power grid measurement data containing missing values is self-cleaned to identify suspected spoof data injection attack data points in the power grid measurement data and generate a cleanup mask. Based on the topology of the power grid, the self-cleaned power grid measurement data, and the cleanup mask, the spatiotemporal embedding features of the power grid measurement data are extracted. Data completion is performed using an adversarial generative network consisting of a generator and a dual-task discriminator. The generator generates completed data based on the self-cleaned power grid measurement data and the spatiotemporal embedding features. The dual-task discriminator is used to determine the authenticity of the input data and the cleanliness of the data points in the input data. The completed data generated by the generator is filled into the missing positions of the power grid measurement data to obtain the completed power grid data.
2. The power grid data missing completion method based on adversarial generative networks and graph embedding as described in claim 1, characterized in that, The self-cleaning process includes: By combining the physical consistency test results based on the physical operation law of the power grid and the topology smoothness test results based on the power grid topology, the confidence level of each data point is evaluated, and the cleanup mask is generated according to the preset rules.
3. The power grid data missing completion method based on adversarial generative networks and graph embedding as described in claim 2, characterized in that, The physical consistency test result is obtained by calculating the residual between the power grid measurement data and the predicted value of at least one power flow model.
4. The power grid data missing completion method based on adversarial generative networks and graph embedding as described in claim 2, characterized in that, The topology smoothness test result is obtained by quantifying the difference between the data of any node and the data of at least one neighboring node in the power grid topology diagram.
5. The power grid data missing completion method based on adversarial generative networks and graph embedding as described in claim 1, characterized in that, Extracting the spatiotemporal embedding features of the power grid measurement data includes: Graph convolutional networks are used to aggregate neighborhood information of nodes at each time step to extract spatial features; The time series of the spatial features is input into a gated recurrent unit to capture the time dependence and form the spatiotemporal embedding features.
6. The power grid data missing completion method using adversarial generative networks and graph embedding as described in claim 5, characterized in that, The input information for the graph convolutional network also includes: A cleanup mask and a missing mask used to identify the locations of missing raw power grid data.
7. The power grid data missing completion method using adversarial generative networks and graph embedding as described in claim 1, characterized in that, The training process of the adversarial generative network includes: The dual-task discriminator is trained using training samples containing real power grid measurement data and their corresponding cleanliness masks, enabling the discriminator to learn the ability to distinguish the authenticity of data and the ability to reconstruct the cleanliness probability of data points.
8. The method for power grid data missing completion using adversarial generative networks and graph embedding as described in claim 1 or 7, characterized in that, The training process of the adversarial generative network also includes: The generator is trained using a joint loss function that causes the completed data generated by the generator to be judged as real in the realism judgment of the dual-task discriminator, as clean in the cleanliness judgment, and to be consistent with the original value on known data points that are identified as trustworthy by the cleanliness mask.
9. The power grid data missing completion method using adversarial generative networks and graph embedding as described in claim 8, characterized in that, The joint loss function includes a cleanliness adversarial loss term, which is used during the training of the adversarial generative network to drive the cleanliness judgment result corresponding to the completed data generated by the generator to tend toward a preset target value representing the cleanliness state of the data.
10. The power grid data missing completion method using adversarial generative networks and graph embedding as described in claim 1, characterized in that, The inputs to the generator include: Only the original power grid measurement data of known observations, the missing mask indicating the missing data locations, the cleanup mask, and the spatiotemporal embedding features are retained.