A power distribution network state estimation method based on a double-track residual and admittance clamping physical isolation map attention architecture
Patent Information
- Application Number
- CN202610929499.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-25
AI Technical Summary
当N-1断线突然发生时,深度GAT会将断线区域的高阻抗异常信号经注意力机制平均化扩散至全网,输出虚假的正常状态估计,严重威胁电网安全
[0036]第一、本发明的核心创新在于:将物理拓扑约束从PINN的损失函数级提升至模型前向传播级——通过双轨残差注意力机制将物理知识硬编码进图卷积的消息传递计算图中,同时配合动态梯度平衡机制实现数据驱动与物理约束的自适应协同训练。该方法达到零样本(Zero-shot)拓扑自适应的高精度状态估计。
Smart Images

Figure CN122818007A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to, but is not limited to, the field of power system distribution network operation and control technology, and particularly relates to a distribution network state estimation method based on a dual-track residual and admittance clamp physical isolation graph attention architecture. Background Technology
[0002] Distribution System State Estimation (DSSE) is a core functional module of modern Distribution Management Systems (DMS). Its task is to infer the operating status of each node in the entire network, such as voltage amplitude and phase angle, by integrating limited measurement data with the system topology model. Unlike transmission networks, distribution networks face severe challenges, including extremely sparse deployment of measurement equipment (typical coverage of only 10%–30%), high R / X ratios, significant three-phase imbalances, and frequent topology changes (such as line switching, distributed generation connection / disconnection, and N-1 fault line breaks). Traditional Weighted Least Squares (WLS) and Extended Kalman Filter (EKF) methods suffer from problems such as ill-conditioned Jacobian matrix and iterative divergence when insufficient measurement leads to unobservable systems.
[0003] In recent years, Graph Neural Networks (GNNs) have become a research hotspot in the field of Data Sequencing and Grid Security (DSSE) due to their natural suitability for modeling the non-Euclidean topology of power distribution networks. The DSS2 model proposed by Delft University of Technology (IEEE TSG, 2024) has achieved a breakthrough in unobservable scenarios through hypergraph modeling and a weakly supervised WLS strategy; EleGNN introduces node edge feature propagation guided by electrical centrality; and SKNet-GAT explores state estimation schemes based on multi-source data fusion.
[0004] However, existing GNN-based DSSE schemes reveal two types of structural defects when faced with extreme operating conditions in actual power distribution networks.
[0005] (I) Over-smoothing and Attention Collapse in Deep Graph Networks. Distribution networks typically have deep topological layers. The multi-layer message passing in deep GNNs leads to a uniformity in the hidden layer representations of nodes (over-smoothing), making it impossible to distinguish between disconnected and normal nodes in the embedding space. For Graph Attention Networks (GAT), over-smoothing causes attention collapse, where the attention distribution of all nodes becomes uniform, losing the ability to discern key neighborhood information. When N-1 disconnections suddenly occur, deep GATs will average and spread the high-impedance abnormal signal from the disconnected area to the entire network through the attention mechanism, outputting false normal state estimates, seriously threatening power grid safety.
[0006] (II) Fundamental Limitations of Physical Information Soft Constraint Mechanisms. Physical Information Graph Networks (PINN), represented by Physics-Constrained GAT (MDPIProcesses, 2025), attempt to make the model output conform to physical laws by embedding the power flow equation residuals as soft constraint regularization terms in the loss function. However, this post-punishment soft constraint has two major drawbacks: First, the soft constraint only applies gradient guidance during the backpropagation of the training phase and cannot physically block the information transmission of broken branches during the forward propagation of the model's lower layers. In the message aggregation stage of graph convolution, the model will still transmit false power information along physically broken branches, causing the state estimation of the broken region to be incorrectly neutralized by the normal signal of the unbroken region. Second, when the broken position is at the edge of the receptive field of the deep network, the physical gradient and the data gradient are significantly inconsistent in direction on the shared parameters, causing gradient tearing. The two optimization objectives are at odds, causing the training to fail to converge or converge to a pseudo solution that violates physical laws. At a deeper level, the fundamental problem with PINN is that physical knowledge only indirectly affects network parameters as an additional term in the loss function, never entering the model's forward inference computation graph. Forward propagation is entirely driven by statistical associations learned from the training data, lacking hard encoding of physical topological constraints. When the test topology differs from the training topology (e.g., an N-1 disconnection not seen in the training set occurs), the model's forward inference still follows the message passing path learned from the normal topology, leading to systematic estimation bias.
[0007] In summary, there is an urgent need for a novel architecture that can physically block abnormal propagation paths in the low-level forward propagation while maintaining topological sensitivity and physical consistency in deep networks.
[0008] Based on the above analysis, the urgent technical problems that need to be solved in the existing technology are:
[0009] Existing technologies suffer from four major defects when dealing with topological abrupt changes in distribution networks, such as N-1 line breaks, including oversmoothing, attention collapse, gradient tearing, and lack of physical information embedded in forward propagation. Summary of the Invention
[0010] To address the problems of existing technologies, this invention provides a power distribution network state estimation method based on a dual-track residual and admittance clamped physical isolation graph attention architecture. The core innovation of this invention lies in elevating physical topology constraints from the loss function level of PINN to the model forward propagation level—by hard-coding physical knowledge into the message-passing computation graph of graph convolution through a dual-track residual attention mechanism, and simultaneously achieving adaptive collaborative training of data-driven and physical constraints through a dynamic gradient balancing mechanism. This method achieves high-precision state estimation with zero-shot topology adaptation.
[0011] This invention is implemented as follows: A distribution network state estimation method based on a dual-track residual and admittance clamping physical isolation graph attention architecture includes:
[0012] Step S1: Obtain the distribution network topology information and measurement data, and model the distribution network as a graph structure. The input features of each node include voltage amplitude, active power, reactive power, node type identifier and measurement availability flag. The input features of each edge include line resistance R, reactance X and branch closure status flag closed_status.
[0013] Step S2: Embed the physical isolation mask during the message aggregation phase of the forward propagation;
[0014] The admittance magnitude Y_mag = 1 / sqrt(R^2 + X^2) for each branch is calculated. The global baseline admittance Y_max is used for clamping normalization and multiplied by the closed status flag to generate a physical isolation mask P_edge = clamp(Y_mag / Y_max, 0, 1) * closed_status. This ensures that the mask of the disconnected branch (closed_status = 0) is precisely zeroed, thus physically blocking the information transmission of the branch during the attention energy calculation stage. This is different from the ex-post penalty mechanism of Physical Information Neural Network (PINN) which uses physical constraints as an additional term in the loss function.
[0015] Step S3: Construct a deep network consisting of stacked L layers of dual-track residual attention map convolutional layers, with each layer's forward propagation containing both vertical feature residual tracks and horizontal attention residual tracks.
[0016] The vertical feature residual track is used to protect shallow measurement signals from being smoothed by deep convolutions. This is achieved through cross-layer identity residual connections h_new=NonLin(Norm(W_proj*Aggregate(H)))+h_old. The horizontal attention residual track is used to embed physical topology constraints from the forward propagation attention energy level into the network computation graph. During attention energy calculation, the historical attention energy matrix e_prev is maintained across layers. The current layer energy e_ij=s_ij+softplus(beta)*e_prev is injected after being weighted by the learnable constant positive weight softplus(beta). This ensures that the deep attention distribution smoothly inherits the physical topology structure encoded in the shallow layer, preventing the attention distribution from collapsing during topology abrupt changes. At the same time, the physical isolation mask from step S2 is injected into the current layer energy e_ij=e_ij+softplus(gamma)*P_edge in an additive form. Here, beta and gamma are both learnable parameters, and constant positive constraints are achieved through the softplus function.
[0017] Step S4: Embed the voltage magnitude estimate and phase angle estimate of each node in the final layer node through the output network mapping;
[0018] Step S5: During the training phase, a hybrid objective function combining data loss and physical regularization loss is adopted, and a two-layer adaptive weight scheduling mechanism is introduced.
[0019] The first layer uses the physical loss scheduling coefficient phi. In the early stages of training, phi=0 for pure data-driven warm-up, and then phi gradually increases to 1.0 to introduce physical constraints. The second layer is a GradNorm-style dynamic gradient balancing. The shared weight matrix before the output layer is selected as the gradient monitoring anchor point. The ratio of the gradient L2 norm of the data loss and physical loss to the shared weight is calculated and clamping constraints are applied. The physical loss weight lambda_phy is dynamically updated with an exponential moving average to achieve batch-wise adaptive balancing between data-driven and physical constraints. When facing unknown broken topology during the inference stage, only pre-generated broken topology data needs to be loaded. The closed_status field in step S2 automatically sets the broken branch mask to zero, freezes the model parameters, and completes the entire network state estimation in a single forward propagation without fine-tuning.
[0020] Furthermore, the global reference admittance Y_max in step S2 is preset to 20.0. This value ensures a reasonable distribution of the P_edge value range of normal lines in the IEEE 33-node distribution network, while avoiding batch-to-batch numerical collapse caused by dynamically calculating Y_max batch by batch.
[0021] Furthermore, in step S3, the lateral residual fusion weight beta and the physical prior weight gamma are both constrained to be positive through the softplus function: softplus(x)=log(1+e^x), ensuring that historical attention memory and physical admittance always participate in energy calculation with positive impetus, preventing the weights from degenerating to negative values during the learning process; where beta is initialized to 0.5 and gamma is initialized to 1.0, both are tensors of shape [1,K,1] (K is the number of attention heads), and are learned independently among each attention head.
[0022] Furthermore, the attention mechanism in step S3 adopts K>=4 parallel attention heads, each head independently calculates attention score and energy, and after aggregation, the K*d-dimensional spliced features are compressed back to d dimensions through the linear projection layer W_proj (dimensional d rows K*d columns), as a dimensional firewall to prevent the feature dimension from expanding layer by layer, and connected with the vertical feature residual to form a closed loop.
[0023] Furthermore, the GradNorm-style dynamic gradient balancing mechanism in step S5 is specifically implemented as follows: the shared weight W_out1 before the output layer is selected as the gradient monitoring anchor point, and the L2 norms G_data and G_phy of the gradients of the data loss and physical loss with respect to the weight are calculated respectively; the target ratio lambda_hat=clamp(G_data / (G_phy+epsilon),0,100) is calculated, and the physical loss weight is updated with the exponential moving average lambda_phy=(1-eta)*lambda_phy+eta*lambda_hat, and finally the joint loss L=L_data+phi*lambda_phy*L_phy.
[0024] Furthermore, the scheduling strategy for the physical loss scheduling coefficient phi in step S5 is as follows: before training, in the N_warmup rounds (N_warmup>=100), phi=0 is set, and training is driven only by data loss, so that the model first learns the basic distribution pattern of the measurement data; thereafter, phi increases linearly from 0 to 1.0 in the N_ramp rounds, gradually introducing physical constraints;
[0025] In the zero-sample inference stage of step S6, the preprocessing of the disconnected topology test data includes a dimension alignment operation: first, the test set features are inversely normalized to the original physical dimensions using the mean and standard deviation of their own training set, and then renormalized using the mean and standard deviation of the normal topology training set to ensure that the feature distribution of the disconnected data is strictly consistent with that of the training model.
[0026] Another objective of this invention is to provide a distribution network state estimation system based on a dual-track residual and admittance clamping physical isolation graph attention architecture, comprising:
[0027] The data acquisition module is used to acquire distribution network topology information and measurement data, and to model the distribution network as a graph structure. The input features of each node include voltage amplitude, active power, reactive power, node type identifier and measurement availability flag. The input features of each edge include line resistance R, reactance X and branch closure status flag closed_status.
[0028] An embedding module is used to embed physical isolation masks during the message aggregation phase of forward propagation.
[0029] The deep network construction module is used to build a deep network consisting of stacked L layers of dual-track residual attention map convolutional layers. Each forward propagation layer contains both vertical feature residual tracks and horizontal attention residual tracks.
[0030] The mapping module is used to embed the final layer nodes into voltage magnitude estimates and phase angle estimates mapped to each node through the output network;
[0031] The training module employs a hybrid objective function that combines data loss and physical regularization loss during the training phase, and introduces a two-layer adaptive weight scheduling mechanism.
[0032] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the distribution network state estimation method based on the dual-track residual and admittance clamp physical isolation graph attention architecture.
[0033] Another object of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the distribution network state estimation method based on the dual-track residual and admittance clamp physical isolation graph attention architecture.
[0034] Another objective of this invention is to provide an information data processing terminal for implementing the power distribution network state estimation system based on the dual-track residual and admittance clamp physical isolation graph attention architecture.
[0035] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:
[0036] First, the core innovation of this invention lies in elevating physical topological constraints from the loss function level of PINN to the model forward propagation level—by hard-coding physical knowledge into the message passing computation graph of graph convolution through a dual-track residual attention mechanism, while simultaneously achieving adaptive collaborative training of data-driven and physical constraints through a dynamic gradient balancing mechanism. This method achieves high-precision state estimation with zero-shot topological adaptation.
[0037] 1. Physics-Constrained GAT and other PINN methods employ forward propagation-level physical embedding, distinct from the post-processing penalties of traditional PINN. This architecture hard-encodes physical topology constraints into the message-passing computation graph of graph convolution through a dual-track residual mechanism—the P_edge mask directly blocks the transmission of information from broken branches during the attention energy calculation stage, and prev_e is propagated between layers and smoothly integrated with historical physical states—allowing physical constraints to play a role in the forward propagation of network inference. In contrast, PINN methods such as Physics-Constrained GAT only apply soft physical constraints to the loss function, and forward propagation is entirely driven by statistical associations learned from data, resulting in systematic bias when the test topology differs from the training topology. This architecture's forward propagation-level physical embedding fundamentally solves this problem.
[0038] 2. Dual-track residual collaboration breaks down oversmoothing and attention collapse. The vertical feature residual track protects shallow high-frequency measurement signals from reaching deeper layers without transformation through a simple identity addition x=x+x_res, maintaining the distinctiveness of node embeddings even when the network depth is expanded to 8 layers. The horizontal attention residual track (prev_e mechanism) transmits historical physical state information across layers at the attention energy level, and is fused by learnable softplus (beta) weighted fusion. The synergistic effect of the two is: the vertical track protects signal fidelity, and the horizontal track protects topological consistency—when an N-1 disconnection occurs, the P_edge zeroing effect is strengthened layer by layer through prev_e, resulting in a gradual transfer of deep attention rather than abrupt collapse.
[0039] 3. Admittance clamping with hard physical isolation eliminates gradient tearing. `P_edge=clamp(Y_mag / Y_max,0,1)*closed_status` uses the `closed_status` switch to ensure precise zeroing of the forward propagation weights of disconnected branches. Because physical isolation occurs before message aggregation (rather than as a post-hoc penalty of the loss function), the directional consistency between physical and data gradients on shared parameters is guaranteed, fundamentally eliminating the gradient tearing problem common in PINN. The clamping design with global baseline admittance `Y_max=20.0` also avoids batch-to-batch numerical collapse caused by batch-by-batch dynamic calculation of `Y_max`.
[0040] 4. A dynamic gradient balancing mechanism enables adaptive collaborative training between data and physics. Unlike fixed-weight PINN, which requires repeated manual adjustments to the physics loss coefficients, this architecture employs a hierarchical adaptive scheduling mechanism—phi preheating scheduling (purely data-driven for the first 100 rounds) + GradNorm dynamic lambda_phy (automatically adjusted based on the gradient norm ratio)—to automatically adapt the strength of physical constraints as training progresses. The design of replacing L2 with L1 norm in the physics loss further alleviates the gradient direction distortion caused by the significant differences in the dimensions of voltage, phase angle, and power.
[0041] 5. Zero-shot, high-precision topology-adaptive inference. Physical topology constraints are hard-coded into the forward propagation computation graph. Therefore, when faced with broken topologies not seen in the training set, only the `closed_status` field needs to be updated, and the model can automatically adapt in a single inference iteration—unlike solutions such as EleGNN, which require training multiple models separately for different topology conditions. Dimensional alignment preprocessing ensures that the feature distribution of the broken data is strictly consistent with the trained model, a key engineering guarantee for the correctness of zero-shot inference.
[0042] 6. Strong physical interpretability. The P_edge mask directly maps line admittance to information transmission reliability, prev_e records and transmits historical physical states, and beta and gamma, after being constrained by softplus constant positivity, serve as learnable physical thrust weights—every step of the model's aggregation operation has a traceable physical interpretation, meeting the stringent safety and reliability requirements of power systems. Attached Figure Description
[0043] Figure 1 This is a flowchart of the power distribution network state estimation method based on the dual-track residual and admittance clamping physical isolation graph attention architecture provided in the embodiments of the present invention.
[0044] Figure 2 This is a block diagram of a power distribution network state estimation system based on a dual-track residual and admittance clamping physical isolation graph attention architecture, provided in an embodiment of the present invention.
[0045] Figure 3 This is a flowchart of the overall process of the distribution network state estimation method based on the dual-track residual and admittance clamp physical isolation graph attention architecture provided in this embodiment of the invention. It shows the end-to-end process from input data construction, physical mask generation, dual-track RAA forward propagation (including prev_e inter-layer transfer and P_edge additive injection) to state estimation output, and marks the correspondence between steps one to six.
[0046] Figure 4 This is a breakdown diagram of the internal architecture of the RAA attention convolutional layer (RAAGATConv) provided in this embodiment of the invention. It details three sub-modules: (a) Calculation of data-driven attention energy s_ij—linear transformation, multi-head reshaping, source-target attention score summation, and LeakyReLU activation; (b) Lateral attention residual trajectory—prev_e is added to e_ij after being weighted by softplus (beta); (c) Physical admittance prior injection—P_edge is added to e_ij after being weighted by softplus (gamma). Different colored arrows distinguish the data stream (black), physical mask stream (red), and residual memory stream (blue).
[0047] Figure 5 This is a pipeline diagram for generating a physical isolation mask provided in an embodiment of the present invention. It shows the complete pipeline from R,X through Z_mag, Y_mag, Y_mag / Y_max, clamp to multiplied by closed_status. The precise zeroing effect when closed_status=0 and the global clamping design when Y_max=20.0 are highlighted.
[0048] Figure 6This is a schematic diagram of the dynamic gradient balancing mechanism provided in an embodiment of the present invention. It shows a GradNorm-style two-layer adaptive scheduling: the upper layer phi preheating scheduling curve (0 to 1.0, 100 rounds of preheating + 100 rounds of growth), and the lower layer lambda_phy EMA update mechanism (G_data / G_phy ratio calculation, clamping, smoothing).
[0049] Figure 7 This is a schematic diagram of the measurement configuration and disconnection dataset of the IEEE 33-node distribution network test system provided in this embodiment of the invention. The locations of 11 voltage measurement nodes (solid circles), 6 power measurement branches (thick lines), and typical disconnection branches in the disconnection dataset (red crosses) are marked.
[0050] Figure 8 This is a diagram showing the results of the zero-sample disconnection topology robustness test of the RAA architecture provided in this embodiment of the invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0052] like Figure 1 As shown in the figure, the power distribution network state estimation method based on the dual-track residual and admittance clamp physical isolation graph attention architecture provided by the present invention includes the following steps:
[0053] Step S1: Obtain the distribution network topology information and measurement data, and model the distribution network as a graph structure. The input features of each node include voltage amplitude, active power, reactive power, node type identifier and measurement availability flag. The input features of each edge include line resistance R, reactance X and branch closure status flag closed_status.
[0054] Step S2: Embed the physical isolation mask during the message aggregation phase of the forward propagation;
[0055] The admittance magnitude Y_mag = 1 / sqrt(R^2 + X^2) for each branch is calculated. The global baseline admittance Y_max is used for clamping normalization and multiplied by the closed status flag to generate a physical isolation mask P_edge = clamp(Y_mag / Y_max, 0, 1) * closed_status. This ensures that the mask of the disconnected branch (closed_status = 0) is precisely zeroed, thus physically blocking the information transmission of the branch during the attention energy calculation stage. This is different from the ex-post penalty mechanism of Physical Information Neural Network (PINN) which uses physical constraints as an additional term in the loss function.
[0056] Step S3: Construct a deep network consisting of stacked L layers of dual-track residual attention map convolutional layers, with each layer's forward propagation containing both vertical feature residual tracks and horizontal attention residual tracks.
[0057] The vertical feature residual track is used to protect shallow measurement signals from being smoothed by deep convolutions. This is achieved through cross-layer identity residual connections h_new=NonLin(Norm(W_proj*Aggregate(H)))+h_old. The horizontal attention residual track is used to embed physical topology constraints from the forward propagation attention energy level into the network computation graph. During attention energy calculation, the historical attention energy matrix e_prev is maintained across layers. The current layer energy e_ij=s_ij+softplus(beta)*e_prev is injected after being weighted by the learnable constant positive weight softplus(beta). This ensures that the deep attention distribution smoothly inherits the physical topology structure encoded in the shallow layer, preventing the attention distribution from collapsing during topology abrupt changes. At the same time, the physical isolation mask from step S2 is injected into the current layer energy e_ij=e_ij+softplus(gamma)*P_edge in an additive form. Here, beta and gamma are both learnable parameters, and constant positive constraints are achieved through the softplus function.
[0058] Step S4: Embed the voltage magnitude estimate and phase angle estimate of each node in the final layer node through the output network mapping;
[0059] Step S5: During the training phase, a hybrid objective function combining data loss and physical regularization loss is adopted, and a two-layer adaptive weight scheduling mechanism is introduced.
[0060] The first layer uses the physical loss scheduling coefficient phi. In the early stages of training, phi=0 for pure data-driven warm-up, and then phi gradually increases to 1.0 to introduce physical constraints. The second layer is a GradNorm-style dynamic gradient balancing. The shared weight matrix before the output layer is selected as the gradient monitoring anchor point. The ratio of the gradient L2 norm of the data loss and physical loss to the shared weight is calculated and clamping constraints are applied. The physical loss weight lambda_phy is dynamically updated with an exponential moving average to achieve batch-wise adaptive balancing between data-driven and physical constraints. When facing unknown broken topology during the inference stage, only pre-generated broken topology data needs to be loaded. The closed_status field in step S2 automatically sets the broken branch mask to zero, freezes the model parameters, and completes the entire network state estimation in a single forward propagation without fine-tuning.
[0061] like Figure 2 As shown, an embodiment of the present invention provides a distribution network state estimation system based on a dual-track residual and admittance clamping physical isolation graph attention architecture, comprising:
[0062] The data acquisition module is used to acquire distribution network topology information and measurement data, and to model the distribution network as a graph structure. The input features of each node include voltage amplitude, active power, reactive power, node type identifier and measurement availability flag. The input features of each edge include line resistance R, reactance X and branch closure status flag closed_status.
[0063] An embedding module is used to embed physical isolation masks during the message aggregation phase of forward propagation.
[0064] The deep network construction module is used to build a deep network consisting of stacked L layers of dual-track residual attention map convolutional layers. Each forward propagation layer contains both vertical feature residual tracks and horizontal attention residual tracks.
[0065] The mapping module is used to embed the final layer nodes into voltage magnitude estimates and phase angle estimates mapped to each node through the output network;
[0066] The training module employs a hybrid objective function that combines data loss and physical regularization loss during the training phase, and introduces a two-layer adaptive weight scheduling mechanism.
[0067] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the distribution network state estimation method based on the dual-track residual and admittance clamp physical isolation graph attention architecture.
[0068] Another object of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the distribution network state estimation method based on the dual-track residual and admittance clamp physical isolation graph attention architecture.
[0069] Another objective of this invention is to provide an information data processing terminal for implementing the power distribution network state estimation system based on the dual-track residual and admittance clamp physical isolation graph attention architecture.
[0070] Specific implementation of the present invention:
[0071] like Figure 3The purpose of this invention is to overcome the shortcomings of the prior art and provide a distribution network state estimation method based on a dual-track residual and admittance clamped physical isolation graph attention architecture. The core innovation of this invention lies in elevating physical topology constraints from the loss function level of PINN to the model forward propagation level—by hard-coding physical knowledge into the message passing computation graph of graph convolution through a dual-track residual attention mechanism, while simultaneously achieving adaptive collaborative training of data-driven and physical constraints through a dynamic gradient balancing mechanism. This method achieves high-precision state estimation with zero-shot topology adaptation.
[0072] Step 1: Constructing the data structure and input features of the power distribution network diagram
[0073] The distribution network is modeled as a directed graph, with the node set corresponding to the bus (N=33 in the IEEE 33-node system) and the edge set corresponding to the distribution lines (32 branches plus 1 tie switch). Each node's input feature vector contains 8-dimensional node features: voltage magnitude, voltage phase angle, injected active and reactive power, measurement availability flag, and auxiliary identifier features. Each edge's input features include 6-dimensional edge features (line parameters R, X, closed_status flag, etc.), plus 2-dimensional extended edge parameters (real part R and imaginary part X of impedance) for physical mask calculation. For measurement configuration, the 33-node system includes 11 voltage measurement nodes (node numbers 0, 1, 2, 3, 5, 7, 12, 13, 25, 28, 29) and 6 power measurement branches (branch numbers 0, 3, 8, 9, 15, 28).
[0074] Step 2: Generation of Physical Isolation Mask Based on Admittance Clamp
[0075] The physical isolation mask is the first mechanism in this architecture that embeds physical knowledge into the forward propagation. For each branch, the resistance R and reactance X are extracted from its extended edge parameters, and the branch impedance magnitude and admittance magnitude are calculated:
[0076]
[0077] Physical prior weights are generated using a global baseline admittance clamp normalization method.
[0078]
[0079] Here, Y_max=20.0 is the globally preset upper bound of the baseline admittance (to avoid batch-to-batch numerical collapse caused by batch-by-batch max calculation), and closed_status (0 or 1) is the branch closure status indicator. This mask is directly embedded in the message passing stage of graph convolution—unlike PINN, which treats it as a post-hoc penalty for the loss term, this mask plays a role in the attention energy calculation during forward propagation. When a branch is disconnected, closed_status=0 causes P_edge to be precisely zeroed, physically blocking the information transmission of that branch when aggregating neighborhood messages, thus achieving physical circuit breaking in the forward propagation stage.
[0080] Figure 5 Admittance clamp physical isolation mask generation pipeline and P_edge distribution diagram
[0081] Step 3: Forward Propagation of the Dual-Track Residual Attention Map Convolutional Layer
[0082] An L=8 layer DeepSkipRAA network is constructed, with each layer containing a RAA attention convolution operator (RAAGATConv), a linear projection layer, and a LayerNorm normalization layer. The dual-track residual mechanism is the core innovation of this invention; its design philosophy is to elevate the intervention of physical information from the loss function level of PINN to the internal computational graph of the model's forward propagation. The forward propagation process of each layer is detailed below.
[0083] (a) Input transformation and initial attention energy
[0084] The input node features are linearly transformed and reshaped into a multi-head form H = W_in*h, with a shape of [N,K,d], where K=4 is the number of attention heads and d=64 is the dimension of the hidden layer per head. For each edge, the attention scores of the source and target nodes are calculated, and the initial data-driven attention energy is obtained by LeakyReLU activation.
[0085]
[0086] The s_ij generated at this stage is a purely data-driven attention signal, and physical information has not yet been introduced.
[0087] (b) Lateral attention residual trajectory (prev_e mechanism) - Physical smoothing injection
[0088] The lateral attention residual track is the key difference between this architecture and conventional PINN and standard GAT. Its core function is to directly inject historical physical state information into the attention energy computation layer during forward propagation. It maintains the historical attention energy matrix e_prev (the original energy of the previous layer, unnormalized) passed across layers, and fuses the effective attention energy of the current layer with historical information.
[0089]
[0090] Where beta is a learnable parameter (shape [1,K,1], initial value 0.5), the softplus function ensures that beta is always positive. This mechanism has a dual physical meaning: First, when the distribution network topology does not change abruptly, e_prev carries the learning memory of the shallow network for the normal physical topology. After being weighted by softplus(beta), it is smoothly integrated into the deep energy calculation, so that the deep attention distribution smoothly inherits the physical topology structure encoded by the shallow layer, preventing the amplification of attention noise caused by the increase in depth; Second, when N-1 disconnection occurs, since the physical isolation mask P_edge in step two accurately zeros the disconnected branch, this zeroing effect is strengthened layer by layer through the inter-layer transmission of prev_e - the disconnection signal detected by the shallow layer is transmitted to the deep layer as historical memory. The deep attention will not collapse due to the topology change, but will gradually transfer the attention from the disconnected branch to the normal branch. This is essentially a physical state smoothing filter in forward propagation, which is fundamentally different from PINN, which only applies physical constraints at the loss level: the former directly encodes the physical state in the inference computation graph, while the latter only indirectly affects the parameters in the training backpropagation.
[0091] (c) Physical admittance prior injection – hard physical constraint additive embedding
[0092] Inject the attention energy additively into the physical isolation mask from step two:
[0093]
[0094] Where gamma is a learnable parameter (shape [1,K,1], initial value 1.0), subject to constant positive constraint by softplus. This design uses additive injection instead of multiplicative gating because: multiplicative gating will completely reduce the attention of the entire edge to zero and rescale the softmax denominator when P_edge=0, which may cause distortion of the attention distribution of the remaining edges; additive injection adds energy with the zero mask as a negative offset, so that softmax naturally decays the attention of the broken branch to near zero, while keeping the attention distribution ratio of the remaining normal branches unchanged. Thus, physical information has been embedded forward propagated through two paths: (1) P_edge hard mask directly acts on energy calculation (physical isolation); (2) prev_e is passed between layers and smoothly merged (physical state memory).
[0095] (d) Message aggregation and vertical feature residuals
[0096] Perform softmax normalization on all incoming edges of the target node:
[0097]
[0098] After aggregating neighborhood features and compressing the dimensions using linear projection, apply vertical feature residual connections:
[0099]
[0100] W_proj compresses the expanded dimension (K*d=256) after multi-head splicing back to d=64 dimensions, acting as a dimensional firewall. The residual connection h_new=...+h_old is a simple identity addition with no additional learnable parameters—this design ensures that the shallow layer encodes the original measurement signal directly to the deep layer without any transformation, fundamentally breaking the oversmoothing effect. Vertical and lateral residuals work together: vertically protecting high-frequency measurement signals from distortion, and laterally protecting the physical topology memory from collapse.
[0101] Figure 4 RAA Attention Convolutional Layer (RAAGATConv) Internal Architecture Decomposition Diagram
[0102] Step 4: Output Layer and State Estimation
[0103] After L=8 layers of dual-track residual graph convolution, the final layer nodes are embedded through a two-layer fully connected network to output the voltage amplitude and phase angle estimates of each node.
[0104] Step 5: Dynamic gradient balancing mixed loss training
[0105] This step explains how to train the above dual-track residual architecture through a dynamic gradient balancing mechanism—this is also the third innovation that distinguishes this architecture from fixed-weight PINN.
[0106] Data loss is calculated using SmoothL1 weighted data loss:
[0107]
[0108] The physical loss is calculated using the weighted least squares edge residual (GSP-WLS Edge) based on graph signal processing. The weighted residual is then compared with the measured values after back-calculating the power distribution from the model-estimated state.
[0109]
[0110] The core calculation of physical loss uses the L1 norm (absolute value) instead of the traditional L2 squared form:
[0111]
[0112] Where delta is the residual between the measured value and the physical derivation value, R^{-1} is the measurement accuracy weight (inverse covariance), and w is the regularization coefficient. The reason for using the L1 norm instead of L2 is that the dimensions of voltage (~1.0 pu), phase angle (~0.01 rad), and power (~MW / Mvar) in the distribution network differ significantly. The L2 squared operation would exponentially amplify this difference, causing gradient updates to be dominated by large numerical dimensions—the problem of dimensional distortion. The L1 norm linearly maps the residual to the loss contribution, effectively mitigating the distortion of the gradient direction caused by dimensional imbalance. Additional regularization terms include safety constraints such as voltage over-limit penalties, excessive phase angle difference penalties, and line overload penalties.
[0113] Dynamic gradient balancing mechanism (GradNorm style): In the fixed-weight PINN method, the weight coefficients of data loss and physical loss need to be manually and repeatedly adjusted, and the same set of weights cannot adapt to the needs of different training stages—in the early stages of training, the model has not yet learned the basic data distribution, and excessively large physical loss will dominate the gradient direction, leading to non-convergence; in the later stages of training, insufficient physical constraints cannot guarantee physical consistency. This invention introduces an adaptive weight scheduling mechanism, designed in two layers:
[0114] Layer 1: Physical loss scheduling coefficient phi (stage-level scheduling). Before training, for N_warmup=100 rounds, phi is set to 0, and training is driven solely by data loss, allowing the model to learn the basic distribution pattern of the measurement data. Subsequently, phi linearly increases from 0 to 1.0 over N_ramp=100 rounds, gradually introducing physical constraints. This warmup scheduling ensures that physical constraints only take effect after the model has a basic estimation capability.
[0115] The second layer uses GradNorm dynamic weights lambda_phy (batch-adaptive). The shared weight matrix W_out1 before the output layer is selected as the gradient monitoring anchor point, and the L2 norm of the gradients of the data loss and physical loss with respect to this shared weight is calculated respectively.
[0116]
[0117] The physical loss weights are updated using an exponential moving average smoothing method:
[0118]
[0119] Where epsilon is a small constant to prevent division by zero, and eta is the EMA smoothing coefficient. The final joint loss is:
[0120]
[0121] The core advantage of this dynamic gradient balancing mechanism lies in the fact that lambda_phy is automatically adjusted based on the gradient norm ratio of each training batch—when the data gradient is much larger than the physical gradient, the physical weights are automatically increased to strengthen the physical constraints; conversely, when the physical gradient is too large, the upper limit clamp (100) of lambda_hat prevents the physical loss from dominating and causing data fitting degradation. In training practice, lambda_phy converged to approximately 1.34 after 600 training rounds.
[0122] Step Six: Zero-Shot Topology Adaptive Inference
[0123] After training, the model parameters are frozen. When an unknown open-circuit fault occurs in the distribution network, only a pre-generated open-circuit topology dataset (such as 33bus_broken) needs to be loaded, where the closed_status field of the open-circuit branch is set to 0. In step two, multiplying P_edge by closed_status automatically sets the weight of the open-circuit branch to zero. The key preprocessing in the inference stage is dimensional alignment: first, the test set features are denormalized to the original physical dimensions, and then renormalized using the mean / standard deviation of the training set. The model can complete the full network state estimation under the new topology in a single forward propagation without any fine-tuning—that is, zero-sample topology adaptive inference. The key to its success is that the physical topology constraints have been hard-coded into the forward propagation computation graph through P_edge masking and the prev_e mechanism, rather than relying on statistical associations learned during training.
[0124] Figure 3 Overall Flowchart of Distribution Network State Estimation Method Based on Dual-Track Residual and Admittance Clamping Physical Isolation Graph Attention Architecture
[0125] 7.1 Test System and Measurement Configuration
[0126] Taking the IEEE 33-node distribution network standard test system (reference voltage 12.66kV, total active load 3715kW, total reactive load 2300kvar) as an example, there are 11 voltage measurement nodes: nodes 0, 1, 2, 3, 5, 7, 12, 13, 25, 28, 29 (coverage 33.3%); and 6 power measurement branches: branches 0, 3, 8, 9, 15, 28. The node input feature dimension is 8 (V, theta, P, Q, and measurement availability flags, etc.), and the edge input feature dimension is 6 (R, X, closed_status, etc.) plus 2-dimensional extended edge parameters (R, X).
[0127] Figure 7 Schematic diagram of measurement configuration and disconnection dataset for IEEE 33-node distribution network test system
[0128] 7.2 Data Generation and Preprocessing
[0129] Training Data (Normal Topology): Based on a normal 33-node topology, 10,000 power flow samples with random load fluctuations are generated. The true values of node voltages are obtained using an AC power flow solver. The code first divides the full training set and test set at a 9:1 ratio, then uses labeled training graphs from the full training set at a 30% labeling rate for training, stored in the data / 33bus / directory. Disconnection Test Data: A single disconnection topology dataset is pre-generated and stored in the data / 33bus_broken / directory, where some branches have closed_status=0, indicating that the disconnection condition never occurred during training. Normalization Strategy: The mean and standard deviation of each feature are calculated on the training set (standard ruler). After loading the disconnection test set, it is first denormalized back to the original physical dimensions, and then renormalized using the training set mean / standard deviation—dimensional alignment is a key prerequisite for the correctness of zero-shot inference.
[0130] 7.3 Model Hyperparameter Configuration
[0131] The network has 8 layers (L=8), 64 hidden dimensions (d=64), 4 attention heads (K=4), and a global baseline admittance (Y_max=20.0). Initial beta is 0.5 (learnable, with a softplus constant positive constraint), and initial gamma is 1.0 (learnable, with a softplus constant positive constraint). Dropout is 0.0, and the activation function is LeakyReLU (slope=0.2). The optimizer is Adamax (lr=3e-3), trained for 600 epochs. The physical loss scheduler phi is 0 for epochs 1-100 (purely data-driven warm-up), linearly increasing to 1.0 for epochs 101-200, and remaining at 1.0 for epochs 201-600. GradNorm has a lambda_hat upper limit clamp of 100 and a gradient clipping max_norm of 1.0.
[0132] 7.4 Training Process and Results
[0133] Load normal topology training data from data / 33bus / and calculate P_edge=clamp(Y_mag / 20,0,1). Construct an 8-layer DeepSkipRAA model and initialize it with Xavier. Train for 600 epochs using the hierarchical adaptive scheduling from step five. The validation set V_MAE converges to approximately 0.0059 pu, Th_MAE converges to approximately 0.0060 rad (approximately 0.34 degrees), and lambda_phy converges to approximately 1.34.
[0134] Figure 6 RAA Model Training Process: Error Convergence Curve and Two-Layer Adaptive Weight Scheduling Mechanism
[0135] 7.5 Zero-Sample Disconnection Inference
[0136] Load the trained model (parameters frozen). Load the broken test set from data / 33bus_broken / and perform dimensional alignment. For each edge, calculate P_edge = clamp(Y_mag / 20,0,1)*closed_status—the broken branch's P_edge is zeroed. In a single forward propagation, the zeroing of P_edge, through additive injection, significantly attenuates the corresponding attention energy, while the normal topological memory carried by prev_e smoothly guides the attention distribution to migrate to the surviving branch. During evaluation, only the MAE and branch power flow error of the surviving nodes (V_true>0.5 pu) are counted.
[0137] Figure 8 RAA architecture zero-sample disconnection topology robustness test results
[0138] In the description of this invention, unless otherwise stated, "a plurality of" means two or more; the terms "upper," "lower," "left," "right," "inner," "outer," "front end," "rear end," "head," "tail," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0139] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.
[0140] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A distribution network state estimation method based on a dual-track residual and admittance clamping physical isolation graph attention architecture, characterized in that, The power distribution network state estimation method based on the dual-track residual and admittance clamp physical isolation graph attention architecture includes the following steps: Step S1: Obtain the distribution network topology information and measurement data, and model the distribution network as a graph structure. The input features of each node include voltage amplitude, active power, reactive power, node type identifier and measurement availability flag. The input features of each edge include line resistance R, reactance X and branch closure status flag closed_status. Step S2: Embed the physical isolation mask during the message aggregation phase of the forward propagation; The admittance magnitude Y_mag = 1 / sqrt(R^2 + X^2) for each branch is calculated. The global baseline admittance Y_max is used for clamping normalization and multiplied by the closed status flag to generate a physical isolation mask P_edge = clamp(Y_mag / Y_max, 0, 1) * closed_status. This ensures that the mask of the disconnected branch (closed_status = 0) is precisely zeroed, thus physically blocking the information transmission of the branch during the attention energy calculation stage. This is different from the ex-post penalty mechanism of Physical Information Neural Network (PINN) which uses physical constraints as an additional term in the loss function. Step S3: Construct a deep network consisting of stacked L layers of dual-track residual attention map convolutional layers, with each layer's forward propagation containing both vertical feature residual tracks and horizontal attention residual tracks. The vertical feature residual track is used to protect shallow measurement signals from being smoothed by deep convolutions. This is achieved through cross-layer identity residual connections h_new=NonLin(Norm(W_proj*Aggregate(H)))+h_old. The horizontal attention residual track is used to embed physical topology constraints from the forward propagation attention energy level into the network computation graph. During attention energy calculation, the historical attention energy matrix e_prev is maintained across layers. The current layer energy e_ij=s_ij+softplus(beta)*e_prev is injected after being weighted by the learnable constant positive weight softplus(beta). This ensures that the deep attention distribution smoothly inherits the physical topology structure encoded in the shallow layer, preventing the attention distribution from collapsing during topology abrupt changes. At the same time, the physical isolation mask from step S2 is injected into the current layer energy e_ij=e_ij+softplus(gamma)*P_edge in an additive form. Here, beta and gamma are both learnable parameters, and constant positive constraints are achieved through the softplus function. Step S4: Embed the voltage magnitude estimate and phase angle estimate of each node in the final layer node through the output network mapping; Step S5: During the training phase, a hybrid objective function combining data loss and physical regularization loss is adopted, and a two-layer adaptive weight scheduling mechanism is introduced. The first layer uses the physical loss scheduling coefficient phi. In the early stages of training, phi=0 for pure data-driven warm-up, and then phi gradually increases to 1.0 to introduce physical constraints. The second layer is a GradNorm-style dynamic gradient balancing. The shared weight matrix before the output layer is selected as the gradient monitoring anchor point. The ratio of the gradient L2 norm of the data loss and physical loss to the shared weight is calculated and clamping constraints are applied. The physical loss weight lambda_phy is dynamically updated with an exponential moving average to achieve batch-wise adaptive balancing between data-driven and physical constraints. When facing unknown broken topology during the inference stage, only pre-generated broken topology data needs to be loaded. The closed_status field in step S2 automatically sets the broken branch mask to zero, freezes the model parameters, and completes the entire network state estimation in a single forward propagation without fine-tuning.
2. The distribution network state estimation method based on the dual-track residual and admittance clamping physical isolation graph attention architecture as described in claim 1, characterized in that, In step S2, the global reference admittance Y_max is preset to 20.
0. This value ensures a reasonable distribution of the P_edge value range of normal lines in the IEEE 33-node distribution network, while avoiding batch-to-batch numerical collapse caused by dynamically calculating Y_max batch by batch.
3. The distribution network state estimation method based on the dual-track residual and admittance clamping physical isolation graph attention architecture as described in claim 1, characterized in that, In step S3, the lateral residual fusion weight beta and the physical prior weight gamma are both constrained by the softplus function: softplus(x)=log(1+e^x), which ensures that historical attention memory and physical admittance always participate in energy calculation with positive impetus, and prevents the weights from degenerating into negative values during the learning process; where beta is initialized to 0.5 and gamma is initialized to 1.0, both are tensors of shape [1,K,1] (K is the number of attention heads), and are learned independently among each attention head.
4. The distribution network state estimation method based on the dual-track residual and admittance clamping physical isolation graph attention architecture as described in claim 1, characterized in that, The attention mechanism in step S3 uses K>=4 parallel attention heads, each head independently calculates attention score and energy. After aggregation, the K*d-dimensional spliced features are compressed back to d dimensions through a linear projection layer W_proj (dimensional d rows K*d columns), serving as a dimensional firewall to prevent the feature dimensions from expanding layer by layer, and forming a closed loop with the vertical feature residual.
5. The distribution network state estimation method based on the dual-track residual and admittance clamping physical isolation graph attention architecture as described in claim 1, characterized in that, The GradNorm-style dynamic gradient balancing mechanism in step S5 is specifically implemented as follows: the shared weight W_out1 before the output layer is selected as the gradient monitoring anchor point, and the L2 norms G_data and G_phy of the gradients of the data loss and physical loss with respect to the weights are calculated respectively; the target ratio lambda_hat=clamp(G_data / (G_phy+epsilon),0,100) is calculated, and the physical loss weights are updated with the exponential moving average lambda_phy=(1-eta)*lambda_phy+eta*lambda_hat, and finally the joint loss L=L_data+phi*lambda_phy*L_phy.
6. The distribution network state estimation method based on the dual-track residual and admittance clamping physical isolation graph attention architecture as described in claim 1, characterized in that, The scheduling strategy for the physical loss scheduling coefficient phi in step S5 is as follows: before training, in the N_warmup rounds (N_warmup>=100), phi=0 is set, and training is driven only by data loss, so that the model first learns the basic distribution pattern of the measurement data; thereafter, phi increases linearly from 0 to 1.0 in the N_ramp rounds, gradually introducing physical constraints. In the zero-sample inference stage of step S6, the preprocessing of the disconnected topology test data includes a dimension alignment operation: first, the test set features are inversely normalized to the original physical dimensions using the mean and standard deviation of their own training set, and then renormalized using the mean and standard deviation of the normal topology training set to ensure that the feature distribution of the disconnected data is strictly consistent with that of the training model.
7. A distribution network state estimation system based on a dual-track residual and admittance clamping physical isolation graph attention architecture, implementing the distribution network state estimation method based on any one of claims 1-6, characterized in that, The distribution network state estimation system based on the dual-track residual and admittance clamp physical isolation graph attention architecture includes: The data acquisition module is used to acquire distribution network topology information and measurement data, and to model the distribution network as a graph structure. The input features of each node include voltage amplitude, active power, reactive power, node type identifier and measurement availability flag. The input features of each edge include line resistance R, reactance X and branch closure status flag closed_status. An embedding module is used to embed physical isolation masks during the message aggregation phase of forward propagation. The deep network construction module is used to build a deep network consisting of stacked L layers of dual-track residual attention map convolutional layers. Each forward propagation layer contains both vertical feature residual tracks and horizontal attention residual tracks. The mapping module is used to embed the final layer nodes into voltage magnitude estimates and phase angle estimates mapped to each node through the output network; The training module employs a hybrid objective function that combines data loss and physical regularization loss during the training phase, and introduces a two-layer adaptive weight scheduling mechanism.
8. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the distribution network state estimation method based on the dual-track residual and admittance clamp physical isolation graph attention architecture as described in any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the distribution network state estimation method based on a dual-track residual and admittance clamp physical isolation graph attention architecture as described in any one of claims 1-6.
10. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the power distribution network state estimation system based on the dual-track residual and admittance clamp physical isolation graph attention architecture as described in claim 7.