A topology sequential self-repairing method and system based on an active power distribution network
Patent Information
- Application Number
- CN202610932741.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-15
AI Technical Summary
[0004]本发明提供一种基于有源配电网的拓扑序贯自修复方法及系统,解决了现有基于GNN的重构方法因缺乏对故障后拓扑动态演变的感知机制而导致的电网安全问题
[0015]相比于现有技术,本发明的有益效果在于以下所述中的至少一点:
Smart Images

Figure CN122763433A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power distribution control technology, and in particular to a topology sequential self-repair method and system based on active power distribution networks. Background Technology
[0002] With the large-scale integration of distributed power sources, energy storage, and flexible loads into the active distribution network, the topology and power flow distribution of the distribution network are becoming increasingly complex.
[0003] In existing technologies, the field of topology reconfiguration generally adopts deep learning algorithms represented by graph neural networks (GNN). However, existing GNN-based methods have insufficient perception of the dynamic evolution of the distribution network topology after a fault during the reconfiguration process. This can lead to infeasible situations such as power exceeding limits in the generated reconfiguration schemes, making it difficult to quickly output the optimal topology decision that balances network security constraints and maximizes the recovery of power loss loads in complex scenarios. Summary of the Invention
[0004] This invention provides a topology sequential self-repair method and system based on active distribution networks, which solves the power grid security problem caused by the lack of a perception mechanism for the dynamic evolution of the topology after a fault in existing GNN-based reconstruction methods.
[0005] To address the aforementioned technical problems, this invention provides a topology sequential self-healing method based on an active power distribution network, the method comprising: Acquire the electrical and physical attribute data of the target active power distribution network; The electrical attribute data and the physical attribute data are mapped to obtain a power grid heterogeneity diagram; The heterogeneous graph network update technique driven by physical mechanism is used to perform feature extraction and state derivation processing on the heterogeneous graph of the power grid to obtain the state representation vector; Using the state representation vector as the state set and the closed-loop-open-loop pairing of the target active distribution network as the action set, a sequential reconfiguration model is constructed. The sequential reconfiguration model is optimized using near-end strategy optimization technology to obtain a topology sequential self-repair strategy. The optimization process is configured to perform interactive processing on the sequential reconfiguration model and the heterogeneous power grid diagram, and to update the parameters of the sequential reconfiguration model based on the obtained environmental feedback data to obtain the topology sequential self-repair strategy.
[0006] As one preferred embodiment, the electrical attribute data and the physical attribute data are mapped to obtain a power grid heterogeneity diagram, including: The electrical attribute data is sequentially subjected to node feature extraction and type encoding to obtain node heterogeneous features; the physical attribute data is sequentially subjected to edge feature extraction and state identification to obtain edge heterogeneous features. The heterogeneous graph of the power grid is obtained by processing the heterogeneous features of the nodes and the heterogeneous features of the edges using heterogeneous graph modeling technology.
[0007] As one preferred approach, a physical mechanism-driven heterogeneous graph network update technique is used to extract features and derive states from the heterogeneous power grid graph, resulting in a state representation vector, including: Obtain the hidden states of several nodes in the heterogeneous power grid graph; Based on the physical attribute data and the hidden states of several nodes, the heterogeneous power grid graph is processed by message passing based on physical constraints to obtain message vectors of several edges in the heterogeneous power grid graph. Based on the type data of several nodes, a specific weight matrix corresponding to the type data of the nodes is determined; based on the specific weight matrix, a nonlinear transformation is performed on the hidden states of several nodes and the message vectors of several edges to obtain the updated hidden states of several nodes. The updated hidden states of several nodes are globally aggregated to obtain the state representation vector.
[0008] As one preferred embodiment, the process of constructing the closed-loop-open-loop pairing includes: Obtain the current topology and switching status of the target active distribution network; From the current topology state, select a normally open tie switch as the switch to be closed; Based on the current topology, a graph search algorithm is used to identify the loop formed by closing the switch to be closed, and the physical loop is obtained. In the physical loop, a normally closed sectionalizing switch is selected as the switch to be disconnected; The switch to be closed and the switch to be opened are combined to obtain a closed-loop-open-loop pairing action.
[0009] As one preferred embodiment, the process of constructing the sequential reconstruction model includes: Using the topological safety boundary and electrical quantity safety boundary of the target active distribution network as constraints, and with the optimization objective of maximizing the restoration of lost power loads, a reward function for evaluating the set of actions is defined. The sequential reconstruction model is constructed by processing the state set, the action set, and the reward function.
[0010] As one preferred embodiment, the interactive processing of the sequential reconstruction model and the power grid heterogeneity diagram includes: The closed-loop-open-loop pairing action is input into the simulation environment corresponding to the heterogeneous power grid diagram to update the switching state in the heterogeneous power grid diagram, thereby obtaining the updated heterogeneous power grid diagram. Based on the updated grid heterogeneity diagram, a preset reward function is used to process the data to obtain a real-time reward. Based on the real-time reward, it is determined whether the interaction termination condition is met. The real-time reward, the updated grid heterogeneity diagram, and the state representation vector when the interaction termination condition is met are combined to output the environmental feedback data.
[0011] As one preferred embodiment, the step of updating the parameters of the sequential reconstruction model based on the obtained environmental feedback data to obtain the topology sequential self-healing strategy includes: Based on the environmental feedback data, the dominance function estimate is determined; Using the advantage function estimate as input, the policy parameters of the sequential reconstruction model are updated using the proximal policy optimization technique to obtain the updated policy parameters. The updated strategy parameters are used as the topology sequential self-healing strategy.
[0012] The present invention also provides a topology sequential self-healing system based on an active distribution network, comprising: The acquisition module is used to acquire electrical and physical attribute data of the target active power distribution network. The mapping module is used to map the electrical attribute data and the physical attribute data to obtain a power grid heterogeneity diagram. The update module is used to perform feature extraction and state derivation processing on the heterogeneous graph of the power grid using physical mechanism-driven heterogeneous graph network update technology to obtain a state representation vector. The construction module is used to construct a sequential reconfiguration model using the state representation vector as the state set and the closed-loop-open-loop pairing of the target active distribution network as the action set. An optimization module is used to optimize the sequential reconfiguration model using near-end strategy optimization technology to obtain a topology sequential self-repair strategy. The optimization process is configured to perform interactive processing on the sequential reconfiguration model and the heterogeneous power grid diagram, and to update the parameters of the sequential reconfiguration model based on the obtained environmental feedback data to obtain the topology sequential self-repair strategy.
[0013] As one preferred embodiment, the process of constructing the closed-loop-open-loop pairing includes: Obtain the current topology and switching status of the target active distribution network; From the current topology state, select a normally open tie switch as the switch to be closed; Based on the current topology, a graph search algorithm is used to identify the loop formed by closing the switch to be closed, and the physical loop is obtained. In the physical loop, a normally closed sectionalizing switch is selected as the switch to be disconnected; The switch to be closed and the switch to be opened are combined to obtain a closed-loop-open-loop pairing action.
[0014] As one preferred embodiment, the interactive processing of the sequential reconstruction model and the power grid heterogeneity diagram includes: The closed-loop-open-loop pairing action is input into the simulation environment corresponding to the heterogeneous power grid diagram to update the switching state in the heterogeneous power grid diagram, thereby obtaining the updated heterogeneous power grid diagram. Based on the updated grid heterogeneity diagram, a preset reward function is used to process the data to obtain a real-time reward. Based on the real-time reward, it is determined whether the interaction termination condition is met. The real-time reward, the updated grid heterogeneity diagram, and the state representation vector when the interaction termination condition is met are combined to output the environmental feedback data.
[0015] Compared with the prior art, the beneficial effects of the present invention are at least one of the following: This invention acquires electrical and physical attribute data of a target active distribution network; maps the electrical and physical attribute data to obtain a heterogeneous network graph; uses a physical mechanism-driven heterogeneous graph network update technique to extract features and derive states from the heterogeneous graph to obtain state representation vectors; constructs a sequential reconfiguration model using the state representation vectors as the state set and the closed-loop-open-loop pairings of the target active distribution network as the action set; and optimizes the sequential reconfiguration model using a near-end strategy optimization technique to obtain a topology sequential self-repair strategy. The optimization process is configured to interactively process the sequential reconfiguration model and the heterogeneous network graph, and update the parameters of the sequential reconfiguration model based on the obtained environmental feedback data to obtain the topology sequential self-repair strategy.
[0016] Compared with existing technologies, this invention addresses the problems of insufficient perception of the dynamic evolution of the topology after a fault and the tendency to generate unsafe solutions in existing GNNs during distribution network reconfiguration. By introducing a heterogeneous graph network driven by physical mechanisms, it achieves accurate perception of the electrical state and topological relationship of the power grid after a fault. Furthermore, it constructs a sequential decision model and uses near-end policy optimization technology for training, enabling the model to directly output the optimal reconfiguration strategy that is dynamically safe and feasible. This significantly improves the decision-making speed, safety, and reliability of self-recovery after a fault in complex active distribution networks. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a topology sequential self-healing method based on an active power distribution network in one embodiment of the present invention. Figure 2 This is a schematic diagram of the topology sequential self-healing system based on an active power distribution network in one embodiment of the present invention; Figure label: Among them, 11. Acquisition module; 12. Mapping module; 13. Update module; 14. Construction module; 15. Optimization module; 21. Processor; 22. Memory. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0019] In the description of this invention, it should be noted that, unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this specification is for the purpose of describing specific embodiments only and is not intended to limit the invention. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0020] One embodiment of the present invention provides a topology sequential self-healing method based on an active distribution network. For details, please refer to [link to relevant documentation]. Figure 1 , Figure 1 The diagram shown is a flowchart illustrating a topology sequential self-healing method based on an active distribution network according to one embodiment of the present invention. The method includes: S1: Obtain the electrical and physical attribute data of the target active power distribution network; S2: Map the electrical attribute data and the physical attribute data to obtain a power grid heterogeneity diagram; S3: The heterogeneous graph network update technology driven by physical mechanism is used to perform feature extraction and state derivation processing on the heterogeneous graph of the power grid to obtain the state representation vector; S4: Using the state representation vector as the state set and the closed-loop-open-loop pairing of the target active distribution network as the action set, construct a sequential reconfiguration model; S5: The sequential reconfiguration model is optimized using near-end strategy optimization technology to obtain a topology sequential self-repair strategy. The optimization process is configured to perform interactive processing on the sequential reconfiguration model and the heterogeneous power grid diagram, and to update the parameters of the sequential reconfiguration model based on the obtained environmental feedback data to obtain the topology sequential self-repair strategy.
[0021] Specifically, the acquisition of electrical attribute data mainly relies on the real-time acquisition capabilities of the distribution automation system. Specifically, operating electrical quantities such as node voltage amplitude and phase angle, active and reactive power and current of branches are acquired at second-level cycles through the SCADA system or PMU, and snapshot values at the time of fault are taken when a fault occurs; the real-time output of distributed power sources, the SOC of energy storage and charging and discharging power are reported by the DG inverter communication interface and the BMS battery management system, respectively; the adjustable capacity and demand-side response status of flexible loads are obtained through load management terminals or demand response platforms; and the current opening and closing status of each switch is pushed in real time by the distribution automation terminal DTU or FTU in the form of status change messages.
[0022] The acquisition of physical attribute data mainly relies on offline ledgers and GIS systems, because parameters such as line impedance, transformer rated capacity and ratio, conductor type, line length, switch type and node classification do not change under normal operation and fault conditions. These data are exported from the distribution network GIS system or equipment ledger and are static parameters that do not need to be updated in real time after acquisition.
[0023] After acquisition, a crucial alignment process is required: establishing a one-to-one mapping between the device numbers in the SCADA system and the topology node numbers in the GIS system to ensure that each electrical quantity can be accurately attached to the corresponding graph node and edge; simultaneously, missing value completion, outlier filtering, and timestamp alignment are performed on the collected electrical quantities, and the physical parameters are standardized in terms of dimensions. Finally, the cleaned electrical quantities are used as the dynamic features of nodes and edges, and the physical parameters are used as the static features of nodes and edges, mapped into a heterogeneous power grid graph containing electrical edges and physical edges according to a two-level node-branch structure, which serves as the input for subsequent models.
[0024] The process of mapping the electrical attribute data and the physical attribute data to obtain a heterogeneous power grid graph includes: sequentially extracting node features and encoding types from the electrical attribute data to obtain node heterogeneous features; sequentially extracting edge features and identifying states from the physical attribute data to obtain edge heterogeneous features; and processing the node heterogeneous features and edge heterogeneous features using heterogeneous graph modeling techniques to obtain the heterogeneous power grid graph.
[0025] Specifically, node feature extraction and type encoding are performed on the electrical attribute data. Specifically, the operating status quantities of each node, such as buses, distributed generation, energy storage, and flexible loads, are extracted as numerical features from the electrical attribute data. These include node voltage amplitude, voltage phase angle, active power injection, reactive power injection, current output of distributed generation, SOC and charging / discharging power of energy storage, and adjustable capacity of flexible loads. These features reflect the real-time electrical status of the node after a fault. Subsequently, the node type is encoded. One-hot encoding or learnable embedding vectors are used to map different types such as PQ nodes, PV nodes, balancing nodes, DG nodes, energy storage nodes, and flexible load nodes into mutually unambiguous type identifier vectors, enabling the model to distinguish the physical roles of different nodes.
[0026] The physical attribute data is processed for edge feature extraction and status identification. Specifically, the inherent parameters of each branch are extracted from the physical attribute data as numerical features, including line resistance, reactance, line length, transformer rated capacity and turns ratio, conductor type, etc. These parameters reflect the physical transmission capacity of the branch. At the same time, the current operating status of the branch is identified by encoding information such as switch status, whether the branch carries power flow, and whether it is in a fault isolation zone into a status identification vector, which is then concatenated with the physical parameters to form a complete edge feature.
[0027] Heterogeneous graph modeling technology is used to fuse node heterogeneous features and edge heterogeneous features to obtain a power grid heterogeneous graph. Specifically, the topological skeleton of the graph is constructed with all nodes as the vertex set and all branches as the edge set. Then, the node heterogeneous features obtained in the first step are assigned to the corresponding vertices, and the edge heterogeneous features obtained in the second step are assigned to the corresponding edges. At the same time, various edge types are defined according to different feature sources. For example, electrical edges carry power flow features for power flow inference in message passing, and physical edges carry impedance and topological features for constraint propagation. Various associations can also be established between nodes based on electrical connection relationships and physical affiliation relationships.
[0028] By unifying the scattered node and edge features into a heterogeneous graph with a clear structure and rich semantics, subsequent physical mechanism-driven network updates can simultaneously utilize electrical and physical information for feature aggregation and state deduction on the same graph structure.
[0029] Specifically, a physical mechanism-driven heterogeneous graph network update technique is used to perform feature extraction and state derivation processing on the heterogeneous graph of the power grid to obtain a state representation vector. This includes: obtaining the hidden states of several nodes in the heterogeneous graph of the power grid; performing message passing processing based on physical constraints on the heterogeneous graph of the power grid based on the physical attribute data and the hidden states of several nodes to obtain message vectors of several edges in the heterogeneous graph of the power grid; determining a specific weight matrix corresponding to the type data of several nodes based on the type data of several nodes; performing nonlinear transformation processing on the hidden states of several nodes and the message vectors of several edges based on the specific weight matrix to obtain the updated hidden states of several nodes; and performing global aggregation processing on the updated hidden states of several nodes to obtain the state representation vector.
[0030] Specifically, the hidden states of several nodes in the heterogeneous power grid graph are obtained. Specifically, the heterogeneous features of the nodes are used as the initial input. Through a learnable linear transformation or multilayer perceptron, the feature vector of each node is mapped to a low-dimensional hidden state vector. This hidden state vector is a compressed representation of the node in the current topology. Initially, it is equivalent to the embedding of node features and will be continuously updated in subsequent message passing.
[0031] Based on physical attribute data and the hidden states of several nodes, a message passing process based on physical constraints is performed on the heterogeneous graph of the power grid to obtain message vectors for several edges. Specifically, for each edge in the graph, the current hidden states of the two nodes it connects are used as the message source, and the physical attribute data corresponding to the edge is used as the physical constraint. The message vector transmitted from one end node to another is calculated through a message function modulated by physical parameters. This message function is not an arbitrary aggregation function in a normal GNN, but embeds physical relationships such as Ohm's law and Kirchhoff's law into the message calculation. For example, the voltage difference between nodes is weighted by the line impedance to generate the message, so that the content of the message itself reflects the distribution law of physical power flow.
[0032] Then, based on the type encoding of each node, such as PQ node, PV node, balancing node, DG node, energy storage node, flexible load node, etc., the matrix corresponding to the node type is selected from a set of predefined or learnable weight matrices. Different types of nodes correspond to different weight matrices. For example, the weight matrix of PQ node focuses on the transformation of active and reactive power injection, the weight matrix of PV node focuses on the transformation of voltage amplitude maintenance, and the weight matrix of DG node focuses on the transformation of power output fluctuation.
[0033] Next, the message vectors of all edges received by the node are aggregated, and then fed together with the node's current hidden state into a nonlinear transformation function parameterized by a specific weight matrix to obtain the node's new hidden state after this round of updates.
[0034] Preferably, the nonlinear transformation function can also be a fully connected layer with an activation function.
[0035] By employing a global readout function, such as global average pooling, global max pooling, or attention-based weighted summation, the updated hidden states of all nodes are compressed into a fixed-dimensional vector. This vector represents the global state of the entire distribution network under the current topology after a fault.
[0036] Using the state representation vector as the state set and the closed-loop-open-loop pairing of the target active distribution network as the action set, a sequential reconfiguration model is constructed.
[0037] The process of constructing closed-loop-open-loop pairing includes: obtaining the current topology and switch status of the target active distribution network; selecting a normally open tie switch from the current topology as the switch to be closed; based on the current topology, using a graph search algorithm to identify the loop formed by closing the switch to be closed, and obtaining the physical loop; selecting a normally closed sectionalizing switch from the physical loop as the switch to be opened; and combining the switch to be closed and the switch to be opened to obtain the closed-loop-open-loop pairing action.
[0038] In traditional solutions, static one-time output of reconstruction results often leads to transient ring networks or secondary faults. This step transforms it into a safe stepwise decision-making (MDP) process.
[0039] Specifically, in order to mathematically guarantee that the distribution network always maintains a radial shape and is free of loops during the refactoring process, the actions at each step are defined. For a pairing operation combination, that is: ①The system algorithm first selects a currently open tie switch and closes it; ②Use the depth-first search (DFS) algorithm to quickly locate the only physical loop formed in the distribution network due to the closure of the switch; ③ The system simultaneously selects another normally closed sectionalizing switch in this loop to disconnect it.
[0040] Through this closed-loop-open-loop pairing binding operation, the security risks of short-term loop networks generated during the reconstruction intermediate state can be eliminated, ensuring that the number of grid edges and nodes before and after reconstruction can meet the requirements. From a mathematical and logical perspective, the radial shape of the distribution network topology is strictly guaranteed.
[0041] in, Indicates the operation of closing the switch. This indicates a disconnection operation, E represents the set of edges in the power grid, and V represents the set of nodes in the power grid.
[0042] The process of constructing the sequential reconfiguration model includes: defining a reward function to evaluate the set of actions, with the topological safety boundary and electrical quantity safety boundary of the target active distribution network as constraints and maximizing the recovery of lost power load as the optimization objective; and processing the state set, the set of actions, and the reward function to construct the sequential reconfiguration model.
[0043] Specifically, the state representation vector is used as the state set, the closed-loop-open-loop pairing of the target active distribution network is used as the action set, and the topology safety boundary and electrical quantity safety boundary are used as constraints. The goal is to maximize the recovery of lost power load, and a reward function is defined to evaluate the action set.
[0044] Specifically, two types of safety boundaries are first defined: the topology safety boundary requires that the reconstructed network must meet radial constraints, that is, there are no closed loops, and all load nodes must be connected to the power supply nodes; the electrical quantity safety boundary requires that the active power, reactive power and current amplitude of each branch after reconstruction do not exceed their rated capacity, and the voltage amplitude of each node is kept within the allowable deviation range.
[0045] Based on this, the reward function is designed as a sum of three parts: the first part is the amount of power restoration of the lost load. For each successful restoration of power to a load node, the reward is positive, and the larger the restoration amount, the higher the reward; the second part is a safety constraint penalty. If the action causes a loop to form, the branch power to exceed the limit, or the node voltage to exceed the limit, a large negative reward is immediately applied; the third part is an action efficiency penalty. To avoid the strategy from using too many unnecessary switching operations, a small negative reward is given for each action, encouraging the strategy to complete the reconstruction in the fewest steps.
[0046] The state representation vector obtained in the previous step is used as the state space of the sequential decision, that is, the environmental description observed by the policy at each time step; the closed-loop-open-loop pairing is used as the action space, that is, the operation options that the policy can execute at each time step. Each action corresponds to closing one connection switch and opening one segment switch at the same time, ensuring that the network is still radial after reconstruction; the reward function defined in the previous step is used as the environmental feedback signal, that is, the immediate evaluation returned by the environment after each action is executed.
[0047] Then, based on the Markov decision process framework, the state space, action space, and reward function are combined to construct a sequential decision model with PPO as the optimization algorithm. The policy network of this model takes the state representation vector as input and outputs the probability distribution of each closed-loop-open-loop paired action, while the value network takes the state representation vector as input and outputs the value estimate of the current state.
[0048] By transforming the distribution network topology reconfiguration problem from a traditional one-time combinatorial optimization problem into a sequential decision problem, the strategy can execute switching operations step by step. Each step makes a decision based on the current real topology state and obtains immediate feedback, thus naturally adapting to the scenario of dynamic topology evolution after a fault. This avoids the infeasibility problem caused by outdated topology information in a one-time global decision.
[0049] The sequential reconfiguration model is optimized using near-end strategy optimization technology to obtain a topology sequential self-repair strategy. The optimization process is configured to perform interactive processing on the sequential reconfiguration model and the heterogeneous power grid diagram, and to update the parameters of the sequential reconfiguration model based on the obtained environmental feedback data to obtain the topology sequential self-repair strategy.
[0050] The sequential reconstruction model and the heterogeneous power grid graph are interactively processed to obtain environmental feedback data, including: inputting the closed-loop-open-loop pairing action into the simulation environment corresponding to the heterogeneous power grid graph; updating the switch states in the heterogeneous power grid graph to obtain an updated heterogeneous power grid graph; processing the updated heterogeneous power grid graph using a preset reward function to obtain a real-time reward; and determining whether the interaction termination condition is met based on the real-time reward; combining the real-time reward, the updated heterogeneous power grid graph, and the state representation vector when the interaction termination condition is met to output the environmental feedback data.
[0051] Specifically, the strategy network of the sequential reconfiguration model outputs a closed-loop-open-loop pairing action, which selects a tie switch to close and a segment switch to open. This action is sent to the simulation environment constructed by the heterogeneous power grid graph for execution. The simulation environment modifies the switch state identifiers of the corresponding edges in the heterogeneous graph according to this action, changing the state of the selected tie switch from open to closed and the state of the selected segment switch from closed to open. Then, based on the new switch states, it recalculates the power flow distribution of the entire network, updates the power flow feature vector of each edge and the electrical feature vectors such as voltage and power of each node, and finally generates a heterogeneous power grid graph with updated switch states and electrical quantities.
[0052] Then, all node voltages, branch power, switch states, and other information in the updated heterogeneous power grid diagram are substituted into the reward function defined in the first step, and each item is calculated: if the action successfully restores the lost load without triggering any safety constraints, a positive reward is obtained; if the action causes the branch power to exceed the limit, the node voltage to exceed the limit, or a closed loop to be formed, a large negative reward is obtained; at the same time, the number of steps is accumulated as a penalty.
[0053] After the calculation is completed, the obtained scalar value is used as a real-time reward, and it is checked whether the current interaction termination conditions are met. Termination conditions usually include: all network loads have been restored, the maximum number of steps has been reached, or there has been no load restoration for several consecutive steps.
[0054] When the interaction termination condition is met, the real-time reward sequence accumulated in each step of this round of interaction, the finally obtained updated grid heterogeneous diagram, and the state representation vector at the termination time, i.e. the hidden state vector after global aggregation, are packaged into a complete set of environmental feedback data, which are used as training samples for the PPO algorithm to update policy parameters.
[0055] The updated heterogeneous grid diagram includes the latest topology and electrical quantity states. Based on the obtained environmental feedback data, the sequential reconfiguration model is updated with parameters to obtain the topology sequential self-repair strategy. This includes: determining an advantage function estimate based on the environmental feedback data; using the advantage function estimate as input, performing gradient update processing on the strategy parameters of the sequential reconfiguration model using near-end strategy optimization techniques to obtain updated strategy parameters; and using the updated strategy parameters as the topology sequential self-repair strategy.
[0056] Specifically, the environmental feedback data collected during the interaction is arranged in chronological order to form a complete trajectory sequence. Each trajectory includes the state representation vector at each moment, the closed-loop-open-loop paired action executed, the real-time reward obtained, and the final state when the interaction ends. Then, the generalized advantage estimation (GAE) method is used to calculate the advantage value at each step in this trajectory. The calculation method is to subtract the real-time reward of each step from a baseline value based on value network estimation, and then discount and accumulate the rewards of future steps through a decay coefficient to obtain an advantage estimate that takes into account both immediate rewards and long-term returns.
[0057] Substituting the advantage function estimate obtained in the previous step into the objective function of PPO, the core of which is a truncated probability ratio, that is, the ratio of the probability of the new policy performing a certain action to the probability of the old policy performing the same action, multiplied by the advantage value corresponding to that action; then the gradient of this objective function is calculated, but the key to PPO is the introduction of a pruning mechanism. When the probability ratio of the new and old policies exceeds a preset range, the gradient is truncated, preventing the policy from shifting too much in a single update. Subsequently, optimizers such as Adam are used to update the weight parameters of the policy network along the gradient direction, while simultaneously updating the parameters of the value network to improve the accuracy of the baseline estimation.
[0058] After multiple rounds of interactive iterations and policy parameter convergence, the final policy network weights are fixed. This set of weights defines a complete mapping relationship: input the state representation vector after any fault, output the probability distribution of each closed-loop-open-loop pairing action, and this probability distribution is the optimal reconstruction decision under the current topology.
[0059] By solidifying the trained model parameters into a directly deployable decision strategy, when a fault occurs in a real distribution network, the strategy can be input by mapping the electrical and physical quantities after the fault into state representation vectors. This allows the strategy to gradually output a sequence of switching operations that satisfies safety constraints and maximizes the recovery load, without having to solve the optimization problem again each time.
[0060] In another embodiment, the PPO algorithm is used to construct an Actor-Critic reinforcement learning framework to sequentially solve the distribution network topology reconfiguration problem.
[0061] The process iterates continuously according to a closed loop of "state observation - action selection - policy execution - environmental feedback - parameter update" until any of the following termination conditions are met: ① The power loss load recovery rate reached its maximum value; ② There are no legitimate actions in the system that could further improve the power restoration level; ③ Reach the preset maximum number of operation steps limit.
[0062] At this point, a complete optimal topology reconfiguration operation sequence is output. This sequence can achieve optimal self-healing recovery after a fault, provided that the distribution network always maintains a radial operating structure, the node voltage meets the safety boundary, and the current does not exceed the thermal stability limit.
[0063] Through continuous interactive training with the distribution network simulation environment, the system eventually learns the optimal security reconfiguration strategy applicable to actual fault scenarios, realizing rapid, autonomous, and reliable topology self-repair of the active distribution network.
[0064] Another embodiment of the present invention provides a topology sequential self-healing system based on an active distribution network. For details, please refer to [link to relevant documentation]. Figure 2 , Figure 2 The diagram shown is a schematic representation of a topology-sequential self-healing system based on an active power distribution network according to one embodiment of the present invention. The system includes: Module 11 is used to acquire electrical and physical attribute data of the target active power distribution network. Mapping module 12 is used to map the electrical attribute data and the physical attribute data to obtain a power grid heterogeneity diagram; Update module 13 is used to perform feature extraction and state derivation processing on the heterogeneous graph of the power grid using physical mechanism-driven heterogeneous graph network update technology to obtain a state representation vector. Construction module 14 is used to construct a sequential reconfiguration model using the state representation vector as the state set and the closed-loop-open-loop pairing of the target active distribution network as the action set. The optimization module 15 is used to optimize the sequential reconstruction model using near-end strategy optimization technology to obtain a topology sequential self-repair strategy. The optimization process is configured to perform interactive processing on the sequential reconstruction model and the grid heterogeneity diagram, and to perform parameter update processing on the sequential reconstruction model based on the obtained environmental feedback data to obtain the topology sequential self-repair strategy.
[0065] As one preferred embodiment, the process of constructing the closed-loop-open-loop pairing includes: Obtain the current topology and switching status of the target active distribution network; From the current topology state, select a normally open tie switch as the switch to be closed; Based on the current topology, a graph search algorithm is used to identify the loop formed by closing the switch to be closed, and the physical loop is obtained. In the physical loop, a normally closed sectionalizing switch is selected as the switch to be disconnected; The switch to be closed and the switch to be opened are combined to obtain a closed-loop-open-loop pairing action.
[0066] As one preferred embodiment, the interactive processing of the sequential reconstruction model and the power grid heterogeneity diagram includes: The closed-loop-open-loop pairing action is input into the simulation environment corresponding to the heterogeneous power grid diagram to update the switching state in the heterogeneous power grid diagram, thereby obtaining the updated heterogeneous power grid diagram. Based on the updated grid heterogeneity diagram, a preset reward function is used to process the data to obtain a real-time reward. Based on the real-time reward, it is determined whether the interaction termination condition is met. The real-time reward, the updated grid heterogeneity diagram, and the state representation vector when the interaction termination condition is met are combined to output the environmental feedback data.
[0067] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A topology sequential self-healing method based on an active power distribution network, characterized in that, include: Acquire the electrical and physical attribute data of the target active power distribution network; The electrical attribute data and the physical attribute data are mapped to obtain a power grid heterogeneity diagram; The heterogeneous graph network update technique driven by physical mechanism is used to perform feature extraction and state derivation processing on the heterogeneous graph of the power grid to obtain the state representation vector; Using the state representation vector as the state set and the closed-loop-open-loop pairing of the target active distribution network as the action set, a sequential reconfiguration model is constructed. The sequential reconfiguration model is optimized using near-end strategy optimization technology to obtain a topology sequential self-repair strategy. The optimization process is configured to perform interactive processing on the sequential reconfiguration model and the heterogeneous power grid diagram, and to update the parameters of the sequential reconfiguration model based on the obtained environmental feedback data to obtain the topology sequential self-repair strategy.
2. The topology sequential self-healing method based on active power distribution network according to claim 1, characterized in that, The mapping process between the electrical attribute data and the physical attribute data to obtain a power grid heterogeneity diagram includes: The electrical attribute data is sequentially subjected to node feature extraction and type encoding to obtain node heterogeneous features; the physical attribute data is sequentially subjected to edge feature extraction and state identification to obtain edge heterogeneous features. The heterogeneous graph of the power grid is obtained by processing the heterogeneous features of the nodes and the heterogeneous features of the edges using heterogeneous graph modeling technology.
3. The topology sequential self-healing method based on active power distribution network according to claim 1, characterized in that, The heterogeneous graph network update technique driven by physical mechanisms is used to extract features and derive states from the heterogeneous graph of the power grid, resulting in a state representation vector, including: Obtain the hidden states of several nodes in the heterogeneous power grid graph; Based on the physical attribute data and the hidden states of several nodes, the heterogeneous power grid graph is processed by message passing based on physical constraints to obtain message vectors of several edges in the heterogeneous power grid graph. Based on the type data of several nodes, a specific weight matrix corresponding to the type data of the nodes is determined; based on the specific weight matrix, a nonlinear transformation is performed on the hidden states of several nodes and the message vectors of several edges to obtain the updated hidden states of several nodes. The updated hidden states of several nodes are globally aggregated to obtain the state representation vector.
4. The topology sequential self-healing method based on active power distribution network according to claim 1, characterized in that, The process of constructing the closed-loop-open-loop pairing includes: Obtain the current topology and switching status of the target active distribution network; From the current topology state, select a normally open tie switch as the switch to be closed; Based on the current topology, a graph search algorithm is used to identify the loop formed by closing the switch to be closed, and the physical loop is obtained. In the physical loop, a normally closed sectionalizing switch is selected as the switch to be disconnected; The switch to be closed and the switch to be opened are combined to obtain a closed-loop-open-loop pairing action.
5. The topology sequential self-healing method based on active power distribution network according to claim 1, characterized in that, The process of constructing the sequential reconstruction model includes: Using the topological safety boundary and electrical quantity safety boundary of the target active distribution network as constraints, and with the optimization objective of maximizing the restoration of lost power loads, a reward function for evaluating the set of actions is defined. The sequential reconstruction model is constructed by processing the state set, the action set, and the reward function.
6. The topology sequential self-healing method based on active distribution networks as described in claim 4, characterized in that, The interactive processing of the sequential reconstruction model and the power grid heterogeneity diagram includes: The closed-loop-open-loop pairing action is input into the simulation environment corresponding to the heterogeneous power grid diagram to update the switching state in the heterogeneous power grid diagram, thereby obtaining the updated heterogeneous power grid diagram. Based on the updated grid heterogeneity diagram, a preset reward function is used to process the data to obtain a real-time reward. Based on the real-time reward, it is determined whether the interaction termination condition is met. The real-time reward, the updated grid heterogeneity diagram, and the state representation vector when the interaction termination condition is met are combined to output the environmental feedback data.
7. The topology sequential self-healing method based on active distribution networks as described in claim 1, characterized in that, The sequential reconstruction model is updated based on the obtained environmental feedback data to obtain the topology sequential self-healing strategy, including: Based on the environmental feedback data, the dominance function estimate is determined; Using the advantage function estimate as input, the policy parameters of the sequential reconstruction model are updated using the proximal policy optimization technique to obtain the updated policy parameters. The updated strategy parameters are used as the topology sequential self-healing strategy.
8. A topology-sequential self-healing system based on an active power distribution network, characterized in that, include: The acquisition module is used to acquire electrical and physical attribute data of the target active power distribution network. The mapping module is used to map the electrical attribute data and the physical attribute data to obtain a power grid heterogeneity diagram. The update module is used to perform feature extraction and state derivation processing on the heterogeneous graph of the power grid using physical mechanism-driven heterogeneous graph network update technology to obtain a state representation vector. The construction module is used to construct a sequential reconfiguration model using the state representation vector as the state set and the closed-loop-open-loop pairing of the target active distribution network as the action set. An optimization module is used to optimize the sequential reconfiguration model using near-end strategy optimization technology to obtain a topology sequential self-repair strategy. The optimization process is configured to perform interactive processing on the sequential reconfiguration model and the heterogeneous power grid diagram, and to update the parameters of the sequential reconfiguration model based on the obtained environmental feedback data to obtain the topology sequential self-repair strategy.
9. The topology sequential self-healing system based on active distribution network as described in claim 8, characterized in that, The process of constructing the closed-loop-open-loop pairing includes: Obtain the current topology and switching status of the target active distribution network; From the current topology state, select a normally open tie switch as the switch to be closed; Based on the current topology, a graph search algorithm is used to identify the loop formed by closing the switch to be closed, and the physical loop is obtained. In the physical loop, a normally closed sectionalizing switch is selected as the switch to be disconnected; The switch to be closed and the switch to be opened are combined to obtain a closed-loop-open-loop pairing action.
10. The topology sequential self-healing system based on active distribution network as described in claim 9, characterized in that, The interactive processing of the sequential reconstruction model and the power grid heterogeneity diagram includes: The closed-loop-open-loop pairing action is input into the simulation environment corresponding to the heterogeneous power grid diagram to update the switching state in the heterogeneous power grid diagram, thereby obtaining the updated heterogeneous power grid diagram. Based on the updated grid heterogeneity diagram, a preset reward function is used to process the data to obtain a real-time reward. Based on the real-time reward, it is determined whether the interaction termination condition is met. The real-time reward, the updated grid heterogeneity diagram, and the state representation vector when the interaction termination condition is met are combined to output the environmental feedback data.