Power distribution system fault restoration method based on hierarchical multi-agent deep reinforcement learning
By employing a hierarchical multi-agent deep reinforcement learning method, this study identifies grid-type and grid-following power sources in power distribution systems, divides fault recovery stages, and constructs a multi-agent model. This approach solves the challenge of collaborative recovery of high-proportion distributed resources, achieves adaptive recovery and resilience enhancement of power distribution systems, addresses the issues of action space explosion and sparse rewards, and ensures the safety and real-time performance of the recovery process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN UNIV
- Filing Date
- 2026-02-12
- Publication Date
- 2026-06-02
Smart Images

Figure CN122136835A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power distribution system fault recovery, and in particular to a power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning. Background Technology
[0002] With the deepening of the global energy transition and the intensification of extreme climate change, the frequency of large-scale power outages caused by natural disasters is increasing. Enhancing the resilience of distribution networks has become crucial to ensuring power supply security. The widespread integration of high-proportion distributed generation sources provides a new approach for autonomous recovery after distribution network faults. However, current fault recovery research mostly focuses on static islanding or single network reconstruction methods, lacking in-depth consideration of the dynamic supporting roles and response capabilities of heterogeneous resources during the recovery process. How to efficiently coordinate massive heterogeneous resources for adaptive islanding construction and multi-stage sequential recovery based on different resource characteristics such as network structure and network integration has become a bottleneck problem in improving the resilience of modern power grids. On the other hand, traditional distribution network fault recovery methods are mostly based on mathematical programming or heuristic search rules. Their deterministic analysis paradigm is difficult to effectively cope with the high-dimensional uncertainties on both the source and load sides, and faces a serious "curse of dimensionality" when applied to large-scale systems, resulting in limited real-time decision-making and generalization capabilities. In recent years, while deep reinforcement learning technology has shown potential in complex decision-making, its direct application to the restoration of new power distribution networks has significant drawbacks: First, the large number of controlled devices leads to an exponential expansion of the joint action space, making it extremely difficult for the model to converge within a finite timeframe; second, due to the long execution cycle of restoration tasks, traditional algorithms face a severe sparse reward problem, making it difficult for the agent to capture effective multi-source collaborative restoration logic. Therefore, there is an urgent need to develop a novel adaptive fault recovery method that can achieve decoupling of long and short time sequences, dynamic topology awareness, and multiple physical constraint guarantees to ensure that power distribution systems with a high proportion of distributed resource access can safely, orderly, and efficiently restore power supply after extreme faults. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning.
[0004] The objective of this invention is achieved through the following technical solution: a fault recovery method for power distribution systems based on hierarchical multi-agent deep reinforcement learning, the method comprising,
[0005] S1. Identify and classify grid-connected and grid-linked power sources in the power distribution system, and divide the fault recovery process into the black start stage, the photovoltaic-storage collaborative recovery stage, the island expansion stage, and the island merging stage.
[0006] S2. Construct a multi-stage fault recovery model for the power distribution system with the goal of maximizing weighted load recovery;
[0007] S3. Construct a fault recovery decision model based on hierarchical multi-agent deep reinforcement learning, including an upper-layer single agent that performs macro-stage planning based on a semi-Markov decision process, and a lower-layer heterogeneous multi-agent group that performs micro-device collaborative control based on a Markov decision process.
[0008] S4. The upper-layer single agent uses a graph attention network to model the topological relationship of the multi-island system. By calculating the attention weights between island nodes, it generates a dynamic merging strategy for deciding the optimal island merging time and objects.
[0009] S5. Through a dual action masking mechanism consisting of task phase isolation mask and physical security shielding mask, invalid actions of lower-layer heterogeneous multi-agents are dynamically filtered out, and the fault recovery decision model is trained using a progressive course learning and training paradigm from easy to difficult.
[0010] Specifically, the objective function of the multi-stage fault recovery model is:
[0011] ;
[0012] In the formula, f is the target value for full system recovery; K is the island set index; N k Let K be the set of load nodes within the k-th isolated island. is the priority weight coefficient for load node i; x is the rated active power of load node i; i For a binary variable representing the power supply state of load i, when x i =1 indicates that power has been restored; otherwise, power has been lost.
[0013] Specifically, the constraints of the multi-stage fault recovery model include:
[0014] Island capacity constraints:
[0015] ;
[0016] In the formula, G k For the set of distributed power sources within island k; For the active power output of each distributed power source; The power output for the grid-type energy storage within this isolated island;
[0017] Node voltage constraints:
[0018] ;
[0019] In the formula, , These are the upper and lower limits of the system node voltage amplitude, respectively;
[0020] Branch power constraints:
[0021] ;
[0022] In the formula, The active power transmitted by branch ij; Its maximum permissible transmission power limit;
[0023] Current flow constraints:
[0024] ;
[0025] In the formula, , These represent the active and reactive power injected into node i, respectively; j is the imaginary unit; V i For node voltage phasors; It is the conjugate of the voltages of adjacent nodes; G is the set of nodes connected to node i; ij B ij Let be the real and imaginary parts of the nodal admittance matrix elements;
[0026] Network topology constraints:
[0027] ;
[0028] In the formula, g represents the restored network topology; The set of all topologies that satisfy the radial operating conditions;
[0029] Operational constraints of grid-based energy storage:
[0030] ;
[0031] In the formula, The energy stored at time t is the output power; , These are the upper and lower limits of output, respectively; Let be the state of charge at time t; Rated capacity; For charge and discharge efficiency; To control the cycle.
[0032] Specifically, the upper-layer single agent is modeled based on a semi-Markov decision process as follows:
[0033] state space Action space External reward function The external reward function The calculation is as follows:
[0034] ;
[0035] In the formula, , , These are the weighting coefficients for each reward and penalty item; To restore incremental weighted load; Rewards are given for completing tasks during specific recovery phases; This is a penalty for violating system operation security constraints.
[0036] Specifically, the upper-layer single agent uses a dual-confrontation deep Q-network to generate the Q-values of each macroscopic action:
[0037] ;
[0038] In the formula, s represents the current state; O represents the action space set; For any action in the action space; State-value stream; For action-advantage flow; These are the learnable parameters of the neural network;
[0039] The dual-battle deep Q-network introduces a priority experience replay mechanism during training, utilizing a SumTree structure to perform importance sampling based on the absolute difference error of the samples, with the sampling probability satisfying:
[0040] ;
[0041] In the formula, The sampling probability; The priority weight of sample i after adjustment by the smoothing factor; The priority weight of the j-th sample in the experience replay pool after adjustment by the smoothing factor; This is a smoothing factor used to adjust the uniformity of the sampling distribution.
[0042] Specifically, the lower-layer heterogeneous multi-agent group is modeled based on a Markov decision process as follows:
[0043] Observation space Action space External reward function The external reward function The calculation is as follows:
[0044] ;
[0045] In the formula, , , , These are the weighting coefficients for the corresponding reward and penalty items, respectively; Rewards are given for voltage frequency stability; To provide consistent incentives for achieving macroeconomic recovery goals; Penalty for the cost of the action; This is a phased incentive designed to target specific important actions.
[0046] Specifically, the lower-layer heterogeneous multi-agent group uses a multi-agent soft actor-critic algorithm for control decisions. The centralized critic network in the multi-agent soft actor-critic algorithm is updated by minimizing the squared Bellman error, and its target value is represented as the expected value on the probability distribution of the full action space at the next time step.
[0047] ;
[0048] In the formula, Evaluation network for target; Discount factor; This is the observation state for the next moment; For any action in the action space at the next moment; The action probability distribution output by the joint policy of the agents; Temperature coefficient;
[0049] The distributed actor network in the multi-agent soft actor-critic algorithm is optimized by minimizing the KL divergence between the local policies of each agent and their corresponding Q-values. Its loss function expression is:
[0050] ;
[0051] In the formula, s represents the current state; For the action of agent i; Let i be the action space of agent i; Select a probability vector for the action of agent i;
[0052] The heterogeneous multi-agent group includes a switching agent, a network-building agent, and a network-following agent;
[0053] The switch agent uses a graph convolutional network (GCN) as its backbone encoding structure, and its feature propagation process satisfies:
[0054] ;
[0055] In the formula, This is the feature matrix of the nodes in the (l+1)th layer; It is a non-linear activation function; This is the local power grid adjacency matrix; W is the degree matrix; (l) This is the weight matrix.
[0056] The network-building agent and the network-following agent adopt a multilayer perceptron (MLP) as the backbone coding structure.
[0057] Specifically, S4 includes,
[0058] Abstract each active power supply island in the distribution network into a node and construct an island relationship diagram:
[0059] ;
[0060] In the formula, V is the set of nodes; E is the set of edges;
[0061] Using the attention mechanism in graph attention networks, calculate the normalized attention coefficients between each isolated node and its neighboring nodes:
[0062] ;
[0063] In the formula; This is the normalized attention coefficient; is the shared feature linear transformation matrix; 'a' is the attention weight vector; , , These are the input feature vectors for isolated nodes i, j, and k, respectively. This is a vector concatenation operation; LeakyReLU is a non-linear activation function; N i This is the set of neighboring islands that can be physically connected to node i.
[0064] The neighbor features are weighted and aggregated based on the normalized attention coefficients to produce graph embedding vectors that contain topological context information;
[0065] The optimal merging timing and target are determined from multiple candidate merging objects based on graph embedding vectors.
[0066] Specifically, the dual-action masking mechanism is as follows:
[0067] Based on the current macroscopic instructions issued by the upper-level single agent, a task phase isolation mask is generated to shield device actions that are unrelated to the goal of that phase.
[0068] By combining the real-time topology of the power grid and the status of equipment, a physical security shielding mask is generated to forcibly shield dangerous actions that may lead to network loops, illegal grid connection or reverse redundancy operations.
[0069] By performing a logical AND operation between the original action space and the isolation mask and physical security mask, the dynamically reduced effective action space is obtained:
[0070] ;
[0071] In the formula, For effective action space; For isolation mask; For physical security shielding; o T These are macro-level instructions.
[0072] Specifically, the progressive course learning and training paradigm is implemented by constructing multiple sub-task courses of increasing difficulty, as follows:
[0073] ;
[0074] In the formula, , , The courses are divided into initial, intermediate, and advanced levels; the skills are progressively improved by unlocking tasks in subsequent recovery stages.
[0075] Initial Course First, enable black start and photovoltaic-storage collaborative recovery tasks, limit the action space, and train the model to master the basic skills of voltage and frequency support and core load recovery;
[0076] Intermediate Course :exist Building upon the existing tasks, further unlock the island expansion phase tasks and train the model to master the ability to reconfigure the grid and expand the power supply range using line switches while maintaining system stability.
[0077] Advanced courses :exist Based on the mission, the full-stage recovery mission, including island merging, is fully opened up, and the training model is trained to master the closed-loop control capability of multi-source collaboration and global optimal merging decision-making in multi-island systems.
[0078] When the load recovery rate or the cumulative reward obtained by the agent in performing a complete recovery task in the current course environment reaches the preset stability index, a smooth switching mechanism is used to advance to the next sub-course with increasing difficulty.
[0079] The present invention has the following advantages:
[0080] 1. This invention constructs an active distribution network fault recovery framework that takes into account the heterogeneous support characteristics of grid-connected and grid-linked resources. By fully exploring the dynamic complementary mechanism between grid-connected and grid-linked power sources at each stage of recovery, it achieves deep coupling between recovery paths and resource characteristics, significantly improving the adaptive recovery capability and system resilience of the distribution system under high-proportion distributed resource access.
[0081] 2. This invention constructs a hierarchical multi-agent decision-making architecture, achieving effective decoupling between macro-planning and micro-execution. By modeling long-sequence recovery strategies and short-sequence device control separately, it effectively solves the problems of "action space explosion" and "curse of dimensionality" caused by the massive number of controlled devices, and alleviates the reward sparsity problem caused by long-sequence tasks, significantly improving the model's learning efficiency and real-time decision-making in large-scale power distribution networks;
[0082] 3. This invention proposes a dynamic island merging strategy based on graph attention networks. By constructing an island relationship graph and autonomously learning the mutual influence weights between islands, it achieves accurate perception of the global topological evolution of a multi-island system. Compared with traditional methods, this invention can quantitatively evaluate the benefits and risks of merging, autonomously decide the optimal merging timing and targets, and effectively enhance the operational stability during the coexistence of multiple islands.
[0083] 4. This invention introduces a dual action masking mechanism based on task phase isolation and physical security rules. By shielding illegal actions that violate the physical operation rules of the power grid (such as ring network operation) and reverse redundancy operations in real time at the decision-making level, it ensures that the agent can conduct safe exploration under strict constraints, guarantees the sequential progress and irreversibility of the recovery process, and solves the security bottleneck of deep reinforcement learning in the power industry application;
[0084] 5. This invention designs a progressive learning and training paradigm to accelerate model convergence. By constructing a sequence of tasks from easy to difficult, the model is guided to master recovery strategies from local skills to global collaboration in stages. Utilizing the knowledge transfer effect, the convergence speed of the agent in complex source-load uncertainty environments is significantly accelerated, and the generalization performance and robustness of the final generated strategy under different failure scenarios are enhanced. Attached Figure Description
[0085] Figure 1 This is a schematic diagram of the fault recovery method of the present invention;
[0086] Figure 2 This is a schematic diagram of the multi-stage adaptive fault recovery framework of the present invention;
[0087] Figure 3 This is a diagram of the fault recovery decision model architecture of the present invention;
[0088] Figure 4 This is a schematic diagram of the dynamic island merging strategy based on graph attention network of the present invention.
[0089] Figure 5 This is a schematic diagram of the dual action mask and course learning training paradigm for the invention. Detailed Implementation
[0090] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention; that is, the described embodiments are merely some embodiments of the invention, and not all embodiments. The components of the embodiments of the invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0091] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0092] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0093] The present invention will be further described below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.
[0094] like Figures 1 to 5 As shown, a fault recovery method for power distribution systems based on hierarchical multi-agent deep reinforcement learning is proposed. This method includes:
[0095] S1. Construct a multi-stage adaptive recovery framework for a high-proportion distributed resource active distribution network, identify and classify grid-connected and grid-following power sources in the distribution system, and construct fault recovery as a progressive four-stage recovery sequence, divided into black start stage, photovoltaic-storage collaborative recovery stage, island expansion stage and island merging stage.
[0096] Phase 1 is the black start phase of grid-connected energy storage, in which grid-connected energy storage with autonomous voltage building capability serves as the sole voltage source to establish the initial voltage and frequency reference of the island and restore the most priority critical load of the node.
[0097] Phase two is the photovoltaic-storage coordinated recovery phase. In the already stable island, grid-connected photovoltaics configured at the same node are connected after meeting the synchronization conditions, switch to maximum power point tracking mode and undertake the main active power output, while the original grid-connected energy storage becomes a regulating unit, dynamically compensating for the difference between photovoltaic output fluctuations and load demand, and jointly maintaining the stability of the island.
[0098] Phase 3 is the island expansion phase. Based on the current power margin of the island, independent grid-connected distributed power sources in the network are evaluated and gradually integrated. By orderly closing the corresponding line switches, the power supply range is extended to the end of the network, more general loads are restored, and the island range is expanded.
[0099] Phase four is the island merging phase. During the stable operation of multiple islands, when there are multiple stable islands in the system and one island experiences a power deficit, the crisis island is merged with the selected healthy island through the interconnection switch to achieve power mutual assistance and improve overall resilience.
[0100] Grid-based energy storage, grid-connected photovoltaic systems, and independent grid-connected distributed power sources possess the following recovery characteristics and work synergistically in each recovery stage:
[0101] Grid-based energy storage, as a black-start power source, has inverters capable of instantaneous startup and autonomously establishing and maintaining islanded voltage. and The baseline capability serves as the sole voltage and frequency support point for the isolated system. Frequency regulation in grid-connected energy storage is achieved through virtual synchronous generator (VSG) technology, satisfying the following dynamic relationships:
[0102] ;
[0103] In the formula, J is the system inertial time constant; Given active power; This represents the actual output active power. f is the damping coefficient; f is the current frequency; f ref This is the reference frequency.
[0104] Its voltage regulation is achieved through droop control and reactive power regulation, satisfying the following relationship:
[0105] ;
[0106] In the formula, V is the output voltage; V ref Reference voltage value; k Q Where is the voltage droop factor; Q is the reactive power;
[0107] The system combines grid-connected photovoltaic (PV) and grid-connected energy storage to form an AC-side coupled grid-connected PV-energy storage system. After the islanding stabilizes, the grid-connected PV is integrated into the AC microgrid and operates in maximum power point tracking (MPPT) mode to provide active power. ,in Equal to available power The output active power of the grid-type energy storage Responsible for dynamically compensating the grid-connected photovoltaic power output With load demand The difference between them, that is:
[0108] ;
[0109] Independent grid-connected distributed power sources employ current source control. Their inverters cannot autonomously establish or support voltage and frequency. They must be connected to the islanded system under the following synchronous grid connection conditions and operate in maximum power point tracking mode to provide active power support. :
[0110] ;
[0111] ;
[0112] ;
[0113] In the formula, It is the voltage difference; It is the voltage difference; Synchronization voltage threshold; The synchronization frequency threshold; Available active power for maximum power point tracking.
[0114] In the Phase 4 island merging phase, the merging process adopts master-slave control, in which healthy islands act as master stations to maintain voltage and frequency references, and crisis islands act as slave stations to synchronize with the grid.
[0115] S2. Based on a multi-stage adaptive recovery framework, a multi-stage fault recovery model for the power distribution system is constructed with the goal of maximizing weighted load recovery.
[0116] The objective function of the multi-stage fault recovery model is:
[0117] ;
[0118] In the formula, f is the target value for full system recovery; K is the island set index; N k Let K be the set of load nodes within the k-th isolated island. The priority weight coefficient for load node i is preset based on the load importance level; x is the rated active power of load node i; i For a binary variable representing the power supply state of load i, when x i =1 indicates that power has been restored; otherwise, power has been lost.
[0119] The constraints of the multi-stage fault recovery model include:
[0120] Island capacity constraints:
[0121] ;
[0122] In the formula, G k For the set of distributed power sources within island k; For the active power output of each distributed power source; The power output for the grid-type energy storage within this isolated island;
[0123] Node voltage constraints:
[0124] ;
[0125] In the formula, , These are the upper and lower limits of the system node voltage amplitude, respectively;
[0126] Branch power constraints:
[0127] ;
[0128] In the formula, The active power transmitted by branch ij; Its maximum permissible transmission power limit;
[0129] Current flow constraints:
[0130] ;
[0131] In the formula, , These represent the active and reactive power injected into node i, respectively; j is the imaginary unit; V i For node voltage phasors; It is the conjugate of the voltages of adjacent nodes; G is the set of nodes connected to node i; ij B ij Let be the real and imaginary parts of the nodal admittance matrix elements;
[0132] Network topology constraints:
[0133] ;
[0134] In the formula, g represents the restored network topology; The set of all topologies that satisfy the radial operating conditions;
[0135] Operational constraints of grid-based energy storage:
[0136] ;
[0137] In the formula, The energy stored at time t is the output power; , These are the upper and lower limits of output, respectively; Let be the state of charge at time t; Rated capacity; For charge and discharge efficiency; To control the cycle.
[0138] S3. Construct a fault recovery decision model based on hierarchical multi-agent deep reinforcement learning, including an upper-layer single agent that performs macro-stage planning based on a semi-Markov decision process, and a lower-layer heterogeneous multi-agent group that performs micro-device collaborative control based on a Markov decision process.
[0139] The fault recovery decision model adopts a two-layer hierarchical architecture, including an upper-layer single agent and a lower-layer heterogeneous multi-agent group, and achieves collaborative optimization through instruction transmission and reward feedback between the upper and lower layers;
[0140] The upper-level single agent is modeled based on a semi-Markov decision process as follows:
[0141] state space : From the global summary vector Graph structure data representing the topological connections of isolated islands Together, they form a graph structure containing the system's total load recovery rate, critical load recovery rate, and average state of charge. The node features of the island topology connection relationship graph structure include the island's active state, current execution stage, and power margin.
[0142] Action space Defined as a discrete set of macroscopic instructions for each recovery island, the instruction set directly corresponds to the five recovery options in the multi-stage recovery framework: black start, collaborative recovery, external expansion, island merging, and hold-and-wait.
[0143] External reward function This includes sparse external rewards, comprising progress rewards based on load recovery, milestone rewards for completing specific phases of tasks, and penalties for violating safety constraints, used to evaluate the long-term effects of macro-level decisions; the external reward function is calculated as follows:
[0144] ;
[0145] In the formula, , , These are the weighting coefficients for each reward and penalty item; To restore incremental weighted load; Rewards are given for completing tasks during specific recovery phases; This is a penalty for violating system operation security constraints.
[0146] The upper-level single agent uses a dual-battle deep Q-network to generate the Q-values for each macroscopic action:
[0147] ;
[0148] In the formula, s represents the current state; O represents the action space set; For any action in the action space; State-value stream; For action-advantage flow; These are the learnable parameters of the neural network;
[0149] The dual-battle deep Q-network introduces a priority experience replay mechanism during training, utilizing a SumTree structure to perform importance sampling based on the absolute difference error of the samples, with the sampling probability satisfying:
[0150] ;
[0151] In the formula, The sampling probability; The priority weight of sample i after adjustment by the smoothing factor; The priority weight of the j-th sample in the experience replay pool after adjustment by the smoothing factor; This is a smoothing factor used to adjust the uniformity of the sampling distribution.
[0152] The lower-level heterogeneous multi-agent group is modeled based on Markov decision processes as follows:
[0153] Observation space This includes real-time observation of local electrical measurements by each heterogeneous actuator and the local topological characteristics of the island to which it belongs.
[0154] Action space This refers to the mode switching actions of power supply equipment, the opening and closing actions of line and load switches, and the grid connection actions of distributed power sources.
[0155] External reward function This includes a dense signal weighted by system stability gain, task progress guidance gain, and key operation guidance reward; its calculation is as follows:
[0156] ;
[0157] In the formula, , , , These are the weighting coefficients for the corresponding reward and penalty items, respectively; Rewards are given for voltage frequency stability; To provide consistent incentives for achieving macroeconomic recovery goals; Penalty for the cost of the action; This is a phased incentive designed to target specific important actions.
[0158] The lower-level heterogeneous multi-agent group employs a multi-agent soft actor-critic algorithm for control decisions, following a centralized training and distributed execution architecture. The centralized critic network in this algorithm is updated by minimizing the squared Bellman error, and its target value is represented as the expected value on the full action space probability distribution at the next time step.
[0159] ;
[0160] In the formula, Evaluation network for target; Discount factor; This is the observation state for the next moment; For any action in the action space at the next moment; The action probability distribution output by the joint policy of the agents; Temperature coefficient;
[0161] The distributed actor network in the multi-agent soft actor-critic algorithm is optimized by minimizing the KL divergence between the local policies of each agent and their corresponding Q-values. Its loss function is expressed as:
[0162] ;
[0163] In the formula, s represents the current state; For the action of agent i; Let i be the action space of agent i; Select a probability vector for the action of agent i;
[0164] The heterogeneous multi-agent group includes a switching agent, a network-building agent, and a network-following agent;
[0165] The switch agent uses a graph convolutional network (GCN) as its backbone encoding structure, and its feature propagation process satisfies:
[0166] ;
[0167] In the formula, This is the feature matrix of the nodes in the (l+1)th layer; It is a non-linear activation function; This is the local power grid adjacency matrix; W is the degree matrix; (l) This is the weight matrix.
[0168] The network-building agent and the network-following agent adopt a multilayer perceptron (MLP) as the backbone coding structure.
[0169] The multi-agent soft actor-critic algorithm uses a soft parameter update mechanism to maintain the target network, and its parameter update process satisfies: ;
[0170] In the formula, For target network parameters; These are the current network parameters; It is a very small smoothing operator.
[0171] S4. The upper-layer single agent uses a graph attention network to model the topological relationship of the multi-island system. By calculating the attention weights between island nodes, it generates a dynamic merging strategy for deciding the optimal island merging time and objects.
[0172] The upper-level single agent abstracts each active power supply island in the distribution network into nodes and constructs an island relationship graph:
[0173] ;
[0174] In the formula, V is the set of nodes, which are currently active islands; E is the set of edges, which are the physical connections between islands.
[0175] The upper-layer single agent utilizes the attention mechanism in graph attention networks to quantify the mutual influence weight between isolated node i and its neighbor j by calculating the normalized attention coefficient between each isolated node and its neighbor nodes. The calculation formula is as follows:
[0176] ;
[0177] In the formula; This is the normalized attention coefficient; is the shared feature linear transformation matrix; 'a' is the attention weight vector; , , These are the input feature vectors for isolated nodes i, j, and k, respectively. This is a vector concatenation operation; LeakyReLU is a non-linear activation function; N i Let i be the set of neighboring islands that may have physical connections with node i.
[0178] The upper-layer single agent performs weighted aggregation of neighbor features based on normalized attention coefficients, generates a graph embedding vector containing topological context information, and uses it as the state space. Components; through the attention coefficient The relative size reflects the priority of different neighboring islands in providing power support or risk sharing, thereby autonomously deciding the optimal timing and target for merging among multiple candidate merge targets.
[0179] S5. Through a dual action masking mechanism consisting of task phase isolation mask and physical security shielding mask, invalid actions of lower-layer heterogeneous multi-agents are dynamically filtered out, and the fault recovery decision model is trained using a progressive course learning and training paradigm from easy to difficult.
[0180] The dual-action masking mechanism is as follows:
[0181] Based on the current macroscopic instructions issued by the upper-layer single agent, a task phase isolation mask is generated to shield device actions that are unrelated to the goal of this phase, so as to ensure the consistency between the lower-layer actuator action space and the global recovery logic.
[0182] By combining the real-time topology of the power grid and the status of equipment, a physical security shielding mask is generated to forcibly shield dangerous actions that may lead to network loops, illegal grid connection or reverse redundancy operations.
[0183] By performing a logical AND operation between the original action space and the isolation mask and physical security mask, the dynamically reduced effective action space is obtained:
[0184] ;
[0185] In the formula, For effective action space; For isolation mask; For physical security shielding; o T These are macro-level instructions.
[0186] When an agent outputs an illegal action command, the fault recovery decision model forces the lower-level actuator to explore the optimal recovery path within the physical safety boundary by either setting the probability of that action in the policy distribution to zero or assigning a negative infinity penalty to the corresponding Q value.
[0187] The progressive learning and training paradigm is implemented by constructing multiple sub-task courses of increasing difficulty, represented as follows:
[0188] ;
[0189] In the formula, , , The courses are divided into initial, intermediate, and advanced levels; the skills are progressively improved by unlocking tasks in subsequent recovery stages.
[0190] Initial Course First, enable black start and photovoltaic-storage collaborative recovery tasks, limit the action space, and train the model to master the basic skills of voltage and frequency support and core load recovery;
[0191] Intermediate Course :exist Building upon the existing tasks, further unlock the island expansion phase tasks and train the model to master the ability to reconfigure the grid and expand the power supply range using line switches while maintaining system stability.
[0192] Advanced courses :exist Based on the mission, the full-stage recovery mission, including island merging, is fully opened up, and the training model is trained to master the closed-loop control capability of multi-source collaboration and global optimal merging decision-making in multi-island systems.
[0193] When the load recovery rate or the cumulative reward obtained by the agent in performing a complete recovery task in the current course environment reaches the preset stability index, the course will be smoothly switched to the next sub-course with increasing difficulty through a smooth switching mechanism.
[0194] The above description is merely a preferred embodiment of the present invention and does not constitute any limitation on the present invention. Any person skilled in the art can make many possible variations and modifications to the technical solution of the present invention, or modify it into equivalent embodiments, without departing from the scope of the present invention. Therefore, any modifications, equivalent changes, and alterations made to the above embodiments based on the technology of the present invention without departing from the scope of the present invention are within the protection scope of the present invention.
Claims
1. A fault recovery method for power distribution systems based on hierarchical multi-agent deep reinforcement learning, characterized in that: The method includes, S1. Identify and classify grid-connected and grid-linked power sources in the power distribution system, and divide the fault recovery process into the black start stage, the photovoltaic-storage collaborative recovery stage, the island expansion stage, and the island merging stage. S2. Construct a multi-stage fault recovery model for the power distribution system with the goal of maximizing weighted load recovery; S3. Construct a fault recovery decision model based on hierarchical multi-agent deep reinforcement learning, including an upper-layer single agent that performs macro-stage planning based on a semi-Markov decision process, and a lower-layer heterogeneous multi-agent group that performs micro-device collaborative control based on a Markov decision process. S4. The upper-layer single agent uses a graph attention network to model the topological relationship of the multi-island system. By calculating the attention weights between island nodes, it generates a dynamic merging strategy for deciding the optimal island merging time and objects. S5. Through a dual action masking mechanism consisting of task phase isolation mask and physical security shielding mask, invalid actions of lower-layer heterogeneous multi-agents are dynamically filtered out, and the fault recovery decision model is trained using a progressive course learning and training paradigm from easy to difficult.
2. The power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning according to claim 1, characterized in that: The objective function of the multi-stage fault recovery model is: ; In the formula, f is the target value for full system recovery; K is the island set index; N k Let K be the set of load nodes within the k-th isolated island. is the priority weight coefficient for load node i; x is the rated active power of load node i; i For a binary variable representing the power supply state of load i, when x i =1 indicates that power has been restored; otherwise, power has been lost.
3. The power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning according to claim 2, characterized in that: The constraints of the multi-stage fault recovery model include: Island capacity constraints: ; In the formula, G k For the set of distributed power sources within island k; For the active power output of each distributed power source; The power output for the grid-type energy storage within this isolated island; Node voltage constraints: ; In the formula, , These are the upper and lower limits of the system node voltage amplitude, respectively; Branch power constraints: ; In the formula, The active power transmitted by branch ij; Its maximum permissible transmission power limit; Current flow constraints: ; In the formula, , These represent the active and reactive power injected into node i, respectively; j is the imaginary unit; V i For node voltage phasors; It is the conjugate of the voltages of adjacent nodes; G is the set of nodes connected to node i; ij B ij Let be the real and imaginary parts of the nodal admittance matrix elements; Network topology constraints: ; In the formula, g represents the restored network topology; The set of all topologies that satisfy the radial operating conditions; Operational constraints of grid-based energy storage: ; In the formula, The energy stored at time t is the output power; , These are the upper and lower limits of output, respectively; Let be the state of charge at time t; Rated capacity; For charge and discharge efficiency; To control the cycle.
4. The power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning according to claim 1, characterized in that: The upper-layer single agent is modeled based on a semi-Markov decision process as follows: state space Action space External reward function The external reward function The calculation is as follows: ; In the formula, , , These are the weighting coefficients for each reward and penalty item; To restore incremental weighted load; Rewards are given for completing tasks during specific recovery phases; This is a penalty for violating system operation security constraints.
5. The power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning according to claim 4, characterized in that: The upper-layer single agent uses a dual-confrontation deep Q-network to generate the Q-values for each macroscopic action: ; In the formula, s represents the current state; O represents the action space set; For any action in the action space; State-value stream; For action-advantage flow; These are the learnable parameters of the neural network. The dual-battle deep Q-network introduces a priority experience replay mechanism during training, utilizing a SumTree structure to perform importance sampling based on the absolute difference error of the samples, with the sampling probability satisfying: ; In the formula, The sampling probability; The priority weight of sample i after adjustment by the smoothing factor; The priority weight of the j-th sample in the experience replay pool after adjustment by the smoothing factor; This is a smoothing factor used to adjust the uniformity of the sampling distribution.
6. The power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning according to claim 4, characterized in that: The lower-layer heterogeneous multi-agent group is modeled based on a Markov decision process as follows: Observation space Action space External reward function The external reward function The calculation is as follows: ; In the formula, , , , These are the weighting coefficients for the corresponding reward and penalty items, respectively; Rewards are given for voltage frequency stability; To provide consistent incentives for achieving macroeconomic recovery goals; Penalty for the cost of the action; This is a phased incentive designed to target specific important actions.
7. The power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning according to claim 6, characterized in that: The lower-layer heterogeneous multi-agent group uses a multi-agent soft actor-critic algorithm for control decisions. The centralized critic network in the multi-agent soft actor-critic algorithm is updated by minimizing the squared Bellman error, and its target value is represented as the expected value on the probability distribution of the full action space at the next time step. ; In the formula, Evaluation network for target; Discount factor; This is the observation state for the next moment; For any action in the action space at the next moment; The action probability distribution output by the joint policy of the agents; Temperature coefficient; The distributed actor network in the multi-agent soft actor-critic algorithm is optimized by minimizing the KL divergence between the local policies of each agent and their corresponding Q-values. Its loss function expression is: ; In the formula, s represents the current state; For the action of agent i; Let i be the action space of agent i; Select a probability vector for the action of agent i; The heterogeneous multi-agent group includes a switching agent, a network-building agent, and a network-following agent; The switch agent uses a graph convolutional network (GCN) as its backbone encoding structure, and its feature propagation process satisfies: ; In the formula, This is the feature matrix of the nodes in the (l+1)th layer; It is a non-linear activation function; This is the local power grid adjacency matrix; W is the degree matrix; (l) This is the weight matrix. The network-building agent and the network-following agent adopt a multilayer perceptron (MLP) as the backbone coding structure.
8. The power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning according to claim 6, characterized in that: S4 specifically includes, Abstract each active power supply island in the distribution network into a node and construct an island relationship diagram: ; In the formula, V is the set of nodes; E is the set of edges; Using the attention mechanism in graph attention networks, calculate the normalized attention coefficients between each isolated node and its neighboring nodes: ; In the formula; This is the normalized attention coefficient; is the shared feature linear transformation matrix; 'a' is the attention weight vector; , , These are the input feature vectors for isolated nodes i, j, and k, respectively. This is a vector concatenation operation; LeakyReLU is a non-linear activation function; N i This is the set of neighboring islands that can be physically connected to node i. The neighbor features are weighted and aggregated based on the normalized attention coefficients to produce graph embedding vectors that contain topological context information; The optimal merging timing and target are determined from multiple candidate merging objects based on graph embedding vectors.
9. The power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning according to claim 1, characterized in that: The dual-action masking mechanism is as follows: Based on the current macroscopic instructions issued by the upper-level single agent, a task phase isolation mask is generated to shield device actions that are unrelated to the goal of that phase. By combining the real-time power grid topology and equipment status, a physical security shielding mask is generated to forcibly shield dangerous actions; By performing a logical AND operation between the original action space and the isolation mask and physical security mask, the dynamically reduced effective action space is obtained: ; In the formula, For effective action space; For isolation mask; For physical security shielding; o T These are macro-level instructions.
10. The power distribution system fault recovery method based on hierarchical multi-agent deep reinforcement learning according to claim 9, characterized in that: The advanced course learning and training paradigm is implemented by constructing multiple sub-task courses with increasing difficulty, as follows: ; In the formula, , , The courses are divided into initial, intermediate, and advanced levels; skill progression is achieved by unlocking tasks in each subsequent recovery phase. Initial Course First, enable black start and photovoltaic-storage collaborative recovery tasks, limit the action space, and train the model to master the basic skills of voltage and frequency support and core load recovery; Intermediate Course :exist Building upon the existing tasks, further unlock the island expansion phase tasks and train the model to master the ability to reconfigure the grid and expand the power supply range using line switches while maintaining system stability. Advanced courses :exist Based on the mission, the full-stage recovery mission, including island merging, is fully opened up, and the training model is trained to master the closed-loop control capability of multi-source collaboration and global optimal merging decision-making in multi-island systems. When the load recovery rate or the cumulative reward obtained by the agent in performing a complete recovery task in the current course environment reaches the preset stability index, the course will be smoothly switched to the next sub-course with increasing difficulty through a smooth switching mechanism.