Active power distribution network toughness improving method based on multi-agent deep reinforcement learning

Through a method based on deep reinforcement learning, a problem model for mobile energy storage pre-distribution and microgrid group dynamic division is constructed, and a multi-agent deep Q network algorithm (MADQN) is used to solve the problem, which solves the problem that the power system is difficult to recover key loads in extreme disasters, and achieves the resilience of the distribution network.

CN120046648APending Publication Date: 2025-05-27HOHAI UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510032984.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing power system is difficult to effectively restore critical loads in extreme disasters, resulting in insufficient resilience of the distribution network.

Method used

Using a method based on deep reinforcement learning, a problem model for mobile energy storage pre-allocation and microgrid group dynamic division is constructed, and a multi-agent deep Q network algorithm (MADQN) is used to solve problems to achieve post-disaster recovery and resilience improvement.

Benefits of technology

Through rapid decision-making and effective recovery of critical loads, the distribution network's resilience and resilience in extreme disasters has been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046648A_ABST
    Figure CN120046648A_ABST
Patent Text Reader

Abstract

The invention relates to the field of power system optimization scheduling, in particular to an active power distribution network toughness improving method based on multi-agent deep reinforcement learning. The method comprises the following steps: (1) constructing a toughness improvement problem model: taking a distribution network internal key load recovery amount and a mobile energy storage pre-distribution cost as target functions, wherein the model follows a power distribution system safety operation constraint condition; (2) constructing a Markov decision process: defining Markov decision variables, and establishing a state space, an action space and a corresponding reward function; the strategy solving link is allocated to a plurality of agents to improve the training performance; (3) constructing an intelligent agent-environment interaction interface, and establishing a power flow analysis and topology check module to realize communication between the environment and the intelligent agent; and (4) model training and online application: combining an improved CTCE-MADQN algorithm with a strategy solving process, training the model to obtain a mobile energy storage configuration and microgrid group optimal division strategy, and improving the toughness of the power distribution network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power system optimal dispatch, and specifically to a method for improving the resilience of active distribution networks based on deep reinforcement learning. Background Art

[0002] In order to effectively improve the ability of power systems to resist extreme disasters, a method for improving the resilience of active distribution networks based on deep reinforcement learning is proposed. First, with the key load recovery amount of the distribution network as the objective function, comprehensively considering the safe operation constraints of the distribution network and the microgrid cluster, a problem model for the pre-allocation of mobile energy storage and the dynamic partitioning of the microgrid cluster is constructed, and an agent-environment interaction interface and a distribution system simulation environment are established based on this model. Then, the problem model for the pre-allocation of mobile energy storage and the dynamic partitioning of the microgrid cluster is constructed as a Markov decision process, and the improved multi-agent deep Q-network algorithm (MADQN) is used to solve the problem, and the average episode return is defined as the training performance index to improve the convergence and decision-making performance of the algorithm. The results of the numerical example show that the pre-configuration of mobile energy storage and the dynamic partitioning strategy of the microgrid given by this method can make quick decisions, effectively recover key loads, and improve the resilience of the distribution network. Summary of the Invention

[0003] Object of the Invention: The object of the present invention is to provide a method for improving the resilience of active distribution networks based on deep reinforcement learning. An active distribution network post-disaster recovery model under extreme disasters is established. By pre-allocating mobile energy storage before the disaster and forming a microgrid cluster after the disaster, the power supply of key loads is restored, and this decision-making process is solved using a multi-agent deep reinforcement learning algorithm, improving the strategy solution speed, effectively recovering key loads, and enhancing the ability of the distribution network to resist extreme disasters.

[0004] Technical Solution: The specific steps of the present invention are as follows:

[0005] S1: Construct a problem model for the pre-allocation of mobile energy storage and the dynamic partitioning of the microgrid cluster: With the key load recovery amount within the distribution network and the pre-allocation cost of mobile energy storage as the objective function, the model follows six types of constraints, namely, distribution system power flow constraints, node voltage constraints, line current constraints, distribution network topology constraints, mobile energy storage allocation constraints, and generator capacity constraints.

[0006] S2: Construct a Markov decision process: Define Markov decision variables, establish the state space, action space of the problem model for the pre-allocation of mobile energy storage and the dynamic partitioning of the microgrid cluster, and establish the corresponding reward function. Assign the strategy solution link to multiple agents to improve the training performance.

[0007] S3: Construct the agent-environment interaction interface: It is divided into a power flow analysis and topology check module. Establish a power flow model for the distribution system, where distributed power sources are modeled as constant voltage sources, and their outputs will be part of the power flow results. In addition, establish and save the topology map of the distribution network for topology analysis. Then, after receiving the actions from the agent, check the radial topology constraints by searching for the paths between load nodes and DGs based on depth-first search. For control operations that do not violate the topology constraints, continue to execute the actions and perform power flow calculations. When the verification of all constraint conditions and the power flow calculations are completed, obtain the current state of the system at the next stage and calculate the reward to feedback to the agent.

[0008] S4: Model training and online application: Combine the CTCE-MADQN algorithm with the policy solving process to train the model to obtain the optimal microgrid cluster division strategy for improving the resilience of the distribution network. And save the agent model obtained in the training session and use it in the application session to improve the decision-making speed during application.

[0009] In step S1, with the goal of maximizing the weighted recovery of loads, combined with the operating constraints of the distribution network and microgrids, construct a dynamic microgrid cluster division problem model for improving the resilience of the distribution network. The distribution network is represented by a directed graph P=(N, E, G, S). N and E respectively represent the sets of all nodes and edges (transmission lines, etc.) in the distribution network, distributed power sources are described by the set G, , representing the distribution network nodes where all distributed power sources are connected. The set of remote control switches used to perform the microgrid division operation is represented by the set S. The specific model of the above problem can be expressed as:

[0010] (1) Objective function

[0011] At each moment t, the goal is to maximize the recovery of critical loads. Its objective function can be expressed as:

[0012]

[0013] In the formula: E s [·] is the mathematical expectation function, ω i is the load weight coefficient, used to represent the importance of the load. p t,i =a t,i p t,i , where a t,i is used to represent whether the load is recovered. For recovered loads, its value is 1, otherwise 0, α j is the unit capacity energy storage configuration cost, c jCapacity configuration for energy storage. Considering the output limitations of distributed power sources within the distribution network and the high generation costs, it is impossible to use distributed power sources to restore all loads within the distribution network. Therefore, critical loads have a higher restoration priority than non-critical loads.

[0014] (2) Constraint conditions

[0015] To ensure the safe operation of the distribution network during the implementation of microgrid group division and load restoration, six types of constraints need to be followed, namely distribution system power flow constraint, node voltage constraint, line current constraint, distribution network topology constraint, mobile energy storage allocation constraint, and generator capacity constraint.

[0016] 1) Distribution system power flow constraint. The normal operation of the distribution system needs to satisfy the power flow constraint, and the specific expression is as follows:

[0017]

[0018] 2) Node voltage constraint. When performing dynamic microgrid division in the distribution network, it is necessary to ensure that the node voltage remains within a reasonable range:

[0019] V t,i,min ≤V t,i ≤V t,i,max , i ∈ N, t ∈ T

[0020] In the formula: V t,i,min is the lower limit of the voltage amplitude of node i at time t, V t,i,max is the upper limit of the voltage amplitude of node i at time t, V t,i is the voltage amplitude of node i at time t.

[0021] 3) Line current constraint. To prevent line overload, the line current in the distribution network system should be within its limit:

[0022] I t,n,min ≤I t,n ≤I t,n,max , n ∈ E, t ∈ T

[0023] In the formula: I t,n,min is the lower bound of the current of line n at time t, I t,n,max is the upper bound of the current of line n at time t, I t,n is the current of line n at time t.

[0024] 4) Topological constraints. During the post-disaster recovery stage, through dynamic division of the microgrid group, the internal microgrid can utilize its own power generation resources to restore critical loads. For the load restoration problem, a single-source - single-microgrid control strategy is adopted. That is, only one distributed power supply supplies power within one microgrid, and the critical loads within the microgrid are powered by one distributed power supply via one path. Each microgrid operates independently, and there is no connection between any two microgrids. Considering that when the distribution network is decomposed into multiple island microgrids, it is necessary to ensure that the network always maintains a radial structure. For any originally connected distribution network, the topological constraints can be expressed as:

[0025]

[0026] where: l i,j is a binary variable indicating whether nodes i and j are connected. When nodes i and j are connected, l i,j = 1; when nodes i and j are not connected, l i,j = 0. For the radial distribution network structure, the topological constraints can be simplified to:

[0027]

[0028] 5) Mobile energy storage allocation constraints. The mobile energy storage resources are limited, and the number of mobile energy storage units arranged in the pre-allocation stage should not exceed the upper limit of the energy storage allocation quantity:

[0029]

[0030] where: λ i is a binary variable indicating whether node i is connected to the mobile energy storage, and N ME,max is the upper limit of the number of mobile energy storage units arranged in the pre-allocation stage.

[0031] 6) Distributed power supply capacity constraints. As an emergency response resource, during the entire recovery period, the output of the distributed power supply is limited by its capacity, and its output should be within the limit range:

[0032] P t,i,min ≤ P t,i ≤ P t,i,max t ∈ T, i ∈ G

[0033] Q t,i,min ≤ Q t,i ≤ Q t,i,max t ∈ T, i ∈ G

[0034] where: P t,i,min and Q t,i,min are respectively the lower limits of the active power and reactive power outputs of power supply i at time t, P t,i,max and Q t,i,maxThey are the upper limits of the active power and reactive power output of the power supply i at time t, respectively.

[0035] In step S2, the state space and action space of the mobile energy storage pre-allocation and microgrid group dynamic partitioning problem are established, and the corresponding reward function is set to train the intelligent agent. By defining the Markov decision variables, a complete Markov decision process is established. The learning and decision-making capabilities of the intelligent agent largely come from the design of the Markov decision process. The decision variables are refined to improve the learning ability and learning efficiency of the intelligent agent. The above-mentioned Markov decision variables are as follows:

[0036] 1) Action space. The action space is the decision variable of the intelligent agent, including the energy storage allocation location and the execution status action space of the remote control switch. It can be expressed as:

[0037] A = {α i , β t,s ∣t ∈ T, s ∈ S, i ∈ N}

[0038] In the formula: α i represents the location of the mobile energy storage pre-allocated before the disaster, and β t,s is a binary variable representing the operation of the remote control switch. a t,s = 1 indicates that the switch s performs a disconnection operation at time t, and vice versa for closing. Therefore, the action space A is discrete and consists of a finite number of binary quantities.

[0039] 2) State space. The state space is divided into two categories. One is the load data, that is, the load data restored after the decision is executed; the other is the distributed power source data, that is, the output limit and actual output data of the distributed power source:

[0040] S = {p t,i , q t,i , P t,k,min , P t,k,max P t,k , Q t,k,min , Q t,k,max , Q t,k , U t,i , I t,i , ∣t ∈ T, i ∈ N, k ∈ M}

[0041] According to the definition of the state space S, its internal parameters contain continuous variables (continuous quantities such as load, power of the power source, voltage, current, etc.), so the state space S is continuous.

[0042] 3) Reward function. Classify the constraints according to the severity of the penalties generated by violating the constraints. Strong constraints include power flow constraints of the distribution system, line current constraints, topology constraints, and generator capacity constraints, while weak constraints include node voltage constraints. The reward factors for the decisions that violate the corresponding constraints and are fed back to the agent are set as follows:

[0043]

[0044] In the formula: α 1,t represents the load restored at time t, and α 2,t is the penalty coefficient generated by violating the voltage constraint, and its value is related to the voltage over-limit value. α 3,t is the energy storage allocation cost in the pre-layout stage. k 0 has a value much larger than k 1 and k 2 , indicating that actions that violate strong constraints will be severely punished. If the decision made by the agent causes the node voltage amplitude to exceed the range of ±10% p.u., this decision will be considered a restoration failure, and a reward of -k 0 will be given instead of being recorded as an action that violates weak constraints.

[0045] In step S3, construct an agent-environment interaction interface (AEI). The environment needs to receive actions from the agent and then feedback the system state and rewards. To achieve this, it is necessary to ensure that all parameters related to actions, states, and rewards are available, and all constraint violation situations can be detected in the environment. The construction process of AEI is as follows:

[0046] 1) Build a distribution network system in Python, initialize the time as t, and the state as s t .

[0047] 2) Receive an action a t from the agent, and decompose this action set into an energy storage allocation action c t and a line action l t .

[0048] 3) According to the line action l t , perform a topology analysis on the distribution network structure to determine whether the action a violates the topology constraint. If it violates the topology constraint, set s t+1 as the termination state, r t = -k 0 , and jump to step 2), otherwise go to step 4).

[0049] 4) Execute the energy storage action c t and run a power flow analysis. If it violates the power flow constraint, set s t+1 as the termination state, r t = -k0 And jump to step 2), otherwise go to step 5).

[0050] 5) In the distribution system simulation environment, perform the load restoration operation and calculate α 1,t α 2,t α 3,t , and generate the final state s t+1 ={p t,i ,q t,i ,P t,k,min ,P t,k,max P t,k ,Q t,k,min ,Q t,k,max ,Q t,k ,U t ,I t}, calculate the final reward r t =k 1 α 1,t -k 2 α 2,t -k 3 α 3,t .

[0051] 6) Feed s t+1 and r t back to the agent.

[0052] In step S4, in order to improve the rapidity and accuracy of formulating the distribution system resilience improvement strategy, and considering that the problem model has a continuous state space and a discrete action space, and the action space is relatively large and has independent characteristics, an improved MADQN is selected as the deep reinforcement learning algorithm for policy solution. Compared with the traditional deep reinforcement learning algorithm, by introducing experience replay and target Q network, and using multiple agents for centralized training and centralized decision-making (CTCEModel) to improve the decision-making and convergence of the algorithm.

[0053] The centralized training and centralized execution steps of CTCE-MADQN are as follows: In the training stage, the states, actions, and rewards of all agents are collected and used to train the global Q network. This centralized method allows agents to utilize each other's information, thus accelerating the learning process and improving the policy quality. In the execution stage, the agents still share the global information and generate actions according to the unified decision network. This can ensure the consistency of the execution policy and avoid conflicts between individuals. The specific principle of the improved CTCE-MADQN algorithm used in this patent is as follows:

[0054] 1) Joint state representation: Use the joint state to better reflect the overall information of the system.

[0055] The local states of all agents are combined to form the global state S t。Assume there are N agents in the system, and the local state of each agent is s i , then S t = [s 1 , s 2 ,..., s N

[0056] 2) Joint action selection:

[0057] The action combinations of the agents form the global action A t , that is, A t = [a 1 , a 2 ,..., a n , where a i is the local action of the i-th agent.

[0058] 3) Joint reward design:

[0059] The system provides a joint reward R t according to the actions A t of all agents and the global state S t to encourage cooperation. The reward signal can reflect the global goal, such as the load recovery rate, the stability of the distribution system, etc.

[0060] 4) Centralized training of the Q-network:

[0061] During the training process, a global deep Q-network is used, with the joint state S t and the joint action A t as inputs, and the Q-value of the joint action is output: Q(S t , A t ). The Q-network is optimized by minimizing the temporal difference (TD) error:

[0062]

[0063] where the target value y is expressed as:

[0064] y = R t + γmax A′ Q(S t+1 , A′; θ′)

[0065] where γ is the discount factor and θ′ is the parameter of the target network.

[0066] 5) Centralized decision-making in the execution phase

[0067] In the decision-making phase, all agents share the same Q-network and determine the optimal policy by calculating the maximum Q-value of the joint action:

[0068] ​

[0069] Then decode the joint action and distribute it to each agent.

[0070] 6) Update the parameters of the Deep Q - Network:

[0071] For each Q - Network, define the optimal action - value function:

[0072] Q * (s,a) = max π E τ~π [G τ ∣s t = s,a t = a]

[0073] Assume that the state at the next moment t + 1 is s', all possible actions that each agent can take are a' and their value Q * (s′,a′) is known, then the optimal policy is to choose the action a' that maximizes the expected action value r+γQ * (s′,a′):

[0074] Q * (s,a) = E s′ [r+γmax a′ Q * (s′,a′)∣s t = s,a t = a]

[0075] The action - value function can be estimated by using the Bellman equation as an iterative update:

[0076] Q i+1 (s,a) = E s′ [r+γmax a′ Q i (s′,a′)∣s t = s,a t = a]

[0077] The above value iteration algorithm converges to the optimal action - value function. As i→∞, Q i →Q * . When the scale of the state - action space is large, it is difficult for general function approximation methods to estimate each state - action value. Therefore, the MADQN algorithm uses a neural network function (Deep Q - Network) with weights θ as a function approximator, denoted as Q(s,a;θ)≈Q * (s,a), to estimate the action - value function. The Q - Network takes the state as input and the action - value as output.

[0078] During the training process, the training parameter θ can be gradually adjusted i, to reduce the mean square error in the Bellman equation. Among them, the optimal target value is r + γmax a′ Q * (s′, a′) uses the approximate target value y = r + γmax a′ Q(s′, a′; θi) instead, and uses the parameters θ of the previous iteration i . At each iteration, the loss function L(θ) of each agent is expressed as:

[0079] L(θ i ) = E[(y - Q(s, a; θ i )) 2 )

[0080] By taking the partial derivative of the loss function with respect to the weights, the gradient of the loss function with respect to the weights θ i is obtained:

[0081] ▽ θi L(θ i ) = E[(y - Q(s, a; θ))▽ θi Q(s, a; θ i )]

[0082] For the loss function L(θ i ), use mini-batch stochastic gradient descent to optimize the loss function and update the weights, and perform iterative updates on θ i .

[0083] The specific training process is as follows:

[0084] (1) Initialize the distribution network: parameters such as load, line, and topology for simulating the power system.

[0085] (2) Initialize the experience pool, Q-network, and target Q-network of each agent. Initialize the environmental state s t , and the time step t 0 .

[0086] (3) Use the greedy algorithm to select the action a t from the Q-network.

[0087] (4) Interact with the distribution system simulation environment through the AEI interface to perform corresponding topology analysis and power flow analysis, generate a new state s t+1 , calculate the reward r and other relevant information, and store the experience (s t , a, r, s t+1 ) in the experience pool.

[0088] (5) Sample a mini-batch of data samples from the experience pool.

[0089] The approximate target action value y tIf action a violates the strong constraint, then y t = -c 0 ; otherwise, y t = r + γmax a′ Q(s′, a′; θ i ).

[0090] (6) Use the stochastic gradient descent algorithm to minimize the loss function and update the Q-network parameters θ. After every C updates, replace the target Q-network with the Q-network.

[0091] (7) If the maximum training cycle is reached, end the training and save the deep Q-network; otherwise, go to (3).

[0092] The agent interacts with the environment through the AEI interface to obtain state-value information, and uses the approximate target action value y t and uses the stochastic gradient descent algorithm to update and iterate the Q-network. Among them, the greedy algorithm is used to limit the free exploration and decision-making of the agent:

[0093] The greedy algorithm means that when the agent selects action a t , there is a probability of ε to select a random action, and a probability of 1 - ε to select the action with the maximum value. The greedy algorithm used in this patent is as follows:

[0094]

[0095] In the formula: ε 0 is the initial exploration rate, k is the exploration rate decay factor, and its value is related to the sizes of the state and action spaces. t i and t max represent the current step number and the maximum training step number respectively. Adjust the free exploration rate of the agent in the training stage through exponential decay. Give the agent a high degree of freedom to explore at the initial stage of training to fully explore potential decision-making situations, and gradually reduce the exploration rate in the later stage of training to drive the agent to select the optimal decision and improve the convergence of the training process.

[0096] The calculation formula of MER is as follows:

[0097]

[0098] In the formula: n represents the current number of episodes, T i represents the length of the i-th episode, and n.agents represents the number of agents. After the training is completed, define the mean episode reward (MER) metric to evaluate the training performance. MER represents the average of the episode rewards from the start of training to the current step, and it can effectively reflect the convergence performance of the training process.

[0099] The beneficial effects of the present invention are as follows:

[0100] Based on the deep reinforcement learning algorithm and aiming to improve the resilience of the distribution system, combined with the constraints such as the safe operation of the distribution network and microgrid, a method for dividing a microgrid group for improving the resilience of the distribution network based on MA-DRL is proposed, which can restore some critical loads for the distribution system when the power supply capacity of the main grid is insufficient due to extreme events and improve the resilience of the distribution network.

[0101] The CTCE-MADQN algorithm has good convergence performance. The agent can quickly learn the load restoration strategy without additional prior knowledge, effectively improving the rapidity and accuracy of decision-making. Compared with the traditional MILP algorithm, the method based on MA-DRL reconstructs the distribution network into multiple microgrids, realizes the refined division of the microgrid group, restores the power supply of critical loads, and effectively improves the resilience of the distribution system. Using a model-free algorithm can avoid the modeling of complex problems, effectively improve the solving speed of the algorithm, and has obvious advantages in terms of computational efficiency in the application link.

[0102] Explanation of the attached drawings

[0103] Figure 1 It is the schematic diagram of establishing the agent-environment interface

[0104] Figure 2 It is the training flowchart of CTCE-MADQN Specific implementation manners

[0105] To make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0106] A method for improving the resilience of an active distribution network based on deep reinforcement learning includes the following steps:

[0107] S1: Construct a problem model for pre-distribution of mobile energy storage and dynamic division of microgrid groups: taking the recovery amount of critical loads inside the distribution network and the pre-distribution cost of mobile energy storage as the objective function, and the model follows six types of constraints, namely, distribution system power flow constraint, node voltage constraint, line current constraint, distribution network topology constraint, mobile energy storage distribution constraint, and generator capacity constraint.

[0108] S2: Construct a Markov decision process: Define Markov decision variables, establish the state space and action space of the problem model for pre-allocation of mobile energy storage and dynamic partitioning of microgrid clusters, and establish the corresponding reward function. Assign the policy solution process to multiple agents to improve training performance.

[0109] S3: Construct an agent-environment interaction interface: It is divided into a power flow analysis and topology check module. Establish a power flow model of the distribution system, where distributed power sources are modeled as constant voltage sources, and their outputs will be part of the power flow results. In addition, establish and save the topology diagram of the distribution network for topology analysis. Then, after receiving the action from the agent, check the radial topology constraint by searching for the path between the load node and the DG based on depth-first search. For control operations that do not violate the topology constraint, continue to execute the action and perform power flow calculation. When the verification of all constraint conditions and the power flow calculation are completed, obtain the current state of the next stage of the system and calculate the reward to feedback to the agent.

[0110] S4: Model training and online application: Combine the deep Q-network algorithm with the policy solution process, train the model to obtain the optimal partitioning strategy for microgrid clusters, and improve the resilience of the distribution network. Save the agent model obtained in the training process and use it in the application process to improve the decision-making speed during application.

[0111] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for improving the resilience of active distribution networks based on multi-agent deep reinforcement learning, characterized in that: The following steps are involved: S1: Construct a model for the problem of mobile energy storage pre-allocation and dynamic division of microgrid groups: Taking the key load recovery amount within the distribution network and the mobile energy storage pre-allocation cost as the objective function, the model follows six types of constraints, including distribution system flow constraints, node voltage constraints, line current constraints, distribution network topology constraints, mobile energy storage allocation constraints, and generator capacity constraints. S2: Construct Markov decision process: define Markov decision variables, establish the state space and action space of the mobile energy storage pre-allocation and microgrid group dynamic partition problem model, and establish the corresponding reward function. Assign the strategy solving link to multiple agents to improve training performance. S3: Construct the agent-environment interaction interface: divided into power flow analysis and topology check modules. A power flow model of the distribution system is established, in which the distributed power source is modeled as a constant voltage source, and its output will be used as part of the power flow result. In addition, a topological diagram of the distribution network is established and saved for topological analysis. Then, after receiving the action from the agent, the radial topological constraints are checked by searching the path between the load node and the DG based on a depth-first search. For control operations that do not violate the topological constraints, the action continues to be executed and the power flow calculation is performed. When all constraints are checked and the power flow calculation is completed, the current state of the system in the next stage is obtained, and the reward is calculated and fed back to the agent. S4: Model training and online application: Combine the CTCE-MADQN algorithm with the strategy solving process, train the model to obtain the optimal partitioning strategy for the microgrid group, and improve the resilience of the distribution network. Save the intelligent agent model obtained in the training phase and use it in the application phase to improve the decision-making speed during application.

2. The method for improving the resilience of an active distribution network based on deep reinforcement learning according to claim 1, characterized in that: In step S1, the objective function is to maximize the weighted recovery of the load, and the distribution network and microgrid operation constraints are combined to construct a dynamic partition problem model for the microgrid group for improving the resilience of the distribution network. The distribution network is represented by a directed graph P = (N, E, G, S). N and E represent the set of all nodes and edges (transmission lines, etc.) in the distribution network, respectively. The distributed power source is described by the set G. Represents the distribution network nodes to which all distributed generation sources are connected. The set of remote control switches used to perform microgrid partitioning operations is represented by set S. The specific model of the above problem can be expressed as: (1) Objective function At each moment t, the goal is to restore the critical load to the maximum extent. The objective function can be expressed as: Where: E s [·] is the mathematical expectation function, ω i is the load weight coefficient, which is used to indicate the importance of the load. where a t,i It is used to indicate whether the load has been restored. If the load has been restored, its value is 1, otherwise it is 0. j is the energy storage configuration cost per unit capacity, c j Allocate capacity for energy storage. Considering the output limitations of distributed power sources within the distribution network and the high cost of power generation, it is impossible to use distributed power sources to restore all loads within the distribution network. Therefore, critical loads have a higher restoration priority than non-critical loads. (2) Constraints In order to ensure the safe operation of the distribution network during the execution of microgrid group division and load recovery, six types of constraints need to be followed, namely, distribution system flow constraints, node voltage constraints, line current constraints, distribution network topology constraints, mobile energy storage allocation constraints, and generator capacity constraints. 1) Power flow constraints of the distribution system. The normal operation of the distribution system needs to meet the power flow constraints. The specific expression is as follows: 2) Node voltage constraints. When the distribution network is dynamically divided into microgrids, it is necessary to ensure that the node voltage remains within a reasonable range: In t,i,min ≤V t,i ≤V t,i,max ,i∈N,t∈T Where: V t,i,min is the lower limit of the voltage amplitude of node i at time t, V t,i,max is the upper limit of the voltage amplitude of node i at time t, V t,i is the voltage amplitude of node i at time t. 3) Line current constraints. To prevent line overload, the line current in the distribution network system should be within its limit: I t,n,min ≤I t,n ≤I t,n,max ,n∈E,t∈T Where: I t,n,min is the lower limit of the current of line n at time t, I t,n,max is the upper limit of the current of line n at time t, I t,n is the current in line n at time t. 4) Topological constraints. In the post-disaster recovery phase, through dynamic microgrid group division, the microgrid can use its own power generation resources to restore critical loads. For the load recovery problem, a single-source-single-microgrid control strategy is adopted. That is, a microgrid is powered by only one distributed power source, and the critical loads in the microgrid are powered by one distributed power source through one path. Each microgrid operates independently, and there is no path between any two microgrids. Considering that when the distribution network is decomposed into multiple isolated microgrids, it is necessary to ensure that the network always maintains a radial structure. For any originally connected distribution network, the topological constraint can be expressed as: Where: l i,j is a binary quantity indicating whether nodes i and j are connected. When nodes i and j are connected, l i,j =1, when nodes i and j are not connected, l i,j = 0. For the radial distribution network structure, the topological constraint can be simplified as: 5) Mobile energy storage allocation constraints. Mobile energy storage resources are limited, and the number of mobile energy storage deployed in the pre-allocation stage should not exceed the upper limit of energy storage allocation: Where: i is a binary quantity indicating that node i is connected to mobile energy storage, N ME,max The upper limit of the number of mobile energy storage that can be deployed in the pre-allocation stage. 6) Distributed power capacity constraints. As an emergency response resource, the output of distributed power is limited by its capacity during the entire recovery period, and its output should be within the limit: P t,i,min ≤P t,i ≤P t,i,max t∈T,i∈G Q t,i,min ≤Q t,i ≤Q t,i,max t∈T,i∈G Where: P t,i,min and Q t,i,min are the lower limits of active power and reactive power output of power source i at time t, P t,i,max and Q t,i,max They are the upper limits of the active power and reactive power output of power source i at time t respectively.

3. The method for improving the resilience of an active distribution network based on deep reinforcement learning according to claim 1, characterized in that: In step S2, the state space and action space of the mobile energy storage pre-allocation and microgrid group dynamic partitioning problem are established, and the corresponding reward function is set to train the intelligent agent. By defining the Markov decision variables, a complete Markov decision process is established. The learning and decision-making ability of the intelligent agent is largely derived from the design of the Markov decision process. The decision variables are refined to improve the learning ability and learning efficiency of the intelligent agent. The above-mentioned Markov decision variables are as follows: 1) Action space. The action space is the agent’s decision variable, including the energy storage allocation location and the remote control switch execution state. The action space can be expressed as: A={a i ,b t,s ∣t∈T,s∈S,i∈N} Where: α i represents the pre-disaster pre-allocated mobile energy storage location, β t,s is a binary variable representing the remote control switch operation, a t,s =1 means that the switch s performs an opening operation at time t, otherwise it performs a closing operation. Therefore, the action space A is discrete, which consists of a finite number of binary quantities. 2) State space. The state space is divided into two categories. One is load data, that is, the load data restored after the decision is executed; the other is distributed power data, that is, the output limit and actual output data of distributed power: S={p t,i ,q t,i ,P t,k,min ,P t,k,max P t,k ,Q t,k,min ,Q t,k,max ,Q t,k ,U t,i ,I t,i ,∣t∈T,i∈N,k∈M} According to the definition of the state space S, its internal parameters contain continuous variables (load, power supply power, voltage, current and other continuous quantities) because the state space S is continuous. 3) Reward function. Constraints are classified according to the severity of the penalty for violating them. Strong constraints include power flow constraints, line current constraints, topology constraints, and generator capacity constraints, while weak constraints include node voltage constraints. The reward factor for the decision that violates the corresponding constraint is set as follows: Where: α 1,t represents the load restored at time t, α 2,t is the penalty coefficient for violating the voltage constraint, and its value is related to the voltage limit. 3,t The cost of energy storage allocation in the pre-layout phase. The value of k0 is much larger than k1 and k2, indicating that actions that violate strong constraints will be punished extremely highly. If the decision made by the agent causes the node voltage amplitude to exceed the range of ±10% pu, the decision will be considered a recovery failure and a reward of -k0 will be given instead of being recorded as an action that violates weak constraints.

4. According to claim 1, a method for improving the resilience of an active distribution network based on multi-agent deep reinforcement learning is characterized in that: In step S3, the agent-environment interaction interface (AEI) is constructed. The environment needs to accept actions from the agent and then feedback system status and rewards. To achieve this, it is necessary to ensure that all parameters related to actions, status, and rewards are available and that all constraint violations can be detected in the environment. The construction process of AEI is as follows: 1) Build a distribution network system in Python, initialize the time to t and the state to s t . 2) Receive an action a from the agent t , decompose the action set into energy storage allocation action c t With line action t . 3)According to the line action l t , perform topological analysis on the distribution network structure and determine whether action a violates the topological constraint. If it violates the topological constraint, let s t+1 is the terminal state, r t =-k0, and jump to step 2), otherwise jump to step 4). 4) Perform energy storage action c t Run the power flow analysis and if the power flow constraint is violated, set s t+1 is the terminal state, r t =-k0 and jump to step 2), otherwise jump to step 5). 5) In the power distribution system simulation environment, perform load recovery operations and calculate α 1,t α 2,t α 3,t , and generate the final state s t+1 ={p t,i ,q t,i ,P t,k,min ,P t,k,max P t,k ,Q t,k,min ,Q t,k,max ,Q t,k ,U t,i ,I t,i }, and calculate the final reward r t =k1α 1,t -k2α 2,t -k3α 3,t . 6) t+1 With r t Feedback to the agent.

5. The method for improving the resilience of active distribution network based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: In step S4, in order to improve the speed and accuracy of the distribution system resilience improvement strategy, and considering that the problem model has a continuous state space and a discrete action space, and the action space is relatively large and has its own independent characteristics, the improved MADQN is selected as the deep reinforcement learning algorithm for strategy solving. Compared with the traditional deep reinforcement learning algorithm, the algorithm improves its decision-making and convergence by introducing experience replay and target Q network, and using multiple agents for centralized training and centralized decision-making (CTCE-Model). The centralized training of CTCE-MADQN is performed as follows: During the training phase, the states, actions, and rewards of all agents are collected and used to train the global Q network. This centralized approach allows agents to use each other's information, thereby accelerating the learning process and improving the quality of the strategy. During the execution phase, agents still share global information and generate actions based on a unified decision network. This ensures consistency in the execution strategy and avoids conflicts between individuals. The specific principles of the improved CTCE-MADQN algorithm used in this patent are as follows: 1) Joint state representation: Use joint state to better reflect the overall information of the system. The local states of all agents are combined to form the global state S t Assume that there are N agents in the system and the local state of each agent is s i , then S t =[s1,s2,...,s N ] 2) Joint action selection: The actions of the agents are combined to form a global action A t , that is, A t =[a1,a2,...,a n ], where a i is the local action of the ith agent. 3) Joint reward design: The system is based on the actions A of all agents t and the global state S t Provide a joint reward R t , to encourage collaboration. The reward signal can reflect global goals, such as load recovery rate, distribution system stability, etc. 4) Q network centralized training: During the training process, a global deep Q network is used, with the input joint state S t and joint action A t , output the Q value of the joint action: Q(S t ,A t ), the Q network is optimized by minimizing the temporal difference (TD) error: L=E[(y-Q(S t ,A t ;θ)) 2 ] The target value y is expressed as: y=R t +γmax A′ Q(S t+1 ,A′;θ′) where γ is the discount factor and θ′ is the parameter of the target network. 5) Centralized decision-making during the execution phase In the decision-making stage, all agents share the same Q network and determine the optimal strategy by calculating the maximum Q value of the joint action: Then decode the joint action Assigned to each agent. 6) Deep Q network parameter update: For each Q network, define the optimal action-value function: Q * (s,a)=max π E τ~π [G τ ∣s t =s,a t =a] Assume that the state at the next moment t+1 is s', and all possible actions that each agent may take are a' and their value is Q * (s′, a′) is known, then the optimal strategy is to choose an action that makes the expected value r+γQ * (s′,a′) maximizes the action a': Q * (s,a)=E s′ [r+γmax a′ Q * (s′,a′)∣s t =s,a t =a] The action-value function can be estimated by using the Bellman equation as an iterative update: Q i+1 (s,a)=E s′ [r+γmax a′ Q i (s′,a′)∣s t =s,a t =a] The above value iteration algorithm converges to the optimal action-value function. As i→∞, Q i →Q * When the state-action space is large, it is difficult for general function approximation methods to estimate the value of each state-action. Therefore, the MADQN algorithm uses a neural network function (deep Q network) with weight θ as a function approximator, expressed as Q(s,a;θ)≈Q * (s,a), used to estimate the action-value function. The Q network takes the state as input and the action-value as output. During the training process, the training parameters θ can be gradually adjusted i , to reduce the mean square error in the Bellman equation. The optimal target value is r+γmax a′ Q * (s′, a′) uses the approximate target value y=r+γmax a′ Q(s′,a′;θi) is used instead, using the parameters θ of the previous iteration i At each iteration, the loss function L(θ) of each agent is expressed as: L(θ i )=E[(yQ(s,a;θ i )) 2 ] By taking the partial derivative of the loss function with respect to the weight, we can get the loss function with respect to the weight θ i The gradient is: For the loss function L(θ i ), use mini-batch stochastic gradient descent to optimize the loss function and update the weights, for θ i Perform iterative updates. The specific training process is as follows: (1) Initialize the distribution network: load, line, topology and other parameters to simulate the power system. (2) Initialize each agent's experience pool, Q network, and target Q network. Initialize the environment state s t , time t0. (3) Use the greedy algorithm to select action a from the Q network t . (4) Interact with the power distribution system simulation environment through the AEI interface to perform corresponding topology analysis and power flow analysis to generate new state s t+1 , and calculate the reward r and other related information, store the experience (s t ,a,r,s t+1 ) to the experience pool. (5) Sample small batches of data from the experience pool. Approximate target action value y t If action a violates a strong constraint, then y t = -c0, otherwise y t = r + γmax a′ Q(s′,a′;θ i ). (6) Use the stochastic gradient descent algorithm to minimize the loss function and update the Q network parameters θ. After every C updates, the target Q network is replaced by the Q network. (7) If the maximum training cycle is reached, end the training and save the deep Q network; otherwise go to (3). The agent interacts with the environment through the AEI interface to obtain state-value information and approximates the target action value y t And use the stochastic gradient descent algorithm to update the Q network iteratively. A greedy algorithm is used to restrict the agent's free exploration and decision-making: The greedy algorithm indicates that the agent chooses action a t When , there is a probability of ε to select a random action, and a probability of 1-ε to select the action with the greatest value. The greedy algorithm used in this patent is as follows: Where: ε0 is the initial exploration rate, k is the exploration rate attenuation factor, and its value is related to the state and action space size. i and t max Represent the current step number and the maximum training step number respectively. The free exploration rate of the agent in the training stage is adjusted through exponential decay. The agent is given a high degree of freedom to explore in the initial stage of training to fully explore potential decision-making situations. The exploration rate is gradually reduced in the later stage of training to drive the agent to choose the optimal decision and improve the convergence of the training process. The calculation formula of MER is as follows: Where: n represents the current number of fragments, T i Represents the length of the i-th segment, and n.agents represents the number of agents. After the training is completed, the average segment return index (MER) is defined to evaluate the training performance. MER is represented by the average segment return from the start of training to the current step, which can effectively reflect the convergence performance of the training process. The beneficial effects of the present invention are as follows: Based on the deep reinforcement learning algorithm, this paper aims to improve the resilience of the distribution system. Combined with the constraints such as the safe operation of the distribution network and microgrid, a microgrid group division method based on MA-DRL for improving the resilience of the distribution network is proposed. When extreme events cause insufficient power supply capacity of the main grid, some key loads can be restored for the distribution system, thereby improving the resilience of the distribution network. The CTCE-MADQN algorithm has good convergence performance, and the intelligent agent can quickly learn the load recovery strategy without additional prior knowledge, which effectively improves the speed and accuracy of decision-making. Compared with the traditional MILP algorithm, the MA-DRL-based method reconstructs the distribution network into multiple microgrids, realizes the refined division of microgrid groups, realizes the power supply restoration of key loads, and effectively improves the resilience of the distribution system. The use of model-free algorithms can avoid the modeling of complex problems and effectively improve the speed of algorithm solution. In the application link, this method has obvious advantages in computational efficiency.

Citation Information

Cited By

  • Multi-microgrid multivariate scheduling method, system and device based on virtual energy storage and medium

    CN120675204A

  • Power distribution network reconstruction method and system based on multi-agent safety reinforcement learning

    CN121566598A

  • Distributed energy agent regulation and control method and system based on deep reinforcement learning

    CN122001026A