An active power distribution grid resilience self-healing method and controller considering dynamic microgrids
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2026-08-11
AI Technical Summary
在配电网故障自愈问题上得到广泛应用的数学优化方法有分支定界、动态规划等算法,然而,数学优化算法不适用于处理系统规模大、复杂性高的问题,往往会导致维数灾难问题、计算量大、时间长、时效性不好等问题
[0025]Beneficial Effects: Compared with existing technologies, this invention has the following advantages: First, this invention considers the role of dynamic microgrids in the self-healing of active distribution networks and establishes a reinforcement learning framework for real-time management and control of active distribution networks. This enables the self-healing method to perceive the changing state of the microgrid, more closely approximating the actual situation of the dynamic microgrid, thus improving the self-healing effect. Second, this invention designs a hierarchical reinforcement learning mechanism, modeling the dynamic recovery strategy of distribution network faults according to a hierarchical decision-making mechanism. This more accurately reflects the control process of the dynamic microgrid, thereby improving the self-healing effect. Finally, the reward mechanism of reinforcement learning is improved for dynamic microgrids, cleverly enhancing the reinforcement learning effect and efficiency. This invention uses a hierarchical reinforcement learning model to manage the elastic self-healing process of active distribution networks in real time, generating appropriate and effective intelligent distribution network fault self-healing strategies to reasonably resolve distribution network faults.
Smart Images

Figure CN117791560B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a self-healing technology for distribution network faults, and more particularly to an active distribution network resilient self-healing method and controller that takes into account dynamic microgrids. Background Technology
[0002] Distribution networks incorporating distributed generation (DG) represent the future direction of distribution network development, with self-healing being their core function and most important characteristic. Due to the "closed-loop design, open-loop operation" characteristic of distribution networks, when a power outage is caused by a permanent fault, the switches in the active distribution network automatically isolate the fault, dividing the network into outage and energized areas. When a fault occurs in the distribution network, it is promptly located and isolated. Then, through the rational allocation and operation of sectionalizing switches and tie switches in the distribution network, the load in the outage area is quickly restored using the optimal power supply strategy, achieving the ultimate goal of uninterrupted power supply. With the large-scale integration of DG into the distribution network, its inherent characteristics cause fluctuations and randomness in voltage, network losses, power flow distribution, and network structure. If the relationship between DG and the distribution network is not properly managed, not only will the advantages of DG be largely lost, but the distribution network will also become extremely unstable, even leading to an expansion of the outage area. Therefore, it is necessary to develop appropriate and effective smart distribution network fault self-healing strategies to rationally address distribution network faults.
[0003] Existing smart distribution network fault self-healing strategies mainly fall into three categories: mathematical optimization methods, heuristic search methods, and intelligent optimization methods. Mathematical optimization methods, such as branch and bound and dynamic programming, are widely used in distribution network fault self-healing. However, these algorithms are unsuitable for handling large-scale and highly complex systems, often leading to problems such as the curse of dimensionality, high computational cost, long processing time, and poor timeliness. Heuristic search methods are based on the radial tree-like topology of the distribution network, simplifying it into binary trees, minimum spanning trees, etc. By incorporating heuristic information into the search process, the fault recovery problem can be optimized towards the optimal solution. However, heuristic search lacks stability, depends on the initial state, and struggles to find the optimal solution. Common intelligent optimization methods, such as particle swarm optimization and genetic algorithms, also suffer from problems such as generating a large number of infeasible solutions and easily getting trapped in local optima. Summary of the Invention
[0004] Purpose of the invention: To address the above problems, this invention proposes an active distribution network elastic self-healing method and controller that considers dynamic microgrids. Based on hierarchical reinforcement learning, a dynamic fault recovery strategy model is obtained, and a dynamic fault recovery scheme for the distribution network is generated. This avoids the problems of instability, low fault repair rate, and easy getting trapped in local optima in traditional self-healing methods, and realizes the elastic self-healing of the active distribution network in dynamic microgrids.
[0005] Technical Solution: The technical solution adopted in this invention is a resilient self-healing method for active distribution networks considering dynamic microgrids. It includes: generating a dynamic fault recovery scheme for the distribution network based on fault line information and a dynamic fault recovery strategy model, thereby achieving resilient self-healing of the active distribution network in a dynamic microgrid. The dynamic fault recovery strategy model employs a reinforcement learning framework and a hierarchical reinforcement learning algorithm for model training. The high-level strategies of the hierarchical reinforcement learning algorithm include distribution network islanding strategies, and the low-level strategies include distribution network reconfiguration strategies. The architecture of the intelligent agent module under the reinforcement learning framework is as follows: taking the state of the distribution network after a fault as input, and then... A high-level strategy obtains the target corresponding to the islanding action for power loss regions, and then a low-level strategy obtains the distribution network action. The distribution network action is output to the environment module under the reinforcement learning framework. The environment module under the reinforcement learning framework includes an active distribution network simulation model. The distribution network action is input to the active distribution network simulation model to simulate the formation of islands and output the distribution network operation state after the action. The environment module outputs a reward signal based on the distribution network operation state after the action and the target reward function, and feeds it back to the agent module. The distribution network state includes the topology and operating parameters of the distribution network model, and the distribution network action includes the switching state of each line. During the reinforcement learning training process, all states, targets, actions, and rewards appear sequentially, for example: state 1 --> action 1 --> reward 1 --> state 2 --> action 2 --> reward 2... and so on. In this process, the agent and the environment continuously interact.
[0006] The target reward function is a hierarchical target reward function, which includes: detecting the state of the high-level strategy target according to the high-level strategy target detection algorithm; if a valid island is formed, the reward is calculated according to the corresponding state; and detecting whether all the faults currently occurring in the distribution network have been restored according to the termination detection algorithm, that is, whether the load of the nodes affected by the fault in the distribution network has resumed operation. If the distribution network fault has been restored or all available switches have been used, the network enters the termination state, and the reward is calculated according to the termination state.
[0007] The target reward function for the state where effective islands are formed is:
[0008]
[0009] In the formula, H t To form an effective island, the target reward function for the corresponding state is LW when the goal is to restore power. k μ is the priority weight of load k. k Let Γ be the state variable of load k, L be all nodes on the line, and Γ be the state variable of load k. LFor the set of all nodes on the line; if the goal is to restore the load to the maximum extent, LW k Let μ be the load amount of load k. k Let Γ be the state variable of load k, L be all nodes on the line, and Γ be the state variable of load k. L This is the set of all nodes on the line.
[0010] The objective reward function for calculating rewards based on the termination state is:
[0011]
[0012] In the formula, r t The objective reward function is used to calculate rewards based on the termination state, where λ1, λ2, and λ3 are the weighting coefficients for distributed power source operating costs, network losses after self-healing recovery, and voltage offset after self-healing recovery, respectively. Γ DG C represents the set of all distributed power sources on the line. DG,k P represents the cost factor for the power output of the k-th distributed generation source in the distribution network. DG,k x represents the active power output of the k-th distributed generation in the distribution network, where n is the number of branches; i R represents the state of the i-th switch node; i Let I be the resistance of the i-th branch; i V represents the magnitude of the current flowing through the i-th branch. i V ri These are the actual voltage value and the rated voltage value at point i, respectively.
[0013] The distribution network islanding strategy obtains islanding actions for power loss areas based on the overall state characteristics of the distribution network; the distribution network reconfiguration strategy obtains switch closing actions within islands based on the state characteristics of power loss areas, forming microgrids operating in islands; the distribution network reconfiguration strategy is based on the islanding scheme transmitted by the distribution network islanding strategy.
[0014] The active distribution network simulation model constructs the distribution network simulation environment by using the distribution network topology, network parameters, and distributed power source parameters; the active distribution network simulation model simulates the fault state through a fault line generator and outputs the state of the distribution network after the fault; the fault line generator generates fault line information based on the fault distribution law of the distribution network by using the input fault probability parameters.
[0015] The training of the fault dynamic recovery strategy model includes the following process:
[0016] Step 5.1: Generation of target g from high-level distribution network strategy t To represent the ideal islanding in state-observed distribution network restoration, the high-level strategy π hi Targets are generated sequentially based on the priority order of power loss regions;
[0017] Step 5.2: The low-level strategy takes actions to make the action after c time steps as close as possible to the target state g. t ;
[0018] Step 5.3: Training of the low-level policies is completed using standard methods, adding g to the value and policy models. t As additional input;
[0019] Step 5.4: In the off-policy training part, hierarchical reinforcement learning regenerates a new high-level action. To correct the problem that the actions generated by the low-level strategy are different from the actions of historical samples, the high-level actions that maximize the trajectory conditional probability of the samples are selected as new experience for retraining.
[0020] The preferred error loss function for training the fault dynamic recovery strategy model is:
[0021]
[0022] In the formula, Loss is the error loss function. For low-level strategies, H t To form the target reward function for the corresponding state in the case of effective island formation, r t The objective reward function is used to calculate the reward based on the termination state. Let s be the action value function. t Let g be the environmental state at time t. t For the target state of the low-level policy, a t Let t be the action at time t, and γ be the discount factor.
[0023] This invention provides a microgrid controller, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned active distribution network resilient self-healing method considering dynamic microgrids.
[0024] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned method for resilient self-healing of active distribution networks considering dynamic microgrids.
[0025] Beneficial Effects: Compared with existing technologies, this invention has the following advantages: First, this invention considers the role of dynamic microgrids in the self-healing of active distribution networks and establishes a reinforcement learning framework for real-time management and control of active distribution networks. This enables the self-healing method to perceive the changing state of the microgrid, more closely approximating the actual situation of the dynamic microgrid, thus improving the self-healing effect. Second, this invention designs a hierarchical reinforcement learning mechanism, modeling the dynamic recovery strategy of distribution network faults according to a hierarchical decision-making mechanism. This more accurately reflects the control process of the dynamic microgrid, thereby improving the self-healing effect. Finally, the reward mechanism of reinforcement learning is improved for dynamic microgrids, cleverly enhancing the reinforcement learning effect and efficiency. This invention uses a hierarchical reinforcement learning model to manage the elastic self-healing process of active distribution networks in real time, generating appropriate and effective intelligent distribution network fault self-healing strategies to reasonably resolve distribution network faults. Attached Figure Description
[0026] Figure 1 This is a flowchart of the resilient self-healing method for active distribution networks considering dynamic microgrids, as described in this invention, following the sequence from model establishment to model execution.
[0027] Figure 2 This is a diagram of the reinforcement learning framework for real-time management and control of active power distribution networks as described in this invention. Detailed Implementation
[0028] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0029] The active distribution network resilient self-healing method considering dynamic microgrids described in this invention includes: generating a dynamic fault recovery scheme for the distribution network based on the actual fault line information and a dynamic fault recovery strategy model, thereby realizing the active distribution network resilient self-healing of dynamic microgrids.
[0030] The flowchart for establishing a dynamic fault recovery strategy model is as follows: Figure 1 As shown, taking a specific active power distribution network system as an example, the process of establishing and training a dynamic fault recovery strategy model is described in detail:
[0031] Step 1: Obtain the topology, network parameters, and distributed generation parameters of the active distribution network, and construct an active distribution network simulation model.
[0032] In this step, the distribution network topology, network parameters, and distributed generation parameters are obtained to simplify the distribution network topology and construct a distribution network simulation environment. Then, a fault line generator is constructed, fault probability parameters are input, and fault line information is generated based on the fault distribution patterns of the distribution network.
[0033] Step 2: Build a reinforcement learning framework for real-time management and control of active power distribution networks.
[0034] Step 2.1: Establish the state space S.
[0035] All system information acquired by the agent collectively constitutes the system state space. System state information represents the environmental information perceived by the agent, including changes in the environment after the agent executes its output actions. It serves as the basis for the agent to make decisions and evaluate its long-term benefits. In the distribution network model, the system state can be represented by two parts: the distribution network topology and the distribution network operating parameters. The topology state space S is constructed separately for each. G and the state space of the running parameters S O :
[0036] S={S G S O}(1)
[0037] S G ={ <V i,t V j,t >}(2)
[0038]
[0039] in, <V i,t V j,t > indicates node V in the distribution network at time t i and V j The edges formed represent the state of the distribution network topology by describing the information of all edges in the distribution network topology; These represent the active and reactive power consumed by node i at time t, respectively. These represent the active and reactive power outputs of the photovoltaic device at node i at time t;
[0040] Step 2.2: Establish the action space.
[0041] The agent's actions are designed as state transitions for each circuit switch, with each action corresponding to a circuit switch. The action space is designed as follows:
[0042] In the formula, A represents the action space; N1 represents the number of branches in the system; a i This indicates changing the switching state of the i-th line in the system. Specifically, if the i-th line is currently open, its switch is closed to reconnect the line; if the i-th line is currently closed, its switch is opened to disconnect the line and terminate the operation. This design effectively avoids invalid action selection; N j This is the set of lines that have been operated on in the j-th round. This design can effectively avoid the invalidity of actions.
[0043] Step 2.3: Establish the state transition function and relevant constraints of the distribution network.
[0044] Under the following distribution network-related constraints, distribution network state simulation is performed based on the system's state transition function. The constraints include equality constraints and inequality constraints. The constraints corresponding to the distribution system in this embodiment include:
[0045] (1) Equality constraints:
[0046] ① This mainly refers to power flow constraints, which can be understood as the power demand at node t being equal to the difference between the inflow power and the outflow power:
[0047]
[0048] In the formula, IE t OE t Let each be a set of edges flowing into and out of node t, respectively. These represent the total inflow power and total outflow power at node t, respectively; P t Let be the power requirement of node t.
[0049]
[0050] In the formula, Let be the imaginary part of the current. and The real and imaginary parts of voltage; P DG,k and Q DG,k Let represent the active and reactive power of the distributed power source k.
[0051] ② Kirchhoff constraint: The Kirchhoff current equation at node k is as follows:
[0052]
[0053]
[0054] In the formula, and These are the imaginary and real parts of the load current flowing out of node k, respectively; and These are the imaginary and real parts of the distributed power source current flowing into node k, respectively; and These are the imaginary and real parts of the branch current in the branch line km, respectively.
[0055] ③ The radial constraint condition is based on the following characteristics of the spanning tree: By introducing a binary variable γ, it can be ensured that the distribution network topology corresponds to the spanning tree connected to the main substation. Each node except the root node (substation node) has only one parent node, which can be expressed by the following formula:
[0056]
[0057]
[0058]
[0059]
[0060]
[0061] In the formula: γ is a binary variable, Γ B / Γ SW To remove the set of branches that have been disconnected; Γ sub For the set of substation nodes in the system; Γ N / Γ sub Let s(k) be the set of system nodes excluding substation nodes, and let s(k) be the set of all nodes connected to node k.
[0062] (2) Inequality constraints:
[0063] ① Branch current constraint: The magnitude of the branch current cannot exceed the maximum allowable current of the branch.
[0064] I i ≤I i max (14) In formula I i I max These represent the magnitude of the current flowing through branch i and the maximum allowable current for branch i, respectively.
[0065] ② Node voltage constraints: that is, the node voltage value must be maintained at the minimum and maximum allowable voltage values of the node.
[0066] U i min ≤U i ≤U i max (15)
[0067] In the formula U i U is the voltage value at node i; imin U imax These are the minimum and maximum allowable voltage values for node i, respectively.
[0068] ③ Feeder capacity constraint: The power of a branch cannot exceed the maximum allowable power value of the branch.
[0069] S i ≤S i max (16)
[0070] In the formula S i S imax These are the power of branch i and the maximum allowable power, respectively.
[0071] ④ Transformer overload constraint: This means that the power carried by the transformer cannot exceed the maximum allowable power value.
[0072] S j ≤S jmax (17)
[0073] In the formula S j S jmax These are the power and maximum allowable power of transformer j, respectively.
[0074] ⑤ Distributed power supply capacity constraint: That is, the output of the DG on node i cannot exceed the maximum capacity of the connected DG.
[0075]
[0076] Step 2.4: Based on the distribution network simulation model, fault line generator, state space, action space, and state transition function, build a reinforcement learning framework for real-time control of active distribution networks.
[0077] As attached Figure 2 As shown, the reinforcement learning module mainly comprises four elements: state, action, transition probability, and reward function. The state is the machine's perception of the environment; all possible states constitute the state space. Actions are the actions the machine takes; all possible actions constitute the action space. Transition probability is the probability that the current state will transition to another state after performing an action. The reward function is the reward given to the machine by the environment during state transitions. Reinforcement learning continuously tries different approaches in the environment, adjusts its policy based on feedback, and generates a final policy. Based on this policy, the machine knows which action to perform in which state.
[0078] Step 3: Based on the hierarchical reinforcement learning mechanism, model the dynamic fault recovery strategy of the distribution network according to the hierarchical decision-making mechanism.
[0079] Step 3.1: Determine the load restoration priority based on the importance of the loads in the distribution network, and then sort the power loss areas after calculating them;
[0080] Step 3.2: Model the distribution network islanding strategy. As a high-level strategy, input the overall state characteristics of the distribution network to obtain the islanding actions for power loss areas;
[0081] Step 3.3: Model the distribution network reconfiguration strategy. As a low-level strategy, input the state characteristics of the power loss area to obtain the switching closing actions within the island, forming a microgrid operating in an island. The low-level strategy is based on the island division scheme transmitted from the high-level strategy, and is combined with the restored load of the upper level to obtain the total power of important loads restored by the distribution network.
[0082] Step 3.4: Construct a high-level strategy target detection algorithm to detect the state of high-level strategy targets and detect the current state of the distribution network. If an effective island is formed, calculate the reward according to the corresponding state.
[0083] Step 3.5: Construct a low-level strategy termination detection algorithm to detect the termination status of the low-level strategy. If the distribution network fault is recovered or all available switches are in use, the strategy enters the termination state, and the reward is calculated according to the termination state.
[0084] Step 4: Model the hierarchical target reward function and generate reward signals based on the operating status of the distribution network.
[0085] Step 4.1: Establish the reward function for the target state. When a fault occurs, the distribution network performs dynamic repair. If an island is formed, the load in the power loss area is restored. A reward is given for this process. The target state reward function is as follows:
[0086]
[0087] In the formula: LW k The priority weight for load k; μ k Let k be the state variable of the load, and L be the total number of nodes on the line. If the objective is simply to restore the load to its maximum extent, then the priority weights can be replaced with the load quantity.
[0088] Step 4.2: Establish the reward function for the termination state. During the distribution network self-healing process, as much load as possible is restored, the system loss after self-healing is minimized, and the voltage deviation after restoration is minimized compared to the voltage before restoration. Based on this, a reward is given after the distribution network fault is restored, and the reward function for the termination state is established:
[0089] (1) Distributed power generation operating cost F PV Assuming a linear relationship with the power it provides, the power supply's output power should be kept as low as possible during normal operation.
[0090]
[0091] In the formula: C DG,k P represents the cost factor for the power output of the k-th distributed generation source in the distribution network. DG,k This represents the active power output by the k-th distributed power source in the distribution network.
[0092] (2) Network loss F after self-healing recovery loss To ensure that network loss is minimized during fault recovery.
[0093]
[0094] In the formula: n is the number of branches; x iR represents the state of the i-th switch node; i Let I be the resistance of the i-th branch; i Let be the magnitude of the current flowing through the i-th branch.
[0095] (3) Ensure voltage deviation F after self-healing recovery V Minimum.
[0096] F V =|V i -V ri | (22)
[0097] In the formula: V i V ri These are the actual voltage value and the rated voltage value at point i, respectively.
[0098] When using different control variables, these objective functions are linearly weighted and finally combined to form a single objective function. The reward function is as follows:
[0099]
[0100] In the formula: s t Let a be the environmental state at time t. t Let λ1, λ2, λ3, and λ4 be the action at time t, and let λ1, λ2, λ3, and λ4 be the weighting coefficients.
[0101] Step 5: The hierarchical reinforcement learning algorithm HIRO is used to train the high-level and low-level strategies simultaneously in the power distribution network simulation environment to obtain the fault dynamic recovery strategy model.
[0102] Step 5.1: Generation of target g from high-level distribution network strategy t To represent the ideal islanding in state-observed distribution network restoration, the high-level strategy π hi Targets are generated sequentially based on the priority order of power loss regions;
[0103] Step 5.2: The low-level strategy takes actions to make the action after c time steps as close as possible to the target state g. t ;
[0104] Step 5.3: Training of the low-level policies is completed using standard methods, adding g to the value and policy models. t As additional input The error loss function for the value is as follows:
[0105]
[0106] In the formula, g t The target state of the low-level strategy. For low-level strategies, r(s) t gt a t s t+1 ) is the reward function.
[0107] Step 5.4: Refine the objective of the high-level policy. In the off-policy training part, hierarchical reinforcement learning regenerates a new high-level action. This addresses the issue of actions generated by low-level strategies differing from historical sample actions. By maximizing the trajectory conditional probability of the samples, the high-level actions that maximize this probability are selected as new experience for retraining.
[0108] In one embodiment, the microgrid controller of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the active distribution network resilient self-healing method considering dynamic microgrids.
[0109] In one embodiment, the computer-readable storage medium of the present invention stores a computer program thereon, which, when executed by a processor, implements the above-described method for resilient self-healing of active distribution networks considering dynamic microgrids.
[0110] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0111] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0112] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0113] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
Claims
1. A method for resilient self-healing of active distribution networks considering dynamic microgrids, characterized in that, include: Based on the faulty line information and the fault dynamic recovery strategy model, a distribution network fault dynamic recovery scheme is generated to realize the active distribution network elastic self-healing of the dynamic microgrid. The fault dynamic recovery strategy model adopts a reinforcement learning framework and uses a hierarchical reinforcement learning algorithm for model training. The high-level strategy of the hierarchical reinforcement learning algorithm includes the distribution network islanding strategy, and the low-level strategy includes the distribution network reconfiguration strategy. The architecture of the agent module under the reinforcement learning framework is as follows: taking the state of the distribution network after the fault as input, the target corresponding to the islanding action for the power loss area is obtained through the high-level strategy, and the action of the distribution network is obtained through the low-level strategy. The action of the distribution network is output to the environment module under the reinforcement learning framework. The environment module under the reinforcement learning framework includes an active distribution network simulation model. The actions of the distribution network are input to the active distribution network simulation model to simulate the formation of islands and output the operating state of the distribution network after the actions. The environment module outputs a reward signal and feeds it back to the agent module based on the operating state of the distribution network after the actions and the target reward function. The state of the distribution network includes the topology and operating parameters of the distribution network model, and the actions of the distribution network include the switching status of each line. The target reward function is a hierarchical target reward function, which includes: detecting the state of the high-level strategy target according to the high-level strategy target detection algorithm; if an effective island is formed, then calculating the reward according to the corresponding state. According to the termination detection algorithm, it is detected whether all the faults currently occurring in the distribution network have been restored, that is, whether the loads of the nodes affected by the faults in the distribution network have resumed operation. If the distribution network faults have been restored or all available switches have been used, the system enters the termination state, and the reward is calculated according to the termination state.
2. The method for resilient self-healing of active distribution networks considering dynamic microgrids according to claim 1, characterized in that: The target reward function for the state where effective islands are formed is: ; In the formula, H t To form an effective island, the target reward function for the corresponding state, when the goal is to restore power. LW k For load k Priority weights, For load k The state variables, Γ L The set of all nodes on the line; If the goal is to restore the load to the maximum extent possible LW k For load k The load capacity, For load k The state variables, Γ L This is the set of all nodes on the line.
3. The method for resilient self-healing of active distribution networks considering dynamic microgrids according to claim 1, characterized in that: The objective reward function for calculating rewards based on the termination state is: ; In the formula, The objective reward function is used to calculate the reward based on the termination state. λ 1. λ 2. λ 3 represents the weighting coefficients for distributed power source operating costs, network losses after self-healing recovery, and voltage deviation after self-healing recovery. Γ DG This refers to the collection of all distributed power sources on the line. C DG,k For the distribution network k Cost coefficient of distributed power output; P DG,k For the distribution network k The active power output of a distributed power source n The number of branches; For the first i The status of the switching nodes of the branch circuit; R i For the first i The resistance of each branch; I i For the flow through the first i The magnitude of the current in each branch, V i , V ri The first i The actual voltage value and rated voltage value of each branch circuit.
4. The method for resilient self-healing of active distribution networks considering dynamic microgrids according to claim 1, characterized in that: The distribution network islanding strategy obtains islanding actions for power loss areas based on the overall state characteristics of the distribution network; the distribution network reconfiguration strategy obtains switch closing actions within islands based on the state characteristics of power loss areas, forming microgrids operating in islands; the distribution network reconfiguration strategy is based on the islanding scheme transmitted by the distribution network islanding strategy.
5. The method for resilient self-healing of active distribution networks considering dynamic microgrids according to claim 1, characterized in that: The active distribution network simulation model constructs the distribution network simulation environment by using the distribution network topology, network parameters, and distributed power source parameters; the active distribution network simulation model simulates the fault state through a fault line generator and outputs the state of the distribution network after the fault; the fault line generator generates fault line information based on the fault distribution law of the distribution network by using the input fault probability parameters.
6. The method for resilient self-healing of active distribution networks considering dynamic microgrids according to claim 1, characterized in that, The training of the fault dynamic recovery strategy model includes the following process: Step 5.1: Generation of High-Level Strategy Targets for Distribution Network To represent the ideal islanding in state-observed distribution network restoration, high-level strategies Targets are generated sequentially based on the priority order of power loss regions; Step 5.2: Low-level strategies take actions to achieve... c The actions taken after each time step should be as close as possible to the target state. ; Step 5.3: Training of the low-level policies is completed using standard methods, adding [the following to the value and policy models]... As additional input; Step 5.4: In the off-policy training part, hierarchical reinforcement learning regenerates a new high-level action. To correct the problem that the actions generated by the low-level strategy are different from the actions of historical samples, the high-level actions that maximize the trajectory conditional probability of the samples are selected as new experience for retraining.
7. The method for resilient self-healing of active distribution networks considering dynamic microgrids according to claim 6, characterized in that: The error loss function for training the fault dynamic recovery strategy model is: ; In the formula, Loss Let be the error loss function. This is a low-level strategy. H t To form the target reward function for the corresponding state in the case of effective island formation, The objective reward function is used to calculate the reward based on the termination state. Let the action value function be... for t The environmental state at any given time, for t The low-level policy objective state at any given moment. a t for t The action at time t, where γ is the discount factor.
8. A microgrid controller, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the active distribution network resilient self-healing method considering dynamic microgrids as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the active distribution network resilient self-healing method considering dynamic microgrids as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Active power distribution network fault recovery method based on reinforcement learning method
CN113872198A
Multi-microgrid system layered reinforcement learning optimization method and system, and storage medium
CN115115211A