Multi - Microgrid Distributed Control Method and Device Based on Multi - Agent Reinforcement Learning

Through the multi-agent reinforcement learning method, the problems of high computational complexity and poor adaptability in the formation of multi-micro grids are solved, and the rapid and robust multi-micro grid recovery is achieved, which improves the resilience of the power system and renewable energy utilization rate.

CN120110023BActive Publication Date: 2025-08-05NAT UNIV OF DEFENSE TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510590882.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-05
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Traditional power systems have high computational complexity and poor adaptability in the formation of multi-micro grids, making it difficult to deal with uncertain disaster scenarios. Centralized control has privacy leakage and network security risks, and it is difficult to operate efficiently in large-scale systems.

Method used

The multi-agent reinforcement learning method is adopted to model the multi-micro grid formation process as a distributed partially observable Markov decision-making process, and the connection-only mechanism and the self-organizing network mechanism are used for environmental interaction, and the strategy network is trained in combination with DQN algorithm and action mask technology to achieve local decision-making and global optimization.

Benefits of technology

Reduces the computational complexity, improves the system's adaptability to dynamic uncertain scenarios, avoids privacy leakage and single point failure risks, achieves rapid recovery of key load power supply in minute-level, and improves renewable energy utilization and system resilience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120110023B_ABST
    Figure CN120110023B_ABST
Patent Text Reader

Abstract

This application relates to a multi-microgrid distributed control method and device based on multi-agent reinforcement learning. The method includes: modeling the power network system, and modeling the formation process of multi-microgrids as a distributed partially observable Markov decision process; deploying a multi-agent reinforcement learning algorithm according to the partially observable Markov decision process, and each agent respectively interacts with the power network system for environmental interaction sampling information and stores it into its own experience replay pool; based on the DQN algorithm, each agent trains its corresponding policy network to learn the optimal behavior by sampling the experience samples in the experience replay pool and according to the action masking technology and experience screening technology; using the trained policy network model to control the switch actions of each area of the power system respectively to form a microgrid and restore the load power supply. Using this method can effectively improve the resilience of the power system and achieve the rapid formation and restoration of multi-microgrids.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of smart grid optimization, and particularly to a multi-microgrid distributed control method and device based on multi-agent reinforcement learning. Background Art

[0002] Traditional power systems adopt a top-down centralized energy production mode. However, with the decentralization and digitalization of energy production, power systems are gradually transforming into a distributed and diversified microgrid mode. Microgrids can provide power supply to critical infrastructure after disasters by integrating distributed energy resources, thereby enhancing the resilience of the power system. Especially in extreme disaster events, forming multi-microgrids to restore critical loads has become an effective strategy. The problem of forming multi-microgrids faces many challenges. First, the decision-making space is too large, resulting in high computational complexity. Especially in large-scale power systems, the number of decision variables grows exponentially. Second, most existing research methods are based on deterministic scenarios, while actual disaster scenarios are often uncertain and dynamic, making it difficult to effectively respond to them through traditional optimization methods. In addition, traditional centralized control methods rely on global information, which not only poses risks of privacy leakage and network security, but also is difficult to operate efficiently in large-scale systems.

[0003] Existing technologies have problems such as high computational complexity, poor adaptability, and low training efficiency when dealing with the problem of forming multi-microgrids. Therefore, there is an urgent need for a new method that can effectively enhance the resilience of the power system and achieve the rapid formation and restoration of multi-microgrids under the condition of incomplete measurement and control information. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a multi-microgrid distributed control method and device based on multi-agent reinforcement learning that can effectively enhance the resilience of the power system and achieve the rapid formation and restoration of multi-microgrids.

[0005] A multi-microgrid distributed control method based on multi-agent reinforcement learning, the method includes:

[0006] Model the power network system, and model the process of forming multi-microgrids as a distributed partially observable Markov decision process;

[0007] Deploy a multi-agent reinforcement learning algorithm according to the partially observable Markov decision process. Each agent respectively interacts with the power network system through a connection-only mechanism and a self-organizing network mechanism to sample information from the environment and store it in its own experience replay pool. Based on the DQN algorithm, each agent trains its corresponding policy network to learn the optimal behavior by sampling experience samples from the experience replay pool and according to the action masking technology and experience screening technology;

[0008] The trained policy network model is used to control the switching actions of each area of the power system to form a microgrid and restore the power supply to the load.

[0009] A multi-microgrid distributed control device based on multi-agent reinforcement learning, the device includes:

[0010] A system modeling module for modeling the power network system and modeling the formation process of the multi-microgrid as a distributed partially observable Markov decision process;

[0011] A model training module for deploying a multi-agent reinforcement learning algorithm according to the partially observable Markov decision process. Each agent respectively interacts with the power network system through a connection-only mechanism and a self-organizing network mechanism to sample information from the environment and then store it in its own experience replay pool; based on the DQN algorithm, each agent trains its corresponding policy network to learn the optimal behavior by sampling the experience samples in the experience replay pool and according to the action masking technology and the experience screening technology;

[0012] A multi-microgrid distributed control module for using the trained policy network model to control the switching actions of each area of the power system to form a microgrid and restore the power supply to the load.

[0013] The above multi-microgrid distributed control method and device based on multi-agent reinforcement learning first model the problem as a multi-agent partially observable Markov decision process. Each agent makes decisions only relying on local observation information without global data concentration, solving the problems of communication interruption and incomplete measurement in disasters, enhancing the system's adaptability to dynamic uncertain scenarios, and avoiding the privacy leakage and single-point failure risks of centralized control. The connection-only mechanism is used to limit the communication range, reduce the data transmission volume and link dependence, reduce the privacy risk and enhance the anti-destruction ability; the self-organizing network mechanism supports the dynamic reconstruction of the communication topology, automatically maintaining local cooperation when a link failure is caused by a disaster, ensuring that the control instruction is not interrupted, and enhancing the disaster preparedness resilience. Finally, according to the action masking technology, the physical constraints of the power system are embedded to mask invalid actions, compressing the decision space from exponential level to polynomial level, greatly reducing the computational complexity; the experience screening technology preferentially retains high-value disaster preparedness samples, alleviating the problem of scarce disaster data, enhancing the learning efficiency of the agent for extreme scenarios, and achieving rapid recovery within minutes. This application supports regional parallel decision-making through a multi-agent distributed architecture, avoiding the centralized computing delay, and is applicable to distribution networks with hundreds of nodes; the DQN-based policy network captures nonlinear relationships through end-to-end learning, realizing the joint optimization of topology optimization and energy scheduling, and enhancing the critical load recovery rate and renewable energy utilization rate. Each agent operates independently and cooperates through wired communication. Even if some areas are isolated from the main network, it can still optimize the policy independently based on local experience, ensuring the continuous and stable operation of the island microgrid, and having strong robustness to extreme events such as multiple faults and network attacks. Brief Description of the Drawings

[0014] Figure 1 It is a schematic flow diagram of a multi - microgrid distributed control method based on multi - agent reinforcement learning in an embodiment;

[0015] Figure 2 It is a schematic diagram of the interaction between multi - agents and the power grid environment in an embodiment;

[0016] Figure 3 It is a schematic diagram of the specific decision - making process of a single agent in an embodiment;

[0017] Figure 4 It is a schematic diagram of the action masking technology in another embodiment;

[0018] Figure 5 It is a schematic flow diagram of the experience screening process in an embodiment;

[0019] Figure 6 It is a schematic diagram of the solution result formed by the multi - microgrid in an embodiment;

[0020] Figure 7 It is a structural block diagram of a multi - microgrid distributed control device based on multi - agent reinforcement learning in an embodiment. Detailed Embodiments

[0021] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0022] In one embodiment, as Figure 1 shown, a multi - microgrid distributed control method based on multi - agent reinforcement learning is provided, including the following steps:

[0023] Step 102, model the power network system, and model the multi - microgrid formation process as a distributed partially observable Markov decision process.

[0024] The power network system model includes bus nodes and feed - line circuits, denoted as where represents the bus, represents the feed - line circuit; at each node there is a load consuming active power and reactive power . The load has priorities, with weights Representation. With the help of information and communication technology, each line and load in the system is equipped with a controllable switch, which can be connected and disconnected through remote control. When the large power grid loses service due to extreme natural disasters, multiple microgrids are formed by distributed generators installed at certain nodes to supply power to critical loads, and each microgrid is powered by a distributed generator.

[0025] Model the process of forming multi - microgrids as a distributed partially observable Markov decision process, where the partially observable Markov decision process includes the state information of the environment , the observation information of the agent , the action space of the agent , the reward function , the state transition process .

[0026] Specifically, contains the characteristic data of all nodes and lines in the power grid system, and what each agent observes is a subset of ; Suppose there are agents in total, and the observation of the th agent, and are the active power and reactive power output by the distributed generator respectively, and represent the active and reactive power consumed by the load respectively, is a 0 - 1 type variable representing the connection state of the load, 0 represents disconnected, and 1 represents connected.

[0027] The action space of the agent contains the switches of all loads within the observed range of the agent. In the initial state, it is assumed that all switches in the system are in the disconnected state. The agent selects a load switch to turn it into the connected state, and the already - connected switches are not allowed to be turned into the disconnected state again, which is the only - connection mechanism; when a load is turned into the connected state, all lines between the node where the load is located and the node where the distributed generator is located are found through the graph - search method and switched to the connected state, which is the self - organizing network mechanism; similarly, the line switches also follow the only - connection mechanism.

[0028] In the initial state, , and are both set to 0. When a load is selected to be connected, its device parameters are extracted from the corresponding parameters of the test - case system and assigned to and , It is assigned a value of 1. Due to the existence of only the connection mechanism, these data remain unchanged during a training round after being changed until the start of the next training round.

[0029] This application adopts a distributed multi-agent framework. The environmental rewards obtained by each agent not only include individual rewards but also global rewards that reflect the overall performance. The purpose of forming a microgrid is to maximize the power supply load of priority loads. Thus, the individual reward is related to the change in the load of the current microgrid, and its formula is shown in Equation 1.

[0030] (1);

[0031] where is the priority weight of the relevant load. When it means that a newly added load to the microgrid has obtained power supply. At this time, the environmental reward is set to itself; when it means that the agent has selected an unprofitable action. At this time, the environmental reward is set to a negative value. Thus, the individual reward function of the agent can be expressed by Equation 2.

[0032] (2)

[0033] In particular, the operation of the power grid system needs to meet many operation constraints to ensure the stable and safe operation of the system. These constraints are shown in Equations (3)-(7).

[0034] (3)

[0035] (4)

[0036] (5)

[0037] (6)

[0038] (7)

[0039] Constraints 3 and 4 are the upper and lower bounds of the generator output power. Constraint 5 ensures the power balance between the generators and loads in the system. is the line loss; Constraint 6 is the line transmission capacity constraint; Constraint 7 is that the operating voltage of each bus is within the safe range. When these constraint conditions are not met, the individual reward is also a negative value.

[0040] For the global reward, if the actions of all agents result in an increase in the cumulative weighted load of the entire power network, all agents will receive a positive value as the global reward. The formula is shown in Equation 8.

[0041] (8)

[0042] The goal of the reward function is to maximize the expected total reward obtained by each agent, as shown in Equation (9).

[0043] (9)

[0044] Step 104: Deploy the multi-agent reinforcement learning algorithm according to the partially observable Markov decision process. Each agent interacts with the power network system through the connection-only mechanism and the self-organizing network mechanism to sample environmental information, and then stores it in its own experience replay pool. Based on the DQN algorithm, each agent trains its corresponding policy network to learn the optimal behavior by sampling the experience samples in the experience replay pool and according to the action masking technique and the experience screening technique.

[0045] Deploy the multi-agent reinforcement learning method from the perspective of distributed computing. As Figure 3 shown, each agent obtains the information within its observation range and then inputs into its policy network to output the action selected in the current state , that is, each agent separately selects the load to supply power to and turns on the switch. These individual actions are aggregated into a joint action and input into the power system to change the state of the system . Then, each agent obtains the corresponding environmental feedback information, including the reward value and the state at the next moment . In this process, the current state, the actions taken, the environmental rewards, and the state at the next moment of the agent are saved in their respective experience pools in the form of tuples for subsequent training.

[0046] In the design of the policy network, the hidden layer of the policy network consists of an adaptive attention mechanism, a graph convolutional network, and a fully connected layer. In the input data of the policy network, is the node feature data, is the edge connection information. Since each agent only changes the connection state of one load at a time, the changing part in the feature data is very small, which may make it difficult for the policy network to identify valuable features. Therefore, an adaptive attention mechanism is introduced. First, the node feature data is input into several fully connected layers, and then the normalized feature data is obtained through the softmax layer. Then, it is weighted with the original node feature data to calculate the new node feature data after being processed by the adaptive attention mechanism. The new node feature data and are jointly used as the input of the graph convolutional network. Then, it is mapped to the action value distribution through the fully connected layer. Then, the agent action The greedy strategy selects the next action to be executed. The activation functions used in the policy network are all ReLU.

[0047] This application uses the DQN algorithm framework to train each agent to learn the policy. The DQN framework has a dual-policy network structure. In addition to the Q-network used to generate action values at the current moment, there is also a target Q-network used to generate action values at the next moment. The loss function of the Q-network is as shown in Equation 10.

[0048] (10)

[0049] (11)

[0050] where and represent the parameters of the Q-network and the target Q-network respectively, is the discount factor, is the experience pool; represents the action value of the agent executing action under observation ; is the reward value generated by taking action and transferring to the next moment state .

[0051] The action mask technique and the experience screening technique are adopted to optimize the training process. Since selecting repeated actions will result in ineffective gains, when the agent selects actions within one round of training, the actions that have been selected before need to be masked. As Figure 4 shown, the action mask technique multiplies the action value distribution output by the Q-network by a mask sequence consisting of 0s and 1s. The length of the mask sequence is equal to the number of actions output by the Q-network. Let the positions of the actions to be discarded in the mask be set to 0, and the positions of the actions to be retained be set to 1.

[0052] The experience screening mechanism controls the ratio of positive reward samples to negative reward samples in the agent's experience pool to be 1:1. As Figure 5 shown, that is, when the number of positive reward samples in the experience pool is greater than the number of negative reward samples, only negative reward samples are allowed to enter the experience pool; when the number of negative reward samples is greater than the number of positive reward samples, only positive reward samples are allowed to enter the experience pool.

[0053] Step 106, use the trained policy network model to control the switch actions of each area of the power system to form a microgrid and restore the load power supply.

[0054] In a known power system scenario, multiple regions are divided as the observation inputs for the regional agents. Multiple agents interact with the power system environment respectively to obtain observation and feedback information, save this information in an experience replay buffer, and then sample the sample data in the experience pool by means of random mini-batch sampling to train the Q network. The target Q network adopts a soft update method. Each time it is updated, some parameters in the Q network are assigned to the target Q network. The soft update formula is shown in Equation (12).

[0055] (12)

[0056] Where is a coefficient, taking a number between 0 and 1.

[0057] During the training process, when the average reward value reaches a relatively stable state, the training is stopped, and the trained Q network model is saved. The power system parameters of the training scenario are input into the Q network model to obtain decision-making actions, and finally the networking schemes of multiple microgrids are solved.

[0058] The above multi-microgrid distributed control method based on multi-agent reinforcement learning first models the problem as a multi-agent partially observable Markov decision process. Each agent makes decisions only relying on local observation information without global data concentration, solves the problems of communication interruption and incomplete measurement in disasters, improves the adaptability of the system to dynamic uncertain scenarios, and avoids the privacy leakage and single-point failure risks of centralized control. The only connection mechanism is used to limit the communication range, reduce the data transmission volume and link dependence, reduce the privacy risk and improve the anti-destruction ability; the self-organizing network mechanism supports dynamic reconstruction of the communication topology, automatically maintains local cooperation when disasters cause link failures, ensures that control instructions are not interrupted, and enhances disaster preparedness resilience. Finally, according to the action mask technology, physical constraints of the power system are embedded to mask invalid actions, compress the decision space from exponential level to polynomial level, greatly reduce the computational complexity; the experience screening technology preferentially retains high-value disaster preparedness samples, alleviates the problem of scarce disaster data, improves the learning efficiency of agents for extreme scenarios, and realizes rapid recovery in minutes. This application supports regional parallel decision-making through a multi-agent distributed architecture, avoids centralized computing delay, and is applicable to distribution networks with hundreds of nodes; the policy network based on DQN captures non-linear relationships through end-to-end learning, realizes the joint optimization of topology optimization and energy scheduling, and improves the critical load recovery rate and renewable energy utilization rate. Each agent runs independently and cooperates through wired communication. Even if some regions are isolated from the main network, it can still autonomously optimize the strategy based on local experience, ensure the continuous and stable operation of the island microgrid, and has strong robustness to extreme events such as multiple faults and cyber attacks.

[0059] In one embodiment, the power network system is modeled, including:

[0060] The power grid system model includes bus nodes and feed lines, denoted as where represents a bus, and represents a feed line;

[0061] At each node there is a load that consumes active power and reactive power ;

[0062] Controllable switches are equipped for each line and load in the power system and are connected and disconnected through remote control. When the large power grid loses service due to extreme natural disasters, multiple microgrids are formed by distributed generators installed at some nodes to supply power to critical loads, and each microgrid is powered by a distributed generator.

[0063] In one embodiment, the partially observable Markov decision process includes the state information of the environment, the observation information of the agent, the action space of the agent, the reward function, and the state transition process; the state information of the environment contains the characteristic data of all nodes and lines in the power grid system, and the observation information of each agent is a subset of the state information of the environment ; the observation information of the agent , and are the active power and reactive power output by the distributed generator respectively, and represent the active and reactive power consumed by the load respectively, is the edge connection information, is a 0-1 type variable representing the connection state of the load, 0 means disconnected, and 1 means connected, is the th agent; the action space of the agent contains the switches of all loads within the observable range of the agent.

[0064] In one embodiment, the reward function includes an individual reward function and a global reward; the model of the reward function is:

[0065] .

[0066] Constraint conditions:

[0067] ;

[0068] where is the line loss, is the individual reward, represents the global reward, and are the active power and reactive power output by the distributed generator respectively, and represent the upper and lower limits of the active power output by the distributed generator respectively, represents the bus, represents the bus node, represents the set of bus nodes, represents the feeder line, represents the set of feeder lines, represents the maximum capacity of the feeder line, j represents the imaginary unit, represents the reactive power transmitted on the feeder line, represents the bus node where the active power consumed by the connected load is, represents the active power transmitted on the feeder line, represents the bus node where the voltage amplitude is.

[0069] In one embodiment, the connection mechanism only means that the agent selects a load switch to turn to the connected state, and the already connected switch is not allowed to turn to the disconnected state again; the self-organizing network mechanism means that when a load is turned to the connected state, all the lines between the node where the load is located and the node where the distributed generator is located are found through the graph search method and switched to the connected state.

[0070] In one embodiment, the hidden layer of the policy network consists of an adaptive attention mechanism, a graph convolutional network and a fully connected layer; the policy network has a dual-policy network structure, including a Q network for generating action values at the current moment and a target Q network for generating action values at the next moment.

[0071] In one embodiment, the loss function of the Q network is:

[0072] ;

[0073] where, and represent the parameters of the Q network and the target Q network respectively, is the discount factor, is the experience pool; represents the action value of the agent executing the action under the observation ; is the reward value generated by taking the action and transferring to the next moment state .

[0074] In one embodiment, the target Q-network is updated as follows:

[0075] ;

[0076] where is a coefficient, taking a value between 0 and 1.

[0077] In one embodiment, the action masking technique means multiplying the action value distribution output by the Q-network with a masking sequence consisting of 0s and 1s. The length of the masking sequence is equal to the number of actions output by the Q-network. Among them, the positions of the actions to be discarded in the mask are set to 0, and the positions of the actions to be retained are set to 1; the experience screening mechanism means controlling the ratio of positive reward samples to negative reward samples in the agent's experience pool to be 1:1. When the number of positive reward samples in the experience pool is greater than the number of negative reward samples, only negative reward samples are allowed to enter the experience pool; when the number of negative reward samples is greater than the number of positive reward samples, only positive reward samples are allowed to enter the experience pool.

[0078] In a specific embodiment, it is assumed that after an extreme disaster causes the public power grid to lose service, all loads in the power grid lose power supply, and only the distributed generators installed at some nodes can be relied on to supply power to the loads. As Figure 6 shown, Figure 2 the power grid environment exemplified in

[0079] should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover,

[0080] In one embodiment, as Figure 7 shown, a multi-microgrid distributed control device based on multi-agent reinforcement learning is provided, including: a system modeling module 702, a model training module 704, and a multi-microgrid distributed control module 706, where:

[0081] The system modeling module 702 is used to model the power network system and model the formation process of multiple microgrids as a distributed partially observable Markov decision process;

[0082] The model training module 704 is used to deploy a multi-agent reinforcement learning algorithm according to the partially observable Markov decision process. Each agent interacts with the power network system through the only connection mechanism and the self-organizing network mechanism to sample environmental information and then store it in its own experience replay pool; Based on the DQN algorithm, each agent trains its corresponding policy network to learn the optimal behavior by sampling the experience samples in the experience replay pool and according to the action masking technology and the experience screening technology;

[0083] The multi-microgrid distributed control module 706 is used to use the trained policy network model to control the switch actions of each area of the power system to form a microgrid and restore the load power supply.

[0084] For the specific limitations of the multi-microgrid distributed control device based on multi-agent reinforcement learning, reference can be made to the limitations of the multi-microgrid distributed control method based on multi-agent reinforcement learning in the above text, which will not be elaborated here. Each module in the above multi-microgrid distributed control device based on multi-agent reinforcement learning can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0085] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0086] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0087] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A multi-microgrid distributed control method based on multi-agent reinforcement learning, characterized in that: The method comprises: Modeling the power network system, the multi-microgrid formation process is modeled as a distributed partially observable Markov decision process; According to the partially observable Markov decision process, a multi-agent reinforcement learning algorithm is deployed. Each agent interacts with the power grid system through the connection-only mechanism and the self-organizing network mechanism to sample information and then stores it in its own experience replay pool. Based on the DQN algorithm, each agent learns the optimal behavior by sampling experience samples in the experience replay pool and training its corresponding strategy network based on the action masking technology and the experience screening technology. The connection-only mechanism means that the agent selects a load switch to be connected, and the already connected switch is not allowed to be converted to the disconnected state again. The self-organizing network mechanism means that when a load is converted to the connected state, the graph search method is used to find the location of the load. All lines between the node and the node where the distributed generator is located are switched to a connected state; the action masking technology means multiplying the Q network output action value distribution by a mask sequence composed of 0 and 1, and the length of the mask sequence is equal to the number of Q network output actions, wherein the action position to be discarded in the mask is set to 0, and the action position to be retained is set to 1; the experience screening mechanism means that the ratio of positive and negative return samples in the experience pool of the control agent is 1:

1. When the number of positive return samples in the experience pool is greater than the number of negative return samples, only negative return samples are allowed to enter the experience pool; when the number of negative return samples is greater than the number of positive return samples, only positive return samples are allowed to enter the experience pool; The trained strategy network model is used to control the switching actions in each area of the power system to form a microgrid and restore power supply to the load.

2. The method according to claim 1, characterized in that Modeling of power network systems, including: The power network system model includes busbar nodes and Feeder lines, represented by ,in represents the busbar, Indicates a feeder line; At each node There is a consumption of active power and reactive power load; Each line and load in the power system is equipped with a controllable switch to connect and disconnect through remote control. When the main power grid loses service due to extreme natural disasters, multiple microgrids are formed by distributed generators installed at certain nodes to supply power to critical loads. Each microgrid is powered by a distributed generator.

3. The method according to claim 1, characterized in that The partially observable Markov decision process includes the state information of the environment, the observation information of the agent, the action space of the agent, the reward function and the state transition process; the state information of the environment contains the characteristic data of all nodes and lines in the power grid system, and the observation information of each agent Is the status information of the environment A subset of the agent's observation information , and are the active power and reactive power output by the distributed generators, and Represent the active and reactive power consumed by the load, For edge information, It is a 0-1 type variable representing the connection status of the load, 0 represents disconnection and 1 represents connection. For the An agent; the action space of the agent includes the switches of all loads within the observation range of the agent.

4. The method according to claim 3, characterized in that The reward function includes an individual reward function and a global reward; the model of the reward function is: Constraints: in, is the line loss, For individual rewards, represents the global reward, and are the active power and reactive power output by the distributed generators, and Respectively represent the upper and lower limits of the active power output of the distributed generator, represents the busbar, represents a busbar node, represents the busbar node set, Indicates the feeder line, represents the set of feeder lines, Indicates the maximum capacity of the feeder line, j represents the imaginary unit, Represents the reactive power transmitted on the feeder line, Indicates a busbar node The active power consumed by the connected load at Indicates the active power transmitted on the feeder line, Indicates a busbar node The voltage amplitude at .

5. The method according to claim 1, wherein The hidden layer of the policy network consists of an adaptive attention mechanism, a graph convolutional network and a fully connected layer; the policy network has a dual-policy network structure, including a Q network for generating action values at the current moment and a target Q network for generating action values at the next moment.

6. The method according to claim 5, characterized in that The loss function of the Q network is: in, and represent the parameters of the Q network and the target Q network respectively, is the discount factor, For the experience pool; Indicates that the agent is observing Next action The action value of To take action Transfer to the next state The reward value generated.

7. The method according to claim 6, characterized in that The target Q network is updated as follows: in, is the coefficient, which is a number between 0 and 1.

8. A multi-microgrid distributed control device based on multi-agent reinforcement learning, characterized in that: The device comprises: System modeling module, used to model the power network system and model the multi-microgrid formation process as a distributed partially observable Markov decision process; The model training module is used to deploy a multi-agent reinforcement learning algorithm based on the partially observable Markov decision process. Each agent interacts with the power grid system through the connection-only mechanism and the self-organizing network mechanism to sample information and then stores it in its own experience replay pool; based on the DQN algorithm, each agent samples the experience samples in the experience replay pool and trains its corresponding strategy network to learn the optimal behavior based on the action masking technology and the experience screening technology; the connection-only mechanism means that the agent selects a load switch to be connected, and the already connected switch is not allowed to be converted to the disconnected state again; the self-organizing network mechanism means that when a load is converted to the connected state, it is found through the graph search method All lines between the node where the load is located and the node where the distributed generator is located are switched to a connected state; the action masking technology means multiplying the Q network output action value distribution by a mask sequence composed of 0 and 1, and the length of the mask sequence is equal to the number of Q network output actions, wherein the action position to be discarded in the mask is set to 0, and the action position to be retained is set to 1; the experience screening mechanism means that the ratio of positive and negative return samples in the experience pool of the control agent is 1:

1. When the number of positive return samples in the experience pool is greater than the number of negative return samples, only negative return samples are allowed to enter the experience pool; when the number of negative return samples is greater than the number of positive return samples, only positive return samples are allowed to enter the experience pool; The multi-microgrid distributed control module is used to use the trained strategy network model to control the switching actions in each area of the power system to form a microgrid and restore power supply to the load.

Citation Information

Patent Citations

  • Power distribution network fault recovery method and system based on multi-agent reinforcement learning

    CN119482715A