A microgrid group power control method and system based on multi-agent reinforcement learning
By applying multi-agent reinforcement learning methods in the microgrid group system and optimizing the energy storage scheduling strategy, the problem of fluctuations in the new energy output in the microgrid group system is solved, system stability is improved and the impact on the power grid is reduced.
Patent Information
- Application Number
- CN202311670616.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-12-07
AI Technical Summary
The existing technology is difficult to effectively suppress the long-term and large-scale fluctuations in the microgrid cluster system, resulting in a great impact on the safe and stable operation of the power grid.
The micronet group power control method based on multi-agent reinforcement learning is adopted. By establishing a simulation system for the micronet group, each micronet is simulated as an independent agent, and the preset multi-agent reinforcement learning algorithm is used to optimize and update the policy network of each agent to generate an energy storage scheduling strategy that is more conducive to the stable operation of the micronet group.
It reduces the power fluctuations of the microgrid group network connection points, improves the operating stability of the microgrid group system, and reduces the impact on the uncertainty on the power grid.
Smart Images

Figure CN117674233B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of renewable energy power generation technology, and in particular to a microgrid group power control method and system based on multi-agent reinforcement learning. Background Art
[0002] In recent years, microgrids, as a small source-grid-load-storage system that organically integrates renewable energy generation, loads, energy storage devices, and monitoring and protection devices, have become an effective way to absorb renewable energy generation and improve the reliability of distribution networks. Due to the shortcomings of a single microgrid, such as poor anti-disturbance capability, limited energy storage capacity, and lack of backup, an effective way is to interconnect multiple adjacent microgrid entities to form a microgrid group, and coordinate scheduling through effective optimization scheduling strategies, thereby improving the ability to adapt to the uncertainty of renewable energy and loads.
[0003] For the grid-connected microgrid system, as its power generation and load capacity increase, it brings challenges to the safe and stable operation of the power grid: when the microgrid is connected to the grid and exchanges energy with the large power grid, it is necessary to ensure the instantaneous balance of the power supply and load power of the microgrid system; during energy exchange, on the one hand, more power fed back to the grid will bring more new energy power generation benefits to the microgrid, on the other hand, the strong randomness and volatility of new energy will reduce the power quality of the grid connection point and cause energy loss of the transmission line; if the exchange power of the grid connection point fluctuates frequently, the node voltage fluctuation will also increase, causing voltage stability problems such as voltage flicker; exchange power fluctuations will also affect the operating frequency of the power grid, and in extreme cases will cause the frequency to drop, affecting the stability and safety of the large power grid. Therefore, it is necessary to coordinate and optimize the dispatching strategy of the microgrid group, and at the same time optimize the exchange power and power fluctuation between the microgrid group and the power grid, so as to increase the new energy power generation benefits of the microgrid group and reduce its impact on the safe and stable operation of the power grid.
[0004] Multi-agent reinforcement learning is to let multiple agents that make autonomous decisions and actions continuously improve their behaviors through interaction with the environment to maximize their respective objective functions. In multi-agent reinforcement learning, each agent is a learner that continuously takes actions to collect data to discover behaviors that maximize the objective function. Among them, the joint actions of all agents will affect the feedback rewards of each agent, and also affect the system environment state at the next moment, and thus affect all subsequent rewards. In cooperative multi-agent reinforcement learning, the rewards of all agents are a common system-level reward. All agents will improve their action strategies through learning to maximize the common cumulative rewards. However, the common cooperative multi-agent reinforcement learning method is not suitable for the optimization problem that comprehensively considers the economy and safety of power exchange between microgrids and large power grids.
[0005] The existing energy storage dispatching strategies generally only consider the microgrid renewable energy consumption or the maximization of new energy power generation revenue within the dispatching cycle, or smooth the new energy output on a time scale of seconds or minutes based on the characteristics of new energy output fluctuations to make it meet the technical specifications for new energy grid connection. The latter generally considers real-time regulation or smoothing based on short-term power forecasts to passively improve the smoothness of the existing power curve in terms of smoothing fluctuations in new energy output. Due to the limited capacity of energy storage devices, this type of method will lead to frequent changes in energy storage operation modes and cannot effectively suppress long-term and large-scale fluctuations in new energy output. Summary of the invention
[0006] The present invention proposes a microgrid power control method and system based on multi-agent reinforcement learning, which reduces power fluctuations at grid connection points and improves the stability of microgrid system operation by real-time scheduling of energy storage devices in the microgrid system.
[0007] In a first aspect, an embodiment of the present invention provides a microgrid group power control method based on multi-agent reinforcement learning, comprising:
[0008] Establishing a microgrid group simulation system for each microgrid in the microgrid group, wherein the microgrid group simulation system is used to simulate the operation process of the microgrid group, and each intelligent agent in the microgrid group simulation system corresponds to each microgrid in the microgrid group;
[0009] According to the training sample set in the cache area, the strategy network of each agent is iteratively updated using a preset multi-agent reinforcement learning algorithm until the cumulative number of iterative updates reaches a preset value, and the current strategy network of each agent is extracted and input into the corresponding microgrid as the energy storage scheduling strategy; wherein, in each iteration, the training sample set in the cache area is updated according to the real-time operation data of the microgrid group simulation system during operation;
[0010] The energy storage charging and discharging actions of each microgrid during operation are controlled according to the energy storage scheduling strategy, thereby performing power control.
[0011] The embodiment of the present invention provides a microgrid group power control method based on multi-agent reinforcement learning. By establishing a microgrid group simulation system for the microgrid group, each microgrid in the microgrid group is simulated as an independent agent, and the operation process of the microgrid group is simulated through the operation process of the agent, the policy network of each agent can be adjusted in the microgrid group simulation system, and then an energy storage scheduling strategy that is more conducive to the stable operation of the microgrid group is generated; during the operation of the microgrid group simulation system, real-time operation data of the microgrid group simulation system is continuously collected as training samples, and based on these training samples, a preset multi-agent reinforcement learning algorithm is used to optimize and update the policy network of each agent, thereby improving the overall benefit of the policy network; after completing one update, the microgrid group simulation system continues to operate using the updated policy network to generate training samples, and the multi-agent reinforcement learning algorithm optimizes the policy network again according to the new training samples to achieve automatic iterative update of the policy network; after stopping the iteration, a final energy storage scheduling strategy is generated according to the policy network, and the charging and discharging actions of the actual running microgrid group are controlled based on the energy storage scheduling strategy, thereby reducing the power fluctuation of the microgrid group grid connection point and improving the stability of the microgrid group system operation.
[0012] In a possible implementation, the microgrid group simulation system is established for each microgrid in the microgrid group, specifically:
[0013] Constructing the status information of each intelligent agent and the microgrid group simulation system according to the current status information of each microgrid;
[0014] According to the normal operating conditions of each microgrid, construct the operating constraints of each intelligent agent;
[0015] According to the operation constraints of each intelligent agent, construct the average exchange power function and power variance fluctuation function of the microgrid group simulation system under the long-term operation environment;
[0016] Constructing the objective function of the microgrid simulation system according to the average exchange power function and the power variance fluctuation function;
[0017] A deep neural network is used to construct and initialize the strategy network of each intelligent agent and the mean-variance value function network of the microgrid group simulation system.
[0018] The embodiment of the present invention provides a method for establishing a microgrid group simulation system for each microgrid in a microgrid group, determining state information, operation constraints, average exchange power function and power variance fluctuation function and objective function in the microgrid group simulation system, and finally determining the update direction of the policy network of each intelligent agent through the objective function, so as to ensure the efficiency of policy update during training and avoid invalid update; in the objective function, the average exchange power function corresponds to the energy feedback benefit, and the power variance fluctuation function corresponds to the influence of power fluctuation on the power grid, and the balance ratio of the two objectives can be adjusted according to the actual situation, so that the method can flexibly adjust the weight coefficient in the objective function according to the power consumption characteristics and power consumption scenarios of the region in the actual application process to cope with different scenarios; through the simulation method, the energy storage scheduling strategy of the microgrid group can be updated online without stopping the machine, without affecting the normal use and daily power demand of users; based on the simulation system, the update optimization can effectively avoid the large power fluctuation of the microgrid group caused by multiple iterations of updating the energy storage scheduling strategy, and only the final stable strategy network needs to be extracted as the energy storage scheduling strategy to ensure the stable operation of the microgrid during the training process.
[0019] Furthermore, the state information of each intelligent agent and the microgrid group simulation system is constructed according to the current state information of each microgrid, specifically:
[0020] Using Expressions To represent the state information of the i-th agent in the microgrid simulation system at time t, where and They represent the renewable energy output power, load power and energy storage charge level of the ith agent at time t, respectively, and N is the number of agents;
[0021] According to the state information expression of each intelligent agent, the system state expression of the microgrid group simulation system is determined where s t Represents the system state of the microgrid simulation system at time t.
[0022] Furthermore, the operation constraints of each intelligent agent are constructed according to the normal operation conditions of each microgrid, specifically:
[0023] According to the energy balance requirements of the microgrid group simulation system, the energy balance equation of the microgrid group simulation system is constructed:
[0024]
[0025] Among them, r(s t ,a t ) is the feedback reward of the microgrid simulation system at time t, is the total output power of the microgrid simulation system, is the energy storage charging and discharging action controlled by the strategy network at time t for agent i, Indicates energy storage discharge action;
[0026] According to the energy storage capacity and power limit of the energy storage device of each intelligent agent, the charge and discharge constraint equation of the energy storage device of each intelligent agent is constructed:
[0027]
[0028] in, is the maximum energy storage capacity of the energy storage device of intelligent agent i, P_ch^max is the maximum charging power of the energy storage device, and P_dis^max is the maximum discharging power of the energy storage device.
[0029] Furthermore, according to the operation constraints of each intelligent agent, the average exchange power function and power variance fluctuation function of the microgrid simulation system under the long-term operation environment are constructed, and the specific formula is:
[0030]
[0031] Among them, η π is the average exchange power of the microgrid simulation system under long-term operation environment, ζ π The power variance fluctuation of the microgrid simulation system under long-term operation environment, is the set of charging and discharging actions of each intelligent agent in the microgrid group simulation system at time t, and T is the running time of the microgrid group simulation system.
[0032] Furthermore, the objective function in the microgrid simulation system is determined according to the average exchange power function and the power variance fluctuation function, and the specific formula is:
[0033]
[0034] in, is the objective function in the microgrid simulation system, and β is the mean-variance weight coefficient.
[0035] In a possible implementation, the real-time operation data of the microgrid group simulation system includes the current state information of each intelligent body, the current energy storage charging and discharging action, the state information at the next moment, and the current feedback reward of the microgrid group simulation system;
[0036] The current state information of each intelligent body includes the new energy output power, load power and energy storage charge level of each intelligent body at the current moment;
[0037] The current energy storage charging and discharging actions of each intelligent body include the charging action or discharging action of the energy storage device of each intelligent body at the current moment;
[0038] The current feedback reward of the microgrid group simulation system is the total output power of the microgrid group simulation system after each intelligent agent performs the corresponding current energy storage charging and discharging action.
[0039] The embodiment of the present invention further illustrates the specific content of the real-time operation data of the microgrid group simulation system. By obtaining the current state information of each intelligent body, the current energy storage charging and discharging action, the state information at the next moment and the current feedback reward of the microgrid group simulation system, the impact of the current strategy network on the overall power and power fluctuation of the microgrid group simulation system can be calculated, and then the strategy network can be adjusted and optimized according to these data to reduce power fluctuations while improving the overall power of the microgrid group simulation system.
[0040] In a possible implementation, the training sample set in the cache area is updated according to the real-time operation data of the microgrid simulation system during operation, specifically:
[0041] Clearing the training samples in the buffer area;
[0042] During the operation of the microgrid group simulation system, each intelligent agent performs the corresponding current energy storage charging and discharging action according to the corresponding current state information and strategy network;
[0043] Acquire the current state information of each intelligent agent, the current energy storage charging and discharging action, and the state information of each intelligent agent at the next moment after executing the corresponding current energy storage charging and discharging action;
[0044] Calculate the current feedback reward of the microgrid group simulation system according to the current energy storage charging and discharging actions of each intelligent agent;
[0045] The current state information of each intelligent agent, the current energy storage charging and discharging action, the state information at the next moment and the current feedback reward of the microgrid group simulation system are combined and stored in the cache as a training sample;
[0046] The training samples are continuously acquired and stored in a buffer area until the buffer area is full, thereby completing the updating of the training sample set.
[0047] The embodiment of the present invention provides a method for updating a training sample set according to the real-time operation data of the microgrid simulation system. Each intelligent agent, as an independent decision-making individual, has its own energy storage scheduling strategy, namely, a strategy network. Each intelligent agent performs the corresponding current energy storage charging and discharging action according to its corresponding current state information and the strategy network and deduces the state information of the next moment based on the energy storage charging and discharging action to simulate the actual microgrid operation process, and calculates the corresponding current feedback reward after each charging and discharging action is executed. The operation data generated during the simulation operation is recorded and stored in a cache area as a training sample, thereby realizing the automatic generation and collection of training samples, omitting the steps of manually collecting historical data and labeling in the conventional machine learning process, optimizing the training process, and improving the optimization efficiency of the energy storage scheduling strategy.
[0048] In a possible implementation, the strategy network of each agent is iteratively updated using a preset multi-agent reinforcement learning algorithm according to the training sample set in the buffer area, wherein the specific process of each iteration is:
[0049] Estimate the estimated average exchange power of the microgrid simulation system according to the sample path in the cache area and the values of all training samples, wherein the sample path is a sample track formed by the training samples in chronological order;
[0050] Correcting the feedback reward of the microgrid group simulation system to each training sample according to the estimated average exchange power to obtain a corrected feedback reward of the microgrid group simulation system to each training sample;
[0051] Estimate the mean-variance advantage function of the microgrid group simulation system according to the modified feedback reward of the microgrid group simulation system to each training sample;
[0052] Constructing a policy network loss function according to the mean-variance advantage function, wherein the policy network loss function is used to optimize and update the policy network of each intelligent agent;
[0053] The agents are randomly sorted, and the policy networks of the agents are optimized and updated in turn according to the policy network loss function and the deep neural network.
[0054] The embodiment of the present invention provides a method for iteratively updating the policy network of each agent using a preset multi-agent reinforcement learning algorithm, constructing a policy network loss function for updating the policy network based on a training sample set, and then optimizing and updating the policy network of each agent in turn based on the policy network loss function and a deep neural network after randomly sorting each agent. Compared with the currently commonly used random dynamic programming algorithm, the neural network approximation avoids the computational difficulties caused by the large state dimension, and can learn the optimal strategy for large-scale problems when the physical model is unknown.
[0055] In the second aspect, accordingly, an embodiment of the present invention provides a microgrid group power control system based on multi-agent reinforcement learning, including a simulation module, a strategy update module and a control module;
[0056] The simulation module is used to establish a microgrid group simulation system for each microgrid in the microgrid group, wherein the microgrid group simulation system is used to simulate the operation process of the microgrid group, and each intelligent agent in the microgrid group simulation system corresponds to each microgrid in the microgrid group;
[0057] The strategy update module is used to iteratively update the strategy network of each agent using a preset multi-agent reinforcement learning algorithm according to the training sample set in the cache area, until the cumulative number of iterative updates reaches a preset value, extract the current strategy network of each agent and input it into the corresponding microgrid as the energy storage scheduling strategy; wherein, in each iteration, the training sample set in the cache area is updated according to the real-time operation data of the microgrid group simulation system during operation;
[0058] The control module is used to control the energy storage charging and discharging actions of each microgrid during operation according to the energy storage scheduling strategy, thereby performing power control.
[0059] In a possible implementation, the simulation module includes a state information building unit, an operation constraint building unit, a power function building unit, an objective function building unit and a neural network unit:
[0060] Wherein, the state information construction unit is used to construct the state information of each intelligent agent and the microgrid group simulation system according to the current state information of each microgrid;
[0061] The operation constraint construction unit is used to construct the operation constraints of each intelligent agent according to the normal operation conditions of each microgrid;
[0062] The power function construction unit is used to construct the average exchange power function and the power variance fluctuation function of the microgrid group simulation system under the long-term operation environment according to the operation constraints of each intelligent agent;
[0063] The objective function construction unit is used to construct the objective function of the microgrid simulation system according to the average exchange power function and the power variance fluctuation function;
[0064] The neural network unit is used to construct and initialize the strategy network of each intelligent agent and the mean-variance value function network of the microgrid group simulation system using a deep neural network.
[0065] Furthermore, the state information construction unit constructs the state information of each intelligent agent and the microgrid group simulation system according to the current state information of each microgrid, specifically:
[0066] Using Expressions To represent the state information of the i-th agent in the microgrid simulation system at time t, where and They represent the renewable energy output power, load power and energy storage charge level of the ith agent at time t, respectively, and N is the number of agents;
[0067] According to the state information expression of each intelligent agent, the system state expression of the microgrid group simulation system is determined where s t Represents the system state of the microgrid simulation system at time t.
[0068] Furthermore, the operation constraint construction unit constructs the operation constraints of each intelligent agent according to the normal operation conditions of each microgrid, specifically:
[0069] According to the energy balance requirements of the microgrid group simulation system, the energy balance equation of the microgrid group simulation system is constructed:
[0070]
[0071] Among them, r(s t ,a t ) is the feedback reward of the microgrid simulation system at time t, is the total output power of the microgrid simulation system, is the energy storage charging and discharging action controlled by the strategy network at time t for agent i, Indicates energy storage discharge action;
[0072] According to the energy storage capacity and power limit of the energy storage device of each intelligent agent, the charge and discharge constraint equation of the energy storage device of each intelligent agent is constructed:
[0073]
[0074] in, is the maximum energy storage capacity of the energy storage device of agent i, is the maximum charging power of the energy storage device, is the maximum discharge power of the energy storage device.
[0075] Furthermore, the power function construction unit constructs the average exchange power function and power variance fluctuation function of the microgrid simulation system under a long-term operation environment according to the operation constraints of each intelligent agent. The specific formula is:
[0076]
[0077] Among them, η π is the average exchange power of the microgrid simulation system under long-term operation environment, ζ π The power variance fluctuation of the microgrid simulation system under long-term operation environment, is the set of charging and discharging actions of each intelligent agent in the microgrid group simulation system at time t, and T is the running time of the microgrid group simulation system.
[0078] Furthermore, the objective function construction unit determines the objective function in the microgrid simulation system according to the average exchange power function and the power variance fluctuation function. The specific formula is:
[0079]
[0080] in, is the objective function in the microgrid simulation system, and β is the mean-variance weight coefficient.
[0081] In a possible implementation, the real-time operation data of the microgrid group simulation system includes the current state information of each intelligent body, the current energy storage charging and discharging action, the state information at the next moment, and the current feedback reward of the microgrid group simulation system;
[0082] The current state information of each intelligent body includes the new energy output power, load power and energy storage charge level of each intelligent body at the current moment;
[0083] The current energy storage charging and discharging actions of each intelligent body include the charging action or discharging action of the energy storage device of each intelligent body at the current moment;
[0084] The current feedback reward of the microgrid group simulation system is the total output power of the microgrid group simulation system after each intelligent agent performs the corresponding current energy storage charging and discharging action.
[0085] In a possible implementation, the training sample set in the cache area is updated according to the real-time operation data of the microgrid simulation system during operation, specifically:
[0086] Clearing the training samples in the buffer area;
[0087] During the operation of the microgrid group simulation system, each intelligent agent performs the corresponding current energy storage charging and discharging action according to the corresponding current state information and strategy network;
[0088] Acquire the current state information of each intelligent agent, the current energy storage charging and discharging action, and the state information of each intelligent agent at the next moment after executing the corresponding current energy storage charging and discharging action;
[0089] Calculate the current feedback reward of the microgrid group simulation system according to the current energy storage charging and discharging actions of each intelligent agent;
[0090] The current state information of each intelligent agent, the current energy storage charging and discharging action, the state information at the next moment and the current feedback reward of the microgrid group simulation system are combined and stored in the cache as a training sample;
[0091] The training samples are continuously acquired and stored in a buffer area until the buffer area is full, thereby completing the updating of the training sample set.
[0092] In a possible implementation, the strategy update module includes an average exchange power estimation unit, a feedback reward correction unit, an advantage function estimation unit, a loss function construction unit, and an optimization update unit;
[0093] The average exchange power estimation unit is used to estimate the estimated average exchange power of the microgrid simulation system according to the sample path in the cache area and the values of all training samples, and the sample path is a sample track formed by the training samples in chronological order;
[0094] The feedback reward correction unit is used to correct the feedback reward of the microgrid group simulation system for each training sample according to the estimated average exchange power, so as to obtain the corrected feedback reward of the microgrid group simulation system for each training sample;
[0095] The advantage function estimation unit is used to estimate the mean-variance advantage function of the microgrid group simulation system according to the modified feedback reward of the microgrid group simulation system to each training sample;
[0096] The loss function construction unit is used to construct a policy network loss function according to the mean-variance advantage function, and the policy network loss function is used to optimize and update the policy network of each intelligent agent;
[0097] The optimization and updating unit is used to randomly sort the agents and optimize and update the policy networks of the agents in turn according to the policy network loss function and the deep neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] Figure 1 : A flow chart of a microgrid group power control method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0099] Figure 2 : A schematic diagram of a scenario in which a microgrid system exchanges power with a power grid provided in an embodiment of the present invention.
[0100] Figure 3 : A flow chart of establishing a microgrid group simulation system in a microgrid group power control method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0101] Figure 4 : A schematic diagram of the process of iteratively updating the strategy network of each intelligent agent in a microgrid power control method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0102] Figure 5 : A structural schematic diagram of iteratively updating the strategy network of each agent in a microgrid power control method based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0103] Figure 6 : A structural schematic diagram of a microgrid group power control system based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0104] Figure 7 : A structural schematic diagram of a simulation module in a microgrid power control system based on multi-agent reinforcement learning provided in an embodiment of the present invention.
[0105] Figure 8 : A structural diagram of a strategy update module in a microgrid power control system based on multi-agent reinforcement learning provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0106] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0107] It should be noted that the step numbers in the text are only for the convenience of explaining the specific embodiments and are not used to limit the order in which the steps are executed.
[0108] Throughout this specification, the microgrid group described in the embodiment of the present invention is composed of a plurality of microgrids connected to the grid. The microgrid refers to a microgrid, and the ultimate source of power supply for the microgrid is new energy power generation (including photovoltaic power generation, wind power generation, etc.). For example, by configuring a photovoltaic power generation system, the direct current generated by the solar photovoltaic panel is converted into alternating current and input into the grid, and then the grid supplies power to the load. In some scenarios, the microgrid will also match some energy storage systems. When there is a surplus of photovoltaic power generation, the energy storage system will store the surplus power. When the photovoltaic power generation is insufficient or zero, the energy storage system will supply power to the microgrid to achieve stable power supply to the microgrid. This mutual dispatching of power supply needs to be coordinated by a control center (such as a supervision and management system), and the normal operation of the microgrid needs to be based on the normal operation of the control center. The energy storage scheduling strategy described in the embodiment of the present invention is the scheduling strategy of the control center for the energy storage system of the microgrid. The scenarios of the multi-microgrid system connected to the grid described in the embodiment of the present invention include but are not limited to the scenarios connected to the grid.
[0109] Embodiment 1:
[0110] like Figure 1 As shown, embodiment 1 provides a microgrid power control method based on multi-agent reinforcement learning, including steps S1 to S3:
[0111] Step S1, establishing a microgrid group simulation system for each microgrid in the microgrid group, wherein the microgrid group simulation system is used to simulate the operation process of the microgrid group, and each intelligent agent in the microgrid group simulation system corresponds to each microgrid in the microgrid group;
[0112] Step S2: According to the training sample set in the buffer area, the strategy network of each agent is iteratively updated using a preset multi-agent reinforcement learning algorithm until the cumulative number of iterative updates reaches a preset value, and the current strategy network of each agent is extracted and input into the corresponding microgrid as the energy storage scheduling strategy; wherein, in each iteration, the training sample set in the buffer area is updated according to the real-time operation data of the microgrid group simulation system during operation;
[0113] Step S3: controlling the energy storage charging and discharging actions of each microgrid during operation according to the energy storage scheduling strategy, thereby performing power control.
[0114] The embodiment of the present invention provides a microgrid group power control method based on multi-agent reinforcement learning. By establishing a microgrid group simulation system for the microgrid group, each microgrid in the microgrid group is simulated as an independent agent, and the operation process of the microgrid group is simulated through the operation process of the agent, the policy network of each agent can be adjusted in the microgrid group simulation system, and then an energy storage scheduling strategy that is more conducive to the stable operation of the microgrid group is generated; during the operation of the microgrid group simulation system, real-time operation data of the microgrid group simulation system is continuously collected as training samples, and based on these training samples, a preset multi-agent reinforcement learning algorithm is used to optimize and update the policy network of each agent, thereby improving the overall benefit of the policy network; after completing one update, the microgrid group simulation system continues to operate using the updated policy network to generate training samples, and the multi-agent reinforcement learning algorithm optimizes the policy network again according to the new training samples to achieve automatic iterative update of the policy network; after stopping the iteration, a final energy storage scheduling strategy is generated according to the policy network, and the charging and discharging actions of the actual running microgrid group are controlled based on the energy storage scheduling strategy, thereby reducing the power fluctuation of the microgrid group grid connection point and improving the stability of the microgrid group system operation.
[0115] Figure 2 The present invention is a schematic diagram of a scenario in which a microgrid system exchanges power with a power grid in an embodiment of the present invention.
[0116] In a possible implementation, in step S1, a microgrid group simulation system is established for each microgrid in the microgrid group, such as Figure 3 As shown, it includes steps S101 to S105:
[0117] Step S101, constructing the status information of each intelligent agent and the microgrid group simulation system according to the current status information of each microgrid;
[0118] Step S102: constructing operation constraints of each intelligent agent according to the normal operation conditions of each microgrid;
[0119] Step S103: constructing an average exchange power function and a power variance fluctuation function of the microgrid simulation system under a long-term operation environment according to the operation constraints of each intelligent agent;
[0120] Step S104: constructing the objective function of the microgrid simulation system according to the average exchange power function and the power variance fluctuation function;
[0121] Step S105: construct and initialize the strategy network of each intelligent agent and the mean-variance value function network of the microgrid group simulation system using a deep neural network.
[0122] The embodiment of the present invention provides a method for establishing a microgrid group simulation system for each microgrid in a microgrid group, determining state information, operation constraints, average exchange power function and power variance fluctuation function and objective function in the microgrid group simulation system, and finally determining the update direction of the policy network of each intelligent agent through the objective function, so as to ensure the efficiency of policy update during training and avoid invalid update; in the objective function, the average exchange power function corresponds to energy feedback benefit, and the power variance fluctuation function corresponds to the influence of power fluctuation on the power grid, and the balance ratio of the two objectives can be adjusted according to the actual situation, so that the method can flexibly adjust the weight coefficient in the objective function according to the power consumption characteristics and power consumption scenarios of the region in the actual application process to cope with different scenarios; through the simulation method, the energy storage scheduling strategy of the microgrid group can be updated online without stopping the machine, without affecting the normal use and daily power demand of users; based on the simulation system, the update optimization can effectively avoid the large power fluctuation of the microgrid group caused by multiple iterations of updating the energy storage scheduling strategy, and only the final stable strategy network needs to be extracted as the energy storage scheduling strategy to ensure the stable operation of the microgrid during the training process.
[0123] Furthermore, in step S101, the state information of each intelligent agent and the microgrid group simulation system is constructed according to the current state information of each microgrid, specifically:
[0124] Using Expressions To represent the state information of the i-th agent in the microgrid simulation system at time t, where and They represent the renewable energy output power, load power and energy storage charge level of the ith agent at time t, respectively, and N is the number of agents;
[0125] According to the state information expression of each intelligent agent, the system state expression of the microgrid group simulation system is determined where s t Represents the system state of the microgrid simulation system at time t.
[0126] It should be noted that the energy storage state of the ith agent at the next moment, i.e., moment t+1, is is the charging and discharging action of agent i at time t, and the charging and discharging action is determined by the strategy network corresponding to agent i The control is performed according to the current state information of the intelligent agent i, and the new energy output of the intelligent agent at time t+1 is and load level Changes according to random environmental factors such as weather.
[0127] Furthermore, in step S102, the operation constraints of each intelligent agent are constructed according to the normal operation conditions of each microgrid, specifically:
[0128] According to the energy balance requirements of the microgrid group simulation system, the energy balance equation of the microgrid group simulation system is constructed:
[0129]
[0130] Among them, r(s t ,a t ) is the feedback reward of the microgrid simulation system at time t, is the total output power of the microgrid simulation system, is the energy storage charging and discharging action controlled by the strategy network at time t for agent i, Indicates energy storage discharge action;
[0131] According to the energy storage capacity and power limit of the energy storage device of each intelligent agent, the charge and discharge constraint equation of the energy storage device of each intelligent agent is constructed:
[0132]
[0133] in, is the maximum energy storage capacity of the energy storage device of agent i, is the maximum charging power of the energy storage device, is the maximum discharge power of the energy storage device.
[0134] Furthermore, in step S103, according to the operation constraints of each intelligent agent, the average exchange power function and power variance fluctuation function of the microgrid simulation system under the long-term operation environment are constructed, and the specific formula is:
[0135]
[0136] Among them, η π is the average exchange power of the microgrid simulation system under long-term operation environment, ζ π is the power variance fluctuation of the microgrid simulation system under a long-term operation environment, and T is the operation time of the microgrid simulation system.
[0137] Further, in step S104, the objective function in the microgrid simulation system is determined according to the average exchange power function and the power variance fluctuation function, and the specific formula is:
[0138]
[0139] in, is the objective function in the microgrid simulation system, and β is the mean-variance weight coefficient.
[0140] In a possible implementation, in step S2, the real-time operation data of the microgrid group simulation system includes the current state information of each intelligent body, the current energy storage charging and discharging action, the state information at the next moment, and the current feedback reward of the microgrid group simulation system;
[0141] The current state information of each intelligent body includes the new energy output power, load power and energy storage charge level of each intelligent body at the current moment;
[0142] The current energy storage charging and discharging actions of each intelligent body include the charging action or discharging action of the energy storage device of each intelligent body at the current moment;
[0143] The current feedback reward of the microgrid group simulation system is the total output power of the microgrid group simulation system after each intelligent agent performs the corresponding current energy storage charging and discharging action.
[0144] The embodiment of the present invention further illustrates the specific content of the real-time operation data of the microgrid group simulation system. By obtaining the current state information of each intelligent body, the current energy storage charging and discharging action, the state information at the next moment and the current feedback reward of the microgrid group simulation system, the impact of the current strategy network on the overall power and power fluctuation of the microgrid group simulation system can be calculated, and then the strategy network can be adjusted and optimized according to these data to reduce power fluctuations while improving the overall power of the microgrid group simulation system.
[0145] In a possible implementation, in step S2, the training sample set in the cache area is updated according to the real-time operation data of the microgrid simulation system during operation, specifically:
[0146] Clearing the training samples in the buffer area;
[0147] During the operation of the microgrid group simulation system, each intelligent agent performs the corresponding current energy storage charging and discharging action according to the corresponding current state information and strategy network;
[0148] Acquire the current state information of each intelligent agent, the current energy storage charging and discharging action, and the state information of each intelligent agent at the next moment after executing the corresponding current energy storage charging and discharging action;
[0149] Calculate the current feedback reward of the microgrid group simulation system according to the current energy storage charging and discharging actions of each intelligent agent;
[0150] The current state information of each intelligent agent, the current energy storage charging and discharging action, the state information at the next moment and the current feedback reward of the microgrid group simulation system are combined and stored in the cache as a training sample;
[0151] The training samples are continuously acquired and stored in a buffer area until the buffer area is full, thereby completing the updating of the training sample set.
[0152] The embodiment of the present invention provides a method for updating a training sample set according to the real-time operation data of the microgrid simulation system. Each intelligent agent, as an independent decision-making individual, has its own energy storage scheduling strategy, namely, a strategy network. Each intelligent agent performs the corresponding current energy storage charging and discharging action according to its corresponding current state information and the strategy network and deduces the state information of the next moment based on the energy storage charging and discharging action to simulate the actual microgrid operation process, and calculates the corresponding current feedback reward after each charging and discharging action is executed. The operation data generated during the simulation operation is recorded and stored in a cache area as a training sample, thereby realizing the automatic generation and collection of training samples, omitting the steps of manually collecting historical data and labeling in the conventional machine learning process, optimizing the training process, and improving the optimization efficiency of the energy storage scheduling strategy.
[0153] In a possible implementation, the strategy network of each agent is iteratively updated using a preset multi-agent reinforcement learning algorithm according to the training sample set in the buffer area, such as Figure 4 As shown, the specific process of each iteration includes steps S201 to S205:
[0154] Step S201, estimating the estimated average exchange power of the microgrid simulation system according to the sample path in the cache area and the values of all training samples, wherein the sample path is a sample track formed by the training samples in chronological order;
[0155] Step S202: correcting the feedback reward of the microgrid group simulation system for each training sample according to the estimated average exchange power to obtain a corrected feedback reward of the microgrid group simulation system for each training sample;
[0156] Step S203, estimating the mean-variance advantage function of the microgrid group simulation system according to the modified feedback reward of the microgrid group simulation system for each training sample;
[0157] Step S204: constructing a policy network loss function according to the mean-variance advantage function, wherein the policy network loss function is used to optimize and update the policy network of each agent;
[0158] Step S205: randomly sort the agents, and optimize and update the policy networks of the agents in turn according to the policy network loss function and the deep neural network.
[0159] The embodiment of the present invention provides a method for iteratively updating the policy network of each agent using a preset multi-agent reinforcement learning algorithm, constructing a policy network loss function for updating the policy network based on a training sample set, and then optimizing and updating the policy network of each agent in turn based on the policy network loss function and a deep neural network after randomly sorting each agent. Compared with the currently commonly used random dynamic programming algorithm, the neural network approximation avoids the computational difficulties caused by the large state dimension, and can learn the optimal strategy for large-scale problems when the physical model is unknown.
[0160] Preferably, in step S201, the average current feedback reward in each training sample is averaged to estimate the average exchange power η of the microgrid simulation system. π .
[0161] In step S202, the formula r f,π (s t ,a t )=r(s t ,a t )-β(r(s t ,a t )-η π ) 2 Modify the feedback reward of the microgrid simulation system for each training sample, where r f,π (s t ,a t ) is the corrected feedback reward of the training sample stored at time t, r(s t ,a t ) is the feedback reward of the training sample stored at time t before correction.
[0162] In step S203, first use the formula Estimate the objective function of the microgrid simulation system, where W is the number of training samples in the training sample set; then use the formula Estimate the mean-variance function corresponding to different training samples in is the mean-variance value function network of the microgrid group simulation system constructed in step S105; at the kth iteration, the loss function based on which the mean-variance value function network of the microgrid group simulation system is updated is
[0163]
[0164] in, v is a hyperparameter that needs to be manually adjusted according to the complexity of the problem; finally, the mean-variance advantage function of the microgrid simulation system is estimated according to the following formula: in
[0165] In step S204 and step S205, the strategy network loss function corresponding to each agent is constructed according to the mean-variance advantage function, and then the agents are randomly sorted. The agents after random sorting are numbered as j. 1 ,…,j N Finally, according to the corresponding strategy network loss function, agent j is 1 ,…,j N The policy network of agent j is updated n The policy network loss function expression is:
[0166]
[0167] in and They are respectively 1 ,…,j n-1 The policy network before and after the update. In an embodiment of the present invention, each agent has its own independent policy network. After collecting a series of samples, the policy network is updated in sequence according to the randomly generated arrangement order. The number of neurons in the input layer of the policy network corresponds to the dimension of the system state, the activation function used in the hidden layer of the deep neural network is the relu function, the number of neurons in the output layer is 1, corresponding to the energy storage output level, and the activation function used in the output layer is the softmax function. The common value function network uses a fully linked network.
[0168] Figure 5 This is a schematic diagram of a structure for iteratively updating the strategy network of each agent in an embodiment of the present invention. It should be noted that the unit names and connection order of each unit indicated in the figure are only a specific embodiment of the present invention and are not intended to limit the scope of protection of the present invention. Figure 5 As shown, at each iterative update, the central controller of the system sends a corresponding strategy network to each microgrid strategy unit (i.e., each intelligent agent in the microgrid group simulation system), and each microgrid strategy unit performs normal operation according to the sent strategy, thereby generating real-time operation data; the storage unit collects the real-time operation data generated by each microgrid strategy unit and stores it as a training sample; the processing unit updates the strategy network corresponding to each microgrid strategy unit in the central controller according to the training sample set and the preset strategy update formula to complete this iterative update.
[0169] Embodiment 2:
[0170] Correspondingly, such as Figure 6 As shown, an embodiment of the present invention provides a microgrid power control system based on multi-agent reinforcement learning, including a simulation module 10, a strategy update module 20 and a control module 30;
[0171] The simulation module 10 is used to establish a microgrid group simulation system for each microgrid in the microgrid group, wherein the microgrid group simulation system is used to simulate the operation process of the microgrid group, and each intelligent agent in the microgrid group simulation system corresponds to each microgrid in the microgrid group;
[0172] The strategy update module 20 is used to iteratively update the strategy network of each agent using a preset multi-agent reinforcement learning algorithm according to the training sample set in the cache area, until the cumulative number of iterative updates reaches a preset value, extract the current strategy network of each agent and input it into the corresponding microgrid as the energy storage scheduling strategy; wherein, in each iteration, the training sample set in the cache area is updated according to the real-time operation data of the microgrid group simulation system during operation;
[0173] The control module 30 is used to control the energy storage charging and discharging actions of each microgrid during operation according to the energy storage scheduling strategy, thereby performing power control.
[0174] In one possible implementation, Figure 7 As shown, the simulation module 10 includes a state information construction unit 101, an operation constraint construction unit 102, a power function construction unit 103, an objective function construction unit 104 and a neural network unit 105:
[0175] The state information construction unit 101 is used to construct the state information of each agent and the microgrid group simulation system according to the current state information of each microgrid;
[0176] The operation constraint construction unit 102 is used to construct the operation constraints of each intelligent agent according to the normal operation conditions of each microgrid;
[0177] The power function construction unit 103 is used to construct the average exchange power function and the power variance fluctuation function of the microgrid group simulation system under the long-term operation environment according to the operation constraints of each intelligent agent;
[0178] The objective function construction unit 104 is used to construct the objective function of the microgrid simulation system according to the average exchange power function and the power variance fluctuation function;
[0179] The neural network unit 105 is used to construct and initialize the strategy network of each intelligent agent and the mean-variance value function network of the microgrid group simulation system using a deep neural network.
[0180] Furthermore, the state information construction unit 101 constructs the state information of each agent and the microgrid group simulation system according to the current state information of each microgrid, specifically:
[0181] Using Expressions To represent the state information of the i-th agent in the microgrid simulation system at time t, where and They represent the renewable energy output power, load power and energy storage charge level of the ith agent at time t, respectively, and N is the number of agents;
[0182] According to the state information expression of each intelligent agent, the system state expression of the microgrid group simulation system is determined where s t Represents the system state of the microgrid simulation system at time t.
[0183] Furthermore, the operation constraint construction unit 102 constructs the operation constraints of each intelligent agent according to the normal operation conditions of each microgrid, specifically:
[0184] According to the energy balance requirements of the microgrid group simulation system, the energy balance equation of the microgrid group simulation system is constructed:
[0185]
[0186] Among them, r(s t ,a t ) is the feedback reward of the microgrid simulation system at time t, is the total output power of the microgrid simulation system, is the energy storage charging and discharging action controlled by the strategy network at time t for agent i, Indicates energy storage discharge action;
[0187] According to the energy storage capacity and power limit of the energy storage device of each intelligent agent, the charge and discharge constraint equation of the energy storage device of each intelligent agent is constructed:
[0188]
[0189] in, is the maximum energy storage capacity of the energy storage device of intelligent agent i, P_ch^max is the maximum charging power of the energy storage device, and P_dis^max is the maximum discharging power of the energy storage device.
[0190] Furthermore, the power function construction unit 103 constructs the average exchange power function and power variance fluctuation function of the microgrid simulation system under a long-term operation environment according to the operation constraints of each intelligent agent. The specific formula is:
[0191]
[0192] Among them, η π is the average exchange power of the microgrid simulation system under long-term operation environment, ζ π The power variance fluctuation of the microgrid simulation system under long-term operation environment, is the set of charging and discharging actions of each intelligent agent in the microgrid group simulation system at time t, and T is the running time of the microgrid group simulation system.
[0193] Furthermore, the objective function construction unit 104 determines the objective function in the microgrid simulation system according to the average exchange power function and the power variance fluctuation function. The specific formula is:
[0194]
[0195] in, is the objective function in the microgrid simulation system, and β is the mean-variance weight coefficient.
[0196] In a possible implementation, the real-time operation data of the microgrid group simulation system includes the current state information of each intelligent body, the current energy storage charging and discharging action, the state information at the next moment, and the current feedback reward of the microgrid group simulation system;
[0197] The current state information of each intelligent body includes the new energy output power, load power and energy storage charge level of each intelligent body at the current moment;
[0198] The current energy storage charging and discharging actions of each intelligent body include the charging action or discharging action of the energy storage device of each intelligent body at the current moment;
[0199] The current feedback reward of the microgrid group simulation system is the total output power of the microgrid group simulation system after each intelligent agent performs the corresponding current energy storage charging and discharging action.
[0200] In a possible implementation, the training sample set in the cache area is updated according to the real-time operation data of the microgrid simulation system during operation, specifically:
[0201] Clearing the training samples in the buffer area;
[0202] During the operation of the microgrid group simulation system, each intelligent agent performs the corresponding current energy storage charging and discharging action according to the corresponding current state information and strategy network;
[0203] Acquire the current state information of each intelligent agent, the current energy storage charging and discharging action, and the state information of each intelligent agent at the next moment after executing the corresponding current energy storage charging and discharging action;
[0204] Calculate the current feedback reward of the microgrid group simulation system according to the current energy storage charging and discharging actions of each intelligent agent;
[0205] The current state information of each intelligent agent, the current energy storage charging and discharging action, the state information at the next moment and the current feedback reward of the microgrid group simulation system are combined and stored in the cache as a training sample;
[0206] The training samples are continuously acquired and stored in a buffer area until the buffer area is full, thereby completing the updating of the training sample set.
[0207] In one possible implementation, Figure 8 As shown, the strategy updating module 20 includes an average exchange power estimation unit 201, a feedback reward correction unit 202, an advantage function estimation unit 203, a loss function construction unit 204 and an optimization updating unit 205;
[0208] The average exchange power estimation unit 201 is used to estimate the estimated average exchange power of the microgrid simulation system according to the sample path in the cache area and the values of all training samples, and the sample path is a sample track formed by the training samples in chronological order;
[0209] The feedback reward correction unit 202 is used to correct the feedback reward of the microgrid group simulation system for each training sample according to the estimated average exchange power, and obtain the corrected feedback reward of the microgrid group simulation system for each training sample;
[0210] The advantage function estimation unit 203 is used to estimate the mean-variance advantage function of the microgrid group simulation system according to the modified feedback reward of the microgrid group simulation system for each training sample;
[0211] The loss function construction unit 204 is used to construct a policy network loss function according to the mean-variance advantage function, and the policy network loss function is used to optimize and update the policy network of each intelligent agent;
[0212] The optimization and updating unit 205 is used to randomly sort the agents and optimize and update the policy networks of the agents in turn according to the policy network loss function and the deep neural network.
[0213] The embodiment of the present invention provides a microgrid power control system based on multi-agent reinforcement learning. By establishing a microgrid simulation system for the microgrid, each microgrid in the microgrid is simulated as an independent agent. The operation process of the microgrid is simulated through the operation process of the agent. The policy network of each agent can be adjusted in the microgrid simulation system, thereby generating an energy storage scheduling strategy that is more conducive to the stable operation of the microgrid. During the operation of the microgrid simulation system, real-time operation data of the microgrid simulation system is continuously collected as training samples, and based on these training samples, a preset multi-agent reinforcement learning algorithm is used to optimize and update the policy network of each agent, thereby improving the overall benefit of the policy network. After completing one update, the microgrid simulation system continues to operate using the updated policy network to generate training samples, and the multi-agent reinforcement learning algorithm optimizes the policy network again according to the new training samples to achieve automatic iterative update of the policy network. After stopping the iteration, the final energy storage scheduling strategy is generated according to the policy network, and the charging and discharging actions of the actual running microgrid are controlled based on the energy storage scheduling strategy, thereby reducing the power fluctuation of the microgrid grid connection point and improving the stability of the microgrid system operation.
[0214] The more detailed working principle and step flow of this embodiment can refer to, but are not limited to, the relevant records of Embodiment 1.
[0215] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. It is particularly pointed out that for those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of protection of the present invention.
Claims
1. A microgrid power control method based on multi-agent reinforcement learning, It is characterized in that include: A microgrid group simulation system is established for each microgrid in the microgrid group, specifically: based on the current state information of each microgrid, the state information of each intelligent agent and the microgrid group simulation system is constructed; based on the normal operating conditions of each microgrid, the operation constraints of each intelligent agent are constructed; based on the operation constraints of each intelligent agent, the average exchange power function and power variance fluctuation function of the microgrid group simulation system under a long-term operating environment are constructed; based on the average exchange power function and the power variance fluctuation function, the objective function of the microgrid group simulation system is constructed; using a deep neural network to construct and initialize the strategy network of each intelligent agent and the mean-variance value function network of the microgrid group simulation system, wherein the microgrid group simulation system is used to simulate the operation process of the microgrid group, and each intelligent agent in the microgrid group simulation system corresponds to each microgrid in the microgrid group; According to the training sample set in the cache area, the strategy network of each agent is iteratively updated using a preset multi-agent reinforcement learning algorithm until the cumulative number of iterative updates reaches a preset value, and the current strategy network of each agent is extracted and input into the corresponding microgrid as the energy storage scheduling strategy; wherein, in each iteration, the training sample set in the cache area is updated according to the real-time operation data of the microgrid group simulation system during operation; The energy storage charging and discharging actions of each microgrid during operation are controlled according to the energy storage scheduling strategy, thereby performing power control.
2. A microgrid power control method based on multi-agent reinforcement learning as claimed in claim 1, It is characterized in that The state information of each intelligent agent and the microgrid group simulation system is constructed according to the current state information of each microgrid, specifically: Using Expressions i∈{1,…,N} represents the state information of the i-th agent in the microgrid simulation system at time t, where and They represent the renewable energy output power, load power and energy storage charge level of the ith agent at time t, respectively, and N is the number of agents; According to the state information expression of each intelligent agent, the system state expression of the microgrid group simulation system is determined where s t Represents the system state of the microgrid simulation system at time t.
3. A microgrid power control method based on multi-agent reinforcement learning as claimed in claim 1, It is characterized in that The operation constraints of each intelligent agent are constructed according to the normal operation conditions of each microgrid, specifically: According to the energy balance requirements of the microgrid group simulation system, the energy balance equation of the microgrid group simulation system is constructed: Among them, r(s t ,a t ) is the feedback reward of the microgrid simulation system at time t, is the total output power of the microgrid simulation system, is the energy storage charging and discharging action controlled by the strategy network at time t for agent i, Indicates energy storage discharge action; According to the energy storage capacity and power limit of the energy storage device of each intelligent agent, the charge and discharge constraint equation of the energy storage device of each intelligent agent is constructed: in, is the maximum energy storage capacity of the energy storage device of agent i, is the maximum charging power of the energy storage device, is the maximum discharge power of the energy storage device, represents the energy storage charge level of the ith agent at time t.
4. A microgrid power control method based on multi-agent reinforcement learning as claimed in claim 1, It is characterized in that According to the operation constraints of each intelligent agent, the average exchange power function and power variance fluctuation function of the microgrid simulation system under the long-term operation environment are constructed. The specific formula is: Among them, η π is the average exchange power of the microgrid simulation system in the long-term operation environment when the joint strategy π is adopted, ζ π is the power variance fluctuation of the microgrid simulation system under the long-term operation environment when the joint strategy π is adopted, is the set of charging and discharging actions of each agent in the microgrid simulation system at time t, T is the running time of the microgrid simulation system, Take the expected value of the random path guided by the joint strategy π.
5. A microgrid power control method based on multi-agent reinforcement learning as claimed in claim 1, It is characterized in that The objective function in the microgrid simulation system is determined according to the average exchange power function and the power variance fluctuation function, and the specific formula is: in, is the objective function in the microgrid simulation system, β is the mean-variance weight coefficient, η π is the average exchange power of the microgrid simulation system in the long-term operation environment when the joint strategy π is adopted, ζ π The power variance fluctuation of the microgrid simulation system in a long-term operation environment when the joint strategy π is adopted.
6. A microgrid power control method based on multi-agent reinforcement learning as claimed in claim 1, It is characterized in that The real-time operation data of the microgrid simulation system includes the current state information of each intelligent body, the current energy storage charging and discharging action, the state information at the next moment, and the current feedback reward of the microgrid simulation system; The current state information of each intelligent body includes the new energy output power, load power and energy storage charge level of each intelligent body at the current moment; The current energy storage charging and discharging actions of each intelligent body include the charging action or discharging action of the energy storage device of each intelligent body at the current moment; The current feedback reward of the microgrid group simulation system is the total output power of the microgrid group simulation system after each intelligent agent performs the corresponding current energy storage charging and discharging action.
7. A microgrid power control method based on multi-agent reinforcement learning as claimed in claim 1, It is characterized in that The training sample set in the buffer area is updated according to the real-time operation data of the microgrid simulation system during operation, specifically: Clearing the training samples in the buffer area; During the operation of the microgrid group simulation system, each intelligent agent performs the corresponding current energy storage charging and discharging action according to the corresponding current state information and strategy network; Acquire the current state information of each intelligent agent, the current energy storage charging and discharging action, and the state information of each intelligent agent at the next moment after executing the corresponding current energy storage charging and discharging action; Calculate the current feedback reward of the microgrid group simulation system according to the current energy storage charging and discharging actions of each intelligent agent; The current state information of each intelligent agent, the current energy storage charging and discharging action, the state information at the next moment and the current feedback reward of the microgrid group simulation system are combined and stored in the cache as a training sample; The training samples are continuously acquired and stored in a buffer area until the buffer area is full, thereby completing the updating of the training sample set.
8. A microgrid power control method based on multi-agent reinforcement learning as claimed in claim 1, It is characterized in that According to the training sample set in the buffer area, the strategy network of each agent is iteratively updated using a preset multi-agent reinforcement learning algorithm, wherein the specific process of each iteration is: Estimate the estimated average exchange power of the microgrid simulation system according to the sample path in the cache area and the values of all training samples, wherein the sample path is a sample track formed by the training samples in chronological order; Correcting the feedback reward of the microgrid group simulation system to each training sample according to the estimated average exchange power to obtain a corrected feedback reward of the microgrid group simulation system to each training sample; Estimate the mean-variance advantage function of the microgrid group simulation system according to the modified feedback reward of the microgrid group simulation system to each training sample; Constructing a policy network loss function according to the mean-variance advantage function, wherein the policy network loss function is used to optimize and update the policy network of each intelligent agent; The agents are randomly sorted, and the policy networks of the agents are optimized and updated in turn according to the policy network loss function and the deep neural network.
9. A microgrid power control system based on multi-agent reinforcement learning. It is characterized in that It includes simulation module, strategy update module and control module; The simulation module is used to establish a microgrid group simulation system for each microgrid in the microgrid group, wherein the microgrid group simulation system is used to simulate the operation process of the microgrid group, and each intelligent agent in the microgrid group simulation system corresponds to each microgrid in the microgrid group; The simulation module includes a state information construction unit, an operation constraint construction unit, a power function construction unit, an objective function construction unit and a neural network unit: wherein the state information construction unit is used to construct the state information of each intelligent agent and the microgrid group simulation system according to the current state information of each microgrid; the operation constraint construction unit is used to construct the operation constraints of each intelligent agent according to the normal operating conditions of each microgrid; the power function construction unit is used to construct the average exchange power function and power variance fluctuation function of the microgrid group simulation system under a long-term operating environment according to the operation constraints of each intelligent agent; the objective function construction unit is used to construct the objective function of the microgrid group simulation system according to the average exchange power function and the power variance fluctuation function; the neural network unit is used to construct and initialize the strategy network of each intelligent agent and the mean-variance value function network of the microgrid group simulation system using a deep neural network; The strategy update module is used to iteratively update the strategy network of each agent using a preset multi-agent reinforcement learning algorithm according to the training sample set in the cache area, until the cumulative number of iterative updates reaches a preset value, extract the current strategy network of each agent and input it into the corresponding microgrid as the energy storage scheduling strategy; wherein, in each iteration, the training sample set in the cache area is updated according to the real-time operation data of the microgrid group simulation system during operation; The control module is used to control the energy storage charging and discharging actions of each microgrid during operation according to the energy storage scheduling strategy, thereby performing power control.
10. A microgrid power control system based on multi-agent reinforcement learning as claimed in claim 9, It is characterized in that The state information construction unit constructs the state information of each intelligent agent and the microgrid group simulation system according to the current state information of each microgrid, specifically: Using Expressions i∈{1,…,N} represents the state information of the i-th agent in the microgrid simulation system at time t, where and They represent the renewable energy output power, load power and energy storage charge level of the ith agent at time t, respectively, and N is the number of agents; According to the state information expression of each intelligent agent, the system state expression of the microgrid group simulation system is determined where s t Represents the system state of the microgrid simulation system at time t.
11. A microgrid power control system based on multi-agent reinforcement learning as claimed in claim 9, It is characterized in that The operation constraint construction unit constructs the operation constraints of each intelligent agent according to the normal operation conditions of each microgrid, specifically: According to the energy balance requirements of the microgrid group simulation system, the energy balance equation of the microgrid group simulation system is constructed: Among them, r(s t ,a t ) is the feedback reward of the microgrid simulation system at time t, is the total output power of the microgrid simulation system, is the energy storage charging and discharging action controlled by the strategy network at time t for agent i, Indicates energy storage discharge action; According to the energy storage capacity and power limit of the energy storage device of each intelligent agent, the charge and discharge constraint equation of the energy storage device of each intelligent agent is constructed: in, is the maximum energy storage capacity of the energy storage device of agent i, P_ch^max is the maximum charging power of the energy storage device, P_dis^max is the maximum discharge power of the energy storage device, represents the energy storage charge level of the ith agent at time t.
12. A microgrid power control system based on multi-agent reinforcement learning as claimed in claim 9, It is characterized in that The power function construction unit constructs the average exchange power function and power variance fluctuation function of the microgrid simulation system under a long-term operation environment according to the operation constraints of each intelligent agent. The specific formula is: Among them, η π is the average exchange power of the microgrid simulation system in the long-term operation environment when the joint strategy π is adopted, ζ π is the power variance fluctuation of the microgrid simulation system under the long-term operation environment when the joint strategy π is adopted, is the set of charging and discharging actions of each agent in the microgrid simulation system at time t, T is the running time of the microgrid simulation system, Take the expected value of the random path guided by the joint strategy π.
13. A microgrid power control system based on multi-agent reinforcement learning as claimed in claim 9, It is characterized in that The objective function construction unit determines the objective function in the microgrid simulation system according to the average exchange power function and the power variance fluctuation function. The specific formula is: in, is the objective function in the microgrid simulation system, β is the mean-variance weight coefficient, η π is the average exchange power of the microgrid simulation system in the long-term operation environment when the joint strategy π is adopted, ζ π The power variance fluctuation of the microgrid simulation system in a long-term operation environment when the joint strategy π is adopted.
14. A microgrid power control system based on multi-agent reinforcement learning as claimed in claim 9, It is characterized in that The real-time operation data of the microgrid simulation system includes the current state information of each intelligent body, the current energy storage charging and discharging action, the state information at the next moment, and the current feedback reward of the microgrid simulation system; The current state information of each intelligent body includes the new energy output power, load power and energy storage charge level of each intelligent body at the current moment; The current energy storage charging and discharging actions of each intelligent body include the charging action or discharging action of the energy storage device of each intelligent body at the current moment; The current feedback reward of the microgrid group simulation system is the total output power of the microgrid group simulation system after each intelligent agent performs the corresponding current energy storage charging and discharging action.
15. A microgrid power control system based on multi-agent reinforcement learning as claimed in claim 9, It is characterized in that The training sample set in the buffer area is updated according to the real-time operation data of the microgrid simulation system during operation, specifically: Clearing the training samples in the buffer area; During the operation of the microgrid group simulation system, each intelligent agent performs the corresponding current energy storage charging and discharging action according to the corresponding current state information and strategy network; Acquire the current state information of each intelligent agent, the current energy storage charging and discharging action, and the state information of each intelligent agent at the next moment after executing the corresponding current energy storage charging and discharging action; Calculate the current feedback reward of the microgrid group simulation system according to the current energy storage charging and discharging actions of each intelligent agent; The current state information of each intelligent agent, the current energy storage charging and discharging action, the state information at the next moment and the current feedback reward of the microgrid group simulation system are combined and stored in the cache as a training sample; The training samples are continuously acquired and stored in a buffer area until the buffer area is full, thereby completing the updating of the training sample set.
16. A microgrid power control system based on multi-agent reinforcement learning as claimed in claim 9, It is characterized in that The strategy update module includes an average exchange power estimation unit, a feedback reward correction unit, an advantage function estimation unit, a loss function construction unit and an optimization update unit; The average exchange power estimation unit is used to estimate the estimated average exchange power of the microgrid simulation system according to the sample path in the cache area and the values of all training samples, and the sample path is a sample track formed by the training samples in chronological order; The feedback reward correction unit is used to correct the feedback reward of the microgrid group simulation system for each training sample according to the estimated average exchange power, so as to obtain the corrected feedback reward of the microgrid group simulation system for each training sample; The advantage function estimation unit is used to estimate the mean-variance advantage function of the microgrid group simulation system according to the modified feedback reward of the microgrid group simulation system to each training sample; The loss function construction unit is used to construct a policy network loss function according to the mean-variance advantage function, and the policy network loss function is used to optimize and update the policy network of each intelligent agent; The optimization and updating unit is used to randomly sort the agents and optimize and update the policy networks of the agents in turn according to the policy network loss function and the deep neural network.
Citation Information
Patent Citations
Multi-microgrid power distribution system distributed scheduling method based on multi-agent reinforcement learning
CN113780622A