Multi-stage Multi-energy Management Method for Multiple Microgrids Based on Deep Reinforcement Learning

Through a multi-stage multi-energy management method based on deep reinforcement learning, a Markov decision-making process model is built and an intelligent network is trained, and a complex problem of energy management of multiple microgrids is solved, efficient coordination and real-time adjustments are achieved inside and outside the microgrid, and the energy management effect is improved.

CN120198253BActive Publication Date: 2025-07-18UESTC (SHENZHEN) ADVANCED RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510679360.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-07-18
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

In the prior art, the energy management of multiple microgrids is complex, resulting in poor management effects and it is difficult to efficiently and stably utilize renewable energy.

Method used

Using a multi-stage multi-energy management method based on deep reinforcement learning, a Markov decision-making process model is constructed, the status, action and reward values of the microgrid are obtained in stages, the network structure of the agent is trained, and the energy management strategy is obtained.

Benefits of technology

It realizes efficient real-time collaboration within and between multiple microgrids, improves energy management effects, and can adjust strategies in real time according to the regulatory results at different stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198253B_ABST
    Figure CN120198253B_ABST
Patent Text Reader

Abstract

The present application discloses a multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning, including: constructing a Markov decision process model including multiple agents according to the state space and action space of the multiple microgrids; based on the Markov decision process model, respectively obtaining the states, actions of each microgrid in multiple energy regulation stages in the microgrid environment, and the reward value of each microgrid at each time step, training the agent network structure of the agent according to the states, actions and reward values until convergence to obtain the final agent network structure, and the agent network structure is used to obtain the energy management strategy of the microgrid in the actual execution process. The present application enables multiple microgrids to interact with the environment independently in stages during the multi-stage energy regulation process, and adjusts the energy management strategy in real time according to the regulation results of different stages, so as to achieve efficient real-time coordination within and between microgrids.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent microgrids, and particularly to a multi-stage and multi-energy management method for multiple microgrids based on deep reinforcement learning. Background Art

[0002] With the gradual development of the global consensus on energy conservation and emission reduction in various countries, the installed capacity of global renewable energy has increased significantly, and renewable energy has become a new force for China to ensure power supply. However, the large-scale substitution of renewable energy has brought challenges to the operation of the power system. How to use renewable energy more efficiently, stably, and reliably is the key to promoting the green and low-carbon transformation of energy in the current power system. The wide application of microgrids mainly based on renewable energy such as solar energy and wind energy provides higher economic benefits and flexibility for the smart grid. However, problems such as the unique local high-share and uncertainty of renewable energy, the complex coupling characteristics of various energy management technologies, and the uneven energy distribution among multiple microgrids make the multi-energy management of multiple microgrids face many difficulties and challenges. How to design an energy management method that takes into account both multi-energy regulation within the microgrid and multi-energy coordination among microgrids has become an urgent problem to be solved. Summary of the Invention

[0003] The purpose of this application is to provide a multi-stage and multi-energy management method for multiple microgrids based on deep reinforcement learning to solve the technical problem that the energy management of multiple microgrids in the same area is complex and the energy management effect is poor in the prior art. The many technical effects that can be produced by the preferred technical solutions provided in this application are described in detail below.

[0004] To achieve the above purpose, this application provides the following technical solutions:

[0005] In a first aspect, a multi-stage and multi-energy management method for multiple microgrids based on deep reinforcement learning provided by the present application includes: constructing a Markov decision process model including multiple agents according to the state space and action space of each microgrid, where the state space includes observable states of each microgrid being in the IDR (Integrated Demand Response) stage, P2P (Peer-to-Peer) trading stage, and EC (Energy Conversion) stage respectively, and the action space includes executable actions of each microgrid being in the IDR stage, the P2P trading stage, and the EC stage respectively; based on the Markov decision process model, obtaining the states, actions, and reward values of each microgrid in multiple energy regulation stages in the microgrid environment at each time step, and training the agent network structure of the agent to convergence according to the states, actions, and reward values to obtain a final agent network structure, where the agent network structure is used to obtain the energy management strategy of the microgrid in the actual execution process.

[0006] In some embodiments, each microgrid includes an IDR agent, a P2P agent, and an EC agent. The obtaining of the states, actions, and reward values of each microgrid in multiple energy regulation stages in the microgrid environment at each time step includes: each microgrid obtains its own IDR state from the microgrid environment, and processes the IDR state through its own IDR agent to obtain the action in the IDR stage; each microgrid executes the action in the IDR stage in the microgrid environment to obtain the dynamic adjustment result in the IDR stage and its own P2P trading state, and processes the P2P trading state through its own P2P agent to obtain the action in the P2P trading stage; each microgrid executes the action in the P2P trading stage in the microgrid environment to obtain the dynamic adjustment result in the P2P trading stage and its own EC state, and processes the EC state through its own EC agent to obtain the action in the EC stage, and each microgrid executes the action in the EC stage in the microgrid environment to obtain the dynamic adjustment result in the EC stage; according to the dynamic adjustment result in the IDR stage, the dynamic adjustment result in the P2P trading stage, and the dynamic adjustment result in the EC stage, obtaining the reward value of each microgrid at the current time step and the IDR state at the next time step from the microgrid environment.

[0007] In some embodiments, the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning further includes: selecting the final action in the IDR stage from the actions in the IDR stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result and the P2P trading status in the IDR stage according to the final action in the IDR stage; and / or selecting the final action in the P2P trading stage from the actions in the P2P trading stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result and the EC status in the P2P trading stage according to the final action in the P2P trading stage; and / or selecting the final action in the EC stage from the actions in the EC stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result in the EC stage according to the final action in the EC stage.

[0008] In some embodiments, training the agent network structure of the agent based on the state, the action, and the reward value until convergence to obtain the final agent network structure includes: sampling at least one training batch, and sequentially inputting the at least one training batch into the actor target network and the critic target network of the agent network structure to obtain the target Q value, where the training batch at least includes the state, action, reward value of the microgrid, and the state corresponding to the next time step; calculating the loss value of the actor network and the loss value of the critic network of the agent network structure based on the target Q value, and updating the parameters of the actor network and the parameters of the critic network.

[0009] In some embodiments, the IDR state includes the time step, electrical energy demand, thermal energy demand, indoor temperature, outdoor temperature, solar power generation, grid electricity price, natural gas price, cold energy storage state, and thermal energy storage state. The P2P trading state includes the time step, the solar power generation, the grid electricity price, the natural gas price, the cold energy storage state, the thermal energy storage state, the demand result in the IDR stage, the P2P trading price of electrical energy, and the P2P trading price of heating energy. The EC state includes the time step, the solar power generation, the grid electricity price, the natural gas price, the cold energy storage state, the thermal energy storage state, the demand result in the IDR stage, the actual energy quantity of electrical energy trading, and the actual energy quantity of thermal energy trading.

[0010] In some embodiments, the IDR action space of the microgrid includes the electricity curtailment rate, the heat curtailment rate, the electricity transfer state, the heat transfer state, and the desired room temperature. The P2P trading action space of the microgrid includes the P2P electricity price and electricity trading volume, and the P2P heat price and heat trading volume. The EC action space of the microgrid includes the GT utilization rate, the electric-thermal-cool ratio, the heat storage charge-discharge ratio, and the cold storage charge-discharge ratio.

[0011] In some embodiments, the reward space of the microgrid includes transaction costs with the main power grid and the natural gas grid, P2P energy transaction costs, user discomfort penalties for demand reduction, user temperature discomfort penalties, transferable demand incentives, and GT (Gas Turbine) incentives.

[0012] In a second aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, it implements the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described above.

[0013] In a third aspect, the present application provides a processing device, including: one or more processors; a memory for storing one or more computer programs, and one or more of the processors are used to execute the one or more computer programs stored in the memory, so that one or more of the processors execute the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described above.

[0014] In a fourth aspect, the present application provides a computer program product, which is stored on a data carrier and is designed to execute the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described above.

[0015] Implementing one of the above technical solutions of the present application has the following advantages or beneficial effects: In the present application, the state space of each microgrid is divided into the IDR stage, the P2P transaction stage, and the EC stage, and then the states, actions, and reward values of each microgrid in each stage are obtained for training the intelligent agent network structure to obtain the energy management strategy of the microgrid. In this case, it is possible to realize the whole process multi-energy overall coordination of the internal energy supply and demand process of multiple microgrids and the energy transactions between each microgrid, enabling multiple microgrids to interact with the environment independently in different stages during the multi-stage energy regulation process, and being able to adjust the energy management strategy in real time according to the regulation results of different stages, achieving efficient real-time coordination within and between microgrids, thereby greatly improving the effect of energy management. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0017] Figure 1 is a flowchart of the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to the embodiments of the present application;

[0018] Figure 2 is an interaction schematic diagram of the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to an embodiment of the present application;

[0019] Figure 3 is a structural block diagram of a processing device according to an embodiment of the present application.

[0020] In the figure: 1. Processing device; 10. Memory; 11. Processor. Detailed implementation manners

[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, various exemplary embodiments to be described below will refer to the corresponding drawings, which form a part of the exemplary embodiments and describe various exemplary embodiments that may be adopted to implement the present application. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. It should be understood that they are merely examples of processes, methods, devices, etc. consistent with some aspects of the present application disclosed in detail in the appended claims. Other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present application.

[0022] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", etc. indicate the orientation or positional relationship based on the orientation or position shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the indicated elements must have a specific orientation, be constructed and operated in a specific orientation. The terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. The meaning of the term "plurality" is two or more. The terms "connected" and "connected" should be understood in a broad sense. For example, they may be fixedly connected, detachably connected, integrally connected, mechanically connected, electrically connected, communicatively connected, directly connected, indirectly connected through an intermediate medium, and may be the communication inside two elements or the interaction relationship between two elements. The term " / and" includes any and all combinations of one or more of the related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0023] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments, and only the parts related to the embodiments of the present application are shown.

[0024] Such as Figure 1 and Figure 2As shown in the figure, the present application provides a multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning, including the following steps (step S1 to step S2):

[0025] S1. According to the state space and action space of multiple microgrids, construct a Markov decision process model including multiple agents, where the state space includes the observable states of each microgrid in the IDR stage, P2P trading stage, and EC stage respectively, and the action space includes the executable actions of each microgrid in the IDR stage, P2P trading stage, and EC stage respectively.

[0026] Each microgrid can sequentially follow the processes of the IDR stage, P2P trading stage, and EC stage. That is, at the current time step, each microgrid can enter the P2P trading stage from the IDR stage in sequence, then enter the EC stage from the P2P trading stage, and finally enter the IDR stage of the next time step from the EC stage of the current time step. The IDR stage can refer to the stage of integrated demand response (IDR), the P2P trading stage can refer to the stage of peer-to-peer (P2P) energy trading, and the EC stage can refer to the stage of energy conversion (EC) based on an energy hub (EH).

[0027] In some embodiments, the microgrids can include residential microgrids, commercial microgrids, and industrial microgrids. For example, commercial microgrids can be shopping malls, office buildings, etc., industrial microgrids can be factories, etc., and residential microgrids can be residential communities, apartments, etc.

[0028] In some embodiments, the tuple of the Markov decision process model can be represented by the following formula:

[0029] ,

[0030] where, is the system state space, which is composed of the union of all agent state spaces of all microgrids; is the set of action spaces executable by each agent; is the set of observable state spaces of each agent, which is derived from the system state; is the transition probability function, which gives the probability of transitioning from state to the next state , which is obtained from the real-time feedback of the actual energy management system; is the set of reward functions of each agent.

[0031] In some embodiments, the state space of the microgrid can be represented by the following formula:

[0032]

[0033] Wherein, represents the time step, represents the observable state of each microgrid in the IDR phase, represents the observable state of each microgrid in the P2P trading phase, represents the observable state of each microgrid in the EC phase, represents the electrical energy demand, represents the thermal energy demand, represents the indoor temperature, represents the outdoor temperature, represents the solar power generation, represents the grid electricity price, represents the natural gas price, represents the thermal energy storage state, represents the cold energy storage state, represents the demand result in the IDR phase, represents the P2P trading price of electrical energy, represents the P2P trading price of heating energy, represents the actual energy quantity of electrical energy P2P trading, represents the actual energy quantity of thermal energy P2P trading.

[0034] The observable state of each microgrid in the IDR phase can be simply referred to as the IDR state. The IDR state can include the time step, electrical energy demand, thermal energy demand, indoor temperature, outdoor temperature, solar power generation, grid electricity price, natural gas price, cold energy storage state, and thermal energy storage state. When the microgrid is in the IDR phase, the microgrid can adjust the execution of various flexible demands according to the energy supply and demand inside and between microgrids. The main goal is to optimize the demand distribution by adjusting the demand curve and the renewable energy supply curve.

[0035] The observable state of each microgrid in the P2P trading phase can be simply referred to as the P2P trading state. The P2P trading state can include the time step, solar power generation, grid electricity price, natural gas price, cold energy storage state, thermal energy storage state, demand result in the IDR phase, P2P trading price of electrical energy, and P2P trading price of heating energy. When the microgrid is in the P2P trading phase, multiple microgrids can use the P2P energy method to conduct electrical energy and thermal energy trading within the region.

[0036] The observable state of each microgrid in the EC phase can be simply referred to as the EC state, and the EC state can include the time step, solar power generation, grid electricity price, natural gas price, cold energy storage state, thermal energy storage state, demand results in the IDR phase, actual energy quantity of electricity trading, and actual energy quantity of heat trading. When the microgrid is in the EC phase, the microgrid can achieve the mutual conversion and supply-demand balance among electric energy, natural gas, thermal energy, and cold energy through various energy conversion and energy storage devices in the energy hub.

[0037] In some embodiments, the IDR action space of the microgrid can include the electricity curtailment rate, heat curtailment rate, electricity transfer state, heat transfer state, and desired room temperature. The P2P trading action space can include the P2P electricity price and electricity trading volume and the P2P heat price and heat trading volume. The EC action space can include the GT utilization rate, electric-to-thermal-to-cooling ratio, charge-discharge ratio of thermal energy storage, and charge-discharge ratio of cold energy storage. Specifically, the electricity curtailment rate can refer to the reduction ratio of the electricity demand that can be curtailed. The heat curtailment rate can refer to the reduction ratio of the heating demand that can be curtailed. The electricity transfer state can refer to the execution state of the electricity demand that can be transferred. The heat transfer state can refer to the execution state of the heating demand that can be transferred. The desired room temperature can refer to the desired indoor temperature. The P2P electricity price and electricity trading volume can refer to the energy trading volume of electric energy under multiple P2P electricity prices. The P2P heat price and heat trading volume can refer to the energy trading volume of thermal energy under multiple P2P heat prices. The GT utilization rate can refer to the utilization rate of the Gas Turbine (GT) device based on the maximum capacity. The electric-to-thermal-to-cooling ratio can refer to the ratio of electric energy used to meet the heating and cooling demands. The charge-discharge ratio of thermal energy storage can refer to the charge-discharge ratio of the thermal energy storage device. The charge-discharge ratio of cold energy storage can refer to the charge-discharge ratio of the cold energy storage device.

[0038] The executable actions of each microgrid in the IDR phase can be simply referred to as IDR actions. The executable actions of each microgrid in the P2P trading phase can be simply referred to as P2P trading actions. The executable actions of each microgrid in the EC phase can be simply referred to as EC actions. Among them, the IDR actions can include the electricity curtailment rate, heat curtailment rate, electricity transfer state, heat transfer state, and desired room temperature. The P2P trading actions can include the P2P electricity price and electricity trading volume and the P2P heat price and heat trading volume. The EC actions can include the GT utilization rate, electric-to-thermal-to-cooling ratio, charge-discharge ratio of thermal energy storage, and charge-discharge ratio of cold energy storage.

[0039] In some embodiments, the action space of the microgrid can be represented by the following formula:

[0040]

[0041] Among them, represents the executable actions of each microgrid in the IDR phase, represents the executable actions of each microgrid in the P2P trading phase, Represents the executable actions when each microgrid is in the EC stage, Represents the reduction ratio of the electricity demand that can be curtailed, Represents the reduction ratio of the heating demand that can be curtailed, Represents the execution status of the electricity demand that can be shifted, Represents the execution status of the heating demand that can be shifted, Represents the desired indoor temperature, Represents the energy trading volume of electrical energy under multiple P2P electricity prices, Represents the energy trading volume of thermal energy under multiple P2P heat prices, Represents the utilization rate of the GT device based on its maximum capacity, Represents the proportion of electrical energy used to meet heating and cooling demands, Represents the charge-discharge ratio of the thermal energy storage device, Represents the charge-discharge ratio of the cold energy storage device.

[0042] In some embodiments, the reward space of the microgrid may include transaction costs with the main power grid and the natural gas grid, P2P energy transaction costs, user discomfort penalties for demand curtailment, user temperature discomfort penalties, incentives for shiftable demand, and GT incentives. Specifically, the transaction costs with the main power grid and the natural gas grid may refer to the energy transaction costs of the main power grid and the natural gas network, the user discomfort penalty for demand curtailment may refer to the user discomfort penalty caused by the curtailable demand, and the GT incentive may refer to the incentive for using the GT device.

[0043] In some embodiments, the microgrid At The reward function at time Can be defined as follows:

[0044]

[0045] Wherein, Represents the energy transaction costs of the main power grid and the natural gas network, Represents the P2P energy transaction costs, Represents the user discomfort penalty caused by the curtailable demand, Represents the user temperature discomfort penalty, Represents the incentive for shiftable demand, Represents the incentive for using the GT device.

[0046] Specifically, the reward function of the microgrid aims to minimize the energy cost and fully absorb renewable energy while considering the user's comfort. The reward function of the microgrid is independent among different microgrids but shared among the agents in each stage of the same microgrid.

[0047] In some embodiments, the energy trading cost between the main power grid and the natural gas network can be expressed by the following formula:

[0048]

[0049] Wherein, 、 represents the quantity of energy traded with the power grid, 、 represents the price of energy traded with the power grid, represents the quantity of energy traded with the natural gas network, represents the price of energy traded with the natural gas network.

[0050] In some embodiments, the P2P energy trading cost can be expressed by the following formula:

[0051]

[0052] Wherein, 、 represent the energy trading volume, 、 represent the average trading price.

[0053] In some embodiments, the user discomfort penalty caused by the curtailable demand can be expressed by the following formula:

[0054]

[0055] Wherein, 、 respectively represent the reduction ratios of the curtailable electricity and heating demands, 、 are preset coefficients. Due to the curtailable demand, the penalty for user discomfort increases with the increase of the reduction ratio.

[0056] In some embodiments, the user temperature discomfort penalty can be expressed by the following formula:

[0057]

[0058] Wherein, represents the desired indoor temperature, represents the temperature penalty coefficient. The penalty function for user discomfort caused by the indoor temperature deviation is determined by the difference between the desired indoor temperature and the optimal temperature.

[0059] In some embodiments, the transferable demand incentive can be expressed by the following formula:

[0060]

[0061] Among them, and represent the execution status of transferable electricity and heating demands, and are preset coefficients. The purpose of transferable demand incentives is to encourage the microgrid to consider long-term interests and actively select the execution time of transferable demands.

[0062] In some embodiments, the incentive for using GT devices can be expressed by the following formula:

[0063]

[0064] Among them, represents the utilization rate of the GT device based on its maximum capacity. If is greater than 0.1, an incentive for using the GT device will be generated to encourage the microgrid to explore the potential utilization and expected benefits of the GT device.

[0065] In some embodiments, the goal of a microgrid system including multiple microgrids can be to obtain the optimal joint strategy of all agents to maximize the total expected reward of all microgrids over a period of time Therefore, the optimal joint strategy of the microgrid can be expressed by the following formula:

[0066]

[0067] Among them, represents the optimal joint strategy, is the discount factor, is the set of all joint strategies

[0068] S2. Based on the Markov decision process model, at each time step, the state, action, and reward value of each microgrid in multiple energy regulation stages in the microgrid environment are obtained respectively. The agent network structure of the agent is trained based on the state, action, and reward value until convergence to obtain the final agent network structure, which is used to obtain the energy management strategy of the microgrid during the actual execution process. The energy regulation stages include the IDR stage, the P2P trading stage, and the EC stage.

[0069] ​In some embodiments, each microgrid may include an IDR agent, a P2P agent, and an EC agent. Specifically, each microgrid may interact with the microgrid environment three times through the IDR agent, the P2P agent, and the EC agent to sequentially complete the supervision processes of the IDR phase, the P2P trading phase, and the EC phase. In each phase, the agent can obtain the state of the phase through observation, as well as the dynamic adjustment results generated by the actions executed in the previous phase. At the same time, the agent can determine the supervision behavior corresponding to the phase. The IDR phase, the P2P trading phase, and the EC phase within a single time step interact with each other as a whole for evaluation. At the end of each time step, the microgrid environment provides a reward value for each microgrid and continues with the steps of the next time step.

[0070] In some embodiments, obtaining the state, actions, and reward value of each microgrid in multiple energy regulation phases in the microgrid environment in each time step may include: each microgrid obtains its own IDR state from the microgrid environment and processes the IDR state through its own IDR agent to obtain the actions in the IDR phase; each microgrid executes the actions in the IDR phase in the microgrid environment to obtain the dynamic adjustment results in the IDR phase and its own P2P trading state, and processes the P2P trading state through its own P2P agent to obtain the actions in the P2P trading phase; each microgrid executes the actions in the P2P trading phase in the microgrid environment to obtain the dynamic adjustment results in the P2P trading phase and its own EC state, and processes the EC state through its own EC agent to obtain the actions in the EC phase. Each microgrid executes the actions in the EC phase in the microgrid environment to obtain the dynamic adjustment results in the EC phase; according to the dynamic adjustment results in the IDR phase, the dynamic adjustment results in the P2P trading phase, and the dynamic adjustment results in the EC phase, obtain the reward value of each microgrid at the current time step and the IDR state of the next time step from the microgrid environment.

[0071] In some embodiments, the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning may further include: selecting the final action in the IDR phase from the actions in the IDR phase through the ε-greedy greedy policy, and obtaining the dynamic adjustment results in the IDR phase and the P2P trading state according to the final action in the IDR phase; and / or selecting the final action in the P2P trading phase through the ε-greedy greedy policy, and obtaining the dynamic adjustment results in the P2P trading phase and the EC state according to the final action in the P2P trading phase; and / or selecting the final action in the EC phase from the actions in the EC phase through the ε-greedy greedy policy, and obtaining the dynamic adjustment results in the EC phase according to the final action in the EC phase.

[0072] Specifically, with a probability of 1 - ε, the actions in the IDR phase output by the IDR agent can be used as the final actions in the IDR phase, and with a probability of ε, a random action in the IDR actions can be selected as the final action in the IDR phase. Correspondingly, with a probability of 1 - ε, the actions in the P2P transaction phase output by the P2P agent can be used as the final actions in the P2P transaction phase, and with a probability of ε, a random action in the P2P transaction actions can be selected as the final action in the P2P transaction phase. With a probability of 1 - ε, the actions in the EC phase output by the EC agent can be used as the final actions in the EC phase, and with a probability of ε, a random action in the EC actions can be selected as the final action in the EC phase.

[0073] In some embodiments, all the states and , actions and reward values of each microgrid in the IDR phase, P2P transaction phase, and EC phase can be used as experience and stored in the replay buffer . Then, a small batch of experiences can be randomly sampled from the replay buffer to update the agent policies in the IDR phase, P2P transaction phase, and EC phase of each microgrid, that is, to train each agent.

[0074] In some embodiments, the agent network structure of each agent can include a critic network, an actor network, a critic target network, and an actor target network. Specifically, the actor network can refer to a policy function that can output actions according to the state; the critic network can refer to a value function that can output the Q value, that is, the expected cumulative reward, the reward value, according to the state and actions; the actor target network can output target actions according to the state corresponding to the next time step, and the critic target network can output the target Q value according to the state corresponding to the next time step and the target actions from the actor target network.

[0075] In some embodiments, training the agent network structure of the agent based on the state, actions, and reward values until convergence to obtain the final agent network structure can include: sampling at least one training batch, sequentially inputting at least one training batch into the actor target network and the critic target network of the agent network structure, calculating the target Q value, where the training batch includes at least the state, actions, reward values of the microgrid, and the state corresponding to the next time step; calculating the loss value of the actor network and the loss value of the critic network of the agent network structure based on the target Q value, and updating the parameters of the actor network and the parameters of the critic network. Specifically, at least one training batch can be sampled from the replay buffer for training the agent network structure.

[0076] In some embodiments, the loss function of the critic network can be expressed by the following formula:

[0077]

[0078] ,

[0079] where are the parameters of the critic network.

[0080] In some embodiments, for the parameters of the actor network, policy gradient can be used for updating, and the gradient can be expressed by the following formula:

[0081]

[0082] where are the parameters of the actor network, is the centralized action-value function, which outputs the Q-value of the microgrid agent at the stage considering the states and actions of all agents.

[0083] In some embodiments, the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning may further include: soft-updating the parameters of the actor target network and the parameters of the critic target network. The soft-update can be expressed by the following formula:

[0084]

[0085]

[0086] where is the soft-update coefficient.

[0087] In some embodiments, the sampling stage and the training stage can be alternated until all network parameters of the agent network structure converge to the optimal total reward of the whole process. Finally, each actor network can be combined into an energy management strategy in the actual execution process, and can output the optimal energy management actions according to the real-time microgrid state.

[0088] In this application, the state space of each microgrid is divided into the IDR stage, the P2P trading stage, and the EC stage, and then the states, actions, and reward values of each microgrid in each stage are obtained for training the intelligent agent network structure to obtain the energy management strategy of the microgrid. In this case, it is possible to realize the internal energy supply and demand process of multiple microgrids and the whole-process multi-energy overall coordination of energy trading between each microgrid, enabling multiple microgrids to interact with the environment independently in stages during the multi-stage energy regulation process, and being able to adjust the energy management strategy in real time according to the regulation results of different stages, achieving efficient real-time coordination within and between microgrids, thereby greatly improving the effect of energy management.

[0089] Those of ordinary skill in the art can understand that all or part of the features / steps of implementing the above method embodiments can be realized by a method, a data processing system, or a computer program. These features can be implemented without using hardware, entirely using software, or using a combination of hardware and software. The aforementioned computer program can be stored in one or more computer-readable storage media. When the computer program stored on the storage medium is executed (such as by a processor), it executes the steps of the method embodiments of the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described above.

[0090] The aforementioned storage media that can store program codes include: a static hard disk, a solid-state drive, a random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), an optical storage device, a magnetic storage device, a flash memory, a magnetic disk or an optical disc, and / or a combination of the above devices, that is, it can be realized by any type of volatile or non-volatile storage device or a combination thereof.

[0091] As Figure 3 shown, this application also provides an embodiment of a processing device 1, including one or more processors 11 and a memory 10; wherein, the memory 10 is used to store one or more computer programs, and one or more processors 11 are used to execute the one or more computer programs stored in the memory 10 so that the processor 11 executes the features / steps of the method embodiments of the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described above.

[0092] The present application also provides a computer program product, which is stored on a data carrier and designed to execute the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described above. Therefore, the computer program product according to the present application has the same advantages as those described in detail with reference to the device according to the present application. The computer program product can be executed as computer-readable instruction codes in each appropriate programming language such as JAVA, C++, etc. In addition, the computer program product can be provided on a network, such as the Internet, or a network user can download the computer program product from a network, such as the Internet, when needed. The computer program product can be implemented either by means of a computer program, i.e., software, or by means of one or more dedicated electronic circuits, i.e., hardware, or in any hybrid form, i.e., by means of software components and hardware components, or in a hybrid form of software, hardware, or software and hardware.

[0093] The above are only the preferred embodiments of the present application. Those skilled in the art will know that various changes or equivalent replacements can be made to these features and embodiments without departing from the spirit and scope of the present application. Additionally, under the teaching of the present application, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present application. Therefore, the present application is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the protection scope of the present application.

Claims

1. A multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning, characterized in that, Including: Construct a Markov decision process model including multiple agents according to the state space and action space of multiple microgrids, where the state space includes the observable states of each microgrid in the IDR stage, P2P trading stage, and EC stage respectively, and the action space includes the executable actions of each microgrid in the IDR stage, P2P trading stage, and EC stage respectively; Based on the Markov decision process model, at each time step, obtain the states, actions, and reward values of each microgrid in multiple energy regulation stages in the microgrid environment respectively. Train the agent network structure of the agent according to the states, actions, and reward values until convergence to obtain the final agent network structure, and the agent network structure is used to obtain the energy management strategy of the microgrid during the actual execution process; Each microgrid includes an IDR agent, a P2P agent, and an EC agent. Obtaining the states, actions, and reward values of each microgrid in multiple energy regulation stages in the microgrid environment at each time step respectively includes: each microgrid obtains its own IDR state from the microgrid environment, and processes the IDR state through its own IDR agent to obtain the action in the IDR stage; Each microgrid executes the action in the IDR stage in the microgrid environment to obtain the dynamic regulation result in the IDR stage and its own P2P trading state, and processes the P2P trading state through its own P2P agent to obtain the action in the P2P trading stage; Each microgrid executes the action in the P2P trading stage in the microgrid environment to obtain the dynamic regulation result in the P2P trading stage and its own EC state, and processes the EC state through its own EC agent to obtain the action in the EC stage. Each microgrid executes the action in the EC stage in the microgrid environment to obtain the dynamic regulation result in the EC stage; According to the dynamic regulation result in the IDR stage, the dynamic regulation result in the P2P trading stage, and the dynamic regulation result in the EC stage, obtain the reward value of each microgrid at the current time step and the IDR state at the next time step from the microgrid environment.

2. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 1, wherein The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning further includes: selecting the final action in the IDR stage from the actions in the IDR stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result and the P2P trading status in the IDR stage according to the final action in the IDR stage; and / or selecting the final action in the P2P trading stage from the actions in the P2P trading stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result and the EC status in the P2P trading stage according to the final action in the P2P trading stage; and / or selecting the final action in the EC stage from the actions in the EC stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result in the EC stage according to the final action in the EC stage.

3. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 1, characterized in that, Training the agent network structure of the agent based on the state, the action, and the reward value until convergence to obtain the final agent network structure includes: sampling at least one training batch, and sequentially inputting the at least one training batch into the actor target network and the critic target network of the agent network structure to obtain the target Q value, where the training batch at least includes the state, action, reward value of the microgrid, and the state corresponding to the next time step; calculating the loss value of the actor network and the loss value of the critic network of the agent network structure based on the target Q value, and updating the parameters of the actor network and the parameters of the critic network.

4. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 1, characterized in that, The IDR state includes the time step, electricity demand, heat demand, indoor temperature, outdoor temperature, solar power generation, grid electricity price, natural gas price, cold energy storage state, and heat energy storage state. The P2P trading state includes the time step, the solar power generation, the grid electricity price, the natural gas price, the cold energy storage state, the heat energy storage state, the demand result in the IDR stage, the P2P trading price of electric energy, and the P2P trading price of heating energy. The EC state includes the time step, the solar power generation, the grid electricity price, the natural gas price, the cold energy storage state, the heat energy storage state, the demand result in the IDR stage, the actual energy quantity of electric energy trading, and the actual energy quantity of heat energy trading.

5. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 1, characterized in that, The IDR action space of the microgrid includes the electricity curtailment rate, the heat curtailment rate, the electricity transfer state, the heat transfer state, and the desired room temperature. The P2P trading action space of the microgrid includes the P2P electricity price and electricity transaction volume and the P2P heat price and heat transaction volume. The EC action space of the microgrid includes the GT utilization rate, the electric-heat-cool ratio, the heat storage charge-discharge ratio, and the cold storage charge-discharge ratio.

6. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 1, characterized in that, The reward space of the microgrid includes the transaction costs with the main power grid and the natural gas grid, the P2P energy transaction costs, the user discomfort penalty for demand curtailment, the user temperature discomfort penalty, the incentive for transferable demand, and the GT incentive.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed, it implements the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described in any one of claims 1-6.

8. A processing device, characterized in that, Including: One or more processors; A memory for storing one or more computer programs, and one or more of the processors are configured to execute the one or more computer programs stored in the memory, so that the one or more processors execute the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described in any one of claims 1-6.

9. A computer program product, characterized in that, The computer program product is stored on a data carrier and is designed to execute the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Interconnected micro-grid group optimization scheduling method based on multi-agent deep learning

    CN119362534A

  • Agent policy learning method with privacy protection in mobile edge computing

    WO2024254892A1