Multi-stage multi-energy management method for multiple micro-grids based on deep reinforcement learning
Through a method based on deep reinforcement learning, a Markov decision-making process model of multiple agents is constructed, which solves the complexity of energy management of multiple microgrids, realizes efficient energy management strategies within and between microgrids, and improves the energy management effect.
Patent Information
- Application Number
- CN202510679360.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
In the prior art, the energy management of multiple microgrids is complex, resulting in poor energy management effects. Especially in the problems of local high share and uncertainty of renewable energy, the complex coupling characteristics of multiple energy management technologies, and the uneven energy distribution of energy between multiple microgrids, it is difficult to design an effective energy management method that takes into account the multi-energy regulation within the microgrid and the multi-energy coordination between microgrids.
Using a multi-stage multi-energy management method based on deep reinforcement learning, a Markov decision-making process model including multiple agents is constructed, and independently interacted with the environment in stages, and the status, actions and reward values of each microgrid in the IDR, P2P transaction and EC stages are obtained, and the agent network structure is trained to obtain energy management strategies.
The energy supply and demand process within multiple microgrids and the entire process of energy transactions between the multiple microgrids has been realized. The energy management strategy can be adjusted in real time according to the regulation results of different stages, and efficient real-time coordination within the microgrid and between the microgrids, thereby greatly improving the effect of energy management.
Smart Images

Figure CN120198253A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent microgrids, and particularly to a multi-stage and multi-energy management method for multiple microgrids based on deep reinforcement learning. Background Art
[0002] With the gradual development of the global consensus on energy conservation and emission reduction in various countries, the installed capacity of global renewable energy has increased significantly, and renewable energy has become a new force for ensuring power supply in China. However, the large-scale substitution of renewable energy has brought challenges to the operation of the power system. How to utilize renewable energy more efficiently, stably, and reliably is the key to promoting the green and low-carbon transformation of energy in the current power system. The wide application of microgrids mainly based on renewable energy such as solar energy and wind energy provides higher economic benefits and flexibility for the smart grid. However, problems such as the unique local high-share and uncertainty of renewable energy, the complex coupling characteristics of various energy management technologies, and the uneven energy distribution among multiple microgrids make the multi-energy management of multiple microgrids face many difficulties and challenges. How to design an energy management method that takes into account both multi-energy regulation within the microgrid and multi-energy coordination among microgrids has become an urgent problem to be solved. Summary of the Invention
[0003] The purpose of this application is to provide a multi-stage and multi-energy management method for multiple microgrids based on deep reinforcement learning to solve the technical problem in the prior art that the energy management of multiple microgrids in the same area is complex and the energy management effect is poor. The many technical effects that can be produced by the preferred technical solutions provided in this application are described in detail below.
[0004] To achieve the above purpose, this application provides the following technical solutions: In a first aspect, a multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning provided by the present application includes: constructing a Markov decision process model including multiple agents according to the state space and action space of each microgrid, where the state space includes observable states of each microgrid in the IDR (Integrated Demand Response) stage, P2P (Peer-to-Peer) trading stage, and EC (Energy Conversion) stage, and the action space includes executable actions of each microgrid in the IDR stage, the P2P trading stage, and the EC stage; based on the Markov decision process model, obtaining the states, actions, and reward values of each microgrid in multiple energy regulation stages in the microgrid environment at each time step, and training the agent network structure of the agent to convergence according to the states, actions, and reward values to obtain the final agent network structure, where the agent network structure is used to obtain the energy management strategy of the microgrid in the actual execution process.
[0005] In some embodiments, each microgrid includes an IDR agent, a P2P agent, and an EC agent. Obtaining the states, actions, and reward values of each microgrid in multiple energy regulation stages in the microgrid environment at each time step includes: each microgrid obtains its own IDR state from the microgrid environment, and processes the IDR state through its own IDR agent to obtain the action in the IDR stage; each microgrid executes the action in the IDR stage in the microgrid environment to obtain the dynamic adjustment result in the IDR stage and its own P2P trading state, and processes the P2P trading state through its own P2P agent to obtain the action in the P2P trading stage; each microgrid executes the action in the P2P trading stage in the microgrid environment to obtain the dynamic adjustment result in the P2P trading stage and its own EC state, and processes the EC state through its own EC agent to obtain the action in the EC stage, and each microgrid executes the action in the EC stage in the microgrid environment to obtain the dynamic adjustment result in the EC stage; according to the dynamic adjustment result in the IDR stage, the dynamic adjustment result in the P2P trading stage, and the dynamic adjustment result in the EC stage, obtaining the reward value of each microgrid at the current time step and the IDR state at the next time step from the microgrid environment.
[0006] In some embodiments, the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning further includes: selecting the final action in the IDR stage from the actions in the IDR stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result and the P2P trading status in the IDR stage according to the final action in the IDR stage; and / or selecting the final action in the P2P trading stage from the actions in the P2P trading stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result and the EC status in the P2P trading stage according to the final action in the P2P trading stage; and / or selecting the final action in the EC stage from the actions in the EC stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result in the EC stage according to the final action in the EC stage.
[0007] In some embodiments, training the agent network structure of the agent according to the state, the action, and the reward value until convergence to obtain the final agent network structure includes: sampling at least one training batch, and sequentially inputting the at least one training batch into the actor target network and the critic target network of the agent network structure to obtain the target Q value, where the training batch at least includes the state, action, reward value of the microgrid, and the state corresponding to the next time step; calculating the loss value of the actor network and the loss value of the critic network of the agent network structure based on the target Q value, and updating the parameters of the actor network and the parameters of the critic network.
[0008] In some embodiments, the IDR state includes the time step, electricity demand, heat demand, indoor temperature, outdoor temperature, solar power generation, grid electricity price, natural gas price, cold energy storage state, and heat energy storage state. The P2P trading status includes the time step, the solar power generation, the grid electricity price, the natural gas price, the cold energy storage state, the heat energy storage state, the demand result in the IDR stage, the P2P trading price of electric energy, and the P2P trading price of heating energy. The EC state includes the time step, the solar power generation, the grid electricity price, the natural gas price, the cold energy storage state, the heat energy storage state, the demand result in the IDR stage, the actual energy quantity of electric energy trading, and the actual energy quantity of heat energy trading.
[0009] In some embodiments, the IDR action space of the microgrid includes the electricity curtailment rate, the heat curtailment rate, the electricity transfer state, the heat transfer state, and the desired room temperature. The P2P trading action space of the microgrid includes the P2P electricity price and electricity trading volume and the P2P heat price and heat trading volume. The EC action space of the microgrid includes the GT utilization rate, the electric-heat-cool ratio, the heat storage charge-discharge ratio, and the cold storage charge-discharge ratio.
[0010] In some embodiments, the reward space of the microgrid includes transaction costs with the main power grid and the natural gas grid, P2P energy transaction costs, user discomfort penalties for demand reduction, user temperature discomfort penalties, transferable demand incentives, and GT (Gas Turbine) incentives.
[0011] In a second aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, it implements the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described above.
[0012] In a third aspect, the present application provides a processing device, including: one or more processors; a memory for storing one or more computer programs, and one or more of the processors for executing the one or more computer programs stored in the memory, so that one or more of the processors execute the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described above.
[0013] In a fourth aspect, the present application provides a computer program product, which is stored on a data carrier and is designed to execute the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described above.
[0014] Implementing one of the above technical solutions of the present application has the following advantages or beneficial effects: In the present application, the state space of each microgrid is divided into the IDR stage, the P2P trading stage, and the EC stage, and then the states, actions, and reward values of each microgrid in each stage are obtained for training the intelligent agent network structure to obtain the energy management strategy of the microgrid. In this case, it is possible to achieve the whole-process multi-energy overall coordination of the internal energy supply and demand process of multiple microgrids and the energy transactions between each microgrid. It enables multiple microgrids to interact with the environment independently in stages during the multi-stage energy regulation process, and can adjust the energy management strategy in real time according to the regulation results of different stages, achieving efficient real-time coordination within and between microgrids, thereby greatly improving the effect of energy management. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings: Figure 1 is a flowchart of the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to the embodiments of the present application; Figure 2It is an interaction schematic diagram of a multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to an embodiment of the present application; Figure 3 It is a structural block diagram of a processing device according to an embodiment of the present application.
[0016] In the figure: 1. Processing device; 10. Memory; 11. Processor. Detailed implementation manners
[0017] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, various exemplary embodiments to be described below will refer to the corresponding drawings, which form a part of the exemplary embodiments and describe various exemplary embodiments that may be adopted to implement the present application. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. It should be understood that they are only examples of processes, methods, devices, etc. consistent with some aspects of the present application disclosed in detail in the appended claims. Other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present application.
[0018] In the description of the present application, it should be understood that terms such as "center", "longitudinal", "lateral", etc. indicate the orientation or positional relationship based on the orientation or position shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the elements referred to must have a specific orientation, be constructed and operated in a specific orientation. Terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. The meaning of the term "plurality" is two or more. The terms "connected" and "coupled" should be understood in a broad sense. For example, they may be fixedly connected, detachably connected, integrally connected, mechanically connected, electrically connected, communicatively connected, directly connected, indirectly connected through an intermediate medium, and may be the internal communication of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0019] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments, and only the parts related to the embodiments of the present application are shown.
[0020] As Figure 1 and Figure 2 shown, the present application provides a multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning, including the following steps (step S1 to step S2): S1. Construct a Markov decision process model including multiple agents according to the state space and action space of multiple microgrids, where the state space includes the observable states of each microgrid being in the IDR stage, P2P trading stage, and EC stage respectively, and the action space includes the executable actions of each microgrid being in the IDR stage, P2P trading stage, and EC stage respectively.
[0021] Each microgrid can sequentially follow the processes of the IDR stage, P2P trading stage, and EC stage. That is, at the current time step, each microgrid can enter the P2P trading stage in sequence from the IDR stage, then enter the EC stage from the P2P trading stage, and finally enter the IDR stage of the next time step from the EC stage of the current time step. The IDR stage can refer to the stage of Integrated Demand Response (IDR), the P2P trading stage can refer to the stage of Peer-to-Peer (P2P) energy trading, and the EC stage can refer to the stage of Energy Conversion (EC) based on an EnergyHub (EH).
[0022] In some embodiments, the microgrid can include residential microgrids, commercial microgrids, and industrial microgrids. For example, commercial microgrids can be shopping malls, office buildings, etc., industrial microgrids can be factories, etc., and residential microgrids can be residential communities, apartments, etc.
[0023] In some embodiments, the tuple of the Markov decision process model can be represented by the following formula: , where, is the system state space, which is composed of the union of the state spaces of all agents of all microgrids; is the set of action spaces executable by each agent; is the set of observable state spaces of each agent, which is derived from the system state; is the transition probability function, which gives the probability of transitioning from state to the next state , and is obtained from the real-time feedback of the actual energy management system; is the set of reward functions of each agent.
[0024] In some embodiments, the state space of the microgrid can be represented by the following formula: where, represents the time step, Indicates the observable state of each microgrid during the IDR phase, Indicates the observable state of each microgrid during the P2P trading phase, Indicates the observable state of each microgrid during the EC phase, Indicates the electrical energy demand, Indicates the thermal energy demand, Indicates the indoor temperature, Indicates the outdoor temperature, Indicates the solar power generation, Indicates the grid electricity price, Indicates the natural gas price, Indicates the thermal energy storage state, Indicates the cold energy storage state, Indicates the demand result during the IDR phase, Indicates the P2P trading price of electrical energy, Indicates the P2P trading price of heating energy, Indicates the actual energy quantity of electrical energy P2P trading, Indicates the actual energy quantity of thermal energy P2P trading.
[0025] The observable state of each microgrid during the IDR phase can be simply referred to as the IDR state. The IDR state can include the time step, electrical energy demand, thermal energy demand, indoor temperature, outdoor temperature, solar power generation, grid electricity price, natural gas price, cold energy storage state, and thermal energy storage state. When the microgrid is in the IDR phase, the microgrid can adjust the execution of various flexible demands according to the energy supply and demand within and between microgrids. The main goal is to optimize the demand distribution by adjusting the demand curve and the renewable energy supply curve.
[0026] The observable state of each microgrid during the P2P trading phase can be simply referred to as the P2P trading state. The P2P trading state can include the time step, solar power generation, grid electricity price, natural gas price, cold energy storage state, thermal energy storage state, demand result during the IDR phase, P2P trading price of electrical energy, and P2P trading price of heating energy. When the microgrid is in the P2P trading phase, multiple microgrids can use the P2P energy method to trade electrical energy and thermal energy within the region.
[0027] The observable state of each microgrid during the EC phase can be simply referred to as the EC state. The EC state can include the time step, solar power generation, grid electricity price, natural gas price, cold energy storage state, thermal energy storage state, demand result during the IDR phase, actual energy quantity of electrical energy trading, and actual energy quantity of thermal energy trading. When the microgrid is in the EC phase, the microgrid can achieve the mutual conversion and supply-demand balance between electrical energy, natural gas, thermal energy, and cold energy through various energy conversion and energy storage devices of the energy hub.
[0028] In some embodiments, the IDR action space of the microgrid may include the electricity curtailment rate, the heat curtailment rate, the electricity transfer state, the heat transfer state, and the desired room temperature. The P2P trading action space may include the P2P electricity price and electricity transaction volume and the P2P heat price and heat transaction volume. The EC action space may include the GT utilization rate, the electricity-to-heat-to-cooling ratio, the heat energy storage charge-discharge ratio, and the cold energy storage charge-discharge ratio. Specifically, the electricity curtailment rate may refer to the reduction ratio of the electricity demand that can be curtailed. The heat curtailment rate may refer to the reduction ratio of the heating demand that can be curtailed. The electricity transfer state may refer to the execution state of the electricity demand that can be transferred. The heat transfer state may refer to the execution state of the heating demand that can be transferred. The desired room temperature may refer to the desired indoor temperature. The P2P electricity price and electricity transaction volume may refer to the energy transaction volume of electric energy under multiple P2P electricity prices. The P2P heat price and heat transaction volume may refer to the energy transaction volume of heat energy under multiple P2P heat energy prices. The GT utilization rate may refer to the utilization rate of the gas turbine (GT) equipment based on the maximum capacity. The electricity-to-heat-to-cooling ratio may refer to the ratio of electric energy used to meet the heating and cooling demands. The heat energy storage charge-discharge ratio may refer to the charge-discharge ratio of the heat energy storage device. The cold energy storage charge-discharge ratio may refer to the charge-discharge ratio of the cold energy storage device.
[0029] The executable actions of each microgrid in the IDR stage can be abbreviated as IDR actions. The executable actions of each microgrid in the P2P trading stage can be abbreviated as P2P trading actions. The executable actions of each microgrid in the EC stage can be abbreviated as EC actions. Among them, the IDR actions may include the electricity curtailment rate, the heat curtailment rate, the electricity transfer state, the heat transfer state, and the desired room temperature. The P2P trading actions may include the P2P electricity price and electricity transaction volume and the P2P heat price and heat transaction volume. The EC actions may include the GT utilization rate, the electricity-to-heat-to-cooling ratio, the heat energy storage charge-discharge ratio, and the cold energy storage charge-discharge ratio.
[0030] In some embodiments, the action space of the microgrid can be represented by the following formula: Among them, represents the executable actions of each microgrid in the IDR stage, represents the executable actions of each microgrid in the P2P trading stage, represents the executable actions of each microgrid in the EC stage, represents the reduction ratio of the electricity demand that can be curtailed, represents the reduction ratio of the heating demand that can be curtailed, represents the execution state of the electricity demand that can be transferred, represents the execution state of the heating demand that can be transferred, represents the desired indoor temperature, Represents the energy trading volume of electric energy under multiple P2P electricity prices, Represents the energy trading volume of heat energy under multiple P2P heat prices, Represents the utilization rate of the GT device based on its maximum capacity, Represents the proportion of electric energy used to meet heat and cooling demands, Represents the charge-discharge ratio of the heat energy storage device, Represents the charge-discharge ratio of the cold energy storage device.
[0031] In some embodiments, the reward space of the microgrid may include transaction costs with the main grid and the natural gas grid, P2P energy transaction costs, user discomfort penalties for demand reduction, user temperature discomfort penalties, transferable demand incentives, and GT incentives. Specifically, the transaction costs with the main grid and the natural gas grid may refer to the energy transaction costs of the main grid and the natural gas network, the user discomfort penalty for demand reduction may refer to the user discomfort penalty caused by the reducible demand, and the GT incentive may refer to the incentive for using the GT device.
[0032] In some embodiments, the microgrid At The reward function at time Can be defined as follows: Wherein, Represents the energy transaction costs of the main grid and the natural gas network, Represents the P2P energy transaction costs, Represents the user discomfort penalty caused by the reducible demand, Represents the user temperature discomfort penalty, Represents the transferable demand incentive, Represents the incentive for using the GT device.
[0033] Specifically, the reward function of the microgrid aims to minimize the energy cost and fully absorb renewable energy while considering the user's comfort. The reward function of the microgrid is independent among different microgrids, but is shared among the agents at different stages of the same microgrid.
[0034] In some embodiments, the energy transaction costs of the main grid and the natural gas network can be expressed by the following formula: Wherein, , Represents the amount of energy traded with the power grid, , Represents the price of the energy traded with the power grid, Represents the amount of energy traded with the natural gas network, Represents the energy price for transactions with the natural gas network.
[0035] In some embodiments, the P2P energy trading cost can be expressed by the following formula: Where, 、 Represents the energy trading volume, 、 Represents the average trading price.
[0036] In some embodiments, the user discomfort penalty due to curtailable demand can be expressed by the following formula: Where, 、 Represent the reduction ratios of curtailable electricity and heating demand respectively, 、 Are preset coefficients. Due to the curtailable demand, the penalty for user discomfort increases with the increase of the reduction ratio.
[0037] In some embodiments, the user temperature discomfort penalty can be expressed by the following formula: Where, Represents the desired indoor temperature, Represents the temperature penalty coefficient. The penalty function for user discomfort caused by indoor temperature deviation is determined by the difference between the desired indoor temperature and the optimal temperature.
[0038] In some embodiments, the transferable demand incentive can be expressed by the following formula: Where, 、 Represent the execution status of transferable electricity and heating demand, 、 Are preset coefficients. The purpose of the transferable demand incentive is to encourage the microgrid to consider long-term interests and actively select the execution time of transferable demand.
[0039] In some embodiments, the incentive for using GT equipment can be expressed by the following formula: Where, Represents the utilization rate of GT equipment based on the maximum capacity. If Is greater than 0.1, then an incentive for using GT equipment will be generated to encourage the microgrid to explore the potential utilization and expected benefits of GT equipment.
[0040] In some embodiments, the goal of a microgrid system including multiple microgrids can be to obtain the optimal joint strategy of all agents , maximizing the sum of expected rewards of all microgrids over a period of time . Therefore, the optimal joint strategy of the microgrid can be expressed by the following formula: where represents the optimal joint strategy, is the discount factor, is the set of all joint strategies .
[0041] S2. Based on the Markov decision process model, at each time step, the states, actions of each microgrid in multiple energy regulation stages in the microgrid environment, and the reward value of each microgrid are obtained respectively. The agent network structure of the agent is trained based on the states, actions and reward values until convergence, and the final agent network structure is obtained. The agent network structure is used to obtain the energy management strategy of the microgrid during the actual execution process. The energy regulation stages include the IDR stage, the P2P trading stage and the EC stage.
[0042] In some embodiments, each microgrid can include an IDR agent, a P2P agent and an EC agent. Specifically, each microgrid can interact with the microgrid environment 3 times through the IDR agent, the P2P agent and the EC agent to complete the supervision processes of the IDR stage, the P2P trading stage and the EC stage in sequence. In each stage, the agent can obtain the state of this stage through observation, as well as the dynamic adjustment result generated by the action executed in the previous stage. At the same time, the agent can determine the supervision behavior of the corresponding stage. The IDR stage, the P2P trading stage and the EC stage within a single time step interact with each other as a whole for evaluation. At the end of each time step, the microgrid environment provides a reward value for each microgrid and continues with the steps of the next time step.
[0043] In some embodiments, obtaining the states, actions of each microgrid in multiple energy regulation stages in the microgrid environment, and the reward value of each microgrid in each time step may include: each microgrid obtains its own IDR state from the microgrid environment, and processes the IDR state through its own IDR agent to obtain the action in the IDR stage; each microgrid executes the action in the IDR stage in the microgrid environment to obtain the dynamic regulation result in the IDR stage and its own P2P trading state, and processes the P2P trading state through its own P2P agent to obtain the action in the P2P trading stage; each microgrid executes the action in the P2P trading stage in the microgrid environment to obtain the dynamic regulation result in the P2P trading stage and its own EC state, and processes the EC state through its own EC agent to obtain the action in the EC stage, and each microgrid executes the action in the EC stage in the microgrid environment to obtain the dynamic regulation result in the EC stage; according to the dynamic regulation results in the IDR stage, the dynamic regulation results in the P2P trading stage, and the dynamic regulation results in the EC stage, obtain the reward value of each microgrid at the current time step and the IDR state at the next time step from the microgrid environment.
[0044] In some embodiments, the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning may further include: selecting the final action in the IDR stage from the actions in the IDR stage through the ε-greedy greedy policy, and obtaining the dynamic regulation result in the IDR stage and the P2P trading state according to the final action in the IDR stage; and / or selecting the final action in the P2P trading stage through the ε-greedy greedy policy, and obtaining the dynamic regulation result in the P2P trading stage and the EC state according to the final action in the P2P trading stage; and / or selecting the final action in the EC stage from the actions in the EC stage through the ε-greedy greedy policy, and obtaining the dynamic regulation result in the EC stage according to the final action in the EC stage.
[0045] Specifically, the action in the IDR stage output by the IDR agent may be used as the final action in the IDR stage with a probability of 1 - ε, and a random action in the IDR actions may be selected as the final action in the IDR stage with a probability of ε. Correspondingly, the action in the P2P trading stage output by the P2P agent may be used as the final action in the P2P trading stage with a probability of 1 - ε, and a random action in the P2P trading actions may be selected as the final action in the P2P trading stage with a probability of ε. The action in the EC stage output by the EC agent may be used as the final action in the EC stage with a probability of 1 - ε, and a random action in the EC actions may be selected as the final action in the EC stage with a probability of ε.
[0046] In some embodiments, all the states of each microgrid in the IDR stage, the P2P trading stage, and the EC stage may be and action and reward value are stored as experience in the replay buffer , and then a mini-batch of experiences is randomly sampled from the replay buffer to update the agent policies in the IDR phase, P2P trading phase, and EC phase of each microgrid, that is, to train each agent.
[0047] In some embodiments, the agent network structure of each agent may include a critic network, an actor network, a critic target network, and an actor target network. Specifically, the actor network may refer to a policy function that can output an action according to the state; the critic network may refer to a value function that can output a Q value, that is, the expected cumulative reward and reward value, according to the state and action; the actor target network can output a target action according to the state corresponding to the next time step, and the critic target network can output a target Q value according to the state corresponding to the next time step and the target action from the actor target network.
[0048] In some embodiments, training the agent network structure of the agent to convergence according to the state, action, and reward value may include: sampling at least one training batch, sequentially inputting at least one training batch into the actor target network and the critic target network of the agent network structure, calculating the target Q value, where the training batch at least includes the state, action, reward value of the microgrid, and the state corresponding to the next time step; calculating the loss value of the actor network and the loss value of the critic network of the agent network structure based on the target Q value, and updating the parameters of the actor network and the parameters of the critic network. Specifically, at least one training batch can be sampled from the replay buffer for training the agent network structure.
[0049] In some embodiments, the loss function of the critic network can be expressed by the following formula: , where are the parameters of the critic network.
[0050] In some embodiments, for the parameters of the actor network where are the parameters of the actor network, is a centralized action value function that outputs the Q-value of the microgrid agent at the stage considering the states and actions of all agents. of all agents.
[0051] In some embodiments, the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning may further include: softly updating the parameters of the actor target network and the parameters of the critic target network. The soft update can be represented by the following formula: where
[0052] is the soft update coefficient.
[0053] In some embodiments, the sampling stage and the training stage can be alternated until all network parameters of the agent network structure converge to the optimal total reward for the whole process. Finally, each actor network can be combined into an energy management strategy in the actual execution process, and can output the optimal energy management action according to the real-time microgrid state.
[0054] In the present application, the state space of each microgrid is divided into the IDR stage, the P2P trading stage, and the EC stage, and then the states, actions, and reward values of each microgrid in each stage are obtained for training the agent network structure to obtain the energy management strategy of the microgrid. In this case, it is possible to realize the whole-process multi-energy overall coordination of the internal energy supply and demand process of multiple microgrids and the energy trading between each microgrid, and enable multiple microgrids to interact with the environment independently in stages during the multi-stage energy regulation process, and be able to adjust the energy management strategy in real time according to the regulation results of different stages, achieving efficient real-time coordination within and between microgrids, thereby greatly improving the effect of energy management.
[0055] The aforementioned storage medium that can store program code includes: a static hard disk, a solid-state drive, a random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), an optical storage device, a magnetic storage device, a flash memory, a magnetic disk or an optical disc and / or a combination of the above devices, that is, it can be implemented by any type of volatile or non-volatile storage device or a combination thereof.
[0056] As Figure 3 shown, an embodiment of a processing device 1 according to the present application includes one or more processors 11 and a memory 10; wherein, the memory 10 is used to store one or more computer programs, and the one or more processors 11 are used to execute the one or more computer programs stored in the memory 10, so that the processors 11 execute the features / steps of the above-mentioned embodiment of the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning.
[0057] The present application also provides a computer program product, which is stored on a data carrier and is designed to execute the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning as described above. Therefore, the computer program product according to the present application has the same advantages as those described in detail with reference to the device according to the present application. The computer program product can be executed as computer-readable instruction codes in each appropriate programming language such as JAVA, C++. In addition, the computer program product can be provided on a network, such as the Internet, or a network, such as Internet users can download the computer program product from the network, such as the Internet when needed. The computer program product can be implemented either by means of a computer program, i.e., software, or by means of one or more dedicated electronic circuits, i.e., hardware, or in any mixed form, i.e., by means of software components and hardware components, or in a mixed form of software, hardware or software and hardware.
[0058] The above are only the preferred embodiments of the present application. Those skilled in the art know that various changes or equivalent replacements can be made to these features and embodiments without departing from the spirit and scope of the present application. In addition, under the teaching of the present application, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present application. Therefore, the present application is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the protection scope of the present application.
Claims
1. A multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning, characterized in that Including: Construct a Markov decision process model including multiple agents according to the state space and action space of multiple microgrids, where the state space includes the observable states of each microgrid in the IDR stage, P2P trading stage, and EC stage respectively, and the action space includes the executable actions of each microgrid in the IDR stage, P2P trading stage, and EC stage respectively; Based on the Markov decision process model, at each time step, obtain the states, actions, and reward values of each microgrid in multiple energy regulation stages in the microgrid environment respectively. Train the agent network structure of the agent according to the states, actions, and reward values until convergence to obtain the final agent network structure, and the agent network structure is used to obtain the energy management strategy of the microgrid during the actual execution process.
2. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 1, characterized in that Each microgrid includes an IDR agent, a P2P agent, and an EC agent. Obtaining the states, actions, and reward values of each microgrid in multiple energy regulation stages in the microgrid environment at each time step respectively includes: each microgrid obtains its own IDR state from the microgrid environment, and processes the IDR state through its own IDR agent to obtain the action in the IDR stage; Each microgrid executes the action in the IDR stage in the microgrid environment to obtain the dynamic regulation result in the IDR stage and its own P2P trading state, and processes the P2P trading state through its own P2P agent to obtain the action in the P2P trading stage; Each microgrid executes the action in the P2P trading stage in the microgrid environment to obtain the dynamic regulation result in the P2P trading stage and its own EC state, and processes the EC state through its own EC agent to obtain the action in the EC stage. Each microgrid executes the action in the EC stage in the microgrid environment to obtain the dynamic regulation result in the EC stage; According to the dynamic regulation result in the IDR stage, the dynamic regulation result in the P2P trading stage, and the dynamic regulation result in the EC stage, obtain the reward value of each microgrid at the current time step and the IDR state at the next time step from the microgrid environment.
3. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 2, characterized in that, The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning further includes: selecting the final action in the IDR stage from the actions in the IDR stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result and the P2P trading status in the IDR stage according to the final action in the IDR stage; and / or selecting the final action in the P2P trading stage from the actions in the P2P trading stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result and the EC status in the P2P trading stage according to the final action in the P2P trading stage; and / or selecting the final action in the EC stage from the actions in the EC stage through the ε-greedy greedy policy, and obtaining the dynamic adjustment result in the EC stage according to the final action in the EC stage.
4. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 1, characterized in that Training the agent network structure of the agent to convergence according to the state, the action, and the reward value to obtain the final agent network structure includes: sampling at least one training batch, and sequentially inputting the at least one training batch into the actor target network and the critic target network of the agent network structure to obtain the target Q value, where the training batch at least includes the state, action, reward value of the microgrid, and the state corresponding to the next time step; calculating the loss value of the actor network and the loss value of the critic network of the agent network structure based on the target Q value, and updating the parameters of the actor network and the parameters of the critic network.
5. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 1, characterized in that The IDR state includes the time step, electricity demand, heat demand, indoor temperature, outdoor temperature, solar power generation, grid electricity price, natural gas price, cold energy storage state, and heat energy storage state. The P2P trading state includes the time step, the solar power generation, the grid electricity price, the natural gas price, the cold energy storage state, the heat energy storage state, the demand result in the IDR stage, the P2P trading price of electric energy, and the P2P trading price of heating energy. The EC state includes the time step, the solar power generation, the grid electricity price, the natural gas price, the cold energy storage state, the heat energy storage state, the demand result in the IDR stage, the actual energy quantity of electricity trading, and the actual energy quantity of heat trading.
6. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 1, characterized in that, The IDR action space of the microgrid includes the electricity curtailment rate, the heat curtailment rate, the electricity transfer state, the heat transfer state, and the desired room temperature. The P2P trading action space of the microgrid includes the P2P electricity price and electricity transaction volume and the P2P heat price and heat transaction volume. The EC action space of the microgrid includes the GT utilization rate, the electric-heat-cool ratio, the heat storage charge-discharge ratio, and the cold storage charge-discharge ratio.
7. The multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to claim 1, characterized in that The reward space of the microgrid includes the transaction costs with the main grid and the natural gas grid, the P2P energy transaction costs, the user discomfort penalty for demand curtailment, the user temperature discomfort penalty, the incentive for transferable demand, and the GT incentive.
8. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed, it implements the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to any one of claims 1-7.
9. A processing device, characterized in that, Including: One or more processors; A memory for storing one or more computer programs, and one or more of the processors are configured to execute the one or more computer programs stored in the memory, so that one or more of the processors execute the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to any one of claims 1-7.
10. A computer program product, characterized in that, The computer program product is stored on a data carrier and is designed to execute the multi-stage multi-energy management method for multiple microgrids based on deep reinforcement learning according to any one of claims 1-7.
Citation Information
Patent Citations
Interconnected micro-grid group optimization scheduling method based on multi-agent deep learning
CN119362534A
Agent policy learning method with privacy protection in mobile edge computing
WO2024254892A1