A multi-energy micro-grid control method and system under extreme weather

By optimizing the scheduling strategy of multi-energy microgrids through multi-agent federated transfer reinforcement learning, the resilience and efficiency issues of multi-energy microgrids under extreme weather conditions are solved, enabling autonomous decision-making and economical operation of the system, and improving the system's stability and adaptability.

CN120511775BActive Publication Date: 2026-06-09SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2025-05-09
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Traditional centralized control methods are ill-suited to the complexity and resilience of multi-energy microgrids under extreme weather conditions. Existing multi-agent reinforcement learning algorithms suffer from low learning efficiency and poor cross-scenario adaptability under extreme weather conditions, fail to effectively quantify the cost of load reduction, and lack safe and efficient scheduling strategies.

Method used

A multi-agent federated transfer reinforcement learning approach is adopted. By establishing a multi-energy microgrid model, combining photovoltaic elements, energy storage systems, generators, loads, energy management systems, power and heat networks, cogeneration units, and gas boiler units, the PPO algorithm is used to optimize the scheduling strategy, thereby achieving autonomous decision-making and economical operation of the system.

Benefits of technology

It improves the overall energy utilization efficiency of multi-energy microgrids under extreme weather conditions, enhances the autonomy and resilience of the system, reduces the need for human intervention, and ensures the stable and economical operation of the system under extreme conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120511775B_ABST
    Figure CN120511775B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-energy micro-grid control method and system under extreme weather, by constructing the multi-agent federal migration reinforcement learning framework of multi-energy micro-grid energy management system under extreme weather, including selecting the state variable, action variable and reward function of characterizing multi-energy management system;Through interaction with real-time data, the system can adapt to changing environmental conditions, cope with the uncertainty of renewable energy, energy storage system, heat network supply side and power load side, and realize the optimization of energy management strategy.The application of the present application covers the scheduling and control of multi-energy system, multi-agent reinforcement learning, federal learning, transfer learning and other fields, and can realize the economic operation and resilience improvement of multi-energy micro-grid under extreme weather.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of integrated energy technology, specifically relating to a multi-energy microgrid control method and system under extreme weather conditions. Background Technology

[0002] With the increasing severity of global climate change, the frequency and intensity of extreme weather events have significantly increased, posing a serious threat to the stable operation of power systems. To address this challenge, improving system resilience has become a key research focus in the power system field. Traditional power systems primarily rely on centralized generation and large-scale transmission networks, an architecture that often exhibits vulnerability under extreme weather conditions. In contrast, distributed energy systems and microgrid technologies, due to their flexibility and reliability advantages, demonstrate significant potential in enhancing system resilience.

[0003] In the operation and control of microgrids, traditional centralized control methods struggle to cope with the increasing scale and complexity of systems. Single-agent reinforcement learning methods, due to their high action space dimensionality, often suffer from low learning efficiency and convergence difficulties. Multi-agent methods, however, can significantly reduce the action space dimensionality of individual agents by decomposing complex systems into multiple cooperating agents, thereby improving learning efficiency and system performance. Currently, the application of multi-agent reinforcement learning-based control technology in multi-energy microgrids under extreme weather conditions still requires further development, particularly in optimizing system economic operation and improving resilience. Therefore, an efficient method is urgently needed for optimizing the scheduling strategy of multi-energy microgrids under extreme weather conditions.

[0004] In the emergency control method and system for multi-energy storage coordinated urban integrated energy systems (CN202311156257.1), although setting the weights of each node in each subsystem model can prioritize the restoration of critical loads and ensure effective load restoration under insufficient energy supply, thereby improving the overall utilization efficiency of the energy system, it fails to effectively quantify the cost of load reduction under extreme weather conditions and provide a calculable and optimizable resilience assessment standard for energy management systems. Meanwhile, while single multi-agent reinforcement learning algorithms can solve most scheduling and control problems in integrated energy systems and multi-energy microgrids, they urgently need to be combined with other learning algorithms to address core issues such as insufficient resilience to extreme weather, poor cross-scenario adaptability, and multi-objective dynamic optimization in extreme weather scenarios, in order to achieve safe, efficient, and adaptive intelligent scheduling. Summary of the Invention

[0005] The purpose of this invention is to propose a multi-agent federated transfer reinforcement learning-based control method for multi-energy microgrids under extreme weather conditions, thereby optimizing the joint scheduling strategy of multi-energy microgrids and improving the overall energy utilization efficiency of the system.

[0006] The present invention solves the above-mentioned technical problems through the following technical solutions.

[0007] A multi-energy microgrid control method under extreme weather conditions includes:

[0008] Step 1: Establish a multi-energy microgrid model. The multi-energy microgrid model includes photovoltaic elements, energy storage systems, generators, loads, energy management systems, power and heat networks, combined heat and power units, and gas boiler unit models.

[0009] Step 2: Based on the multi-energy microgrid model, establish an economic operation optimization model for the microgrid under extreme weather conditions. Solve the economic operation optimization model for the microgrid under extreme weather conditions using a reinforcement learning algorithm to obtain the control strategy for the multi-energy microgrid under extreme weather conditions.

[0010] Furthermore, the objective of the economic operation optimization model is to minimize the total operating cost of different types of energy systems. The total operating cost at time t is shown in the following formula:

[0011]

[0012] Where, λ ESS λ represents the price per unit of stored electrical energy. gen The price represented by λ is the price per unit of generator output. load This indicates the penalty pricing incurred due to the removal of a unit of load. It is the heat production cost of the heating network. It is the operating and fuel cost of the combined heat and power unit. It is the electricity generation cost of the power grid. It is the operating cost of the gas-fired boiler unit. It is the charging and discharging energy of the energy storage system. β is the amount of electricity generated by generator i at time t. l,t It's the load reduction ratio. It is the load response of load node l at time t, where Δt is the duration of a unit of time.

[0013] Furthermore, the reinforcement learning algorithm is a multi-agent proximal policy optimization algorithm (PPO algorithm) based on federated transfer learning. The multi-agent proximal policy optimization algorithm based on federated transfer learning is built on the Markov model framework according to the set of state variables, the set of action variables, and the reward function.

[0014] Furthermore, the state variables include the current energy state at time t and the energy state of the energy storage system. Counter c for microgrid de-splitting time steps n,t The feature vector v of time information t Heating network load Electricity price Time step t, the state variable s of the nth agent n,t as follows:

[0015]

[0016] Where, β l,t It's the load reduction ratio. L is the load response of load node l at time t, and L is the index set of the load.

[0017] Furthermore, the action variables are selected from those that directly affect the reward and state. Therefore, the current state is input, and the control quantity of the charging and discharging energy amplitude of the energy storage system is adjusted accordingly. Control quantity of output power amplitude of combined heat and power units and gas boiler units As a variable representing the actions taken by the agent:

[0018]

[0019] a n,t Represents action variables.

[0020] Furthermore, the reward function includes:

[0021] Penalty term Δ for disruption of power grid balance t As shown below:

[0022]

[0023] Where L is the index set of load elements, β l,t It's the load reduction ratio. It is the load response of load node l at time t, J is the index set of the energy storage system, P t c It is the electricity generated by the combined heat and power unit. It is the charging and discharging energy of the energy storage system.

[0024] The agent's optimization objective is to find the economically optimal solution over time steps within the feasible region; therefore, the reward r is... n,t The settings are as follows:

[0025]

[0026] Among them, k0, k1, k2, k3, k4, and k5 are used as weighting factors to balance the various terms. It is the heat production cost of the heating network. It is the operating and fuel cost of the combined heat and power unit. It is the electricity generation cost of the power grid. It is the operating cost of the gas-fired boiler unit, λ ESS It is the price per unit of stored electrical energy. It is the charging and discharging energy of the energy storage system, β l,t Yes, λ gen It is the price per unit of generator output. λ is the amount of electricity generated by generator i at time t. load It is the penalty pricing resulting from the removal of unit load. It is the load response of load node l at time t, where Δt is the duration of a unit of time. t It is the penalty term for the disruption of the power grid balance, and K, L and J are the index sets of generators, loads and energy storage systems, respectively.

[0027] Furthermore, the network structure model of the multi-agent proximal policy optimization algorithm based on federated transfer learning includes an action network and an evaluation network.

[0028] Furthermore, the network structure model of the agents based on the federated transfer learning multi-agent proximal policy optimization algorithm is loaded from the CityLearn environment and transferred to each client to train the agents, update and save the local models of the agents. Based on the characteristics of federated learning, the global model of the central cloud server is shown in the following formula:

[0029]

[0030] The global model from the central cloud server is migrated to each client, and the network model of each client is shown in the following formula:

[0031]

[0032] in, For the next global model on the central cloud server, I represents the number of agents. For each client, there is a network model.

[0033] The system for implementing a multi-energy microgrid control method under extreme weather conditions includes:

[0034] A generalized extreme weather multi-energy microgrid construction module is used to construct multi-energy microgrid models under extreme weather conditions based on photovoltaic elements, energy storage systems, generators, loads, energy management systems, power and heat networks, cogeneration units, and gas boiler units.

[0035] The subsystem model building module is used to build subsystem models based on different energy types, energy usage, energy storage, load status, and the collaborative relationships between energy sources in a multi-energy microgrid.

[0036] The economic operation optimization target model construction module is used to establish an economic operation optimization model of the microgrid under extreme weather conditions based on the multi-energy microgrid model. The objective of the economic operation optimization model of the microgrid under extreme weather conditions is to minimize the operating costs of different types of energy systems.

[0037] The module for solving the economic operation optimization objective model is used to solve the economic operation optimization objective model through a multi-agent federated transfer reinforcement learning algorithm, and to obtain emergency control strategies for multi-energy microgrids under extreme weather conditions.

[0038] A computer device of the invention includes: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements the multi-energy microgrid control method under extreme weather conditions.

[0039] Compared with existing technologies, the beneficial effects of the present invention are as follows:

[0040] (1) This invention optimizes the strategy of joint scheduling of multi-energy microgrids by establishing a multi-energy microgrid model and using a multi-agent federated transfer reinforcement learning method for multi-energy microgrid control under extreme weather conditions, thereby improving the overall energy utilization efficiency of the system.

[0041] (2) The multi-agent PPO algorithm in this invention enables the system to make autonomous decisions. By interacting with the multi-energy microgrid model environment, the agents continuously learn and improve their decisions to adapt to different operating conditions, which reduces the need for human intervention and improves the autonomy of the system.

[0042] (3) Through intelligent decision-making, the system can learn and optimize the strategy of joint dispatch of multi-energy microgrids to ensure the stable and economical operation of the system and improve the system reliability and resilience. Attached Figure Description

[0043] Figure 1 This is a flowchart of the multi-agent reinforcement learning training process in an embodiment.

[0044] Figure 2 This is an example of an architecture diagram for a multi-energy microgrid control method under extreme weather conditions.

[0045] Figure 3 This is a flowchart illustrating a multi-energy microgrid control method under extreme weather conditions, as an example. Detailed Implementation

[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0048] like Figures 2-3 As shown, this embodiment provides a multi-energy microgrid control method under extreme weather conditions based on multi-agent federated transfer reinforcement learning. The method includes:

[0049] Step 1: Establish a multi-energy microgrid model, including photovoltaic elements, energy storage system, generator, load, energy management system, power and heat network, combined heat and power unit, and gas boiler unit model; the models are built as follows:

[0050] The constraints of photovoltaic elements are shown in equation (1) below:

[0051]

[0052] in, This represents the power generation capacity of photovoltaic element i. This represents the power output of photovoltaic element i.

[0053] The energy storage system model is shown in equation (2) below:

[0054]

[0055] in, and Let $\mathbf$ represent the energy state, minimum energy level, and maximum energy level of the energy storage system, respectively. The constraints of the energy storage system are shown in equation (3) below:

[0056]

[0057] in, This indicates the charging and discharging energy of the energy storage system. Indicates the charging process. This indicates the discharge process. and These represent the threshold values ​​for charging and discharging energy, respectively. The dynamics of energy storage are shown in equation (4) below:

[0058]

[0059] Where Δt represents the duration of one unit of time; Indicates electrical energy capacity; φ ch and φ dis These represent the charging and discharging efficiencies of the energy storage system, respectively:

[0060] The diesel generator model is shown in equation (6) below:

[0061]

[0062] in, This represents the amount of electricity generated by generator i at time t. and These represent the minimum and maximum output forces, respectively. The load model is shown in equations (7) to (9) below:

[0063]

[0064] β l,t ∈[0,1](8)

[0065]

[0066] in, This represents the load response of load node l at time t. α l,t This represents the weight of load offloading. K, L, J, and I are the index sets for generators, loads, energy storage systems, and photovoltaic elements, respectively.

[0067] The energy management system model is shown in equation (10) below:

[0068]

[0069] in, This represents the probability of the power grid being disconnected.

[0070] The power balance equation at time t is shown in equation (11):

[0071]

[0072] Among them, P t E >0 indicates the amount of power flowing from the main grid to the microgrid; P t E <0 indicates the amount of power flowing from the microgrid back to the main grid; β l,t Indicates the load reduction ratio.

[0073] The power grid and heating network models are shown in equations (12)-(13) below:

[0074]

[0075] in, and P E P t E The maximum and minimum values ​​of the operating range. This represents the value of heat energy traded from the heating grid to the multi-energy microgrid. P represents the maximum value of heat energy trading. t E This indicates the power exchange capacity with the main grid.

[0076] The electricity and heat generation costs of the power grid and heating network are shown in equations (14)-(15):

[0077]

[0078] in, It's the electricity price. This is the electricity purchase price. H This indicates the price of heat. This represents the cost of heat.

[0079] The model of the combined heat and power unit is shown in equations (17)-(18):

[0080]

[0081] in, These are heat production, electricity production, and the natural gas consumed. and δ c These represent the thermoelectric ratio and the gas-to-thermal conversion efficiency, respectively.

[0082] The output capacity and slope of the power generation are set as shown in equations (19)-(20):

[0083]

[0084] Among them, P c and Let P be the lower and upper bounds of the power generation, respectively. t c R represents the amount of electricity generated. c It is a slope constraint.

[0085] Operating and fuel costs of combined heat and power units As shown in equation (21):

[0086]

[0087] Where, λ g It refers to the price of natural gas. c It is the operating price of a combined heat and power (CHP) unit. It represents thermal energy.

[0088] The gas-fired boiler unit model is shown in equation (22) below:

[0089]

[0090] Where, δ g This indicates the operating price of a gas-fired boiler unit. This indicates the amount of natural gas consumed by the gas-fired boiler unit. This indicates the generated heat power.

[0091] The constraints that the gas-fired boiler unit must meet during operation are shown in equations (23)-(25):

[0092]

[0093] Where, λ gb This indicates a gas-fired boiler unit. H g and R represents the minimum and maximum values ​​of the heat power generated, respectively. g express The slope constraint. λ g and λ gb These represent the price of natural gas and the operating price of gas-fired boilers, respectively.

[0094] The goal of an energy management system is to meet the electricity load while minimizing operating costs in order to improve the system's resilience in the face of extreme events. The mathematical definition of the resilience index is shown in the following equation (26):

[0095]

[0096] Where R represents the resilience of the microgrid. α l,t Indicates the load reduction ratio. λ lood This indicates a fine for reducing the load. This indicates the amount of load reduction. L This indicates the load index set. ∝ indicates an inverse relationship.

[0097] The total operating cost at time t is shown in equation (27):

[0098]

[0099] Where, λ ESS λ represents the price per unit of stored electrical energy. gen The price represented by λ is the price per unit of generator output. load This indicates the penalty pricing incurred due to the removal of a unit of load. It is the heat production cost of the heating network. It is the operating and fuel cost of the combined heat and power unit. It is the electricity generation cost of the power grid. It is the operating cost of the gas-fired boiler unit. It is the charging and discharging energy of the energy storage system. β is the amount of electricity generated by generator i at time t. l,t It's the load reduction ratio. It is the load response of load node l at time t, where Δt is the duration of a unit of time.

[0100] Step 2: Based on the microgrid model, establish an economic operation optimization model for the microgrid under extreme weather conditions. The objective of the economic operation optimization model is to minimize the operating costs of different types of energy systems. As an example, the economic operation optimization model for multi-energy microgrids under extreme weather conditions is solved using a multi-agent federated transfer reinforcement learning algorithm to obtain the control strategy for multi-energy microgrids under extreme weather conditions. A Markov model framework is constructed based on state variables, constraint variables, and indicators to design the set of state variables, action variables, and reward function in the federated transfer learning multi-agent deep reinforcement learning model, i.e., designing the set of state variables, action variables, and reward function in multi-agent deep reinforcement learning.

[0101] Reinforcement learning models involve four key elements: state, action, reward, and environment. The state represents the agent's situation, the action is the operation the agent can perform, the reward is immediate feedback, and the environment is the external world. In deep reinforcement learning, the agent interacts with the environment by observing the state, selecting actions, receiving rewards, and so on, to learn how to formulate strategies to maximize cumulative rewards. Specifically:

[0102] (1) State Variable Design. In multi-energy microgrid control tasks under extreme weather conditions, the state should be selected based on environmental indicators that best reflect the current operating status of the system and are directly related to actions. This invention selects the energy state of the nth agent at the current time t. counter c n,t , eigenvector v t Heating network load Electricity price Time step t.

[0103] The state variables are shown in the following equation:

[0104]

[0105] (2) Action Variable Design. Action variables should be selected that directly affect the reward and state. Therefore, the current state is input, and the amplitude of the output power of the cogeneration unit and the gas boiler unit, as well as the degree of change of the charging and discharging amplitude of the energy storage system, are used as the action variables for the agent:

[0106]

[0107] Among them, an,t This represents the action variable of the nth agent at the current time t. This represents the degree of change in the charging and discharging amplitude of the energy storage system at the current time t for the nth intelligent agent. This represents the degree of change in the magnitude of the output power of the combined heat and power unit at the current time t for the nth intelligent agent. This indicates the degree of change in the amplitude of the output power of the gas-fired boiler unit at the current time t for the nth intelligent agent.

[0108] (3) Reward Function Design. As an example, the optimization objective of the agent is to find the economically optimal solution in the feasible region over time steps. Therefore, the reward is set as shown in equation (28):

[0109]

[0110] The multi-agent federated transfer reinforcement learning algorithm is based on a multi-agent PPO training network structure, which sets the hidden layer size, number of hidden layers, activation function, learning rate, batch size, discount factor, and replay buffer size of the action network and evaluation network of the PPO network structure.

[0111] like Figure 1 This describes the set of states of a multi-energy microgrid under extreme weather conditions at time step t. These states are mapped to the energy system decision variable a via a policy function. n,t The agent interacts with the environment in the next moment, generating a new state s′. n,t And the economic cost required for the next phase of multi-energy microgrids -r n,t This information is stored in an experience pool and randomly selected as training samples during network training.

[0112] Figure 1 The action network used in the decision network selects action variables based on state variables. The objective function of the decision network is L. n (θ n As shown in equation (29):

[0113]

[0114] in, It is the dominance function calculated based on the network's output and reward; δ represents the policy cutoff ratio; ρ t,n θ represents the ratio of the probabilities of the new and old strategies. n Indicates the strategy parameter; E t Represents the mathematical expectation; clip() represents the expected value of ρ. t,n Limited to the range [1-δ, 1+δ]; r n,tγ represents the immediate reward at time step t in the nth round of network updates; γ represents the discount factor; T represents the maximum number of steps the agent can interact with the environment in a training round; t represents the current time step; φ represents the current time step. n This represents the network parameters for updating the evaluation network in the nth round; s T,n This indicates the state of the network update in the nth round at time step T; This represents the function that updates the value in the nth round of the network; s t,n π represents the state of the network in the nth round of updates at time step t; θ Indicates the current strategy; 'a' represents the action; 's' represents the state. This indicates the old strategy.

[0115] Load the agent's network model from the CityLearn environment and migrate it to various clients. Train the agent, update and save the agent's local model. Based on the characteristics of federated learning, the global model of the central cloud server is shown in the following formula:

[0116]

[0117] in, The next global model is the central cloud server, where I represents the number of agents.

[0118] The global model of the central cloud server is migrated to each client, and the network model of each client is shown in equation (31) below:

[0119]

[0120] Table 1. Multi-Agent SAC Training Network Structure and Training Parameter Settings

[0121]

[0122] By interacting with the microgrid model, the agent is trained to make optimal decisions considering the uncertainties of photovoltaic output, power load, and generator output, in order to maximize the reward function and achieve cost optimization, thus reaching an economical operating level.

[0123] The agent was trained using Python, with 20,000 training epochs. The agent was saved after every 10,000 training epochs. Training was stopped when the maximum number of training epochs was reached. The experiment consisted of two steps: First, the agent was trained to avoid network constraints and achieve optimal decision-making; after convergence, the updated network parameters were saved. Second, the agent's decision-making performance was tested on a test dataset after loading the network parameters.

[0124] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A control method for multi-energy microgrids under extreme weather conditions, characterized in that, include: Step 1: Establish a multi-energy microgrid model. The multi-energy microgrid model includes photovoltaic elements, energy storage systems, generators, loads, energy management systems, power and heat networks, combined heat and power units, and gas boiler unit models. Step 2: Based on the multi-energy microgrid model, establish an economic operation optimization model for the microgrid under extreme weather conditions. Solve the economic operation optimization model for the microgrid under extreme weather conditions using a reinforcement learning algorithm to obtain the control strategy for the multi-energy microgrid under extreme weather conditions. The reinforcement learning algorithm is a multi-agent proximal policy optimization algorithm based on federated transfer learning. The multi-agent proximal policy optimization algorithm based on federated transfer learning is built on the Markov model framework according to the set of state variables, the set of action variables and the reward function. The state variables include the current time. Status of charge, status of charge of energy storage system Counter for microgrid disconnection time steps Feature vectors of time information Heating network load Electricity price Time step , No. The state variables of the agent as follows: in, It's the load reduction ratio. It is a load node In time Load response, It is the index set of the load; The selected action variables directly affect the reward and state. Therefore, the current state is input, and the control quantity of the charging and discharging energy amplitude of the energy storage system is adjusted accordingly. Control quantity of output power amplitude of combined heat and power units and gas boiler units As a variable representing the actions taken by the agent: Represents action variables; The reward function includes: Penalties for disrupting the power grid balance As shown below: in, It is a set of indexes for load elements. It's the load reduction ratio. It is a load node In time Load response, It is an index set for energy storage systems. It is the electricity generated by the combined heat and power unit. It is the charging and discharging energy of the energy storage system; The agent's optimization objective is to find the economically optimal solution over time steps within the feasible region; therefore, the reward... The settings are as follows: in, Used as a weighting factor to balance the various terms. It is the heat production cost of the heating network. It is the operating and fuel cost of the combined heat and power unit. It is the electricity generation cost of the power grid. It is the operating cost of the gas-fired boiler unit. It is the price per unit of stored electrical energy. It is the charging and discharging energy of the energy storage system. yes, It is the price per unit of generator output. It is a generator In time Electricity generation, It is the penalty pricing resulting from the removal of unit load. Yes, it is a load node. In time Load response, It is the duration of a unit of time. It is a penalty for disrupting the power grid's energy balance. , and These are index sets for generators, loads, and energy storage systems, respectively.

2. The multi-energy microgrid control method under extreme weather conditions according to claim 1, characterized in that, The objective of the economic operation optimization model is to minimize the total operating cost of different types of energy systems within a given time frame. The total operating cost is shown in the following formula: in, This indicates the price per unit of stored electrical energy. This indicates the price per unit of generator output. This indicates the penalty pricing incurred due to the removal of a unit of load. It is the heat production cost of the heating network. It is the operating and fuel cost of the combined heat and power unit. It is the electricity generation cost of the power grid. It is the operating cost of the gas-fired boiler unit. It is the charging and discharging energy of the energy storage system. It is a generator In time Electricity generation, It's the load reduction ratio. It is a load node In time Load response, It is the duration of a unit of time.

3. The multi-energy microgrid control method under extreme weather conditions according to claim 1, characterized in that, The network structure model of the multi-agent proximal policy optimization algorithm based on federated transfer learning includes an action network and an evaluation network.

4. The multi-energy microgrid control method under extreme weather conditions according to claim 3, characterized in that, Load the network structure model of the agent-based multi-agent proximal policy optimization algorithm from the CityLearn environment, transfer it to each client, train the agents, update and save the local agent models. Based on the characteristics of federated learning, the global model of the central cloud server is shown in the following formula: The global model from the central cloud server is migrated to each client, and the network model of each client is shown in the following formula: in, The next global model for the central cloud server, Indicates the number of agents. For each client, there is a network model.

5. A system for implementing the multi-energy microgrid control method under extreme weather conditions as described in claim 1, characterized in that, include: A generalized extreme weather multi-energy microgrid construction module is used to construct multi-energy microgrid models under extreme weather conditions based on photovoltaic elements, energy storage systems, generators, loads, energy management systems, power and heat networks, cogeneration units, and gas boiler units. The subsystem model building module is used to build subsystem models based on different energy types, energy usage, energy storage, load status, and the collaborative relationships between energy sources in a multi-energy microgrid. The economic operation optimization target model construction module is used to establish an economic operation optimization model of the microgrid under extreme weather conditions based on the multi-energy microgrid model. The objective of the economic operation optimization model of the microgrid under extreme weather conditions is to minimize the operating costs of different types of energy systems. The module for solving the economic operation optimization objective model is used to solve the economic operation optimization objective model through a multi-agent federated transfer reinforcement learning algorithm, and to obtain emergency control strategies for multi-energy microgrids under extreme weather conditions.

6. A computer device, characterized in that, It includes: a memory and a processor, and a computer program stored in the memory, which, when executed on the processor, implements a multi-energy microgrid control method under extreme weather conditions as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Emergency regulation and control method and system for urban integrated energy system with multi-energy storage cooperation

    CN117196211A

  • Micro-grid energy storage scheduling optimization method based on value distribution depth Q network

    CN117060386A

  • Comprehensive energy system optimization scheduling method based on deep reinforcement learning

    CN117455183A