A double-layer collaborative control method for an electric-thermal-gas integrated energy system
By constructing a two-layer decision control model and combining multi-agent reinforcement learning and power flow calculation, the problems of reward function design complexity and model convergence in integrated energy systems are solved, realizing the optimal control of the electric-heat-gas system and improving computational efficiency and model convergence.
Patent Information
- Application Number
- CN202210657453.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-06-10
AI Technical Summary
Existing technologies struggle to effectively address the complexity of reward function design and model convergence issues in multi-agent reinforcement learning within integrated energy systems. This is particularly true in electric-thermal-gas systems, where the large dimensionality and complex coupling effects of multi-agent systems lead to computational complexity and difficulty in achieving convergence.
A two-layer decision control model is constructed, consisting of an upper-layer multi-agent reinforcement learning model and a lower-layer power flow calculation model. The lower-layer model provides feedback on the power flow calculation results to simplify the design of the reward function. Deep reinforcement learning is performed using the MADDPG framework algorithm, and the agent actions are designed by combining the coupling relationship between electricity, heat, and natural gas.
It effectively simplifies the design of the reward function, improves the convergence and computational efficiency of the model, reduces computational complexity, and realizes the optimized control of the integrated energy system of electricity, heat and gas.
Smart Images

Figure CN115102158B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of integrated energy system optimization control, specifically relating to a two-layer collaborative control method for an integrated energy system of electricity, heat and gas. Background Technology
[0002] Under the national strategy of "dual carbon," the National Development and Reform Commission issued the "Improved Plan for Dual Control of Energy Consumption Intensity and Total Amount" in September 2021, proposing a more comprehensive indicator setting and implementation mechanism for "dual control of energy consumption," and resolutely controlling high-energy-consuming and high-emission projects. Under the trend of energy transition, the main energy source of Integrated Energy Systems (IES) will also change. On the one hand, the integration of various clean energy sources such as wind power, photovoltaics, and natural gas on the energy side increases the uncertainty in power supply and heating; on the load side, new loads that can deeply participate in energy interaction, such as electric vehicles, smart homes, and other forms of energy loads, also inject new uncertainties into the demand side. On the other hand, IES integrates multiple energy sources within a region, such as coal, oil, natural gas, electricity, and heat, and the increasing number of existing market investment entities means that in the current market environment, power grid companies have transformed from the dominant players in energy network construction to important participants. Unlike traditional integrated energy coordination decisions that are based on a holistic perspective, IES faces multiple competing entities, and there are significant game-theoretic relationships between different entities. Against this backdrop, how to consider the game relationship among diversified investment entities in an integrated energy system and propose a control strategy for integrated energy systems has become an urgent problem to be solved.
[0003] Furthermore, considering the close coupling relationships between different energy networks such as electricity and natural gas in integrated energy systems, solving the control strategy problem of integrated energy systems is a nonlinear and non-convex optimization problem, which is difficult to solve using only traditional mathematical modeling methods. In recent years, reinforcement learning (RL) has made significant breakthroughs in data resolution, learning power, and computational power, and has been applied in fields such as intelligent manufacturing and smart healthcare, demonstrating good application results. However, on the one hand, for multi-agent reinforcement learning, as the number of agents increases, the state space becomes larger, and the action space also grows exponentially, which leads to a very large dimensionality and computational complexity in multi-agent systems; on the other hand, each agent in a multi-agent system is mutually coupled and influences each other, and the quality of the reward design directly affects the quality of the learned strategy. Therefore, setting the reward function is a challenge in multi-agent reinforcement learning. Summary of the Invention
[0004] The purpose of this invention is to address the aforementioned problems in the existing technology by providing a two-layer collaborative control method for an integrated energy system of electricity, heat and gas that simplifies reward function design and improves model convergence.
[0005] To achieve the above objectives, the technical solution of the present invention is as follows:
[0006] A two-layer coordinated control method for an integrated electric-heat-gas energy system includes the following steps:
[0007] Step A: Construct a two-layer decision control model, which includes a multi-agent reinforcement learning model at the upper layer and a power flow calculation model at the lower layer. The multi-agents in the multi-agent reinforcement learning model include a power grid agent, a heating network agent, and a gas network agent.
[0008] Step B: Input the collected historical status information data of the integrated energy system into the two-layer decision control model for iterative training. The status information data includes electricity load demand, heat load demand, gas load demand, renewable energy power generation, real-time electricity sales price, and gas purchase price.
[0009] Step C: Collect real-time status information data of the integrated energy system and input it into the trained two-level decision control model to obtain the optimal output of the integrated energy system, namely the electrical output of the cogeneration unit, the thermal output of the gas boiler, the electrical output of the gas turbine, the electrical power transmitted from the grid to the system, and the amount of natural gas transmitted from the natural gas grid to the system.
[0010] In step A, the reward function of the multi-agent reinforcement learning model is:
[0011] r=(γF-r pun )
[0012]
[0013]
[0014]
[0015] c cost (t)=(c be P buy (t)+c bg V buy (t))
[0016]
[0017] i h (t)=C sh H sell (t)
[0018] i g (t)=C sg V sell (t)
[0019] In the above formula, r is the reward value, γ is the reward scaling factor, F is the objective function, and r pun The threshold penalty coefficient is set to m when the multi-agent reinforcement learning model's actions cause the power flow to exceed the threshold; otherwise, it is set to 0. I(t) and C(t) represent the system's energy sales cost and operating cost during time period t, respectively, and T is the total number of time periods. e (t), i h (t), i g (t) represents the cost of selling electricity, heat, and natural gas, respectively, and c represents the cost of selling natural gas. cost (t) represents the cost of energy purchased by the system in time period t, N is the number of agents, and c be c bg These are the purchase prices per unit for electricity and natural gas, respectively, P buy (t) represents the electrical power supplied by the power grid to the system during time period t, V. buy (t) represents the amount of natural gas supplied from the natural gas network to the system during time period t, C se C sh C sg These are the unit prices for electricity, heat, and natural gas, respectively. H represents the electrical power delivered by the system to the user during time period t. sell (t) represents the heat power (V) supplied by the system to the user during time period t. sell (t) represents the volume of natural gas required by the user during time period t.
[0020] In step A, the agent actions of the multi-agent reinforcement learning model are as follows:
[0021]
[0022]
[0023]
[0024]
[0025]
[0026]
[0027] p GT (t)=η GT V GT (t)
[0028]
[0029]
[0030] p grid (t)+p RES (t)+p GT (t)+p CHP (t)-p EB (t)=p load (t)
[0031] h CHP (t)+h GB (t)+h EB (t)=h load (t)
[0032] In the above formula, A t A collection of actions of an intelligent agent. The actions of the smart grid combined heat and power unit, the smart heating network gas boiler, and the smart gas grid gas turbine are respectively, p CHP (t), h GB (t), p GT (t), p EB (t) represents the electrical output of the combined heat and power unit, the thermal output of the gas boiler, the electrical output of the gas turbine, and the electrical output of the electric boiler, respectively, during time period t. CHP (t), h GB (t), h EB (t) represents the thermal output of the combined heat and power unit, gas-fired boiler, and electric boiler during time period t, respectively, and η CHP η GE These are the stator-to-heat ratio and gas-to-electricity conversion efficiency of the combined heat and power unit, respectively, V. CHP (t), V GB (t), V GT (t) represents the natural gas consumption of the combined heat and power unit, gas boiler, and gas turbine during time period t, respectively, where Δt is the duration of each time period, and Q is the value of Q. LHV For natural gas with low calorific value, η GT η GH η EB These represent the gas-to-electric conversion coefficient of a gas turbine, the gas-to-heat conversion efficiency of a gas-fired boiler, and the electro-thermal conversion coefficient of an electric boiler, respectively. grid (t) represents the electrical power supplied by the power grid to the system during time period t, p RES (t) represents the output power of renewable energy during time period t, p load (t), h load (t) represents the electrical load demand and heat load demand during time period t, respectively.
[0033] In step A, the power flow calculation model includes an electricity power flow calculation model and a gas grid power flow calculation model;
[0034] The power flow calculation model is as follows:
[0035]
[0036]
[0037] In the above formula, ΔP and ΔQ are the active and reactive losses of the line, respectively; P and Q are the active and reactive power of the line, respectively; U is the node voltage; and R and X are the resistance and reactance of the line, respectively.
[0038] The gas flow calculation model is as follows:
[0039]
[0040] w ij,t +w ji,t =0
[0041]
[0042]
[0043] ψ min ≤ψ i,t ≤ψ max
[0044] w ij,min ≤w ij,t ≤w ij,max
[0045] In the above formula, c be c bg These are the purchase prices per unit for electricity and natural gas, respectively, P buy (t) represents the electrical power supplied by the power grid to the system during time period t, V. buy (t) represents the amount of natural gas supplied from the natural gas network to the system during time period t, w j,t For the injection airflow at pipe node j during time period t, w ij,t w ji,t These represent the airflow transmitted from node i to j and from j to i during time period t, respectively. jk,t Let Z(j) be the airflow from node j to node k in the pipeline, and let Z(j) and v(j) be the sets of pipelines with node j as the end node and the set with node j as the beginning node, respectively. For the gas source point j during time period t, Let be the natural gas consumption of the multiple agents at pipeline node j during time period t. Let C be the gas load at pipeline node j during time period t. ij Let ψ be a constant related to the pipe length, operating temperature, and pressure difference between nodes. i,t ψ j,tψ represents the air pressure at pipeline nodes i and j during time period t. min ψ max These represent the lower and upper limits of the gas pressure at the pipeline node, respectively. ij,min w ij,max These represent the lower and upper limits of the airflow transmitted through pipe ij, respectively.
[0046] Step B includes the following steps in sequence:
[0047] Step B1: First, input the collected historical state information data of the integrated energy system into the multi-agent reinforcement learning model. Then, the multi-agent reinforcement learning model calculates the multi-agent action plan based on its policy network. The multi-agent action plan includes the electrical output of the grid intelligent agent cogeneration unit, the thermal output of the heating network intelligent agent gas boiler, and the electrical output of the gas network intelligent agent gas turbine.
[0048] Step B2: After receiving the multi-agent action plan, the power flow calculation model performs energy flow calculation on it and determines whether a power flow overshoot occurs. If so, the determination result is fed back to the multi-agent reinforcement learning model; if not, the determination result, along with the calculated power output from the power grid to the system and the amount of natural gas delivered from the natural gas grid to the system, is fed back to the multi-agent reinforcement learning model.
[0049] Step B3: The multi-agent reinforcement learning model calculates the reward value based on the received feedback information and the reward function, and stores its input and output data and reward value in the experience pool;
[0050] Step B4: Repeat steps B1-B3 until the loss function and reward function reach stability.
[0051] The multi-agent reinforcement learning model is a deep reinforcement learning model based on the MADDPG framework algorithm.
[0052] Step C involves matching real-time status information data with data in the experience pool, and selecting the action plan with the highest average reward value over a period of time as the optimal output of the integrated energy system.
[0053] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0054] 1. This invention discloses a two-layer collaborative control method for an integrated electric-heat-gas energy system. First, a two-layer decision-making control model is constructed, consisting of an upper-layer multi-agent reinforcement learning model and a lower-layer power flow calculation model. Then, historical state information data of the integrated energy system is input into the two-layer decision-making control model for iterative training. Subsequently, real-time state information data of the integrated energy system is collected and input into the trained two-layer decision-making control model to obtain the optimal output of the integrated energy system. On one hand, this method feeds back the energy flow calculation results of the lower-layer power flow calculation model to the upper-layer multi-agent reinforcement learning model to calculate the reward value. When the upper-layer multi-agent reinforcement learning model... When the actions output by the learning model violate power flow limits, a penalty coefficient is added to the reward function, thereby reducing the design complexity of the reward function when the multi-agent reinforcement learning model's output actions violate constraints. On the other hand, this method incorporates balancing units into the lower-level power flow calculation model, calculating the electrical power supplied by the grid to the system and the amount of natural gas supplied by the natural gas grid to the system. Feeding this back to the upper-level multi-agent reinforcement learning model further reduces the design complexity of the reward function. The introduction of this lower-level power flow calculation model avoids slow or non-convergent convergence caused by errors in designing the penalty term of the reward function. Therefore, this invention effectively simplifies the design of the reward function and improves the model's convergence.
[0055] 2. In the two-layer collaborative control method for an integrated electric-heat-gas energy system of this invention, the agent action design of the multi-agent reinforcement learning model considers the coupling and balance relationship between integrated energy devices. Based on this relationship, the constraints on the actions of the reinforcement agents can be effectively strengthened, the action space can be reduced, and thus the computational complexity can be reduced. Therefore, this invention reduces computational complexity. Attached Figure Description
[0056] Figure 1 This is a mechanism diagram of the two-level decision control model in this invention.
[0057] Figure 2 This is a diagram of the integrated energy system in Example 1.
[0058] Figure 3 This is a coupled network diagram of the integrated energy system in Example 1. Detailed Implementation
[0059] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings.
[0060] This invention provides a two-layer collaborative control method for an integrated energy system of electricity, heat, and gas. This method is based on a two-layer decision control model consisting of an upper-layer multi-agent reinforcement learning model and a lower-layer power flow calculation model. The lower-layer model provides the upper-layer model with energy flow data of the integrated energy system, and the upper-layer model rewards the decision-making behavior of the agents based on the data provided by the lower-layer model. This method can improve the convergence and training speed of the model and effectively solve the high-dimensional nonlinear optimization problem in complex coupled networks.
[0061] In deep reinforcement learning, the design of the reward function is crucial. Rewards guide the agent to extract decision-related factors from state information and refine them for action selection in the action space. When a multi-agent reinforcement learning model learns an optimal scheduling strategy for an integrated energy system, it may choose actions that do not conform to the system's operational constraints. In such cases, it is usually necessary to define a penalty for the agent when it takes an action that deviates from the correct path—that is, to add a penalty term to the reward function to guide the agent to make the correct decision. The constraints, power flows, and couplings in an integrated energy system are highly complex. If guidance is only provided by setting a penalty term in the reward function, an incorrect or inappropriate penalty term could lead to slow learning speeds for the agent, or even prevent the entire model from converging. Therefore, this invention adopts a two-layer model, with the lower layer calculating the balance constraints, output limits, and power flow conditions of the integrated energy system. This avoids slow or unconvergent convergence caused by errors in the design of the penalty term in the reward function.
[0062] This invention employs a deep reinforcement learning model within the MADDPG framework. MADDPG enables a fully cooperative hybrid game process through distributed multi-agent interactions. Each agent has an optimization module consisting of a policy network (Actor) and a value function network (Critic). The observed state variables of the multi-agents include electricity load, heat load, and gas load in the integrated energy system. Their control actions are the output of each energy device, and the reward value function is designed based on the costs and benefits of the integrated energy system. Local observations are acquired through sensors and communication networks, and action plans are calculated based on the policy network. After each action, the system updates its operating state and feeds back the reward value. Through a "centralized training - distributed execution" training and deployment framework, each agent evaluates the policy based on the received reward signal. This iterative process ultimately yields the policy that maximizes the long-term reward. As the multi-agent reinforcement learning algorithm is continuously trained and updated, its action-reward curve continuously improves and eventually converges, enabling decision-making regarding the time-varying state and control strategies of the integrated energy system.
[0063] Example 1:
[0064] See Figure 1A two-layer coordinated control method for an integrated electric-heat-gas energy system is implemented according to the following steps:
[0065] 1. Construct a two-layer decision control model, which includes a multi-agent reinforcement learning model at the upper layer and a power flow calculation model at the lower layer. The multi-agent reinforcement learning model is a deep reinforcement learning model under the MADDPG framework algorithm, which models the output decision problem of the integrated energy system as a Markov decision process and uses DRL to solve the problem.
[0066] 1.1 Reward Function Design
[0067] The control objective of this system is to maximize the system's net profit within period T by controlling the output of controllable equipment, using the following objective function:
[0068]
[0069]
[0070]
[0071] c cost (t)=(c be P buy (t)+c bg V buy (t))
[0072]
[0073] i h (t)=C sh H sell (t)
[0074] i g (t)=C sg V sell (t)
[0075] In the above formula, F is the objective function, I(t) and C(t) are the energy sales cost and operating cost of the system in time period t, respectively, T is the total number of time periods, and i e (t), i h (t), i g (t) represents the cost of selling electricity, heat, and natural gas, respectively, and c represents the cost of selling natural gas. cost (t) represents the cost of energy purchased by the system in time period t, N is the number of agents, and c be c bg These are the purchase prices per unit for electricity and natural gas, respectively, P buy (t) represents the electrical power supplied by the power grid to the system during time period t, V. buy(t) represents the amount of natural gas supplied from the natural gas network to the system during time period t, C se C sh C sg These are the unit prices for electricity, heat, and natural gas, respectively. H represents the electrical power delivered by the system to the user during time period t. sell (t) represents the heat power (V) supplied by the system to the user during time period t. sell (t) represents the volume of natural gas required by the user during time period t.
[0076] For a heating network, its power generation is transmitted internally from within the integrated energy system and is not purchased from external sources. Its cost is converted from the existing gas source output and electricity output. Therefore, the reward function is designed as follows:
[0077] r=(γF-r pun )
[0078] In the above formula, r is the reward value, γ is the reward scaling factor, and r pun This is the penalty coefficient for exceeding the limit. When the actions of the multi-agent reinforcement learning model exceed the limit, its value is set to a positive integer of 200; otherwise, its value is 0.
[0079] 1.2 Motion Space Design
[0080] The actions in the IES can be represented by the output of three types of equipment: the combined heat and power (CHP) system of the grid intelligent agent, the gas-fired boiler of the heating network intelligent agent, and the gas turbine of the gas network intelligent agent. When the electrical output p of CHP... CHP Once (t) is determined, the CHP thermal output h can be determined according to the formula. CHP (t) and gas consumption V CHP (t); when GB's thermal output h GB Once (t) is determined, the gas consumption V of GB can be determined. GB (t), at which point the thermal output h of EB can also be determined. EB (t) and EB electricity consumption p EB (t). Therefore, the agent's action can be represented as:
[0081]
[0082]
[0083]
[0084]
[0085] For CHP units, there is a coupling relationship between their output electrical power and thermal power. Based on whether their electrothermal ratio changes, they are classified into two types: constant heat source ratio and variable heat source ratio. This invention uses a constant heat source ratio, denoted by η.CHP express:
[0086]
[0087]
[0088] Gas turbines generate electricity by burning natural gas, as shown in the following expression:
[0089] p GT (t)=η GT V GT (t)
[0090] The mathematical model for an electric boiler connected to the power grid as a load is as follows:
[0091]
[0092] The mathematical model of a gas-fired boiler connected to the gas grid as a load is as follows:
[0093]
[0094] During time period t, the electrical power consumed by the electric boiler can be determined through electrical power balance constraints and thermal power balance constraints:
[0095] p grid (t)+p RES (t)+p GT (t)+p CHP (t)-p EB (t)=p load (t)
[0096] h CHP (t)+h GB (t)+h EB (t)=h load (t)
[0097] In the above formula, A t A collection of actions of an intelligent agent. The actions of the smart grid combined heat and power unit, the smart heating network gas boiler, and the smart gas grid gas turbine are respectively, p CHP (t), h GB (t), p GT (t), p EB (t) represents the electrical output of the combined heat and power unit, the thermal output of the gas boiler, the electrical output of the gas turbine, and the electrical output of the electric boiler, respectively, during time period t. CHP (t), h GB (t), h EB (t) represents the thermal output of the combined heat and power unit, gas-fired boiler, and electric boiler during time period t, respectively, and η CHP ηGE These are the stator-to-heat ratio and gas-to-electricity conversion efficiency of the combined heat and power unit, respectively, V. CHP (t), V GB (t), V GT (t) represents the natural gas consumption of the combined heat and power unit, gas boiler, and gas turbine during time period t, respectively, where Δt is the duration of each time period, and Q is the value of Q. LHV For natural gas with low calorific value, η GT η GH η EB These represent the gas-to-electric conversion coefficient of a gas turbine, the gas-to-heat conversion efficiency of a gas-fired boiler, and the electro-thermal conversion coefficient of an electric boiler, respectively. grid (t) represents the electrical power supplied by the power grid to the system during time period t, p RES (t) represents the output power of renewable energy during time period t, p load (t), h load (t) represents the electrical load demand and heat load demand during time period t, respectively.
[0098] The power flow calculation model includes an electricity power flow calculation model and a gas grid power flow calculation model. The electricity power flow calculation model is as follows:
[0099]
[0100]
[0101] In the above formula, ΔP and ΔQ are the active and reactive losses of the line, respectively; P and Q are the active and reactive power of the line, respectively; U is the node voltage; and R and X are the resistance and reactance of the line, respectively.
[0102] The gas flow calculation model is as follows:
[0103] F = min(c be P buy (t)+c bg V buy (t))
[0104]
[0105] w ij,t +w ji,t =0
[0106]
[0107]
[0108] ψ min ≤ψ i,t ≤ψ max
[0109] w ij,min≤w ij,t ≤w ij,max
[0110] In the above formula, F is the objective function, and c be c bg These are the purchase prices per unit for electricity and natural gas, respectively, P buy (t) represents the electrical power supplied by the power grid to the system during time period t, V. buy (t) represents the amount of natural gas supplied from the natural gas network to the system during time period t, w j,t For the injection airflow at pipe node j during time period t, w ij,t w ji,t These represent the airflow transmitted from node i to j and from j to i during time period t, respectively. jk,t Let Z(j) be the airflow from node j to node k in the pipeline, and let Z(j) and v(j) be the sets of pipelines with node j as the end node and the set with node j as the beginning node, respectively. For the gas source point j during time period t, Let be the natural gas consumption of the multiple agents at pipeline node j during time period t. Let C be the gas load at pipeline node j during time period t. ij Let ψ be a constant related to the pipe length, operating temperature, and pressure difference between nodes. i,t ψ j,t ψ represents the air pressure at pipeline nodes i and j during time period t. min ψ max These represent the lower and upper limits of the gas pressure at the pipeline node, respectively. ij,min w ij,max These represent the lower and upper limits of the airflow transmitted through pipe ij, respectively.
[0111] 2. First, the historical state information data of the integrated energy system is collected and input into the multi-agent reinforcement learning model. Then, the multi-agent reinforcement learning model calculates the multi-agent action plan based on its policy network. The state information data includes electricity load demand, heat load demand, gas load demand, renewable energy power generation, real-time electricity sales price and gas purchase price. The multi-agent action plan includes the power output of the grid intelligent agent's cogeneration unit, the heat output of the heating network intelligent agent's gas boiler, and the power output of the gas network intelligent agent's gas turbine.
[0112] 3. After receiving the multi-agent action plan, the power flow calculation model performs energy flow calculation on it and determines whether a power flow overshoot occurs. If so, the determination result is fed back to the multi-agent reinforcement learning model; if not, the determination result, along with the calculated power output from the power grid to the system and the amount of natural gas delivered from the natural gas grid to the system, is fed back to the multi-agent reinforcement learning model.
[0113] 4. The multi-agent reinforcement learning model calculates the reward value through the reward function based on the received feedback information, and stores the state of the current time period, the state of the next time period, and the action in the experience pool.
[0114] 5. Repeat steps 2-4 until the loss function L(θ) and reward function reach a stable state:
[0115]
[0116] In the above formula, S is the number of training samples per round, and y j Let Q be the target Q value, and θ be the estimated network parameter value. μ (s j ,a j ) represents the network value of the value function.
[0117] 6. Collect real-time status information data of the integrated energy system and input it into the trained two-level decision control model. By matching the real-time status information data with the data in the experience pool, the action plan with the highest average reward value over a period of time is taken as the optimal output of the integrated energy system, namely the electrical output of the cogeneration unit, the thermal output of the gas boiler, the electrical output of the gas turbine, the electrical power transmitted from the grid to the system, and the amount of natural gas transmitted from the natural gas grid to the system. At the same time, the input status data and the matched actions are stored in the experience pool so that the strategy can be continuously updated.
[0118] The integrated energy system used in this embodiment consists of a 33-node power system and a 20-node natural gas system (see system diagram). Figure 2 See the coupled network diagram. Figure 3 The heating network is connected to the power grid in the form of electrical load. The power system includes two CHP power sources, two wind power sources, and one photovoltaic power source. The heating system includes four heat sources (one electric boiler and three gas boilers), and the natural gas system has six gas sources. The system scheduling duration is 24 hours, with a 1-hour interval between two adjacent time periods. To test the proposed model's ability to handle system uncertainties, the electric, heat, and gas loads in the IES and the output of renewable energy sources are considered. The existence of uncertainties leads to a large number of different scenarios when performing IES control. In this embodiment, one year's worth of renewable energy (wind turbines and photovoltaics), electric load data, heat load data, and gas load data are input into the model for strategy result verification. The search time step is 24, and 50,000 rounds are set.
[0119] Simulation results show that when using the aforementioned two-layer model to control the integrated energy system, initially, due to the agent's incomplete exploration of action strategies, the agent may choose to sacrifice profit to ensure that operational constraints are met, resulting in a low reward function value. Through extensive learning in the later stages, the agent can make effective decisions for different state scenarios, and the reward function continuously increases until convergence after approximately 26,000 rounds.
[0120] The above results demonstrate that the proposed two-layer model, by fully considering the dynamic game among the three main components (electricity, heat, and gas), can effectively improve the energy utilization efficiency of the integrated energy system. Furthermore, the model combines data-driven multi-agent reinforcement learning with traditional power flow algorithms, resulting in higher solution efficiency and enabling coordinated control of the integrated energy system.
Claims
1. A two-layer coordinated control method for an integrated electric-heat-gas energy system, characterized in that: The control method includes the following steps in sequence: Step A: Construct a two-layer decision control model, which includes a multi-agent reinforcement learning model at the upper layer and a power flow calculation model at the lower layer. The multi-agent agents in the multi-agent reinforcement learning model include a power grid agent, a heating network agent, and a gas network agent. The reward function of the multi-agent reinforcement learning model is: ; ; ; ; ; ; ; ; In the above formula, As a reward value, Let F be the reward scaling factor, and F be the objective function. This is the penalty coefficient for exceeding the limit. When the actions of the multi-agent reinforcement learning model exceed the limit, its value is a set positive integer m; otherwise, its value is 0. , Let $T$ be the energy sales cost and operating cost of the system during time period $t$, respectively, where $T$ is the total number of time periods. , , These are the costs of selling electricity, heat, and natural gas, respectively. Let N be the cost of energy purchased by the system during time period t, and N be the number of agents. , These are the purchase prices per unit for electricity and natural gas, respectively. Let t be the electrical power supplied by the power grid to the system during time period t. Let t represent the amount of natural gas supplied from the natural gas network to the system during time period t. , , These are the unit prices for electricity, heat, and natural gas, respectively. Let t be the electrical power delivered by the system to the user during time period t. The heat output of the system to users during time period t is the heat output of the system. The volume of natural gas required by the user during time period t; Step B: Input the collected historical status information data of the integrated energy system into the two-layer decision control model for iterative training. The status information data includes electricity load demand, heat load demand, gas load demand, renewable energy power generation, real-time electricity sales price, and gas purchase price. Step C: Collect real-time status information data of the integrated energy system and input it into the trained two-level decision control model to obtain the optimal output of the integrated energy system, namely the electrical output of the cogeneration unit, the thermal output of the gas boiler, the electrical output of the gas turbine, the electrical power transmitted from the grid to the system, and the amount of natural gas transmitted from the natural gas grid to the system.
2. The dual-layer coordinated control method for an integrated electric-heat-gas energy system according to claim 1, characterized in that: In step A, the agent actions of the multi-agent reinforcement learning model are as follows: ; ; ; ; ; ; ; ; ; ; ; In the above formula, A collection of actions of an intelligent agent. , , These refer to the operations of the smart grid combined heat and power unit, the smart heating network gas boiler, and the smart gas network gas turbine. , , , These represent the electrical output of the combined heat and power unit, the thermal output of the gas-fired boiler, the electrical output of the gas turbine, and the electrical output of the electric boiler during time period t. , , These represent the heat output of the combined heat and power unit, gas-fired boiler, and electric boiler during time period t, respectively. , These are the stator-to-heat ratio and gas-to-electricity efficiency of a combined heat and power (CHP) unit. , , These represent the natural gas consumption of the combined heat and power unit, gas boiler, and gas turbine during time period t. The duration of each time period, It is a low-calorific-value natural gas. , , These are the gas-to-electric conversion coefficient of a gas turbine, the gas-to-heat conversion efficiency of a gas-fired boiler, and the electro-thermal conversion coefficient of an electric boiler. Let t be the electrical power supplied by the power grid to the system during time period t. Let t be the output power of renewable energy during time period t. , These represent the electrical load demand and heat load demand for time period t, respectively.
3. The dual-layer coordinated control method for an integrated electric-heat-gas energy system according to claim 1, characterized in that: In step A, the power flow calculation model includes an electricity power flow calculation model and a gas grid power flow calculation model; The power flow calculation model is as follows: ; ; In the above formula, , These are the active and reactive power losses of the line, respectively. , These are the active and reactive power of the line, respectively. For node voltage, , These are the resistance and reactance of the circuit, respectively. The gas flow calculation model is as follows: ; ; ; ; ; ; In the above formula, For the injected airflow at pipe node j during time period t, , These represent the airflow transmitted from pipe node i to j and from j to i during time period t, respectively. Let J be the airflow flowing from node J to node K in the pipeline. For the gas source point j during time period t, Let be the natural gas consumption of the multiple agents at pipeline node j during time period t. Let be the gas load at pipeline node j during time period t. Let ij be a constant related to the pipe length, operating temperature, and pressure difference between nodes. , Let be the air pressures at pipeline nodes i and j during time period t. , These represent the lower and upper limits of the air pressure at the pipeline node, respectively. , These represent the lower and upper limits of the airflow transmitted through pipe ij, respectively.
4. The dual-layer coordinated control method for an integrated electric-heat-gas energy system according to claim 1, characterized in that: Step B includes the following steps in sequence: Step B1: First, input the collected historical state information data of the integrated energy system into the multi-agent reinforcement learning model. Then, the multi-agent reinforcement learning model calculates the multi-agent action plan based on its policy network. The multi-agent action plan includes the electrical output of the grid intelligent agent cogeneration unit, the thermal output of the heating network intelligent agent gas boiler, and the electrical output of the gas network intelligent agent gas turbine. Step B2: After receiving the multi-agent action plan, the power flow calculation model performs energy flow calculation on it and determines whether a power flow overshoot occurs. If so, the determination result is fed back to the multi-agent reinforcement learning model; if not, the determination result, along with the calculated power output from the power grid to the system and the amount of natural gas delivered from the natural gas grid to the system, is fed back to the multi-agent reinforcement learning model. Step B3: The multi-agent reinforcement learning model calculates the reward value based on the received feedback information through the reward function, and stores its input and output data in the experience pool; Step B4: Repeat steps B1-B3 until the loss function and reward function reach stability.
5. The dual-layer coordinated control method for an integrated electric-heat-gas energy system according to claim 4, characterized in that: The multi-agent reinforcement learning model is a deep reinforcement learning model based on the MADDPG framework algorithm.
6. The dual-layer coordinated control method for an integrated electric-heat-gas energy system according to claim 5, characterized in that: Step C involves matching real-time status information data with data in the experience pool, and selecting the action plan with the highest average reward value over a period of time as the optimal output of the integrated energy system.