Scheduling Method for Electrothermal Coupled Systems Based on Twin Delay DDPG Algorithm
By improving the scheduling of electrothermal coupling systems using the twin-delay DDPG algorithm, the hyperparameter sensitivity and single-objective optimization shortcomings of the DDPG algorithm are addressed. This enables real-time response and online optimization to renewable energy and load uncertainties, thereby enhancing the low-carbon and economical scheduling performance of electrothermal coupling systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2026-03-13
AI Technical Summary
Existing reinforcement learning DDPG algorithms suffer from problems in scheduling electro-thermal coupled systems, such as sensitivity to hyperparameters, insufficient single-objective optimization, and difficulty in real-time response to uncertainties in renewable energy and electricity/heat loads.
We employ the Siamese Delayed DDPG algorithm, combined with truncated double Q learning, delayed policy update, and target policy smoothing techniques, to design a multi-objective reward function, establish a reinforcement learning environment for the electrothermal coupled system, and conduct interactive training to optimize scheduling.
It enables low-carbon economic dispatch of electrothermal coupling systems, can respond in real time to the uncertainties of renewable energy and electricity/heat loads, provides online optimization capabilities, and improves the effectiveness of dispatch and multi-objective optimization capabilities.
Smart Images

Figure CN116502751B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power transmission and distribution, and relates to a scheduling method for electrothermal coupled systems based on the twin delay DDPG algorithm. Background Technology
[0002] In recent years, with the increasing demand for electricity and heating, the coupling between power systems and heating systems has become increasingly close. Compared with traditional coal-fired units, combined heat and power (CHP) units can not only supply electricity and heat simultaneously with higher energy efficiency but also produce less carbon emissions. Meanwhile, the high energy conversion rate of electric boilers and the low carbon emissions of gas-fired boilers have also been widely applied in CHP systems. Researching the low-carbon economic dispatch of CHP systems helps break down operational barriers between different energy systems, enabling efficient synergy and complementarity between electricity and heat, and contributing to energy conservation and emission reduction applications.
[0003] With the trend of large-scale renewable energy grid integration, the research on low-carbon economic dispatch of electro-thermal coupling systems containing renewable energy has also become a major focus. However, the uncertainty of source load caused by the intermittency and volatility of energy and errors in electricity / heat load forecasting during the dispatch process poses a significant challenge to the low-carbon economic dispatch of electro-thermal coupling systems. Methods for handling source load uncertainty mainly include stochastic optimization and robust optimization. Robust optimization methods are a type of uncertainty decision-making method based on interval perturbation information. Because they consider the optimal solution under the worst-case scenario, the results may be relatively conservative. Stochastic optimization methods generally require prior assumptions about the probability distribution of random variables, and their performance heavily depends on the skill and experience of the modeler, resulting in insufficient accuracy in characterizing uncertainty.
[0004] In recent years, reinforcement learning has attracted considerable attention from researchers due to its advantages of avoiding the modeling of complex uncertainties, adapting to uncertainties, and making rapid decisions. In the power sector, some studies have applied reinforcement learning methods to solve optimal scheduling problems. For example, some literature uses the Deep Q-Network (DQN) algorithm based on reinforcement learning to manage and optimize energy in microgrids with economic efficiency as the goal; other literature addresses the volatility of renewable energy output and the uncertainty of load demand by using the DQN algorithm to achieve energy optimization management of integrated electric and thermal energy systems; and still others address the problems of large power fluctuations on both the load and power supply sides, as well as the complexity of various uncertainties by adding energy storage systems to ensure real-time supply and demand balance, and using the DQN algorithm to achieve coordinated control of composite energy storage in micro-external power grids.
[0005] However, the reinforcement learning DQN algorithm is implemented based on a discrete action space. During the optimization scheduling process, the action values of the devices need to be discretized. Excessive discretization can lead to an explosion of action space dimensions, while insufficient discretization reduces optimization accuracy. Therefore, many scholars are focusing on the reinforcement learning DDPG (Deep Deterministic Policy Gradient) algorithm, which is implemented based on a continuous action space. Some literature addresses the intermittency of renewable energy and the uncertainty on the energy consumption side by transforming the dynamic economic scheduling of the integrated energy system into a reinforcement learning process and verifying it using the DDPG algorithm. Other literature, based on the reinforcement learning DDPG algorithm, can achieve dynamic control of the external power grid's maximum transmission capacity in about one second. Still other literature uses the reinforcement learning DDPG algorithm to achieve autonomous control of the external power grid's safe operation based on real-time system data collected by the SCADA system and phasor measurement unit. Finally, other literature uses reinforcement learning to model the dynamic energy conversion and management problem of the electro-pneumatic coupling system as an MDP (Markov Decision Process) and solves it using the DDPG algorithm. Summary of the Invention
[0006] While the reinforcement learning-based Directed Delayed DDPG algorithm has been extensively studied and effectively applied in power optimization decision-making, it still has certain drawbacks, such as extreme sensitivity to hyperparameters and the unfavorable effect of deterministic policies on action exploration. Therefore, this invention proposes an improved version of the DDPG algorithm, the Siamese Delayed DDPG algorithm, employing truncated double Q-learning, delayed policy updates, and target policy smoothing techniques. This effectively addresses the shortcomings of the DDPG algorithm and has been successfully applied in numerous fields, including robot control and UAV edge computing.
[0007] Meanwhile, most existing studies employ single-objective optimization, with few utilizing reinforcement learning methods to solve multi-objective optimization problems. Therefore, this invention comprehensively considers operating costs and carbon emission targets in the optimization process, transforming the low-carbon economic scheduling process of the electrothermal coupling system into a Markov decision process. It designs action spaces and state spaces, incorporates constraint and penalty mechanisms to design a multi-objective reward function, and establishes a reinforcement learning environment for the low-carbon economic scheduling of the electrothermal coupling system. The Siamese Delayed DDPG algorithm is used for interactive training with the environment, and the trained reinforcement learning model can make online decisions. Finally, through case studies, the effectiveness and superiority of the Siamese Delayed DDPG algorithm in achieving low-carbon economic scheduling of the electrothermal coupling system are demonstrated.
[0008] This invention is achieved through the following technical solution: a scheduling method for an electrothermal coupled system based on the twin-delay DDPG algorithm, comprising the following steps:
[0009] Step 1: Establish a low-carbon economic dispatch model for the electrothermal coupling system that takes into account economic efficiency and carbon emissions;
[0010] Step 2: Transform the low-carbon economic dispatch process of the electrothermal coupling system containing renewable energy into a Markov decision process;
[0011] Step 3: Design a multi-objective reward function for reinforcement learning based on the low-carbon economic scheduling model and penalty constraint mechanism of the electrothermal coupling system;
[0012] Step 4: Use the twin-delayed DDPG algorithm to interactively train the reinforcement learning agent; the trained reinforcement learning agent performs low-carbon economic scheduling of the electrothermal coupling system.
[0013] More specifically, in step one, the low-carbon economic dispatch model of the electrothermal coupling system includes system operating costs and carbon emissions, as well as constraints;
[0014]
[0015]
[0016] In the formula, F represents the system operating cost; P t,grid Q represents the power purchased from the external power grid at time t. t,gas Let $ be the gas purchase capacity of the natural gas source at time t; t,grid Let $ be the electricity price at time t; gas The price of natural gas; α n P represents the unit maintenance cost of the nth device; t,n Let be the output of the nth device at time t; T be the operating time; N be the total number of devices; C be the carbon emissions; λ be the carbon emission factor of the external power grid; z t t represents the natural gas consumption at time t; w represents the lower heating value of natural gas; y represents the carbon content per unit calorific value of natural gas; d represents the carbon oxidation rate of natural gas.
[0017] Power balance constraints:
[0018]
[0019] In the formula, P t,pv The output of the photovoltaic power station at time t; P t,chp P represents the electrical output of the cogeneration unit at time t. t,load Let P be the electrical load at time t; t,eb H represents the electrical power consumed by the electric boiler at time t. t,chp H represents the thermal output of the cogeneration unit at time t; t,eb H represents the thermal output of the electric boiler at time t. t,gb H represents the thermal output of the gas-fired boiler at time t. t,load The heat load at time t; Q t,chp Q represents the gas power consumed by the cogeneration unit at time t.t,gb Let t be the gas power consumed by the gas boiler;
[0020] Coupling device constraints:
[0021]
[0022] In the formula, γ chp For the gas-to-electric conversion efficiency of a combined heat and power unit; κ chp The heat-to-power ratio of a combined heat and power (CHP) unit; κ eb The electrothermal conversion efficiency of the electric boiler; κ gb The gas-heat conversion efficiency of a gas-fired boiler;
[0023] Equipment output upper and lower limit constraints:
[0024]
[0025] In the formula, and These are the upper and lower limits of the electrical output of a combined heat and power unit; and These are the upper and lower limits of the thermal output of the electric boiler, respectively. and For the upper and lower limits of the thermal output of the gas-fired boiler; g t,pv Let t be the predicted output value of the photovoltaic power station;
[0026] Equipment ramping constraints:
[0027]
[0028] In the formula, For combined heat and power (CHP) units; For combined heat and power (CHP) units, the landslide rate is used. The ramp-up rate of the electric boiler; The landslide rate of the electric boiler.
[0029] More specifically, in step two, the Markov decision process is defined by five variables: [s, a, r, P, γ]. Here, s represents the state space, which is the set of all state variables of the environment that the reinforcement learning agent can perceive; a represents the action space, which is the set of actions that the reinforcement learning agent can take in response to the environment; r represents the reward function, which is the immediate reward returned by the environment to the decision-maker based on the state and action; γ represents the reward discount rate; and P represents the state transition probability.
[0030] More specifically, the state space is represented as:
[0031] s t =[g t+1,load ,h t+1,load ,g t+1,pv ,$t,grid ,P t,chp H t,eb (7)
[0032] In the formula, g t+1,load The predicted electricity load at time t+1; h t+1,load The predicted heat load value at time t+1; g t+1,pv This represents the predicted output of the photovoltaic power station at time t+1.
[0033] More specifically, the action space is represented as:
[0034] a t =[ΔP t,chp ,ΔH t,eb ,P t,pv (8)
[0035] In the formula, ΔP t,chp ΔH represents the change in electrical output of the cogeneration unit at time t. t,eb This refers to the change in the thermal output of the electric boiler, while also satisfying the equipment ramp-up constraint, namely:
[0036]
[0037] In the formula, This represents the minimum ramp rate for a combined heat and power (CHP) unit. This represents the maximum ramp rate for a combined heat and power (CHP) unit. This represents the minimum ramp rate for the electric boiler. This represents the maximum ramp rate for the electric boiler.
[0038] More specifically, in step three, the reward function for the electrothermal coupling system includes: system operating cost, carbon emissions, and equipment over-limit penalty; the system operating cost and carbon emissions are normalized.
[0039]
[0040] In the formula, f t The normalized system operating cost target; F t,max F represents the upper limit of the system operating cost at time t; t,min F represents the lower bound of the system operating cost at time t; t Let c be the system operating cost at time t; t For normalized carbon emission targets; C t C represents carbon emissions at time t; t,max C represents the upper limit of carbon emissions at time t; t,min β represents the lower limit of carbon emissions at time t; t Let be the device violation penalty factor at time t. It is 1 if any device violates the limit, and 0 if no device violates the limit. Therefore, the reinforcement learning multi-objective reward function is as follows:
[0041] r t =f t +c t -β t (11)
[0042] In the formula, r t To reinforce learning the multi-objective reward function at time t.
[0043] More specifically, in step four, the Siamese Delayed DDPG algorithm learns two Q-value functions and updates the critic network using a double Q-value learning approach, thus establishing two Q-value networks to estimate the value of the next state:
[0044]
[0045] In the formula, Let Q represent the Q value of policy Ⅰθ1 in state space s and action space a; Represents the policy network I in state space s; Let Q represent the Q value of policy IIθ2 in state space s and action space a; Represents the policy network II in state space s;
[0046] Compile the Bellman equation using the minimum of the two Q values:
[0047]
[0048] In the formula, Y(s) represents the value of the current state space s; r(s) is the immediate reward in the state space s; and γ is the reward discount rate. This represents the total discount for future rewards.
[0049] More specifically, the twin-delay DDPG algorithm delays policy updates: policy networks I and II are updated at a lower frequency than the Q-value network.
[0050] More specifically, the twin-delayed DDPG algorithm modifies and updates equation (13) as follows:
[0051]
[0052] In the formula, ε represents the added noise, sampled from a normal distribution, i.e., ε ~ N(0,σ); clip() is the cutoff function; c is the cutoff factor; σ is the standard deviation of the normal distribution; N(0,σ) represents a normal distribution with a standard deviation of σ; Q θ Let Q be the Q-value under strategy θ.
[0053] This invention first establishes a low-carbon economic scheduling model for an electrothermal coupling system that considers both economic efficiency and carbon emissions. Then, it transforms the low-carbon economic scheduling process of the electrothermal coupling system containing renewable energy into a Markov decision process, aiming to minimize both economic efficiency and carbon emissions. A multi-objective reward function is designed with a penalty constraint mechanism, and an improved algorithm based on deep deterministic policy gradients is used. The Siamese Delayed Generation (DDPG) algorithm is employed to interactively train the reinforcement learning agent. Finally, the reinforcement learning agent trained by the proposed method can respond in real-time to the uncertainties of renewable energy and electricity / heat loads, optimizing the low-carbon economic scheduling of the electrothermal coupling system containing renewable energy online.
[0054] This invention achieves low-carbon economic scheduling of electrothermal coupling systems based on reinforcement learning. A low-carbon economic scheduling model for electrothermal coupling systems, taking into account carbon emissions, is established. A reinforcement learning environment model is designed for this model, and the Siamese Delayed DDPG algorithm is used for verification. The invention has the following advantages: 1) The Siamese Delayed DDPG algorithm used in this invention overcomes the shortcomings of the DDPG algorithm. 2) This invention solves the problem that reinforcement learning can only be used for single-objective optimization. 3) This invention can respond in real time to the uncertainties of renewable energy and electricity / heat load. 4) Compared with traditional optimization scheduling methods, this invention can optimize electrothermal coupling systems online. Attached Figure Description
[0055] Figure 1 This is a reinforcement learning process for electrothermal coupling systems.
[0056] Figure 2 The training results represent multi-objective reward values.
[0057] Figure 3 The training results represent the penalty for exceeding the limit.
[0058] Figure 4 The output curves are for electrical load, heat load, and photovoltaic output.
[0059] Figure 5 The result is the optimization of the power system.
[0060] Figure 6 The results are from the optimization of the thermal system.
[0061] Figure 7 To reinforce learning methods and robust optimization methods. Detailed Implementation
[0062] A scheduling method for an electrothermal coupled system based on the twin-delay DDPG algorithm, comprising the following steps:
[0063] Step 1: Establish a low-carbon economic dispatch model for the electro-thermal coupling system, taking into account economic efficiency and carbon emissions. The energy supply side of the photovoltaic-integrated electro-thermal coupling system includes the external power grid, natural gas source, and photovoltaic power plant; the energy conversion equipment includes combined heat and power (CHP) units, electric boilers, and gas-fired boilers; and the energy demand side includes electrical load and heat load. The low-carbon economic dispatch model for the electro-thermal coupling system designed in this invention includes system operating costs and carbon emissions. The system operating costs include the cost of electricity traded with the external power grid, the cost of purchasing natural gas, and the cost of equipment operation and maintenance; carbon emissions are referenced in the "Guidelines for Enterprise Greenhouse Gas Emission Accounting and Reporting - Power Generation Facilities," and include carbon emissions from purchased electricity and CHP units.
[0064]
[0065]
[0066] In the formula, F represents the system operating cost; P t,grid Q represents the power purchased from the external power grid at time t. t,gas Let $ be the gas purchase capacity of the natural gas source at time t; t,grid Let $ be the electricity price at time t; gas The price of natural gas; α n P represents the unit maintenance cost of the nth device; t,n Let be the output of the nth device at time t; T be the operating time; N be the total number of devices; C be the carbon emissions; λ be the carbon emission factor of the external power grid; z t t represents the natural gas consumption at time t; w represents the lower heating value of natural gas; y represents the carbon content per unit calorific value of natural gas; d represents the carbon oxidation rate of natural gas; 44 / 12 represents the relative molecular mass ratio of carbon dioxide to carbon.
[0067] Power balance constraints:
[0068]
[0069] In the formula, P t,pv The output of the photovoltaic power station at time t; P t,chp P represents the electrical output of the cogeneration unit at time t. t,load Let P be the electrical load at time t; t,eb H represents the electrical power consumed by the electric boiler at time t. t,chp H represents the thermal output of the cogeneration unit at time t; t,eb H represents the thermal output of the electric boiler at time t. t,gb H represents the thermal output of the gas-fired boiler at time t. t,load The heat load at time t; Q t,chp Q represents the gas power consumed by the cogeneration unit at time t. t,gb Let t be the gas power consumed by the gas boiler.
[0070] Coupling device constraints:
[0071]
[0072] In the formula, γ chp For the gas-to-electric conversion efficiency of a combined heat and power unit; κ chp The heat-to-power ratio of a combined heat and power (CHP) unit; κ eb The electrothermal conversion efficiency of the electric boiler; κ gb This refers to the gas-heat conversion efficiency of a gas-fired boiler.
[0073] Equipment output upper and lower limit constraints:
[0074]
[0075] In the formula, and These are the upper and lower limits of the electrical output of a combined heat and power unit; and These are the upper and lower limits of the thermal output of the electric boiler, respectively. and For the upper and lower limits of the thermal output of the gas-fired boiler; g t,pv Let t be the predicted output of the photovoltaic power station.
[0076] Equipment ramping constraints: In this invention, the gas boiler is a heat supply balance unit and is not subject to ramping constraints.
[0077]
[0078] In the formula, For combined heat and power (CHP) units; For combined heat and power (CHP) units, the landslide rate is used. The ramp-up rate of the electric boiler; The landslide rate of the electric boiler.
[0079] In this embodiment, the electrothermal coupling system is coupled to a combined heat and power (CHP) unit and an electric boiler. The CHP unit has a gas-to-electricity conversion efficiency of 0.3 and a heat-to-electricity ratio of 1.2. The electric boiler has an electrothermal conversion efficiency of 0.9. The power system achieves power balance by trading electricity with the external power grid, and the thermal system achieves heat-power balance through a gas-fired boiler with a gas-to-thermal conversion efficiency of 0.7. The upper and lower limits of output and ramp-up constraints for each device are shown in Table 1. The external power grid carbon emission factor is 0.6101 tCO2 / MWh, the natural gas carbon emission factor (w×y×d×44 / 12) is 0.413 tCO2 / MWh, the natural gas price is 695 yuan / MWh, and the electricity price is shown in Table 2.
[0080] Table 1. Operating parameters of each device
[0081]
[0082] Table 2. Time-of-use electricity pricing
[0083] Time period time Electricity price (RMB / MWh) Peak hours 17-18 1092 Peak hours 9-11,19-20 732 Normal period 6-8,12-16,21-23 497.16 Valley period 0-5 294
[0084] Step 2: Transform the low-carbon economic dispatch process of the electrothermal coupling system containing renewable energy into a Markov decision process: Reinforcement learning problems can be modeled using Markov decision processes. A typical Markov decision process can be defined by five variables: [s, a, r, P, γ]. Here, s represents the state space, which is the set of all state variables that the reinforcement learning agent can perceive in the environment; a represents the action space, which represents the set of actions that the reinforcement learning agent can take in response to the environment. If the action space is a continuous variable, it is called a continuous action; if the action space is a discrete variable, it is called a discrete action; r is the reward function, which is the immediate reward returned by the environment to the decision-maker based on the state and action, and is an indicator for evaluating the state and action; γ is the reward discount rate, which represents the discount factor for future rewards; P represents the state transition probability, determined by the environment. Furthermore, in a Markov decision process, the state transition probability satisfies the Markov property, i.e., the state space s at time t is constant. t With action space a t Only the state space s at time t+1 t+1 It has an impact, but it does not affect the subsequent state; that is, the Markov decision process is without aftereffects.
[0085] The state space should consider factors that will influence decision-making as much as possible. For the problem solved by this invention, the state space includes the predicted values of electricity load, heat load, and photovoltaic power plant output for the next time step, and the current electricity price, combined heat and power unit output, and electric boiler heat output, i.e.:
[0086] s t =[g t+1,load ,h t+1,load ,g t+1,pv ,$ t,grid ,P t,chp H t,eb (7)
[0087] In the formula, g t+1,load The predicted electricity load at time t+1; h t+1,load The predicted heat load value at time t+1; g t+1,pv This represents the predicted output of the photovoltaic power station at time t+1.
[0088] The action space represents the set of changes in action variables at each time step. Action variables include: changes in electrical output of combined heat and power (CHP) units, changes in thermal output of electric boilers, and output of photovoltaic power plants. It can be represented as:
[0089] a t =[ΔP t,chp ,ΔH t,eb ,P t,pv (8)
[0090] In the formula, ΔP t,chp ΔH represents the change in electrical output of the cogeneration unit at time t. t,eb This represents the change in the thermal output of the electric boiler, while also satisfying the equipment ramp-up constraint, namely:
[0091]
[0092] In the formula, This represents the minimum ramp rate for a combined heat and power (CHP) unit. This represents the maximum ramp rate for a combined heat and power (CHP) unit. This represents the minimum ramp rate for the electric boiler. This represents the maximum ramp rate for the electric boiler.
[0093] Step 3: Designing a multi-objective reward function for reinforcement learning based on a low-carbon economic scheduling model and penalty constraint mechanism for electrothermal coupling systems: The reward function designed for electrothermal coupling systems in this invention includes: system operating cost, carbon emissions, and equipment exceeding limits penalty. Since the units for system operating cost and carbon emissions are not uniform, they are normalized:
[0094]
[0095] In the formula, f t The normalized system operating cost target; F t,max F represents the upper limit of the system operating cost at time t; t,min F represents the lower bound of the system operating cost at time t; t Let c be the system operating cost at time t; t For normalized carbon emission targets; C t C represents carbon emissions at time t; t,max C represents the upper limit of carbon emissions at time t; t,min β represents the lower limit of carbon emissions at time t; t Let be the device violation penalty factor at time t. It is 1 if any device violates the limit, and 0 if no device violates the limit. Therefore, the reinforcement learning multi-objective reward function is as follows:
[0096] r t =f t +c t -β t (11)
[0097] In the formula, r t To reinforce learning the multi-objective reward function at time t.
[0098] Step four: The Siamese Delayed DDPG algorithm is used to interactively train the reinforcement learning agent. The trained agent then performs low-carbon economic scheduling of the electrothermal coupling system. The DDPG algorithm can effectively handle most reinforcement learning processes in continuous action spaces, but it is exceptionally sensitive to hyperparameter and other types of adjustments. Therefore, the Siamese Delayed DDPG algorithm introduces three key technologies to address this issue.
[0099] The first key technology is truncated double Q-learning. The Siamese Delayed DDPG algorithm learns two Q-value functions and updates the critic network using double Q-learning, establishing two Q-value networks to estimate the value of the next state.
[0100]
[0101] In the formula, Let Q represent the Q value of policy Ⅰθ1 in state space s and action space a; Represents the policy network I in state space s; Let Q represent the Q value of policy IIθ2 in state space s and action space a; Let represent the policy network II in state space s.
[0102] Compile the Bellman equation using the minimum of the two Q values:
[0103]
[0104] In the formula, Y(s) represents the value of the current state space s; r(s) is the immediate reward in the state space s; and γ is the reward discount rate. This represents the total discount for future rewards.
[0105] The second key technique is delayed policy updates. The target network (Q-value network) is a powerful tool for achieving stable updates in reinforcement learning algorithms because function approximation requires multiple gradient updates to converge. The target network (Q-value network) provides the algorithm with a stable update target during the learning process. Therefore, if the target network (Q-value network) can be used to reduce the error of multi-step updates, and policy updates under erroneous state estimation lead to divergent policy updates, then the policy networks (Policy Network I and Policy Network II) should be updated at a lower frequency than the Q-value network to minimize the error in value estimation before performing policy updates.
[0106] The Siamese Delayed DDPG algorithm reduces the update frequency of the policy network, which is designed to be updated only after the value network has been updated d times. This policy update method can result in a smaller variance in the estimation of the Q-value function, thus achieving higher quality policy updates.
[0107] The third key technology is target policy smoothing. The DDPG algorithm may suffer from overfitting when estimating narrow peaks in the value space. In the Siamese Delayed DDPG algorithm, the idea that similar actions have similar value estimation is used to perform fuzzy fitting on the value of a small region around the target action. That is, truncated normal distribution noise is added to each action as regularization to smooth the calculation of Q value, thereby avoiding overfitting. Equation (13) is modified and updated as follows:
[0108]
[0109] In the formula, ε represents the added noise, sampled from a normal distribution, i.e., ε ~ N(0,σ); clip() is the cutoff function; c is the cutoff factor; σ is the standard deviation of the normal distribution; N(0,σ) represents a normal distribution with a standard deviation of σ; Q θ Let Q be the Q-value under strategy θ.
[0110] The trained reinforcement learning agent can respond in real time to the uncertainties of renewable energy and electricity / heat load, and perform low-carbon economic scheduling of electro-thermal coupling systems containing renewable energy online.
[0111] In this embodiment, the twin-delay DDPG algorithm in continuous action space is used for training based on the low-carbon economic scheduling environment model of the electrothermal coupling system. A total of 5000 training iterations are performed. The reinforcement learning training hardware environment is a personal computer with a CPU i7-12700H and a GPU 3060. The multi-objective reward value training results are as follows: Figure 2 As shown, the penalty for exceeding the limit constraint is as follows: Figure 3 As shown.
[0112] from Figure 2 As can be seen, the reinforcement learning agent converges around 2000 iterations. Figure 3 The results show that the penalty term for exceeding the limit gradually converges and decreases to 0 starting from 2000 iterations. Since 500 random actions were set at the beginning of training to store more action schemes in the reinforcement learning agent's cache pool, the curve shows a large jump around 500 iterations. Figure 2 The reason for the fluctuations and oscillations is that reinforcement learning agents need to continuously try and fail under the action policy to avoid getting trapped in local optima.
[0113] Call the trained reinforcement learning agent and test as follows Figure 4 The diagram shows the electricity load, heat load, and photovoltaic power station output scenarios.
[0114] The power system optimization results based on reinforcement learning are as follows: Figure 5 As shown, the optimization results of the thermal system are as follows: Figure 6As shown, the combined heat and power (CHP) units have the highest output during peak electricity price periods to achieve optimal operating costs and carbon emissions for the entire system. Electric boilers operate at full capacity except during peak periods due to their high energy conversion efficiency, thus satisfying a compromise among multiple objectives. The output of photovoltaic power plants is also fully utilized.
[0115] To further verify the optimization effect of the twin-delay DDPG algorithm used in this invention on the low-carbon economic scheduling of electrothermal coupling systems, the heuristic algorithm particle swarm optimization, the step-by-step optimal solver, and the multi-step optimal solver were used for optimization. The optimization comparison results are shown in Table 3.
[0116] Table 3 Comparison of different optimization methods
[0117]
[0118]
[0119] As shown in Table 3, the traditional heuristic Particle Swarm Optimization (PSO) algorithm yields poor optimization results. This is because PSO is prone to getting trapped in local optima, and the iterative computation required results in a long optimization time. Single-step optimal solvers, which do not consider the global optimum and use a greedy strategy to find the optimal solution at the current moment, are easily limited by constraints such as hill climbing, resulting in the worst optimization results, although they are very fast. Multi-step optimal solvers perform a large number of iterations to obtain the global optimum, resulting in the longest solution time and poor versatility, suitable only for the current scenario. The reinforcement learning-based DDPG and Siamese Delayed DDPG algorithms have the fastest optimization time and good versatility, and can also solve quickly in electro-thermal coupling systems with random source-load fluctuations, making them suitable for online optimization scenarios. The Siamese Delayed DDPG algorithm used in this invention is an improved version of the DDPG algorithm, so its optimization results are even better.
[0120] Addressing Source Load Uncertainty: The Siamese Delayed DDPG algorithm can adapt to source load uncertainty and respond online to source load fluctuations. To further verify this advantage of the Siamese Delayed DDPG algorithm, it is compared with the robust optimization method. The electrical load, heat load, and photovoltaic power plant output are fluctuated at 5%, 10%, 15%, and 20% of their normal distribution, respectively. The average optimization results of the Siamese Delayed DDPG algorithm and the robust optimization method under these random fluctuations are tested. The optimization results are as follows: Figure 7 As shown. From Figure 7As can be seen, the larger the range of random fluctuations in source load, the lower the multi-objective reward value becomes. Compared with robust optimization methods, the reinforcement learning method used in this paper has a higher multi-objective reward value. Since the Siamese Delayed DDPG algorithm does not need to specifically model uncertainty during training and can gradually learn the source and load probability distributions in the environment, the Siamese Delayed DDPG algorithm can better cope with random fluctuations in source load.
[0121] The above embodiments are merely explanations of the present invention and are not intended to limit the present invention. After reading this specification, those skilled in the art can make modifications to these embodiments without contributing any inventive step, but as long as they are within the scope of the claims of the present invention, they are protected by patent law.
Claims
1. A scheduling method for an electrothermal coupled system based on the twin-delay DDPG algorithm, characterized in that, The steps are as follows: Step 1: Establish a low-carbon economic dispatch model for the electrothermal coupling system that takes into account economic efficiency and carbon emissions, with constraints including equipment ramp-up constraints; Step 2: Transform the low-carbon economic dispatch process of the electro-thermal coupling system containing renewable energy into a Markov decision process. The Markov decision process is defined by five variables [s, a, r, P, γ]. Here, s represents the state space, which is the set of all state variables of the environment that the reinforcement learning agent can perceive; a represents the action space, which represents the set of actions that the reinforcement learning agent can take against the environment; r is the reward function, which is the immediate reward returned by the environment to the decision-making agent based on the state and action; γ is the reward discount rate; and P represents the state transition probability. The state space includes the predicted electricity load, predicted heat load, and predicted output of the photovoltaic power station at time t+1, and also includes the current electricity price, the electrical output of the cogeneration unit, and the thermal output of the electric boiler. The action space is defined as the change in electrical output of the cogeneration unit and the change in thermal output of the electric boiler at time t, and the changes satisfy the equipment ramp-up constraint. Step 3: Design a reinforcement learning multi-objective reward function based on the low-carbon economic scheduling model and penalty constraint mechanism of the electrothermal coupling system; normalize the system operating cost and carbon emissions to eliminate unit differences, and introduce a device over-limit penalty factor. When there is a device over-limit, the penalty factor is 1, and when all devices do not over-limit, the penalty factor is 0. Step 4: Use the twin-delayed DDPG algorithm to interactively train the reinforcement learning agent. The trained reinforcement learning agent will then perform low-carbon and economical scheduling of the electrothermal coupling system.
2. The scheduling method for an electrothermal coupled system based on the twin-delay DDPG algorithm according to claim 1, characterized in that, In step one, the low-carbon economic dispatch model of the electrothermal coupling system includes system operating costs and carbon emissions, as well as constraints. (1); (2); In the formula, F represents the system operating cost; Let t be the power purchased from the external power grid at time t; Let t be the gas purchase capacity of the natural gas source at time t; Let be the electricity price at time t; The price of natural gas. The unit maintenance cost of the nth device; Let n be the output of the nth device at time t; T represents operating time; N represents the total number of devices; C represents carbon emissions; Carbon emission factor of external power grid; t represents the natural gas consumption at time t; w represents the lower heating value of natural gas; y represents the carbon content per unit calorific value of natural gas; d represents the carbon oxidation rate of natural gas. Power balance constraints: (3); In the formula, The photovoltaic power station outputs power at time t; Let t be the electrical output of the cogeneration unit; Let t be the electrical load at time t; Let t be the electrical power consumed by the electric boiler. Let t be the thermal output of the cogeneration unit; Let t be the thermal output of the electric boiler. Let t be the thermal output of the gas-fired boiler. The heat load at time t; Let t be the gas power consumed by the cogeneration unit; Let t be the gas power consumed by the gas boiler; Coupling device constraints: (4); In the formula, For the gas-to-electric conversion efficiency of combined heat and power units; The heat-to-power ratio of a combined heat and power (CHP) unit; The electrothermal conversion efficiency of the electric boiler; The gas-heat conversion efficiency of a gas-fired boiler; Equipment output upper and lower limit constraints: (5); In the formula, and These are the upper and lower limits of the electrical output of a combined heat and power unit; and These are the upper and lower limits of the thermal output of the electric boiler, respectively. and These are the upper and lower limits of the thermal output of the gas-fired boiler. Let t be the predicted output value of the photovoltaic power station; Equipment ramping constraints: (6); In the formula, For combined heat and power (CHP) units; For combined heat and power (CHP) units, the landslide rate is used. The ramp-up rate of the electric boiler; The landslide rate of the electric boiler.
3. The scheduling method for an electrothermal coupled system based on the twin-delay DDPG algorithm according to claim 2, characterized in that, The state space is represented as: (7); In the formula, This represents the predicted electricity load at time t+1. The predicted heat load value at time t+1; This represents the predicted output of the photovoltaic power station at time t+1.
4. The scheduling method for an electrothermal coupled system based on the twin-delay DDPG algorithm according to claim 3, characterized in that, The action space is represented as: (8); In the formula, Let t be the change in electrical output of the cogeneration unit at time t; This refers to the change in the thermal output of the electric boiler, while also satisfying the equipment ramp-up constraint, namely: (9); In the formula, This represents the minimum ramp rate for a combined heat and power (CHP) unit. This represents the maximum ramp rate for a combined heat and power (CHP) unit. This represents the minimum ramp rate for the electric boiler. This represents the maximum ramp rate for the electric boiler.
5. The scheduling method for an electrothermal coupled system based on the twin-delay DDPG algorithm according to claim 4, characterized in that, In step three, the reward function for the electrothermal coupling system includes: system operating cost, carbon emissions, and equipment over-limit penalty; the system operating cost and carbon emissions are normalized. (10); In the formula, The normalized system operating cost target; This represents the upper limit of the system operating cost at time t. This represents the lower bound of the system operating cost at time t. Let t be the system operating cost at time t; The normalized carbon emission target; Let t be the carbon emissions at time t; The upper limit of carbon emissions at time t; The lower limit of carbon emissions at time t; Let be the device violation penalty factor at time t, which is 1 if any device violates the limit, and 0 if no device violates the limit; the reinforcement learning multi-objective reward function is shown in the following formula: (11); In the formula, To reinforce learning the multi-objective reward function at time t.
6. The scheduling method for an electrothermal coupled system based on the twin-delay DDPG algorithm according to claim 5, characterized in that, In step four, the Siamese Delayed DDPG algorithm learns two Q-value functions and updates the critic network using a double Q-value learning approach, thus establishing two Q-value networks to estimate the value of the next state: (12); In the formula, Represent policy I in state space s and action space a. Q value; Represents the policy network I in state space s; Indicates policy II in state space s and action space a. Q value; Represents the policy network II in state space s; Compile the Bellman equation using the minimum of the two Q values: (13); In the formula, This represents the value of the current state space s; The immediate reward in state space s; For reward discount rates; This represents the total discount for future rewards.
7. The scheduling method for an electrothermal coupled system based on the twin-delay DDPG algorithm according to claim 6, characterized in that, The twin-delay DDPG algorithm delays policy updates: Policy network I and policy network II are updated at a lower frequency than the Q-value network.
8. The scheduling method for an electrothermal coupled system based on the twin-delay DDPG algorithm according to claim 6, characterized in that, The twin-delayed DDPG algorithm modifies and updates equation (13) as follows: (14); In the formula, The noise to be added is sampled from a normal distribution, i.e. ; clip() is the cutoff function; c is the cutoff factor; The standard deviation of the normal distribution; The standard deviation is expressed as The normal distribution; For strategy The Q value.
Citation Information
Patent Citations
Industrial micro-grid load optimization scheduling method and system containing combined heat and power generation
CN112821465A
Optimized scheduling method for integrated energy system
CN115759604A