Integrated energy system low-carbon optimization scheduling method based on action adjustment reinforcement learning

By introducing carbon emission flow analysis method and action space adjustment mechanism in the integrated energy system, combined with the SAC algorithm, the problems of dynamic changes in carbon emission factors and low equipment operation constraint efficiency in the existing technology are solved, and the economic low-carbon optimization scheduling and efficient training of the system are realized.

CN120046784APending Publication Date: 2025-05-27HANGZHOU DIANZI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510118185.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing low-carbon optimization scheduling methods for integrated energy systems cannot effectively consider the dynamic changes of carbon emission factors over time and space, and the reinforcement learning-based method is inefficient when dealing with equipment operation constraints.

Method used

Using a reinforcement learning method based on action adjustment, a carbon emission flow analysis model of the comprehensive energy system is established by introducing a carbon emission flow analysis method, and the SAC algorithm is used for optimization scheduling. The action space adjustment phase adjusts the action space through power balance constraints, and feedbacks the regular terms of the action offset to the loss function of the policy network.

Benefits of technology

The energy consumption adjustment of the integrated energy system under carbon potential awareness is realized, ensuring the economic and low-carbon operation of the system, improving the efficiency of reinforcement learning methods in equipment operation constraint processing, and avoiding local optimal solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046784A_ABST
    Figure CN120046784A_ABST
Patent Text Reader

Abstract

The invention discloses an integrated energy system low-carbon optimization scheduling method based on action adjustment reinforcement learning. According to the method, firstly, a carbon emission flow model of an integrated energy system is established based on an EH coupling model, and the aim of minimizing the system operation cost and the carbon transaction cost is achieved; and designing reinforcement learning by taking the known state quantity of the system at the t moment as a state space and taking the output of each device in the day as an action space, and designing reward functions including cost reward and violation of system constraint penalty. In the earlier-stage exploration stage, a strategy network is used for outputting actions according to the current system state, and network parameters are updated. And when the training test reaches a preset threshold value, entering an action space adjustment stage, adjusting the action output by the strategy network at the t moment according to a constraint condition, and introducing a regular term of an action offset into a loss function of the strategy network. And performing low-carbon optimization scheduling on the integrated energy system by using the updated network, and outputting an output scheme of each device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of reinforcement learning, and relates to the low-carbon optimal scheduling of integrated energy systems. Specifically, it relates to a low-carbon optimal scheduling method for integrated energy systems based on action-adjusted reinforcement learning. Background Art

[0002] Integrated energy systems achieve efficient energy utilization through multi-energy complementarity, thereby reducing the system operation cost and carbon emissions, and contributing to the global energy transition and carbon neutrality goals. In the existing low-carbon optimal scheduling of integrated energy systems, the calculation of carbon emissions fails to fully consider the dynamic changes of carbon emission factors over time and space. Therefore, it is impossible to guide the energy consumption adjustment of integrated energy systems under the perception of carbon potential through the dynamically changing carbon emission factors.

[0003] In addition, since the low-carbon optimization model of integrated energy systems is a non-convex and non-linear mathematical model, although heuristic algorithms can effectively solve it, there are also problems such as low solution accuracy and slow convergence speed. The integrated energy system optimization scheduling method based on reinforcement learning does not require accurate model information. During training, historical data is used to train a neural network to establish a function mapping network from "state" to "action", which can effectively solve the non-linear optimization model and obtain a low-carbon optimal scheduling scheme for integrated energy systems. However, when dealing with the operation constraints of each device in integrated energy optimization scheduling, existing methods usually simply use a reward function penalty for soft constraints, resulting in low intelligent training efficiency of reinforcement learning methods. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention proposes a low-carbon optimal scheduling method for integrated energy systems based on action-adjusted reinforcement learning. The carbon emission flow analysis method in the power system is introduced into the integrated energy system, and an integrated energy system carbon emission flow analysis model is established based on the carbon emission virtual equivalent network of the energy hub (EH). On this basis, with the minimum of the integrated energy system operation cost and carbon trading cost as the optimization goal, the SAC (Soft Actor-Critic) algorithm in reinforcement learning is used for solution. During the training process of the policy network parameters of the SAC algorithm, the action space is adjusted through power balance constraints, and the action space adjustment process is fed back to the loss function of the policy network through a regularization term of the action offset. An accurate and efficient low-carbon optimal scheduling scheme is output using the trained policy network.

[0005] The low-carbon optimal scheduling method for integrated energy systems based on action-adjusted reinforcement learning specifically includes the following steps:

[0006] Step 1: Establish an integrated energy system model

[0007] Step 1.1: Establish a carbon emission flow model of the integrated energy system based on the EH carbon emission virtual equivalent network

[0008] The devices in the integrated energy system are divided into single-input single-output devices, single-input multi-output devices, and energy storage devices. According to the principle of "carbon emission conservation", a port carbon intensity (PCI) model of the single-input single-output device i is constructed:

[0009]

[0010] where and η i are the PCI of the single-input single-output device i, the generator carbon intensity (GCI) of the energy supply side corresponding to the input end energy, and the energy conversion efficiency, respectively.

[0011] For the single-input multi-output device j, according to the principle of "carbon emission conservation", a PCI model of the output end is constructed:

[0012]

[0013] where and are the PCI of the two energy output ends of the single-input multi-output device j at time t; ρ j,t is the GCI of the energy corresponding to the input end of device j; and and are the energy conversion efficiencies of device j, respectively.

[0014] Referring to the calculation method of the SOC of the energy storage device, a carbon holding rate SOCS model of the m-th type of energy storage device at time t is constructed m,t model.

[0015] Step 1.2, Construct a low-carbon optimal scheduling model of the integrated energy system based on carbon flow analysis

[0016] Taking the minimum of the operating cost C O and the carbon trading cost of the integrated energy system as the goal, a target function is constructed:

[0017]

[0018] where ω o and are the operating cost coefficient and the carbon trading cost coefficient of the integrated energy system, respectively. The constraint conditions of the model are established, including power balance constraints, device operation constraints, and carbon trading constraints.

[0019] Step 2, Design reinforcement learning

[0020] Taking the equipment operation status, output information, and consumer-side demand of the integrated energy system in period t as the state space s t , and the output of each device in the system within a day as the action space a t . The reward function r at time t t includes the cost reward r t c and the penalty r for violating system constraints t p .

[0021] Obtain the historical data of wind, light, and load of the integrated energy system for K days, and set the maximum number of iterations N. Initialize the iteration number n = 0, the training days k = 0, the time period t = 0, initialize the experience buffer D, the policy (Actor) network parameters ψ, and the evaluation (Critic) network parameters θ. Use 80% of the historical data as the training set and 20% as the test set.

[0022] Step 3. Train the reinforcement learning algorithm

[0023] Step 3.1. Preliminary exploration stage

[0024] Take the current period state s of the integrated energy system t as the input of the policy network, and output the action a through the policy network t . And calculate the reward function r at the current time period t , and form a quadruple (s t+1 , a t , r t , s t , s t+1 ), store it in the experience buffer D, and t = t + 1.

[0025] When the size B of the experience buffer > B max , randomly extract a sample of size L batches from the experience buffer, calculate the loss value, and update the parameters ψ and θ of the policy network and the evaluation network.

[0026] Repeat Step 3.1. When t = T, let k = k + 1, t = 1, and continue to repeat Step 3.1 until k = K to complete one training, and let n = n + 1. Where T represents the data length of a natural day.

[0027] Step 3.2. Action space adjustment stage

[0028] As the exploration stage progresses, in order to enable the agent to learn the operation constraints of each device in the integrated energy system more efficiently, an action space adjustment stage is required. When the training times n reach the set threshold, set the process of updating the evaluation network unchanged. When the policy network outputs an action, add an action space adjustment layer, and change the action at time t from Adjusted to

[0029]

[0030] Among them, P t m and are the power of the output energy type τ of the single-input multi-output device j at time t, the power output by the energy storage device m, and the power of the output energy type τ of the single-input single-output device i; P t m,safe and are the power values after action adjustment respectively; is the load demand for the energy type τ; is the G-th adjusted power of the single-input multi-output device j; and are the upper and lower limits of the power of the single-input single-output device i outputting the energy type τ; The energy types include electricity, heat, and cold.

[0031] In order to enhance the performance of the agent, the process of action adjustment is fed back into the policy network, and the regularization terms reg1 to regA of the action offset are added to the loss function L(ψ) of the policy network to enhance the training efficiency:

[0032] L(ψ) safe = L(ψ) + ω 1 reg1 + ω 2 reg2 + ω 3 reg3 +... + ω A regA (8)

[0033]

[0034] Among them, ω 1 ~ω A are the weights of the offset regularization terms, and A is the dimension of the action space; L(ψ) safe represents the loss function of the policy network in the action space adjustment stage. a l,t (A) and are the elements in the A-th column of the action space and their action adjustment values in the l-th sample in batch L.

[0035] Step 3.3: Continuously update the parameters of the evaluation network and the policy network until n = N. Save the network parameters of the last training.

[0036] Step 4: Low-carbon optimal scheduling of the integrated energy system

[0037] Input the state data of the integrated energy system into the trained policy network to output the day-ahead scheduling plan.

[0038] The present invention has the following beneficial effects:

[0039] (1) Introduce the carbon emission flow analysis method of the power system into the integrated energy system, and establish an integrated energy system carbon emission flow analysis model based on the carbon emission virtual equivalent network of the energy hub (EH). In the low-carbon optimal scheduling model, the coordinated complementary output of multiple devices in the system is realized by minimizing the operating cost and carbon trading cost, and the carbon emission responsibility is shared. Ensure the economic and low-carbon operation of the integrated energy system.

[0040] (2) Use the SAC algorithm considering action adjustment to solve the non-convex optimization problem. Aiming at the problem that the conventional SAC algorithm cannot meet the device operation constraints in the penalty soft constraint added to the reward function, an action adjustment stage is added in the update of the policy network parameters, and the action adjustment layer is combined with the power balance constraint to further limit the value range of the action space. So that it can meet the operation constraints of the device, reduce the number of ineffective explorations of the intelligent agent, ensure the stability of the intelligent agent training process, and avoid falling into local optima. Further, the training efficiency of the algorithm can be improved.

[0041] (3) Add a regular term of action offset as feedback information to the policy network loss function in the safe return stage, so that the policy network can better learn the behavior of adjusting the action space. The policy network updates its parameters through backpropagation, reducing the dependence of the intelligent agent on action adjustment. Description of the Drawings

[0042] Figure 1 It is a schematic diagram of the integrated energy system structure.

[0043] Figure 2 It is a flowchart of the low-carbon optimal scheduling method for the integrated energy system based on action-adjusted reinforcement learning. Detailed Embodiments

[0044] The following further explains and illustrates the present invention with reference to the drawings;

[0045] The low-carbon optimal scheduling method for the integrated energy system based on action-adjusted reinforcement learning specifically includes the following steps:

[0046] Step 1. Establish an integrated energy system model

[0047] Step 1.1. Establish a carbon emission flow model of the integrated energy system based on the EH carbon emission virtual equivalent network

[0048] In this embodiment, low-carbon optimal scheduling is carried out for an integrated energy system with electricity, cooling, and heating load demands on the consumer side, such as Figure 1As shown in the figure, the power system of the integrated energy system includes photovoltaic (PV), wind turbine (WT), battery storage system (BS), and the superior power grid. The refrigeration system includes absorption chiller (AC) and electric chiller (EC). The thermal system consists of heat supply and heat storage device (HS). The coupling devices include combined heat and power (CHP) and heat pump (HP).

[0049] According to the principle of "carbon emission conservation", establish the port carbon intensity (PCI) models for single-input single-output devices HP, AC, and EC, and single-input multi-output device CHP:

[0050] ρ HP,t =ρ e,t / η HP (10)

[0051] ρ AC,t =ρ g,t / η AC (11)

[0052] ρ EC,t =ρ e,t / η EC (12)

[0053]

[0054] Among them, ρ HP,t 、ρ AC,t and ρ EC,t are the PCI of the output ports of HP, AC, and EC at time t respectively. and are the PCI of the power generation side and the heat generation side of CHP at time t respectively, with the unit of kgCO2 / (kW·h). ρ e,t 、ρ g,t are the generator carbon intensities (GCI) of electric energy and natural gas at time t respectively; η HP 、η AC and η EC are the electro-thermal conversion efficiency of HP, the thermal-cooling conversion efficiency of AC, and the electro-cooling conversion efficiency of EC respectively. and are the power generation efficiency and heat generation efficiency of CHP respectively.

[0055] Step 1.2: Construct a carbon emission flow model for the energy storage equipment in the integrated energy system

[0056] Referring to the calculation method of the state of charge (SOC) of the energy storage device, the state of charge of carbon (SOCS) of the m-th type of energy storage equipment at time t is expressed by the "carbon charge rate". m,t Model:

[0057]

[0058] Among them, and respectively represent the carbon emissions charged and released by the m-th type of energy storage equipment at time t; SOC m,t is the state of charge of the m-th type of energy storage equipment at time t; E m,cap is the rated capacity of the m-th type of energy storage equipment; and are respectively the charging and discharging powers of the m-th type of energy storage equipment at time t; and are respectively the carbon flow densities at the charging and discharging ports of the m-th type of energy storage equipment at time t; Δt is the scheduling interval time.

[0059] Step 1.3: Construct a low-carbon optimal scheduling model for the integrated energy system based on carbon flow analysis

[0060] Taking the minimum of the operating cost C O of the integrated energy system and the carbon trading cost as the objective, the following objective function is constructed:

[0061]

[0062] Among them, ω o and are respectively the operating cost coefficient and the carbon trading cost coefficient of the integrated energy system.

[0063] The operating cost C O of the integrated energy system includes the cost C b of purchasing electricity from the power grid, the revenue C s from selling electricity to the power grid, and the gas cost C CHP of the CHP:

[0064]

[0065] Among them, P t buy and P t sell are the power of purchasing electricity from the power grid and the power of selling electricity to the power grid by the integrated energy system at time t; c sub,t , c sell,t and c gasThey are the electricity purchase price from the power grid, the electricity sale price to the power grid, and the natural gas price during period t; Q LHV represents the lower calorific value of natural gas; represents the CHP power generation efficiency, and T represents the total dispatching duration; is the generated power output by the CHP during period t.

[0066] Use a stepped carbon trading model to model the carbon trading cost as follows:

[0067]

[0068] V IES,t = V IES - V IES,a (20)

[0069]

[0070] where ρ is the carbon trading floor price; x is the length of the carbon emission range; z is the price growth rate; and are the electric power consumed by the EC, the thermal power consumed by the AC, the electric power consumed by the HP, the thermal power generated by the CHP, the electric power released by the BS, and the thermal power released by the HS during period t, respectively; SOCS BS,t and SOCS HS,t are the carbon load rates of the BS and HS during period t, respectively. V IES,t is the carbon emission trading volume of the integrated energy system during period t; V IES is the actual carbon emission of the integrated energy system. V IES,a is the carbon emission quota of the integrated energy system, and the initial quota is determined by the baseline method:

[0071] V IES,a = V grid + V CHP + V Gas (22)

[0072]

[0073] where V grid 、V CHP and V Gas are the carbon emission right allocation quotas for purchasing electricity from the power grid, CHP heat generation, and natural gas consumption, respectively; χ e is the carbon emission right quota coefficient for generating unit electric power, χ h is the carbon emission right quota coefficient for generating unit thermal power, χ g is the carbon emission right quota coefficient for consuming unit electric power, represents the natural gas power consumed by the CHP during period t.

[0074] Step 1.4, Construct Constraint Conditions

[0075] Establish the constraint conditions of the integrated energy system, where the power balance constraint is:

[0076]

[0077] Among them, P t w and P t s are the photovoltaic power generation output and wind power generation output at time t; and are the electricity, heat, and cooling loads of the integrated energy system respectively; and are the charging power of the BS, the heat storage power of the HS, the heat power output by the HP, the cooling power output by the AC, and the cooling power output by the EC respectively.

[0078] The constraint on the power interaction of the integrated energy system with the superior power grid is:

[0079]

[0080] Among them, and are the upper limits of the power purchase and power sale of the integrated energy system from / to the superior power grid; and are the state variables of power purchase and sale at time t.

[0081] The CHP operation constraint is:

[0082]

[0083] Among them, η loss and are the heat loss rate and heat production coefficient of the CHP respectively; and are the upper and lower limits of the electrical output of the CHP; r is the upper ramp limit of the CHP.

[0084] The wind and light output constraint is:

[0085]

[0086] Among them, P t pw 、P t ps are the maximum values of photovoltaic power generation and wind power generation

[0087] The BS operation constraint is:

[0088]

[0089] Among them, E t BS is the electricity storage capacity of the BS at time t; is the rated capacity of the BS; is the self-loss rate of the BS's electricity storage, taking 0.01; and are the charging and discharging efficiencies of the HS, both taking 0.95; and are the upper and lower limits of the BS's electricity storage state, with values of 0.1 and 0.95 respectively. and are the state 0-1 variables of charging and discharging respectively. and are the maximum charging and discharging powers of the ES respectively.

[0090] The operation constraint of the HP is:

[0091]

[0092] Among them, and are the upper and lower limits of the HP's thermal output.

[0093] The operation constraint of the HS is:

[0094]

[0095] Among them, is the heat storage capacity of the HS at time t; is the rated capacity of the HS. is the self-loss rate of the HS's heat storage, taking 0.01; and are the heat charging and discharging efficiencies of the HS, both taking 0.9; and are the upper and lower limits of the HS's heat storage state, with values of 0.1 and 0.9 respectively. and are the state 0-1 variables of heat charging and discharging respectively; and are the maximum heat charging and discharging powers of the ES respectively.

[0096] The operation constraint of the EC is:

[0097]

[0098] Among them, and are the minimum and maximum powers of the EC's output cold power.

[0099] The operation constraint of the AC is:

[0100]

[0101] Among them, and are the minimum and maximum powers of the AC output cooling.

[0102] Step 2: Design reinforcement learning

[0103] Step 2.1: Use the known information that the integrated energy system can obtain at time period t as the state space s t , including wind and light output, electric, heat, and cooling loads, electricity price information, the electricity storage capacity of the ES in the previous time period, the heat storage capacity of the HS in the previous time period, and the output of the CHP in the previous time period:

[0104]

[0105] Step 2.2: Use the output of each device in the integrated energy system within day at time period t as the action space a t . To reduce the dimension of the action space, only select the electric power output by the CHP, the electric power output by the BS, the heat power output by the HS, and the cooling power output by the AC P t BS , P t HS , to form the action space a t . The output information of the remaining devices is obtained through constraint relationships:

[0106]

[0107] Among them,

[0108] Step 2.3: The reward function r t includes the cost reward r t c and the penalty r t p for violating system constraints:

[0109] r t = r t c + r t p (39)

[0110] r t c = -C O (40)

[0111] r t p = r t grid + r tCHP +r t BS +r t HS +r t HP +r t EC (41)

[0112] where r t grid , r t CHP , r t BS , r t HS , r t HP , and r t EC are the penalties for violating the interaction constraints of the superior power grid, the penalty for violating the CHP operation, the penalty for violating the BS operation, the penalty for violating the HS operation, the penalty for violating the HP operation, and the penalty for violating the EC operation, respectively:

[0113]

[0114] and are the penalty coefficients for violating the power grid interaction constraints, the penalty coefficient for violating the MT operation, the penalty coefficient for violating the ES operation, the penalty coefficient for violating the HP operation, and the penalty coefficient for violating the EC operation, respectively. min() and max() represent taking the minimum value and the maximum value, respectively, and || represents calculating the absolute value.

[0115] Step 3. Train the reinforcement learning algorithm

[0116] As Figure 2 shown, training the SAC algorithm is divided into the preliminary exploration and action space adjustment stages. After convergence, the network parameters are saved. The specific steps are as follows:

[0117] Step 3.1. Obtain the historical data of wind, light, and load of the integrated energy system with K = 30. Use the historical data of the first 24 days as the training set and the data of the last 6 days as the test set.

[0118] Set the maximum number of iterations N. Initialize the iteration number n = 0, the training days k = 0, the time period t = 0, initialize the experience buffer D, the policy (Actor) network parameters ψ, and the evaluation (Critic) network parameters θ.

[0119] Step 3.2. In the preliminary exploration stage, take the current time period state s t of the integrated energy system as the input of the policy network, and output the action a t through the policy network. And calculate the reward function r of the current time periodt , together with the state s of the integrated energy system at the next time step t+1 to form a quadruple (s t , a t , r t , s t+1 ), which is stored in the experience buffer D, and t = t + 1.

[0120] Step 3.3: When the size B of the experience buffer is greater than B max , randomly sample a batch of size L from the experience buffer, calculate the loss value, and update the parameters ψ and θ of the policy network and the evaluation network.

[0121] To enhance the exploration of the action space, the SAC algorithm adopts the maximum entropy policy to maximize the entropy of the action and improve the global search ability of the algorithm. Its objective function maxJ(π) is:[[]]

[0122]

[0123] where ρ π is the state-action distribution generated by the policy π; J(π) is the expected cumulative reward of the policy, α is the temperature coefficient of entropy; π(·|s t ) is the probability distribution of the action a t generated by the policy network in the state s t ; is the policy entropy:[[]]

[0124]

[0125] Maximize the cumulative return J(π) by updating the parameters of the policy network and the evaluation network. Use the evaluation network to evaluate the current policy and estimate the Q value Q(s t , a t ) of each state. Update the parameters of the policy network through the modified Bellman equation:[[]]

[0126]

[0127]

[0128] where y t represents the target value of the policy network, V(·) represents the value output by the target value network for the state s t , γ is the discount factor; p is the state transition probability distribution of the environment.

[0129] Define the loss function L(θ) of the evaluation network as:[[]]

[0130]

[0131] where α cis the learning rate of the evaluation network. represents the gradient of L(θ) with respect to the parameter θ. To ensure the stability of training, the parameters of the evaluation network are used to softly update the parameters θ of the target value network ^ as follows:

[0132] θ ^ t+1 ←hθ t+1 +(1 - h)θ ^ t (54)

[0133] where h is the soft update parameter.

[0134] The loss function L(ψ) of the policy network is defined using the Kullback-Leibler (KL) divergence between the target policy and the policy:

[0135]

[0136] where π ψ (a|s) represents the parameterized policy network, which is the probability distribution of taking action a given state s; is the target policy; D KL (·) is the KL divergence, which is used to measure the difference between two probability distributions; E[·] is the expectation symbol.

[0137] To enhance the exploration of actions, Gaussian-distributed random noise ε is added to the output of the policy network t :

[0138]

[0139] where f(·) is the mapping function of the policy network, is the mean and standard deviation of the action a output by the policy network. The loss function L(ψ) of the policy network is defined as: t The policy network parameters are updated through backpropagation using the Adam optimizer:

[0140]

[0141] where α

[0142]

[0143] is the learning rate of the policy network; a is the gradient of the loss function L(ψ) with respect to the parameter ψ.

[0144] ​Step 3.4: Repeat Steps 3.2 and 3.3. When \(t = T\), let \(k=k + 1\) and \(t = 1\), and then continue to repeat the steps until \(k = 24\) to complete one training session. Then let \(n=n + 1\). Here, \(T\) represents the data length of one natural day.

[0145] Step 3.5: As the exploration phase progresses, to more efficiently learn the operating constraints of each device in the integrated energy system, when the number of training sessions \(n\) reaches the set threshold, enter the action space adjustment phase. Specifically, after the policy network outputs an action, evaluate whether the action satisfies the device operating constraints, and according to the evaluation results, adjust the action at time \(t\) from to

[0146]

[0147] where, P t BS,safe and \(P t HS,safe and are the power values after action adjustment respectively. is the cooling load demand; and are the powers of two CHP adjustments respectively; is the upper limit of the HP output heat power; \(P t grid and and \(P t grid,safe are the power of interaction with the superior power grid, the upper limit of the power of interaction with the superior power grid, the lower limit of the power of interaction with the superior power grid, and the power of interaction with the superior power grid after action adjustment respectively.

[0148] To enable the policy network to better learn the adjustment process and enhance the performance of the policy network, feedback the action adjustment process to the policy network, and add regularization terms \(reg1\sim reg4\) of the offset to the loss function of the policy network to enhance the training efficiency:

[0149]

[0150] where, and are the \(l\)-th samples in batch \(L\) respectively.

[0151] The loss function \(L(\psi)\) of the policy network in the action space adjustment phase safe is:

[0152] \(L(\psi) safe =L(\psi)+\omega 1 reg1+\omega 2 reg2+\omega3 reg3 + ω 4 reg4(71)

[0153] where ω 1 ~ω 4 is the weight of the offset regularization term.

[0154] Step 3.6: Repeat Steps 3.4 and 3.5 to continuously update the parameters of the evaluation network and the policy network until n = N. Save the network parameters of the last training.

[0155] Use the test set data to perform integrated energy system scheduling on the network saved in Step 3.6. The device parameters of the system are shown in Table 1:

[0156] Table 1

[0157]

[0158]

[0159] The specific parameters of the policy network and the evaluation network are shown in Table 2:

[0160] Table 2

[0161]

[0162]

[0163] By comparing the traditional integrated energy system optimization scheduling model with the low-carbon optimization scheduling model of the integrated energy system proposed by this method. Select one-day data from the validation set to verify the feasibility and superiority of the integrated energy system model proposed by the invention. The optimization results of the two optimization scheduling models are shown in Table 3.

[0164] Table 3

[0165]

[0166] It can be seen that the model proposed by this method has a 0.6% higher operating cost, a 22.2% lower carbon trading cost, and a 21.5% lower carbon emission than the traditional model. It can further optimize the selection of the energy types and usage times required by different loads of the integrated energy system from the perspective of the carbon potential of each device, and the carbon emissions and carbon emission costs of the system are significantly reduced, indicating that the model proposed by this method can ensure the effective promotion of overall carbon emission reduction under the operating economy of the integrated energy system.

[0167] To verify the superiority of the reinforcement learning process of this method compared with traditional intelligent algorithms for solving the low-carbon optimization scheduling of integrated energy systems. Compare the traditional intelligent optimization algorithm (particle swarm algorithm) with this method, and the average solution time is shown in Table 4:

[0168] Table 4

[0169]

[0170] It can be seen that this method greatly reduces the average solution time compared with the method using the intelligent optimization algorithm. Although offline training takes a long time, saving the trained model parameters will improve the efficiency of low-carbon optimal scheduling of the integrated energy system during day-ahead optimal scheduling.

[0171] Compare the operating costs of this method and the particle swarm algorithm for scheduling the integrated energy system on the test set data. The results are shown in Table 5:

[0172] Table 5

[0173]

[0174] As can be seen from Table 5, the cost of this method is lower than that of the particle swarm algorithm, and the solution accuracy is higher than that of the particle swarm algorithm, verifying the effectiveness and superiority of the algorithm.

[0175] To prove the superiority of the action space adjustment proposed by this method, it is compared with the DQN algorithm, DDPG algorithm, and SAC algorithm. The results are shown in Table 6:

[0176] Table 6

[0177]

[0178] As can be seen from Table 6, this method requires fewer training times to converge, has high training efficiency, and has the minimum cost at convergence. DQN uses a discrete action space, and the accuracy of its solution is related to the discreteness of the action space; DDPG is a policy-based reinforcement learning algorithm that uses a continuous action space, but because it uses a deterministic policy, it is easy to fall into a local optimum; the SAC algorithm uses the maximum entropy criterion to make the agent further explore. However, these three algorithms have soft constraints on the device operation constraints and cannot guarantee that they fully meet the constraints during training, and violations of the constraints will occur. After introducing the action space adjustment and feedback mechanism, this method will not have violations of the constraints, enhancing the performance of the algorithm.

Claims

1. A low-carbon optimization scheduling method for an integrated energy system based on action adjustment reinforcement learning, characterized by: The specific steps include: Step 1: According to the principle of conservation of carbon emissions, a port carbon flow density model of each device in the integrated energy system is constructed; Calculate the SOCS of various energy storage devices in period t m,t ; The operating cost of the comprehensive energy system is C O and carbon trading costs Minimum is the goal, construct the objective function: Among them, ω o and They are the operating cost coefficient of the integrated energy system and the carbon trading cost coefficient; the constraints for establishing the model include power balance constraints and equipment operation constraints; Step 2: Take the equipment operation status, output information and consumer demand of the integrated energy system at time t as the state space s t , the output of each device in the system is taken as the action space a t ; Reward function r at time t t Including cost reward t c and the penalty r for violating the system constraint t p ; Obtain historical data of the integrated energy system and set the maximum number of iterations N; initialize the number of iterations n = 0, the number of training days k = 0, the time t = 0, and initialize the experience buffer D, the strategy network parameter ψ, and the evaluation network parameter θ; Step 3: Train the reinforcement learning algorithm, including the early exploration phase and the action space adjustment phase. Save the network parameters after convergence. The specific steps are as follows: Step 3.1: The current state of the integrated energy system s t As the input of the policy network, the policy network outputs action a t ; and calculate the reward function r at the current moment t , and the next state s of the integrated energy system t+1 Form a four-tuple (s t ,a t ,r t ,s t+1 ), stored in the experience buffer D, t = t + 1; When the size of the experience buffer B>B max , randomly extract samples of size L batches from the experience buffer, calculate the loss value, and update the parameters ψ and θ of the policy network and the evaluation network; When t=T, let k=k+1, t=0, repeat until k=K, complete one training, let n=n+1; where T represents the data length of a natural day, K represents the historical data length of the integrated energy system; when the number of training times n reaches the set threshold, go to step 3.2; Step 3.2: Set the evaluation network update process unchanged, and set the action a output by the strategy network at time t according to the constraints. t Adjust to And add the regularization terms reg1~regA of action offset to the loss function L(ψ) of the policy network: L(ψ) safe =L(ψ)+ω1reg1+ω2reg2+ω3reg3+...+ω A regA (2) Among them, ω1~ω A is the weight of the offset regularization term, A is the dimension of the action space, L(ψ) safe represents the policy network loss function in the action space adjustment phase; Continue to update the parameters of the evaluation network and the policy network until n=N; save the network parameters of the last training; Step 4: Low-carbon optimization scheduling of integrated energy systems The state data of the integrated energy system is input into the trained strategy network, and the day-ahead scheduling plan is output.

2. The low-carbon optimization scheduling method for an integrated energy system based on action adjustment reinforcement learning as claimed in claim 1, characterized in that: The energy conversion in the integrated energy system is divided into single-input-single-output devices and single-input-multiple-output devices. The port carbon flow density of the single-input-single-output device i is: in, and η i They are the port carbon flow density of single-input-single-output device i, the carbon emission intensity of the power generation side of the energy corresponding to the input end, and the energy conversion efficiency; The port carbon flow density of a single-input-multiple-output device j that can output electrical energy and thermal energy is: in, and are the port carbon flow density on the power generation side and the port carbon flow density on the heating side of the single-input-multiple-output device j in period t respectively; ρ j,t is the carbon emission intensity of the energy corresponding to the input end of device j; and and are the power generation efficiency and heating efficiency of device j, respectively.

3. The low-carbon optimization scheduling method for an integrated energy system based on action adjustment reinforcement learning as claimed in claim 1, characterized in that: Modeling is performed on the integrated energy system with three load demands, namely electricity, cooling and heat, on the consumer side. The power system of the integrated energy system includes photovoltaic PV, wind turbine WT, battery energy storage system BSS and upper power grid, the refrigeration system includes absorption chiller AC and electric chiller EC, the thermal system includes heat supply and heat storage device HS, and the coupling equipment includes cogeneration device CHP and heat pump HP; The operating cost of the comprehensive energy system C O Including the cost of purchasing electricity from the grid C b , Revenue from selling electricity to the grid C s and CHP gas cost C CHP : Among them, P t buy , P t sell is the power purchased from the grid and the power sold to the grid by the integrated energy system during period t; c sub,t 、c sell,t and c gas are the electricity purchase price from the grid, the electricity sale price to the grid and the natural gas price during period t, Q LHV Indicates the lower calorific value of natural gas; Indicates the CHP power generation efficiency; is the power generated by CHP during period t; Using the Tiered Carbon Trading Model to Calculate Carbon Trading Costs To model: V IES,t =V IES -V IES,a (7) Among them, ρ is the carbon trading floor price; x is the length of the carbon emission range; z is the price growth rate; and They are the electric power consumed by EC, the thermal power consumed by AC, the electric power consumed by HP, the thermal power generated by CHP, the electric power released by BS and the thermal power released by HS during period t; SOCS BS,t and SOCS HS,t are the carbon charging rates of BS and HS in period t; ρ HP,t , AC,t are the port carbon flow densities of the HP and AC output ports at time t, is the port carbon flow density on the heating side of CHP during period t; V IES,t V is the carbon emission trading volume of the comprehensive energy system at time t; IES is the actual carbon emissions of the comprehensive energy system; V IES,a For the carbon emission quota of the comprehensive energy system, the initial quota is determined by the baseline method: V IES,a =V grid +V CHP +V Gas (9) Among them, V grid 、V CHP and V Gas are the carbon emission rights allocation quotas for purchasing electricity from the grid, producing heat with CHP, and consuming natural gas; e is the carbon emission quota coefficient for generating unit electric power, χ h is the carbon emission quota coefficient for generating unit thermal power, χ g is the carbon emission quota coefficient per unit of electric power consumed, It represents the natural gas power consumed by CHP during period t.

4. The low-carbon optimization scheduling method for an integrated energy system based on action adjustment reinforcement learning as claimed in claim 3, characterized in that: The wind and solar output P at time t t WT , P t PV , Electric heating and cooling load Electricity price information sub,t 、c sell,t , ES power storage at the last moment HS heat storage capacity at the last moment and CHP output at the previous moment As the state space s t : The electric power output of CHP, the electric power output of BS, the thermal power output of HS and the cooling power output of AC at time t P t BS , P t HS , As the action space a t ; The reward function r t for: r t =r t c +r t p (12) r t c =-C O (13) r t p =r t grid +r t CHP +r t BS +r t HS +r t HP +r t EC (14) Among them, r t grid 、r t CHP 、r t BS 、r t HS 、r t HP and r t EC They are respectively the penalty for violating the interaction constraint of the upper power grid, the penalty for violating the CHP operation, the penalty for violating the BS operation, the penalty for violating the HS operation, the penalty for violating the HP operation and the penalty for violating the EC operation: and They are the penalty coefficients for violating grid interaction constraints, violating MT operation, violating ES operation, violating HP operation, and violating EC operation; min() and max() represent taking the minimum and maximum values, respectively, and || represents calculating the absolute value; and It is the upper limit of the power that the integrated energy system can purchase and sell to the upper-level power grid; is the power storage capacity of BS in period t; is the rated capacity of the BS; and They are the upper and lower limits of BS power storage state respectively; and The upper and lower limits of HP thermal output; is the heat storage capacity of HS in period t; is the rated capacity of HS; and are the upper and lower limits of HS heat storage state respectively.

5. The low-carbon optimization scheduling method for an integrated energy system based on action adjustment reinforcement learning as claimed in claim 1, characterized in that: In the early exploration stage, the objective function maxJ(π) is set as: Among them, ρ π is the state-action distribution generated by strategy π; J(π) is the expected cumulative reward of the strategy, α is the temperature coefficient of entropy; π(·|s t ) is the policy network in state s t The action a generated by t The probability distribution of is the strategy entropy; Use the evaluation network to evaluate the current strategy and estimate the Q value Q(s) of each state t ,a t ), update the parameters of the policy network: Among them, y t represents the target value of the policy network, V(·) represents the target value network for state s t The output value, γ is the discount factor; p is the state transition probability distribution of the environment; The loss function L(θ) of the evaluation network is defined as: Among them, α c is the learning rate of the evaluation network; represents the gradient of L(θ) with respect to the parameter θ; Use the evaluation network parameters to soft update the parameters θ^ of the target value network: θ^ t+1 ←hθ t+1 +(1-h)θ^ t (26) Among them, h is the soft update parameter; Add random noise ε that conforms to the Gaussian distribution to the output of the policy network t : Where f(·) is the mapping function of the policy network, is the policy network output action a t The mean and standard deviation of ; the loss function L(ψ) of the policy network is defined as: Where E[·] is the expected value symbol; back propagation is performed through the Adam optimizer to update the policy network parameters: Among them, α a is the learning rate of the policy network; is the gradient of the loss function L(ψ) with respect to the parameter ψ.

6. The low-carbon optimization scheduling method for an integrated energy system based on action adjustment reinforcement learning as claimed in claim 3, characterized in that: In the action space adjustment phase, the action output by the policy network is evaluated to see whether it meets the device operation constraints, and the action at time t is changed from Adjust to in, P t BS , P t HS and They are the electrical power output by CHP, the electrical power output by BS, the thermal power output by HS and the cooling power output by AC; P t BS,safe , P t HS,safe and They are the power values ​​after action adjustment; and They are the upper limit of BS power storage state, the rated capacity of BS, the power storage amount of BS in period t and the lower limit of BS power storage state; and They are the upper limit of the HS heat storage state, the rated capacity of the HS, the heat storage amount of the HS in period t, and the lower limit of the HS heat storage state; is the cooling load demand; and They are the powers adjusted twice by CHP; and is the upper limit of the thermal power output by HP and the thermal power output by HP in period t; P t grid , and P t grid,safe They are the power of interaction with the upper-level power grid, the upper limit of the power of interaction with the upper-level power grid, the lower limit of the power of interaction with the upper-level power grid and the power of interaction with the upper-level power grid after action adjustment.

7. The low-carbon optimization scheduling method for an integrated energy system based on action adjustment reinforcement learning as claimed in claim 6, characterized in that: The regular terms reg1 to reg4 of the offset are: in, and are the lth sample in batch L respectively.

8. The low-carbon optimization scheduling method for an integrated energy system based on action adjustment reinforcement learning as claimed in claim 6, characterized in that: The loss function L(ψ) of the policy network in the action space adjustment phase safe for: L(ψ) safe =L(ψ)+ω1reg1+ω2reg2+ω3reg3+ω4reg4 (42) Among them, ω1~ω4 are the weights of the offset regularization term.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Comprehensive energy system low-carbon optimization scheduling method based on deep reinforcement learning

    CN121745696A

  • Supply chain carbon footprint real-time pricing method and system based on multi-modal data

    CN122243558A