Multi-time scale power regulation method for integrated energy system based on deep reinforcement learning

Through a method based on deep reinforcement learning, an integrated energy system model of electricity, thermal energy, natural gas and hydrogen energy was constructed, and a two-layer multi-time-scale power control strategy was established. This solved the problems of source-load uncertainty and insufficient multi-time-scale control in existing technologies, and achieved stable and low-carbon operation of the integrated energy system.

CN118763740BActive Publication Date: 2025-10-10HEFEI UNIV OF TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411135369.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2025-10-10
Estimated Expiration
2044-08-19

AI Technical Summary

Technical Problem

Existing technologies are insufficient in multi-time-scale control of power control in integrated energy systems under source-load uncertainty environments, making it difficult to achieve stable and low-carbon operation of the system. In addition, random optimization and robust optimization methods have problems with data accuracy and solution complexity in practical applications.

Method used

Using a method based on deep reinforcement learning, a comprehensive energy system model of electricity, thermal energy, natural gas and hydrogen energy is constructed, and a two-layer multi-time-scale power control strategy is established. The upper and lower layer intelligent agents are trained through deep reinforcement learning algorithms, and the operation of the system is dynamically adjusted to cope with source and load fluctuations. The power control strategy model is trained using deep reinforcement learning algorithms to coordinate the operation of energy supply and energy storage devices.

Benefits of technology

It achieves stable and low-carbon operation of the integrated energy system under uncertain environment, reduces the impact of power fluctuations of the system source and load, improves the reliability and low-carbon nature of the system, and dynamically responds to changes in renewable energy and load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118763740B_ABST
    Figure CN118763740B_ABST
Patent Text Reader

Abstract

The application discloses a kind of comprehensive energy system multi-time scale power regulation methods based on deep reinforcement learning, comprising: 1, based on energy coupling relationship, the comprehensive energy system containing electric energy, heat energy, natural gas, hydrogen energy is mathematically modeled;2, according to the difference of equipment adjustment speed in system, propose upper layer long time scale power regulation strategy and lower layer short time scale power regulation strategy, and construct the constraint condition of system;3, the power regulation problem is converted into Markov decision process, and the upper layer intelligent agent and lower layer intelligent agent are trained using deep reinforcement learning method, and the power regulation decision under multi-time scale is made using the strategy network after training.The application does not depend on accurate prediction of renewable energy and load, can dynamically make quick regulation decision to the fluctuation of source and load, meet the demand of comprehensive energy system power supply and demand balance, reduce system carbon emission level, and has important significance to the operation regulation of comprehensive energy system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of power regulation of integrated energy systems, in particular to a multi-time scale power regulation method for integrated energy systems based on deep reinforcement learning. BACKGROUND

[0002] An integrated energy system breaks through the barriers of the traditional independent energy system by integrating various energy forms and equipment, and optimizes the operation of various heterogeneous energy subsystems according to the principle of energy complementarity and energy cascade utilization, which is of great significance to improve the energy supply reliability and low-carbon environmental protection of the system.

[0003] At present, the research on power regulation of integrated energy systems under source-load uncertainty mainly focuses on stochastic optimization and robust optimization methods. The stochastic optimization method uses the probability distribution information of uncertain parameters to find the optimal solution by optimizing the expected value or other statistical characteristics. The robust optimization method finds the optimal solution in the worst case within the uncertainty range to ensure the reliability and robustness of the solution. The above research still has some significant shortcomings: first, the construction of a stochastic optimization model requires quite accurate probability distribution data, which is difficult to obtain in actual application scenarios, and the direct solution of robust optimization is difficult, which requires dual transformation, and the process is complex and tedious. The accuracy of the model is affected by the nonlinear relationship. Second, due to the difference in the management time scale of different devices, there is a lack of regulation in multiple time scales, which affects the source-load balance and low-carbon operation of the system. SUMMARY

[0004] In order to overcome the shortcomings in the prior art, the present application proposes a multi-time scale power regulation method for integrated energy systems based on deep reinforcement learning, in order to reasonably regulate the operation of each device of the integrated energy system, so as to effectively realize the stable and low-carbon operation of the integrated energy system.

[0005] In order to achieve the above purpose, the technical scheme adopted by the present application is:

[0006] The multi-time scale power regulation method for integrated energy systems based on deep reinforcement learning of the present application is characterized in that it is performed according to the following steps:

[0007] Step 1, a mathematical model of an integrated energy system containing electric energy e, thermal energy h, natural gas g and hydrogen energy H2 is constructed;

[0008] Step 1.1, a coupling device model of an integrated energy system containing electric energy e, thermal energy h, natural gas g and hydrogen energy H2 is established;

[0009] Step 1.2, an energy storage device model of an integrated energy system containing electric energy e, thermal energy h, natural gas g and hydrogen energy H2 is established;

[0010] Step two, according to the difference of the equipment adjustment speed in the integrated energy system, a double-layer multi-time scale power regulation strategy is established;

[0011] Step 2.1, for electric energy e, heat energy h, natural gas g, hydrogen energy H2, based on the hourly wind and light output, electric / thermal load and equipment operation state within a day, a long time scale power regulation strategy of the upper layer of the system is established;

[0012] Step 2.2, according to the power fluctuation of photovoltaic and wind turbine, the external power purchase / sale power, battery charge / discharge power and electric boiler operation power of the system are adjusted, and a short time scale power regulation strategy of the lower layer of the system is established;

[0013] Step 2.3, the energy balance constraint and external energy interaction constraint of electric energy e, heat energy h, natural gas g and hydrogen energy H2 are constructed;

[0014] Step three, the upper and lower intelligent agents are trained by using deep reinforcement learning method;

[0015] Step 3.1, the operation parameters of each device in the integrated energy system are collected, the historical power generation data of photovoltaic power station and wind power station are collected, and the historical electric load and thermal load of users are collected;

[0016] Step 3.2, the state space, action space and reward function of the upper intelligent agent and the lower intelligent agent are established:

[0017] Step 3.3, the upper intelligent agent and the lower intelligent agent are trained by using deep reinforcement learning method respectively, the long time scale power regulation model of the upper layer and the short time scale power regulation model of the lower layer are obtained, which are used for outputting the corresponding lower action space from the collected state space, so that the upper action space and the lower action space are used as the double-layer multi-time scale power regulation result.

[0018] The integrated energy system multi-time scale power regulation method based on deep reinforcement learning has the characteristics that the step 1.1 is performed as follows:

[0019] Step 1.1.1, the mathematical model and operation constraint of the electrolytic cell EL are established by using formula (1):

[0020] (1)

[0021] In formula (1): is the electric power consumed by the electrolytic cell EL at t time; is the equivalent power of hydrogen energy H2 output by the electrolytic cell EL at t time; is the energy conversion efficiency of the electrolytic cell EL; , respectively upper and lower limits of the input power of the electrolyzer EL; , respectively upper and lower limits of the ramping power of the electrolyzer EL;

[0022] Step 1.1.2, establishing the mathematical model and operating constraints of the hydrogen fuel cell HFC using equation (2):

[0023] (2)

[0024] In equation (2): is the equivalent power of the hydrogen fuel cell HFC input hydrogen energy H2 at time t; , respectively the electric power and the thermal power output by the hydrogen fuel cell HFC at time t; , respectively the efficiencies of the hydrogen fuel cell HFC converting into electric energy e and thermal energy h; , respectively upper and lower limits of the input power of the hydrogen fuel cell HFC; , respectively upper and lower limits of the ramping power of the hydrogen fuel cell HFC; , respectively upper and lower limits of the electric-thermal ratio of the hydrogen fuel cell HFC;

[0025] Step 1.1.3, establishing the mathematical model and operating constraints of the electric boiler EB using equation (3):

[0026] (3)

[0027] In equation (3): is the electric power consumed by the electric boiler EB at time t; is the thermal power output by the electric boiler EB at time t; is the energy conversion efficiency of the electric boiler EB; , respectively upper and lower limits of the input power of the electric boiler EB; , respectively upper and lower limits of the ramping power of the electric boiler EB;

[0028] Step 1.1.4, establishing the mathematical model and constraints of the gas turbine GT using equation (4):

[0029] (4)

[0030] In equation (4): is the natural gas g consumption power of the gas turbine GT at time t; , are the electrical power and thermal power output by the gas turbine GT at time t respectively; 、 are the efficiency of gas turbine GT in converting into electrical energy e and thermal energy h, respectively; 、 are the upper and lower limits of the input power of the gas turbine GT, respectively; 、 are the upper and lower limits of the climbing power of the gas turbine GT respectively; 、 are the upper and lower limits of the power-to-heat ratio of the gas turbine, respectively;

[0031] Step 1.1.5: Use equation (5) to establish the mathematical model and constraints of the gas boiler GB:

[0032] (5)

[0033] In formula (5): is the natural gas g power consumption of the gas boiler GB at time t; is the thermal power output of the gas boiler GB at time t; is the energy conversion efficiency of the gas boiler GB; 、 are the upper and lower limits of the input power of the gas boiler GB respectively; 、 They are the upper and lower limits of the ramp power of the gas boiler GB respectively.

[0034] The step 1.2 is carried out as follows:

[0035] Step 1.2.1: Use formula (6) to establish the mathematical model and operation constraints of the battery ES:

[0036] (6)

[0037] In formula (6): is the storage capacity of the battery ES at time t; is the maximum storage capacity of the battery ES; is the state of charge of the battery ES at time t; and are the charging and discharging powers of the battery ES at time t-1 respectively; and are the charging and discharging efficiencies of the battery ES respectively; is the self-consumption coefficient of the battery ES; Δt is the unit time slot length; 、 are the charge and discharge states of the battery ES at time t-1, and are 0-1 variables; is the maximum charge and discharge power of the battery ES; , SoCmaxand SoCminare the maximum and minimum values of the state of charge of the battery ES, respectively;

[0038] Step 1.2.2, establishing a mathematical model of the hydrogen storage tank HS and operating constraints using formula (7):

[0039] (7)

[0040] In formula (7): is the hydrogen storage amount of the hydrogen storage tank HS at time t; is the maximum energy storage amount of the hydrogen storage tank HS; is the energy storage state of the hydrogen storage tank HS at time t; and are the charging and discharging equivalent powers of the hydrogen storage tank HS at time t-1; and are the charging and discharging efficiencies of the hydrogen storage tank HS; is the self-loss coefficient of the hydrogen storage tank HS; , are the energy storage states of the hydrogen storage tank HS at time t-1, and are 0-1 variables; is the maximum value of the charging and discharging equivalent powers of the hydrogen storage tank HS; , are the upper and lower limits of the hydrogen storage amount percentage value of the hydrogen storage tank HS.

[0041] The step 2.1 is performed as follows:

[0042] Step 2.1.1, constructing the energy supply imbalance in the upper layer regulation target using formula (8) ;

[0043] (8)

[0044] In formula (8): , are the output powers of the photovoltaic PV and the wind power wind at time t; is the output consumption amount of the wind and light at time t; is the power purchased from the upper-level power grid at time t; is the maximum power supply capacity of the upper-level power grid; is the thermal energy redundancy amount of the integrated energy system at time t, and max(·) is the maximum value function;

[0045] Step 2.1.2, constructing the carbon emissions contained in the upper-level power purchase and gas purchase using formula (9);

[0046] (9)

[0047] In formula (9): 、 are the carbon emissions contained in the electricity and gas purchased by the superior at time t; 、 are the carbon emissions generated by consuming unit electricity e and natural gas g respectively;

[0048] Formula (10) is used to construct the carbon emissions of the integrated energy system at time t in the upper-level control target. :

[0049] (10)

[0050] Step 2.1.3: Use equations (11) and (12) to 、 Normalization:

[0051] (11)

[0052] (12)

[0053] In formulas (11) and (12): and are the normalized energy supply imbalance and carbon emissions at time t, is the adjustment parameter;

[0054] Step 2.1.4: Use Equation (13) to construct the upper-level long-time scale objective function :

[0055] (13)

[0056] In formula (13): 、 They are the upper-level control targets at time t The Euclidean distance from the optimal target (0,0) and the worst target (1,1) is the comprehensive evaluation value of the upper-level long-time scale target at time t, and T is the control period.

[0057] The step 2.2 is carried out as follows:

[0058] Step 2.2.1: Use formula (14) to construct the unbalanced amount of electricity and heat supply in the lower-level control target ;

[0059] Unb t short = [max( P t EB + P t ch − P t dis + P t e , pur +Δ P t e , pur − P max grid , 0 ) + ( P t PV + P t wind − P t con ) + | P t h , EB + P t h ,sup − L t h |] Δ t 2 (14)

[0060] In formula (14): The power supply of the upper power grid at time t during upper-level regulation; is the adjustment amount of the power supply of the upper power grid at time t; is the heat load at time t; where, is the heat supply power, calculated by formula (15);

[0061] (15)

[0062] Step 2.2.2: Use formula (16) to construct the carbon emissions of the power system in the lower-level control target :

[0063] (16)

[0064] In formula (16): is the carbon emissions generated by consuming unit electricity e, The time interval for lower-level regulation;

[0065] Step 2.2.3: Use equations (17) and (18) to 、 Normalization:

[0066] (17)

[0067] (18)

[0068] In formulas (17) and (18): and are the normalized imbalance of electric heat supply and carbon emissions of the power system at time t, is the adjustment parameter;

[0069] Step 2.2.4: Use Equation (19) to construct the objective function of the lower short time scale :

[0070] (19)

[0071] In formula (19): 、 They are the lower-level control targets at time t The Euclidean distance from the optimal target (0,0) and the worst target (1,1) is the comprehensive evaluation value of the upper-level long-time scale target at time t, and T is the control period.

[0072] In step 2.3, the energy balance constraints of electric energy e, thermal energy h, natural gas g, and hydrogen energy H2 are constructed using formula (20):

[0073] (20)

[0074] In formula (20): is the electrical load at time t, is the power sold by the new energy to the external power grid at time t, is the heat power purchased by the integrated energy system from the external heat grid at time t, is the heat load at time t, is the natural gas purchase amount of the integrated energy system at time t, is the hydrogen purchase amount at time t.

[0075] An external energy interaction constraint is constructed using formula (21):

[0076] (21)

[0077] In formula (21): , , are the maximum carrying capacity of the new energy into the external power grid, the maximum heat supply of the external heat grid, and the maximum gas supply of the external gas grid per unit time, respectively.

[0078] The step 3.2 is performed as follows:

[0079] Step 3.2.1, the action space of the upper agent at time t is established using formula (22):

[0080] (22)

[0081] Step 3.2.2, the state space of the upper agent at time t is established using formula (23):

[0082] (23)

[0083] In formula (23): is the action of the upper agent at time t-1;

[0084] Step 3.2.3, the reward function of the upper agent at time t is established using formula (24):

[0085] (24)

[0086] In formula (24): is the first scaling coefficient;

[0087] Step 3.2.4, the action space of the lower agent at time t is established using formula (25): ​​​​

[0088] (25)

[0089] In formula (25), is the charge / discharge power of the battery ES at time t;

[0090] Step 3.2.5, establishing the state space of the lower agent at time t by using formula (26) :

[0091] (26)

[0092] Step 3.2.6, establishing the reward function of the lower agent at time t by using formula (27) :

[0093] (27)

[0094] In formula (27), is the second scaling coefficient.

[0095] The step 3.3 is performed as follows:

[0096] Step 3.3.1, taking the integrated energy system containing electric energy e, thermal energy h, natural gas g and hydrogen energy H2 as the environment of the agent;

[0097] The upper agent is constructed, including: 2 upper Q networks, 2 upper target Q networks and 1 upper policy network; wherein the upper target Q network has the same structure as the upper Q network;

[0098] The lower agent is constructed, including: 2 lower Q networks, 2 lower target Q networks and 1 lower policy network; wherein the lower target Q network has the same structure as the lower Q network;

[0099] Step 3.3.2, initializing the parameters of the 2 upper Q networks of the upper agent as , , initializing the parameters of the 2 upper target Q networks of the upper agent as , , and initializing the parameters of the upper policy network of the upper agent as ;

[0100] Initializing the parameters of the 2 lower Q networks of the lower agent as , , initializing the parameters of the 2 lower target Q networks of the lower agent as , , and initializing the parameters of the lower policy network of the lower agent as ;

[0101] Initialize the length of upper layer and lower layer experience replay pool as 0, set the control period as T, set the total training times as N, t = 0;

[0102] Step 3.3.3, the upper layer strategy network outputs the upper layer mean of the action at time t and the upper layer variance , , and After Gaussian distribution sampling and linear mapping, the action space of the upper layer agent at time t is obtained ;

[0103] The lower layer strategy network outputs the lower layer mean of the action at time t and the lower layer variance , , and After Gaussian distribution sampling and linear mapping, the action space of the lower layer agent at time t is obtained ;

[0104] Step 3.3.4, according to , , the corresponding reward , is obtained, so that the state space of the upper layer agent at time t+1 and the state space of the lower layer agent at time t+1 are obtained;

[0105] Store an experience produced by the upper layer agent in the transition process into the upper layer experience replay pool;

[0106] Store an experience produced by the lower layer agent in the transition process into the lower layer experience replay pool;

[0107] Step 3.3.5, after assigning to , judge whether is true, if true, it means that the environment exploration of the upper and lower layer agents under a control period is completed, otherwise, return to step 3.3.3 for execution;

[0108] Step 3.3.6: If the length of the upper and lower experience replay pools reaches the preset value L, K experiences are randomly sampled from each of the upper and lower experience replay pools as sampling sets, which are used to train the upper and lower agents respectively. The parameters are updated using the stochastic gradient descent algorithm until the maximum number of training times is reached, thereby obtaining the trained upper-layer policy network in the upper-layer agent and the trained lower-layer policy network in the lower-layer agent. Otherwise, t = 0, and the process returns to step 3.3.3 to perform the next control cycle of environmental exploration.

[0109] Step 3.3.7. Use the upper-level policy network trained in the upper-level intelligent agent as the upper-level long-time-scale power control model to output the upper-level action space corresponding to the collected state space, and use the lower-level policy network trained in the lower-level intelligent agent as the lower-level short-time-scale power control model.

[0110] The electronic device of the present invention includes a memory and a processor, and is characterized in that the memory is used to store a program that supports the processor to execute the multi-time scale control method of the integrated energy system, and the processor is configured to execute the program stored in the memory.

[0111] The present invention provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, which is characterized in that when the computer program is run by a processor, the steps of the multi-time-scale control method of the integrated energy system are executed.

[0112] Compared with the prior art, the beneficial effects of the present invention are embodied in:

[0113] 1. This invention aims to solve the problems of energy supply reliability and low-carbon environmental protection in the integrated energy system. It takes the integrated energy system including electricity, heat, gas and hydrogen as the research object, establishes a mathematical model of the system, and uses the approximate ideal solution sorting method to deal with multi-objective problems, thereby optimizing and improving the reliability and low-carbon nature of the integrated energy system operation.

[0114] 2. To address the problem of differences in control speeds among different devices, the present invention proposes a dual-layer, multi-time-scale power control method that can coordinate the operation of energy supply and energy storage devices. The introduction of a dynamic adjustment stage can quickly respond to changes in renewable energy and system load on a small time scale, reducing the impact of power fluctuations in the system source and load.

[0115] 3. Aiming at the difficulty in solving the uncertainty and nonlinear model on both the source and load sides, the present invention uses a deep reinforcement learning algorithm to train the power control strategy model, effectively overcoming the difficulties in model training caused by the large number of reinforcement learning control objects, hyperparameter sensitivity and time scale differences, which is of great significance to the power control of integrated energy systems under uncertain environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0116] Figure 1 This is a flow chart of the multi-time-scale power control method for an integrated energy system based on deep reinforcement learning of the present invention;

[0117] Figure 2 This is a structural diagram of the integrated energy system model of the present invention;

[0118] Figure 3 This is a model training flowchart of the deep reinforcement learning algorithm of the present invention. DETAILED DESCRIPTION

[0119] In this embodiment, a two-layer power control strategy is established considering the adjustment speed characteristics of the equipment in the integrated energy system. On this basis, according to the Markov theory, a multi-time-scale power control method of the integrated energy system based on deep reinforcement learning is proposed. The method first mathematically models the integrated energy system based on the energy coupling relationship; secondly, according to the differences in the adjustment speeds of the system equipment, a two-layer multi-time-scale power control strategy is proposed, and the approximate ideal solution sorting method is used to deal with the multi-objective problem; then, the control decision problem is expressed as a deep reinforcement learning framework. For the control equipment of different time scales, the state space and action space of the upper and lower layers are defined respectively, and the reward functions of the upper and lower layers are designed. Finally, the deep reinforcement learning algorithm is used to train the upper and lower intelligent agents, and the trained policy network is used to make power control decisions under multiple time scales. The proposed method can make fast control decisions on source-load fluctuations dynamically without relying on the accurate prediction of renewable energy and load, so as to meet the energy supply and demand balance of the system and reduce the carbon emission level of the system. It is of great significance to the power control of the integrated energy system. This method is as follows Figure 1 As shown, the specific steps are as follows:

[0120] Step 1: Construct a mathematical model of the integrated energy system containing electric energy e, thermal energy h, natural gas g, and hydrogen energy H2, including a coupling device model and an energy storage device model. The structure of the integrated energy system model is as follows: Figure 2 As shown;

[0121] Step 1.1: Establish a coupled device model of an integrated energy system containing electric energy e, thermal energy h, natural gas g, and hydrogen energy H2;

[0122] Step 1.1.1: Hydrogen production by water electrolysis is powered by the external power grid and internal wind and solar power generation. The mathematical model and operation constraints of the electrolyzer EL are established using formula (1):

[0123] (1)

[0124] In formula (1): is the electric power consumed by the electrolytic cell EL at time t; is the equivalent power of hydrogen energy H2 input to the electrolyzer EL at time t; is the energy conversion efficiency of the electrolyzer EL; , are the upper and lower limits of the input power of the electrolyzer EL, respectively; , are the upper and lower limits of the ramping power of the electrolyzer EL, respectively;

[0125] Step 1.1.2, the hydrogen fuel cell HFC converts hydrogen energy into electrical energy and thermal energy, and a mathematical model and operating constraints of the hydrogen fuel cell HFC are established by using formula (2):

[0126] (2)

[0127] In formula (2): is the equivalent power of hydrogen energy H2 input to the hydrogen fuel cell HFC at time t; , are the electrical power and thermal power output by the hydrogen fuel cell HFC at time t, respectively; , are the efficiencies of the hydrogen fuel cell HFC in converting hydrogen energy into electrical energy e and thermal energy h, respectively; , are the upper and lower limits of the input power of the hydrogen fuel cell HFC, respectively; , are the upper and lower limits of the ramping power of the hydrogen fuel cell HFC, respectively; , are the upper and lower limits of the electrical-thermal ratio of the hydrogen fuel cell HFC, respectively;

[0128] Step 1.1.3, a mathematical model and operating constraints of the electric boiler EB are established by using formula (3):

[0129] (3)

[0130] In formula (3): is the electrical power consumed by the electric boiler EB at time t; is the thermal power output by the electric boiler EB at time t; is the energy conversion efficiency of the electric boiler EB; , are the upper and lower limits of the input power of the electric boiler EB, respectively; , are the upper and lower limits of the ramping power of the electric boiler EB, respectively;

[0131] Step 1.1.4, a mathematical model and constraints of the gas turbine GT are established by using formula (4):

[0132] (4)

[0133] In formula (4): is the natural gas g consumption power of the gas turbine GT at time t; , are the electric power and the thermal power output by the gas turbine GT at time t, respectively; , are the efficiencies of the gas turbine GT in converting into electric energy e and thermal energy h, respectively; , are the upper and lower limits of the input power of the gas turbine GT, respectively; , are the upper and lower limits of the climbing power of the gas turbine GT, respectively; , are the upper and lower limits of the electric-thermal ratio of the gas turbine, respectively;

[0134] Step 1.1.5, a mathematical model and constraint conditions of the gas boiler GB are established by using formula (5):

[0135] (5)

[0136] In formula (5): is the natural gas g consumption power of the gas boiler GB at time t; is the thermal power output by the gas boiler GB at time t; is the energy conversion efficiency of the gas boiler GB; , are the upper and lower limits of the input power of the gas boiler GB, respectively; , are the upper and lower limits of the climbing power of the gas boiler GB, respectively;

[0137] Step 1.2, an energy storage device model of a comprehensive energy system containing electric energy e, thermal energy h, natural gas g, and hydrogen energy H2 is established;

[0138] Step 1.2.1, a mathematical model and operation constraints of the battery ES are established by using formula (6):

[0139] (6)

[0140] In formula (6): is the storage capacity of the battery ES at time t; is the maximum storage capacity of the battery ES; is the state of charge of the battery ES at time t; and are the charging and discharging power of the battery ES at time t-1, respectively; and are the charging and discharging efficiencies of the battery ES, respectively; is the self-loss coefficient of the battery ES; Δt is the unit time slot length; 、 are the charging and discharging states of the battery ES at t-1 time, and are 0-1 variables; is the maximum charging and discharging power of the battery ES; 、 are the maximum and minimum values of the state of charge of the battery ES;

[0141] Step 1.2.2, the hydrogen storage tank HS can store hydrogen purchased from the outside, and a mathematical model and operation constraints of the hydrogen storage tank HS are established by using formula (7):

[0142] (7)

[0143] In formula (7): is the hydrogen storage amount of the hydrogen storage tank HS at t time; is the maximum energy storage capacity of the hydrogen storage tank HS; is the energy storage state of the hydrogen storage tank HS at t time; and are the charging and discharging equivalent powers of the hydrogen storage tank HS at t-1 time; and are the charging and discharging efficiencies of the hydrogen storage tank HS; is the self-loss coefficient of the hydrogen storage tank HS; 、 are the energy storage states of the hydrogen storage tank HS at t-1 time, and are 0-1 variables; is the maximum value of the charging and discharging equivalent powers of the hydrogen storage tank HS; 、 are the upper and lower limits of the hydrogen storage amount percentage of the hydrogen storage tank HS;

[0144] Step two, according to the difference of the adjustment speed of the equipment in the integrated energy system, a double-layer multi-time scale power regulation strategy is established;

[0145] Step 2.1, for electric energy e, thermal energy h, natural gas g, hydrogen energy H2, based on the hourly wind and light output, electric / thermal load and equipment operation state in a day, a long time scale power regulation strategy of the upper layer of the system is established to meet the electric and thermal load demand in a long time scale.

[0146] Step 2.1.1, the energy supply imbalance in the upper layer regulation target is constructed by using formula (8) ;

[0147] (8)

[0148] In formula (8): 、 PV and wind power output at time t, respectively; is the wind-PV power output consumption at time t; is the power purchased from the upper-level power grid at time t; is the maximum power supply capacity of the upper-level power grid; is the thermal energy redundancy of the integrated energy system at time t, and max(·) is the maximum function;

[0149] Step 2.1.2, two main carbon emission sources in the integrated energy system: virtual carbon emissions of upper-level power purchase and CO2 produced by burning natural gas, and the carbon emissions contained in the upper-level power purchase and gas purchase are constructed by using formula (9);

[0150] (9)

[0151] In formula (9), , are the carbon emissions contained in the upper-level power purchase and gas purchase at time t, respectively; , are the carbon emissions produced by consuming unit electric energy e and natural gas g, respectively;

[0152] The carbon emissions of the integrated energy system at time t in the upper-layer control target are constructed by using formula (10) :

[0153] (10)

[0154] Step 2.1.3, since the dimensions of each target value are different, formula (11) and formula (12) are used to normalize , , respectively:

[0155] (11)

[0156] (12)

[0157] In formula (11) and (12), and are the energy supply imbalance and carbon emissions at time t after normalization, respectively, is the adjustment parameter;

[0158] Step 2.1.4, based on the approximation ideal solution ranking method, the upper-layer long-time scale target function is constructed by using formula (13) :

[0159] (13)

[0160] In formula (13), , upper layer control target at time t Euclidean distance to the optimal target (0, 0) and the worst target (1, 1), comprehensive evaluation value of the upper layer long time scale target at time t, T is the control period;

[0161] Step 2.2, for the fast regulating equipment in the electric energy e, adjust the system external power purchase / sell power, battery charge / discharge power, electric boiler operation power according to photovoltaic, fan power fluctuation, and establish the system lower layer short time scale power regulation strategy to carry out dynamic adjustment on smaller time scale;

[0162] Step 2.2.1, use formula (14) to construct the electric heat supply imbalance in the lower layer control target ;

[0163] Unb t short = [max( P t EB + P t ch − P t dis + P t e , pur +Δ P t e , pur − P max grid , 0 ) + ( P t PV + P t wind − P t con ) + | P t h , EB + P t h ,sup − L t h |] Δ t 2 (14)

[0164] In formula (14): is the upper layer power supply power at time t in the upper layer control; is the adjustment amount of the upper layer power supply power at time t; is the heat load at time t; wherein, is the heat supply power, which is calculated by formula (15);

[0165] (15)

[0166] Step 2.2.2, use formula (16) to construct the electric system carbon emission in the lower layer control target :

[0167] (16)

[0168] In formula (16): is the carbon emission generated by consuming unit electric energy e, is the time interval of the lower layer control.

[0169] Step 2.2.3, since the dimensions of each target value are different, use formula (17) and formula (18) to normalize , respectively:

[0170] (17)

[0171] (18)

[0172] In formula (17) and (18): and respectively are the normalized electric heat supply imbalance, the electric system carbon emission at time t, is the adjustment parameter;

[0173] Step 2.2.4, based on the approximation ideal solution ranking method, the objective function of the lower short time scale is constructed by using formula (19) :

[0174] (19)

[0175] In formula (19), , respectively are the lower regulation target at time t and the Euclidean distance of the optimal target (0, 0) and the worst target (1, 1), is the comprehensive evaluation value of the upper long time scale target at time t, and T is the regulation period;

[0176] Step 2.3, the energy balance constraints of electric energy e, heat energy h, natural gas g and hydrogen energy H2 are constructed by using formula (20):

[0177] (20)

[0178] In formula (20), is the electric load at time t, is the power of new energy sold to the external power grid at time t, is the heat power purchased by the comprehensive energy system from the external heat network at time t, is the heat load at time t, is the natural gas purchase amount of the comprehensive energy system at time t, is the hydrogen purchase amount at time t;

[0179] Step 2.4, the external energy interaction constraint is constructed by using formula (21):

[0180] (21)

[0181] In formula (21), , , respectively are the maximum carrying capacity of new energy into the external power grid, the maximum heat supply of the external heat network, and the maximum gas supply of the external gas network per unit time;

[0182] Step three, deep reinforcement learning combines the technologies of reinforcement learning and deep learning, and learns how to take actions to maximize cumulative rewards through interaction with the environment. As shown in Figure 3 , the upper and lower intelligent agents are trained by using the deep reinforcement learning method;

[0183] Step 3.1, collect the operation parameters of each device in the integrated energy system, collect the historical power generation data of the photovoltaic power station and the wind power station, and collect the historical electrical load and thermal load of the user;

[0184] Step 3.2, according to the Markov decision process, the state space, the action space and the reward function of the upper agent and the lower agent are established:

[0185] Step 3.2.1, the action space of the upper agent at time t is established by using formula (22) :

[0186] (22)

[0187] Step 3.2.2, the state space of the upper agent at time t is established by using formula (23) :

[0188] (23)

[0189] In formula (23), is the action of the upper agent at time t-1;

[0190] Step 3.2.3, the reward function of the upper agent at time t is established by using formula (24) :

[0191] (24)

[0192] In formula (24), is the first scaling coefficient, which aims to improve the convergence of the algorithm.

[0193] Step 3.2.4, the action space of the lower agent at time t is established by using formula (25) :

[0194] (25)

[0195] In formula (25), is the charging / discharging power of the battery ES at time t;

[0196] Step 3.2.5, the state space of the lower agent at time t is established by using formula (26) :

[0197] (26)

[0198] Step 3.2.6, the reward function of the lower agent at time t is established by using formula (27) :

[0199] (27)

[0200] In formula (27): is a second scaling coefficient, and the purpose is to improve the convergence of the algorithm.

[0201] Step 3.3, the goal of the deep reinforcement learning algorithm is to update the Critic network by minimizing the error of the value network Critic, and to update the Actor network by maximizing the expected cumulative reward of the policy network Actor, to find the optimal policy So that the reward function with an entropy regularization term is maximized. The upper layer intelligent agent and the lower layer intelligent agent are trained respectively by using the deep reinforcement learning method, to obtain the upper layer power regulation model of long time scale and the lower layer power regulation model of short time scale;

[0202] Step 3.3.1, the integrated energy system containing electrical energy e, thermal energy h, natural gas g and hydrogen energy H2 is taken as the environment of the intelligent agent;

[0203] The upper layer intelligent agent is constructed, including: 2 upper layer Q networks, 2 upper layer target Q networks and 1 upper layer policy network; wherein the structure of the upper layer target Q network is the same as that of the upper layer Q network; the structure of the 2 upper layer Q networks is as follows: 1 layer of input layer containing 16 neurons, respectively containing layer of hidden layer containing neurons, 1 layer of output layer containing 1 neuron; the policy network of the upper layer is as follows: 1 layer of input layer containing 5 neurons, respectively containing layer of hidden layer containing neurons, 1 layer of output layer containing 11 neurons;

[0204] The lower layer intelligent agent is constructed, including: 2 lower layer Q networks, 2 lower layer target Q networks and 1 lower layer policy network; wherein the structure of the lower layer target Q network is the same as that of the lower layer Q network; the structure of the 2 lower layer Q networks is as follows: 1 layer of input layer containing 15 neurons, respectively containing layer of hidden layer containing neurons, 1 layer of output layer containing 1 neuron; the policy network of the lower layer is as follows: 1 layer of input layer containing 2 neurons, respectively containing layer of hidden layer containing neurons, 1 layer of output layer containing 13 neurons;

[0205] Step 3.3.2, the parameters of the 2 upper layer Q networks of the upper layer intelligent agent are initialized as , , the parameters of the 2 lower layer target Q networks of the upper layer intelligent agent are initialized as , , and the parameters of the upper layer policy network of the upper layer intelligent agent are initialized as ;

[0206] Initialize the parameters of the two lower Q networks of the lower agent as 、 , the parameters of the two lower-level target Q networks of the initial lower-level agent are 、 , initialize the parameters of the lower policy network of the lower agent to ;

[0207] Initialize the length of the upper and lower experience replay pools to 0, set the control period to T, set the total number of training times to N, and t=0;

[0208] Step 3.3.3, upper layer strategy network based on , output the upper mean of the action at time t and upper variance , and After Gaussian distribution sampling and linear mapping, the action space of the upper agent at time t is obtained ;

[0209] The lower-level policy network is based on , output the lower layer mean of the action at time t and the lower variance , and After Gaussian distribution sampling and linear mapping, the action space of the lower-level agent at time t is obtained ;

[0210] Step 3.3.4, according to 、 , and receive corresponding rewards 、 , thus obtaining the state space of the upper agent at time t+1 , the state space of the lower-level agent at time t+1 ;

[0211] The experience generated by the upper-level agent during the transfer process Deposit into the upper level experience replay pool;

[0212] The experience generated by the lower-level agent during the transfer process Stored in the lower level experience replay pool;

[0213] Step 3.3.5, Assign to After that, judge Is it true? If so, it means that the upper and lower layer agents have completed the environmental exploration under one control cycle. Otherwise, return to step 3.3.3 to execute;

[0214] Step 3.3.6, if the lengths of the upper and lower experience replay pools reach the preset value L, K experiences are randomly sampled from the upper and lower experience replay pools as sampling sets, respectively, for training the upper and lower agents, and the parameters are updated through the stochastic gradient descent algorithm, in order to reduce the overestimation bias, the two Q networks of the upper and lower agents are respectively taken as the smaller value during the training process, until the maximum training times are reached, so as to obtain the trained upper strategy network in the upper agent and the trained lower strategy network in the lower agent;Otherwise, t=0, return to step 3.3.3 to perform the environment exploration of the next control period;

[0215] Step 3.3.7, the trained upper strategy network in the upper agent is used as the upper long-time-scale power regulation model for outputting the corresponding upper action space from the collected state space, and the trained lower strategy network in the lower agent is used as the lower short-time-scale power regulation model for outputting the corresponding lower action space from the collected state space, so as to obtain the double-layer multi-time-scale power regulation results of the upper action space and the lower action space.

[0216] In this embodiment, an electronic device includes a memory for storing a program supporting a processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0217] In this embodiment, a computer readable storage medium has a computer program stored thereon, and the computer program is executed by a processor to perform the steps of the above method.

[0218] In summary, the present application firstly performs mathematical modeling on the comprehensive energy system based on the energy coupling relationship;Secondly, according to the difference in the adjustment speed of the system equipment, a double-layer multi-time-scale power regulation strategy is proposed, and the technique of approximation ideal solution sorting is used to process the multi-objective problem;Then, the regulation and decision problem is described as a deep reinforcement learning framework, the state space and action space of the upper and lower layers are set for the regulation equipment of different time scales, and the reward functions of the upper and lower layers are constructed;Finally, the deep reinforcement learning method is used to train the agents of the upper and lower layers, and the long-time-scale power regulation model of the upper layer and the short-time-scale power regulation model of the lower layer are obtained, so as to make the power regulation decision under the multi-time-scale.

Claims

1. A multi-timescale power control method for an integrated energy system based on deep reinforcement learning, characterized in that: The steps are as follows: Step 1: Construct a mathematical model of an integrated energy system containing electric energy e, thermal energy h, natural gas g, and hydrogen energy H2; Step 1.1: Establish a coupled device model of an integrated energy system containing electric energy e, thermal energy h, natural gas g, and hydrogen energy H2; Step 1.2: Establish an energy storage device model for a comprehensive energy system containing electric energy e, thermal energy h, natural gas g, and hydrogen energy H2; Step 2: Based on the differences in the regulation speed of equipment in the integrated energy system, a two-layer multi-time-scale power regulation strategy is established; Step 2.1: For electricity e, thermal energy h, natural gas g, and hydrogen energy H2, establish a long-term power control strategy for the upper layer of the system based on the hourly wind and solar power output, electrical / thermal load, and equipment operating status. Step 2.1.1: Use formula (8) to construct the energy supply imbalance in the upper-level control target ; (8) In formula (8): 、 are the output power of photovoltaic PV and wind power wind at time t respectively; is the wind and solar power output consumption at time t; is the power purchased from the upper grid at time t; The maximum power supply capacity of the upper power grid; is the thermal energy redundancy of the integrated energy system at time t, and max(·) is the maximum value function; Step 2.1.2: Use Equation (9) to construct the carbon emissions contained in the upper-level electricity and gas purchases; (9) In formula (9): 、 are the carbon emissions contained in the electricity and gas purchased by the superior at time t; 、 are the carbon emissions generated by consuming unit electricity e and natural gas g respectively; is the natural gas g power consumption of the gas turbine GT at time t; is the natural gas g power consumption of the gas boiler GB at time t; Formula (10) is used to construct the carbon emissions of the integrated energy system at time t in the upper-level control target. : (10) Step 2.1.3: Use equations (11) and (12) to 、 Normalization: (11) (12) In formulas (11) and (12): and are the normalized energy supply imbalance and carbon emissions at time t, To adjust the parameters; Step 2.1.4: Use Equation (13) to construct the upper-level long-time scale objective function : (13) In formula (13): 、 They are the upper-level control targets at time t The Euclidean distance from the optimal target (0,0) and the worst target (1,1) is the comprehensive evaluation value of the upper-level long-time-scale target at time t, and T is the regulation period; Step 2.2: Adjust the system's external power purchase / sale, battery charge / discharge, and electric boiler operating power based on fluctuations in photovoltaic and wind turbine power, and establish a short-term power control strategy for the system's lower layer. Step 2.2.1: Use formula (14) to construct the unbalanced amount of electricity and heat supply in the lower-level control target ; (14) In formula (14): is the power supply of the upper power grid at time t during upper-level regulation, is the adjustment amount of the power supply of the upper power grid at time t; is the heat load at time t; where, is the heat supply power, calculated by formula (15); (15) In formula (15), is the thermal power output of the gas turbine GT at time t, is the thermal power output of the gas boiler GB at time t; is the thermal power output of the hydrogen fuel cell HFC at time t; Step 2.2.2: Use formula (16) to construct the carbon emissions of the power system in the lower-level control target : (16) In formula (16): is the carbon emissions generated by consuming unit electricity e, The time interval for lower-level regulation; Step 2.2.3: Use equations (17) and (18) to 、 Normalization: (17) (18) In formulas (17) and (18): and are the normalized imbalance of electric heat supply and carbon emissions of the power system at time t, To adjust the parameters; Step 2.2.4: Use Equation (19) to construct the objective function of the lower short time scale : (19) In formula (19): 、 They are the lower-level control targets at time t The Euclidean distance from the optimal target (0,0) and the worst target (1,1) is the comprehensive evaluation value of the upper-level long-time-scale target at time t, and T is the regulation period; Step 2.3: Construct energy balance constraints and external energy interaction constraints for electric energy e, thermal energy h, natural gas g, and hydrogen energy H2; Step 3: Use deep reinforcement learning methods to train the upper and lower layer agents; Step 3.1: Collect the operating parameters of each device in the integrated energy system, collect historical power generation data of photovoltaic power plants and wind power plants, and collect historical electricity and heat loads of users; Step 3.2: Establish the state space, action space, and reward function of the upper and lower agents: Step 3.3: Use deep reinforcement learning methods to train the upper-layer intelligent agent and the lower-layer intelligent agent respectively to obtain an upper-layer long-time-scale power control model and a lower-layer short-time-scale power control model, which are used to output the corresponding lower-layer action space for the collected state space, thereby using the upper-layer action space and the lower-layer action space as a double-layer multi-time-scale power control result.

2. The multi-time-scale power control method for an integrated energy system based on deep reinforcement learning according to claim 1 is characterized in that: The step 1.1 is carried out as follows: Step 1.1.

1. Use formula (1) to establish the mathematical model and operation constraints of the electrolytic cell EL: (1) In formula (1): is the electric power consumed by the electrolytic cell EL at time t; is the equivalent power of hydrogen energy H2 output by electrolyzer EL at time t; is the energy conversion efficiency of the electrolyzer EL; 、 are the upper and lower limits of the input power of the electrolyzer EL respectively; 、 They are the upper and lower limits of the climbing power of the electrolyzer EL respectively; Step 1.1.2: Use equation (2) to establish the mathematical model and operation constraints of the hydrogen fuel cell HFC: (2) In formula (2): is the equivalent power of hydrogen energy H2 input to the hydrogen fuel cell HFC at time t; is the electric power output by the hydrogen fuel cell HFC at time t; 、 are the efficiency of hydrogen fuel cell HFC converting into electrical energy e and thermal energy h respectively; 、 They are the upper and lower limits of the hydrogen fuel cell HFC input power respectively; 、 They are the upper and lower limits of the ramp power of the hydrogen fuel cell HFC respectively; 、 They are the upper and lower limits of the electric-to-heat ratio of hydrogen fuel cells HFC; Step 1.1.3: Use formula (3) to establish the mathematical model and operation constraints of the electric boiler EB: (3) In formula (3): is the electric power consumed by the electric boiler EB at time t; is the thermal power output of the electric boiler EB at time t; is the energy conversion efficiency of the electric boiler EB; 、 are the upper and lower limits of the input power of the electric boiler EB respectively; 、 They are the upper and lower limits of the climbing power of the electric boiler EB respectively; Step 1.1.4: Use the mathematical model and constraints of the gas turbine GT using equation (4): (4) In formula (4): is the electric power output by the gas turbine GT at time t; 、 are the efficiency of gas turbine GT in converting into electrical energy e and thermal energy h, respectively; 、 are the upper and lower limits of the input power of the gas turbine GT, respectively; 、 are the upper and lower limits of the climbing power of the gas turbine GT respectively; 、 are the upper and lower limits of the power-to-heat ratio of the gas turbine, respectively; Step 1.1.5: Use equation (5) to establish the mathematical model and constraints of the gas boiler GB: (5) In formula (5): is the energy conversion efficiency of the gas boiler GB; 、 are the upper and lower limits of the input power of the gas boiler GB respectively; 、 They are the upper and lower limits of the ramp power of the gas boiler GB respectively.

3. The multi-time-scale power control method for an integrated energy system based on deep reinforcement learning according to claim 2 is characterized in that: The step 1.2 is carried out as follows: Step 1.2.1: Use formula (6) to establish the mathematical model and operation constraints of the battery ES: (6) In formula (6): is the storage capacity of the battery ES at time t; is the maximum storage capacity of the battery ES; is the state of charge of the battery ES at time t; and are the charging and discharging powers of the battery ES at time t-1 respectively; and are the charging and discharging efficiencies of the battery ES respectively; is the self-consumption coefficient of the battery ES; Δt is the unit time slot length; 、 are the charge and discharge states of the battery ES at time t-1, and are 0-1 variables; is the maximum charge and discharge power of the battery ES; 、 are the maximum and minimum values ​​of the state of charge of the battery ES respectively; Step 1.2.2: Use formula (7) to establish the mathematical model and operation constraints of the hydrogen storage tank HS: (7) In formula (7): is the hydrogen storage capacity of the hydrogen storage tank HS at time t; is the maximum storage capacity of the hydrogen storage tank HS; is the energy storage state of the hydrogen storage tank HS at time t; and are the equivalent powers of gas storage and gas discharge of the gas storage tank HS at time t-1 respectively; and are the charging and discharging energy efficiencies of the gas storage tank HS; is the self-consumption coefficient of the gas storage tank HS; 、 are the energy storage states of the gas storage tank HS at time t-1, and are 0-1 variables; is the maximum value of the equivalent power of gas storage and discharge of the gas storage tank HS; 、 They are the upper and lower limits of the percentage of the gas storage capacity of the gas storage tank HS respectively.

4. The multi-time-scale power control method for an integrated energy system based on deep reinforcement learning according to claim 3 is characterized in that: In step 2.3, the energy balance constraints of electric energy e, thermal energy h, natural gas g, and hydrogen energy H2 are constructed using formula (20): (20) In formula (20): is the electrical load at time t, is the power sold by renewable energy to the external power grid at time t, The thermal power purchased by the integrated energy system from the external heat network at time t, is the heat load at time t, is the natural gas purchase volume of the integrated energy system at time t, is the hydrogen at time t Gas purchase volume; Use formula (21) to construct external energy interaction constraints: (21) In formula (21): 、 、 They are the maximum carrying capacity of new energy integrated into the external power grid per unit time, the maximum heat supply of the external heating network, and the maximum gas supply of the external gas network.

5. The multi-time-scale power control method for an integrated energy system based on deep reinforcement learning according to claim 4 is characterized in that: The step 3.2 is carried out as follows: Step 3.2.1: Use formula (22) to establish the action space of the upper agent at time t : (23) Step 3.2.2: Use formula (23) to establish the state space of the upper agent at time t : (23) In formula (23): is the action of the upper agent at time t-1; Step 3.2.3: Use formula (24) to establish the reward function of the upper agent at time t : (24) In formula (24): is the first scaling factor; Step 3.2.4: Use formula (25) to establish the action space of the lower-level agent at time t : (25) In formula (25), is the charge / discharge power of the battery ES at time t; Step 3.2.5: Use Equation (26) to establish the state space of the lower-level agent at time t : (26) Step 3.2.6: Use formula (27) to establish the reward function of the lower-level agent at time t : (27) In formula (27): is the second scaling factor.

6. The multi-time-scale power control method for an integrated energy system based on deep reinforcement learning according to claim 5 is characterized in that: The step 3.3 is carried out as follows: Step 3.3.

1. Take the integrated energy system containing electric energy e, thermal energy h, natural gas g, and hydrogen energy H2 as the environment of the intelligent agent; Construct an upper-layer agent, including two upper-layer Q networks, two upper-layer target Q networks, and one upper-layer policy network; the upper-layer target Q network has the same structure as the upper-layer Q network; Construct a lower-layer agent, including two lower-layer Q networks, two lower-layer target Q networks, and one lower-layer policy network; the lower-layer target Q network has the same structure as the lower-layer Q network; Step 3.3.2: Initialize the two upper Q network parameters of the upper agent. 、 , the parameters of the two upper target Q networks of the upper agent are initialized as 、 , initialize the parameters of the upper policy network of the upper agent to ; Initialize the parameters of the two lower Q networks of the lower agent as 、 , the parameters of the two lower-level target Q networks of the initial lower-level agent are 、 , initialize the parameters of the lower policy network of the lower agent to ; Initialize the length of the upper and lower experience replay pools to 0, set the control period to T, set the total number of training times to N, and t=0; Step 3.3.3, upper layer strategy network based on , output the upper mean of the action at time t and upper variance , and After Gaussian distribution sampling and linear mapping, the action space of the upper agent at time t is obtained ; The lower-level policy network is based on , output the lower layer mean of the action at time t and the lower variance , and After Gaussian distribution sampling and linear mapping, the action space of the lower-level agent at time t is obtained ; Step 3.3.4, according to 、 , and receive corresponding rewards 、 , thus obtaining the state space of the upper agent at time t+1 , the state space of the lower-level agent at time t+1 ; The experience generated by the upper-level agent during the transfer process Deposit into the upper level experience replay pool; The experience generated by the lower-level agent during the transfer process Stored in the lower level experience replay pool; Step 3.3.5, Assign to After that, judge Is it true? If so, it means that the upper and lower layer agents have completed the environmental exploration under one control cycle. Otherwise, return to step 3.3.3 to execute; Step 3.3.6: If the length of the upper and lower experience replay pools reaches the preset value L, K experiences are randomly sampled from each of the upper and lower experience replay pools as sampling sets, which are used to train the upper and lower agents respectively. The parameters are updated using the stochastic gradient descent algorithm until the maximum number of training times is reached, thereby obtaining the trained upper-layer policy network in the upper-layer agent and the trained lower-layer policy network in the lower-layer agent. Otherwise, t = 0, and the process returns to step 3.3.3 to perform the next control cycle of environmental exploration. Step 3.3.

7. Use the upper-level policy network trained in the upper-level intelligent agent as the upper-level long-time-scale power control model to output the upper-level action space corresponding to the collected state space, and use the lower-level policy network trained in the lower-level intelligent agent as the lower-level short-time-scale power control model.

7. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the multi-time-scale control method for the integrated energy system according to any one of claims 1 to 6, and the processor is configured to execute the program stored in the memory.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multi-time-scale control method for an integrated energy system according to any one of claims 1 to 6 are executed.

Citation Information

Patent Citations

  • Active power distribution network deep reinforcement learning real-time scheduling method and system

    CN115207977A

  • Power distribution network double-layer optimization scheduling method based on deep reinforcement learning

    CN115986845A