Collaborative scheduling method for integrated energy system based on data driving

By constructing an IES model and combining the TD3-AC algorithm, the efficiency and scheduling strategy problems of compressed air energy storage system are solved, the triple supply of hot and hot electricity is realized, the scheduling accuracy and economy of the system are improved, and stability is enhanced, providing an innovative solution for the intelligent management of large-scale energy storage systems.

CN120494387APending Publication Date: 2025-08-15INNER MONGOLIA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510588357.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing compressed air energy storage technology and energy scheduling strategies have problems such as low overall system efficiency, insufficient thermal energy recovery and storage technology, and equipment coupling effects. It is difficult to effectively cope with the intermittent and volatility of renewable energy. The traditional method has a large amount of calculation and is difficult to cope with the needs of large-scale system scheduling. The deep reinforcement learning algorithm has low learning efficiency and insufficient strategy convergence in a sparse reward environment.

Method used

IES model including renewable energy power generation equipment, heat pumps and advanced adiabatic compressed air energy storage system is built, interactive modular modeling method is adopted, combined with TD3-AC algorithm for optimization, self-attention mechanism and behavioral cloning model for imitation learning is introduced, and scheduling strategies are optimized through Markov decision-making process to realize the triple supply of hot and hot electricity and reduce the system operation cost.

Benefits of technology

It improves the scheduling accuracy and economy of the system, enhances the stability of the system, and achieves the dual goal of triple supply of hot and hot electricity, provides an innovative path for the intelligent management of large-scale energy storage systems, and improves the adaptability and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494387A_ABST
    Figure CN120494387A_ABST
Patent Text Reader

Abstract

The invention relates to an integrated energy system collaborative scheduling method based on data driving in the technical field of integrated energy system operation optimization, and the method comprises the steps: converting the cold, heat and electricity demands of a user side into schedulable coupling constraints; an IES model comprising renewable energy power generation equipment, a heat pump (HP), an absorption refrigerator (AC) and an advanced adiabatic compressed air energy storage system (AA-CAES) is constructed, a scheduling problem of the IES model is converted into a Markov decision process, a Bellman equation is introduced, a complex multi-stage decision problem is converted into a solvable optimal substructure problem, and the optimal substructure problem is solved. And the optimal solution of the IES model is solved by adopting a TD3-AC algorithm. According to the invention, the dual goals of combined cooling heating and power supply and system operation cost reduction can be realized, the scheduling precision of the system can be improved, the stability and economy of the system are enhanced, and an innovative technical path is provided for intelligent management of a large-scale energy storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of integrated energy system operation optimization, and relates to a data-driven integrated energy system collaborative scheduling method. Background Art

[0002] Renewable energy sources such as wind and solar are characterized by significant intermittency and volatility, posing significant challenges to the stable operation of power grids. To address the challenges of renewable energy grid integration, large-scale energy storage technology has become a key solution. Compressed air energy storage (CAES) systems, with their low cost, high efficiency, long life, and large capacity, are poised to become a key development direction in the future energy storage market. CAES systems can store energy when there is excess renewable energy generation and release it when demand rises, enabling peak and frequency regulation of the power grid, effectively improving the utilization efficiency of renewable energy and promoting the coordinated deployment of renewable energy and the power grid.

[0003] At present, advanced compressed air energy storage systems include isothermal compressed air energy storage (I-CAES), thermal storage compressed air energy storage (TS-CAES), advanced adiabatic compressed air energy storage (AA-CAES) and supercritical compressed air energy storage (SC-CAES). Taking the AA-CAES system as an example, as an advanced technology, it does not rely on fossil fuels and reduces carbon emissions through heat recovery and storage technology, with both high efficiency and environmental benefits. However, it cannot be ignored that compressed air energy storage technology still faces many challenges, such as low overall system efficiency, insufficient heat recovery and storage technology, and high coupling effects of equipment within the system. There is an urgent need to study more effective system design solutions and energy scheduling strategies to improve its performance and efficiency and promote its practical application in large-scale renewable energy grid-connected systems.

[0004] Energy scheduling optimization strategies currently primarily utilize traditional methods, heuristic algorithms, and deep reinforcement learning (DRL) algorithms. Traditional optimization methods transform energy scheduling problems into constrained mathematical problems through mathematical modeling and employ methods such as linear programming, mixed integer programming, or dynamic programming to solve them. However, these methods suffer from high computational complexity and difficulty coping with uncertain scenarios and large-scale system scheduling requirements. While heuristic algorithms, such as genetic algorithms, particle swarm algorithms, and ant colony algorithms, are highly adaptable when dealing with complex problems, they are prone to falling into local optimality and lack the ability to optimize large-scale data.

[0005] With the development of artificial intelligence (AI) technology, DRL algorithms are increasingly being used in energy scheduling. The Dual-Delayed Deep Deterministic Policy Gradient (TD3) algorithm (hereafter referred to as the TD3 algorithm) stands out for its ability to handle non-convex optimization problems and its effectiveness in both discrete and continuous action spaces. The TD3 algorithm improves upon the Deep Deterministic Policy Gradient (DDPG) algorithm by using a dual Q-network to mitigate Q-value overestimation. This algorithm has demonstrated promising results in energy management strategy formulation, such as improving training efficiency and reducing operating costs. However, the TD3 algorithm still suffers from low learning efficiency and insufficient policy convergence in sparse reward environments.

[0006] In summary, the existing compressed air energy storage technology and energy scheduling strategies have certain defects, and new technical solutions are urgently needed to overcome the above problems and improve the performance of compressed air energy storage systems and energy scheduling efficiency. Summary of the Invention

[0007] The purpose of this invention is to propose a data-driven coordinated scheduling method for an integrated energy system, aiming to achieve the dual goals of trigeneration of cooling, heating and power and reduce system operating costs, while improving the system's scheduling accuracy, enhancing the system's stability and economy, and providing an innovative technical path for the intelligent management of large-scale energy storage systems.

[0008] Based on the above invention objectives, this application proposes a data-driven integrated energy system coordinated scheduling method, which includes the following steps:

[0009] S1. Convert the cooling, heating, and electricity demands of the user side into dispatchable coupling constraints and construct an IES model that includes renewable energy generation equipment, heat pumps (HP), absorption chillers (AC), and advanced adiabatic compressed air energy storage systems (AA-CAES). The AA-CAES is modeled using an interactive modular modeling approach.

[0010] S2. Convert the scheduling problem of the IES model into a Markov decision process and introduce the Bellman equation to convert the complex multi-stage decision problem into a solvable optimal substructure problem to help find the optimal solution;

[0011] S3. Use the TD3-AC algorithm to solve the optimal solution of the IES model. The TD3-AC algorithm is obtained by combining the self-attention mechanism and the behavioral cloning model on the basis of the double-delay deep deterministic gradient algorithm (TD3).

[0012] Furthermore, in step S1, the renewable energy power generation equipment includes a wind power station and a photovoltaic power station. The IES model architecture includes a power grid and multiple energy systems, including a wind turbine (WT) with a variable-pitch system, a photovoltaic (PV) plant with a dual-axis tracker, an advanced adiabatic compressed air energy storage system (AA-CAES), a heat pump (HP), and an absorption chiller (AC). The wind turbine (WT) with a variable-pitch system and the photovoltaic (PV) plant with a dual-axis tracker serve as renewable energy generation equipment, providing power support for the IES model. The wind turbine (WT) uses a wind vane to measure wind speed and direction in real time, enabling the variable-pitch system to adjust the angle of the wind turbine blades in real time. The power grid utilizes a bidirectional power exchange mechanism. The advanced adiabatic compressed air energy storage system (AA-CAES) serves as an energy storage unit, providing balancing support for the power grid and generating heat energy during charging and cooling energy based on ambient temperature during discharge. The absorption chiller (AC) and heat pump (HP) serve as energy conversion devices, providing cooling and heating functions. Through the synergistic effect of energy flows, the various devices in the IES model effectively achieve conversion and scheduling of the three primary energy sources of electricity, heat, and cooling.

[0013] Furthermore, the coupling constraints in step S1 include energy balance constraints, renewable energy constraints, and operational constraints. The goal of the energy balance constraint is to balance the supply and demand of electricity, heating, and cooling energy at all times. The renewable energy constraint includes the output power constraint of a photovoltaic power station (PV) with a dual-axis tracker due to light intensity, and the output power constraint of a wind power station (WT) with a variable pitch system due to wind speed. The operational constraint is that each device in the architecture of the IES model must limit its output power to a range that ensures the safe operation of the entire system.

[0014] Furthermore, the architecture of the advanced adiabatic compressed air energy storage system (AA-CAES) includes a compression system, an air storage tank, an expansion system, a heat exchanger, a high-temperature storage tank and a low-temperature storage tank. The compression system uses two or more compressors to achieve multi-stage compression, and interstage cooling is used between two adjacent stages of compressors to control temperature and improve performance. The expansion system uses two or more expanders to achieve multi-stage expansion, and interstage heating is used between the two-stage expanders to obtain energy recovery during the expansion process.

[0015] Furthermore, in step S1, the interactive modular modeling method decomposes the complex system into multiple modules through the coupling matrix and the distribution matrix, describes the interaction between the modules, and effectively reduces the uncertainty and complexity in the modeling process.

[0016] Furthermore, in step S1, by combining the multi-energy complementary characteristics of the heat pump and the absorption chiller (AC), the advanced adiabatic compressed air energy storage system (AA-CAES) can realize the combined heat, power and cooling (CCHP system for short). Its objective function is expressed as follows:

[0017]

[0018] Where: C total is the total operating cost of the CCHP system, C grid (t) is the cost of purchasing electricity from the grid, C em (t) is the total operating cost of each device in the CCHP system, C em The calculation formula of (t) is as follows:

[0019] C em (t) = C AA-CAES (t)+C HP (t)+C AC (t)

[0020] In the above formula, C AA-CAES (t), C HP (t), C AC (t) represents the operating costs of the advanced adiabatic compressed air energy storage system (AA-CAES), heat pump (HP), and absorption chiller (AC), respectively.

[0021] The main working principles of this application are:

[0022] The present invention proposes a cooperative scheduling framework based on a deep reinforcement learning algorithm, which solves the problem of insufficient dynamic optimization and scheduling efficiency in multi-energy coupling systems. By constructing an IES model that includes renewable energy power generation equipment and AA-CAES, and adopting an interactive modular modeling method to model AA-CAES, the computational complexity of the high-dimensional state space is reduced, laying a scalable foundation for the intelligent optimization of complex systems. On this basis, combined with the multi-energy complementary characteristics of heat pumps and absorption chillers (ACs), a CCHP system architecture is designed to convert the cooling, heating, and electricity demands on the user side into schedulable coupling constraints. This method can not only meet the diverse energy needs of users, but also improve energy utilization efficiency.

[0023] To address the policy sparsity and curse of dimensionality challenges faced by traditional deep reinforcement learning in complex energy scenarios, the TD3-AC algorithm was proposed. By embedding a self-attention mechanism, the algorithm enables the agent to dynamically perceive key features of multi-energy interactions, thereby improving the decision-making accuracy of the policy network. Furthermore, a behavioral cloning model based on imitation learning is introduced to pre-train the policy network, and network parameters are initialized using historical scheduling data, effectively alleviating the inefficient exploration caused by sparse reward signals. Results demonstrate that under dynamic perturbations, the TD3-AC algorithm can rapidly respond to external environmental changes, reduce system operating costs, and improve scheduling accuracy, validating its robustness and adaptability in multi-energy coupled systems.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] (1) A holistic methodology of “interaction modeling-architecture design-algorithm optimization” was constructed, providing an applicable framework for the coordinated scheduling of high-dimensional energy systems;

[0026] (2) The Bellman equation is added to the Markov decision process, which helps to find the optimal solution. Then, through the integration of deep reinforcement learning and attention mechanism, the model's ability to pay attention to important information in complex environments is enhanced, thereby more effectively optimizing strategies.

[0027] (3) The introduction of an improved imitation learning algorithm breaks through the application bottleneck of traditional deep reinforcement learning in reward-sparse scenarios, provides a solution that takes into account both economy and reliability for IES models containing a high proportion of renewable energy, and promotes the development of IES models towards a more efficient, smarter, and more sustainable direction.

[0028] In summary, the present invention can achieve the dual goals of trigeneration of cooling, heating and power and reduce system operating costs. At the same time, it can improve the system's scheduling accuracy, enhance the system's stability and economy, and provide an innovative technical path for the intelligent management of large-scale energy storage systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0030] Figure 1 The overall architecture of the IES model in the embodiment of the present invention is shown;

[0031] Figure 2 The structure of the AA-CAES system in the embodiment of the present invention is shown;

[0032] Figure 3 The Markov decision process in an embodiment of the present invention is shown;

[0033] Figure 4 The TD3-AC algorithm framework according to an embodiment of the present invention is shown;

[0034] Figure 5 The reward iteration curves of the four algorithms in the embodiments of the present invention are shown;

[0035] Figure 6 shows the loss curve of the behavioral cloning model in an embodiment of the present invention;

[0036] Figure 7 The relationship between electricity purchase / sale and electricity price in the power grid according to an embodiment of the present invention is shown;

[0037] Figure 8 The diagram shows the scheduling result of the electric load under the TD3-AC method in an embodiment of the present invention;

[0038] Figure 9 The figure shows the scheduling result of heat load under the TD3-AC method in an embodiment of the present invention;

[0039] Figure 10 The diagram shows the cooling load scheduling result under the TD3-AC method according to an embodiment of the present invention;

[0040] Figure 11 A comparison diagram of the charge and discharge states of four AA-CAES algorithms in an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0041] The following will be combined with the Figures 1 to 11 The technical solutions of the present invention are clearly and completely described. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0042] This embodiment proposes a data-driven integrated energy system coordinated scheduling method, including the following steps:

[0043] S1. Convert the cooling, heating, and electricity demands of the user side into dispatchable coupling constraints and construct an IES model that includes renewable energy generation equipment, heat pumps (HP), absorption chillers (AC), and advanced adiabatic compressed air energy storage systems (AA-CAES). The AA-CAES is modeled using an interactive modular modeling approach.

[0044] S2. Convert the scheduling problem of the IES model into a Markov decision process and introduce the Bellman equation to convert the complex multi-stage decision problem into a solvable optimal substructure problem to help find the optimal solution;

[0045] S3. Use the TD3-AC algorithm to solve the optimal solution of the IES model. The TD3-AC algorithm is obtained by combining the self-attention mechanism and the behavioral cloning model on the basis of the double-delay deep deterministic gradient algorithm (TD3).

[0046] Specifically, in step S1, the renewable energy power generation equipment includes a wind power station and a photovoltaic power station. Figure 1 As shown, the IES model architecture includes a power grid, a wind turbine (WT) with a variable-pitch system, a wind vane, a photovoltaic (PV) plant with a dual-axis tracker, an advanced adiabatic compressed air energy storage system (AA-CAES), a heat pump (HP), and an absorption chiller (AC). The wind turbine (WT) with a variable-pitch system and the photovoltaic (PV) plant with a dual-axis tracker serve as renewable energy generation equipment, providing power support for the IES model. The wind vane is used to measure wind speed and direction in real time, enabling the variable-pitch system to adjust the angle of the wind turbine blades in real time. The power grid adopts a bidirectional power interaction mechanism. The advanced adiabatic compressed air energy storage system (AA-CAES) serves as an energy storage unit, providing balancing support for the power grid and generating heat energy during charging and cooling energy based on the ambient temperature during discharge. The absorption chiller (AC) and heat pump (HP) serve as energy conversion devices, providing cooling and heating functions. Through the synergistic effect of energy flow, the various devices in the IES model effectively achieve the conversion and scheduling of the three main energy sources of electricity, heat, and cooling.

[0047] Based on the architecture of the above IES model, the system model is constructed as follows:

[0048] S11. Build a power grid model

[0049] In the IES model, the energy flow between each energy system and the power grid is bidirectional. Each energy system can purchase electricity from the power grid or sell excess electricity to the power grid. The cost of purchasing electricity can be expressed as: C grid (t), the calculation formula is as follows:

[0050]

[0051] The above formula is the constructed power grid model. b (t) and C s (t) respectively represent the electricity purchase and sales prices at time t, P grid (t) represents the bidirectional energy transmission between each energy system and the power grid at time t. grid (t)≥0 means that each energy system purchases electricity from the grid. On the contrary, when P grid (t)<0 means that each energy system sells excess electricity to the grid to make a profit.

[0052] S12. Construction of WT model

[0053] A piecewise function is used to model the power generation of a wind power station. The real-time power generation of a wind power station is closely related to the real-time wind speed. As long as the wind speed is between the rated speed and the cut-out speed of the wind turbine, the wind turbine output is set to the rated power. Once the wind speed reaches the wind turbine cut-in speed, the wind turbine power generation increases. Whenever the wind speed falls below the cut-in speed or exceeds the cut-out speed, the wind turbine stops operating. Based on the above, the WT model is constructed as follows:

[0054]

[0055] Where v(t) represents wind speed, P WT (t) is a function of wind speed and represents the power generated by the wind turbine (WT) during the time period t. c-in ,v c-out and v rated They are the cut-in speed, cut-out speed and rated speed (m / s) of the wind turbine (WT) (speed at start-up and shutdown). WT-rated (t) refers to the rated power of the wind turbine (kW), η WT is the power generation efficiency of the wind power station (WT), A is the area swept by its wind rotor, C p (λ,β) is the power coefficient, which is related to the tip speed ratio λ and the blade angle β.

[0056] S13. Build PV model

[0057] A PV model is constructed to show the functional relationship between photovoltaic output power, solar radiation and ambient air temperature. The functional relationship is described as follows:

[0058] P PV (t) = η PV ·S PV ·G(t)·(1-γ(T(t)-T ref )),t∈H (3)

[0059] Where, P PV (t),S PV ,η PV , G(t) and T(t) are the output power of photovoltaic and the area of photovoltaic panel (m 2 ), efficiency of PV system, solar radiation intensity per unit area (kW / m 2 ) and ambient air temperature (℃). is the reference temperature (usually 25°C), and γ is the temperature coefficient, which represents the effect of temperature change on the output power.

[0060] S14. Constructing AA-CAES model

[0061] like Figure 2 As shown, in this embodiment, the architecture of the advanced adiabatic compressed air energy storage system (AA-CAES) includes a compression system, an air storage tank, an expansion system, a heat exchanger, a high-temperature storage tank, and a low-temperature storage tank. The compression system uses a low-pressure compressor and a high-pressure compressor to achieve two-stage compression. Between the low-pressure compressor and the high-pressure compressor, interstage cooling is used through a heat exchanger to control temperature and improve performance. The expansion system uses a high-pressure expander and a low-pressure expander to achieve two-stage expansion. Between the high-pressure expander and the low-pressure expander, interstage heating is used through a heat exchanger to obtain energy recovery during the expansion process. Specifically, both the high-pressure expander and the low-pressure expander use turbines.

[0062] Based on the aforementioned advanced adiabatic compressed air energy storage system (AA-CAES) architecture, an interactive modular modeling approach is used to model the AA-CAES. This involves decomposing the complex system into multiple modules using coupling and allocation matrices, describing the interactions between modules and effectively reducing the uncertainty and complexity of the modeling process. This approach is analyzed based on the energy input-output relationship and can be expressed as:

[0063] O=f(I) (4)

[0064] Where O represents the input of various energies, I represents the output of various energies, and the function f(·) represents the conversion of these energies within the advanced adiabatic compressed air energy storage system (AA-CAES). This conversion can be represented by a coupling matrix and a distribution matrix. The coupling matrix C represents the input and output relationship between the components of the advanced adiabatic compressed air energy storage system (AA-CAES), and the distribution matrix D represents the distribution of the input energy among different modules. This can be expressed as:

[0065] O=C·D·I (5)

[0066] The energy conversion efficiency between input and output in an advanced adiabatic compressed air energy storage system (AA-CAES) is quantitatively described by the coupling degree η. In the modeling of an advanced adiabatic compressed air energy storage system (AA-CAES), the relationship between the coupling degree and the coupling matrix and the allocation matrix can be expressed mathematically as follows:

[0067]

[0068] Where m is the amount of output energy; n is the amount of input energy; c is the coupling factor; d is the distribution factor; η is the efficiency factor; η ij represents the efficiency of converting i forms of energy into j forms of energy. Therefore, the matrix equation can be further written as:

[0069]

[0070] Following the above embodiment, the input of the advanced adiabatic compressed air energy storage system (AA-CAES) includes the electric energy I generated by the renewable energy power generation equipment. e and ambient air at normal temperature and pressure a ; Output includes power load P Load , heat load H Load and cooling load R Load . Electric energy I e and ambient air energy I a After these two inputs enter the system, I e By assigning the matrix D e Divided into I e1 and I e2 Two parts, I e1 Directly supply external power loads, while I e2 The ambient air energy I a Enter the low-pressure compressor for compression together, meeting condition d e1 +d e2 =1. The matrix is described as follows:

[0071]

[0072] In the low pressure compressor, the ambient air can a and electrical energy e2 In the efficiency matrix η LC Under the action of L h The other part is converted into high pressure air internal energy, forming a temperature of T LC The gas output from the low-pressure compressor, the process can be expressed as follows:

[0073]

[0074] The high-temperature and high-pressure gas output from the low-pressure compressor flows rapidly into the first heat exchanger. In the heat exchanger 1, the gas is combined with the cold energy I from the low-temperature storage tank. c Perform heat exchange and output heat energy L h1 and low temperature high pressure gas, the gas is heated to T HE1 Indicates that the internal conversion process of HE1 is as follows:

[0075]

[0076] The low-temperature, high-pressure gas T flowing out of the heat exchanger 1 HE1Then it enters the high-pressure compressor for secondary compression. The input end of the high-pressure compressor consists of electrical energy and the gas. In the high-pressure compressor, part of the electrical energy is converted into heat energy L h , the other part is used to enhance the internal energy of high pressure gas, the temperature is T HC .

[0077]

[0078] The high-temperature and high-pressure gas output by the high-pressure compressor immediately enters the heat exchanger 2, where it exchanges heat with the cold energy in the low-temperature storage tank and outputs heat energy L h2 and temperature T HE2 The low-temperature and high-pressure gas then enters the air storage chamber.

[0079]

[0080] In the air storage chamber, assuming that the loss during gas transportation and storage is negligible, the input and output matrix of the air storage chamber is formed:

[0081] I ASC-in =I ASC-out =[T HE2 ]=T ASC (13)

[0082] The heat discharged by heat exchanger 1 and heat exchanger 2 during operation is concentrated in the high temperature storage tank. Figure 2 The cold energy discharged by heat exchanger 3 and heat exchanger 4 shown in the figure is concentrated in the low-temperature storage tank, and the formula is as follows:

[0083]

[0084] When the stored high-pressure gas needs to be converted into electricity for external use, it must be expanded through the high-pressure expander and the low-pressure expander's two-stage turbine. In order to improve the turbine efficiency, interstage heating is used. Before entering the high-pressure turbine (i.e., the high-pressure expander), the temperature is T ASC The gas and the heat energy flowing out of the high temperature storage tank are used as the input of heat exchanger 3 (HE3), and then energy exchange is carried out to generate cold energy and temperature T HE3 high-temperature and high-pressure gas.

[0085]

[0086] After being heated by the heat exchanger 3, the generated gas flows into the high-pressure turbine, and the cold energy is sent to the low-temperature storage tank for storage. In the high-pressure turbine, the high-temperature and high-pressure gas is converted into mechanical energy through the efficiency matrix I m1 , this mechanical energy is then input into a generator and converted into electrical energy.

[0087]

[0088] When the low-temperature, low-pressure gas is output from the high-pressure turbine, it flows into heat exchanger 4 for final heating. Heat exchanger 4 functions similarly to heat exchanger 3, with low-temperature, low-pressure gas and heat energy at its input. Here, the low-temperature, low-pressure gas exchanges heat with heat energy from the high-temperature storage tank through an efficiency matrix, generating cold energy and high-temperature, low-pressure gas. The process equation is as follows:

[0089]

[0090] Finally, the high temperature, low pressure gas T HE4 Flows into the low-pressure turbine (i.e., low-pressure expander). In the low-pressure turbine, the internal energy is converted into mechanical energy and released in the form of gas at normal temperature and pressure, and discharged directly into the environment. Mechanical Energy I m2 and Energy I m1 They enter the generator together and are converted into electrical energy through the efficiency matrix.

[0091]

[0092] The above is the operation process and energy flow of the entire advanced adiabatic compressed air energy storage system (AA-CAES). In this process, the work done by the compressor can be expressed as:

[0093]

[0094] Where c p is the specific heat capacity. The charging power P at time step t c When (t) is known, the air mass flow rate m during the inflation of the advanced adiabatic compressed air energy storage system (AA-CAES) can be calculated c (t) and heating power H AA-CAES (t):

[0095]

[0096] H AA-CAES (t) = m c (t)·c p ·(T LC (t)-T HE1 (t)) (21)

[0097] When deflation occurs, the work done by the expander and the air mass flow rate are:

[0098]

[0099] In rated load operation, the air leaving the last expander (i.e., the low-pressure expander) is cool enough, below ambient temperature, to be used for cooling purposes. Otherwise, the AA-CAES can only provide heat and electricity. Calculate the amount of available cooling / heating power of the AA-CAES:

[0100]

[0101] Since the temperature of the air storage chamber in the advanced adiabatic compressed air energy storage system (AA-CAES) is constant, its pressure and the instantaneous state of charge (SOC) of the air storage at each time step can be obtained as follows:

[0102]

[0103] The above formula (26) is the AA-CAES model constructed in this embodiment.

[0104] In order to meet the cooling and heating load requirements of users, an absorption chiller (AC) is used to convert part of the thermal energy into cold energy, and a heat pump (HP) is used to consume part of the electrical energy and generate heat and cooling energy, which can meet the heating needs of consumers.

[0105] S15. Construct AC model

[0106] R AC (t) = H AC (t)·COP AC (27)

[0107] Where R AC (t) is the cooling energy produced by the absorption chiller (AC); COP AC is the cooling coefficient, which represents the ratio of cooling capacity to electrical power consumption; H AC (t) is the thermal energy fed to the AC.

[0108] S16. Constructing HP model

[0109] R HP (t) = P HP (t)·COP HP (28)

[0110] H HP (t) = P HP (t)·COP HP (29)

[0111] Where R HP (t) and H HP (t) represent the cooling energy and heat load generated by the heat pump (HP), P HP(t) is the power consumed by the heat pump (HP), COP HP is the coefficient of performance of the heat pump (HP).

[0112] S17. Determine the objective function

[0113] Continuing with the above example, the Advanced Adiabatic Compressed Air Energy Storage System (AA-CAES) can implement a combined cooling, heating, and power (CCHP) system. It accumulates energy through compressed air during periods of low electricity prices, then uses turbines and generators to convert this internal energy into electricity during periods of high electricity prices, generating revenue. Furthermore, the system can also provide excess heat and cooling, optimizing profits while meeting user needs.

[0114] Combining the multi-energy complementary characteristics of heat pumps (HP) and absorption chillers (AC), an advanced adiabatic compressed air energy storage system (AA-CAES) can achieve combined heat, power, and cooling (CCHP system). Its objective function is expressed as follows:

[0115]

[0116] Where: C total is the total operating cost of the CCHP system, C grid (t) is the cost of purchasing electricity from the grid, C em (t) is the total operating cost of each device in the CCHP system, C em The calculation formula of (t) is as follows:

[0117]

[0118] In the above formula, C AA-CAES (t), C HP (t), C AC (t) represents the operating costs of the advanced adiabatic compressed air energy storage system (AA-CAES), heat pump (HP), and absorption chiller (AC), respectively.

[0119] Specifically, the coupling constraints in step S1 include energy balance constraints, renewable energy constraints, and operation constraints.

[0120] To maintain the stability and security of the power system, it is crucial to achieve a balance between power supply and demand. The main goal of the IES model scheduling in this embodiment is to balance the supply and demand of electricity, heating, and cooling energy at all times. The calculation formula for the energy balance constraint is as follows:

[0121] P Load (t)+P HP (t) = P grid (t)+P PV (t)+P WT(t)+P AA-CAES (t) (32)

[0122] H Load (t)+H AC (t) = H AA-CAES (t)+H HP (t) (33)

[0123] R Load (t) = R AC (t)+R HP (t)+R AA-CAES (t) (34)

[0124] Where: P Load (t), H Load (t), R Load (t) represent the electricity, heating and cooling load demands at time t, respectively.

[0125] Renewable energy generation equipment is primarily used for wind and photovoltaic power generation. The output power of photovoltaic power generation is affected by sunlight intensity, while the output power of wind power generation is affected by wind speed. Therefore, the renewable energy constraint, i.e., the power output constraint between the IES model and wind and photovoltaic power generation, should satisfy the following formula:

[0126]

[0127] Where: P WT (t) and P PV (t) are the output power of wind turbine and photovoltaic power station at time t; P WT,min (t) and P WT,max (t) are the minimum and maximum output power of the wind turbine respectively; P PV,min (t) and P PV,max (t) are the upper and lower limits of the output power of the photovoltaic power station.

[0128] The operational constraints mean that each device in the CCHP system must limit its output power to a specific effective range to ensure safe operation. Specifically, the heat pump (HP), absorption chiller (AC), and advanced adiabatic compressed air energy storage system (AA-CAES) must all meet the following constraints (min represents the upper limit, max represents the lower limit):

[0129]

[0130] Because the advanced adiabatic compressed air energy storage system (AA-CAES) uses a matrix modeling approach, the power, energy allocation ratio, and conversion efficiency of each device must vary within a reasonable range. Furthermore, the sum of these energy allocation ratios must equal 1 to ensure system energy balance. Therefore, the following constraints must be met:

[0131]

[0132] To ensure the safe operation of the Advanced Adiabatic Compressed Air Energy Storage System (AA-CAES), we must prevent the storage system from overcharging or discharging. Therefore, we limit the charge state of the energy storage system to a specific safety range:

[0133] SOC min ≤SOC(t)≤SOC max (38)

[0134] Where: SOC min and SOC max They are the minimum and maximum charging states of compressed air energy storage, respectively.

[0135] Continuing with the above embodiment, in step S2, the Markov decision process (MDP) is a mathematical framework for simulating decision-making problems, which provides a formal method to describe the behavior of an agent in an environment and its consequences. The Markov decision process consists of a four-tuple {s, a, r, P}. In the four-tuple: s is the state set; a is the action set; r is the reward function; P is the state transition probability. At each time step, the agent selects an action based on the current state and executes it. The system then transitions to a new state based on the transition probability and provides the agent with a corresponding reward. The agent will continue to execute this process until the termination condition or the preset maximum number of steps is reached. By solving the optimal strategy, MDP can help the agent select the best action for each state, thereby maximizing the cumulative reward over a period of time. The energy scheduling problem in this embodiment is described as an MDP, and its framework diagram is shown as follows Figure 3 shown.

[0136] Specifically, the Markov decision process in this embodiment can be vividly understood as an "intelligent decision-making game" for a complex system. Within this complex system, this mechanism acts like a carefully designed management center, determining the system's state based on several key factors: real-time electricity demand price; wind power generation; photovoltaic power generation; combined cooling, heating, and power (CCHP) power, heating, and cooling load requirements; and the charge state of the advanced adiabatic compressed air energy storage system (AA-CAES). Therefore, the state space can be represented as the following seven-dimensional vector:

[0137] s t =[P Load (t),HLoad (t),R Load (t),C b (t),P WT (t),P PV (t),SOC(t)] (39)

[0138] Where C b (t) is the real-time electricity demand price; P WT (t) is the wind power generation power; P PV (t) is the photovoltaic power generation power; P Load (t), H Load (t), R Load (t) are the load demands of the combined cooling, heating and power generation system; SOC(t) is the state of charge of the advanced adiabatic compressed air energy storage system (AA-CAES).

[0139] By observing the current state of the system, the agent generates control actions based on the environmental characteristics and scheduling strategy. These operations cover the energy input and output processes of each dispatchable device in the system. Each action must meet specific operating constraints within the device's operating safety boundary. The system's action space can be represented as a four-dimensional vector: P AA-CAES (t) is the power output of the advanced adiabatic compressed air energy storage system (AA-CAES); is the charge and discharge state of the advanced adiabatic compressed air energy storage system (AA-CAES); P HP (t) is the output power of the heat pump (HP); P AC (t) is the output power of the absorption chiller (AC).

[0140]

[0141] The Bellman equation was proposed by American mathematician Richard Bellman in the 1950s. It is essentially a mathematical equation that recursively expresses the value of optimal decisions. It can transform complex multi-stage decision-making problems into solvable optimal substructure problems. This embodiment introduces the Bellman equation to provide a foundation for solving the optimal strategy and optimal value function for the IES model scheduling problem. The Bellman equations for the state value function V(s) and the action value function Q(s,a) can be expressed as:

[0142]

[0143] Where: V(s) is the value function in state s, which is the expected cumulative reward that can be obtained by following policy π starting from state s; Q(s,a) is the value function after taking action a in state s, which is the expected cumulative reward that can be obtained by following policy π after taking action a; π(a|s) is the probability of taking action a in state s; P(s′|s,a) is the probability of transitioning to state s′ after taking action a in state s; R(s,a) represents the immediate reward obtained by taking action a in state s; γ is a discount factor with a value range of [0,1], which is used to weigh the importance of future rewards.

[0144] The state transition function describes how the state of the environment changes when the agent performs an action within time t. These probabilities are used to calculate the expected reward, which guides the choice of strategy. It can be expressed as follows:

[0145] P(s t+1 =s′|s t =s,a t =a) (41)

[0146] State transitions involve not only the current state variables and action variables, but also the random nature of environmental variables. These variables include the variability of load demand and the unpredictability of equipment operation. Accurate modeling of complex systems is hindered by the inadequacy of standard methods to capture uncertainty and dynamic changes. The matrix modeling approach employed in this embodiment effectively addresses this issue.

[0147] The optimization goal of the energy scheduling strategy is to minimize the system's operating costs, which is closely related to the design of the reward function. However, the agent's goal is to optimize the reward, while the objective function aims to reduce the cost. A negative sign must be added to the objective function to convert it into an MDP reward function. Therefore, the reward function is designed as follows:

[0148] r t =-[C grid (t)+C em (t)] (42)

[0149] Based on the above embodiments, the inventors adopted a deep reinforcement learning solution. In the field of deep reinforcement learning, model-free learning methods are broadly divided into value-based and policy-based methods. Value-based learning typically uses neural networks to approximate the optimal action-value function. Policy-based learning directly trains the policy function through a neural network to generate a probability distribution of actions to maximize reward. The TD3 algorithm is a secure extension of traditional reinforcement learning algorithms, combining the advantages of policy-based and value-based learning. Through the collaborative work of dual-value networks, the algorithm's stability and learning performance are effectively improved. Therefore, based on the enhanced TD3 algorithm, the inventors proposed the TD3-AC algorithm. To enhance the model's ability to focus on key information, an attention mechanism is introduced. This mechanism helps adaptively weight different input features, thereby enhancing the model's ability to process complex state information, improving learning efficiency and decision quality. In addition, to address the challenge of sparse reward signals, a behavior cloning model derived from imitation learning is used to guide the search process. By integrating expert experience, this approach can accelerate the acquisition of effective policies in environments with sparse rewards and reduce the risk of convergence to suboptimal solutions during the learning process. In summary, these innovations contribute to a more powerful and efficient reinforcement learning framework.

[0150] Among them, the TD3 algorithm includes three key advancements:

[0151] (1) The TD3 algorithm uses two Q networks to calculate the target Q value. By selecting the smaller value from the two networks, TD3 can mitigate the impact of extreme values, thereby enhancing the robustness of the learning process.

[0152] (2) The TD3 algorithm uses a delayed update strategy, updating after a predetermined number of time steps, rather than updating the policy network and target network after every time step. This approach reduces the update frequency and makes learning more stable.

[0153] (3) The TD3 algorithm uses a soft update mechanism for the target network to ensure a gradual change in the target network parameters. This smooth transition helps maintain stability during training. These optimizations address the major issue of value estimation bias, reduce the negative impact of frequent policy updates, and reduce the impact of noise on training. This makes the learning process more stable and convergent.

[0154] In deep reinforcement learning, the main job of the critic network is to judge whether the action taken in a specific state is worthwhile, thereby helping the actor network make the best decision. To this end, the output of the target critic network is used to calculate the reward for the current state and action:

[0155] R t =r t +γR t+1 (43)

[0156] Among them, R t is the cumulative reward at time step t, r t is the immediate reward obtained at time step t, and γ is the discount factor, which indicates the influence of future rewards.

[0157] The policy network of the TD3 algorithm takes the current state as input and outputs the probability of taking each action in that state. It uses a deterministic strategy to generate a specific action with a given state, rather than outputting the probability distribution of the action. The network is optimized through the policy gradient method to maximize the expected long-term reward. The value network is a dual Q network, which is used to estimate the expected cumulative reward of taking an action in a specific state, thereby solving the problem of over-estimation. Training these two independent Q networks together helps reduce the variance of the model and improve stability. The TD3 algorithm uses the smaller of the two Q values when updating the policy to obtain more reliable action evaluations. The expressions of the policy network and value network are as follows:

[0158] μ(s t |θ μ )=NN(s t θ μ ) (44)

[0159] Q(s t ,a t |θ Q )=NN(s t ,a t θ Q ) (45)

[0160] In the formula, μ represents the policy network, s t is the state, θ μ is the parameter of the policy network; Q represents the value network, a t is the action, θ Q are the parameters of the value network.

[0161] The update of the Q target network parameters adopts a soft update method, which combines the target network parameters with the main network parameters according to specific weights. This gradual smooth update strategy avoids the direct replacement of parameters and effectively reduces oscillations in the learning process. Since the parameter update is gradual, it can effectively reduce the training instability caused by sudden changes in the main network parameters. The purpose of the policy network update is to maximize the expected cumulative reward and increase the Q value by guiding the parameters through gradient ascent. By setting the update frequency of the critic higher than the update frequency of the actor, the certainty of the critic is enhanced, enabling the actor to take action based on more reliable information. In other words, the critic will perform several updates, and then the actor will perform one update. The update of the value and policy networks is as follows:

[0162]

[0163] Where, π(s t ) represents the strategy in state s, that is, the probability distribution of choosing to take a certain action in state s; Q(s t ,a t ) is the action value function, which is the expected cumulative reward when action a is selected in state s; β is a hyperparameter used to adjust the influence of learning rate or gradient; Q' is the Q value of the target network, which is separated from the parameters of the current network to avoid training instability.

[0164] The self-attention mechanism introduced in the TD3 algorithm enhances the ability of the AI to effectively discern key internal state information during decision-making, resulting in a dual-delayed deep deterministic policy gradient-self-attention mechanism, denoted as the TD3-A algorithm. This algorithm dynamically calculates and modifies element weights based on the similarity between elements in the input sequence, increasing the model's sensitivity to important features. This approach can capture interrelationships and long-term dependencies between elements in the sequence.

[0165] The matrix X∈R(n)×d represents the input sequence, each row represents a feature vector of one element, n is the length of the input sequence, and d is the dimension of each feature. First, the attention layer will create query, key, and value matrices by linearly transforming the input:

[0166] Q=XW Q , K=XW K , V=XW V (48)

[0167] Where: W Q , W K , W V is the learned weight matrix.

[0168] Next comes the attention weighting step, which calculates the similarity between the query and the keyword to obtain an attention score. First, the inner product of each row vector of the matrix is calculated. To reduce the size of the inner product, it is divided by the square root of the dimension representing the keyword. Finally, the function normalizes the inner product result based on the previous inner product. The exact formula is as follows:

[0169]

[0170] In the energy scheduling problem of this embodiment, the TD3 algorithm explores by collecting environmental data using random operations. However, the initial exploration strategy and the parameter initialization of the actor network can greatly affect the learning process, especially in a wide decision space. Therefore, it requires continuous trial and error, which may lead to suboptimal solutions with high operating costs or excessive constraint violations. In addition, sparse reward signals also hinder the agent's ability to effectively receive valuable feedback. To address these challenges, the inventors introduced a behavioral cloning (BC) model from imitation learning, which learns directly from expert samples to pre-train the actor network in the TD3 algorithm.

[0171] First, we collect expert trajectories from known expert policies as a dataset, which comes from the previous optimization model. In order to pre-train the actor network, we compare the actual actions a in the expert trajectories with the actual actions a m,t and the action π predicted by the actor network BC (s m,t ), calculate the loss of imitation learning. In order to minimize this loss so that the actor network can better imitate the actions of the expert, the loss function is defined as:

[0172]

[0173] Where M represents the number of trajectories in the dataset; H m represents the number of state-action pairs in the trajectory.

[0174] After pre-training is complete, these networks are integrated into the TD3 algorithm's network structure. The agent will use the pre-trained character network to initialize its policy. Simultaneously, the TD3 algorithm's training steps continue, including updating the target network and the experience replay mechanism.

[0175] Based on the above ideas, the overall framework of the TD3-AC algorithm constructed in this embodiment is as follows Figure 4 Since this embodiment uses a model-free deep reinforcement learning method, it focuses on learning directly from data, does not rely on an accurate state prediction model, and improves the decision-making strategy through practice and experience.

[0176] The main working principles of this application are:

[0177] The present invention proposes a cooperative scheduling framework based on a deep reinforcement learning algorithm, which solves the problem of insufficient dynamic optimization and scheduling efficiency in multi-energy coupling systems. By constructing an IES model that includes renewable energy power generation equipment and AA-CAES, and adopting an interactive modular modeling method to model AA-CAES, the computational complexity of the high-dimensional state space is reduced, laying a scalable foundation for the intelligent optimization of complex systems. On this basis, combined with the multi-energy complementary characteristics of heat pumps and absorption chillers (ACs), a CCHP system architecture is designed to convert the cooling, heating, and electricity demands on the user side into schedulable coupling constraints. This method can not only meet the diverse energy needs of users, but also improve energy utilization efficiency.

[0178] To address the policy sparsity and curse of dimensionality challenges faced by traditional deep reinforcement learning in complex energy scenarios, the TD3-AC algorithm was proposed. By embedding a self-attention mechanism, the algorithm enables the agent to dynamically perceive key features of multi-energy interactions, thereby improving the decision-making accuracy of the policy network. Furthermore, a behavioral cloning model based on imitation learning is introduced to pre-train the policy network, and network parameters are initialized using historical scheduling data, effectively alleviating the inefficient exploration caused by sparse reward signals. Results demonstrate that under dynamic perturbations, the TD3-AC algorithm can rapidly respond to external environmental changes, reduce system operating costs, and improve scheduling accuracy, validating its robustness and adaptability in multi-energy coupled systems.

[0179] Experiments and analysis

[0180] To verify the effectiveness and superiority of the method proposed in this example, this example will be verified through numerical simulation experiments. In the experiment, the compressed air energy storage device's energy storage capacity was set to 50 megawatt-hours. The dataset used included one year's demand data from a mixed-use community (including residential and commercial buildings), as well as a wind and photovoltaic power generation dataset from a power station. During data processing, 80% of the data was assigned as a test set for algorithm training and hyperparameter optimization, and 20% of the data was designated as a training set. To confirm the effectiveness of model training and testing, the data was checked for outliers. A high-quality dataset generated by normalizing, cleaning, and preprocessing historical dispatch data was used as the expert dataset. To ensure the performance of the TD3-AC algorithm, key hyperparameters were carefully adjusted and experimented with. The optimal hyperparameter configuration was ultimately determined, including a learning rate of 0.0005 for the actor network, a learning rate of 0.001 for the critic network, and a discount factor of 0.99. Specific parameters are shown in Table 1. The neural network adopts a three-layer feedforward structure, as shown in Table 2.

[0181] Table 1 Hyperparameters of the TD3-AC algorithm

[0182] Hyperparameters value Actor network learning rate 0.0005 Critic network learning rate 0.001 Discount Factor 0.99 Batch size 128 Soft update coefficient 0.995 Soft update rate 0.001 Exploring Noise 0.1 λ learning rate 0.0001 Training rounds 1500 Step length per round 100 Experience replay area size 10,000

[0183] Table 2 Parameters of the neural network

[0184] Input layer Hidden layer 1 Hidden layer 2 Output layer Actor Network 7 256 128 4 Critics Network 11 256 128 1

[0185] The TD3-AC algorithm showed significant advantages in multi-dimensional experiments, overcoming the limitations of traditional DRL algorithms in energy scheduling scenarios. This experiment compared the cumulative reward performance of the TD3-AC algorithm, TD3-A algorithm, TD3 algorithm, and DDPG algorithm under the same parameters over 1500 rounds of training. Figure 5 As shown in Figure 2, the TD3-AC algorithm achieved rapid convergence in the first 500 rounds of training, then stabilized and achieved the highest final cumulative reward. This demonstrated its ability to identify the optimal policy after training. This performance improvement is attributed to the combination of the attention mechanism and the behavior cloning model, which significantly improves decision-making efficiency.

[0186] In particular, the behavior cloning model improves learning efficiency by reducing the loss function, effectively guiding the behavior network to more accurately replicate the expert's behavior. Figure 6 The loss function is shown in Figure 2. Initially, in the early stages of learning, the loss curve shows high values, indicating that the model has not yet accurately captured the expert's strategy. However, as training progresses, the loss gradually decreases and stabilizes, reflecting the model's effectiveness in learning to mimic the expert's behavior. This direct imitation allows the model to set clear learning goals, thereby reducing uncertainty during exploration, effectively reducing error, and accelerating the algorithm's convergence.

[0187] The TD3-A algorithm converges slightly slower than the TD3-AC algorithm, but significantly better than the traditional TD3 algorithm and the DDPG algorithm. This demonstrates the positive impact of the attention mechanism on policy optimization. In contrast, the DDPG algorithm shows minimal improvement in cumulative reward over the first 500 rounds and exhibits considerable lag. Regarding reward fluctuations, both the TD3-AC and TD3-A algorithms exhibit less fluctuation in the later stages of training, indicating greater stability in their policy updates. Notably, the TD3-AC algorithm maintains a nearly smooth reward curve after convergence. Considering all performance metrics, the TD3-AC algorithm is the best algorithm for solving the IES model scheduling optimization problem.

[0188] Reducing system operating costs is one of the goals of this application. Table 3 further compares the overall operating costs of systems using different algorithms on the same day. The operating cost of the TD3-AC algorithm is 23,357 yuan, which is approximately 8.3% lower than the traditional DDPG algorithm, 5.81% lower than the TD3 algorithm, and 3.58% lower than the TD3-A algorithm. This shows that the improvements in this application's algorithm can significantly reduce system operating costs.

[0189] The maximum and minimum SOC values under the four algorithms are estimated by formula (6). Table 4 provides detailed information on the power error, scheduling accuracy, and SOC of each algorithm. It is worth noting that the maximum SOC value achieved by the DDPG algorithm is 0.924, which exceeds the operating limit and poses a risk of overcharging. The TD3-AC algorithm controls the power deviation to 1.536MW through precise scheduling, while the deviations of the other three algorithms are 6.114MW, 4.222MW, and 8.365MW, respectively. In terms of scheduling accuracy, the TD3-AC algorithm outperforms the unmodified TD3 algorithm and the traditional DDPG algorithm by 13.51% and 29.13%, respectively. This fully demonstrates the superiority of this application in scheduling optimization. However, although the TD3-A algorithm improves computational efficiency through the attention mechanism, its 4.222MW power deviation causes the scheduling accuracy to drop by 9.43% compared to the basic TD3 algorithm. This shows that the attention mechanism sacrifices scheduling accuracy and stability to a certain extent in order to improve computational efficiency, thus reflecting the balance between efficiency and accuracy in the algorithm optimization process.

[0190] Table 3 Operation cost of different strategies

[0191] Strategy Cost (yuan) DDPG 25441 TD3 24800 TD3-A 24225 TD3-AC 23357

[0192] Table 4 Minimum and maximum SOC, power error and scheduling accuracy of each algorithm

[0193] TD3-AC TD3-A TD3 DDPG Power error (MW) 1.536 6.114 4.222 8.365 Scheduling accuracy (%) 92.32% 69.43% 78.81% 63.19% Minimum SOC 0.203 0.254 0.189 0.236 Maximum SOC 0.877 0.743 0.825 0.924

[0194] Results show that the TD3-AC algorithm, by combining feature extraction with strategic guidance, provides a new technical approach for multi-objective optimization of complex energy systems. It also achieves breakthroughs in three key dimensions: convergence speed, economic efficiency, and scheduling accuracy.

[0195] The scheduling method of this application effectively realizes the coordinated optimization of the multi-energy linkage system of electricity, heat and cooling. Its essence lies in the adaptive response of the system to load demand and electricity price fluctuations, which is to achieve economic benefits by utilizing the changes in energy prices while meeting the load demand. Figure 7 As shown, the system promotes two-way interaction during peak and valley electricity prices by adopting a "low storage, high discharge" strategy: charging during low electricity consumption periods (00:00-6:00), and discharging and selling electricity during the morning and evening peak periods (6:00-8:00 / 18:00-20:00). The AA-CAES system takes advantage of rising electricity prices during peak hours, charging when prices are low and discharging when prices are high, thereby achieving economic benefits and minimizing operating expenses.

[0196] In this experiment, the TD3-AC algorithm was used to optimize the IES model, and the optimal scheduling strategy was determined through comprehensive analysis. Given the lack of cooling demand in winter, scheduling primarily focused on meeting heating loads and optimizing power utilization, while the focus shifted to addressing cooling demand in summer. For analysis, data from a typical winter and summer day were selected as datasets. These diverse operating conditions effectively demonstrated the model's adaptability and robustness across different seasons and complex energy demand scenarios.

[0197] Figures 8 to 10 The results of the dispatch of three types of energy in the community are shown. Figure 8 The power balance results are shown, which include power exchange with the grid system, power consumption by the HP, CAES charging and discharging power, and the output power of the PV and WT. During the day, PV power generation increases significantly, becoming the primary power source, while wind power and the grid continue to provide power as auxiliary sources. At night, wind power and the grid primarily power the system, while PV power generation is zero. Furthermore, the AA-CAES discharges during two different time periods. The first period is the peak hour (8:00-11:00), and the second period is the evening peak hour (19:00-21:00), when electricity prices are higher. Due to the lower electricity prices during off-peak hours, the system prioritizes charging the AA-CAES with low-cost electricity while discharging during peak hours to reduce system operating costs. This scheduling strategy not only meets real-time load demand but also maximizes economic benefits through rational energy allocation, demonstrating the system's excellent energy optimization and scheduling capabilities.

[0198] Figure 9 and Figure 10 Together, they illustrate the coupled scheduling characteristics of heat and cooling energy. In terms of heat balance, the HP and AA-CAES waste heat recovery devices dynamically complement each other. They mainly meet the basic heating demand during peak hours in the morning and evening, while AA-CAES provides supplementary heating through waste heat when the heat pump capacity is insufficient (such as after 8 o'clock). In addition, the difference between the daytime heat load curve and the system heating power also highlights the cross-utilization of thermal energy. The increase in ambient temperature between 11:00 and 16:00 is conducive to the conversion of part of the heat energy into cooling energy through air conditioning, thereby establishing a dynamic balance in the joint supply of heat and cooling energy.

[0199] The cooling energy balance further illustrates this energy conversion feature: air conditioning usage increases significantly throughout the day, especially during hotter hours, when demand for air conditioning peaks. When the heat load decreases at midday, the heat pump supplements cooling energy through the output of the heat-driven cooling pathway. Simultaneously, the AA-CAES provides a smoother power supply and serves as a supplementary cooling source during peak hours in the morning and evening. This dynamic coupling and conversion mechanism between heat and cooling effectively improves energy efficiency.

[0200] Figure 11 The comparison of charge and discharge states in the graph shows the dynamic regulation advantage of the TD3-AC algorithm. Within the 600-minute observation period, the algorithm can achieve power regulation from -4MW to +4MW, which is 2.7 times that of DDPG. It can quickly adjust the power by 90% after a sudden change in load. It is worth noting that during the continuous load fluctuation stage of 100 to 500 minutes, the TD3-AC algorithm can still maintain a stable output of more than 2MW, and has obvious advantages in adapting to load changes. It is the "large amplitude-fast response-low oscillation" characteristics of the TD3-AC algorithm that enables it to effectively manage Figure 10 The midday cooling load peak is shown. The coordinated operation of CAES discharge and direct PV power supply enables the TD3-AC algorithm to regulate voltage fluctuations within a stable range, highlighting its advantage in handling complex dynamic load conditions.

[0201] Furthermore, while the DDPG algorithm exhibited relatively stable output throughout its operation, its ability to adjust to negative power variations appeared insufficient. In particular, its ability to quickly readjust power output during transient load drops was limited. Clearly, the DDPG algorithm was unable to respond flexibly to frequent load fluctuations, and its performance was limited by output hysteresis.

[0202] The above experimental results show that under dynamic disturbance scenarios, the TD3-AC algorithm can quickly respond to changes in the external environment, reduce system operating costs by 8.3%, and improve scheduling accuracy by 29.13%, verifying its robustness and adaptability in multi-energy coupled systems.

[0203] It is obvious to those skilled in the art that the present application is not limited to the details of the above-mentioned exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the present application is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and scope of the equivalent elements of the claims are included in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim can also be implemented by one unit or device through software or hardware. Words such as first and second are used to indicate names and do not indicate any particular order.

Claims

1. A data-driven integrated energy system collaborative scheduling method, characterized in that: The steps include: S1. Convert the cooling, heating, and electricity demands of the user side into dispatchable coupling constraints and construct an IES model that includes renewable energy generation equipment, heat pumps (HP), absorption chillers (AC), and advanced adiabatic compressed air energy storage systems (AA-CAES). The AA-CAES is modeled using an interactive modular modeling approach. S2. Convert the scheduling problem of the IES model into a Markov decision process and introduce the Bellman equation to convert the complex multi-stage decision problem into a solvable optimal substructure problem to help find the optimal solution; S3. Use the TD3-AC algorithm to solve the optimal solution of the IES model. The TD3-AC algorithm is obtained by combining the self-attention mechanism and the behavioral cloning model on the basis of the double-delay deep deterministic gradient algorithm (TD3).

2. The data-driven integrated energy system coordinated scheduling method according to claim 1, characterized in that: In step S1, the renewable energy power generation equipment includes a wind power station and a photovoltaic power station.

3. A data-driven integrated energy system coordinated scheduling method according to claim 2, characterized in that: In step S1, the architecture of the IES model includes a power grid, a wind turbine (WT) with a variable pitch system, a wind vane, a photovoltaic power station (PV) with a dual-axis tracker, an advanced adiabatic compressed air energy storage system (AA-CAES), a heat pump (HP), and an absorption chiller (AC). The wind turbine (WT) with a variable pitch system and the photovoltaic power station (PV) with a dual-axis tracker serve as renewable energy generation equipment and provide power support for the IES model. The wind vane is used to measure wind speed and direction data in real time, enabling the variable pitch system to adjust the angle of the wind turbine blades in real time. The power grid adopts a bidirectional power interaction mechanism. The advanced adiabatic compressed air energy storage system (AA-CAES) serves as an energy storage unit to provide balancing support for the power grid and is used to generate heat energy during charging and cold energy based on the ambient temperature during discharge. The absorption chiller (AC) and heat pump (HP) serve as energy conversion devices and provide cooling and heating functions. Through the synergistic effect of energy flow, the various devices in the IES model effectively realize the conversion and scheduling of the three main energy sources of electricity, heat, and cold.

4. The data-driven integrated energy system coordinated scheduling method according to claim 3 is characterized by: The coupling constraints in step S1 include energy balance constraints, renewable energy constraints, and operational constraints. The goal of the energy balance constraint is to balance the supply and demand of electricity, heating, and cooling energy at all times. The renewable energy constraint includes the output power constraint of photovoltaic power stations (PV) with dual-axis trackers due to light intensity, and the output power constraint of wind power stations (WT) with variable pitch systems due to wind speed. The operational constraint is that each device in the architecture of the IES model must limit its output power to a range that ensures the safe operation of the entire system.

5. The data-driven integrated energy system coordinated scheduling method according to claim 4 is characterized in that: The architecture of the advanced adiabatic compressed air energy storage system (AA-CAES) includes a compression system, an air storage tank, an expansion system, a heat exchanger, a high-temperature storage tank, and a low-temperature storage tank. The compression system uses two or more compressors to achieve multi-stage compression. Between two adjacent compressor stages, interstage cooling is used to control temperature and improve performance. The expansion system uses two or more expanders to achieve multi-stage expansion. Between the two expanders, interstage heating is used to recover energy during the expansion process.

6. The data-driven integrated energy system coordinated scheduling method according to claim 5, characterized in that: In step S1, the interactive modular modeling method decomposes the complex system into multiple modules through the coupling matrix and the allocation matrix, describes the interaction between the modules, and effectively reduces the uncertainty and complexity in the modeling process.

7. A data-driven integrated energy system coordinated scheduling method according to any one of claims 1 to 6, characterized in that: In step S1, the multi-energy complementary characteristics of the heat pump (HP) and the absorption chiller (AC) are combined to realize the combined heat, power, and cooling (CCHP system) through the advanced adiabatic compressed air energy storage system (AA-CAES). The objective function is expressed as follows: Where: C total is the total operating cost of the CCHP system, C grid (t) is the cost of purchasing electricity from the grid, C em (t) is the total operating cost of each device in the CCHP system, C em The calculation formula of (t) is as follows: C em (t)=C AA-CAES (t)+C HP (t)+C AC (t) In the above formula, C AA-CAES (t), C HP (t), C AC (t) represents the operating costs of the advanced adiabatic compressed air energy storage system (AA-CAES), heat pump (HP), and absorption chiller (AC), respectively.