Layered optimization method and system for scheduling strategy of multi-park integrated energy system

By using a hierarchical dual-delay deep deterministic gradient reinforcement learning algorithm, a hierarchical optimization model for a multi-park integrated energy system is constructed. This solves the problems of time scale differences and equipment heterogeneity, realizes multi-park collaborative optimization and efficient energy dispatch, and improves the system's renewable energy consumption and economic efficiency.

CN121563079APending Publication Date: 2026-02-24GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511693347.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Multi-park integrated energy systems face challenges such as differences in time scales, coordination difficulties with heterogeneous equipment, complexity of multi-park collaborative scheduling, uncertainty of renewable energy, and high-dimensional nonlinear coupling problems. Existing scheduling methods are unable to effectively coordinate multiple time scales and equipment heterogeneity, resulting in high computational complexity, low solution efficiency, and inability to achieve efficient system optimization.

Method used

A hierarchical dual-delay deep deterministic gradient reinforcement learning algorithm is adopted to construct a two-stage multi-energy collaborative optimization model with a non-electric energy coordination layer and a power energy real-time control layer. Through a hierarchical Actor-Critic network architecture, long-term optimization of slow-response equipment and short-term real-time control of fast-response equipment are realized. Combined with shared energy storage centers and energy transmission networks, energy complementarity and mutual assistance among multiple parks are coordinated.

Benefits of technology

By effectively coordinating the characteristics of equipment across multiple time scales, the system has improved the capacity for renewable energy absorption and the overall economic efficiency of the system, enhanced the flexibility and robustness of system operation, reduced computational complexity, and achieved efficient optimization of multi-park integrated energy systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121563079A_ABST
    Figure CN121563079A_ABST
Patent Text Reader

Abstract

The invention relates to a hierarchical optimization method and system for a scheduling strategy of a multi-park integrated energy system, and the method comprises the steps: constructing the multi-park integrated energy system which comprises an energy park, a shared energy storage center and an energy transmission network; establishing mathematical models of the energy park and the shared energy storage center, and establishing a long-time scale scheduling optimization model of the non-electric energy coordination layer and a short-time scale scheduling optimization model of the electric energy real-time control layer according to the mathematical models; the optimization model comprises a scheduling optimization objective function and constraints thereof, and the long-time scale scheduling optimization model and the short-time scale scheduling optimization model form a two-stage multi-energy collaborative optimization model; and solving the two-stage multi-energy collaborative optimization model by adopting a layered double-delay depth deterministic gradient reinforcement learning algorithm, and respectively obtaining and executing optimal scheduling actions of the non-electric energy coordination layer and the electric energy real-time control layer in each long time step and each short time step. According to the method, the solving efficiency and the economical efficiency and robustness of system operation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence scheduling technology, and in particular to a hierarchical optimization method and system for scheduling strategies of multi-park integrated energy systems. Background Technology

[0002] With the global energy structure transformation, multi-park integrated energy systems have become an important technological path to improve energy utilization efficiency and promote the large-scale consumption of renewable energy. Traditional single-energy supply models can no longer meet the multi-dimensional demands of modern industrial parks for energy supply security, economy, and environmental protection. Through the deep coupling and synergistic complementarity of multiple energy carriers such as electricity, natural gas, and heat, multi-park integrated energy systems can significantly improve the overall system operating efficiency and energy supply reliability. Optimized scheduling of multi-park integrated energy systems faces numerous technical challenges. The primary issue is handling differences in time scales. Different types of energy equipment exhibit significantly different response characteristics: fast-response equipment such as energy storage devices, power-to-gas conversion devices, and electric boilers need to cope with load fluctuations on a minute-level time scale; while slow-response equipment such as combined heat and power units, gas storage tanks, and heat transmission networks in energy transmission networks are more suitable for medium- to long-term scheduling on a 15-minute to hour-level scale. Existing scheduling methods mostly adopt a uniform time scale, making it difficult to effectively coordinate this heterogeneous characteristic. The complexity of multi-park collaborative scheduling constitutes another major challenge. Complex energy interdependencies and equipment coordination constraints exist among various energy parks. While traditional centralized scheduling methods possess global optimization capabilities, their computational complexity increases exponentially with the number of parks. Distributed scheduling methods, though computationally efficient, are prone to getting trapped in local optima. In particular, the optimal allocation of shared energy storage resources requires simultaneous consideration of multiple factors, including capacity constraints, charging / discharging efficiency, and cycle life, to achieve fair and efficient distribution among parks. The high penetration rate of renewable energy brings significant uncertainty issues. The strong randomness and intermittency of wind and solar power output, along with prediction errors in load demand, make traditional deterministic optimization methods ill-suited to actual operating environments. Existing stochastic and robust optimization methods can handle uncertainty to some extent, but they require extensive scenario generation and probability distribution assumptions, resulting in heavy computational burdens and stringent model accuracy requirements. Furthermore, the strong nonlinearity of multi-energy coupled systems further increases the difficulty of scheduling optimization. Energy conversion equipment such as power-to-gas, combined heat and power, and electric boilers have complex conversion characteristics and operational constraints. The coupling relationships between different energy carriers exhibit high nonlinearity, making it difficult for traditional linear or convex optimization models to accurately characterize system characteristics. Existing mathematical programming-based scheduling algorithms have significant limitations in handling high-dimensional, nonlinear, and strongly coupled problems; their efficiency drops sharply as system scale increases. While reinforcement learning methods have shown potential in energy scheduling in recent years, existing algorithms mostly employ single-layer architectures, failing to effectively handle complex scheduling problems across multiple time scales. In multi-park scenarios, the lack of intelligent algorithms designed specifically for hierarchical scheduling makes it difficult to effectively coordinate long-term planning at the upper level with real-time control at the lower level, thus failing to fully realize the collaborative optimization potential of integrated energy systems. Therefore, there is an urgent need to develop an intelligent scheduling optimization method and system. Summary of the Invention

[0003] Based on this, the purpose of this invention is to address the above-mentioned technical problems by providing a hierarchical optimization method and system for scheduling strategies of multi-park integrated energy systems that can effectively handle differences in time scales, achieve collaborative optimization of multiple parks, and adapt to highly uncertain environments.

[0004] To achieve the aforementioned objectives, the first aspect of this application provides a hierarchical optimization method for scheduling strategies of multi-park integrated energy systems, comprising: Construct a multi-park integrated energy system, which includes at least an energy park, a shared energy storage center, and an energy transmission network; Mathematical models of the energy park and the shared energy storage center are established respectively. Based on the mathematical models, a long-term scheduling optimization model of the non-electric energy coordination layer and a short-term scheduling optimization model of the power energy real-time control layer are established. The optimization model includes a scheduling optimization objective function and its constraints, a long-term scheduling optimization model and a short-term scheduling optimization model, which together form a two-stage multi-energy collaborative optimization model. A hierarchical dual-delay deep deterministic gradient reinforcement learning algorithm is used to solve the two-stage multi-energy collaborative optimization model. The optimal scheduling actions of the non-electric energy coordination layer and the real-time control layer of the electric energy are obtained and executed at each long time step and each short time step, respectively.

[0005] Preferably, there are at least three energy parks, and each energy park is equipped with at least a power-to-gas conversion device (P2G), an electric boiler (EB), a combined heat and power unit (CHP), and a gas storage tank (GS). The mathematical model of the energy park includes mathematical models of the P2G power conversion equipment, the EB electric boiler, the CHP cogeneration unit, and the GS gas storage tank. Electricity-to-gas (P2G) equipment converts surplus electricity into natural gas through water electrolysis and methanation, achieving flexible conversion between electricity and gas energy forms. The mathematical model for P2G equipment is as follows:

[0006] in, and They are respectively Energy Park The amount of natural gas produced and the electrical power consumed by the P2G (Power to Gas) equipment at any given time; For conversion efficiency; The lower heating value of natural gas represents the amount of heat released when a unit volume of natural gas is completely burned. This represents the upper limit of electrical power consumption; As a key device for electro-thermal energy conversion, the electric boiler (EB) plays an important regulatory role in multi-energy complementary systems. This device utilizes resistance heating or electromagnetic induction principles to achieve efficient conversion of electrical energy into heat energy. By starting operation during periods of low electricity prices or when renewable energy output is excessive, the electric boiler (EB) can convert surplus electrical energy into heat energy. The mathematical model of the electric boiler (EB) is as follows:

[0007] in, and They are respectively Energy Park The output thermal power and electrical power consumption of the electric boiler at all times; The efficiency of generating heat energy in electric boilers; This is the upper limit of the output heat power of the electric boiler; The combined heat and power (CHP) unit adopts an integrated configuration of a micro gas turbine and a lithium bromide absorption chiller to achieve cascaded energy utilization. The micro gas turbine drives a generator set to produce electricity by burning natural gas, and the waste heat from the high-temperature flue gas is recovered by the lithium bromide unit and converted into heat energy, achieving efficient output of CHP. The mathematical model of the CHP unit is as follows:

[0008] in, and They are respectively Energy Park The power generation and heat generation of the CHP cogeneration unit at all times; and These are the efficiencies of the combined heat and power (CHP) unit in generating electrical and thermal energy, respectively. This refers to the lower heating value of natural gas. for park The intake air volume of the CHP unit at all times; and These are the upper limits for the power supply and heating capacity of the combined heat and power (CHP) unit, respectively. Gas storage tanks, as buffer and regulating devices in natural gas systems, play a crucial role in addressing supply and demand mismatches in time and space. When natural gas supply is abundant or prices are low, the tanks are filled with gas for energy storage; when the system experiences natural gas shortages or price peaks, the tanks release the stored natural gas to smooth supply and demand fluctuations. This storage and release strategy not only improves the flexibility of system operation but also achieves economical natural gas dispatch. The mathematical model of the gas storage tank (GS) is as follows:

[0009] in, for GS gas storage tank in energy park The amount of gas stored at any given time; The loss rate of the gas storage tank GS due to its own physical factors; and These are the gas storage and venting efficiencies, respectively. and The gas storage tank GS is respectively The amount of gas stored and released at any given time; and It is a binary state variable, which prevents the gas tank GS from being filled and discharged at the same time; and These are the maximum and minimum capacities of the gas storage tank, respectively. The scheduling time interval is the length of time between adjacent scheduling moments.

[0010] Preferably, the mathematical model of the shared energy storage center is as follows:

[0011] in, For shared energy storage center ES in The capacity of a given moment; for Energy Park The electrical power that constantly interacts with the shared energy storage center (ES); The scheduling time interval is the length of time between adjacent scheduling moments. and These are the charge and discharge efficiencies, respectively. and They are respectively Energy Park The charging and discharging power of the shared energy storage center (ES) is monitored at all times. and As binary state variables, they prevent simultaneous charging and discharging within the same energy park. and This represents the maximum charge / discharge interaction power for energy parks and shared energy storage centers (ES). and Maximum and minimum storage capacity limits for shared energy storage centers (ES); As a core regulating facility in a multi-park power system, the shared energy storage center plays a crucial role in the spatial and temporal transfer of power across parks. Facing the random fluctuations in renewable energy output and the asynchronous changes in load across different parks, the energy storage center utilizes advanced battery storage technology to achieve efficient storage and on-demand release of surplus energy. This shared energy storage center (ES) adopts a centralized management and distributed service operation mode: during periods of surplus power, it absorbs and stores excess energy from various parks through bidirectional converters; during periods of power shortage, it rapidly responds to the power demands of each energy park and precisely deploys stored energy. This bidirectional power regulation mechanism not only achieves mutual support between energy parks but also effectively mitigates fluctuations in renewable energy output, improving the overall reliability and economy of the system.

[0012] Preferably, the scheduling optimization objectives of the non-electric energy coordination layer include minimizing the total operating cost of the non-electric system and maximizing the non-electric load satisfaction, and the scheduling optimization objectives of the real-time power energy control layer include minimizing the total power system cost, maximizing the renewable energy absorption rate, and maximizing the power load satisfaction. The scheduling optimization objective function of the non-electric energy coordination layer is:

[0013] in, Total operating cost of non-electric systems; This is the reward coefficient; This represents the satisfaction level of non-electrical loads.

[0014] The total operating cost of non-electric systems includes non-electric costs within the park and maintenance costs related to heat exchange between parks:

[0015] in, Non-electric costs for each industrial park; Maintenance costs for thermal energy exchange between industrial parks.

[0016] Non-electric costs of the park:

[0017] in, For gas purchase costs; For equipment operating costs; for Real-time natural gas prices; for The amount of gas purchased by iEnergy Park; , These are the operating cost coefficients for the combined heat and power unit (CHP) and the gas storage tank (GS), respectively. For combined heat and power (CHP) units; This refers to the lower heating value of natural gas. Let GS be the gas storage capacity of the gas storage tank in i Energy Park at time T. Maintenance costs for heat exchange between industrial parks:

[0018] in, Service fee per unit of heat energy; for Thermal power of interaction between the park and other parks; The non-electrical load satisfaction rate is:

[0019] in, , These are the weighting coefficients for air load and heat load, respectively. , These are the amounts of air load and heat load to be removed, respectively. This refers to the amount of natural gas on the demand side, i.e., the gas load. The heat power, or heat load, is the heat power on the demand side. The scheduling optimization objective function of the real-time power energy control layer is:

[0020] in, The total cost of the power system; , This is the reward coefficient; For renewable energy consumption rate; For electrical load satisfaction; The total cost of the power system is:

[0021] in, For electricity purchase costs; For equipment operating costs; To share the operating costs of energy storage;

[0022] in, Let be the electricity price at time t; Let i represent the electricity purchased by the i Energy Park at time t. , These are the operating cost coefficients for P2G (electric to gas conversion equipment) and EB (electric boiler); Service factor per unit electricity of the gas storage tank GS; The energy storage cost per unit of electricity for the gas storage tank GS; The renewable energy integration rate is:

[0023] in, , They are respectively i Photovoltaic units and wind turbines installed in the energy park t Predictable power generation at any given time; Electrical load satisfaction index:

[0024] in, This represents the amount of electrical load that can be cut off.

[0025] Preferably, the scheduling optimization model satisfies the following constraints: Long-term operational constraints include natural gas balance constraints, thermal energy balance constraints, inter-park thermal energy interaction balance constraints, and thermal energy interaction power upper limit constraints. Short-timescale operating constraints, including real-time power balance constraints of the power system; Natural gas balance constraints:

[0026] in, for The demand-side natural gas volume in the energy park at time T; Thermal energy balance constraints:

[0027] in, for Thermal power on the demand side of the energy park at time T; for The thermal power of the energy park interacting with other energy parks at time T; when When an energy park transfers heat energy to other energy parks, For positive, when When an energy park obtains heat energy from other energy parks, Negative; There is thermal energy interaction between the parks, and the thermal power interaction between the parks should be conserved, that is, the algebraic sum of the interactive thermal power is zero. The thermal energy interaction balance constraints between the parks are as follows:

[0028] in, , and These represent the heat power of each of the three energy parks interacting with the other two energy parks. A positive value indicates that heat energy is transferred to other energy parks, while a negative value indicates that heat energy is obtained from other energy parks. Within a certain time period, the thermal energy exchange between different energy parks should meet the upper limit constraint of thermal energy exchange power:

[0029] in, The maximum limit for the interactive thermal power between different parks; Real-time power balance constraints of power systems:

[0030] in, for Energy Park The demand-side electrical power at any given moment is the electrical load demand. , They are respectively Energy Park The charging and discharging power that constantly interacts with the shared energy storage center (ES).

[0031] Preferably, when solving the two-stage multi-energy cooperative optimization model, it is converted into a two-stage Markov decision process for solution, including: Markov Decision Processes (MDPs) consist of quadruples Define a state space. Describe the system's operating state and action space. Define executable control decisions and state transition probabilities. Representing system uncertainty, reward function To evaluate the effectiveness of action execution, this approach employs a model-free reinforcement learning method, which does not directly learn the state transition probability function. Instead, it acquires experiential data through continuous interaction with the environment and dynamically adapts to the uncertainty of the source load. After the agent selects and executes an action, the next state is determined solely by the environment and fed back to the agent. This mechanism enables the system to learn online and adapt to complex and uncertain environments. Define a hierarchical state space: The system in The complete state at any given moment is controlled by a long time scale. and short-timescale control state Together they constitute: The long-time state space contains the operating states of slow-response devices and non-electrical load demand information in the system. The specific definition of the long-time state space is:

[0032] in, for i The gas load demand of the energy park at time T; for i The amount of natural gas produced by the P2G gas-to-electric conversion equipment in the energy park at time T; The natural gas consumption of the CHP cogeneration unit in the i Energy Park at time T; for i The gas heat load demand of the energy park at time T; for i The heat output of the electric boiler EB at time T in the energy park; for i The heat production capacity of the combined heat and power unit CHP in the energy park at time T; for i The gas storage capacity of gas storage tank GS at time T in the energy park; for i Thermal power interaction between the park and other parks; The short-timescale state space reflects the real-time operating status of fast-response equipment and the immediate needs of the power system. The definition of the short-timescale state space is:

[0033] in, for i Energy Park t The electrical load demand at any given time; Let be the power generation capacity of the photovoltaic units in Energy Park at time t; Let be the power generation capacity of the wind turbines in Energy Park i at time t; The electrical power consumed by P2G devices; The power generation capacity of a combined heat and power (CHP) unit; The electric power of the electric boiler; The interaction power between the park and the shared energy storage center; For the energy storage capacity of the shared energy storage center; Define a hierarchical action space: Action space It consists of the flexibility supply and adjustment range of each flexibility resource. Within its respective control time interval, the action to be performed is determined by the values ​​of its respective decision variables; The long-time action space is responsible for the global optimization scheduling of slow-response devices in the non-electric energy coordination layer. The long-time action space is defined as follows:

[0034] in, For the natural gas consumption regulation of combined heat and power units; This refers to the amount of thermal energy exchange power regulation between industrial parks. This is the adjustment amount for the gas storage capacity of the gas storage tank; The short-timescale action space is responsible for the real-time control of fast-response devices in the real-time control layer of power energy. The short-timescale action space is defined as follows:

[0035] in, For the power regulation of P2G equipment; This refers to the power adjustment of the electric boiler. Adjustment of charging and discharging power for shared energy storage centers; Set up a tiered reward function: reward function The scheduling optimization objective function is composed of the non-electric energy coordination layer and the electric energy real-time control layer based on two-stage control, wherein... Indicates the state Next action The instant reward received; The short-time-scale reward function evaluates the real-time operating performance and economy of a power system. The short-time-scale reward function is as follows:

[0036] in, , , These are the weighting coefficients; The total cost of the power system; Renewable energy integration rate; For electrical load satisfaction; The long-term reward function needs to be based on the long-term steps. Inside The cumulative instantaneous reward obtained within a short time step and a short time scale is adjusted, and the adjusted long-time scale reward function is defined as follows:

[0037] in, , , These are the weighting coefficients; Total operating cost of non-electric systems; For non-electrical load satisfaction; Feedback and correction items for accumulating immediate rewards over a short time scale; Solving for the optimal policy through policy evaluation: By defining the action-value function Evaluation strategy Advantages and disadvantages, strategies This represents the mapping from the state space to the action space;

[0038] in, This is the discount factor, representing the decay value of future rewards; The optimal strategy is equivalent to solving for the optimal Q-value function:

[0039] in, This represents the function with the optimal Q value.

[0040] Preferably, the hierarchical dual-delay deep deterministic policy gradient algorithm extends the Actor-Critic architecture of the TD3 algorithm into a hierarchical structure, forming a two-layer decision-making system that matches the physical characteristics of the system. The upper-layer Actor-Critic network is responsible for global optimization decisions on a long time scale and controls slow-response devices; the lower-layer Actor-Critic network is responsible for real-time control on a short time scale and controls fast-response devices. In the upper-layer network, the Actor network The system receives a long-term state vector as input, which includes time information, electricity price information, non-electric load demand, and the operating status of each slow-response device. After processing through a two-layer fully connected neural network, it outputs a control action vector, which corresponds to the equipment adjustment of the cogeneration units (CHP), gas storage tanks (GS), and heat transmission network (HT) in all energy parks. The Critic network in the upper layer adopts a dual-Q network structure, namely the main Q network. and The two networks have the same structure but their parameters are updated independently. The problem of Q-value overestimation is alleviated by taking the minimum value, and an Actor network is also set up. Main Q network and Corresponding target network , and The target network parameters are slowly updated to track the main network, ensuring the stability of the training process. The lower-layer network structure corresponds to the upper-layer structure, but it is optimized for short-term timescale control. The lower-layer network contains an Actor network. It receives complete system status information and outputs action vectors corresponding to the equipment adjustment quantities of all energy park P2G (electric to gas conversion equipment), EB (electric boiler), and ES (shared energy storage center); the lower-layer Critic network also adopts a dual-Q network structure corresponding to the upper-layer network, with the main Q network... and The target network corresponding to the lower layer network , and However, the target network parameters in the lower layer network are updated more frequently to meet the needs of real-time control. The upper and lower layer networks are coupled by variables. Establish contact, among which, The output adjustment amount for the combined heat and power (CHP) unit. This refers to the power adjustment amount of the P2G (Power-to-Gas) converter. For the power adjustment of the electric boiler EB, the upper-level network transmits energy coordination instructions to the lower-level network through these coupling variables. The lower-level network performs local optimization and real-time response based on the coupling variables, and feeds back the execution results and immediate rewards to the upper-level network, realizing information sharing and strategy coordination optimization between the upper and lower-level networks.

[0041] Preferably, the process of training the upper-layer Actor-Critic network and the lower-layer Actor-Critic network includes: The training process employs a double-buffered experience replay mechanism, with the upper buffer... Storing empirical tuples over long time scales Lower buffer Storing empirical tuples on a short timescale This independent buffer design effectively solves the problem of data imbalance at different time scales; During training, small batches of training samples are first uniformly and randomly sampled from the memory pool. For the lower-layer network, in the state... Actions are generated through the lower-level Actor network. ,in To explore noise, perform actions and receive immediate rewards. and the next state The experience is stored in the lower-level buffer, and during training, the target Q-value is calculated through the target network.

[0042] in, The target policy noise is used to smooth the target policy. The parameters of the lower-level Critic network are updated by minimizing the TD error. The loss function is defined as:

[0043] The lower-level Actor network employs a delayed update strategy, updating the policy gradient only every d=2 steps. The formula for calculating the policy gradient of the Actor network is:

[0044] The training process for the upper-layer network is the same as that for the lower-layer network, but a reward feedback mechanism is added. At the end of each long time step T, the instantaneous rewards of all short time steps within that long time step are collected. This is added as a correction term to the reward function for long-running steps:

[0045] in, These are the weighting coefficients. The modified long-term reward is an abstract reward coupling formula used to describe the reward transmission mechanism between upper and lower layers in the algorithm. This feedback mechanism enables upper-level decision-making to perceive the actual effect of lower-level execution, achieving true hierarchical coordination optimization.

[0046] Preferably, the solution process for solving the two-stage multi-energy cooperative optimization model using the hierarchical dual-delay deep deterministic strategy gradient algorithm includes: Initialize the parameters of the upper and lower layer Actor network, Critic network, and corresponding target network parameters. The target network parameters are directly copied from the main network parameters. At the same time, initialize the experience buffers of the upper and lower layer networks to be empty sets. Before the training iteration begins, a certain amount of initial experience data is collected by randomly interacting with the environment to fill the buffers. After entering the main training loop, at the beginning of each long time step T, based on the current state... And add an upper-level Actor network to explore noise, and select long-term scale actions. Then it executes, and then enters the inner loop to execute continuously. In each short time step, based on the state... Choose actions with lower-level Actors And execute it to receive an instant reward. and new status Short-term experience is stored in the lower buffer, and network updates are immediately sampled from the lower buffer. Finish After a short time step, calculate the cumulative timely reward correction term to obtain the corrected long-timescale reward. Long-term experience is stored in the upper-layer buffer. Batch data is sampled from the upper-layer buffer to update the upper-layer network parameters. Every fixed number of steps, the target network parameters of the upper and lower Actor-Critic networks are updated using a soft update formula.

[0047] in This refers to the soft update rate. The training process continues until the convergence condition is met: the loss function value fluctuates stably within the expected range or reaches the preset maximum number of training rounds; The converged network model is used for online decision-making in a multi-park integrated energy system. At each long time step and short time step, the optimal scheduling actions of the non-electric energy coordination layer and the electric energy real-time control layer are generated according to the real-time system status, so as to realize the two-stage optimized scheduling of the multi-park integrated energy system.

[0048] To achieve the purpose of the invention, the second aspect of this application provides a hierarchical optimization system for the scheduling strategy of a multi-park integrated energy system, applying the hierarchical optimization method for the scheduling strategy of a multi-park integrated energy system described above. The hierarchical optimization system includes: At least three energy parks, each equipped with combined heat and power units, power-to-gas conversion equipment, electric boilers and gas storage tanks, as well as wind turbines and photovoltaic power generation systems; A shared energy storage center is connected to each park via the power network to achieve power sharing between the parks; An energy transmission network connects the various parks, including an electrical, gas, and heat transmission network, enabling the bidirectional flow and sharing of electrical, gas, and heat energy between the energy parks; A hierarchical scheduling controller is configured to solve the two-stage multi-energy collaborative optimization model using the hierarchical dual-delay deep deterministic policy gradient algorithm, perform collaborative optimization scheduling on long and short time scales, and generate and execute the optimal scheduling actions for the non-electric energy coordination layer and the electric energy real-time control layer at each long and short time step, respectively.

[0049] Compared with the prior art, the beneficial effects of this invention are: This invention proposes a dual-timescale hierarchical decoupling scheme with both long and short timescales in the scheduling architecture. This effectively coordinates the differentiated characteristics of slow-response and fast-response equipment, solves the problem of equipment heterogeneity coordination, and constructs a system architecture that coordinates multiple energy parks, shared energy storage centers, and energy transmission networks. This enables cross-park complementarity and mutual support of electricity, gas, and heat energy, thereby significantly improving the renewable energy absorption capacity and the overall economic efficiency of the system. The deep reinforcement learning algorithm proposed in this invention, namely the hierarchical dual-delay deep deterministic policy gradient algorithm, achieves implicit modeling and adaptive optimization of high-dimensional, uncertain, and strongly coupled systems through its powerful nonlinear fitting and online learning capabilities. This fundamentally overcomes the bottlenecks of computational complexity and poor adaptability of traditional mathematical programming methods, significantly improving solution efficiency and the economic efficiency and robustness of system operation. Attached Figure Description

[0050] Figure 1 This is a flowchart illustrating a hierarchical optimization method for a multi-park integrated energy system scheduling strategy in one embodiment. Figure 2 This is a schematic diagram of the solution process of the layered double-delay deep deterministic policy gradient algorithm in one embodiment; Figure 3 This is a schematic diagram of the overall structure of a multi-park integrated energy system in one embodiment. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention. The following embodiments are used to illustrate the invention but are not intended to limit its scope.

[0052] Example 1 Embodiment 1 of this application provides a hierarchical optimization method for the scheduling strategy of a multi-park integrated energy system, such as... Figure 1 As shown, it includes: S1: Construct a multi-park integrated energy system, which includes at least an energy park, a shared energy storage center, and an energy transmission network; S2: Establish mathematical models for the energy park and the shared energy storage center respectively, and establish a long-term scheduling optimization model for the non-electric energy coordination layer and a short-term scheduling optimization model for the real-time control layer of the power energy based on the mathematical models. S3: The optimization model includes a scheduling optimization objective function and its constraints, a long-term scheduling optimization model and a short-term scheduling optimization model, which together form a two-stage multi-energy collaborative optimization model. S4: The hierarchical double-delay deep deterministic gradient reinforcement learning algorithm is used to solve the two-stage multi-energy collaborative optimization model, obtain the optimal scheduling action, and execute it.

[0053] To address key technical challenges in the scheduling and optimization of multi-park integrated energy systems, such as hierarchical coordination and multi-park collaboration, Embodiment 1 of this invention provides a hierarchical optimization method for scheduling strategies of multi-park integrated energy systems based on a hierarchical deep reinforcement learning algorithm. The core innovation of this method lies in constructing a hierarchical time-scale decoupled architecture and combining it with the Hierarchical Twin Delayed Deep Deterministic Policy Gradient (H-TD3) algorithm to achieve intelligent optimization. At the system architecture level, power sharing among multiple parks is achieved through a shared energy storage center, and inter-park thermal energy collaborative configuration is achieved through a thermal energy transmission network within the energy transmission network. At the scheduling strategy level, a hierarchical collaborative optimization mechanism covering diverse equipment such as power-to-gas conversion equipment, electric boilers, cogeneration units, and gas storage tanks is established. The upper-level control is responsible for the global planning of long-term scale equipment such as cogeneration units, gas storage tanks, and thermal energy transmission networks; in this embodiment, a 15-minute scheduling cycle is used for medium- to long-term optimization decisions. The lower-level control is responsible for the real-time response of short-term scale equipment such as power-to-gas conversion equipment, electric boilers, and shared energy storage; in this embodiment, a 5-minute scheduling cycle is used for rapid adjustment. By constructing a multi-energy flow coupling model and energy balance constraints between industrial parks, deep synergistic optimization of the three energy carriers—electricity, natural gas, and heat—is achieved.

[0054] The algorithm employs an innovative two-layer Actor-Critic network architecture. The upper-layer network focuses on handling long-term global planning tasks, with its state space encompassing key information such as the operating status of slow-response equipment, non-electrical load demand, and system operating constraints. Its action space corresponds to scheduling decisions for equipment such as combined heat and power (CHP), gas storage, and heat transfer. The lower-layer network is responsible for real-time control on short timescales, processing the state space containing complete system information to achieve precise control of fast-response equipment. Through a carefully designed hierarchical reward function and the TD3 algorithm's unique delayed update mechanism, the two layers are ensured to coordinate and converge stably, ultimately forming an intelligent scheduling strategy adapted to multiple timescales, achieving efficient and optimized operation of multi-park integrated energy systems.

[0055] like Figure 2 As shown, Figure 2 A schematic diagram of the solution process for the hierarchical double-delay deep deterministic policy gradient algorithm (hierarchical TD3 algorithm) is shown, including: The complete execution flow of the hierarchical TD3 algorithm is as follows: At the beginning of the algorithm, the parameters of the upper and lower layer Actor and Critic networks, as well as the corresponding target network parameters, are initialized. The target network parameters are directly copied from the main network parameters. Simultaneously, the upper and lower layer experience buffers are initialized to empty sets. Before the training iteration begins, a certain amount of initial experience data is collected through random interaction with the environment to fill the buffers.

[0056] After entering the main training loop, at the beginning of each long time step T, the algorithm adjusts the current state accordingly. Select long-term action scales with the upper-level Actor network (with added exploration noise). Then it executes. It then enters the inner loop and executes continuously. Short time steps: In each short time step, based on the state And the lower-level Actor network selects actions And execute it to receive an instant reward. and new status Short-term experience is stored in the lower-level buffer, and network updates are immediately performed by sampling from the buffer.

[0057] Finish After a short time step, the algorithm calculates the cumulative reward correction term to obtain the corrected long-timescale reward. Long-term experience is stored in the upper-layer buffer. Batch data is sampled from the upper-layer buffer to update the upper-layer network parameters. Every fixed number of steps, the target network parameters of both layers are updated using a soft update formula: .in This is a soft update rate, which ensures the stability of the training process.

[0058] The training process continues until the convergence condition is met: the loss function value fluctuates stably within a small range, or the preset maximum number of training rounds is reached. The converged network model can then be used for online decision-making, outputting the optimal control action based on the real-time system state to achieve two-stage sequential optimization scheduling of the multi-park integrated energy system.

[0059] Through the above-mentioned hierarchical deep reinforcement learning algorithm design, this invention successfully solves technical problems such as multi-timescale coupling, high-dimensional state-action space, and source-load uncertainty, providing an efficient and reliable solution for intelligent optimization scheduling of multi-park integrated energy systems, and has good practical value and promotion prospects.

[0060] Example 2 Embodiment 2 of this application, based on Embodiment 1, provides a hierarchical optimization system for scheduling strategies of multi-park integrated energy systems. The hierarchical optimization system includes: At least three energy parks, each equipped with combined heat and power units, power-to-gas conversion equipment, electric boilers and gas storage tanks, as well as wind turbines and photovoltaic power generation systems; A shared energy storage center is connected to each park via the power network to achieve power sharing between the parks; An energy transmission network connects the various parks, including an electrical, gas, and heat transmission network, enabling the bidirectional flow and sharing of electrical, gas, and heat energy between the energy parks; A hierarchical scheduling controller is configured to solve the two-stage multi-energy collaborative optimization model using the hierarchical dual-delay deep deterministic policy gradient algorithm, perform collaborative optimization scheduling on long and short time scales, and generate and execute the optimal scheduling actions for the non-electric energy coordination layer and the electric energy real-time control layer at each long and short time step, respectively.

[0061] Example 3 This embodiment 3, based on embodiments 1 and 2, provides a specific embodiment of a multi-park integrated energy system, such as... Figure 3 As shown, the system achieves deep coupling and optimized configuration of three energy carriers—electricity, natural gas, and heat—by constructing a multi-park collaborative architecture.

[0062] The system mainly consists of three independent energy parks, a shared energy storage center, and a heat transmission network connecting the parks. The three parks have different energy configuration characteristics: Park 1, an industrial park, serves as the main renewable energy base, equipped with wind turbines and photovoltaic power generation systems; Park 2, a commercial park, and Park 3, a residential park, primarily rely on photovoltaic power generation for clean electricity, with fewer wind turbines. All parks are equipped with complete multi-energy conversion equipment, including combined heat and power (CHP) units, power-to-gas (EPC) devices, electric boilers, and gas storage tanks, forming a fully functional integrated energy supply unit. The shared energy storage center serves as the power regulation hub of the entire system, providing unified power storage and distribution services to the three parks. Through bidirectional power connections with each park, the shared energy storage center can achieve mutual power surplus and deficit among the parks, effectively balancing load fluctuations and renewable energy output changes. The energy transmission network includes an electricity transmission network (upper-level distribution network), a gas transmission network (natural gas network), and a heat transmission network (heat network). The transmission network is connected between parks via pipelines or lines, enabling flexible allocation and efficient utilization of electricity, gas, and heat loads across different parks.

[0063] The entire system can be divided into five functional modules according to the energy flow direction: the energy supply side is responsible for providing the system with primary energy and renewable energy; the energy conversion equipment realizes the mutual conversion between different energy carriers; the energy storage equipment provides the time shift and regulation functions of energy; the energy demand side corresponds to the actual energy demand of each park; and the energy sharing side realizes energy coordination between parks through the shared energy storage center and the heat energy transmission network.

[0064] This hierarchical and zoned architecture design makes full use of the complementary nature of resources and the differences in load between parks. Through the coordinated optimization of multiple energy flows and spatiotemporal complementarity, it achieves efficient consumption of renewable energy and a significant improvement in the overall economic efficiency of the system.

Claims

1. A hierarchical optimization method for scheduling strategies of multi-park integrated energy systems, characterized in that, Includes the following steps: Construct a multi-park integrated energy system, which includes at least an energy park, a shared energy storage center, and an energy transmission network; Mathematical models of the energy park and the shared energy storage center are established respectively. Based on the mathematical models, a long-term scheduling optimization model of the non-electric energy coordination layer and a short-term scheduling optimization model of the power energy real-time control layer are established. The optimization model includes a scheduling optimization objective function and its constraints, a long-term scheduling optimization model and a short-term scheduling optimization model, which together form a two-stage multi-energy collaborative optimization model. A hierarchical dual-delay deep deterministic gradient reinforcement learning algorithm is used to solve the two-stage multi-energy collaborative optimization model. The optimal scheduling actions of the non-electric energy coordination layer and the real-time control layer of the electric energy are obtained and executed at each long time step and each short time step, respectively.

2. The method according to claim 1, characterized in that, There are at least three energy parks, and each energy park is equipped with at least a power-to-gas (P2G) generator, an electric boiler (EB), a combined heat and power (CHP) unit, and a gas storage tank (GS). The mathematical model of the energy park includes mathematical models of the P2G power conversion equipment, the EB electric boiler, the CHP cogeneration unit, and the GS gas storage tank. Mathematical model of P2G for electro-gas conversion equipment: in, and They are respectively Energy Park The amount of natural gas produced and the electrical power consumed by the P2G (Power to Gas) equipment at any given time; For conversion efficiency; The lower heating value of natural gas represents the amount of heat released when a unit volume of natural gas is completely burned. This represents the upper limit of electrical power consumption; Mathematical model of electric boiler EB: in, and They are respectively Energy Park The output thermal power and electrical power consumption of the electric boiler at all times; The efficiency of generating heat energy in electric boilers; This is the upper limit of the output heat power of the electric boiler; Mathematical model of CHP (Combined Heat and Power) unit: in, and They are respectively Energy Park The power generation and heat generation of the CHP cogeneration unit at all times; and These are the efficiencies of the combined heat and power (CHP) unit in generating electrical and thermal energy, respectively. This refers to the lower heating value of natural gas. for park The intake air volume of the CHP unit at all times; and These are the upper limits for the power supply and heating capacity of the combined heat and power (CHP) unit, respectively. Mathematical model of gas storage tank GS: in, for GS gas storage tank in energy park The amount of gas stored at any given time; The loss rate of the gas storage tank GS due to its own physical factors; and These are the gas storage and venting efficiencies, respectively. and The gas storage tank GS is respectively The amount of gas stored and released at any given time; and It is a binary state variable, which prevents the gas tank GS from being filled and discharged at the same time; and These are the maximum and minimum capacities of the gas storage tank, respectively. The scheduling time interval is the length of time between adjacent scheduling moments.

3. The method according to claim 1, characterized in that, The mathematical model for the shared energy storage center is as follows: in, For shared energy storage center ES in The capacity of a given moment; for Energy Park The electrical power that constantly interacts with the shared energy storage center (ES); The scheduling time interval is the length of time between adjacent scheduling moments. and These are the charge and discharge efficiencies, respectively. and They are respectively Energy Park The charging and discharging power of the shared energy storage center (ES) is monitored at all times. and As binary state variables, they prevent simultaneous charging and discharging within the same energy park. and This represents the maximum charge / discharge interaction power for energy parks and shared energy storage centers (ES). and Maximum and minimum storage capacity limits for shared energy storage centers (ES).

4. The method according to claim 2, characterized in that, The scheduling optimization objectives of the non-electric energy coordination layer include minimizing the total operating cost of the non-electric system and maximizing the non-electric load satisfaction rate; the scheduling optimization objectives of the real-time power energy control layer include minimizing the total power system cost, maximizing the renewable energy absorption rate, and maximizing the power load satisfaction rate. The scheduling optimization objective function of the non-electric energy coordination layer is: in, Total operating cost of non-electric systems; This is the reward coefficient; For non-electrical load satisfaction; The total operating cost of the non-electric system is: in, Non-electric costs for each industrial park; Maintenance costs for heat exchange between parks; Non-electric costs of the park: in, For gas purchase costs; For equipment operating costs; for Real-time natural gas prices; for time i Gas purchases by the energy park; , These are the operating cost coefficients for the combined heat and power unit (CHP) and the gas storage tank (GS), respectively. For combined heat and power (CHP) units; This refers to the lower heating value of natural gas. for i The amount of gas stored in the gas storage tank GS in the energy park at time T; Maintenance costs for heat exchange between industrial parks: in, Service fee per unit of heat energy; for Thermal power of interaction between the park and other parks; The non-electrical load satisfaction rate is: in, , These are the weighting coefficients for air load and heat load, respectively. , These are the amounts of air load and heat load to be removed, respectively. This refers to the amount of natural gas on the demand side, i.e., the gas load. The heat power, or heat load, is the heat power on the demand side. The scheduling optimization objective function of the real-time power energy control layer is: in, The total cost of the power system; , This is the reward coefficient; For renewable energy consumption rate; For electrical load satisfaction; The total cost of the power system is: in, For electricity purchase costs; For equipment operating costs; To share the operating costs of energy storage; in, for t Electricity price at any given time; for i Energy Park t The amount of electricity purchased at any given time; , These are the operating cost coefficients for P2G (electric to gas conversion equipment) and EB (electric boiler); Service factor per unit electricity of the gas storage tank GS; The energy storage cost per unit of electricity in the gas storage tank GS; The renewable energy integration rate is: in, Let be the power generation capacity of the photovoltaic units in Energy Park at time t; Let be the power generation capacity of the wind turbines in Energy Park i at time t. , They are respectively i Photovoltaic units and wind turbines installed in the energy park t Predicted power generation capacity at any given time; Electrical load satisfaction index: in, for Energy Park At any given moment, the demand-side electrical power is the electrical load demand. This represents the amount of electrical load that can be cut off.

5. The method according to claim 2, characterized in that, The scheduling optimization model satisfies the following constraints: Long-term operational constraints include natural gas balance constraints, thermal energy balance constraints, inter-park thermal energy interaction balance constraints, and thermal energy interaction power upper limit constraints. Short-timescale operating constraints, including real-time power balance constraints of the power system; Natural gas balance constraints: in, for The demand-side natural gas volume in the energy park at time T; Thermal energy balance constraints: in, for Thermal power on the demand side of the energy park at time T; for The thermal power of the energy park interacting with other energy parks at time T; when When an energy park transfers heat energy to other energy parks, For positive, when When an energy park obtains heat energy from other energy parks, Negative; Inter-park thermal energy interaction balance constraints: in, , and These represent the heat power of each of the three energy parks interacting with the other two energy parks. A positive value indicates that heat energy is transferred to other energy parks, while a negative value indicates that heat energy is obtained from other energy parks. Upper limit constraint on thermal energy interaction power: in, The maximum limit for the interactive thermal power between different parks; Real-time power balance constraints of power systems: in, for Energy Park The demand-side electrical power at any given moment is the electrical load demand. , They are respectively Energy Park The charging and discharging power that constantly interacts with the shared energy storage center (ES).

6. The method according to claim 4, characterized in that, When solving the two-stage multi-energy cooperative optimization model, it is converted into a two-stage Markov decision process for solution, including: Markov Decision Processes (MDPs) consist of quadruples Define a state space. Describe the system's operating state and action space. Define executable control decisions and state transition probabilities. Representing system uncertainty, reward function To evaluate the effectiveness of action execution, a model-free reinforcement learning method is used, which does not directly learn the state transition probability function. Instead, it acquires experiential data through continuous interaction with the environment, dynamically adapts to the uncertainty of the source load, and after the agent selects and executes an action, the next state is determined only by the environment and fed back to the agent. Define a hierarchical state space: The system in The complete state at any given moment is controlled by a long time scale. and short-timescale control state Together they constitute: The long-time state space contains the operating states of slow-response devices and non-electrical load demand information in the system. The specific definition of the long-time state space is: in, for i The gas load demand of the energy park at time T; for i The amount of natural gas produced by the P2G gas-to-electric conversion equipment in the energy park at time T; The natural gas consumption of the CHP cogeneration unit in the i Energy Park at time T; for i The gas heat load demand of the energy park at time T; for i The heat output of the electric boiler EB at time T in the energy park; for i The heat production capacity of the combined heat and power unit CHP in the energy park at time T; for i The gas storage capacity of gas storage tank GS at time T in the energy park; for i Thermal power interaction between the park and other parks; The short-timescale state space reflects the real-time operating status of fast-response equipment and the immediate needs of the power system. The definition of the short-timescale state space is: in, for i Energy Park t The electrical load demand at any given time; Let be the power generation capacity of the photovoltaic units in Energy Park at time t; Let be the power generation capacity of the wind turbines in Energy Park i at time t; The electrical power consumed by P2G devices; The power generation capacity of a combined heat and power (CHP) unit; The electric power of the electric boiler; The interaction power between the park and the shared energy storage center; For the energy storage capacity of the shared energy storage center; Define a hierarchical action space: The long-time action space is responsible for the global optimization scheduling of slow-response devices in the non-electric energy coordination layer. The long-time action space is defined as follows: in, For the natural gas consumption regulation of combined heat and power units; This refers to the amount of thermal energy exchange power regulation between industrial parks. This is the adjustment amount for the gas storage capacity of the gas storage tank; The short-timescale action space is responsible for the real-time control of fast-response devices in the real-time control layer of power energy. The short-timescale action space is defined as follows: in, For the power regulation of P2G equipment; This refers to the power adjustment of the electric boiler. Adjustment of charging and discharging power for shared energy storage centers; Set up a tiered reward function: reward function The scheduling optimization objective function is composed of the non-electric energy coordination layer and the electric energy real-time control layer based on two-stage control, wherein... Indicates the state Next action The instant reward received; The short-timescale reward function is: in, , , These are the weighting coefficients; The total cost of the power system; Renewable energy integration rate; For electrical load satisfaction; The long-term reward function needs to be based on the long-term steps. Inside The cumulative instantaneous reward obtained within a short time step and a short time scale is adjusted, and the adjusted long-time scale reward function is defined as follows: in, , , These are the weighting coefficients; Total operating cost of non-electric systems; For non-electrical load satisfaction; Feedback and correction items for accumulating immediate rewards over a short time scale; Solving for the optimal policy through policy evaluation: By defining the action-value function Evaluation strategy Advantages and disadvantages, strategies This represents the mapping from the state space to the action space; in, This is the discount factor, representing the decay value of future rewards; The optimal strategy is equivalent to solving for the optimal Q-value function: in, This represents the function with the optimal Q value.

7. The method according to claim 1, characterized in that, The layered dual-delay deep deterministic policy gradient algorithm extends the Actor-Critic architecture of the TD3 algorithm into a layered structure, forming a two-layer decision-making system that matches the physical characteristics of the system. The upper-layer Actor-Critic network is responsible for global optimization decisions on a long time scale and controls slow-response devices; the lower-layer Actor-Critic network is responsible for real-time control on a short time scale and controls fast-response devices. In the upper-layer network, the Actor network The system receives a long-term state vector as input, which includes time information, electricity price information, non-electric load demand, and the operating status of each slow-response device. After processing through a two-layer fully connected neural network, it outputs a control action vector, which corresponds to the equipment adjustment of the cogeneration units (CHP), gas storage tanks (GS), and heat transmission network (HT) in all energy parks. The Critic network in the upper layer adopts a dual-Q network structure, namely the main Q network. and The two networks have the same structure but their parameters are updated independently. The problem of Q-value overestimation is alleviated by taking the minimum value, and an Actor network is also set up. Main Q network and Corresponding target network , and The target network parameters are slowly updated to track the main network, ensuring the stability of the training process. The lower-layer network structure corresponds to the upper-layer structure, but it is optimized for short-term timescale control. The lower-layer network contains an Actor network. It receives complete system status information and outputs action vectors corresponding to the equipment adjustment quantities of all energy park P2G (electric to gas conversion equipment), EB (electric boiler), and ES (shared energy storage center); the lower-layer Critic network also adopts a dual-Q network structure corresponding to the upper-layer network, with the main Q network... and The target network corresponding to the lower layer network , and However, the target network parameters in the lower layer network are updated more frequently to meet the needs of real-time control. The upper and lower layer networks are coupled by variables. Establish contact, among which, The output adjustment amount for the combined heat and power (CHP) unit. This refers to the power adjustment amount of the P2G (Power-to-Gas) converter. For the power adjustment of the electric boiler EB, the upper-level network transmits energy coordination instructions to the lower-level network through these coupling variables. The lower-level network performs local optimization and real-time response based on the coupling variables, and feeds back the execution results and immediate rewards to the upper-level network, realizing information sharing and strategy coordination optimization between the upper and lower-level networks.

8. The method according to claim 7, characterized in that, The training process for the upper-layer Actor-Critic network and the lower-layer Actor-Critic network includes: The training process employs a double-buffered experience replay mechanism, with the upper buffer... Storing empirical tuples over long time scales Lower buffer Storing empirical tuples on a short timescale ; During training, small batches of training samples are first uniformly and randomly sampled from the memory pool. For the lower-layer network, in the state... Actions are generated through the lower-level Actor network. ,in To explore noise, perform actions and receive immediate rewards. and the next state The experience is stored in the lower-level buffer, and during training, the target Q-value is calculated through the target network. in, The target policy noise is used to smooth the target policy. The parameters of the lower-level Critic network are updated by minimizing the TD error. The loss function is defined as: The lower-level Actor network employs a delayed update strategy, updating the policy gradient only every d=2 steps. The formula for calculating the policy gradient of the Actor network is: The training process for the upper-layer network is the same as that for the lower-layer network, but a reward feedback mechanism is added. At the end of each long time step T, the instantaneous rewards of all short time steps within that long time step are collected. This is added as a correction term to the reward function for long-running steps: in, These are the weighting coefficients. The modified long-term reward is an abstract reward coupling formula used to describe the reward transfer mechanism between upper and lower layers in the algorithm.

9. The method according to claim 8, characterized in that, The solution process for solving the two-stage multi-energy cooperative optimization model using the hierarchical dual-delay deep deterministic strategy gradient algorithm includes: Initialize the parameters of the upper and lower layer Actor network, Critic network, and corresponding target network parameters. The target network parameters are directly copied from the main network parameters. At the same time, initialize the experience buffers of the upper and lower layer networks to be empty sets. Before the training iteration begins, a certain amount of initial experience data is collected by randomly interacting with the environment to fill the buffers. After entering the main training loop, at the beginning of each long time step T, based on the current state... And add an upper-level Actor network to explore noise, and select long-term scale actions. Then it executes, and then enters the inner loop to execute continuously. In each short time step, based on the state... Choose actions with lower-level Actors And execute it to receive an instant reward. and new status Short-term experience is stored in the lower buffer, and network updates are immediately sampled from the lower buffer. Finish After a short time step, calculate the cumulative timely reward correction term to obtain the corrected long-timescale reward. Long-term experience is stored in the upper-layer buffer. Batch data is sampled from the upper-layer buffer to update the upper-layer network parameters. Every fixed number of steps, the target network parameters of the upper and lower Actor-Critic networks are updated using a soft update formula. in, The updated target network parameters; Main network parameters, For the parameters of the target network in the previous step, This refers to the soft update rate. The training process continues until the convergence condition is met: the loss function value fluctuates stably within the expected range or reaches the preset maximum number of training rounds; The converged network model is used for online decision-making in a multi-park integrated energy system. At each long time step and short time step, the optimal scheduling actions of the non-electric energy coordination layer and the electric energy real-time control layer are generated according to the real-time system status, so as to realize the two-stage optimized scheduling of the multi-park integrated energy system.

10. A hierarchical optimization system for scheduling strategies of multi-park integrated energy systems, used to implement the method of any one of claims 1 to 8, characterized in that, The hierarchical optimization system includes: At least three energy parks, each equipped with combined heat and power units, power-to-gas conversion equipment, electric boilers and gas storage tanks, as well as wind turbines and photovoltaic power generation systems; A shared energy storage center is connected to each park via the power network to achieve power sharing between the parks; An energy transmission network connects the various parks, including an electrical, gas, and heat transmission network, enabling the bidirectional flow and sharing of electrical, gas, and heat energy between the energy parks; A hierarchical scheduling controller is configured to solve the two-stage multi-energy collaborative optimization model using the hierarchical dual-delay deep deterministic policy gradient algorithm, perform collaborative optimization scheduling on long and short time scales, and generate and execute the optimal scheduling actions for the non-electric energy coordination layer and the electric energy real-time control layer at each long and short time step, respectively.