Improved td3 algorithm-based optimal scheduling method for integrated energy system

By improving the action space exploration strategy and reward function of the TD3 algorithm, the local optimum problem caused by sudden changes in cold and heat loads during seasonal transitions in the integrated energy system was solved, achieving optimal scheduling and robustness improvement of the system at all times and reducing operating costs.

CN122133999APending Publication Date: 2026-06-02SHANGHAI UNIVERSITY OF ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI UNIVERSITY OF ELECTRIC POWER
Filing Date
2026-02-25
Publication Date
2026-06-02

Smart Images

  • Figure CN122133999A_ABST
    Figure CN122133999A_ABST
Patent Text Reader

Abstract

This invention relates to an optimized scheduling method for integrated energy systems based on an improved TD3 algorithm. The method includes: constructing an optimized scheduling model for the integrated energy system, including constraints and an objective function, wherein the objective function aims to minimize the total operating cost of the system; improving the action space exploration strategy of the TD3 algorithm; solving the optimized scheduling model using the improved TD3 algorithm; and outputting the optimal scheduling scheme to control the operating status of each device within the integrated energy system. Compared with existing technologies, this invention leverages the advantage of the TD3 algorithm in handling high-dimensional continuous action spaces, enabling rapid and accurate acquisition of the optimal scheduling scheme. Furthermore, the improved action space strategy effectively addresses the local optima problem caused by sudden changes in cold and heat loads during seasonal transitions, thereby achieving optimal scheduling of the integrated energy system throughout all time periods and improving the optimized operation and robustness of the integrated energy system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of integrated energy system operation and control technology, and in particular to an optimized scheduling method for integrated energy systems based on an improved TD3 algorithm. Background Technology

[0002] Under the dual-carbon goals, integrated energy systems (IES), as important carriers for integrating multiple energy types, improving energy utilization efficiency, and reducing carbon emissions, can effectively improve the overall efficiency of social energy utilization, achieve sustainable energy supply, and enhance the flexibility, security, economy, and self-healing capabilities of energy supply. However, with the continuous increase in the penetration rate of renewable energy sources such as photovoltaics and wind power, and the frequent changes in various load demands, the uncertainty of source and load has increased significantly, posing enormous challenges to the optimized operation of integrated energy systems.

[0003] Existing optimization methods for integrated energy systems mainly include analytical methods and heuristic algorithms. Among them, analytical methods, such as branch and bound methods and dynamic programming algorithms, require simplification of complex models, are prone to computational and modeling errors, and are difficult to handle high-dimensional nonlinear optimization problems. Heuristic algorithms, such as genetic algorithms and particle swarm optimization algorithms, although less dependent on mathematical models, are prone to getting trapped in local optima and have long solution times, making them difficult to meet the requirements of online applications.

[0004] In recent years, deep reinforcement learning methods have been applied in the field of integrated energy system optimization due to their advantages of not requiring prior modeling and effectively handling uncertainty. Among them, the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, as an improved version of the Deep Deterministic Policy Gradient (DDPG) algorithm, features fast convergence speed and high robustness, and can handle optimization problems in high-dimensional continuous action spaces. However, when the integrated energy system undergoes seasonal changes, the TD3 algorithm encounters a significant challenge. The sudden shift in cold and heat loads from a long period of zero to large values ​​leads to drastic changes in the training data. Its batch normalization layer cannot adapt to the new data in time, causing the value function and policy function to collapse, resulting in local optima and affecting the optimization effect. This is detrimental to improving the operational performance and robustness of the integrated energy system. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and provide an optimized scheduling method for integrated energy systems based on the improved TD3 algorithm. This method can effectively solve the local optimum problem caused by sudden changes in cold and heat loads during seasonal transitions, and improve the optimized operation effect and robustness of the integrated energy system.

[0006] The objective of this invention can be achieved through the following technical solution: a comprehensive energy system optimization scheduling method based on an improved TD3 algorithm, comprising the following steps: S1. Construct an integrated energy system optimization scheduling model, including constraints and an objective function, wherein the objective function takes minimizing the total operating cost of the system as the optimization objective; S2. Improve the action space exploration strategy of the TD3 algorithm, and use the improved TD3 algorithm to solve the integrated energy system optimization scheduling model, and output the optimal scheduling scheme to control the working status of each device in the integrated energy system.

[0007] Furthermore, the total system operating cost in step S1 includes electricity purchase cost, gas purchase cost, electric energy storage system operation and maintenance cost, thermal energy storage system operation and maintenance cost, cold energy storage system operation and maintenance cost, refrigeration unit operation and maintenance cost, electric heating equipment operation and maintenance cost, and gas boiler operation and maintenance cost.

[0008] The objective function in step S1 is specifically: min F in, For electricity purchase costs; For gas purchase costs; For the operation and maintenance costs of the energy storage system; For the operation and maintenance costs of thermal energy storage systems; For the operation and maintenance costs of cold energy storage systems; For the operation and maintenance costs of the refrigeration unit; For the operation and maintenance costs of electric heating equipment; For the operation and maintenance costs of gas-fired boilers; for t Power purchased from the power grid at all times; for t The price of electricity purchased from the grid at all times; The scheduling time interval; for t Real-time power grid gas purchase capacity; for t Real-time gas purchase price from the grid; for Real-time charging power of the energy storage system; for Discharge power of the instantaneous energy storage system; This represents the operation and maintenance cost coefficient for the energy storage system. for t The thermal energy storage system's charging power at all times; fort The energy release power of the thermal energy storage system at all times; This represents the operation and maintenance cost coefficient for thermal energy storage systems. for t Real-time cold energy storage system charging power; for t Real-time cold energy storage system energy release power; This represents the operation and maintenance cost coefficient for cold energy storage systems. for t Refrigeration unit electrical power at all times; This represents the operation and maintenance cost coefficient for the refrigeration unit. for t The electrical power of the electric heating equipment at all times; This represents the operation and maintenance cost coefficient for electric heating equipment. for t Real-time gas boiler power; This represents the operation and maintenance cost coefficient for gas-fired boilers.

[0009] Furthermore, the constraints in step S1 include operational constraints of the electric energy storage system, the thermal energy storage system, the cold energy storage system, the refrigeration unit, the electric heating equipment, the photovoltaic equipment, the gas boiler, and the power balance constraint.

[0010] The specific constraints in step S1 are as follows: in, for t Energy storage system stores energy at all times; The energy loss coefficient of the energy storage system; Improve the charging efficiency of energy storage systems; Energy storage system energy release efficiency; These are the energy storage system's charging status parameters, with charging status represented by 1 and discharging status by 0. This represents the maximum charging power of the energy storage system. This represents the maximum energy release power of the energy storage system. This represents the minimum energy storage capacity of the energy storage system. This represents the maximum energy storage capacity of the energy storage system. The subscript... This indicates different types of energy storage systems. For electric energy storage systems, For thermal energy storage systems, It is a cold energy storage system. for t Refrigeration unit cooling capacity at all times; for t The constant power input to the refrigeration unit; This refers to the coefficient of performance (COP) of the refrigeration unit. Minimum input cooling power for the refrigeration unit This represents the maximum input refrigeration power of the refrigeration unit. for t The heating power of the electric heating equipment at all times; for t The electric power input to the electric heating equipment is constantly monitored. The coefficient of performance (COP) for electric heating equipment. Minimum input heating power for electric heating equipment This is the maximum input heating power for the electric heating equipment. for t The output power of photovoltaic equipment at all times; This represents the maximum power output of the photovoltaic equipment. for t The constant output heat power of the gas-fired boiler; for t Real-time conversion efficiency of gas-fired boilers; for t Constantly input natural gas power; Minimum output heat power limit for gas-fired boilers This is the maximum output thermal power limit for gas-fired boilers. for t The electrical load power of the integrated energy system at all times; for t The heat load power of the integrated energy system at all times; for t The cooling load power of the integrated energy system at all times; , and These represent the charge / discharge efficiencies of electrical energy storage, thermal energy storage, and cold energy storage, respectively.

[0011] Further, step S2 includes the following steps: S21. Based on the TD3 algorithm, construct an optimized scheduling framework for a comprehensive energy system, defining the state space, action space, and reward function; S22. Improve the action space exploration strategy of the TD3 algorithm by adjusting the action exploration strategy through environmental change recognition and improving the action exploration speed through action noise weighting correction. S23. Based on the improved TD3 algorithm, the optimal scheduling model of the integrated energy system is solved to obtain the optimal scheduling scheme.

[0012] Furthermore, the state space in step S21 includes the electrical load status, thermal load status, cold load status, photovoltaic power generation status, electrical energy storage capacity, thermal energy storage capacity, cold energy storage capacity, electricity price, and natural gas price during the scheduling period, comprehensively reflecting the system's operating status. The action space encompasses the charging and discharging power of electric energy storage systems, the charging and discharging power of thermal energy storage systems, the charging and discharging power of cold energy storage systems, the power of refrigeration units, the power of electric heating equipment, and the power of gas boilers, all of which are continuous variables; The reward value of the reward function consists of energy consumption cost penalty, energy storage system operation and maintenance cost penalty, equipment operation and maintenance cost penalty, and constraint violation penalty.

[0013] Furthermore, the reward function is specifically as follows: in, The energy cost penalty for the integrated energy system includes penalties for electricity purchase costs and penalties for gas purchase costs; Penalty for the operation and maintenance costs of energy storage systems; Penalties will be imposed on the operation and maintenance costs of refrigeration units, electric heating equipment, and gas-fired boilers. Penalties will be imposed for violating any of the constraints. This is the energy consumption penalty coefficient; This is a penalty coefficient for the operation and maintenance costs of energy storage systems. Penalty coefficients for the operation and maintenance costs of refrigeration units, electric heating equipment, and gas boilers; The penalty coefficient is the amount of punishment for violating each constraint. These are the penalty function coefficients for exceeding the power purchase limit of the power grid, the penalty function coefficients for exceeding the power limit of the gas boiler, the penalty function coefficients for exceeding the power limit of the refrigeration unit, the penalty function coefficients for exceeding the capacity limit of the electric energy storage system, the penalty function coefficients for exceeding the capacity limit of the thermal energy storage system, and the penalty function coefficients for exceeding the capacity limit of the cold energy storage system.

[0014] Furthermore, in step S22, the action exploration strategy is adjusted by identifying environmental changes. Specifically, the first start-up and shutdown of cold and hot loads during seasonal changes is used as an environmental change signal. When this environmental change signal is detected, it indicates that the system has entered a sudden change in operating conditions, and the action exploration strategy is adjusted accordingly.

[0015] Furthermore, the adjustment of the trigger action exploration strategy specifically involves adding a random amount to the action value after detecting an environmental change signal. The added random quantity specifically refers to: in, This is the proportionality coefficient; , This provides action information for typical days before and after the seasonal transition. Specifically, it covers the first day of cooling load activation during the winter-spring transition. This is information on typical daily movements in spring. This provides typical daily operation information for winter; on the first day of heat load activation during the summer-autumn transition, This is typical daily action information for autumn. This represents typical daily activity information during the summer.

[0016] Furthermore, in step S22, the speed of action exploration is improved by weighted correction of action noise. Specifically, after the environmental change signal is activated, noise is added to the output action value, and a large noise variance is used to increase the exploration boundary value.

[0017] Furthermore, the motion value corrected by motion noise weighting is specifically as follows: in, μ The output function of the main action network. Main action network parameters, The variables are random values ​​for a Gaussian variable, with an initial variance of 0.5 that gradually decreases as the number of training iterations increases, reaching a minimum of 0.01. As a weighted proportional adjustment variable, The initial value is 0.5, which is gradually decreased during training. When the number of training steps reaches 5000, it is reduced to 0.

[0018] Furthermore, the specific process of step S23 is as follows: Initialize the parameters of the main action network, target action network, main evaluation network, and target evaluation network of the TD3 algorithm, and set the experience replay pool capacity, batch training data size, action network learning rate, evaluation network learning rate, reward discount value, smooth update parameters, and action network delayed update frequency; By perceiving the operating status of the integrated energy system through state space, and outputting actions based on the improved exploration strategy, the system obtains reward values ​​and new states after acting on the environment. The state, action, reward, and new state are stored in the experience replay pool. When the data volume in the experience replay pool reaches a preset value, data is sampled in batches for network training. A delayed update strategy is adopted, where the action network is updated once after the evaluation network is updated multiple times, while the target network updates its parameters using a smooth update method. Repeat the training process until the algorithm converges and outputs the optimal scheduling scheme.

[0019] Compared with the prior art, the present invention has the following advantages: This invention first constructs an optimized scheduling model for an integrated energy system, with the goal of minimizing the total system operating cost. Furthermore, this invention improves the action space exploration strategy of the TD3 algorithm, and then uses the improved TD3 algorithm to solve the optimized scheduling model for the integrated energy system, outputting the optimal scheduling scheme to control the operating status of each device within the integrated energy system. By leveraging the TD3 algorithm's advantage in handling high-dimensional continuous action spaces, the optimal scheduling scheme can be obtained quickly and accurately. Moreover, the improved action space strategy effectively addresses the local optima problem caused by sudden changes in cold and hot loads during seasonal transitions, thereby achieving optimal scheduling of the integrated energy system throughout all time periods and improving the optimized operation and robustness of the integrated energy system.

[0020] This invention constructs an integrated energy system optimization scheduling framework based on the TD3 algorithm, defining a state space, action space, and reward function. Furthermore, it improves the action space exploration strategy of the TD3 algorithm by adjusting the action exploration strategy through environmental change identification. The environmental change identification uses "first start-up and shutdown of heat load" as a signal, which can accurately capture sudden changes in seasonal operating conditions and trigger timely adjustments to the exploration strategy. The action exploration speed is improved by weighted correction of action noise, which can reduce noise variance and accelerate stable convergence.

[0021] The reward function of the TD3 algorithm designed in this invention includes penalties for the energy consumption cost of the integrated energy system, the operation and maintenance cost of the energy storage system, penalties for violating various constraints, and penalties for the operation and maintenance costs of the refrigeration unit, electric heating equipment, and gas boiler. By combining linear penalties and nonlinear limit-crossing penalties, it can fully reflect the system operating cost and avoid variables from getting stuck in boundary values, which is beneficial to improving the learning effect of the algorithm in the subsequent solution process. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the integrated energy system structure; Figure 3 This is a schematic diagram of the TD3 agent in the embodiment; Figure 4 This is a schematic diagram of the photovoltaic curve in the embodiment; Figures 5a-5c The examples show the optimized results of electrical load, cooling load, and heating load during the transitional season. Figure 6a , 6b The results show the summer electrical load and cooling load optimization in the examples; Figure 7a , 7b The results of winter electrical and thermal load optimization are shown in the examples. Figure 8This is a schematic diagram of the electricity price curve in the example; Figures 9a-9c The example shows the optimization results of the electrical load, cooling load, and heating load on the first day of autumn during the transition from summer to autumn, based on the traditional TD3 algorithm. Figures 10a-10c The above example shows the optimization results of the electrical load, cooling load, and heating load on the first day of autumn during the transition from summer to autumn, based on the improved TD3 algorithm. Figure 11 This is a schematic diagram comparing the convergence process of the reward function values ​​during the training process of data from the same season in the embodiment. Figure 12 This is a schematic diagram comparing the convergence process of the reward function value during the training process when the seasons change in the example; Detailed Implementation

[0023] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0024] Example To address the problem that existing optimization algorithms are prone to getting trapped in local optima during seasonal changes in cold and heat loads in integrated energy systems, this proposal suggests an optimized scheduling method for integrated energy systems based on an improved TD3 algorithm. Figure 1 As shown, it includes the following steps: S1. Construct an integrated energy system optimization scheduling model, including constraints and an objective function, wherein the objective function takes minimizing the total operating cost of the system as the optimization objective; S2. Improve the action space exploration strategy of the TD3 algorithm, and use the improved TD3 algorithm to solve the integrated energy system optimization scheduling model, and output the optimal scheduling scheme to control the working status of each device in the integrated energy system.

[0025] This embodiment applies the above solution, addressing the following: Figure 2 The integrated energy system shown uses the annual photovoltaic output and the power of electricity, heat, and cooling loads as basic data. The effectiveness of this scheme is verified by implementing the above-described methods and processes. In this embodiment, the simulation software environment for the integrated energy system is Python 3.7; TensorFlow 2.1; CUDA 10.1; cuDNN 7.6.5; the hardware configuration is: CPU AMD Ryzen 7 5800H, GPU RTX 3060 6G.

[0026] The specific implementation process includes: I. Constructing an Optimal Scheduling Model for Integrated Energy Systems In this scheme, the objective of integrated energy system optimization scheduling is to minimize the total system operating cost within the scheduling cycle. The total operating cost includes electricity purchase cost, gas purchase cost, and the operation and maintenance costs of various equipment, specifically expressed as follows: In the formula: For electricity purchase costs; For gas purchase costs; For the operation and maintenance costs of the energy storage system; For the operation and maintenance costs of thermal energy storage systems; For the operation and maintenance costs of cold energy storage systems; For the operation and maintenance costs of the refrigeration unit; For the operation and maintenance costs of electric heating equipment; For the operation and maintenance costs of gas-fired boilers; for t Power purchased from the power grid at all times; for t The price of electricity purchased from the grid at all times; The scheduling time interval; for t Real-time power grid gas purchase capacity; for t Real-time gas purchase price from the grid; for Real-time charging power of the energy storage system; for Discharge power of the instantaneous energy storage system; This represents the operation and maintenance cost coefficient for the energy storage system. for t The thermal energy storage system's charging power at all times; for t The energy release power of the thermal energy storage system at all times; This represents the operation and maintenance cost coefficient for thermal energy storage systems. for t Real-time cold energy storage system charging power; for t Real-time cold energy storage system energy release power; This represents the operation and maintenance cost coefficient for cold energy storage systems. for t Refrigeration unit electrical power at all times; This represents the operation and maintenance cost coefficient for the refrigeration unit. for t The electrical power of the electric heating equipment at all times; This represents the operation and maintenance cost coefficient for electric heating equipment. for t Real-time gas boiler power; This represents the operation and maintenance cost coefficient for gas-fired boilers.

[0027] Therefore, the objective function is determined as follows: In addition, the model's constraints include various equipment operation constraints and power balance constraints: In the formula: for t Energy storage system stores energy at all times; The energy loss coefficient of the energy storage system; Improve the charging efficiency of energy storage systems; Energy storage system energy release efficiency; These are the energy storage system's charging status parameters, with charging status represented by 1 and discharging status by 0. This represents the maximum charging power of the energy storage system. This represents the maximum energy release power of the energy storage system. This represents the minimum energy storage capacity of the energy storage system. This represents the maximum energy storage capacity of the energy storage system. The subscript... This indicates different types of energy storage systems. For electric energy storage systems, For thermal energy storage systems, It is a cold energy storage system. for t Refrigeration unit cooling capacity at all times; for t The constant power input to the refrigeration unit; This refers to the coefficient of performance (COP) of the refrigeration unit. Minimum input cooling power for the refrigeration unit This represents the maximum input refrigeration power of the refrigeration unit. for t The heating power of the electric heating equipment at all times; for t The electric power input to the electric heating equipment is constantly monitored. The coefficient of performance (COP) for electric heating equipment. Minimum input heating power for electric heating equipment This is the maximum input heating power for the electric heating equipment. for t The output power of photovoltaic equipment at all times; This represents the maximum power output of the photovoltaic equipment. for t The constant output heat power of the gas-fired boiler; for t Real-time conversion efficiency of gas-fired boilers; for t Constantly input natural gas power; Minimum output heat power limit for gas-fired boilers This is the maximum output thermal power limit for gas-fired boilers. for tThe electrical load power of the integrated energy system at all times; for t The heat load power of the integrated energy system at all times; for t The cooling load power of the integrated energy system at all times. , and These represent the charge / discharge efficiencies of electrical energy storage, thermal energy storage, and cold energy storage, respectively.

[0028] II. Model Parameter Settings like Figure 3 As shown, the TD3 algorithm framework includes one main action network, one target action network, two main evaluation networks, and two target evaluation networks. The main action network and the target action network are mapped to the real action space after passing through hidden layers with ReLU activation functions and then through Tanh layers. The inputs of the two evaluation networks and the two target evaluation networks are concatenated through a hidden layer and then input into a hidden layer with ReLU activation functions. The evaluation network parameters are updated using gradient descent, and the target action value is taken as the minimum value of the outputs of the two target evaluation networks. The action network uses a delayed update strategy, updating the action network once after updating the evaluation network multiple times. The target networks use a smooth update method. In this embodiment, the number of neurons in the action network and the judge network are 1024 and 768 respectively, the experience replay pool capacity is set to 100,000, the batch training data size is set to 1024, the learning rate of the action network is set to 0.0001, the learning rate of the judge network is set to 0.001, the reward discount value is set to 0.95, the action noise variance is set to 3, the smoothing update parameter is set to 0.1, the action network delayed update frequency is set to 3, and the sample size is 10,000. To achieve a soft update of the target network, For soft update coefficients This avoids training oscillations caused by parameter mutations and aligns with the parameter update theory of the TD3 algorithm.

[0029] The maximum charge / discharge power of the electrical energy storage is 15MW, the thermal energy storage is 10MW, and the cold energy storage is 8MW. The EER of the refrigeration unit is 3.5, and the COP of the electric heating unit is 2.8, all of which meet the "equipment operation boundary limit" theory in the constraints, ensuring that the optimization actions are within the physically feasible range; the operation and maintenance cost coefficient... Directly related to the cost term in the objective function This provides data support for the linear penalty part of the reward function.

[0030] III. Scheduling Result Analysis 3.1 Scheduling results of data from the same season Historical data from summer, transitional season, and winter were used respectively (e.g., Figures 4 to 8As shown, the TD3 algorithm model is trained to obtain parameter models for different seasons, and the test data of the current season is used to optimize and solve the problem in real time.

[0031] In this embodiment, the state space of the TD3 algorithm includes nine types of information: electric / heat / cooling load, photovoltaic output, energy storage capacity, electricity price, and gas price. This comprehensively covers the operating status of the integrated energy system, enabling the TD3 agent to accurately perceive environmental dynamics—such as during the transitional season from 1:00 to 7:00 (off-peak electricity hours). The agent perceives low electricity price signals through the state space and triggers optimized actions for "electric / heat / cooling energy storage charging" in the action space, which conforms to the reinforcement learning interaction logic of "state-action-reward".

[0032] The action space of the TD3 algorithm covers six types of continuous variables, including energy storage charging and discharging, and equipment output, which matches the core characteristic of the TD3 algorithm: "handling high-dimensional continuous control problems". During the transition season, energy storage and discharging occur from 8:00 to 10:00 (peak hours) and continuous discharging occurs from 19:00 to 21:00 (peak hours). The algorithm achieves a cost optimization strategy of "storing electricity during off-peak hours and discharging during peak hours" through action space search, which verifies the rationality of the action space design.

[0033] The reward function of the TD3 algorithm adopts a design of "linear cost penalty + nonlinear limit violation penalty", where the linear cost penalty is: energy cost penalty. Operation and maintenance cost penalties Used to guide algorithms and reduce operating costs; nonlinear penalty function The penalty is increased when the variable approaches the boundary to prevent the device from exceeding its limits. Experiments without constraint violations demonstrate that the above reward function effectively guides the agent to find the optimal solution within the constraints.

[0034] This solution improves the action exploration strategy of the TD3 algorithm by designing a strategy of "environmental change recognition + action noise weighting correction". Environmental change identification uses "first start-up and shutdown of heat load" as a signal to accurately capture sudden changes in operating conditions during seasonal transitions and trigger adjustments to exploration strategies; motion noise weighting correction passed Implementation: Initially, a large variance noise of 0.5 was used to expand the exploration range, which was gradually reduced to 0.01 during training. By introducing typical daily activity information before and after seasonal changes, we can clarify the direction of our exploration.

[0035] The improved TD3 algorithm can solve real-time optimization scheduling problems while satisfying system power balance and equipment constraints. Furthermore, it exhibits consistent optimization results under loads with different seasonal characteristics, demonstrating strong compatibility and reliability.

[0036] In this embodiment, under the same seasonal dataset, the optimization results of traditional DDPG, TD3, and the improved TD3 method proposed in this scheme are consistent.

[0037] 3.2 Scheduling Results and Performance Analysis During Seasonal Changes To illustrate the performance of the improved TD3 algorithm, this embodiment selects scenarios with significant changes in hot / cold loads, leading to substantial equipment start-ups and shutdowns, as test scenarios. Taking the typical scenario of summer transitioning to autumn as an example, when the hot load first appears in autumn, the combined load changes from electrical and cold loads to electrical, hot, and cold loads. Simulation results obtained using the TD3 and improved TD3 algorithms are used to test the performance of the proposed strategy.

[0038] The model network parameters for summer were obtained by training the model separately using summer data. Typical daily data from the transition season were used as test data to simulate the transition from summer to autumn. The simulation results obtained by using the TD3 algorithm and the improved TD3 algorithm on the first day of heat load are shown in Figure 9 and Figure 10, respectively.

[0039] The shortcomings of the original TD3 algorithm are exposed: The TD3 algorithm is an "online unsupervised learning that interacts with the environment". When the seasons change, the cold / heat load changes abruptly from a long-term zero state to a large value, which causes drastic changes in the training data. The mean and variance of the batch normalization layer cannot be adapted in time, and the value function and policy function collapse. The electric heating cannot be started in the original TD3 algorithm because the action space exploration is insufficient and the search direction is ambiguous, which leads to getting stuck in a local optimum. This is consistent with the theoretical conclusion that "the unimproved TD3 is prone to failure in optimization in the case of data mutation".

[0040] Table 1 shows the comparison of optimization results on the first day of the seasonal change. The total cost before the improvement was 32,425.46 yuan, and the total cost after the improvement was 32,373.58 yuan, a reduction of 0.16% in daily cost. This is less than the result obtained by the algorithm before the improvement, indicating that the simulation results of the improved TD3 algorithm are better than those of the original TD3 algorithm.

[0041] Table 1 Convergence process analysis 4.1 Analysis of the convergence process of data from the same season Training was performed using DDPG, TD3, and a modified TD3 algorithm, respectively. The convergence of the reward value during training on a stable dataset is as follows: Figure 11 As shown, from Figure 11 It can be seen that: As the number of training iterations increases, the reward value approaches a stable value. Specifically, the DDPG method stabilizes after 2500 training iterations, while the TD3 and improved TD3 methods stabilize after 1000 training iterations. Compared to the DDPG algorithm, the TD3 algorithm achieves better dynamic performance during the learning process. A higher reward value after a certain training period indicates a better policy performance. Under the same seasonal dataset, the reward values ​​of DDPG, TD3, and the improved TD3 methods are basically the same after stabilization, indicating that their optimization results are consistent. The improved TD3 algorithm and the unimproved TD3 algorithm show little difference in dynamic performance during training, and the steady-state results are the same. The results are consistent when training multiple times using data from the same season.

[0042] 4.2 Analysis of Convergence Process During Seasonal Changes Figure 12 The impact of the improved action exploration strategy on the convergence of the TD3 algorithm was compared. In a seasonal change scenario, the combined load suddenly changes from long-term electrical and cooling loads to electrical, heating, and cooling loads. The TD3 algorithm's convergence remained volatile, its learning stalled, and its reward value was lower than that of the improved TD3 algorithm, oscillating near a local optimum and failing to find the optimal solution. The improved TD3 algorithm, with its added exploration direction and decaying exploration mechanism, adapted to the training data and produced better prediction results. The algorithm converged, and the final reward value was essentially consistent with the training results on a stable dataset.

[0043] It should be noted that this embodiment assumes a scenario where the combined load suddenly changes from a long-term electrical and cooling load to a combination of electrical, heating, and cooling loads, representing a relatively extreme change. In extreme scenarios, large data variations can lead to insufficient effective data in the dataset, potentially causing the TD3 algorithm to fail to find the optimal solution without improvement. However, this situation is relatively rare, and as the optimization time increases, the TD3 algorithm may converge to the optimal solution.

[0044] In summary, this scheme applies the TD3 algorithm to the optimal scheduling of integrated energy systems. Leveraging its advantage in handling high-dimensional continuous action spaces, it improves the efficiency and accuracy of optimization solutions. The designed reward function combines linear and nonlinear over-limit penalties, reflecting both system operating costs and preventing variables from getting trapped in boundary values, thus enhancing the algorithm's learning effect. This scheme improves the action space exploration strategy of the TD3 algorithm by identifying environmental changes and weighting action noise, effectively solving the local optimum problem caused by sudden changes in cold and hot loads during seasonal transitions, ensuring optimization performance in extreme scenarios. Applying this scheme to practical applications, it demonstrates rapid convergence and high robustness in solving optimal scheduling schemes, achieving excellent optimization results in both stable and seasonal scenarios, reducing system operating costs, and making it suitable for real-time optimal scheduling of integrated energy systems.

Claims

1. A comprehensive energy system optimization scheduling method based on an improved TD3 algorithm, characterized in that, Includes the following steps: S1. Construct an integrated energy system optimization scheduling model, including constraints and an objective function, wherein the objective function takes minimizing the total operating cost of the system as the optimization objective; S2. Improve the action space exploration strategy of the TD3 algorithm, and use the improved TD3 algorithm to solve the integrated energy system optimization scheduling model, and output the optimal scheduling scheme to control the working status of each device in the integrated energy system.

2. The integrated energy system optimization scheduling method based on the improved TD3 algorithm according to claim 1, characterized in that, The total system operating cost in step S1 includes electricity purchase cost, gas purchase cost, operation and maintenance cost of electric energy storage system, operation and maintenance cost of thermal energy storage system, operation and maintenance cost of cold energy storage system, operation and maintenance cost of refrigeration unit, operation and maintenance cost of electric heating equipment, and operation and maintenance cost of gas boiler. The constraints in step S1 include the operation constraints of the electric energy storage system, the operation constraints of the thermal energy storage system, the operation constraints of the cold energy storage system, the operation constraints of the refrigeration unit, the operation constraints of the electric heating equipment, the operation constraints of the photovoltaic equipment and the gas boiler, and the power balance constraints.

3. The integrated energy system optimization scheduling method based on the improved TD3 algorithm according to claim 1, characterized in that, Step S2 includes the following steps: S21. Based on the TD3 algorithm, construct an optimized scheduling framework for a comprehensive energy system, defining the state space, action space, and reward function; S22. Improve the action space exploration strategy of the TD3 algorithm by adjusting the action exploration strategy through environmental change recognition and improving the action exploration speed through action noise weighting correction. S23. Based on the improved TD3 algorithm, the optimal scheduling model of the integrated energy system is solved to obtain the optimal scheduling scheme.

4. The integrated energy system optimization scheduling method based on the improved TD3 algorithm according to claim 3, characterized in that, The state space in step S21 includes the electrical load status, thermal load status, cold load status, photovoltaic power generation status, electrical energy storage capacity, thermal energy storage capacity, cold energy storage capacity, electricity price, and natural gas price during the scheduling period, comprehensively reflecting the system's operating status. The action space encompasses the charging and discharging power of electric energy storage systems, the charging and discharging power of thermal energy storage systems, the charging and discharging power of cold energy storage systems, the power of refrigeration units, the power of electric heating equipment, and the power of gas boilers, all of which are continuous variables; The reward value of the reward function consists of energy consumption cost penalty, energy storage system operation and maintenance cost penalty, equipment operation and maintenance cost penalty, and constraint violation penalty.

5. The integrated energy system optimization scheduling method based on the improved TD3 algorithm according to claim 4, characterized in that, The reward function is specifically as follows: in, The energy cost penalty for the integrated energy system includes penalties for electricity purchase costs and penalties for gas purchase costs; Penalty for the operation and maintenance costs of energy storage systems; Penalties will be imposed on the operation and maintenance costs of refrigeration units, electric heating equipment, and gas-fired boilers. Penalties will be imposed for violating any of the constraints. This is the energy consumption penalty coefficient; This is a penalty coefficient for the operation and maintenance costs of energy storage systems. Penalty coefficients for the operation and maintenance costs of refrigeration units, electric heating equipment, and gas boilers; The penalty coefficient is the amount of punishment for violating each constraint. These are the penalty function coefficients for exceeding the power purchase limit of the power grid, the penalty function coefficients for exceeding the power limit of the gas boiler, the penalty function coefficients for exceeding the power limit of the refrigeration unit, the penalty function coefficients for exceeding the capacity limit of the electric energy storage system, the penalty function coefficients for exceeding the capacity limit of the thermal energy storage system, and the penalty function coefficients for exceeding the capacity limit of the cold energy storage system.

6. The integrated energy system optimization scheduling method based on the improved TD3 algorithm according to claim 3, characterized in that, In step S22, the action exploration strategy is adjusted by identifying environmental changes. Specifically, the first start-up and shutdown of cold and hot loads during seasonal changes is used as an environmental change signal. When this environmental change signal is detected, it indicates that the system has entered a sudden change in operating conditions, and the action exploration strategy is adjusted accordingly.

7. A comprehensive energy system optimization scheduling method based on an improved TD3 algorithm according to claim 6, characterized in that, The adjustment of the trigger action exploration strategy specifically involves adding a random amount to the action value after detecting an environmental change signal. The added random quantity specifically refers to: in, This is the proportionality coefficient; , This provides action information for typical days before and after the seasonal transition. Specifically, it covers the first day of cooling load activation during the winter-spring transition. This is information on typical daily movements in spring. This provides typical daily operation information for winter; on the first day of heat load activation during the summer-autumn transition, This is typical daily action information for autumn. This represents typical daily activity information during the summer.

8. The integrated energy system optimization scheduling method based on the improved TD3 algorithm according to claim 7, characterized in that, In step S22, the speed of motion exploration is improved by weighted correction of motion noise. Specifically, after the environmental change signal is activated, noise is added to the output motion value, and a large noise variance is used to increase the exploration boundary value.

9. A comprehensive energy system optimization scheduling method based on an improved TD3 algorithm according to claim 8, characterized in that, The motion value after motion noise weighting correction is specifically: in, μ The output function of the main action network. Main action network parameters, For random values ​​of the Gaussian variable, This is a weighted proportional adjustment variable.

10. A comprehensive energy system optimization scheduling method based on the improved TD3 algorithm according to any one of claims 4 to 9, characterized in that, The specific process of step S23 is as follows: Initialize the parameters of the main action network, target action network, main evaluation network, and target evaluation network of the TD3 algorithm, and set the experience replay pool capacity, batch training data size, action network learning rate, evaluation network learning rate, reward discount value, smooth update parameters, and action network delayed update frequency; By perceiving the operating status of the integrated energy system through state space, and outputting actions based on the improved exploration strategy, the system obtains reward values ​​and new states after acting on the environment. The state, action, reward, and new state are stored in the experience replay pool. When the data volume in the experience replay pool reaches a preset value, data is sampled in batches for network training. A delayed update strategy is adopted, where the action network is updated once after the evaluation network is updated multiple times, while the target network updates its parameters using a smooth update method. Repeat the training process until the algorithm converges and outputs the optimal scheduling scheme.