Comprehensive energy system optimization method and system in multiple uncertain environments
Through the TD3 and MPC collaborative optimization framework and the hybrid priority experience replay mechanism, the problems of conservative operation strategies and decision reliability caused by multiple uncertainties in the integrated energy system are solved, and a more efficient and stable optimization effect is achieved.
Patent Information
- Application Number
- CN202511315858.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-16
AI Technical Summary
Traditional integrated energy system optimization methods are unable to effectively deal with multiple uncertainties, resulting in overly conservative operating strategies, waste of resources, and decision reliability relying on the accuracy of the prediction model. Deep reinforcement learning has problems in the scheduling field, such as difficult parameter adjustment, sparse rewards, and insufficient stability.
The TD3 and MPC collaborative optimization framework is adopted, combined with a hybrid priority experience replay mechanism. Through dynamic adjustment of double-buffer sampling weights, the demonstration experience generated by MPC actors is learned first, multiple uncertainty association rules are independently explored, and the adaptive ability of the intelligent agent is improved.
It improves the operating efficiency and decision-making stability of the integrated energy system, simplifies the difficulty of parameter adjustment, enhances the adaptability to multiple uncertainties, and achieves faster convergence speed and better optimization results.
Smart Images

Figure CN120822667A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of integrated energy systems, and in particular relates to an integrated energy system optimization method and system under multiple uncertain environments. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Compared to traditional independent, distributed energy systems, the Integrated Energy System (IES), by leveraging its fundamental principles of multi-energy complementarity and coordinated optimization, can break down operational barriers between different energy types, thereby improving renewable energy consumption and overall system energy efficiency. However, its multi-physics coupling characteristics also create significant risks in the cross-domain transmission of multiple uncertainties while achieving multi-energy complementarity, directly threatening system operational safety.
[0004] Existing research typically decouples IES external uncertainties into three categories: renewable energy sources such as wind and solar, multiple loads such as cooling, heating, and electricity, and energy prices. Traditional scheduling methods often focus on single or local sources of uncertainty. Traditional mechanisms based on robust optimization, stochastic optimization, and information gap decision-making achieve pseudo-joint modeling through simple linear superposition of constraint boundaries or probabilistic scenario enumeration, which makes it difficult to characterize the inherent connections between multiple uncertainties. The direct consequence is that to cover all possible scenarios, a large number of boundary constraints and redundant scenarios must be introduced, resulting in overly conservative operating strategies and significant waste of resources in responding to low-probability or even unrealizable scenario combinations. Furthermore, these methods typically require mathematical models or predictive information, and the reliability of the final decision depends largely on the accuracy of the system model and prediction method used. This makes it difficult to adapt to the high-dimensional, non-convex uncertainty space of wind, solar, cooling, heating, electricity, and prices.
[0005] Artificial intelligence is reshaping the approach to solving energy problems. Deep reinforcement learning (DRL), a leading example, overcomes the curse of dimensionality inherent in traditional mechanistic modeling through its end-to-end model-free learning capabilities, enabling data-driven discovery of implicit association rules in high-dimensional, non-convex, uncertain spaces. However, while DRL offers significant theoretical advantages for energy management applications, it still suffers from numerous practical drawbacks: First, representative algorithms, such as the double-delayed policy gradient (TD3), are extremely sensitive to hyperparameters, difficult to tune, and have weak generalization capabilities. Although their performance ceiling is high, their floor is also low, resulting in insufficient stability. Secondly, the reward function design is highly complex. If traditional operating costs are directly used as rewards, it is easy to cause the rewards to converge but the agent's actions still violate the system's hard constraints. Even if the constraints are expressed by introducing penalty terms, it is still difficult to effectively ensure the feasibility of the constraints. Third, the source-charge-price environment parameters in IES are partially observable and cannot fully satisfy the Markov decision assumption, resulting in limited learning and unstable performance of the agent. Finally, IES has a high-dimensional joint state-action space, which makes the rewards highly sparse in the early stages of training. Without prior conditions, the intelligent agent tends to converge the action value to the boundary area and cannot fully explore the optimal strategy. At the same time, the training process converges slowly and is inefficient.
[0006] However, this does not mean that DRL has fundamental limitations in the scheduling field. Its powerful adaptive environment representation capabilities provide irreplaceable advantages for coping with multiple cross-domain uncertainties. The core challenge lies in improving DRL agent-environment interaction, retaining the advantages of artificial intelligence while also aligning with the essence of IES optimization problems. Summary of the Invention
[0007] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides an integrated energy system optimization method and system under multiple uncertain environments, proposes a twin-delayed deep deterministic policy gradient (TD3) and model predictive control (MPC) collaborative optimization framework, and on this basis, designs a hybrid priority experience replay mechanism to coordinate the balance between expert experience and autonomous exploration. By dynamically adjusting the double-buffer sampling weights, the demonstration experience generated by MPC actors is preferentially learned at the beginning of training to avoid the blindness of random strategies. As the interaction data accumulates, the autonomous sampling weight is gradually improved, the intelligent agent is encouraged to explore autonomously, and the source-load-price cross-domain uncertainty association rules are adaptively adopted, so as to achieve better performance in the IES environment with multiple uncertainty coupling.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions: A first aspect of the present invention provides a method for optimizing an integrated energy system under multiple uncertain environments, comprising: Obtain the status of the integrated energy system; Based on the state, the multi-energy flow device action is obtained through the actor network; Among them, during the training process of the actor network, the MPC actor is embedded in the dual-delay deep deterministic policy gradient architecture, and the transfer tuples of the actor network and the MPC actor are stored in the agent experience replay pool and the expert experience replay pool respectively. According to the temporal differential error of the samples in the two experience replay pools, each sample is given a priority, and the mixing ratio is determined according to the training time step. Combining the mixing ratio and priority, the sampling probability of the sample is calculated, and then the update of the comment network is controlled to have a high dependence on the expert experience replay pool in the early stage of training and a high dependence on the agent experience replay pool in the later stage of training. The value gradient of the comment network to the action is used to guide the actor network to optimize the output action.
[0009] Furthermore, the status includes wind power generation power, photovoltaic power generation power, thermal load, electric load, gas load, electricity price, output power of cogeneration device, input power of power-to-gas and charge state of energy storage within the time segment composed of the current time period as the midpoint.
[0010] Furthermore, the multi-energy flow device actions include changes in charging power, changes in discharging power, changes in purchased electricity power, changes in gas input power of the cogeneration device, changes in power-to-gas input power, changes in gas boiler input power and changes in gas purchased power.
[0011] Furthermore, during the training of the actor network, for each time step, the following steps are performed: The MPC actor selects an action in the current time step state and interacts with the environment to obtain rewards and the next time step state. The MPC actor's transfer tuple is stored in the expert experience replay pool and the priority of each sample in the expert experience replay pool is set. The actor network selects an action in the current time step state and interacts with the environment to obtain rewards and the next time step state. The actor network's transfer tuple is stored in the agent experience replay pool and the priority of each sample in the agent experience replay pool is set. If the current time step is an integer multiple of the experience replay interval, then based on the expert experience replay pool and the agent experience replay pool, combined with the target network, multiple small batch update operations are performed on the comment network, and the priority of the samples in the experience replay pool is updated; If the current time step is an integer multiple of the set interval, the actor network and target network are updated.
[0012] Furthermore, the reward is the inverse of the sum of the comprehensive energy system operating costs, the wind and solar power curtailment penalty items, and the equipment over-limit penalty items.
[0013] Furthermore, the small batch update operation includes: According to the sampling probability of the sample, the transfer data is sampled from the experience replay pool; For the sampled transfer data, calculate the importance sampling weight based on the sampling probability; At the current time step, the target actor network is used to generate the target action, Gaussian noise is added to the target action, and the temporal difference target is calculated based on the target critic network. The difference between the temporal difference target and the critic network result is calculated to obtain the temporal difference error. Update the priority of samples in the experience replay pool based on the temporal difference error; Based on the temporal difference error and importance sampling weights, the loss of the review network is calculated and the review network parameters are updated.
[0014] Furthermore, the sampling probability is: ; Among them, P j is the sampling probability of the jth sample; linear weight attenuation coefficient , and are the initial and final attenuation coefficients, T is the total number of training steps, l Indicates the current time step; type j = MPC represents that the jth sample comes from the expert experience replay pool, type j = Agen t indicates that the jth sample comes from the agent experience replay pool; D agent and D mpc are the number of samples in the agent experience replay pool and the expert experience replay pool respectively; β is the priority intensity coefficient; p j is the priority of the j-th sample, p z and p d Represent the priorities of samples from the agent experience replay pool and the expert experience replay pool respectively.
[0015] A second aspect of the present invention provides an integrated energy system optimization system under multiple uncertain environments, comprising: A data acquisition module is configured to: acquire the status of the integrated energy system; The optimization module is configured to: obtain the multi-energy flow device action through the actor network based on the state; Among them, during the training process of the actor network, the MPC actor is embedded in the dual-delay deep deterministic policy gradient architecture, and the transfer tuples of the actor network and the MPC actor are stored in the agent experience replay pool and the expert experience replay pool respectively. According to the temporal differential error of the samples in the two experience replay pools, each sample is given a priority, and the mixing ratio is determined according to the training time step. Combining the mixing ratio and priority, the sampling probability of the sample is calculated, and then the update of the comment network is controlled to have a high dependence on the expert experience replay pool in the early stage of training and a high dependence on the agent experience replay pool in the later stage of training. The value gradient of the comment network to the action is used to guide the actor network to optimize the output action.
[0016] Furthermore, the status includes wind power generation power, photovoltaic power generation power, thermal load, electric load, gas load, electricity price, output power of cogeneration device, input power of power-to-gas and charge state of energy storage within the time segment composed of the current time period as the midpoint.
[0017] Furthermore, the multi-energy flow device actions include changes in charging power, changes in discharging power, changes in purchased electricity power, changes in gas input power of the cogeneration device, changes in power-to-gas input power, changes in gas boiler input power and changes in gas purchased power.
[0018] Compared with the prior art, the present invention has the following beneficial effects: This paper proposes a TD3 and MPC collaborative optimization framework. On this basis, in order to coordinate the balance between expert experience and autonomous exploration, a hybrid priority experience replay mechanism is designed. Through dynamic adjustment of double-buffer sampling weights, the demonstration experience generated by MPC actors is prioritized in the early stage of training to avoid the blindness of random strategies. As the interaction data accumulates, the autonomous sampling weight is gradually improved, which encourages the intelligent agent to explore independently and adapts the source-load-price cross-domain uncertainty association rules.
[0019] The TD3 and MPC collaborative optimization framework proposed in this paper inherits DRL's ability to learn and deal with uncertainty from the environment, and MPC's ability to use physical models to provide basic performance. Compared with traditional TD3, TD3MPC has a simpler reward function design, faster convergence speed, better results, and significantly reduced parameter adjustment difficulty.
[0020] The present invention expands the single-moment state into an extended state space covering historical and future information, and models the state transition characteristics of the Markov decision process based on probabilistic scenario technology to enhance the generalization of TD3 for complex energy problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0022] Figure 1 TD3MPC architecture diagram of the first embodiment of the present invention; Figure 2 This is a structural diagram of the IES according to the first embodiment of the present invention; Figure 3 is the TD3MPC convergence result of the first embodiment of the present invention; Figure 4 This is the convergence comparison result of Example 1 of the present invention; Figure 5 This is the multi-scenario scheduling result of the first embodiment of the present invention. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0024] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0025] Example 1 This embodiment provides a comprehensive energy system optimization method under multiple uncertain environments.
[0026] The core of solving the problems described in the background technology is to compress the exploration space of the intelligent agent to the physically feasible domain and the high-value decision-making area, thereby increasing the reward density. One feasible method is to introduce an expert guidance mechanism for DRL. Model predictive control (MPC) has become the first choice of experts due to its open architecture design and mature theoretical system. In fact, MPC and DRL each have their own advantages and disadvantages and complement each other. The collaboration between the two can not only solve the problems of DRL hyperparameter sensitivity, sparse rewards and poor interpretability through physical guidance, but also make up for the inherent defects of MPC itself, which is model dependence and short-sightedness. In view of the multiple uncertainties faced by IES, the synergy between MPC and DRL is expected to play an important role.
[0027] This embodiment provides a method for optimizing an integrated energy system under multiple uncertainties. It proposes a collaborative optimization framework for TD3 and MPC. Within this framework, Model Predictive Control (MPC) is embedded as an expert actor within a Twin Delayed Deep Deterministic Policy Gradient (TD3) architecture. This framework uses rolling-horizon optimization to provide a safe baseline action sequence that complies with physical constraints. TD3, driven by data, further adaptively explores within the neighborhood, compensating for model mismatches and adapting to multiple uncertainties. A hybrid-priority experience replay mechanism is designed to balance expert experience with autonomous exploration. Furthermore, the single-moment state is expanded into an extended state space encompassing historical and future information, and the state transition characteristics of the Markov decision process are modeled using probabilistic scenario techniques to enhance TD3's generalization to complex energy problems.
[0028] This embodiment provides a method for optimizing an integrated energy system under multiple uncertain environments, including the following steps: Step 1: Obtain historical source-load-price data of the integrated energy system, establish a physical model of the park, including the objective function (reward function), equality constraints, and inequality constraints, and convert it into a Markov decision process.
[0029] like Figure 2 As shown in the figure, the northern IES demonstration park under study incorporates three heterogeneous energy sources: electricity, heat, and gas, all centrally managed by a unified operator. The park draws energy from the upstream power grid, gas grid, wind power, and photovoltaic power plants. These energy sources are interconnected through coupled devices such as combined heat and power (CHP), gas boilers (GB), and power-to-gas (P2G), forming a coordinated multi-energy flow network that collectively meets the diverse energy needs of end users. Random fluctuations in renewable energy, multi-energy loads, and price signals propagate through physical coupling and market mechanisms, severely impacting the stable operation of the IES. It is important to note that the integrated energy system under study is a typical low-voltage distribution network architecture, with a compact electricity / heat / gas network topology and short transmission distances. The network power flow distribution is negligibly affected by line parameters and pipeline pressure drops. The integrated energy system is modeled according to the integrated energy system structure diagram.
[0030] (1) Energy conversion equipment group.
[0031] The energy conversion equipment in the integrated energy system includes CHP, P2G and GB. The input and output power balance equation is as follows: (1); in, 、 They are lThe CHP outputs electrical power and thermal power during the time period; 、 are the electricity and gas production efficiencies of CHP, respectively; for l CHP input gas power during the period; 、 They are l The output gas power and input electric power of P2G in the time period, is the power-to-gas efficiency; 、 They are l The output thermal power and input electrical power of the time period EB, is the heating efficiency of the electric boiler (EB); 、 They are l The output heat power and input gas power of GB in this period, is the heat generation efficiency of GB.
[0032] (2) Energy storage system.
[0033] The integrated energy system is equipped with a power storage system, and the charging and discharging balance model is as follows: (2); in, for l The energy storage period, is the self-loss coefficient, and They are l The charging and discharging power of the time period, and are respectively the charging and discharging energy efficiency, E bat is the total capacity of the energy storage device, and is a 0-1 variable, 1 represents working state, 0 represents non-working state, l Indicates a time interval.
[0034] (3) Rated power and climbing power constraints.
[0035] (3); in, 、 、 They are l Time period equipment i Output power and output power upper and lower limits; Changing power for equipment; 、 They are the upper and lower limits of climbing power respectively.
[0036] (4) Power balance constraints.
[0037] (4); Where, 、 、 They are l Changes in electricity, heat, and gas demand during the time period, 、 They are l Changes in electricity purchases and gas purchases during the time period, for l The change in P2G input power during the time period, for l The change of photovoltaic power generation during the period, for l The change of wind power generation during the period, 、 They are l The change in CHP output electrical power and thermal power during the period, 、 They are l The changes in the output heat power and input gas power of GB during the period, for l The change of CHP input gas power during the period, for l The change in P2G output gas power during the period.
[0038] Step 2: Adaptively improve the traditional TD3 state space and design an extended state space that includes historical and future information to ensure the consistency of the Markov decision process; and model the transition probability of random state variables based on probabilistic scenario generation technology to avoid the interference of abnormal probabilistic scenarios on the overall strategy; establish an MPC Actor based on MPC on the basis of the traditional actor network, explicitly consider the physical constraints of state and input, generate a safe and economical benchmark control action sequence and store it together with the state-reward signal in the expert experience replay pool (Demo buffer). Its objective function is co-designed with the improved TD3 reward function to ensure consistency of optimization goals; the improved TD3 agent further explores the global optimal strategy in the neighborhood of the benchmark control action taken from the Demo buffer, and stores the exploration experience in the agent experience replay pool (Agentbuffer).
[0039] The core advantage of MPC is that it can generate the optimal control sequence by explicitly solving the constrained optimization problem online with the help of dynamically updated system state and environmental information. It is worth noting that MPC and DRL have significant isomorphism in mathematical form. Both are based on state variables, control variables and objective functions as the basic framework. The formal similarity is also one of the reasons for choosing MPC for physical guidance. MPC Actor inherits the standard MPC framework, uses a state space model to describe the system dynamics, and generates a benchmark control strategy in a deterministic scenario. Unlike traditional MPC that only uses the first-step control variable, MPCActor stores the optimal control sequence and its corresponding state-reward information in the entire time domain to the Demo buffer to construct a training sample set that matches the extended state space. The following optimization problem is solved in each rolling time domain: (5); in, z To predict the time domain output; z max and z min Respectively represent the maximum and minimum values of the predicted time domain output; w is the reference trajectory, obtained by optimizing once a day ago, and the optimization objective function is consistent with the TD3 reward function; and is the weight factor; is the state variable matrix of the multi-energy flow coupling device; is the coefficient matrix; is the future control increment matrix of the multi-energy flow coupling device, △ u ( l )express l Control increment of time period; H is a lower triangular matrix whose nonzero elements are defined as ; is the future disturbance vector, including wind and solar energy, electricity, heat, gas loads and electricity prices; d ( l )express l The disturbance variables of the time period; in MPC Actor, the source-load-price random variables are assumed to be deterministic and obtained by the prediction model; N p and H d These are the lengths of the prediction time domain and the control time domain, and the selection rules will be introduced later; H d is a lower triangular matrix whose non-empty elements are defined as ; I is the identity matrix; T is a lower triangular matrix whose non-zero elements are all identity matrices;u min and u max are the minimum and maximum values of the control quantity respectively; u min and △ u max are the minimum and maximum values of the control increment respectively; y min and y max are the minimum and maximum values of the system’s predicted output respectively; Q 、 N 、 N d and M are the coefficient matrices of the single-step state-space model respectively.
[0040] The single-step state space model can be expressed as: (6); in, A 、 B 、 B d 、 C is a coefficient matrix, which can be given according to the multi-energy flow coupling device constraints and power balance equation; this embodiment only gives the power system state space model; the state variable matrix ,in, for l The amount of electricity purchased during the period; the control variable matrix ; ; Disturbance variable matrix According to the energy storage SOC constraint, power balance constraint and other constraints, we can get: ; ; ; ;in, , .
[0041] The TD3 algorithm is an improved Actor-Critic framework proposed for the control problem of high-dimensional continuous action space. Compared with the traditional deep deterministic policy gradient, it significantly improves the stability and convergence efficiency of the policy through the triple technical innovations of dual critic network, delayed policy update and target policy smoothing. IES is modeled as a partially observable Markov decision process, which can be represented by the five-tuple Indicates that, except for multiple random variables, the remaining state transition probabilities P will be implicitly defined by the kinetic equations; is the reward discount factor, which takes a fixed value of 0.9; state space S、 Action Space A And the reward functionR is defined as follows: State Space S : DRL usually assumes that the system satisfies the Markov decision principle, that is, the future state depends only on the current state and has nothing to do with the previous state. However, the source-charge-price random variables in IES often show periodic or inertial characteristics. This causes the Markov assumption of the traditional TD3 algorithm to fail, and strategies that only rely on the current state are prone to suboptimal actions. To this end, this embodiment constructs an extended state space, and uses time-series state fragments with past and future information instead of a single state variable to represent the system state, improving the traditional one-to-single state-action mapping to a many-to-single mapping to improve the completeness of state information. l The time period extended state space includes wind and solar power generation, thermal, electrical, and gas loads, electricity prices, and the operating status of multi-energy flow coupling equipment, which can be expressed as: ;in, Status i exist l - n to l + n Status fragments of the period, specifically: Indicates that wind power l-n to l+n The power of the period, Indicates that photovoltaic l-n to l+n The power of the period, Indicates l-n to l+n The electricity price for the time period, Indicates l-n to l+n The electrical load during the period, Indicates l-n to l+ n Heat load during the period, Indicates l-n to l+n Gas load during the period, Indicates that CHP is l-n to l+n Output power of the time period, Indicates P2G in l-n to l+n The input power of the period, Indicates energy storage l-n to l+n SOC (state of charge) during the time period; l - n to l -1 period state variables are all real values; l to l + nThe source-load-price random state variables of the time period are obtained by relying on the external point prediction model; n It can be determined based on the results of autocorrelation analysis; the state variables of the coupled equipment are determined by the MPC optimization results.
[0042] Action Space A : Actions can adjust variables. In this embodiment, the internal power balance of the integrated energy system is ensured by the actions of multiple energy flow devices. The TD3 action space is consistent with the action sequence output by the MPC Actor and can be described as: ,in, Indicates charging power at l-n to l+n The change in time period, Indicates the discharge power at l-n to l+n The change in time period, Indicates the purchased power l-n to l+n The change in time period, Indicates the CHP input gas power l-n to l+n The change in time period, Indicates the P2G input power l-n to l+n The change in time period, Indicates GB input gas power l-n to l+n The change in time period, Indicates the gas purchasing power l-n to l+n The amount of change in the time period; during optimization, all actions in the state space are obtained, but only the first action is executed.
[0043] State transition probability P :The traditional TD3 algorithm is based on the Markov decision process assumption and only models the state transition through a single sample trajectory. It is difficult to describe the randomness of the state transition caused by the multiple uncertainties of source, charge and price in IES. Specifically, l Time period execution action a l back, l The random state variables of the +1 period cannot be accurately obtained. Relying solely on deterministic predictions will result in insufficient adaptability of the policy network to random environments.
[0044] This embodiment simulates the uncertainty in the Markov decision process using probabilistic scenario generation technology, probabilistically modeling the transition characteristics of the random state space. The transition characteristics of the remaining states can be determined by the system dynamics of the MPC actor. The expected value of the action-value function of the random state is used to update the TD target. The update process can be expressed as: (7); in, y l represents the TD target, express i target critic network, is the parameter of the i-th target critic network, is the discount factor, r l for l Reward at every moment; E ( ) is the expected function, clustering the five random scenarios of wind, solar, cooling, heating and power in the future. Since the clustering results in a discrete scenario set, It can be expressed as: (8); Among them, s k,l+1 for l The kth possible scenario in period +1, p ( s k,l+1 | s l , a l ) is the scene s k,l+1 The probability of occurrence, M is the number of clustering scenes, For the kth scenario l +1 period action; for ease of analysis, define the random state space ;exist l time, s l+1 middle l - n to l Random state space at time -1 w l+1 ( l - n , l -1) is known, and w l+1 ( l , l+n+ 1) There are many possibilities; to ensure the matching of multiple scenarios of source-load-price and w l+1 ( l - n , l -1) with w l+1 ( l , l+n+ 1) Scene matching, based on Latin hypercube sampling and error generation based on point prediction resultsl to l+n+1 The set of random state variable scenarios at time W ; Then the initial scene is reduced based on K-means clustering to obtain p( w k , l+1 | w l , a l )and w k,l+1 ; Since the rest of the states are deterministic, p( w k , l+1 | w l , a l ) is , w k,l+1 The combination of the determined state is The natural matching of multivariate prediction results can ensure that the generated scenario set meets the consistency requirements in the temporal and spatial dimensions; if there is a lack of prediction model support, the initial scenario can be obtained by sampling the fitted probability distribution when generating the scenario.
[0045] Reward Function R :DRL maximizes long-term rewards through continuous trial and error-feedback-evaluation, and sets the opposite of the optimization target as the immediate reward r l , which can be expressed as: (9); F 1 is the operating cost of the comprehensive energy system, which can be expressed as: (10); in, 、 They are l The electricity and gas prices at the time; 、 They are l Purchase electricity and gas power at all times; 、 They are l Moment i The operation and maintenance costs and output power of this type of equipment.
[0046] F 2 is the penalty term for curtailing wind and solar power, which can be expressed as: (11); in, 、 They are l The cost coefficients of wind and solar curtailment at the time; 、 are the amount of wind and solar power curtailment, respectively.
[0047] F 3 is the equipment over-limit penalty term. The penalty mechanism is set according to the constraint conditions and added to the immediate reward to obtain the final reward function. In order to reduce the fitting difficulty and improve the accuracy and convergence stability of the results, this embodiment uses linear penalty instead of step penalty. g ( x i )≥0, design the following penalty terms: (12); in, β i is the penalty coefficient, and the corresponding constant coefficient is set according to different limit-crossing penalties; g ( x i ) includes penalties for exceeding the output and climbing limits of each device.
[0048] Step 4: To balance expert experience with autonomous exploration, a hybrid-priority experience replay mechanism is designed. This dynamically adjusts the sampling weights in a double buffer, prioritizing learning from the demonstration experience generated by MPC actors during the initial training phase to mitigate the blindness of random strategies. As interaction data accumulates, the autonomous sampling weight is gradually increased, encouraging the agent to explore autonomously and adapting the source-load-price cross-domain uncertainty association rules.
[0049] Experience replay mechanism: Priority experience replay improves traditional experience replay by introducing a priority and importance weighting mechanism, giving priority to learning samples with high TD errors to improve the efficiency of strategy improvement. Taking into account that the experience replay pool of the proposed framework consists of two parts: Agent buffer and Demo buffer, this embodiment designs a new hybrid priority experience replay mechanism to balance the conflict between strategy exploration and knowledge transfer. It should be pointed out that the priority experience replay mechanism is not necessary when migrating and applying this framework. If the reward signal of the target problem is dense and the environment is dynamically stable, uniform replay can be used directly to avoid additional computational overhead, thereby reducing training complexity.
[0050] Assign priority to each sample based on the TD error: (13); in, is the expert experience priority attenuation coefficient, which is used to reduce the risk of overfitting the expert sample; is a very small constant, avoiding zero priority;Q (*) represents the review network, and are the state and action of the jth sample respectively; are the parameters of the review network, i =1 or 2, i.e., there are two review networks; is the temporal difference (TD) target of the j-th sample.
[0051] Define linear weight decay coefficient , control the mixing ratio of expert experience and autonomous exploration experience: (14); in, and are the initial and final attenuation coefficients respectively; the initial stage completely relies on expert experience, while the final stage mainly relies on independent exploration; T is the total number of training steps; l Indicates the current time step.
[0052] Define samples by combining mixing ratio and priority j The sampling probability of is: (15); in, type j = MPC Indicates that the jth sample comes from the Demo buffer, p j is the priority of the j-th sample, p z and p d Represents the priority of samples from Agent buffer (agent experience replay pool) and Demo buffer (expert experience replay pool), D agent and D mpc Refers to the number of samples in the two types of experience replay pools, Agent buffer and Demo buffer respectively; is an indicator function, which is 1 when the sample source matches, otherwise it is 0; β is the priority strength coefficient, which degenerates into uniform sampling when it is 0. To eliminate the deviation introduced by priority sampling, the importance sampling weight is calculated: (16); in, Represents the total number of samples in the two types of experience replay pools, P j represents the sampling probability of sample j, α It is a hyperparameter used to determine the effect of offsetting the priority experience replay on the convergence results.
[0053] Corresponding: (17); in, represents the loss function that introduces importance weighting, N is the number of samples drawn at a time.
[0054] So far, a collaborative optimization framework based on TD3MPC has been established. The overall framework is as follows: Figure 1 Shown, including: (1) Initialization parameters: mini-batch data size k , Experience playback interval K , total number of training steps T ; (2) Initialize the network: use random parameters 、 and , initialize two Critic networks (comment networks) respectively Q And Actor network π.
[0055] (3) Initialize the target network: Initialize the target network (target review network 1, target review network 2 and target actor network) parameters to 、 and .
[0056] (4) Initialize the experience replay pool: Create two types of experience replay pools D MPC and D Agent .
[0057] (5) Start the training cycle, for each time period (time step) l : (501) Based on MPC strategy interaction and experience storage: Obtaining integrated energy system l Period Status s l , according to MPCActor (MPC actor), in state s l Next select action a l,1 ,Right now a MPC , interact with the environment (through formula (9)) to get rewards r l,1 and the next period status s l+1,1 ; Transfer the MPC actor's transfer tuple ( s l , a l,1 , rl,1 , s l+1,1 ) Deposit into D MPC , priority p l,1 Set as .
[0058] (502) Based on conventional strategy interaction and storage experience: Based on Actor (actor network), in state s l Select an action a l,2 , interacting with the environment to get rewards r l,2 and the next period status s l+1,2 ; Transition tuples of actor networks ( s l , a l,2 , r l,2 , s l+1,2 ) Deposit into D Agent , priority p l,2 Set as .
[0059] (503) Determine whether to trigger an update: If the current step number l Is an integer multiple of the experience playback interval K. If so, execute: (A) Mini-batch update loop: Perform k mini-batch update operations, that is, for j = 1 ~ k, execute: (a) Sampling experience: According to the probability distribution P defined by formula (15) j , sample transfer data from experience replay; (b) Calculate the importance sampling weight: For the sampled transfer data, based on the probability distribution P j , according to formula (16), calculate the importance sampling weight w j , used to correct the deviation caused by priority sampling; (c) Calculate the target action with noise: At the current time step, use the target actor network to generate the target action, add noise to the target action, and calculate the TD target (target value, including target 1 and target 2) according to the target critic network and formula (7) y l ; In calculation y lWhen , Gaussian noise is added to the action to prevent the strategy from overfitting to extreme actions. This can make the target Q value smoother and enhance the robustness of the strategy: ; in, is the target Actor network parameter, The mean is 0 and the standard deviation is Gaussian distribution, c Take 0.5.
[0060] (d) Update experience priority: Calculate the difference between the TD target and the review network result, and calculate the temporal difference (TD) error , measures the difference between the predicted value and the target value; according to formula (13), based on the calculated TD error, updates the priority of the corresponding experience in the replay buffer p j (include p j,1 and p j,2 ), so that subsequent sampling can focus more on high-value experiences; (e) Calculate the Critic network loss: based on the TD target y l and importance sampling weight w j , according to formula (17), calculate the loss L(θ i ) (including L1 and L2), that is, let the current Q network prediction value approach the target value calculated by the reward, discount factor and target network, and the error is the square expectation; (f) Update the critic network parameters: Use the loss function L(θ i ) gradient, update the parameters of the Critic network (including the first review network, referred to as review 1, the second review network, referred to as review 2) θ i , so that the Critic network can more accurately evaluate the action value; (g) End the mini-batch update cycle: complete k mini-batch updates.
[0061] (B) Determine whether to update the Actor and target network: If the current step l is an integer multiple of the set interval d, then execute: (a) Update Actor network parameters: Update Actor network parameters using deterministic policy gradients , the formula is , using the value gradient of the action by the Critic network to guide the Actor to output a better action; (b) Update target network parameters: Use soft update to update target network parameters. The formula is: ; ; Let the target network slowly follow the current network to improve training stability; is the weight coefficient; (c) End Actor and target network update judgment: Complete the operations related to Actor and target network update.
[0062] (C) End trigger update judgment: complete the update operation.
[0063] (504) Prepare for the next state: Add exploration noise to select the next action and change the current state s l Updated to s l+1 , enter the next round of training cycle.
[0064] (6) When l Reach the total number of training steps T The training cycle ends when .
[0065] After training, the actor network is actually used. The other networks are used during training. In actual application, the input state is used to obtain the device action. That is, the state of the integrated energy system is obtained, and the multi-energy flow device action is obtained through the trained actor network.
[0066] To validate the effectiveness of the proposed method, a simulation was conducted at an IES demonstration park in northern China. The simulation hardware environment used an Intel(R) Core(TM) i9-14900HX CPU and a laptop with 32GB of RAM. Model construction and training were performed using the PyTorch deep learning framework.
[0067] The convergence of the agent during training is as follows Figure 3 As shown in the figure, in the initial stages of training, the algorithm is in the exploration phase. Because the initial action selection is based on a deterministic environment, the agent receives small rewards after making decisions. After approximately 3,000 rounds of training, the agent begins to converge to a stable reward range, indicating that it has learned the optimal scheduling strategy. Due to the multiple uncertainties in the source-load-price environment and the algorithm's noise mechanism, the agent's reward values inevitably fluctuate during training.
[0068] To further highlight the convergence performance of TD3MPC, Figure 4The average reward convergence of different algorithms under 10 random seed runs is shown. The shaded area in the figure represents the maximum and minimum values of the training results, and the dark line represents the average of the 10 smoothed training results. To clarify the characteristics of the TD3 algorithm and the advantages of improved state space adaptability, the improved TD3, which introduces state probability transitions and temporal segment representation, is compared with traditional TD3. Convergence results show that the improved TD3 stabilizes after 9,000 steps, while traditional TD3 only gradually stabilizes after 14,000 steps, and the final convergence reward of the improved TD3 is slightly higher. This is because the improved TD3 effectively avoids unreasonable state transitions by modeling the random transition characteristics of the state space, thereby obtaining more accurate Q-value estimates and ensuring the direction of algorithm updates. Furthermore, compared to traditional single-time state variables, the expanded state space provides richer temporal information for the environment-to-policy mapping. To demonstrate the positive benefits of MPC for DRL, the improved TD3 is further compared with TD3MPC. TD3MPC, which incorporates MPC expert experience guidance, converges rapidly after approximately 4,000 episodes. Compared to the initially fully random exploration model, the introduction of MPC Actors significantly improves the starting point of exploration, significantly increasing initial rewards, avoiding blind trial and error within invalid action domains, and accelerating the rate of cumulative reward growth in the early stages. Later, the agent continuously interacts with the environment through hybrid-priority experience replay, further exploring based on expert experience. This allows the agent to explore towards the optimal solution beyond the sample space with small fluctuations when the reward function approaches a local optimum, significantly improving the late-stage convergence reward.
[0069] Since the test site has abundant wind and solar resources, the load is largely met by renewable energy, resulting in energy supply stability being significantly affected by wind and solar fluctuations. In addition, the uncertainty of electricity prices and multiple loads makes power balancing difficult. To verify the effectiveness of the proposed algorithm, one day in winter, summer, and the transition season was randomly selected as a test set, and the scheduling results were analyzed. Figure 5 As can be seen, the proposed strategy is able to maintain a consistent power balance between supply and demand. Since there is no heat load demand in the summer, both energy purchase and operation and maintenance costs are minimal. The randomly selected transition season test day falls close to winter, so the costs for both are similar.
[0070] From the perspective of equipment output, due to the low cost of cogeneration, most of the electricity load shortfall is met by CHP, with the grid acting as a backup resource when electricity prices are low. The energy storage system operates similarly in winter, summer, and autumn: charging when wind and solar resources are abundant and discharging when electricity prices are high. Notably, the TD3MPC agent does not blindly choose energy storage charging only when renewable energy is abundant. Using real-time information, it comprehensively assesses future trends in electricity and gas prices and multiple loads, and coordinates the P2G and energy storage systems to absorb renewable energy, ensuring that batteries complete energy storage before peak demand, demonstrating strong holistic management. Due to the high cost of P2G, during periods without renewable energy demand pressure, the gas grid fills the gap in gas demand. In autumn and winter, the heat load is primarily met by CHP. Given that its operation is influenced by multiple energy flows, including electricity, gas, and heat, the agent flexibly controls its operation to either a heat-based electricity or electricity-based heat mode, with the remaining heat load shortfall being met by the GB.
[0071] Table 1. Optimization scheduling results of different schemes
[0072] Table 1 shows that TD3MPC outperforms all other models across all metrics, approaching the ideal scenario most closely. Compared to TD3, the Improved TD3 reduces wind and solar curtailment costs, load curtailment costs, and operating costs by 53.93%, 42.86%, and 4.02%, respectively. This suggests that TD3 adopts a more conservative operating strategy, investing more controllable equipment to support the load. However, this increased cost does not improve wind and solar curtailment and load curtailment. This is because unreasonable state transitions are virtually impossible to occur in actual operation, and excessive focus on them leads to meaningless cost increases. Furthermore, the strong temporal correlation of state parameters in the IES violates the Markov principle, resulting in insufficient decision-making reference. Improved TD3, however, addresses these deficiencies through improvements to the state space, significantly enhancing the agent's adaptability to the multiple uncertainties of source, load, and price. Compared to ImprovedTD3, TD3MPC further reduced wind and solar curtailment costs, load curtailment costs, and total costs by 17.28%, 23.74%, and 3.97%, respectively. This demonstrates that the MPC actor can guide ImprovedTD3 to converge to an optimal solution by providing a baseline policy that satisfies physical constraints. In fact, during the preliminary research phase, this study attempted to use RMPC and stochastic MPC as expert strategies to guide DRL. However, these MPC variants that account for environmental uncertainty exhibit significant performance degradation in the multi-dimensional uncertainty scenarios of the IES, rendering them incapable of providing a stable and accurate starting point for exploration. In contrast, deterministic MPC is more efficient due to its stability. Without post-adjustment by ImprovedTD3, MPC alone would struggle to achieve a reliable operating strategy. Although its operating cost is lower, its short-sightedness and environmental noise make it impossible to avoid the risks of multiple uncertainties, resulting in extremely high penalties for wind and solar curtailment and load curtailment. In summary, the TD3MPC framework, which integrates ImprovedTD3 and MPC, performs better in the IES environment with coupled uncertainties.
[0073] RMPC and RO, based on worst-case scenarios, have lower wind, solar, and load curtailment, similar to when facing local uncertainty. However, in the context of multiple uncertainties in source, load, and price, the need to model multiple worst-case scenarios separately increases costs, making them unacceptable. Due to the difficulty in characterizing the joint probability distribution of multiple random variables, RO and CCP, despite incurring high costs, still cannot achieve reliable operating strategies. Compared with traditional optimization methods, DRL, leveraging the powerful representation capabilities of deep learning, is more adaptable to multiple uncertain environments, and TD3MPC further enhances this feature.
[0074] In this example, MPC and DRL demonstrate significant complementary properties in IES scheduling. The TD3MPC framework inherits DRL's ability to learn from the environment and handle uncertainty, as well as MPC's ability to leverage physical models to provide fundamental performance. Compared to traditional TD3, TD3MPC offers a simpler reward function design, faster convergence, superior results, and significantly reduced parameter tuning difficulty.
[0075] In this embodiment, the hybrid priority experience replay mechanism proposed based on the priority experience replay mechanism achieves a balance between the efficiency of expert experience utilization and exploration diversity through a dual-weight priority design. It is more suitable for TD3MPC training, shortening the training time by 20.87% and achieving better results.
[0076] In this example, the strong temporal correlation of IES state variables invalidates the Markov decision assumption of traditional DRL. Replacing single-moment random variables with a sequence of state variables with an ACF greater than 0.7 in the extended state space effectively eliminates pseudo-Markov interference. Furthermore, explicitly modeling the transition characteristics of the extended state space in conjunction with probabilistic scenario generation significantly improves the agent's ability to estimate uncertainty.
[0077] Example 2 This embodiment provides a comprehensive energy system optimization system under multiple uncertain environments, which specifically includes: A data acquisition module is configured to: acquire the status of the integrated energy system; The optimization module is configured to: obtain the multi-energy flow device action through the actor network based on the state; Among them, during the training process of the actor network, the MPC actor is embedded in the dual-delay deep deterministic policy gradient architecture, and the transfer tuples of the actor network and the MPC actor are stored in the agent experience replay pool and the expert experience replay pool respectively. According to the temporal differential error of the samples in the two experience replay pools, each sample is given a priority, and the mixing ratio is determined according to the training time step. Combining the mixing ratio and priority, the sampling probability of the sample is calculated, and then the update of the comment network is controlled to have a high dependence on the expert experience replay pool in the early stage of training and a high dependence on the agent experience replay pool in the later stage of training. The value gradient of the comment network to the action is used to guide the actor network to optimize the output action.
[0078] Furthermore, the status includes wind power generation power, photovoltaic power generation power, thermal load, electric load, gas load, electricity price, output power of cogeneration device, input power of power-to-gas and charge state of energy storage within the time segment composed of the current time period as the midpoint.
[0079] Furthermore, the multi-energy flow device actions include changes in charging power, changes in discharging power, changes in purchased electricity power, changes in gas input power of the cogeneration device, changes in power-to-gas input power, changes in gas boiler input power and changes in gas purchased power.
[0080] It should be noted here that the various modules in this embodiment correspond one-to-one to the various steps in Example 1, and the specific implementation processes are the same, which will not be repeated here.
[0081] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A comprehensive energy system optimization method under multiple uncertain environments, characterized by: include: Obtain the status of the integrated energy system; Based on the state, the multi-energy flow device action is obtained through the actor network; Among them, during the training process of the actor network, the MPC actor is embedded in the dual-delay deep deterministic policy gradient architecture, and the transfer tuples of the actor network and the MPC actor are stored in the agent experience replay pool and the expert experience replay pool respectively. According to the temporal differential error of the samples in the two experience replay pools, each sample is given a priority, and the mixing ratio is determined according to the training time step. Combining the mixing ratio and priority, the sampling probability of the sample is calculated, and then the update of the comment network is controlled to have a high dependence on the expert experience replay pool in the early stage of training and a high dependence on the agent experience replay pool in the later stage of training. The value gradient of the comment network to the action is used to guide the actor network to optimize the output action.
2. The method for optimizing an integrated energy system under multiple uncertain environments according to claim 1, wherein: The status includes wind power generation power, photovoltaic power generation power, thermal load, electric load, gas load, electricity price, output power of cogeneration device, input power of power-to-gas and charge state of energy storage within the time segment composed of the current time period as the midpoint.
3. The method for optimizing an integrated energy system under multiple uncertain environments according to claim 1, characterized in that: The multi-energy flow equipment actions include changes in charging power, changes in discharging power, changes in purchased electricity power, changes in gas input power of the cogeneration device, changes in power-to-gas input power, changes in gas boiler input power and changes in gas purchased power.
4. The method for optimizing an integrated energy system under multiple uncertain environments according to claim 1, wherein: During the training of the actor network, for each time step, the following steps are performed: The MPC actor selects an action in the current time step state and interacts with the environment to obtain rewards and the next time step state. The MPC actor's transfer tuple is stored in the expert experience replay pool and the priority of each sample in the expert experience replay pool is set. The actor network selects an action in the current time step state and interacts with the environment to obtain rewards and the next time step state. The actor network's transfer tuple is stored in the agent experience replay pool and the priority of each sample in the agent experience replay pool is set. If the current time step is an integer multiple of the experience replay interval, then based on the expert experience replay pool and the agent experience replay pool, combined with the target network, multiple small batch update operations are performed on the comment network, and the priority of the samples in the experience replay pool is updated; If the current time step is an integer multiple of the set interval, the actor network and target network are updated.
5. The method for optimizing an integrated energy system under multiple uncertain environments according to claim 4, characterized in that: The reward is the inverse of the total energy system operating cost, wind and solar power curtailment penalty items, and equipment overlimit penalty items.
6. The method for optimizing an integrated energy system under multiple uncertain environments according to claim 4, characterized in that: The small batch update operation includes: According to the sampling probability of the sample, the transfer data is sampled from the experience replay pool; For the sampled transfer data, calculate the importance sampling weight based on the sampling probability; At the current time step, the target actor network is used to generate the target action, Gaussian noise is added to the target action, and the temporal difference target is calculated based on the target critic network. The difference between the temporal difference target and the critic network result is calculated to obtain the temporal difference error. Update the priority of samples in the experience replay pool based on the temporal difference error; Based on the temporal difference error and importance sampling weights, the loss of the review network is calculated and the review network parameters are updated.
7. The method for optimizing an integrated energy system under multiple uncertain environments according to claim 1, wherein: The sampling probability is: ; Among them, P j is the sampling probability of the jth sample; linear weight attenuation coefficient , and are the initial and final attenuation coefficients, T is the total number of training steps, l Indicates the current time step; type j = MPC represents that the jth sample comes from the expert experience replay pool, type j = Agen t indicates that the jth sample comes from the agent experience replay pool; D agent and D mpc are the number of samples in the agent experience replay pool and the expert experience replay pool respectively; β is the priority intensity coefficient; p j is the priority of the j-th sample, p z and p d Represent the priorities of samples from the agent experience replay pool and the expert experience replay pool respectively.
8. A comprehensive energy system optimization system under multiple uncertain environments, characterized by: include: A data acquisition module is configured to: acquire the status of the integrated energy system; The optimization module is configured to: obtain the multi-energy flow device action through the actor network based on the state; Among them, during the training process of the actor network, the MPC actor is embedded in the dual-delay deep deterministic policy gradient architecture, and the transfer tuples of the actor network and the MPC actor are stored in the agent experience replay pool and the expert experience replay pool respectively. According to the temporal differential error of the samples in the two experience replay pools, each sample is given a priority, and the mixing ratio is determined according to the training time step. Combining the mixing ratio and priority, the sampling probability of the sample is calculated, and then the update of the comment network is controlled to have a high dependence on the expert experience replay pool in the early stage of training and a high dependence on the agent experience replay pool in the later stage of training. The value gradient of the comment network to the action is used to guide the actor network to optimize the output action.
9. The integrated energy system optimization system under multiple uncertain environments according to claim 8, characterized in that: The status includes wind power generation power, photovoltaic power generation power, thermal load, electric load, gas load, electricity price, output power of cogeneration device, input power of power-to-gas and charge state of energy storage within the time segment composed of the current time period as the midpoint.
10. The integrated energy system optimization system under multiple uncertain environments according to claim 8, characterized in that: The multi-energy flow equipment actions include changes in charging power, changes in discharging power, changes in purchased electricity power, changes in gas input power of the cogeneration device, changes in power-to-gas input power, changes in gas boiler input power and changes in gas purchased power.
Citation Information
Patent Citations
Integrated energy system energy optimization scheduling method and system based on reinforcement learning
CN115186885A
Comprehensive energy system model prediction control method based on deep reinforcement learning
CN119443674A
Multi-agent deep reinforcement learning method based on human guidance
CN119990246A
Day-ahead and intra-day economic dispatching method for wind and light storage system
CN120150103A
Dymamic control of a manufacturing process using deep reinforcement learning
US20220373980A1
Cited By
TD3-based energy system prediction-optimization closed-loop decision method and system
CN122288322A