Multi-scenario Optimization Method and System for Integrated Energy Systems Based on TPE-SAC
By optimizing the equipment configuration of the integrated energy system based on the TPE-SAC method, and utilizing configuration transfer learning and the SAC algorithm, the shortcomings of robust optimization and deep reinforcement learning are overcome, and real-time balance of cooling, electricity and heat loads is achieved, as well as a reduction in total cost.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies in integrated energy systems suffer from several drawbacks: overly conservative robust optimization methods leading to poor economic efficiency; lack of theoretical basis for constructing uncertain sets; difficulty in handling multi-source uncertainty correlations; low training efficiency of deep reinforcement learning and difficulty in balancing exploration and utilization, resulting in load imbalance and high costs; and difficulty in achieving real-time balance of cooling, electricity, and heating loads.
The TPE-SAC-based approach optimizes equipment configuration through the TPE algorithm, sets a similarity score threshold for configuration transfer learning, optimizes the agent using the SAC algorithm, and combines performance evaluation and iterative updates to ensure real-time balance of cooling, electricity, and heating loads.
It achieves real-time balance of cooling, electricity, and heating loads in multiple scenarios, avoids the cost penalty of supply and demand imbalance, and significantly reduces the total cost of the integrated energy system.
Smart Images

Figure CN121119653B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated energy system optimization technology, specifically to a multi-scenario optimization method and system for integrated energy systems based on TPE-SAC. Background Technology
[0002] With societal development and people's increasing pursuit of comfort, building energy consumption is gradually rising, and the construction industry accounts for 30% of global final energy consumption. Integrated Energy Systems (IES), as a highly efficient energy system, directly provide users with multiple energy loads, reduce energy loss, help reduce environmental pollution, and achieve cascaded energy utilization.
[0003] Existing technologies have included research on scheduling optimization for integrated energy systems facing uncertainties. Robust optimization, as a classic method for handling uncertainties, is widely used in IES operation optimization. Existing technologies have also established distributed robust optimization models. For integrated energy systems including electric-hydrogen hybrid energy storage, the proposed TSDRO-based coordinated scheduling model can better balance conservatism and robustness compared with traditional methods. However, robust optimization methods have the following limitations: (1) they are too conservative, which may lead to poor economic efficiency; (2) the construction of the uncertainty set depends on experience and lacks theoretical basis; (3) they are difficult to handle the correlation of multi-source uncertainties.
[0004] Deep reinforcement learning (DRL) is currently being widely researched as a control method for integrated energy systems facing uncertainty. However, DRL methods face significant challenges in training efficiency and cost control. Existing techniques for applying DRL to real-time energy management of integrated electric-thermal energy systems have found that in the early stages of training, due to the agent's limited environmental awareness and exploratory phase, the system incurs significant load imbalance penalties, resulting in low reward function values. Furthermore, the initial model performance is poor, requiring extensive training to reach optimal performance. This reflects the fundamental problem of achieving an effective balance between exploration and utilization in deep reinforcement learning models. Specifically, over-exploration leads to lengthy training times and soaring costs, while insufficient exploration easily traps the model in local optima, preventing it from achieving globally optimal performance.
[0005] A significant advantage of Integrated Energy Systems (IES) is their ability to achieve combined energy supply through the use of multiple advanced technologies. However, achieving combined energy supply using various energy devices and technologies is also a challenging task, as the coupling between different energy forms and devices poses challenges to design and operation. Furthermore, as a relatively independent energy system, an IES needs to ensure energy supply and demand balance within its service area. However, in actual operation, under uncertain scenarios, the supply and demand balance on both the source and load sides of an IES is uncertain. In such situations, simultaneously ensuring real-time load balance for cooling, electricity, and heating becomes a significant challenge. Summary of the Invention
[0006] This invention is made to solve the above problems, and aims to provide a multi-scenario optimization method and system for integrated energy systems based on TPE-SAC.
[0007] This invention provides a multi-scenario optimization method for integrated energy systems based on TPE-SAC, characterized by the following steps: Step S1, optimizing the configuration parameters of equipment in the integrated energy system using the TPE algorithm model to generate a configuration scheme; Step S2, setting a similarity score threshold, generating a similarity score based on the configuration scheme and historical configuration schemes, and when the similarity score is greater than the similarity score threshold, using the historical agent corresponding to the historical configuration scheme as the loaded agent; Step S3, optimizing the loaded agent using the SAC algorithm model to generate an optimized agent; Step S4, evaluating the performance of the optimized agent, obtaining the performance evaluation results, and storing the configuration scheme, optimized agent, and performance evaluation results; Step S5, feeding back the configuration scheme, optimized agent, and performance evaluation results to the TPE algorithm model to update the TPE algorithm model.
[0008] The TPE-SAC-based multi-scenario optimization method for integrated energy systems provided by this invention may also have the following features: Step S1 includes the following sub-steps: Step S10, dividing historical observations into two subsets according to a preset threshold, namely a high-quality set and a low-quality set, establishing conditional probability density models for the high-quality set and the low-quality set respectively, and obtaining the probability density function of the high-quality set and the probability density function of the low-quality set; Step S11, using the expected improvement as the acquisition function, and using Bayes' rule and the law of total probability to express the expected improvement as the ratio of the probability density function of the high-quality set to the probability density function of the low-quality set; Step S12, selecting sampling points in the acquisition function by maximizing the expected improvement, and generating a configuration scheme based on the sampling points.
[0009] The TPE-SAC-based multi-scenario optimization method for integrated energy systems provided by this invention may also have the following features: Step S2 includes the following sub-steps: Step S20, performing Z-score standardization on all historical configuration scheme parameters to obtain standardized historical configuration scheme parameters; Step S21, calculating the Euclidean distance between the configuration scheme parameters corresponding to the configuration scheme and the standardized historical configuration scheme parameters, and obtaining a similarity score by combining the historical performance evaluation results; Step S22, setting a similarity score threshold, and starting the configuration transfer learning mechanism when the similarity score is greater than the similarity score threshold, using the historical agent as the loading agent.
[0010] The TPE-SAC-based multi-scenario optimization method for integrated energy systems provided by this invention may also have the following features: Step S3 includes the following sub-steps: Step S30, defining a quintuple for the loaded agent using an extended state-space Markov decision process, and establishing an objective function based on the quintuple; Step S31, introducing a policy entropy regularization term into the objective function using the SAC algorithm to obtain a first optimization objective function; Step S32, designing the action space and state space of the first optimization objective function according to the DRL environment, and normalizing the state space to obtain a second optimization objective function; Step S33, limiting the action space according to load demand; Step S34, designing the reward function of the second optimization objective function according to the optimization requirements to obtain a third optimization objective function, and obtaining an optimization agent based on the third optimization objective function.
[0011] The multi-scenario optimization method for integrated energy systems based on TPE-SAC provided by this invention may also have the following features: Step S4 includes the following sub-steps: Step S40, based on historical operating data, a Monte Carlo sampling method is used to generate an evaluation scenario set S, which includes 100 scenarios; Step S41, the relative deviation between the peak value of the cooling and heating load and the total value of the evaluation scenario set S is calculated, and an optimization agent is used to perform year-round simulation operation on the evaluation scenario set S to obtain performance evaluation results, which include the total cost; Step S42, the configuration scheme, the optimization agent, and the performance evaluation results are stored.
[0012] The TPE-SAC-based integrated energy system multi-scenario optimization method provided by this invention may also include the following features:
[0013] Step S0: Set the optimization objective for the equipment in the integrated energy system, namely, minimize the total annual cost through the capacity configuration of the ice storage system, dual-condition refrigeration host, base load refrigeration host, thermal energy storage system, and heat pump system. Model the equipment and determine the constraints of the equipment.
[0014] This invention provides a multi-scenario optimization system for integrated energy systems based on TPE-SAC, characterized by the following features: a TPE algorithm optimization module, used to optimize the configuration parameters of equipment in the integrated energy system using a TPE algorithm model to generate configuration schemes; a configuration migration module, used to set a similarity score threshold, generate a similarity score based on the configuration scheme and historical configuration schemes, and when the similarity score is greater than the similarity score threshold, use the historical agent corresponding to the historical configuration scheme as the loaded agent; a SAC algorithm optimization module, used to optimize the loaded agent using a SAC algorithm model to generate an optimized agent; a performance evaluation module, used to evaluate the performance of the optimized agent, obtain the performance evaluation results, and store the configuration scheme, optimized agent, and performance evaluation results; and an iteration module, used to feed back the configuration scheme, optimized agent, and performance evaluation results to the TPE algorithm model to update the TPE algorithm model.
[0015] The technical effects of this invention are as follows:
[0016] The TPE-SAC-based multi-scenario optimization method and system for integrated energy systems of the present invention includes the following steps: Step S1, optimizing the configuration parameters of equipment in the integrated energy system using the TPE algorithm model to generate a configuration scheme; Step S2, setting a similarity score threshold, generating a similarity score based on the configuration scheme and historical configuration schemes, and when the similarity score is greater than the similarity score threshold, using the historical agent corresponding to the historical configuration scheme as the loaded agent; Step S3, optimizing the loaded agent using the SAC algorithm model to generate an optimized agent; Step S4, performing performance evaluation on the optimized agent, obtaining the performance evaluation results, and storing the configuration scheme, optimized agent, and performance evaluation results; Step S5, feeding back the configuration scheme, optimized agent, and performance evaluation results to the TPE algorithm model to update the TPE algorithm model. Therefore, the TPE-SAC-based multi-scenario optimization method and system for integrated energy systems of the present invention, by combining the TPE-SAC algorithm, ensures the real-time balance of cooling, electricity, and heat loads of the integrated energy system in multiple scenarios, avoids the penalty cost caused by supply and demand imbalance, and significantly reduces the total cost of the integrated energy system. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the multi-scenario optimization method for integrated energy systems based on TPE-SAC in an embodiment of the present invention.
[0018] Figure 2 This is a schematic diagram of the integrated energy system in an embodiment of the present invention.
[0019] Figure 3 This is an overall framework diagram of the algorithm in the multi-scenario optimization method for integrated energy systems based on TPE-SAC in the embodiments of the present invention.
[0020] Figure 4 This is a flowchart illustrating the TPE algorithm in an embodiment of the present invention.
[0021] Figure 5 This is a diagram of peak hot and cold loads in a scene generated in an embodiment of the present invention.
[0022] Figure 6 This is an embodiment of the present invention showing the annual heating and cooling load diagram.
[0023] Figure 7 This is an embodiment of the present invention showing the annual electrical load diagram.
[0024] Figure 8 This is a diagram showing the annual photovoltaic power generation in an embodiment of the present invention.
[0025] Figure 9 This is an embodiment of the present invention showing the annual, monthly, and time-of-use electricity price chart.
[0026] Figure 10 This is a comparison chart of the training results of the Load method and the Dir method in an embodiment of the present invention.
[0027] Figure 11 This is a flowchart illustrating the process of directly training the agent 20 times in an embodiment of the present invention.
[0028] Figure 12 This is a flowchart illustrating the preloaded intelligent agent training process for 300 iterations in an embodiment of the present invention.
[0029] Figure 13 This is a flowchart illustrating the preloaded intelligent agent training process 20 times in an embodiment of the present invention.
[0030] Figure 14 This is an iterative optimization diagram of the annual total cost of TPE-SAC in an embodiment of the present invention.
[0031] Figure 15 This is a schematic diagram of the cold penalty in all scenarios for the three configuration schemes in the embodiments of the present invention.
[0032] Figure 16 This is a schematic diagram of the heat penalty for three configuration schemes in all scenarios in the embodiments of the present invention. Detailed Implementation
[0033] To make the technical means, creative features, objectives and effects of this invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, specifically illustrate the multi-scenario optimization method and system of the integrated energy system based on TPE-SAC of this invention.
[0034] Figure 1This is a flowchart illustrating the multi-scenario optimization method for integrated energy systems based on TPE-SAC in an embodiment of the present invention.
[0035] Figure 2 This is a schematic diagram of the integrated energy system in an embodiment of the present invention.
[0036] Figure 3 This is an overall framework diagram of the algorithm in the multi-scenario optimization method for integrated energy systems based on TPE-SAC in the embodiments of the present invention.
[0037] Figure 4 This is a flowchart illustrating the TPE algorithm in an embodiment of the present invention.
[0038] like Figure 1 , Figure 2 , Figure 3 and Figure 4 As shown, this embodiment provides a multi-scenario optimization method for integrated energy systems based on TPE-SAC, including:
[0039] Step S0: Set the optimization objective for the equipment in the integrated energy system, namely, to minimize the total annual cost through the capacity configuration of the ice storage system (IS), dual-condition refrigeration unit (DTC), base load refrigeration unit (PC), thermal energy storage system (HS), and heat pump system (HP), and model the equipment and determine the constraints of the equipment.
[0040] The optimization objective is first defined as follows:
[0041]
[0042] In the formula Indicates in the scene Total annual cost This represents the annualized equipment investment cost. Indicates in the scene Annual operating costs This is the mathematical expectation symbol.
[0043] The annualized equipment investment cost can be expressed by the following formula:
[0044]
[0045] In the formula, Indicates the first in the energy system The capacity of each device A collection of equipment in an energy system. , For the first in the system The linear investment cost per unit capacity of each piece of equipment.
[0046] For the first in the system The formula for calculating the average annual investment coefficient of each piece of equipment is as follows:
[0047]
[0048] In the formula For bank interest rates, For the first in the system The lifespan of each device.
[0049] To represent a certain scene The annual operating cost is calculated using the following formula:
[0050]
[0051] In the formula Indicates the cost of purchasing electricity. The cost of penalties for imbalance between hot and cold temperatures.
[0052] The following constraints are set for the ice storage system (IS):
[0053]
[0054]
[0055]
[0056] in, and They represent the time steps respectively. and The state of charge of the time-limited ice storage system and These represent the amount of electricity the ice storage system charges and discharges in one hour, respectively. and These represent the minimum and maximum ice storage capacities of the ice storage system, respectively. This indicates the real-time ice storage status of the ice storage system. Indicates the ice storage system at any time The charging / discharging power, with positive values indicating charging and negative values indicating discharging. This indicates the capacity of the ice storage system, and the coefficient 0.3 indicates that the maximum charging and discharging power of the ice storage system is 30% of its rated capacity.
[0057] The operation process of a dual-condition chiller (DTC) is as follows:
[0058]
[0059]
[0060]
[0061] This refers to the electrical power of the dual-mode cooling unit. The coefficient of performance (COP) of the dual-condition refrigeration unit. This refers to the cooling capacity of the dual-mode refrigeration unit. The coefficient of performance (COP) of the dual-condition refrigeration unit. This refers to the ice-making capacity of the dual-mode refrigeration unit. The rated capacity of the dual-mode refrigeration unit.
[0062] The operation of the base-mounted cooling unit (PC) is as follows:
[0063]
[0064]
[0065] This refers to the electrical power of the base-load refrigeration unit during operation. The coefficient of performance (COP) of the base-load refrigeration unit. The cooling capacity of the base-load refrigeration unit. The rated capacity of the base-load refrigeration unit.
[0066] The constraint equations for the thermal energy storage system (HS) are as follows:
[0067]
[0068]
[0069]
[0070] in and They represent the time steps respectively. and The state of charge of the thermal energy storage system and These represent the amount of electricity the thermal energy storage system charges and discharges in one hour, respectively. This indicates the real-time thermal storage status of the thermal energy storage system. and These represent the minimum and maximum values of the thermal energy storage system, respectively. , Indicates the thermal energy storage system at time The heat transfer power is represented by a positive value indicating heat transfer and a negative value indicating heat release. This indicates the capacity of the thermal energy storage system. The maximum charge / discharge power of the thermal energy storage system is limited to 30% of its capacity.
[0071] The operation of a heat pump system (HP) is as follows:
[0072]
[0073]
[0074] This refers to the electrical power of the heat pump system during operation. The coefficient of performance (COP) of a heat pump system. To generate heat, This refers to the rated capacity of the heat pump.
[0075] The specific equation for energy balance is as follows:
[0076]
[0077]
[0078]
[0079]
[0080]
[0081] in, To purchase electricity from the power grid, For photovoltaic power generation, To improve the energy efficiency of the ice storage system, For the energy to be injected into the ice storage system, For the user-side cooling load, For the energy release efficiency of the ice storage system, The energy consumed by the ice storage system. To charge the thermal energy storage system, To improve the charging efficiency of thermal energy storage systems, For the user-side heat load, The energy release efficiency of the thermal energy storage system, The energy consumed by the thermal energy storage system.
[0082] Step S1: Optimize the configuration parameters of equipment in the integrated energy system using the TPE algorithm model to generate a configuration scheme.
[0083] Step S1 specifically includes the following sub-steps:
[0084] Step S10: Divide the historical observations into two subsets according to a preset threshold, namely the high-quality set and the low-quality set. Establish conditional probability density models for the high-quality set and the low-quality set respectively, and obtain the probability density function of the high-quality set and the probability density function of the low-quality set.
[0085] Typically, a percentile of all historical observations is selected, such as the top 20%.
[0086]
[0087]
[0088] in, For high-quality collection, It is a low-quality collection. Indicates the classification threshold. This represents the i-th input sample point. Represents sample points The corresponding objective function value.
[0089] Establish conditional probability density models for high-quality and low-quality sets:
[0090]
[0091]
[0092] in, This is shown when the objective function value is less than the classification threshold. Under the condition, input variables The conditional probability density function, This indicates that when the objective function value is greater than or equal to the classification threshold... Under the condition, input variables The conditional probability density function, and Let represent the probability density functions of high-quality samples and low-quality samples in the input space, respectively, from the high-quality set and the low-quality set.
[0093] Step S11: The expected improvement is taken as the acquisition function, and Bayes' rule and the law of total probability are used to express the expected improvement as the ratio of the probability density function of the good set to the probability density function of the bad set.
[0094] set up According to Bayes' theorem:
[0095]
[0096]
[0097] in, , Denotes the prior probability of a high-quality set , Denotes the prior probability of an inferior set , Indicates the given input When the sample belongs to the high-quality set, the posterior probability is... Indicates the given input At that time, the posterior probability that the sample belongs to the inferior set is assumed to be the optimal value after n function evaluations. Then at the candidate point Improvement amount Defined as:
[0098]
[0099] In the formula, This represents the function that takes the maximum value.
[0100] The acquisition function is defined as the improved mathematical expectation:
[0101]
[0102] In the formula, This indicates a desire to improve the acquisition function. This indicates that an existing observation dataset already exists.
[0103] Through mathematical derivation, the desired improvement can be expressed as:
[0104]
[0105] in, For the probability density function of a high-quality set, Let be the probability density function of the inferior set. It indicates a direct proportional relationship.
[0106] Step S12: Select sampling points in the acquisition function by maximizing the expected improvement, and generate a configuration scheme based on the sampling points:
[0107]
[0108] in This indicates the sampling point to be searched next.
[0109] In this embodiment, the TPE algorithm tends to select candidate points that have high probability density in high-quality samples but low probability density in low-quality samples.
[0110] Step S2: Set a similarity score threshold. Generate a similarity score based on the configuration scheme and the historical configuration scheme. When the similarity score is greater than the similarity score threshold, use the historical agent corresponding to the historical configuration scheme as the loaded agent.
[0111] Step S2 includes the following sub-steps:
[0112] Step S20: Perform Z-score standardization on all historical configuration parameters to obtain standardized historical configuration parameters.
[0113]
[0114] in, This represents the standardized configuration scheme parameter vector. This represents the parameter vector of the original configuration scheme. and These are the mean vector and standard deviation vector for each parameter dimension, respectively.
[0115] Step S21: Calculate the Euclidean distance between the configuration scheme parameters corresponding to the configuration scheme and the standardized historical configuration scheme parameters, and obtain the similarity score by combining the historical performance evaluation results.
[0116] Calculate configuration scheme parameters Parameters of historical configuration schemes Euclidean distance between :
[0117]
[0118] This represents a vector of parameters from the current configuration scheme after standardization. This represents the vector of parameters for the i-th historical configuration scheme after standardization. This represents the 2-norm (Euclidean norm).
[0119] The Euclidean distance is converted into a similarity score, and a comprehensive score is generated by combining historical performance.
[0120]
[0121]
[0122] in, This represents the similarity score between the current configuration and the i-th historical configuration. For similarity adjustment parameters, This represents the comprehensive score calculated based on the i-th historical configuration scheme. For the normalized historical cost, This is the similarity weighting coefficient. This is the historical cost weighting coefficient.
[0123] Step 22: Set a similarity score threshold. When the similarity score exceeds the threshold, activate the configuration transfer learning mechanism and use the historical agent as the loaded agent.
[0124]
[0125] in, The learning rate for reinforcement learning when loading the agent. The learning rate for reinforcement learning when the agent is not loaded. Configure a transfer learning factor to reduce the learning rate of the Actor and Critic networks to 1% of their original values, ensuring that the loaded agents can adapt smoothly to the new configuration.
[0126] Step S3: Optimize the loaded agent using the SAC algorithm model to generate an optimized agent.
[0127] Step S3 includes the following sub-steps:
[0128] Step S30: Define a quintuple for the loaded agent using an extended state-space Markov decision process, establish an objective function based on the quintuple, and define the extended state space as follows:
[0129]
[0130] in, This represents an extended state space. This represents the original environmental state space. Extend the state space to the parameter space. It also includes environmental status. and parameter information .
[0131] Step S31: Introduce the policy entropy regularization term into the objective function using the SAC algorithm to obtain the first optimization objective function:
[0132]
[0133] in, Maximum entropy represents the objective function of reinforcement learning, used to evaluate the policy. The advantages and disadvantages, Indicating in strategy Induced state-action distribution The expected value of the following mathematical expression They are time points State space, action space, and reward function. As a discount factor, For policy entropy, It is a temperature parameter used to balance exploration and reward.
[0134] Step S32: Design the action space and state space of the first optimization objective function based on the DRL environment, and normalize the state space to obtain the second optimization objective function.
[0135] Within the extended state-space framework, the state space Not only includes traditional environmental conditions It also incorporates parameter information. :
[0136]
[0137] Among them, environmental status Include:
[0138]
[0139] In the formula, Indicates time The electrical load demand; Indicates time cooling load demand, Indicates time Heat load demand, Indicates time Electricity price, Indicates time Photovoltaic power generation capacity Indicates time The state of charge of the ice storage system Indicates time State of charge of a thermal energy storage system.
[0140] Parameter status Includes key parameters that affect system behavior:
[0141]
[0142] In the formula, Indicates time The parameter state vector, This indicates the capacity configuration of the ice storage system. This indicates the capacity configuration of the thermal energy storage system. This indicates the capacity configuration of the dual-mode cooling unit. This indicates the capacity configuration of the base-mounted chiller. This indicates the capacity configuration of the heat pump system.
[0143] Normalize the state space:
[0144]
[0145] in, This represents the normalized extended state vector. This represents the minimum environmental condition. This represents the maximum value of the environmental state. The minimum value of the parameter information. This represents the maximum value of the parameter information.
[0146] Step S33: Limit the operating space according to load requirements. :
[0147]
[0148] in For ice storage systems Change value, For thermal energy storage systems The change value is loaded, and the action output by the agent is restricted to [-1, 1] by the tanh function.
[0149] To satisfy the boundary constraints of the actions, the generated actions are processed using the following method.
[0150]
[0151]
[0152] In the formula The original action values are the output of the neural network. This is the action scaling factor. These are the intermediate action values after tanh normalization.
[0153] The integrated energy system automatically switches between cooling and heating modes based on load demand. When the cooling mode is on, the maximum cooling output of the ice storage system is limited to not exceed the cooling load at any given moment. When the cooling mode is off, the ice storage system is restricted to stop operating. The same applies to the heating mode.
[0154] In cooling mode:
[0155]
[0156]
[0157] In the formula, This indicates that the cooling mode is on. This indicates that the cooling mode is off. for Always in cooling mode status. express The cooling mode status at the moment before the current time. express The cooling load requirement at any time, for The number of consecutive days without cooling load, for The number of consecutive days without cooling load, the day before yesterday. for The system enters the corresponding mode when it detects the first cooling or heating load of the day. If no corresponding load is detected for three consecutive days, it exits the corresponding mode.
[0158] To better maintain the system's thermal balance during intelligent training, the following constraints are imposed on the motion space: For example:
[0159]
[0160]
[0161] Indicates the agent at time... The output of the original action value, This represents the truncation function. Indicates time The maximum allowable cooling capacity of the ice storage system is limited to the maximum cooling load at any given time when the cooling mode is on. When the cooling mode is off, the ice storage system is prevented from operating. The same applies to the heating mode.
[0162] Step S34: Design the reward function of the second optimization objective function according to the optimization requirements to obtain the third optimization objective function, and obtain the optimized agent based on the third optimization objective function.
[0163] The reward function is defined as:
[0164]
[0165] in, Indicates time Instant reward value, For electricity purchase costs, As a punishment for energy balance, for Restraint and punishment Penalties for energy storage charging and discharging. Basic punishment, Scaling factor As a weighting factor for electricity purchase costs, For energy balance weights, for Constraint weights Weighting for energy storage operation.
[0166] The formula for calculating the cost of electricity purchase is as follows:
[0167]
[0168] In the formula To account for the remaining electricity demand after deducting energy storage supply and photovoltaic power generation, This refers to the real-time electricity price.
[0169] The formula for calculating the energy balance penalty is:
[0170]
[0171] In the formula, As a thermal balance penalty, This is a penalty for cold balance.
[0172] The thermal equilibrium penalty is defined as:
[0173]
[0174] In the formula The heat load supply and demand gap, This represents the corresponding penalty coefficient. The calculation formula and same.
[0175] The formula for calculating constraint penalties is:
[0176]
[0177] In the formula For ice storage systems Restraint and punishment For thermal energy storage systems Restraint and punishment.
[0178] The formula for calculating the energy storage charge / discharge penalty is:
[0179]
[0180] in and These are the charging and discharging power of the ice storage system, respectively. and These represent the charging and discharging power of the thermal energy storage system, respectively.
[0181] The weight coefficients of each penalty term are determined based on system characteristics and optimization objectives. The weighting of electricity purchase costs reflects the importance of economic objectives. Energy balance weighting ensures supply and demand matching. for Constraining weights to maintain the safe operation of energy storage systems. To optimize the charging and discharging strategy for energy storage operation weights.
[0182] In addition, step S3 also includes: monitoring training convergence using a sliding window coefficient of variation, defining the window size as... =30, when there are consecutive P=10 windows All values are less than the threshold. When convergence is achieved:
[0183]
[0184] in, Indicates the end time The calculated coefficient of variation, Indicates recent The reward sequence for each time step. This represents the standard deviation of the data (Reward) within the window. This represents the mean of the data (Reward) within the window.
[0185] For configuration-based transfer learning, the Hyperband pruning mechanism of Optuna (a hyperparameter optimization framework) is enabled after training a certain number of episodes. If the experiment's performance is significantly worse than other experiments conducted concurrently, it is terminated early and the penalty cost is returned.
[0186]
[0187] in, This is a pruning penalty factor (>1), used to impose additional penalties on poor configurations. This represents the total cost after the penalty. This indicates the actual cost already incurred.
[0188] Figure 5 This is a diagram of peak hot and cold loads in a scene generated in an embodiment of the present invention.
[0189] like Figure 5 As shown, in step S4, the performance of the optimized agent is evaluated, the performance evaluation results are obtained, and the configuration scheme, the optimized agent, and the performance evaluation results are stored.
[0190] Step S4 includes the following sub-steps:
[0191] Step S40: Based on historical operation data, use the Monte Carlo sampling method to generate an evaluation scenario set S, which includes 100 scenarios.
[0192] Step S41: Calculate the relative deviation between the peak value of the cooling and heating load and the total value of the evaluation scenario set S, and use the optimization agent to run the evaluation scenario set S for a whole year to obtain the performance evaluation results, which include the total cost.
[0193] Step S42: Store the configuration scheme, optimize the agent, and evaluate the performance results.
[0194] Step S5: Feed the configuration scheme, optimized agent and performance evaluation results back to the TPE algorithm model to update the TPE algorithm model.
[0195] Figure 6 This is an embodiment of the present invention showing the annual heating and cooling load diagram.
[0196] Figure 7 This is an embodiment of the present invention showing the annual electrical load diagram.
[0197] Figure 8 This is a diagram showing the annual photovoltaic power generation in an embodiment of the present invention.
[0198] Figure 9 This is an embodiment of the present invention showing the annual, monthly, and time-of-use electricity price chart.
[0199] like Figure 6 , Figure 7 , Figure 8 and Figure 9 As shown, to verify the multi-scenario optimization method for integrated energy systems based on TPE-SAC, the following method is used: Figure 2 The integrated energy system shown is used as the subject of this study. The load data used is obtained from a simulation of an office building using EnergyPlus. The electricity price data comes from the time-of-use electricity price for 12 months of a certain year in Zhejiang Province, China. The parameters of each unit in the system are shown in Table 1.
[0200] Table 1. Energy System Parameter Diagram
[0201]
[0202] In the table, This represents the average annualized investment coefficient. This indicates the unit capacity investment cost of the ice storage system. This indicates the unit capacity investment cost of the dual-condition refrigeration unit. This indicates the unit capacity investment cost of the base-load refrigeration unit. This indicates the unit capacity investment cost of a thermal energy storage system. This indicates the unit capacity investment cost of a heat pump system. Indicates the upper and lower limits of the state of charge of the ice storage system. Indicates the upper and lower limits of the state of charge of a thermal energy storage system. This indicates the energy efficiency ratio of the dual-mode chiller in cooling mode. This indicates the energy efficiency ratio of the dual-mode refrigeration unit in ice-making mode. Indicates the energy efficiency ratio of the base-load refrigeration unit. This indicates the energy efficiency ratio of the heat pump system.
[0203] Considering the complexity of the annual operation data of the integrated energy system and the computational efficiency requirements, this study uses a typical day selection method based on K-means clustering to preprocess the training data. Since electricity prices vary significantly across different months, the 8760 hours of data for the entire year are decomposed by month, and typical days are extracted for each month.
[0204] The hourly data for the entire year is divided into calendar months, forming 12 monthly datasets. For month m, the data dimension is... ,in The number of days in the month is represented by 24, the number of hours in each day is 24, and the number of features is 4 (electrical load, thermal load, cooling load, and photovoltaic output). The 4-dimensional feature data for each 24-hour period is flattened to construct a feature matrix. , .
[0205]
[0206] In the formula, Indicates the first The characteristic matrix of the month, Indicates the first The original data for the month, reshape(·) represents the data reshaping function, which flattens the three-dimensional tensor into a two-dimensional matrix. Indicates the first The number of days in a month.
[0207] The feature matrix is standardized using the Z-score method. ,in, This represents the standardized feature matrix for the m-th month. and Let be the mean vector and standard deviation vector of the feature data for month m, respectively.
[0208] Combining K-means clustering and an extreme value selection strategy, the following steps are taken: First, select the dates containing the hourly peak values of electrical load, thermal load, cooling load, and photovoltaic output, as well as the dates with the highest daily total output of each type of load and photovoltaic power generation. If there are no relevant loads in a given month, the corresponding dates are not selected. Then, K-means clustering is performed on the monthly data, selecting the dates closest to each cluster center. Three cluster centers are chosen for each month. The K-means clustering algorithm partitions the data by minimizing the sum of squares within each cluster.
[0209]
[0210] in For the first One cluster, For the corresponding cluster centers, For a single sample point in the dataset, This represents the number of clusters. The final clustering result is as follows: Figure 9 As shown, 80 typical days were selected for reinforcement learning training. To enhance the reinforcement learning's ability to handle uncertainty, a 20% perturbation was added to the load data before it was fed into the environment.
[0211] To analyze the impact of directly loading pre-trained historical agents on shortening training time and reward completion when faced with different configuration schemes, we defined a design range for each device, randomly generated 13 configuration schemes based on this range, selected one as the base scheme, and used the formula... The similarity between other configurations and the base scheme was calculated, and they were named Config1-12 in descending order.
[0212] Figure 10 This is a comparison chart of the training results of the Load method and the Dir method in an embodiment of the present invention.
[0213] like Figure 10 As shown, in order to verify the effectiveness of the configuration transfer learning method proposed in this invention, two cases were set up: case 1: when using SAC to optimize the operation of these 12 configs, the agent was randomly initialized and trained; case 2: when using SAC to optimize the operation of these 12 configs, a stable historical agent was loaded and trained in the Config base during the initialization phase.
[0214] Table 2. The 13 generated configuration schemes
[0215]
[0216] Experimental results demonstrate the significant advantages of configuration transfer learning in reinforcement learning. In all 12 configurations, the Load method (orange line) using a pre-trained model consistently outperformed the Dir method (blue line) trained from scratch. The Load method exhibited three key advantages: (1) higher initial performance, with initial rewards typically ranging from -2750 to -2000, compared to approximately 200-500 for the Dir method; (2) faster learning convergence, with most configurations reaching stable performance within 50-100 training steps, while the Dir method required 100-150 steps; and (3) superior final performance in most configurations. These results indicate that training a stable Load agent using a Config base can significantly improve learning efficiency and performance ceiling.
[0217] The performance differences between different configurations mainly stem from the applicability of the pre-training strategy to the target configuration. In Config1, 2, 5, 6, 9, 11, and 12, the Load method ultimately outperforms the Dir method in terms of stability. This is because the optimization strategies learned by the pre-trained agent on the Config base effectively guide the optimization process of these configurations. The pre-trained model has mastered the coordination mechanisms and energy management strategies between devices, enabling the agent to avoid getting trapped in local optima and find better global optimal strategies. Conversely, in Config3, 4, 7, 8, and 10, although the Load method has a significant advantage in the early stages of training, the final performance of the two methods tends to be similar. This indicates that for these configurations, pre-training knowledge mainly plays a role in accelerating early learning and helping the agent quickly move past the exploration phase. However, in the long-term optimization process, a randomly initialized agent can also achieve similar performance levels through sufficient training. This phenomenon confirms the universality of configuration transfer learning methods in improving training efficiency and also illustrates the importance of the quality of pre-training strategies for improving final performance.
[0218] Figure 11 This is a flowchart illustrating the process of directly training the agent 20 times in an embodiment of the present invention.
[0219] Figure 12 This is a flowchart illustrating the preloaded intelligent agent training process for 300 iterations in an embodiment of the present invention.
[0220] Figure 13 This is a flowchart illustrating the preloaded intelligent agent training process 20 times in an embodiment of the present invention.
[0221] like Figure 11 , Figure 12 and Figure 13 As shown, after 20 iterations of direct training, the agent had not yet learned the charging and discharging actions of energy storage. Most of the load was directly supplied by the heating equipment, and energy storage did not play its corresponding role. However, after loading a pre-trained historical agent, in the same 20 iterations, the agent learned the low-charge, high-discharge action of energy storage, completing energy discharging during periods of high daytime load and high electricity prices. This demonstrates that loading a pre-trained historical agent can effectively enable a new agent to quickly learn the relevant strategies, effectively reducing training time.
[0222] in, Figures 11-13 In the figure, the units of the numbers on the horizontal axis are hours, and the units of the numbers on the vertical axis are kW.
[0223] To verify the effectiveness of the proposed method, we designed two comparative cases to compare it with the proposed method:
[0224] Case 1: A Benchmark Method Based on Deterministic Optimization. Based on a fundamental scenario, this case employs the commercial Gurobi solver for deterministic planning. Gurobi is a high-performance mixed-integer linear programming solver capable of efficiently solving large-scale optimization problems and widely used in energy system planning. This case uses historical data from across the year to optimize equipment capacity configuration, obtaining a deterministic configuration scheme. Subsequently, the obtained configuration scheme is trained and optimized using the SAC agent framework. Finally, the trained agent is applied to 100 uncertain scenarios, and the average annual total cost across these 100 scenarios is calculated as the evaluation metric.
[0225] Case 2: Comparative Approach Based on Stochastic Programming. K-means clustering analysis was performed on 100 generated uncertainty scenarios to extract 10 representative typical scenarios. Two-stage stochastic programming was then used for optimization based on these typical scenarios. Two-stage stochastic programming is a classic mathematical programming method for handling uncertain decision-making problems. Its core idea is to divide the decision-making process into two stages: the first stage makes long-term investment decisions such as equipment configuration before the uncertain parameters are realized; the second stage, after the uncertain parameters are realized, makes short-term operational decisions such as operation scheduling based on the determined decisions of the first stage and the actual scenario information. The objective function is to minimize the sum of the first-stage cost and the expected cost of the second stage under all possible scenarios, thereby achieving coordinated optimization of design and operation while considering uncertainty. After obtaining the configuration scheme, the same evaluation method as in Case 1 was used, i.e., SAC was used for operational optimization training, and the scheme was tested in 100 scenarios to calculate the average annual total cost. Through these two comparative cases, the advantages of the proposed method over traditional deterministic programming and classic stochastic programming methods in handling uncertainty can be verified.
[0226] Figure 14 This is an iterative optimization diagram of the annual total cost of TPE-SAC in an embodiment of the present invention.
[0227] like Figure 14 The figure shows the optimization process of the annual total cost obtained by the TPE-SAC method proposed in this paper. As can be observed from the figure, the current optimal annual total cost, represented by the red line, gradually decreases from the initial approximately ¥5,554,655 to the final optimal solution of ¥5,046,574. The entire optimization process can be clearly divided into two stages: a rapid descent stage (the first 70 iterations) and a fine-search stage (after 70 iterations). In the rapid descent stage, the TPE algorithm quickly identifies promising parameter regions by modeling the probability distribution of existing samples, reducing the optimal cost from ¥5,554,655 to approximately ¥5,129,951, achieving a cost saving of approximately ¥420,000. Subsequently, in the fine-search stage, the algorithm conducts a more detailed exploration near the found optimal regions, finally finding the global optimum in the 173rd iteration.
[0228] Table 3. Configuration schemes obtained by the three methods
[0229]
[0230] The configuration scheme table reveals differences in configuration characteristics among different methods. Case 1, based on deterministic optimization of a single basic scenario, yielded a relatively conservative configuration scheme, with ice storage system at only 9479 kWh and thermal energy storage system at 18524 kWh. This aligns with the characteristic of deterministic optimization, which only considers supply and demand balance under a single scenario, and the optimizer tends to select the minimum configuration that meets the specific scenario requirements to minimize investment costs. Case 2, employing a two-stage stochastic programming method that clusters typical scenarios, exhibits a more aggressive energy storage configuration, with ice storage system reaching 26938 kWh and thermal energy storage system remaining at a similar level of 18483 kWh. This reflects the characteristic of two-stage stochastic programming considering the probability distribution of multiple scenarios, identifying the importance of energy storage in coping with fluctuations in renewable energy output through probability weighting of typical scenarios. However, due to the potential loss of some scenario information during the clustering process, the thermal energy storage configuration is relatively conservative. In comparison, the TPE-SAC method achieves a more balanced configuration, with an ice storage system of 15794 kWh and a thermal energy storage system of 24604 kWh. Furthermore, the configurations of the dual-condition chiller, base-load chiller, and heat pump system are all moderately higher than the previous two methods. This configuration characteristic reflects the core advantage of this method: through reinforcement learning training across numerous scenarios, the agent can learn the collaborative mechanisms of different devices under various uncertain conditions. This allows for a more balanced device configuration in a comprehensive energy system coupling electricity, heat, and cooling, avoiding the excessive conservatism of deterministic methods and overcoming the local bias problem that may exist in traditional stochastic programming.
[0231] The evaluation results of the three methods are shown in Table 4. The evaluation results demonstrate the significant advantages of the TPE-SAC method in system performance optimization. In terms of the key total cost indicator, TPE-SAC achieves a cost of 5,046,574 yuan, saving 6.2% compared to Case 1's 5,377,790 yuan and 2.6% compared to Case 2's 5,182,455 yuan. More importantly, TPE-SAC demonstrates superior operational reliability, with an average cold penalty of only 12,131, a reduction of 96.3% compared to Case 1 and 93.2% compared to Case 2; the average hot penalty is further reduced to 365, a reduction of 99.7% compared to Case 1 and 99.4% compared to Case 2. Although TPE-SAC has the highest investment cost (930,628 yuan), it achieves optimal overall economic efficiency by significantly reducing operating costs (4,103,450 yuan) and almost eliminating penalty costs.
[0232] Figure 15 This is a schematic diagram of the cold penalty in all scenarios for the three configuration schemes in the embodiments of the present invention.
[0233] Figure 16 This is a schematic diagram of the heat penalty for three configuration schemes in all scenarios in the embodiments of the present invention.
[0234] like Figure 15 and Figure 16 As shown, the dynamic performance of different methods in actual operation is illustrated. It can be observed that the TPE-SAC method (blue line) maintains a low and stable operating level for most periods, while Case 1 and Case 2 exhibit significant fluctuations and peaks at certain times. This difference indicates that the TPE-SAC operating strategy, optimized through reinforcement learning, can better adapt to changes in demand across different scenarios, effectively avoiding the penalty costs caused by supply-demand imbalances. Overall, the TPE-SAC method achieves optimal economic performance while ensuring system reliability through intelligent configuration design and operational optimization. Figures 15-16 In the figure, the units of the numbers on the horizontal axis are all hours, and the units of the numbers on the vertical axis are all yuan.
[0235] Table 4. Evaluation results of the three configuration schemes
[0236]
[0237] This embodiment also provides a multi-scenario optimization system for integrated energy systems based on TPE-SAC, including: a TPE algorithm optimization module, a configuration migration module, a SAC algorithm optimization module, a performance evaluation module, and an iteration module.
[0238] The TPE algorithm optimization module uses step S1 above to optimize the configuration parameters of equipment in the integrated energy system using the TPE algorithm model, and generates a configuration scheme.
[0239] The configuration migration module adopts the above step S2, sets a similarity score threshold, generates a similarity score based on the configuration scheme and the historical configuration scheme, and when the similarity score is greater than the similarity score threshold, the historical agent corresponding to the historical configuration scheme is used as the loaded agent.
[0240] The SAC algorithm optimization module uses step S3 above to optimize the loaded agent using the SAC algorithm model, generating an optimized agent.
[0241] The performance evaluation module uses step S4 above to evaluate the performance of the optimized agent, obtains the performance evaluation results, and stores the configuration scheme, the optimized agent, and the performance evaluation results.
[0242] The iterative module uses step S5 to feed back the configuration scheme, optimized agent, and performance evaluation results to the TPE algorithm model, thereby updating the TPE algorithm model.
[0243] This embodiment has the following technical effects:
[0244] The TPE-SAC-based multi-scenario optimization method and system for integrated energy systems involved in this embodiment includes: Step S1, optimizing the configuration parameters of equipment in the integrated energy system using the TPE algorithm model to generate a configuration scheme; Step S2, setting a similarity score threshold, generating a similarity score based on the configuration scheme and historical configuration schemes, and when the similarity score is greater than the similarity score threshold, using the historical agent corresponding to the historical configuration scheme as the loaded agent; Step S3, optimizing the loaded agent using the SAC algorithm model to generate an optimized agent; Step S4, performing performance evaluation on the optimized agent, obtaining the performance evaluation results, and storing the configuration scheme, optimized agent, and performance evaluation results; Step S5, feeding back the configuration scheme, optimized agent, and performance evaluation results to the TPE algorithm model to update the TPE algorithm model. Therefore, the TPE-SAC-based multi-scenario optimization method and system for integrated energy systems of this invention, by combining the TPE-SAC algorithm, ensures the real-time balance of cooling, electricity, and heat loads of the integrated energy system in multiple scenarios, avoids the penalty cost caused by supply and demand imbalance, and significantly reduces the total cost of the integrated energy system.
[0245] This embodiment updates the TPE algorithm model based on the performance evaluation results, significantly improving model performance.
[0246] This embodiment also includes a running mode switching mechanism, which enables the system to automatically switch running modes according to load requirements.
[0247] This embodiment also uses the Monte Carlo sampling method to generate evaluation scenarios, comprehensively testing the adaptability and optimization effect of the trained agent under various conditions.
[0248] Those skilled in the art should understand that this invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to this invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A multi-scenario optimization method for integrated energy systems based on TPE-SAC, characterized in that, include: Step S1: Optimize the configuration parameters of equipment in the integrated energy system using the TPE algorithm model to generate a configuration scheme; Step S2: Set a similarity score threshold, generate a similarity score based on the configuration scheme and the historical configuration scheme, and when the similarity score is greater than the similarity score threshold, use the historical agent corresponding to the historical configuration scheme as the loaded agent. Step S2 includes the following sub-steps: Step S20: Perform Z-score standardization on all historical configuration scheme parameters to obtain standardized historical configuration scheme parameters; Step S21: Calculate the Euclidean distance between the configuration scheme parameters corresponding to the configuration scheme and the standardized historical configuration scheme parameters, and obtain the similarity score by combining the historical performance evaluation results; Step S22: Set the similarity score threshold. When the similarity score is greater than the similarity score threshold, start the configuration transfer learning mechanism and use the historical agent as the loaded agent. Step S3: Optimize the loaded agent using the SAC algorithm model to generate an optimized agent; Step S4: Perform performance evaluation on the optimized agent, obtain the performance evaluation results, and store the configuration scheme, the optimized agent, and the performance evaluation results; Step S5: Feed back the configuration scheme, the optimized agent, and the performance evaluation results to the TPE algorithm model to update the TPE algorithm model.
2. The multi-scenario optimization method for integrated energy systems based on TPE-SAC according to claim 1, Its features are: Step S1 includes the following sub-steps: Step S10: Divide the historical observations into two subsets according to a preset threshold, namely a high-quality set and a low-quality set. Establish conditional probability density models for the high-quality set and the low-quality set respectively to obtain the probability density function of the high-quality set and the probability density function of the low-quality set. Step S11: The desired improvement is taken as the acquisition function, and the desired improvement is expressed as the ratio of the probability density function of the high-quality set to the probability density function of the low-quality set using Bayes' rule and the law of total probability. Step S12: Select sampling points in the acquisition function by maximizing the desired improvement, and generate the configuration scheme based on the sampling points.
3. The multi-scenario optimization method for integrated energy systems based on TPE-SAC according to claim 1, Its features are: in, Step S3 includes the following sub-steps: Step S30: Define the five-tuple of the loaded agent using an extended state-space Markov decision process, and establish an objective function based on the five-tuple; Step S31: The policy entropy regularization term is introduced into the objective function using the SAC algorithm to obtain the first optimization objective function; Step S32: Design the action space and state space of the first optimization objective function according to the DRL environment, and normalize the state space to obtain the second optimization objective function; Step S33: Limit the action space according to load requirements; Step S34: Design the reward function of the second optimization objective function according to the optimization requirements to obtain the third optimization objective function, and obtain the optimized agent based on the third optimization objective function.
4. The multi-scenario optimization method for integrated energy systems based on TPE-SAC according to claim 1, Its features are: Step S4 includes the following sub-steps: Step S40: Based on historical operation data, an evaluation scenario set S is generated using the Monte Carlo sampling method. The evaluation scenario set S includes 100 scenarios. Step S41: Calculate the relative deviation between the peak value of the cooling and heating load and the total value of the evaluation scenario set S, and use the optimization agent to perform a year-round simulation of the evaluation scenario set S to obtain the performance evaluation result, which includes the total cost. Step S42: Store the configuration scheme, the optimized agent, and the performance evaluation results.
5. The multi-scenario optimization method for integrated energy systems based on TPE-SAC according to claim 1, characterized in that, Also includes: Step S0: Set the optimization objective for the equipment in the integrated energy system, namely, minimize the total annual cost through the capacity configuration of the ice storage system, dual-condition refrigeration host, base load refrigeration host, thermal energy storage system, and heat pump system, and model the equipment and determine the constraints of the equipment.
6. A multi-scenario optimization system for integrated energy systems based on TPE-SAC, characterized in that, include: The TPE algorithm optimization module is used to optimize the configuration parameters of equipment in the integrated energy system using the TPE algorithm model and generate configuration schemes. The configuration migration module is used to set a similarity score threshold, generate a similarity score based on the configuration scheme and the historical configuration scheme, and when the similarity score is greater than the similarity score threshold, use the historical agent corresponding to the historical configuration scheme as the loaded agent. The configuration migration module includes the following units: The standardization processing unit is used to perform Z-score standardization processing on all historical configuration scheme parameters of the historical configuration scheme to obtain standardized historical configuration scheme parameters. A similarity score generation unit is used to calculate the Euclidean distance between the configuration scheme parameters corresponding to the configuration scheme and the standardized historical configuration scheme parameters, and to obtain the similarity score by combining the historical performance evaluation results; The transfer learning unit is used to set the similarity score threshold. When the similarity score is greater than the similarity score threshold, the configuration transfer learning mechanism is started, and the historical agent is used as the loaded agent. The SAC algorithm optimization module is used to optimize the loaded agent using the SAC algorithm model to generate an optimized agent. The performance evaluation module is used to evaluate the performance of the optimized agent, obtain the performance evaluation results, and store the configuration scheme, the optimized agent, and the performance evaluation results. The iteration module is used to feed back the configuration scheme, the optimized agent, and the performance evaluation results to the TPE algorithm model to update the TPE algorithm model.
Citation Information
Patent Citations
An electric-heat integrated energy system coordination optimization method, system and equipment and a storage medium
CN113902040A
Multi-target planning method and system for regional integrated energy system
CN116957362A
Load characteristic prediction method based on Informer hyper-parameter optimization model
CN119538024A