Multi-objective optimization scheduling method for integrated energy system based on ppo-moma

CN122736196APending Publication Date: 2026-09-11SHENYANG INST OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610890416.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0003]然而,目前多数研究针对经济性与碳排放协同优化的IES调度策略研究仍较为有限

Benefits of technology

本发明提出的基于PPO-MOMA的综合能源系统多目标优化调度方法,通过建立以运行成本最小和碳排放最小为目标并配置系统功率平衡、机组出力及设备运行的约束条件的综合能源系统多目标优化调度模型,能够可完整刻画多能耦合系统实际运行机理,避免模型简化造成调度结果失真;通过粒子编码生成初始化种群并经Pareto非支配排序筛选非支配解集存入外部储备集,能够快速形成初始可行解空间,为后续迭代寻优奠定优质种群基础;采用PPO-MOMA算法由PPO自适应迭代动态输出交叉概率与变异概率,并对种群执行差分进化交叉与自适应变异、更新外部储备集,克服了传统智能算法交叉变异参数固定、无法适配高维复杂调度空间的缺陷,增强全局搜索能力与种群多样性;以超体积HV为奖励信号驱动PPO策略网络和价值网络迭代更新,直至HV提升停滞再经非支配排序、拥挤度选拔及差分进化局部搜索循环迭代并输出Pareto最优解集,以此借助强化学习自主反馈学习机制,有效提升算法收敛速度、寻优精度以及Pareto解集的收敛性与分布均匀性,避免陷入局部最优;通过熵权-TOPSIS法对Pareto最优解集进行归一化、客观赋权并优选最佳调度方案,对应摒弃人为设定目标权重的主观偏差,科学均衡兼顾运行成本与碳排放双目标,最终在满足系统各类运行约束前提下充分消纳风光可再生能源、合理优化各机组及耦合设备出力,有效降低系统运行成本并减少全环节碳排放,实现综合能源系统经济与低碳协同优化调度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736196A_ABST
    Figure CN122736196A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of energy system scheduling, and discloses a kind of multi-objective optimization scheduling method of integrated energy system based on PPO-MOMA, comprising the following steps: the scheduling model with minimum cost and carbon emission as target is established;Algorithm PPO-MOMA is used to solve the model, wherein PPO is dynamically trained and constantly outputs crossover probability pc And mutation probability pm;And this is used to execute difference evolution crossover and adaptive variation on population, generate new population and update external reserve set AC;With hyper volume index HV As reward signal, through PPO strategy network and value network iteration update until HV stagnation, after non-dominated sorting and crowding degree selection, the Pareto optimal solution set is output by local difference evolution search;The Pareto optimal solution set is decided by entropy weight-TOPSIS method, and the best scheduling scheme is determined;The method can effectively deal with the economic problem of integrated energy system, to reduce the operation cost of integrated energy system and reduce carbon emissions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-objective optimization scheduling technology for energy systems, and in particular to a multi-objective cultural gene algorithm (PPO-MOMA) based on proximal policy optimization (PPO) to solve the optimization scheduling model. Background Technology

[0002] Integrated Energy Systems (IES) have become an important way to solve energy sustainability issues by integrating multiple energy forms such as natural gas, solar power, and wind power within a region. Meanwhile, with the large-scale integration of renewable energy sources such as wind and solar power into IESs, economic viability and carbon emissions have become important standards for measuring system stability and regulation capabilities.

[0003] However, current research on IES scheduling strategies that synergistically optimize economic efficiency and carbon emissions remains relatively limited. Meanwhile, the multi-objective optimization scheduling problem of IES is essentially a high-dimensional, complex, and strongly nonlinear combinatorial optimization problem. Conventional weight-allocation multi-objective optimization methods suffer from drawbacks such as strong subjective dependence, difficulty in scientifically quantifying objective weights, and large deviations in human-defined settings, failing to accurately balance the dual scheduling objectives of economic operation and low carbon emissions.

[0004] Specifically, traditional multi-objective cultural gene algorithms often use fixed parameter settings for crossover and mutation probabilities during iterative optimization, lacking a dynamic adjustment mechanism based on changes in population iteration state, external reserve set quality, and hypervolume index. Fixed parameter values ​​cannot adapt to the optimization needs of the high-dimensional and complex scheduling space of integrated energy systems. Furthermore, conventional multi-objective intelligent optimization algorithms lack policy feedback and autonomous learning capabilities in reinforcement learning, resulting in limited convergence performance and search efficiency in high-dimensional decision spaces. They struggle to stably generate high-quality Pareto non-dominated solution sets, failing to meet the scheduling solution requirements for the coordinated optimization of economic and low-carbon aspects of integrated energy systems. Summary of the Invention

[0005] This invention proposes a multi-objective optimization scheduling method for integrated energy systems based on PPO-MOMA to address the shortcomings of the prior art. This method can effectively address the economic issues of integrated energy systems, thereby reducing the operating costs and carbon emissions of integrated energy systems.

[0006] The technical solution of this invention is: a multi-objective optimization scheduling method for integrated energy systems based on PPO-MOMA, comprising the following steps: Establish a multi-objective optimization scheduling model for integrated energy systems that simultaneously aims at minimizing operating costs and carbon emissions, and configure constraints for system power balance, unit output, and equipment operation. An initial population is generated and constrained according to particle coding rules and scheduling principles. The non-dominated solution set is filtered by Pareto non-dominated sorting and stored in an external reserve set. The multi-objective cultural gene algorithm PPO-MOMA based on proximal strategy optimization is used to solve the multi-objective optimization scheduling model. The PPO algorithm takes the state variable representing the current evolutionary state as input and dynamically outputs the crossover probability and mutation probability. Based on the crossover probability and mutation probability, differential evolution crossover and adaptive mutation operations are performed on the population to generate a new population and update the external reserve set. Using the hypervolume index as the reward signal for the PPO algorithm, the parameters are iteratively updated through the PPO policy network and value network, and the current crossover probability, mutation probability and corresponding population are output. The current population is sorted by non-dominated order and crowding selection to obtain a new generation population, and a local search is performed on the new generation population. The steps of solving the problem to the local search are executed iteratively, and the Pareto optimal solution set in the external reserve set is output. By using the entropy-weighted TOPSIS method to make comprehensive decisions on the Pareto optimal solution set, the optimal integrated energy system scheduling scheme that balances operating costs and carbon emissions is selected.

[0007] In at least one embodiment of the present invention, the state variable of the PPO algorithm is a two-dimensional vector containing the current crossover probability and mutation probability; the action space is set with four types of discrete adjustment actions, which correspond to the parameter adjustment methods of increasing or decreasing the crossover probability and increasing or decreasing the mutation probability, respectively.

[0008] In at least one embodiment of the present invention, the differential evolution crossover uses a scaling factor F and a crossover probability CR to generate a mutant population VaPop, and generates a new population through binary crossover.

[0009] In at least one embodiment of the present invention, when performing differential evolution crossover and adaptive mutation on the population using the crossover probability and mutation probability, the mutation amplitude σ is dynamically adjusted according to the current iteration number gen and the maximum iteration number setgen, and mutation is completed by superimposing Gaussian noise NS on the original population.

[0010] In at least one embodiment of the present invention, when performing non-dominated sorting and crowding selection, the mutated population is merged with the original population, and non-dominated sorting plus crowding descending selection of NSGA-II is used to obtain a new generation of population with unchanged population size.

[0011] In at least one embodiment of the present invention, the entropy weight-TOPSIS method includes: normalizing operating costs and carbon emissions; determining objective weights using information entropy; calculating the weighted Euclidean distance between each scheme and the positive and negative ideal solutions, and selecting the scheme with the largest relative proximity as the final scheduling scheme.

[0012] In at least one embodiment of the present invention, the integrated energy system includes thermal power units, combined heat and power, wind power, photovoltaic power, electricity-to-gas conversion, electric heat pumps, and gas boilers.

[0013] In at least one embodiment of the present invention, the operating cost objective function simultaneously takes into account the operating costs of CHP units, TPU units, electricity purchased from the grid, gas purchased from the gas grid, as well as AE equipment, HP equipment, and GB equipment.

[0014] In at least one embodiment of the present invention, the carbon emission objective function simultaneously takes into account the carbon emissions generated by CHP units, TPU units, purchased electricity from the external grid, gas-fired boilers, and the electricity-to-gas conversion process.

[0015] In at least one embodiment of the present invention, the integrated energy system scheduling process simultaneously satisfies the constraints of power balance, heat balance, and gas volume flow balance; it also satisfies the upper and lower limits of output and ramping constraints of CHP units and TPU units, as well as the upper and lower limits of operating output constraints of wind turbine units, photovoltaic units, electric-to-gas devices, electric heat pumps, and gas boilers.

[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention proposes a multi-objective optimization scheduling method for integrated energy systems based on PPO-MOMA. By establishing a multi-objective optimization scheduling model for integrated energy systems with the objectives of minimizing operating costs and carbon emissions, and configuring constraints on system power balance, unit output, and equipment operation, this method can fully characterize the actual operating mechanism of multi-energy coupled systems, avoiding distortion of scheduling results due to model simplification. By generating an initial population through particle encoding and filtering the non-dominated solution set through Pareto non-dominated sorting and storing it in an external reserve set, an initial feasible solution space can be quickly formed, laying a high-quality population foundation for subsequent iterative optimization. The PPO-MOMA algorithm is used to dynamically output crossover and mutation probabilities through adaptive iteration of PPO, and differential evolution crossover and adaptive mutation are performed on the population to update the external reserve set. This overcomes the shortcomings of traditional intelligent algorithms, such as fixed crossover and mutation parameters and inability to adapt to high-dimensional complex scheduling spaces, enhancing global search capabilities and population dynamics. The algorithm employs a population diversity mechanism. It uses hypervolume (HV) as a reward signal to drive iterative updates of the PPO policy network and value network until HV improvement stagnates. Then, it iterates through non-dominated sorting, crowding selection, and differential evolution local search to output the Pareto optimal solution set. This leverages the autonomous feedback learning mechanism of reinforcement learning to effectively improve the algorithm's convergence speed, optimization accuracy, and the convergence and distribution uniformity of the Pareto solution set, avoiding getting trapped in local optima. The entropy-weighted TOPSIS method is used to normalize and objectively weight the Pareto optimal solution set, selecting the best scheduling scheme. This eliminates the subjective bias of manually setting target weights, scientifically balancing the dual objectives of operating cost and carbon emissions. Ultimately, under the premise of meeting various system operating constraints, it fully utilizes wind and solar renewable energy, rationally optimizes the output of each unit and coupled equipment, effectively reduces system operating costs and carbon emissions across all stages, and achieves economical and low-carbon synergistic optimization scheduling of the integrated energy system. Attached Figure Description

[0017] Figure 1 This is a structural diagram of the integrated energy system according to an embodiment of the present invention; Figure 2 This is the overall flowchart of the algorithm of this invention; Figure 3 The Pareto curve for the economic carbon emissions of this invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the described embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains. As used herein, the words “comprising” or “including” and similar terms mean that an element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects.

[0020] Currently, research on IES scheduling strategies that synergistically optimize economic efficiency and carbon emissions remains relatively limited. Meanwhile, the multi-objective optimization scheduling problem of IES is essentially a high-dimensional, complex, and strongly nonlinear combinatorial optimization problem. Traditional mathematical methods are no longer sufficient in terms of accuracy and efficiency, and weight allocation methods also have limitations due to their strong subjectivity and difficulty in quantifying weights.

[0021] In contrast, Pareto-based intelligent optimization methods can effectively overcome the aforementioned problems, exhibiting stronger adaptability and optimization capabilities. The PPO algorithm, as an emerging reinforcement learning method, combines multi-objective intelligent optimization strategies, maintaining global search capabilities while improving learning efficiency and convergence performance in high-dimensional policy spaces, demonstrating broad application prospects.

[0022] In view of this, the present invention proposes a multi-objective optimization scheduling method for integrated energy systems based on PPO-MOMA, which can effectively address the economic problems of integrated energy systems, thereby reducing the operating costs and carbon emissions of integrated energy systems. Combination Figures 1 to 3 As shown, a multi-objective optimization scheduling method for integrated energy systems based on PPO-MOMA includes the following steps: Establish a multi-objective optimization scheduling model for integrated energy systems that simultaneously aims at minimizing operating costs and carbon emissions, and configure constraints for system power balance, unit output, and equipment operation. An initial population is generated and constrained according to particle coding rules and scheduling principles. The non-dominated solution set is filtered by Pareto non-dominated sorting and stored in an external reserve set. The multi-objective cultural gene algorithm PPO-MOMA based on proximal strategy optimization is used to solve the multi-objective optimization scheduling model. The PPO algorithm takes the state variable representing the current evolutionary state as input and dynamically outputs the crossover probability and mutation probability. Based on the crossover probability and mutation probability, differential evolution crossover and adaptive mutation operations are performed on the population to generate a new population and update the external reserve set. Using the hypervolume index as the reward signal for the PPO algorithm, the parameters are iteratively updated through the PPO policy network and value network, and the current crossover probability, mutation probability and corresponding population are output. The current population is sorted by non-dominated order and crowding selection to obtain a new generation population, and a local search is performed on the new generation population. The steps of solving the problem to the local search are executed iteratively, and the Pareto optimal solution set in the external reserve set is output. The optimal Pareto solution set is comprehensively evaluated using the entropy-weighted TOPSIS method to select the optimal integrated energy system scheduling scheme that balances operating costs and carbon emissions. Specifically, the detailed steps of this scheme are as follows: (1) Population initialization. gen=0; Generate an initial population and perform constraint processing according to the particle coding rules and scheduling principles; initialize the population objective function value; select the non-dominated solution set of the initial population using the Pareto non-dominated sorting method and store it in the external reserve set (AC).

[0023] (2) Global search of MOMA algorithm. The global search of MOMA algorithm includes crossover, mutation and selection operations. Crossover and mutation operations are performed in PPO algorithm. PPO algorithm continuously updates the crossover probability pc and mutation probability pm based on historical experience, and then continuously updates the population through crossover and mutation. After training, the best population and crossover and mutation probabilities are returned.

[0024] The specific steps are as follows: 1) Crossover and mutation operations based on the PPO algorithm.

[0025] a. Input data. This includes the parameters of the PPO algorithm itself and the variables passed from the MOMA algorithm. The MOMA algorithm's input variables are the population Pop, the hypervolume index HV, the external reserve set AC, and the current iteration number gen.

[0026] b. Initialization. Load the environment configuration program and initialize the environment; build and initialize the policy network and value network. The policy network takes a 2D state Stat=[pc, pm] as input and outputs a probability distribution of 4 actions after passing through two fully connected layers; the value network takes a 2D state as input and outputs a state value after passing through two fully connected layers.

[0027] c. Reset the environment. Initialize the external reserve set AC in the environment configuration to be the AC passed in from the gen-th iteration of MOMA, the population EnPop in the environment configuration to be the population Pop passed in from the gen-th iteration of MOMA, the pc in State to be any value between 0.6 and 0.9, the pm to be any value between 0.05 and 0.3, the reward to be 0, and the optimal hypervolume index. HV is passed to the MOMA for the gen-th iteration.

[0028] d. Environment Interaction Loop. The environment interaction loop is the core part of the PPO algorithm during training. Through interaction with the environment, it collects data such as State, Action, Reward, and Isdone for subsequent calculation of the advantage function and verifies the current policy's performance in the environment, providing feedback for subsequent policy updates. The specific process is as follows:

[0029] ① Action Selection and Value Prediction. For action selection, firstly, the current State is input into the policy network to obtain the predicted value from the policy network and convert it into the probability of each action; secondly, an Action is randomly selected from the action set a={1,2,3,4} based on the action probability; finally, the natural logarithm of the probability of the selected Action is taken to calculate the log probability of the action, which is used to calculate the loss function and update the policy in the subsequent PPO algorithm. For value prediction, the current State is input into the value network directly to obtain the predicted value of the current state, Value.

[0030] ② Calculate State. Update the crossover and mutation probabilities in Stat based on the selected action, as shown in Equation (1).

[0031] ③ Differential Evolutionary Crossover. Differential evolutionary crossover has high global search capability and fast convergence speed for complex search spaces, and is widely used in problems. First, a new mutant population VaPop is generated by adding a difference vector, as shown in Equation (2); second, binary crossover is used to select information from the target vector to generate a new population, as shown in Equation (3).

[0032] In the formula, Population represents the population in the environmental configuration; Represents three random, non-repeating population indices, r represents the current individual index; F is the scaling factor; CR represents the differential evolution crossover probability; This represents a dimension randomly selected from the population decision variable dimension Dim.

[0033] ④ Adaptive mutation. First, calculate the mutation magnitude based on the current iteration number gen and the maximum iteration number setgen. As shown in Equation (4), the variation amplitude is large in the early stage of the search, which is beneficial for exploration, while the variation amplitude is small in the later stage, which is beneficial for local search; secondly, the noise term NS is generated by Gaussian distribution and variation amplitude, as shown in Equation (5); finally, the mutation operation is performed on the original population, as shown in Equation (6).

[0034] In the formula, T represents the operating period; Dim represents the size of the decision variable dimension.

[0035] ⑤ Calculate the hypervolume index. The hypervolume index (HV) is a commonly used metric for evaluating algorithm performance; a higher value indicates a better quality Pareto solution. First, the population after crossover and mutation... The non-dominated solutions are screened, sorted, stored, and updated in the external reserve set AC; then the HV index of the Pareto solution of the new population is calculated according to Equation (7).

[0036] In the formula, This represents the Lebesgue measure, used to calculate the volume of a high-dimensional space; S represents the volume formed by the reference point and the non-dominated solution point i; S represents the non-dominated solution set.

[0037] ⑥ Calculate the Reward and determine the termination condition. In the PPO algorithm, the Reward defines the learning objective during the interaction with the environment and combines it with the advantage function to improve the accuracy of policy updates. In this paper, the reward value R is defined by comparing the HV index. If the HV index of the population after crossover mutation is higher than the historical best hypervolume index, the reward is awarded. If the value is large, a reward is given; otherwise, a penalty is imposed, as shown in Equation (8). The determination of the termination signal Isdone not only helps to end the environmental interaction cycle in advance and accelerate the running efficiency, but it is also one of the important variables in the calculation of the advantage function. The determination method is based on the method of HV improvement stagnation. If HV does not improve significantly for several consecutive times, the current round of training is terminated, as shown in Equation (9).

[0038] ⑦ Store experience. Store the calculated State, Reward, Isdone, selected Action, and predicted value estimate in the experience pool for use in calculating the advantage function.

[0039] e. Calculate the advantage function. The advantage function represents the improvement of the current action (Action) relative to the current state's value estimate (Value). It measures the advantage of choosing an action in state (State) compared to the average policy, providing directional guidance for the policy gradient and increasing the probability of high-reward actions while decreasing the probability of low-reward actions. The advantage function is calculated using Generalized Advantage Estimation (GAE). Combined with time difference (TD) error calculation and discount factor GAE hyperparameters The calculation is performed in reverse order, starting from the last environment interaction Maxsteps process, to complete the step advantage function for each environment interaction. The estimate.

[0040] f. Calculate the loss function. The loss function is the objective function of the PPO algorithm, which is calculated by weighting the policy loss and value loss using coefficients. Weighted calculation. Policy loss is based on the probability ratio of the old and new policies. Combined with the advantage function Introducing a clipping threshold The update range is limited by a pruning mechanism to prevent excessive policy updates; the value loss is the state value predicted by the value network. and the actual value of the target The mean square error is used to guide the optimization of the value network to approach the target value. Specifically, as shown in equations (10) and (11).

[0041] g. Network parameter optimization. The gradients of the policy network parameter θ and the value network parameter ϕ of the total loss function are calculated using the differential function, and the results are returned to the Adam network optimizer to optimize the parameters of the policy network and value network, thus completing the network update.

[0042] h. Terminate the judgment. Determine whether the condition is met. If so, return the trained crossover and mutation probabilities [pc, pm] and the global search population NewPop.

[0043] 2) Selection method.

[0044] The populations after crossover and mutation are merged with the original population to form a new population, and non-dominated selection based on crowding (NSGA-II Selection) is used for selection. First, the new population is sorted by non-dominated order and the crowding distance is calculated. Second, the non-dominated solution is selected as the new population. Finally, the original population is sorted in descending order of crowding and used to supplement the new population, thus generating the new population.

[0045] (3) Local search. The differential evolution algorithm is used for local search. The differential operator of this algorithm can better guide the exploration of the neighborhood and has stronger directionality. Moreover, the DE operation itself is based on the population to perform differential, which can update the population from multiple dimensions and enhance the regional diversity of the population.

[0046] (4) Calculate the objective function value. Calculate the objective function value of the new population, sort the non-dominated solutions and update the external reserve set, calculate the HV index of the non-dominated solutions in the external reserve set, and prepare for crossover and mutation based on the PPO algorithm in the next cycle.

[0047] (5) The algorithm terminates. Check if the number of iterations is satisfied. If so, output the external storage set.

[0048] Algorithm evaluation metrics: The performance of the PPO-MOMA algorithm is evaluated using two metrics: Inverted Generation Distance (IGD) and Hypervolume (HV). HV focuses on the convergence and distribution diversity of the Pareto solution set, reflecting the breadth of the frontier solution set; a larger HV is better, and its mathematical formula is shown in Equation (7). IGD evaluates the average distance between the solution set and the true Pareto front, mainly reflecting the convergence of the algorithm; a smaller IGD indicates better convergence, as shown in Equation (12).

[0049] In the formula, P is the set of points on the true Pareto front; Let P represent the number of points in set P; Q is the Pareto optimal solution set obtained by the algorithm. This represents the minimum Euclidean distance to a point in the set.

[0050] Comprehensive decision-making: Integrated energy system optimal scheduling is not only a multi-objective optimization problem but also a multi-attribute comprehensive decision-making problem. The entropy-weighted TOPSIS method is used to determine the optimal compromise solution. The specific decision-making process is as follows:

[0051] Step 1: Objective Normalization. Based on the Pareto optimal solution set obtained by the PPO-MOMA algorithm, the operating costs, flexibility-to-supply ratio, and carbon emissions obtained from different scheduling schemes are normalized to eliminate differences in the dimensions and orders of magnitude of different objectives, as shown in the following formula:

[0052] In the formula, This represents the value of the j-th objective function under the i-th scheduling scheme; and Let $\mathbf{j}$ represent the maximum and minimum values ​​of the j-th objective function under all scheduling schemes.

[0053] Step 2: Determine the positive and negative ideal solutions. Determine the positive ideal solution for the multi-objective optimal scheduling of the integrated energy system. and negative ideal solution This provides a reference point for subsequent distance calculations.

[0054] Step 3: Calculate weights using the information entropy method. Weights are calculated based on the differences in the distribution of each objective function value, avoiding subjective bias.

[0055] (15) In the formula, The entropy represents the information entropy of the j-th optimization objective. The larger the entropy, the more uniform the data and the lower its importance. N represents the number of optimization scheduling schemes selected. This represents the redundancy of the j-th optimization objective. The higher the redundancy, the greater the weight. This represents the weight of the j-th optimization objective.

[0056] Step 4: Calculate the weighted Euclidean distance. Calculate the positive ideal solution distance for each scheduling scheme according to equation (50). Distance to negative ideal solution The smaller the distance, the closer the solution is to the positive ideal solution or the farther away it is from the negative ideal solution.

[0057] (16) Step 5: Calculate the decision score of the scheduling scheme and select the optimal solution. If the scheduling scheme is close to the positive ideal solution and far from the negative ideal solution, then... The solution with the highest score is the optimal compromise solution.

[0058] (17) Furthermore, the integrated energy system structure of this embodiment of the invention has the following characteristics: The power supply side includes a combined heat and power (CHP) unit, wind turbine units, and an external power grid, providing the necessary energy foundation for the entire system; the intermediate flexible equipment includes electric heat pumps and electrolyzers, achieving decoupling between electricity-heat and electricity-hydrogen, thus improving flexibility; the load side includes electrical load, hydrogen market, and heat load, meeting industrial, social, and residential needs. An optimized scheduling model for the integrated energy system is derived from this system:

[0059] Objective function: (1) Economic scheduling: The IES economic scheduling objective is operating cost. Minimum, determined by the operating cost of the CHP unit. TPU unit operating costs Costs of purchasing and selling electricity and gas And the operating costs of AE, HP and GB equipment The objective function is as follows: (18) (19) In the formula, T represents the operating cycle; and Let represent the electrical output and thermal output of the i-th CHP unit at time t, respectively; This represents the electrical output of the j-th TPU unit at time t; Indicates the unit's coal consumption coefficient; This indicates the amount of power reduction caused by extracting a unit of heat; and These represent the interaction quantities with the upstream power grid and the upstream gas network at time t, respectively. and Indicates the interaction coefficient; , as well as These represent the electrical energy consumed by AE and HP, and the gas energy consumed by GB, respectively. This represents the operating cost coefficient.

[0060] (2) Low-carbon scheduling: IES low-carbon scheduling target is carbon emissions Minimum carbon emissions from CHP and TPU coal combustion Carbon emissions from the upper-level power grid GB carbon emissions The composition of carbon emissions from electricity-to-gas conversion is shown in the following formula: (20) (twenty one) Constraints: 1) Power balance constraint: (twenty two) In the formula, as well as This indicates the actual output of WT and PV.

[0061] 2) Thermal power balance constraint: (twenty three) In the formula, This represents the heat energy generated by HP at time t; This indicates that GB generates heat energy at time t; This represents the required heat load at time t.

[0062] 3) Gas volumetric flow rate balance constraint: (twenty four) In the formula, This represents the required gas load at time t.

[0063] 4) Combined heat and power (CHP) units: (25) In the formula, and These represent the upper and lower limits of the thermal output of the CHP unit, respectively. and These represent the downhill ramp power and uphill ramp power of the CHP unit, respectively. This indicates the fixed hot spot ratio of the CHP unit.

[0064] 5) Thermal power units: (26) In the formula, and These represent the upper and lower limits of the TPU unit's output, respectively. and These represent the downhill ramp power and uphill ramp power of the TPU unit, respectively.

[0065] 6) Wind turbine units: (27) In the formula, This indicates the maximum output of photovoltaic power.

[0066] 7) Photovoltaic units: (28) 8) Electric to gas conversion: The electro-gas conversion process includes two steps: electro-hydrogen production and methanation. The electro-hydrogen production process is as follows: (29) In the formula, This represents the volume of hydrogen produced by the electro-hydrogen conversion at time t. Indicates the efficiency of the electrolytic cell; Indicates the energy consumption for water electrolysis. ; This indicates the maximum electrical energy that the electrolytic cell can consume.

[0067] Methanation process: (30) In the formula, This represents the volume of methane gas produced by the electro-gas conversion at time t. Indicates methanation efficiency; This indicates the molar ratio of hydrogen to methane. This indicates the maximum volume of gas that the methanation equipment can produce.

[0068] 9) Electric heat pump: (31) In the formula, and These represent the upper and lower limits of the output power of the electric heat pump, respectively. This represents the electrothermal conversion coefficient of an electric heat pump.

[0069] 10) Gas-fired boilers: (32) In the formula, Indicates boiler thermal efficiency; This indicates the lower heating value of the gas, 10.5. ; and These represent the upper and lower limits of the output of the gas-fired boiler.

[0070] Decision variables: Decision variables include CHP unit output. and TPU unit output AE absorbs electrical energy HP thermal output and GB heat output .

[0071] Optimization algorithm: MOMA algorithm based on PPO algorithm.

[0072] This application takes a comprehensive energy system consisting of a 100MW wind farm, a 100MW photovoltaic unit, three thermal power units, two combined heat and power units, a 20MW electric heat pump (electrothermal conversion coefficient of 3.5), a methanation unit with a hydrogen / methane molar ratio of 4 (methanation efficiency of 0.85), a 20MW gas-fired boiler (boiler efficiency of 0.9), and a 30MW electrolyzer [electrolyzer efficiency of 70%, water electrolysis energy consumption of 2.93 kWh / Nm3] as an example. Specific parameters of the thermal power units and combined heat and power units are shown in Table 1. The dispatch cycle is 24 hours, the unit dispatch interval is 1 hour, and the hydrogen sales price is 2.7 yuan / (N·m3). Electricity load data, heat load data, wind power forecast data, and electricity price data are shown in Table 2. The numerical results of the objective function after comprehensive decision-making are shown in Table 3.

[0073] Table 1 Parameters of Thermal Power Units and Cogeneration Units Table 2 Input Data Table 3 Optimization Target Results Substituting the proposed integrated energy system optimization scheduling method based on the PPO algorithm and the MOMA algorithm into the simulation, and after comprehensive decision-making on the Pareto curve using the entropy weight method-TOPSIS, the compromise optimization result of the above embodiment is obtained: the optimal economic cost of the PPO-MOM algorithm is 1.915 × 10⁻⁶. 6 The optimal carbon emission is 1.931 × 10⁻⁶. 6 Nm3. The results obtained using HV and IGD evaluation indicators were 0.01578 and 0.1320, respectively.

[0074] Results Analysis: 1. The PPO-MOMA algorithm demonstrates superior performance in solving the economic-low-carbon multi-objective optimal scheduling model of integrated energy systems. The PPO-MOMA algorithm minimizes carbon emissions and makes fuller use of local resources, reducing operating costs. 2. The proposed PPO-MOMA algorithm effectively solves this multi-objective optimal scheduling model. The superiority of the proposed algorithm is verified through comparisons using Pareto fronts, convergence curves of different objective functions, and algorithm evaluation metrics.

[0075] The overall beneficial effects of this method are as follows: 1. A MOMA algorithm based on the PPO algorithm is designed to solve the optimization scheduling model. The PPO algorithm, through dynamic training of the crossover and mutation factors, can find the optimal global search population, verifying the applicability of the algorithm to the model; 2. The Pareto front-based intelligent optimization method can effectively overcome the difficulties of traditional mathematical methods in terms of solution accuracy and efficiency, demonstrating stronger adaptability and optimization capabilities. As an emerging reinforcement learning method, the PPO algorithm, combined with a multi-objective intelligent optimization strategy, improves learning efficiency and convergence performance in high-dimensional policy spaces while maintaining global search capabilities, showing broad application prospects.

[0076] The above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions implemented in the present invention, and should all be covered within the protection scope of the present invention.

Claims

1. A PPO-MOMA based multi-objective optimization scheduling method for integrated energy systems, characterized in that, Includes the following steps: Establish a multi-objective optimization scheduling model for integrated energy systems that simultaneously aims at minimizing operating costs and carbon emissions, and configure constraints for system power balance, unit output, and equipment operation. An initial population is generated and constrained according to particle coding rules and scheduling principles. The non-dominated solution set is filtered by Pareto non-dominated sorting and stored in an external reserve set. The multi-objective cultural gene algorithm PPO-MOMA based on proximal strategy optimization is used to solve the multi-objective optimization scheduling model. The PPO algorithm takes the state variable representing the current evolutionary state as input and dynamically outputs the crossover probability and mutation probability. Based on the crossover probability and mutation probability, differential evolution crossover and adaptive mutation operations are performed on the population to generate a new population and update the external reserve set. Using the hypervolume index as the reward signal for the PPO algorithm, the parameters are iteratively updated through the PPO policy network and value network, and the current crossover probability, mutation probability and corresponding population are output. The current population is sorted by non-dominated order and crowding selection to obtain a new generation population, and a local search is performed on the new generation population. The steps of solving the problem to the local search are executed iteratively, and the Pareto optimal solution set in the external reserve set is output. By using the entropy-weighted TOPSIS method to make comprehensive decisions on the Pareto optimal solution set, the optimal integrated energy system scheduling scheme that balances operating costs and carbon emissions is selected.

2. The PPO-MOMA-based integrated energy system multi-objective optimization scheduling method of claim 1, wherein, The state variable of the PPO algorithm is a two-dimensional vector containing the current crossover probability and mutation probability. The action space is set with four types of discrete adjustment actions, which correspond to the parameter adjustment methods for increasing or decreasing the crossover probability and increasing or decreasing the mutation probability, respectively.

3. The PPO-MOMA-based integrated energy system multi-objective optimization scheduling method of claim 1, wherein, The differential evolution crossover uses a scaling factor F and a crossover probability CR to generate a mutant population VaPop, and then generates a new population through binary crossover.

4. The PPO-MOMA-based integrated energy system multi-objective optimization scheduling method of claim 1, wherein, When performing differential evolution crossover and adaptive mutation on the population using the crossover probability and mutation probability, the mutation amplitude σ is dynamically adjusted according to the current iteration number gen and the maximum iteration number setgen, and mutation is completed by superimposing Gaussian noise NS on the original population.

5. The PPO-MOMA-based integrated energy system multi-objective optimization scheduling method of claim 1, wherein, When performing non-dominated ordination and crowding selection, the mutated population is merged with the original population, and non-dominated ordination plus crowding descending selection of NSGA-II is used to obtain a new generation of population with unchanged population size.

6. The PPO-MOMA-based integrated energy system multi-objective optimization scheduling method of claim 1, wherein, The entropy-weighted TOPSIS method includes: normalizing operating costs and carbon emissions; determining objective weights using information entropy; calculating the weighted Euclidean distance between each scheme and the positive and negative ideal solutions, and selecting the scheme with the largest relative proximity as the final scheduling scheme.

7. The multi-objective optimization scheduling method for integrated energy systems based on PPO-MOMA as described in claim 1, characterized in that, The integrated energy system includes thermal power units, combined heat and power, wind power, photovoltaic power, electricity-to-gas conversion, electric heat pumps, and gas boilers.

8. The multi-objective optimization scheduling method for integrated energy systems based on PPO-MOMA as described in claim 1, characterized in that, The operating cost objective function simultaneously takes into account the operating costs of CHP units, TPU units, electricity purchased from the grid, gas purchased from the gas grid, as well as AE equipment, HP equipment, and GB equipment.

9. The multi-objective optimization scheduling method for an integrated energy system based on PPO-MOMA as described in claim 1, characterized in that, The carbon emission objective function simultaneously takes into account the carbon emissions generated by CHP units, TPU units, electricity purchased from the external grid, gas boilers, and the electricity-to-gas conversion process.

10. The multi-objective optimization scheduling method for integrated energy systems based on PPO-MOMA as described in claim 1, characterized in that, The integrated energy system scheduling process simultaneously satisfies constraints on power balance, heat balance, and gas volume flow balance; it also satisfies the upper and lower limits of output and ramp-up constraints for CHP units and TPU units, as well as the upper and lower limits of operating output constraints for wind turbines, photovoltaic units, electric-to-gas devices, electric heat pumps, and gas boilers.