Multi-energy optimization scheduling system integrated with multi-parameter system

By integrating the five key indicators in the multi-energy coproduction system into the Energy System Performance Index (ESP), and combining the DRL-MOEA algorithm, the performance of the multi-energy optimization scheduling system is maximized, solving the problems of low energy utilization and high redundancy costs in the existing technology, and improving the system's comprehensive performance and sustainable development capabilities.

CN120235293APending Publication Date: 2025-07-01HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510305409.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to deeply integrate multi-energy collaborative scheduling, multi-dimensional dynamic evaluation and intelligent algorithm optimization in multi-energy co-production systems, resulting in low energy utilization, high redundancy costs, and difficult to adapt to dynamic supply and demand changes.

Method used

A multi-energy optimization scheduling system is adopted that integrates multi-parameter systems. By integrating five indicators of efficiency, economic feasibility, elasticity, evolution potential and environmental impact into a single quantitative indicator - Energy System Performance Index (ESP), and combining the DRL-MOEA fusion algorithm to determine key variables to maximize the performance of the energy system.

Benefits of technology

It significantly improves the comprehensive performance and sustainable development capabilities of the energy system, and achieves the optimal allocation of energy supply through intelligent scheduling strategies, reduces energy waste, improves energy utilization efficiency, and adapts to the dynamic changes in energy supply and demand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235293A_ABST
    Figure CN120235293A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-energy optimal scheduling system integrated with a multi-parameter system, and aims to realize reduction of power operation cost and optimal configuration of energy supply through intelligent management and scheduling strategies. The technology comprises a multi-energy co-production system module and a 5E comprehensive evaluation module, the multi-energy co-production system module is composed of an electric power unit, a thermal power unit and a refrigerating unit, and various energy requirements are met; and five dimensions of efficiency, economic feasibility, elasticity, evolution potential and emission are fused into a single quantitative index, namely an energy system performance index (ESP). Through fusion of deep reinforcement learning (DRL) and a multi-objective evolutionary algorithm (MOEAs), namely, a DRL-MOEA algorithm, dynamic optimization of key parameters is realized so as to maximize the ESP. According to the technology, the economical efficiency, the environmental protection property, the adaptability and the flexibility of an energy system are improved, and sustainable development is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart grid and sustainable energy scheduling, and particularly to a multi-energy optimal scheduling system integrating a multi-parameter system. Background Art

[0002] With the rapid growth of energy demand and the promotion of sustainable development goals, multi-energy combined production systems have gradually become a research hotspot in the energy field due to their advantages in the coordinated supply of electricity, heat, and cooling energy. Traditional energy systems mostly adopt independent optimization of single energy forms. For example, the power system focuses on power generation efficiency, the thermal system pays attention to heat energy distribution, and the refrigeration system relies on the operation of independent equipment. This fragmented optimization mode leads to low energy utilization efficiency, high redundant costs, and difficulty in adapting to dynamic supply and demand changes.

[0003] Existing scheduling technologies often rely on linear programming, heuristic algorithms, or single-objective optimization models, and have significant limitations and error phenomena in dealing with multi-energy coupling, real-time fluctuations, and multi-dimensional performance balance evaluation. There is an urgent need for an innovative scheduling system that can deeply integrate multi-energy coordinated scheduling, multi-dimensional dynamic evaluation, and intelligent algorithm optimization to improve the comprehensive performance and sustainable development ability of the energy system. Summary of the Invention

[0004] Object of the Invention: To solve the problems mentioned in the background art, the present invention discloses a multi-energy optimal scheduling system integrating a multi-parameter system. By fusing five indicators of Efficiency, Economic Viability, Elasticity, Evolutionary Potential, and Emission into a single quantitative indicator, and combining the DRL-MOEA fusion algorithm to determine key variables, the energy system performance index ESP is maximized, and the energy system is comprehensively evaluated and optimized.

[0005] Technical Solution:

[0006] The present invention discloses a multi-energy optimal scheduling system integrating a multi-parameter system, which includes a multi-energy combined production system module and a 5E comprehensive evaluation module;

[0007] The multi-energy combined production system module is used to achieve the combined production and supply of electricity, heat, and cooling energy through intelligent scheduling strategies. The module includes a power network group, a thermal network group, and a refrigeration unit;

[0008] The power network group is used to dynamically adjust the power supply and demand balance;

[0009] The thermal network group is used to optimize heat energy distribution and storage;

[0010] The refrigeration unit is used to be driven by electric energy and waste heat to meet the cooling load demand;

[0011] The multi-energy combined production system module provides the balance variables of the power network group, the thermal network group and the refrigeration unit to the 5E comprehensive evaluation module;

[0012] The 5E comprehensive evaluation module is used to integrate five dimensions of efficiency, economic feasibility, flexibility, evolution potential and emissions into a single quantitative index: the Energy System Performance Index ESP, construct a Maximize ESP objective function by combining balance variables, and obtain the maximum value of ESP;

[0013] Construct a DRL-MOEA fusion algorithm: use the DRL agent to interact and learn in real time to adjust the operation strategy, maximize the cumulative reward to improve ESP; MOEA generates a multi-objective Pareto front, providing an optimal trade-off solution set for the 5E dimension; DRL and MOEA cooperate to optimize and dynamically adjust the balance variables provided by the multi-energy combined production system module.

[0014] Furthermore, the electric energy balance of the power network group is expressed as follows:

[0015] E(t)+E ec =E de +εE plant +E gen -(1-ε)E out (t)

[0016] In the formula, ε represents the state variable interacting with the power grid, that is, ε = 1 means that the multi-energy combined production system purchases electric power from the power grid at time t, and ε = 0 means that the system sells electric power to the power grid at time t; E out refers to the electric energy output at time t;

[0017] The heat energy balance of the thermal network group is expressed as follows:

[0018] H(t)+Q ac (t)+Q ex (t)=Q re (t)+Q hn (t)

[0019] In the formula, Q ac is the input power of the absorption chiller; Q ex is the heat loss of the system; Q re represents the waste heat generated by the generator set; Q hn represents the heat generated in the thermal network;

[0020] The cooling balance of the refrigeration unit is expressed as follows:

[0021] C(t)=C ec (t)+Cac (t)

[0022] Wherein, C ac and C ec are the cooling outputs of the absorption chiller and the electric chiller, respectively;

[0023] The power network group is the core of power supply. Select E gen as the key variable. The thermal network group is the main source of heat energy supply. The heat energy output Q b and Q hn from the boiler are the key variables. The refrigeration outputs C ac and C ec of the refrigeration unit are the key variables, which directly affect the satisfaction of the cooling load. The variables E gen , Q b , Q hn , C ac and C ec are input into the 5E comprehensive evaluation module.

[0024] Furthermore, the construction of the Maximize ESP objective function is specifically as follows:

[0025] Compare the multi-objective functions in terms of efficiency, economic feasibility, flexibility, evolution potential, and emissions between the poly-generation system and the separate production system to improve the comprehensive performance of the system:

[0026] The efficiency dimension considers the degree of minimization of losses during the energy conversion and distribution processes in the system, which can be expressed as:

[0027]

[0028] Wherein, η i is the efficiency of the i-th device, E out,i is the energy output of the i-th device, E in,i is the energy input of the i-th device, and n is the number of energy conversion devices in the system;

[0029] Economic feasibility involves evaluating the cost-benefit ratio of the energy system, including the initial investment, operation and maintenance costs, and the expected economic return, which can be expressed as:

[0030]

[0031] Wherein, R t is the revenue in the t-th year, C t is the cost in the t-th year, r is the discount rate, and T is the evaluation period;

[0032] The flexibility dimension refers to the ability of the system to maintain stable operation in the face of supply-demand fluctuations, price changes, or external shocks, which can be expressed as:

[0033]

[0034] Wherein, S t is the supply at the t-th moment, D t is the demand at the t-th moment, and T is the total number of moments;

[0035] The evolution potential refers to the ability of a system or technology to develop and adapt to new environments or requirements in the future, and can be expressed as:

[0036]

[0037] Wherein, C i,upgrade is the upgradable capacity of the i-th technology, C i is the current capacity of the i-th technology, and m is the number of technology items;

[0038] The environmental impact dimension is an important indicator for evaluating the environmental impact of an energy system, especially the emissions of greenhouse gases and pollutants, and can be expressed as:

[0039]

[0040] Wherein, E j is the total usage of the j-th energy, E j,clean is the usage of clean energy, and p is the number of energy types;

[0041] Combining the above sub-objective functions, the comprehensive objective function is constructed as follows:

[0042] Maximize ESP = ω eff ·E eff + ω eco ·E eco + ω ela ·E ela + ω evo ·E evo + ω emi ·E emi

[0043] Wherein, ω eff , ω eco , ω ela , ω evo , ω emi are the weight factors of efficiency, economic feasibility, elasticity, evolution potential, and environmental impact respectively, and ω eff = ω eco = ω ela = ω evo = ω emi = 0.2.

[0044] Furthermore, the DRL-MOEA fusion algorithm determines the variables E by fusing deep reinforcement learning (DRL) and multi-objective evolutionary algorithms (MOEAs) through selection. gen , Q b , Q hn , C ac and C ec . The DRL agent is responsible for real-time learning and adjusting operation strategies, while the MOEA is used to generate and optimize the Pareto front among multiple objectives. The DRL-MOEA algorithm realizes the dynamic optimization of key parameters in the polygeneration system through initializing the environment and agent, interactive learning, multi-objective optimization, policy update, data sharing, and final iterative optimization, so as to maximize the energy system performance index ESP. The trained model is deployed into the actual system to achieve real-time scheduling and performance monitoring.

[0045] Furthermore, the policy update is specifically as follows:

[0046] The DRL agent updates the weights of its deep neural network according to the state-action-reward data obtained after executing actions based on the current policy in each iteration, combined with the Pareto front provided by the MOEA. This process involves calculating the TD error of the Q value and using the gradient descent method for optimization to approximate the expected return of taking a specific action in a given state. The goal of policy update is to maximize the long-term cumulative reward while maximizing the ESP comprehensive index. The policy update can be expressed as:

[0047] Q(s t , a t ) ← Q(s t , a t ) + α[r t + γ max a Q(s t+1 , a) - Q(s t , a t )]

[0048] In the formula, Q(s t , a t ) is the Q value of taking action a t in state s t . α is the learning rate, γ is the discount factor, r t is the immediate reward, and max a Q(s t+1 , a) is the maximum Q value updated based on the Pareto front information, reflecting the optimal future expected return.

[0049] Furthermore, the data sharing is specifically as follows:

[0050] The data pairs of state-action-reward-new state, i.e., the quadruples, collected by the DRL agent during its interaction with the poly-generation system are not only used for the training of its own neural network but also stored in the experience replay buffer and shared with the MOEA population, providing feedback information in the actual environment, which helps the MOEA population evaluate and optimize the performance of its individuals in terms of efficiency, economic feasibility, resilience, evolution potential, and emissions. The sharing process can be expressed as:

[0051] D shared = D DRL ∪D MOEA

[0052] In the formula, D shared represents the shared dataset, D DRL is the data in the experience replay buffer of the DRL agent, and D MOEA is the historical performance data of the individuals in the MOEA population.

[0053] Furthermore, the DRL-MOEA algorithm introduces an enhanced multi-stage optimization strategy as follows:

[0054] By monitoring the change trend of ESP in real time and the difference from the preset target, an adaptive weight adjustment mechanism is introduced to dynamically adjust the weights of efficiency, economic feasibility, resilience, evolution potential, and environmental impact in ESP. The weight update formula is as follows:

[0055] w i (t + 1) = w i (t) + β·Δw i (t)

[0056] In the formula, w i (t) represents the weight of the i-th objective at time t, β is the adjustment factor, and Δw i (t) is the weight change amount calculated based on the performance feedback. This mechanism enables the algorithm to flexibly adjust the weights of each objective according to the real-time feedback of the system performance to respond to the demands and environmental changes of the energy system at different stages;

[0057] The learning rate of the DRL agent is adaptively adjusted according to the performance feedback of the policy to optimize the learning process. The learning rate update formula is as follows:

[0058] α(t + 1) = α(t)·e γ·ΔESP(t)

[0059] In the formula, α(t) is the learning rate at time t, γ is the adjustment coefficient, and ΔESP(t) is the change amount of ESP at time t. This adaptive learning rate scheduling helps the algorithm reduce the learning rate when the improvement of ESP is small or shows a decline, avoiding over-adjustment; while increasing the learning rate when the improvement is large, accelerating the learning process.

[0060] Adopt the elite strategy to retain. In the MOEA population, a part of the best-performing individuals are retained as "elites" and directly inherited to the next generation to ensure that excellent strategies are retained. The selection of elite individuals is based on their positions on the Pareto front and ESP values;

[0061] Introduce a multi-scale evaluation mechanism. Combine the short-term and long-term ESP performances, as well as the robustness of the strategy, design a composite evaluation index, establish a real-time feedback mechanism, and feedback the real-time performance data of the DRL agent and the MOEA population into the weight adjustment and learning rate update to form a closed-loop optimization system, so that the performance feedback of the DRL agent and the MOEA population can be used to adjust the algorithm parameters in real time to continuously optimize ESP.

[0062] Beneficial effects:

[0063] 1. Through the integrated intelligent management and scheduling strategy, the present invention significantly reduces the power operation cost. Through the collaborative work of the poly-generation system module and the 5E comprehensive evaluation module, the optimal configuration of energy supply is realized, effectively reducing energy waste and improving energy utilization efficiency, thus achieving significant improvement in the dimension of economic feasibility.

[0064] 2. The present invention integrates five key dimensions of efficiency, economic viability, elasticity, evolutionary potential, and emission into a single quantitative index - the energy system performance index (ESP), and further optimizes the energy system scheduling efficiency by maximizing ESP, reducing the energy redundancy cost.

[0065] 3. The present invention combines DRL and MOEAs to construct a fusion algorithm DRL-MOEA, realizes the synergistic effect of real-time learning and global optimization, realizes the dynamic optimization of key variables in the poly-generation system, and finally deploys the trained model to the actual system to realize real-time scheduling and performance monitoring to adapt to the dynamic changes of energy supply and demand and improve the comprehensive performance of the system.

[0066] 4. Through the enhanced multi-stage optimization strategy, including dynamic weight adjustment, adaptive learning rate, elite strategy retention, multi-scale evaluation, and feedback loop optimization, the present invention significantly improves the ability of the DRL-MOEA algorithm to adapt to the changes of the energy system in a dynamic environment, effectively balances the trade-off between multiple objectives, and achieves the best result of ESP. This strategy not only enhances the adaptability and robustness of the algorithm, but also improves its effectiveness and reliability in practical applications. Description of the drawings

[0067] Figure 1 is the energy flow block diagram of the multi-energy co-production system of the present invention;

[0068] Figure 2 is the flow chart of the DRL-MOEA fusion algorithm. Specific embodiments

[0069] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0070] As Figure 1 - Figure 2 shown, the present invention discloses a multi-energy optimal scheduling system integrating a multi-parameter system, which includes a multi-energy co-production system module and a 5E comprehensive evaluation module;

[0071] The multi-energy co-production system module is used to achieve the co-production and co-supply of electricity, heat and cooling energy through an intelligent scheduling strategy. The module includes an electric power network group, a heat network group and a refrigeration unit;

[0072] The electric power network group is composed of a power plant, distributed energy, electric load, electric energy storage and a generating unit. The generated electricity can be used to meet the electricity demand of users and can drive an electric chiller to meet the cooling load of users. The system operates in a grid-connected mode. When redundant power is generated, the electric load is returned to the grid, and the shortage of electricity can also be supplemented by the grid. It is used for dynamically adjusting the balance between power supply and demand;

[0073] The heat network group consists of a boiler, a thermal power plant, a heat load, a heat energy storage, a heat exchanger and an exhaust gas exchanger. The heat energy generated by the boiler and the thermal power plant is distributed to the heat load through the heat network. At the same time, the heat energy storage unit can store the surplus heat energy. The heat exchanger is used to recover the waste heat generated during the power generation process to improve the overall thermal efficiency of the system. The exhaust gas exchanger helps manage the exhaust gas generated during the power generation process, including the cleaning and reuse processes, and is used for optimizing the heat energy distribution and storage;

[0074] The refrigeration unit consists of an electric chiller and an absorption chiller. The waste heat generated by the generating unit, including the jacket water heat and the exhaust heat, is recovered and used to generate cold air and heating in the absorption chiller respectively. The electric chiller directly uses electricity to generate cooling energy to meet the cooling load demand of the building load. The absorption chiller uses heat energy to drive the refrigeration process and can use heat energy more efficiently. It is used to drive with electric energy and waste heat to meet the cooling load demand;

[0075] InFigure 1 Among them, E de and E gen respectively represent the power generation of distributed energy and generating units; E plant is the electric power generated by the power plant flowing into the power grid; E ec is the input power of the electric chiller; E, H, and C respectively represent the user's demands for electricity, heating, and cooling.

[0076] The power balance of the poly-generation system is expressed as follows:

[0077] E(t) + E ec = E de + εE plant + E gen -(1 - ε)E out (t)

[0078] In the formula, ε represents the state variable interacting with the power grid (i.e., ε = 1 means the poly-generation system purchases electricity from the power grid at time t, while ε = 0 means the system sells electricity to the power grid at time t); E out refers to the electric energy output at time t.

[0079] The heat balance of the poly-generation system is expressed as follows:

[0080] H(t) + Q ac (t) + Q ex (t) = Q re (t) + Q hn (t)

[0081] In the formula, Q ac is the input power of the absorption chiller; Q ex is the heat loss of the system; Q re represents the waste heat generated by the generating unit; Q hn represents the heat generated in the heat network.

[0082] The cooling balance of the poly-generation system is expressed as follows:

[0083] C(t) = C ec (t) + C ac (t)

[0084] In the formula, C ac and C ec are respectively the cooling outputs of the absorption chiller and the electric chiller.

[0085] The poly-generation system module provides the balance variables of the power network group, heat network group, and refrigeration unit to the 5E comprehensive evaluation module;

[0086] The 5E comprehensive evaluation module is used to integrate five dimensions, namely efficiency, economic viability, elasticity, evolutionary potential, and emissions, into a single quantitative indicator: the Energy System Performance Index (ESP). By combining balance variables, a Maximize ESP objective function is constructed to obtain the maximum value of ESP.

[0087] Efficiency, which refers to the degree of minimizing losses during the energy conversion and distribution processes in the system. A high-efficiency system can minimize energy losses during conversion and transmission, thereby improving the overall energy utilization efficiency.

[0088] Economic Viability, which involves evaluating the cost-benefit ratio of the energy system, including initial investment, operation and maintenance costs, and expected economic returns. An economically viable system should ensure long-term financial sustainability, that is, the total revenue generated by the system over its life cycle exceeds the total cost. This is usually evaluated through cost-benefit analysis, return on investment, and net present value calculations.

[0089] Elasticity, which refers to the ability of the system to maintain stable operation in the face of supply and demand fluctuations, price changes, or external shocks. A highly elastic system can flexibly adjust its operation mode to adapt to changes in external conditions, such as through energy storage, demand-side response, and management strategies.

[0090] Evolutionary Potential, which refers to the ability of the system or technology to develop and adapt to new environments or demands in the future. A system with high evolutionary potential can be easily upgraded and expanded to adapt to changes in technological progress and market demands, such as integrating renewable energy and smart grid technologies.

[0091] Emission, which is an important indicator for evaluating the environmental impact of the energy system, especially the emissions of greenhouse gases (such as CO2) and pollutants. A low-emission system reduces the negative impact on the environment by adopting clean energy technologies, improving combustion efficiency, and implementing emission control measures, thereby supporting sustainable development goals.

[0092] The Energy System Performance Index (ESP) comprehensively evaluates the performance of the energy system in multiple dimensions, namely efficiency, economic viability, elasticity, evolutionary potential, and emissions, in a quantitative manner. By maximizing ESP, the energy system can be promoted to develop towards higher efficiency, better economy, and stronger environmental sustainability.

[0093] S1: Establish the objective function:

[0094] From Figure 1It can be seen that the generator set is the core of power supply, so E is selected as a key variable. gen Boilers and thermal power plants are the main sources of heat energy supply. Therefore, their heat energy outputs Q and Q are key variables. b and Q hn The refrigeration outputs C and C of electric refrigerators and absorption refrigerators are key parameters because they directly affect the satisfaction of the cooling load. ac and C ec

[0095] Select and consider a multi-objective function in five aspects: Efficiency, Economic Viability, Elasticity, Evolutionary Potential, and Emission. Compare the polygeneration system with the single production system to improve the comprehensive performance of the system.

[0096] The efficiency dimension considers the degree of minimization of losses during the energy conversion and distribution processes of the system, which can be expressed as:

[0097]

[0098] In the formula, η i is the efficiency of the i-th device, E out,i is the energy output of the i-th device, and E in,i is the energy input of the i-th device. n is the number of energy conversion devices in the system.

[0099] Economic viability involves evaluating the cost-benefit ratio of the energy system, including initial investment, operation and maintenance costs, and expected economic returns, which can be expressed as:

[0100]

[0101] In the formula, R t is the revenue in the t-th year, C t is the cost in the t-th year, r is the discount rate, and T is the evaluation period.

[0102] The elasticity dimension refers to the ability of the system to maintain stable operation in the face of supply-demand fluctuations, price changes, or external shocks, which can be expressed as:

[0103]

[0104] In the formula, S t is the supply at the t-th moment, D t is the demand at the t-th moment, and T is the total number of moments.

[0105] ​The evolution potential refers to the ability of a system or technology to develop and adapt to new environments or requirements in the future, which can be expressed as:

[0106]

[0107] where C i,upgrade is the upgradable capacity of the i-th technology, C i is the current capacity of the i-th technology, and m is the number of technology items.

[0108] The environmental impact dimension is an important indicator for evaluating the environmental impact of an energy system, especially the emissions of greenhouse gases and pollutants, which can be expressed as:

[0109]

[0110] where E j is the total usage of the j-th energy source, E j,clean is the usage of clean energy, and p is the number of energy types.

[0111] Therefore, by integrating the above sub-objective functions, the comprehensive objective function is constructed as follows:

[0112] Maximize ESP = ω eff ·E eff + ω eco ·E eco + ω ela ·E ela + ω evo ·E evo + ω emi ·E emi

[0113] where ω eff , ω eco , ω ela , ω evo , ω emi are the weight factors of efficiency, economic feasibility, flexibility, evolution potential, and environmental impact, respectively, and ω eff = ω eco = ω ela = ω evo = ω emi = 0.2.

[0114] S2: Use selection to fuse deep reinforcement learning (DRL) with multi-objective evolutionary algorithms (MOEAs) to determine the variables E gen , Q b , Q hn , C ac and C ecThis fusion algorithm is also known as the DRL-MOEA (Deep Reinforcement Learning with Multi-Objective Evolutionary Algorithms) algorithm. The DRL agent is responsible for real-time learning and adjusting operation strategies, while the MOEAs are used to generate and optimize the Pareto front between multiple objectives, ensuring the best balance in the five E dimensions of efficiency, economic feasibility, resilience, evolution potential, and environmental impact. The DRL-MOEA algorithm realizes the dynamic optimization of key parameters in the poly-generation system through initializing the environment and agent, interactive learning, multi-objective optimization, policy update, data sharing, and final iterative optimization, so as to maximize the Energy System Performance Index (ESP), and finally deploy the trained model to the actual system to achieve real-time scheduling and performance monitoring to adapt to the dynamic changes in energy supply and demand and improve the comprehensive performance of the system.

[0115] S2.1: Initialization: Define the operation environment of the poly-generation system and initialize the population of the Deep Reinforcement Learning (DRL) agent and the Multi-Objective Evolutionary Algorithm (MOEA). The operation environment specifically includes the state space and the action space. The state space is constructed to include the operation parameters of key equipment such as power generation units, boilers, thermal power plants, and refrigeration units, as well as external conditions such as energy demand and price. The action space covers all possible control actions, such as adjusting power generation or heat energy distribution. The following are the mathematical expressions of the state space and the action space:

[0116] S = {s1, s2,..., s n}

[0117] A = {a1, a2,..., a m}

[0118] In the formula, s n represents the nth state of the system, including the operation parameters of all equipment and external conditions; a m represents the mth possible action.

[0119] S2.2: Interactive learning: The DRL agent starts to interact with the actual environment of the poly-generation system, collecting state-action-reward-new state data pairs by executing actions and observing the results. Specifically, the DRL agent selects an action a t according to the current system state s t and the learned probability policy. After executing this action, the system will transfer to a new state s t+1 and generate an immediate reward r t . This reward reflects the impact of the action on the system performance and can be the short-term change based on the Energy System Performance Index (ESP). The reward function can be expressed as:

[0120] r t = f(ESP(s t , a t , s t+1 ))

[0121] where f is a function that calculates the reward based on the state transition and the change in ESP.

[0122] The goal of the DRL agent is to maximize the cumulative reward, i.e., the increase in long-term ESP. Subsequently, the agent uses this empirical data to update its policy by training a deep neural network to approximate the optimal Q-value function, which predicts the expected return of taking a specific action in a given state. This process involves calculating the time difference (TD) error and using the gradient descent method to update the network weights to reduce the error and improve the performance of the policy. In this way, the DRL agent gradually learns how to maximize ESP in the polygeneration system.

[0123] S2.3: Multi-objective optimization: The MOEA searches the solution space and generates the Pareto front by simulating natural selection and genetic mechanisms such as crossover and mutation, so as to achieve the optimal trade-off in the 5E dimension. The Pareto front consists of those individuals for which no other solution in the objective space can improve all objectives simultaneously, i.e., a solution can be considered Pareto-optimal only if it improves at least one objective without deteriorating the others. The MOEA approximates the true Pareto-optimal solution by maintaining a non-dominated solution set (Pareto front), and these solutions provide a view of the trade-off between different objectives, offering diverse choices for decision-makers and ensuring the maximization of ESP under dynamic changes. The MOEA evaluates the fitness of each individual in the population, which is calculated based on the following multi-objective function:

[0124] Minimize F(x) = (f1(x), f2(x),..., f k (x))

[0125] where F(x) represents the multi-objective fitness vector of individual x, and f1(x), f2(x),..., f k (x) correspond to the quantization indicators in the 5E dimension respectively.

[0126] S2.4: Policy Update: The DRL agent updates the weights of its deep neural network based on the state-action-reward data obtained after executing actions according to the current policy in each iteration, combined with the Pareto front provided by the MOEA. This process involves calculating the TD error of the Q-value and using gradient descent for optimization to approximate the expected return of taking a specific action in a given state. The goal of policy update is to maximize the long-term cumulative reward while maximizing the ESP comprehensive metric. The policy update can be expressed as:

[0127] Q(s t ,a t )←Q(s t ,a t )+α[r t +γmax a Q(s t+1 ,a)-Q(s t ,a t )]

[0128] In the formula, Q(s t ,a t ) is the Q-value of taking action a t under state s t , α is the learning rate, γ is the discount factor, r t is the immediate reward, and max a Q(s t+1 ,a) is the maximum Q-value updated based on the Pareto front information, reflecting the optimal future expected return.

[0129] S2.5: Data Sharing: The data pairs of state-action-reward-new state (i.e., quadruples) collected by the DRL agent during the interaction with the polygeneration system are not only used for the training of its own neural network but also stored in the experience replay buffer and shared with the MOEA population. These data provide feedback information in the actual environment, helping the MOEA population evaluate and optimize the fitness of its individuals, especially the performance in the 5E dimension. This sharing process can be expressed as:

[0130] D shared =D DRL ∪D MOEA

[0131] In the formula, D shared represents the shared dataset, D DRL is the data in the experience replay buffer of the DRL agent, and D MOEA is the historical performance data of individuals in the MOEA population.

[0132] S2.6: Iterative Optimization: Repeat steps S2.2 to S2.5, and continuously update and optimize the strategy based on new data and experience. The DRL agent continuously interacts with the polygeneration system, collects state-action-reward data, and uses this data to update its strategy. Meanwhile, the MOEA runs a multi-objective optimization process to generate the Pareto front, and the DRL agent further adjusts its strategy based on this front information to maximize the ESP.

[0133] There may be calculation errors in the process of calculating the energy system performance index by the DRL-MOEA algorithm, making the result not optimal. Therefore, the algorithm needs to be optimized. Thus, an enhanced multi-stage optimization strategy is introduced, which combines the concepts of dynamic weight adjustment and adaptive learning rate.

[0134] First, by real-time monitoring the change trend of the ESP and the difference from the preset target, an adaptive weight adjustment mechanism is introduced to dynamically adjust the weights of efficiency, economic feasibility, flexibility, evolution potential, and environmental impact in the ESP. The weight update formula is as follows:

[0135] w i (t + 1) = w i (t) + β · Δw i (t)

[0136] In the formula, w i (t) represents the weight of the i-th objective at time t, β is the adjustment factor, and Δw i (t) is the weight change calculated based on the performance feedback. This mechanism enables the algorithm to flexibly adjust the weights of each objective according to the real-time feedback of the system performance to respond to the demands and environmental changes of the energy system at different stages.

[0137] Meanwhile, the learning rate of the DRL agent is adaptively adjusted according to the performance feedback of the strategy to optimize the learning process. The learning rate update formula is as follows:

[0138] α(t + 1) = α(t) · e γ·ΔESP(t)

[0139] In the formula, α(t) is the learning rate at time t, γ is the adjustment coefficient, and ΔESP(t) is the change in ESP at time t. This adaptive learning rate scheduling helps the algorithm to reduce the learning rate when the improvement of the ESP is small or there is a decline, avoiding over-adjustment; while increasing the learning rate when the improvement is large to accelerate the learning process.

[0140] In addition, the strategy also includes elite strategy retention, that is, a part of the best-performing individuals in the MOEA population are retained as "elites" and directly inherited to the next generation to ensure that excellent strategies are retained. The selection of elite individuals is based on their positions on the Pareto front and ESP values, which helps the algorithm maintain the diversity and quality of solutions in multi-objective optimization.

[0141] To comprehensively evaluate the algorithm performance, a multi-scale evaluation mechanism is introduced. Combining the short-term and long-term ESP performances and the robustness of the strategy, composite evaluation indicators are designed. This mechanism not only focuses on the immediate ESP improvement but also on the stability and trend of ESP, providing comprehensive performance feedback for the algorithm.

[0142] Finally, by establishing a real-time feedback mechanism, the real-time performance data of the DRL agent and the MOEA population are fed back into the weight adjustment and learning rate update to form a closed-loop optimization system. This system enables the performance feedback of the DRL agent and the MOEA population to be used for real-time adjustment of algorithm parameters to continuously optimize ESP.

[0143] The above description of the embodiments enables those skilled in the art to implement or use the present invention. Various modifications to the embodiments will be obvious to those skilled in the art. The general principles of the present invention can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention should not be limited to the embodiments shown herein but should cover the widest range consistent with the principles and novel features disclosed in the present invention.

Claims

1. A multi-energy optimization dispatching system integrating a multi-parameter system, the system comprising a multi-energy cogeneration system module and a 5E comprehensive evaluation module; The multi-energy cogeneration system module is used to realize the cogeneration and supply of electricity, heat and refrigeration energy through intelligent scheduling strategies, and the module includes an electric power network group, a thermal network group and a refrigeration unit; The power network group is used to dynamically adjust the balance between power supply and demand; The thermal network group is used to optimize the distribution and storage of thermal energy; The refrigeration unit is used to drive with electric energy and waste heat to meet the cooling load demand; The multi-energy cogeneration system module provides the balance variables of the power network group, the thermal network group and the refrigeration unit to the 5E comprehensive evaluation module; The 5E comprehensive evaluation module is used to integrate the five dimensions of efficiency, economic feasibility, resilience, evolution potential and emissions into a single quantitative indicator: the energy system performance index ESP, which combines the balance variables to construct the Maximize ESP objective function and obtain the maximum value of ESP; Construct a DRL-MOEA fusion algorithm: use DRL agents to interactively learn in real time to adjust the operation strategy and maximize the cumulative reward to improve ESP; MOEA generates a multi-objective Pareto frontier and provides an optimal trade-off solution set in 5E dimensions; DRL and MOEA collaborate to optimize and dynamically adjust the balance variables provided by the multi-energy cogeneration system module.

2. The multi-energy optimization scheduling system integrating multi-parameter systems according to claim 1 is characterized in that: The energy balance of the power network group is expressed as follows: E(t)+E ec JE de +εE plant +E gen -(1-ε)E out (t) In the formula, ε represents the state variable that interacts with the power grid, that is, ε = 1 means that the multi-energy cogeneration system purchases electricity from the power grid at time t, and ε = 0 means that the system sells electricity to the power grid at time t; E out Refers to the electrical energy output at time t; The thermal energy balance of the thermal network group is expressed as follows: H(t)+Q ac (t)+Q ex (t)=Q re (t)+Q hn (t) In the formula, Q ac is the input power of the absorption chiller; Q ex is the heat loss of the system; Q re Indicates the waste heat generated by the generator set; Q hn Represents the heat generated in the thermal network; The cooling balance of the refrigeration unit is expressed as follows: C(t)=C ec (t)+C ac (t) In the formula, C ac and C ec are the cooling outputs of the absorption chiller and the electric chiller, respectively; The power grid group is the core of power supply. gen The thermal network group is the main source of heat supply, and the heat output Q generated by the boiler is b and Q hn is the key variable, the cooling output C of the refrigeration unit ac and C ec It is a key variable that directly affects the satisfaction of cooling load. Variable E gen , Q b , Q hn , C ac and C ec Enter the 5E comprehensive evaluation module.

3. The multi-energy optimization scheduling system integrating multi-parameter systems according to claim 1 is characterized in that: The MaximizeESP objective function is constructed as follows: The multi-objective function of five aspects, namely efficiency, economic feasibility, flexibility, evolution potential and emissions, is used to compare the multi-energy cogeneration system with the individual production system to improve the overall performance of the system: The efficiency dimension considers the degree to which the system minimizes losses during energy conversion and distribution, which can be expressed as: Where η i is the efficiency of the ith device, E out,i is the energy output of the ith device, E in,i is the energy input of the ith device, n is the number of energy conversion devices in the system; Economic feasibility involves evaluating the cost-benefit ratio of an energy system, including initial investment, operation and maintenance costs, and expected economic returns, which can be expressed as: In the formula, R t is the return in year t, C t is the cost in year t, r is the discount rate, and T is the evaluation period; The elasticity dimension refers to the system's ability to maintain stable operation in the face of supply and demand fluctuations, price changes, or external shocks, which can be expressed as: In the formula, S t is the supply at time t, D t is the demand at time t, T is the total number of moments; Evolution potential refers to the ability of a system or technology to develop and adapt to new environments or requirements in the future, which can be expressed as: In the formula, C i,upgrade is the scalable capacity of the i-th technology, C i is the current capacity of the i-th technology, and m is the number of technologies; The environmental impact dimension is an important indicator for evaluating the environmental impact of energy systems, especially the emissions of greenhouse gases and pollutants, which can be expressed as: In the formula, E j is the total usage of the jth energy source, E j,clean is the amount of clean energy used, p is the number of energy types; Combining the above sub-objective functions, the comprehensive objective function is constructed as follows: Maximize ESP=ω eff ·E eff +oh eco ·E eco +oh ela ·E ela +oh evo ·E evo +oh emi ·E emi In the formula, ω eff ,ω eco ,ω ela ,ω evo ,ω emi are the weighting factors for efficiency, economic feasibility, resilience, evolution potential and environmental impact, and ω eff =ω eco =ω ela =ω evo =ω emi =0.

2.

4. The multi-energy optimization scheduling system integrating multi-parameter systems according to claim 2 or 3, characterized in that: The DRL-MOEA fusion algorithm uses the selection to fuse deep reinforcement learning DRL with multi-objective evolutionary algorithm MOEAs to determine the variable E gen , Q b , Q hn , C ac and C ec , the DRL agent is responsible for real-time learning and adjusting the operation strategy, while MOEA is used to generate and optimize the Pareto frontier between multiple objectives. The DRL-MOEA algorithm realizes the dynamic optimization of key parameters in the multi-energy cogeneration system through initializing the environment and agents, interactive learning, multi-objective optimization, strategy updating, data sharing and final iterative optimization to maximize the energy system performance index ESP, and deploy the trained model to the actual system to achieve real-time scheduling and performance monitoring.

5. The multi-energy optimization scheduling system integrating multi-parameter systems according to claim 4 is characterized in that: The strategy update is as follows: The DRL agent updates the weights of its deep neural network based on the state-action-reward data obtained after executing actions according to the current strategy in each iteration, combined with the Pareto frontier provided by MOEA. This process involves calculating the TD error of the Q value and optimizing it using the gradient descent method to approximate the expected return of taking a specific action in a given state. The goal of the strategy update is to maximize the long-term cumulative reward while maximizing the ESP comprehensive indicator. The strategy update can be expressed as: Q(s t ,a t )←Q(s t ,a t )+α[r t +γmax a Q(s t+1 ,a)-Q(s t ,a t )] In the formula, Q(s t ,a t ) is the state s t Take action a t Q value, α is the learning rate, γ is the discount factor, r t is an immediate reward, and max a Q(s t+1 ,a) is the maximum Q value after updating the Pareto frontier information, reflecting the optimal expected return in the future.

6. The multi-energy optimization scheduling system integrating multi-parameter systems according to claim 4 is characterized in that: The data sharing is as follows: The state-action-reward-new-state data pairs, i.e., quadruples, collected by the DRL agent during the interaction with the multi-energy cogeneration system are not only used for the training of its own neural network, but also stored in the experience replay buffer and shared with the MOEA population to provide feedback information in the actual environment, which helps the MOEA population to evaluate and optimize the fitness of its individuals in terms of efficiency, economic feasibility, resilience, evolutionary potential, and emission dimensions. The sharing process can be expressed as: D shared =D DRL ∪D MOEA Where D shared represents the shared dataset, D DRL is the data in the experience replay buffer of the DRL agent, and D MOEA It is the historical performance data of individuals in the MOEA population.

7. The multi-energy optimization scheduling system integrating multi-parameter systems according to claim 4 is characterized in that: The DRL-MOEA algorithm introduces an enhanced multi-stage optimization strategy as follows: By real-time monitoring of the changing trend of ESP and the difference from the preset target, an adaptive weight adjustment mechanism is introduced to dynamically adjust the weights of efficiency, economic feasibility, elasticity, evolution potential and environmental impact in ESP. The weight update formula is as follows: w i (t+1)=w i (t)+β·Δw i (t) In the formula, w i (t) represents the weight of the i-th target at time t, β is the adjustment factor, Δw i (t) The weight change calculated based on performance feedback. This mechanism enables the algorithm to flexibly adjust the weights of each objective according to the real-time feedback of system performance to respond to the needs and environmental changes of the energy system at different stages; The learning rate of the DRL agent is adaptively adjusted based on the performance feedback of the strategy to optimize the learning process. The learning rate update formula is as follows: α(t+1)=α(t)·e γ·ΔESP(t) Where α(t) is the learning rate at time t, γ is the adjustment coefficient, and ΔESP(t) is the change in ESP at time t. This adaptive learning rate scheduling helps the algorithm reduce the learning rate when the ESP improvement is small or decreases to avoid over-adjustment; and increase the learning rate when the improvement is large to accelerate the learning process. Adopt elite strategy retention, retain a part of the best performing individuals in the MOEA population as "elites", and directly pass them on to the next generation to ensure that excellent strategies are retained. The selection of elite individuals is based on their position on the Pareto frontier and ESP value; A multi-scale evaluation mechanism is introduced, combining the short-term and long-term ESP performance and the robustness of the strategy, a composite evaluation indicator is designed, and a real-time feedback mechanism is established to feed back the real-time performance data of the DRL agent and MOEA population into the weight adjustment and learning rate update, forming a closed-loop optimization system. The performance feedback of the DRL agent and MOEA population can be used to adjust the algorithm parameters in real time to continuously optimize ESP.