Multi-agent optimal compromise scheduling method for multi-agent hydrogen-containing integrated energy system
By combining the construction of a multi-agent reinforcement learning model with the approximate ideal solution sorting method, the collaborative optimization scheduling problem of a multi-agent hydrogen-containing integrated energy system was solved, efficient coordination among agents and maximization of exergy efficiency were achieved, and the quality and flexibility of scheduling decisions were improved.
Patent Information
- Application Number
- CN202510950958.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Multi-agent hydrogen-containing integrated energy systems face conflicts of interest, information asymmetry, and uncertainty in energy supply and demand in the dynamic coupling of heterogeneous energy sources and multi-agent interactive collaborative optimization scheduling. Existing optimization algorithms are insufficient in dealing with complexity and uncertainty, making system collaborative optimization difficult.
A multi-agent reinforcement learning model is constructed, including hydrogen-containing integrated energy service providers, electricity load aggregators, heat load aggregators, and hydrogen load aggregators. Combined with the multi-agent proximal policy optimization algorithm (MAPPO) and the method of ranking by near-ideal solutions (TOPSIS), efficient coordination and optimal compromise scheduling among agents are achieved through an orderly decision-making mechanism and parallel operation of multiple schemes.
It improves the scheduling decision-making quality and optimization efficiency of the multi-agent hydrogen-containing integrated energy system, solves the problems of gradient instability and slow convergence, and achieves more flexible energy allocation and maximizes system exergy efficiency.
Smart Images

Figure CN120450394B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of integrated energy scheduling, and relates to a multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system. Background Art
[0002] The multi-agent hydrogen-containing integrated energy system (HIES), comprising hydrogen-containing integrated energy providers (HIEPs), electric load aggregators (ELAs), thermal load aggregators (TLAs), and hydrogen load aggregators (HLAs), is a new energy supply model that integrates multiple energy types and operating entities. Due to its multi-energy complementarity, efficient conversion, and flexible scheduling capabilities, it is considered a key path to promoting the green energy transition. However, the optimal scheduling of a multi-agent hydrogen-containing integrated energy system faces numerous challenges in a multi-agent scenario with the coexistence of multiple heterogeneous energy carriers. Conflicting interests among different entities, information asymmetry, and uncertainty in energy supply and demand make the coordinated optimization of the system a typical complex decision-making problem with multiple objectives and constraints.
[0003] Existing research on multi-agent hydrogen-based integrated energy systems mostly uses game theory to construct a framework for distributing benefits among the various stakeholders. This multi-agent game model resolves conflicts of interest while achieving economic operation of the integrated energy system. However, these approaches primarily focus on optimizing traditional energy scenarios and utilize traditional optimization algorithms such as particle swarm optimization, heuristic algorithms, and mixed integer programming. The introduction of hydrogen into the system complicates the internal device coupling mechanisms and significantly increases the uncertainty of system operation. Therefore, it is necessary to develop models that accurately describe the dynamic coupling of energy sources such as electrothermal hydrogen and the interactions among multiple stakeholders, as well as to employ optimization algorithms that are more adaptable to environmental changes and can coordinate the interests of multiple stakeholders.
[0004] Compared to traditional optimization algorithms, reinforcement learning (RL) excels at handling nonlinearity, uncertainty, and complex decision-making in complex system optimization and scheduling. As the problems RL tackles become increasingly complex, its applications have shifted from single-agent RL to multi-agent RL. Summary of the Invention
[0005] In order to fully consider the dynamic coupling of heterogeneous energy and the interactive coordination of multiple agents in hydrogen-containing integrated energy systems and realize the coordinated optimization operation of multiple energy carriers and multiple agents, a multi-agent optimal compromise scheduling method for multi-agent hydrogen-containing integrated energy systems was proposed.
[0006] The present invention is implemented by the following technical solution: A multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system, comprising the following steps:
[0007] Step 1: For the multi-agent hydrogen-containing integrated energy system optimization scheduling model, a multi-agent reinforcement learning model consisting of a hydrogen-containing integrated energy service provider agent (HIEP-Agent), an electric load aggregator agent (ELA-Agent), a thermal load aggregator agent (TLA-Agent), and a hydrogen load aggregator agent (HLA-Agent) is constructed;
[0008] Step 2: Initialize the multi-agent policy network, value network parameters, and environment state, and set hyperparameters such as the TOPSIS weight vector and scheduling period.
[0009] Step 3: Through the ordered decision generation mechanism, the action generation order of each agent is sampled in sequence, and the corresponding scheduling plan is constructed;
[0010] Step 4: Input all scheduling plans into the environment, run the complete scheduling cycle, and accumulate their reward values and status information;
[0011] Step 5: Evaluate the cumulative rewards of all scheduling solutions based on the TOPSIS method and select the optimal compromise solution. The action sequence corresponding to the optimal compromise solution is the final scheduling solution for the current optimization cycle.
[0012] Step 6: Based on the reward signal of the optimal compromise solution, the multi-agent proximal policy optimization (MAPPO) algorithm is used to update the policy network and value network parameters of the agent;
[0013] Step 7: If the set number of training rounds has not been reached, proceed to the next round and return to step 3;
[0014] Step 8: When the maximum number of training rounds is reached or the multi-agent proximal policy optimization algorithm converges, the training is terminated and the final scheduling plan is output.
[0015] Further preferably, the multi-agent hydrogen-containing integrated energy system optimization scheduling model is as follows:
[0016] ;
[0017] ;
[0018] Where: is the system exergy efficiency; t is time, T is the total time; 、 、 are the energy quality coefficients of electrical energy, thermal energy and hydrogen energy respectively; 、 and are the electrical load, thermal load and hydrogen load at time t respectively; Indicates the power of interaction with the upper grid; is the natural gas input during period t; and They are Photovoltaic and wind turbine power output during the time period; is the calorific value of natural gas; 、 are the power generation power and power generation efficiency of the cogeneration unit in period t respectively; 、 are the heating power and heating efficiency of the cogeneration unit during period t respectively; 、 and They are The total amount of electricity load reduced, the total amount of heat load reduced, and the total amount of hydrogen load reduced during the period; Compensate for heat load reduction costs; is the heat load compensation cost coefficient; Compensating for hydrogen load reduction costs; is the hydrogen load compensation cost coefficient; Compensation costs for electrical load reduction; is the electricity load compensation cost coefficient.
[0019] Further preferred, the constraints of the multi-agent hydrogen-containing integrated energy system optimization scheduling model include cogeneration unit constraints, electric boiler equipment constraints, electric hydrogen production equipment constraints, fuel cell equipment constraints, reducible load constraints, electric power balance constraints, thermal power balance constraints, hydrogen power balance constraints, and energy storage equipment constraints.
[0020] Further preferably, the state space of each agent is defined as follows:
[0021] ;
[0022] ;
[0023] ;
[0024] ;
[0025] Where: 、 、 and They represent the state space collections of the hydrogen-containing integrated energy service provider agent, the electric load aggregator agent, the thermal load aggregator agent, and the hydrogen load aggregator agent respectively; 、 and They represent the electricity price, heat price and hydrogen price of each entity in the internal transaction of hydrogen-containing integrated energy service providers; Indicates the electricity purchase price of the upper power grid, Indicates the electricity selling price of the upper-level power grid; express Total electric load power during the period; express Total heat load power during the period; express Total hydrogen load power during the period; for Thermal output of fuel cell equipment during the period; The heat output of the combined heat and power unit; for The electric power consumed by the hydrogen production equipment during the time period; for The electric power consumed by the electric boiler equipment during the period; for The power output of fuel cell equipment during the period; for Thermal output of electric boiler equipment during the time period; for The hydrogen power consumed by the fuel cell equipment during the period; for Hydrogen output of time-slot electric hydrogen production equipment; express Time period energy storage charge state; express Thermal energy storage capacity during the period; express Hydrogen storage capacity during the period.
[0026] Further preferably, the action space of each agent is represented as follows:
[0027] ;
[0028] Where: 、 、 and denote the action spaces of the hydrogen-containing integrated energy service provider agent, the electric load aggregator agent, the thermal load aggregator agent, and the hydrogen load aggregator agent, respectively; express Time period energy storage power output; express Thermal output of thermal energy storage during the period; express Hydrogen storage output during the period.
[0029] Further optimization, the reward function of the hydrogen-containing integrated energy service provider agent is:
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] Where: 、 、 、 The rewards are for the Hydrogen Integrated Energy Service Provider Agent (HIEP-Agent), the Electric Load Aggregator Agent (ELA-Agent), the Thermal Load Aggregator Agent (TLA-Agent), and the Hydrogen Load Aggregator Agent (HLA-Agent); is the cost scaling ratio; Auxiliary value to make the reward function return to positive value; Reduce the unsatisfactory cost of electricity load, The operation and maintenance costs of electric energy storage; Reduce unsatisfactory costs for heat load, The operation and maintenance costs of thermal energy storage; The cost of purchasing heat; Reducing unsatisfactory costs for hydrogen load, The operation and maintenance costs of hydrogen energy storage; The cost of purchasing hydrogen.
[0035] Furthermore, the network update of the multi-agent proximal policy optimization (MAPPO) algorithm is based on the Actor-Critic network, and the update objective function of the Actor network is:
[0036] ;
[0037] Where: It is The total loss of the Actor network of agents; is the shear loss; is the entropy regularization term; is the probability ratio of the new and old strategies; is the generalized advantage estimate; is the shear range; is the entropy regularization coefficient; It is The policy function of an agent, For the Actor network parameters of each agent; is the strategy entropy; It means to find the expectation of the empirical trajectory, and clip is the clipping function.
[0038] The critic network updates its parameters by minimizing the mean square error between the value function's prediction and the actual return:
[0039] ;
[0040] Where: is the loss function of the Critic network, The global state of the Critic network Value prediction; is the target value calculated from the trajectory data.
[0041] Furthermore, an ordered decision generation mechanism is used, whereby the agents generate actions in sequence according to a set processing order. The action generation of the current agent depends on its own state and the action of the previous agent. The previous action is dynamically added to the observation data of the current agent to form an enhanced observation vector, which can be expressed mathematically as follows:
[0042] ;
[0043] Where: For the The enhanced observation vector of each agent, Represents a splicing operation, is the state of the i-th agent at time t, is the action of the previous agent at time t.
[0044] Further optimization is performed, and the electric load aggregator agent takes the enhanced observation vector as input and uses the Actor network to generate actions. The thermal load aggregator agent (TLA-Agent) and the hydrogen load aggregator agent (HLA-Agent) also generate actions according to the ordered decision generation mechanism. The four agents are arranged and combined, and there are 24 action generation sequences, that is, 24 scheduling schemes; each scheduling scheme independently inputs the environment, interacts with the environment to obtain the cumulative reward and state transition sequence within the scheduling cycle, and secondly, all scheduling schemes are executed in parallel without interfering with each other.
[0045] Further preferably, for a given scheduling period, the cumulative reward of the scheduling scheme is calculated as follows:
[0046] ;
[0047] Where: is the cumulative reward of the k-th scheduling solution, is the reward of the kth scheduling plan at time t.
[0048] This paper proposes to achieve more efficient interaction and coordination among agents by constructing a multi-agent reinforcement learning model consisting of four agents: a hydrogen-containing integrated energy provider agent (HIEP-Agent), an electric load aggregator agent (ELA-Agent), a thermal load aggregator agent (TLA-Agent), and a hydrogen load aggregator agent (HLA-Agent). Compared to algorithms like MADDPG, which rely on high-dimensional continuous action spaces, the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm's clipping mechanism limits probability ratios, preventing over-adjustments during policy updates and addressing issues such as gradient instability and slow convergence. However, the MAPPO algorithm relies on independent states, independent actions, and a single policy network to generate action sequences during policy updates. This lack of global optimization leads to insufficient collaboration among agents, easily causing action conflicts, a fixed scheduling order, and inflexible energy allocation, which in turn affects optimization efficiency. To address these shortcomings, this paper proposes a novel multi-agent optimal compromise reinforcement learning algorithm (MAOCPPO) that combines the MAPPO algorithm with the TOPSIS (Toposis) method to solve the multi-agent optimal scheduling problem in integrated energy systems. This algorithm incorporates information about the actions of adjacent agents into the current state space to enhance collaboration. Furthermore, a multi-scheduling evaluation mechanism based on the TOPSIS method is established to select the action combination closest to the ideal solution, effectively balancing the competitive and cooperative relationships among agents and improving the quality of scheduling decisions. By comprehensively considering the goals and constraints of each agent, the algorithm achieves efficient optimal scheduling of multi-agent hydrogen-containing integrated energy systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0050] In order to facilitate a complete understanding of the technical concept of the present invention, the present invention is further explained in detail below.
[0051] The multi-agent hydrogen-containing integrated energy system (HIES) studied in this paper consists of a higher-level power grid, a natural gas grid, a hydrogen-containing integrated energy provider (HIEP), an electric load aggregator (ELA), a thermal load aggregator (TLA), and a hydrogen load aggregator (HLA). The HIEP's key equipment includes combined heat and power (CHP), electric boilers (EB), electric hydrogen generators (EL), and fuel cells (HFC), enabling flexible conversion and supply of electricity, heat, and hydrogen. The load aggregator's key equipment includes electric energy storage, thermal energy storage, and hydrogen storage. The HIEP can interact with the higher-level power grid, purchasing electricity when power generation is insufficient and selling it when power generation is sufficient. The price of electricity sold to the grid is lower than the internal transaction price within the multi-agent hydrogen-containing integrated energy system, but higher than the price sold to the ELA. Therefore, the HIEP prioritizes the needs of the load aggregator. ELA integrates dispersed small and medium-sized electrical loads within a multi-agent hydrogen-based integrated energy system and equips them with electrical energy storage; TLA manages the thermal loads within a multi-agent hydrogen-based integrated energy system and equips them with thermal energy storage; and HLA manages the hydrogen loads within a multi-agent hydrogen-based integrated energy system and equips them with hydrogen storage. Loads are categorized as critical fixed loads and curtailable loads. Load aggregators utilize energy storage and flexible loads through demand response (DR) to optimize energy usage and reduce costs.
[0052] The goal of optimizing the operation of a multi-agent hydrogen-containing integrated energy system is to maximize exergy efficiency by adjusting the operation strategy under the premise of satisfying various constraints. The present invention comprehensively considers the exergy efficiency of the multi-agent hydrogen-containing integrated energy system to accurately evaluate the high-quality utilization of electricity, heat, cooling, gas and other energies. The optimization scheduling model of the multi-agent hydrogen-containing integrated energy system is as follows:
[0053] (1);
[0054] (2);
[0055] Where: is the system exergy efficiency; t is time, T is the total time; 、 、 are the energy quality coefficients of electrical energy, thermal energy and hydrogen energy respectively; 、 and are the electrical load, thermal load and hydrogen load at time t respectively; Indicates the power of interaction with the upper grid; is the natural gas input during period t; and They are Photovoltaic and wind turbine power output during the time period; is the calorific value of natural gas; 、 are the power generation power and power generation efficiency of the cogeneration unit in period t respectively; 、 are the heating power and heating efficiency of the cogeneration unit during period t respectively; 、 and They are The total amount of electricity load reduced, the total amount of heat load reduced, and the total amount of hydrogen load reduced during the period; Compensate for heat load reduction costs; is the heat load compensation cost coefficient; Compensating for hydrogen load reduction costs; is the hydrogen load compensation cost coefficient; Compensation costs for electrical load reduction; is the electricity load compensation cost coefficient.
[0056] The constraints of the multi-agent hydrogen-containing integrated energy system optimization scheduling model include cogeneration unit constraints, electric boiler equipment constraints, electric hydrogen production equipment constraints, fuel cell equipment constraints, reducible load constraints, electric power balance constraints, thermal power balance constraints, hydrogen power balance constraints, and energy storage equipment constraints.
[0057] 1) Constraints on cogeneration units:
[0058] (3);
[0059] Where: for Natural gas power consumed by the cogeneration unit during the period; is the heat-to-electricity ratio of the cogeneration unit; and They represent the lower and upper limits of the electrical output of the cogeneration unit respectively; The heat output of the combined heat and power unit; and They represent the lower and upper limits of the thermal output of the cogeneration unit respectively; and represent the ramp rate and ramp rate of the CHP unit respectively; for The power generation capacity of the cogeneration unit during the period.
[0060] 2) Electric boiler equipment constraints:
[0061] (4);
[0062] Where: and They are The electric power and heat output of the electric boiler equipment during the time period; for Thermal output of electric boiler equipment during the time period; for Energy conversion efficiency of electric boiler equipment during certain time periods; and Respectively represent the lower and upper limits of the thermal output of the electric boiler equipment; and They represent the sliding rate and climbing rate of electric boiler equipment respectively.
[0063] 3) Constraints on electric hydrogen production equipment:
[0064] (5);
[0065] Where: and They are The hydrogen output and electric power consumed by the hydrogen production equipment during each period; Energy conversion efficiency of electricity-to-hydrogen equipment; and They represent the lower and upper limits of hydrogen output of electric hydrogen production equipment respectively; and They represent the sliding rate and climbing rate of the electric hydrogen production equipment respectively, for Hydrogen output of electric hydrogen production equipment during the period.
[0066] 4) Fuel cell equipment constraints:
[0067] (6);
[0068] Where: 、 and They are The hydrogen power, electrical output and thermal output consumed by the fuel cell equipment during the period; and The electrical and thermal efficiencies of the fuel cell device.
[0069] 5) Reducible load constraints:
[0070] In an optimization cycle, the load that can be reduced in each period should not exceed the maximum load that can be reduced in that period.
[0071] (7);
[0072] Where: 、 and Respectively The maximum amount of electricity load that can be reduced, the maximum amount of heat load that can be reduced, and the maximum amount of hydrogen load that can be reduced in a period of time.
[0073] 6) Electric power balance constraints:
[0074] (8);
[0075] (9);
[0076] (10);
[0077] Where: express The amount of electricity exchanged with the upper power grid during each period; express Total renewable energy power output during the period; express The electrical output of equipment within the hydrogen integrated energy provider (HIEP) during the period; express Time period energy storage power output; express Total electrical load power during the period.
[0078] 7) Thermal power balance constraints:
[0079] (11);
[0080] Where: express Thermal output of thermal energy storage during the period; express Total heat load power during the period.
[0081] 8) Hydrogen power balance constraints:
[0082] (12);
[0083] Where: express Hydrogen storage output during the period; express Total hydrogen load power during the period.
[0084] 9) Energy storage constraints:
[0085] Energy storage constraints include energy storage output constraints and energy storage state of charge constraints:
[0086] (13);
[0087] Where: express Time period energy storage charge state; and are the upper and lower limits of the electric energy storage output respectively; and are the upper and lower limits of the state of charge of the energy storage, respectively.
[0088] Thermal energy storage constraints include thermal energy storage output constraints and thermal energy storage capacity constraints:
[0089] (14);
[0090] Where: express Thermal energy storage capacity during the period; and are the upper and lower limits of thermal energy storage output respectively; and are the upper and lower limits of thermal energy storage capacity respectively.
[0091] Hydrogen storage constraints include hydrogen storage output constraints and hydrogen storage capacity constraints:
[0092] (15);
[0093] Where: express Hydrogen storage capacity during the period; and are the upper and lower limits of hydrogen storage output respectively; and are the upper and lower limits of hydrogen storage capacity respectively.
[0094] Aiming at the multi-agent hydrogen-containing integrated energy system optimization scheduling model, the present invention constructs a multi-agent reinforcement learning model consisting of four agents: hydrogen-containing integrated energy service provider agent (HIEP-Agent), electric load aggregator agent (ELA-Agent), thermal load aggregator agent (TLA-Agent) and hydrogen load aggregator agent (HLA-Agent).
[0095] Multi-agent reinforcement learning algorithms need to transform the problem into a multi-agent Markov decision process (MMDP), which takes into account the interactions between multiple agents and mainly includes nine elements: ,in is the state space, For action space, is the number of agents, For rewards, is the transition probability, is the discount factor, For the agent strategy, For observation space, It is a joint agent strategy.
[0096] The state space is the environmental characteristics that the agent can obtain, which is used to describe the environmental information it perceives. The state space of each agent is defined as follows:
[0097] (16);
[0098] (17);
[0099] (18);
[0100] (19);
[0101] Where: 、 、 and They represent the state space collections of the hydrogen-containing integrated energy service provider agent, the electric load aggregator agent, the thermal load aggregator agent, and the hydrogen load aggregator agent respectively; 、 and They represent the electricity price, heat price and hydrogen price of each entity in the internal transaction of hydrogen-containing integrated energy service providers, Indicates the electricity purchase price of the upper power grid, Indicates the electricity selling price of the upper-level power grid.
[0102] Based on the coupling characteristics between decision variables, the action space of each agent can be defined as follows: HIEP-Agent actions are simplified to the thermal output of the CHP unit and the thermal output of the fuel cell HFC; ELA reduces the electric load through demand response and adjusts the energy storage charging and discharging power, and its action space consists of the electric load reduction and the electric energy storage power; TLA coordinates with the thermal storage system through thermal load regulation, and its action space includes the thermal load reduction and the thermal energy storage power; HLA forms an action space consisting of the hydrogen load reduction and the hydrogen energy storage power through hydrogen load reduction and hydrogen storage regulation. The action space of each agent is represented as follows:
[0103] (20);
[0104] Where: 、 、 and They represent the action spaces of the hydrogen-containing integrated energy service provider agent, the electric load aggregator agent, the thermal load aggregator agent, and the hydrogen load aggregator agent, respectively.
[0105] For the hydrogen-containing integrated energy service provider agent (HIEP-Agent), its optimization goal is to maximize the system exergy efficiency. The objective function formula is the reward function, while the optimization goal of the load aggregator agent is to minimize the energy cost, including the unsatisfactory cost of load reduction. , Energy storage operation and maintenance costs and the energy purchase costs of each entity within the multi-agent hydrogen-containing integrated energy system , reward function:
[0106] (twenty one);
[0107] (twenty two);
[0108] (twenty three);
[0109] (twenty four);
[0110] Where: 、 、 、 The rewards are for the Hydrogen Integrated Energy Service Provider Agent (HIEP-Agent), the Electric Load Aggregator Agent (ELA-Agent), the Thermal Load Aggregator Agent (TLA-Agent), and the Hydrogen Load Aggregator Agent (HLA-Agent); is the cost scaling ratio; Auxiliary value to make the reward function return to positive value; Reduce the unsatisfactory cost of electricity load, The operation and maintenance costs of electric energy storage; Reduce unsatisfactory costs for heat load, The operation and maintenance costs of thermal energy storage; The cost of purchasing heat; Reducing unsatisfactory costs for hydrogen load, The operation and maintenance costs of hydrogen energy storage; The cost of purchasing hydrogen.
[0111] Multi-agent reinforcement learning (MARL) algorithms enable multiple agents to learn and optimize policies independently or collaboratively in complex environments through their own experience and interactions. The Multi-Agent Proximal Policy Optimization (MAPPO) algorithm is a type of multi-agent reinforcement learning algorithm, an extension of the Proximal Policy Optimization (PPO) algorithm for multi-agent environments. It ensures the stability of each agent's policy update while allowing a certain degree of variability between policies. Through a pruning mechanism and a multi-agent-specific loss function, it maintains the stability of probability ratios and the diversity of policies.
[0112] The MAPPO algorithm's network update is based on an actor-critic network. The agent uses the actor network to output its policy, while the critic network estimates the value of state-action pairs. The actor network parameters are updated via gradient ascent to maximize the pruned objective function. The updated objective function of the actor network is:
[0113] (25);
[0114] Where: It is The total loss of the Actor network of agents; It is the clipping loss, which is used to limit the amplitude of the policy update to ensure that the update is not too large, thereby maintaining the stability of the training; is the entropy regularization term, which is used to encourage the agent to explore and prevent the strategy from converging to the local optimal solution too early; is the probability ratio of the new and old strategies. The shear mechanism limits the strategy update range. When the difference between the new and old strategies is too large, the boundary value is used to replace the actual probability ratio. is the generalized advantage estimate, calculated by the critic network; To limit the update range of the strategy, is the entropy regularization coefficient; It is The policy function of an agent, For the Actor network parameters of each agent; is the policy entropy, used to encourage exploration; It represents the expectation of the empirical trajectory (i.e., taking the average of all sampled samples), and clip is the clipping function.
[0115] The critic network updates its parameters by minimizing the mean square error between the value function's prediction and the actual return:
[0116] (26);
[0117] Where: is the loss function of the Critic network, The global state of the Critic network Value prediction; is the target value calculated from the trajectory data.
[0118] MAPPO's multi-agent decision-making method has three limitations: 1) MAPPO relies on a policy gradient iterative optimization mechanism, which requires multiple rounds of policy updates and environmental interactions to gradually converge to a steady-state policy. It cannot find the current optimal action sequence within a single-step decision cycle; 2) A single policy network generates a fixed-order action sequence, which solidifies the energy allocation pattern and cannot generate heterogeneous scheduling solutions based on dynamic environmental conditions, resulting in insufficient flexibility in resource supply and demand matching; 3) The online training mechanism limits policy diversity, and complex systems need to explore a better solution space through diversified action sequences.
[0119] To address these issues, this paper proposes a multi-agent optimal compromise mechanism. This mechanism aims to achieve an optimal compromise solution by leveraging the ordered action generation and dependencies between agents, thereby improving both the rewards for each agent and the overall efficiency of the system. This optimal compromise mechanism is primarily achieved by combining an ordered decision-making mechanism, the parallel execution of multiple solutions, a reward accumulation mechanism, and the TOPSIS method.
[0120] (1) Orderly decision-making mechanism:
[0121] This invention designs an ordered decision generation mechanism. Agents generate actions sequentially according to a set processing order. The action generated by the current agent depends on its own state and the action of the previous agent. The previous action is dynamically added to the observation data of the current agent to form an enhanced observation vector. In this way, subsequent agents can adjust their own actions through the neural network based on the decision information of the previous agent, thereby better adapting to the global task objectives. The mathematical expression of enhanced observation is as follows:
[0122] (27);
[0123] Where: For the The enhanced observation vector of each agent, Represents a splicing operation, is the state of the i-th agent at time t, is the action of the previous agent at time t.
[0124] Taking the ordered decision-making action sequence of the hydrogen integrated energy service provider agent (HIEP-Agent), the electric load aggregator agent (ELA-Agent), the thermal load aggregator agent (TLA-Agent), and the hydrogen load aggregator agent (HLA-Agent) as an example, during the time period, the HIEP state observation space remains unchanged, and the electric load aggregator agent enhances the observation vector:
[0125] (28);
[0126] Where: is the enhanced observation vector of the load aggregator agent at time t, is the state of the electric load aggregator agent at time t, is the action of the hydrogen-containing integrated energy service provider agent at time t.
[0127] The electric load aggregator agent takes the augmented observation vector as input and generates actions using an actor network. The thermal load aggregator agent (TLA-Agent) and the hydrogen load aggregator agent (HLA-Agent) also generate actions using an ordered decision-making mechanism. Using permutations (factorial of 4), these four agents generate 24 possible action sequences, or 24 possible scheduling solutions.
[0128] (2) Parallel operation of multiple schemes and reward accumulation mechanism: The ordered decision generation mechanism generates 24 agent action generation sequences, each of which corresponds to a complete action generation scheduling scheme. To comprehensively evaluate the optimization performance of different schemes, the algorithm runs all generated scheduling schemes in parallel. Specifically, each scheduling scheme is independently input into the environment and interacts with the environment to obtain the cumulative reward and state transition sequence within the scheduling cycle. Secondly, all scheduling schemes are executed in parallel without interfering with each other. This design avoids the local optimal problem that may be caused by a single scheduling scheme and accelerates the efficiency of scheduling scheme evaluation.
[0129] For a given scheduling period, the cumulative reward of the scheduling scheme is calculated as follows:
[0130] (29);
[0131] Where: is the cumulative reward of the k-th scheduling solution, is the reward of the kth scheduling plan at time t.
[0132] (3) Top-of-the-Ideal Solution Ranking Method: After obtaining the cumulative rewards for the 24 scheduling solutions, each of which has four agent rewards, the optimal compromise solution among the 24 scheduling solutions is obtained using the Top-of-the-Ideal Solution Ranking Method (TOPSIS) method, taking into account the four rewards. TOPSIS is a multi-objective decision-making method that compares the relative distance of each scheduling solution to the positive and negative ideal solutions and selects the scheduling solution closest to the ideal solution as the final optimal compromise solution.
[0133] The present invention combines the multi-agent reinforcement learning MAPPO algorithm with the optimal compromise mechanism based on TOPSIS to form a new multi-agent optimal compromise reinforcement learning algorithm MAOCPPO. The relevant mechanisms and principles of the multi-agent optimal compromise reinforcement learning algorithm are described in detail above. Figure 1, a multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system, comprising the following steps:
[0134] Step 1. Construct a multi-agent reinforcement learning model: For the multi-agent hydrogen-containing integrated energy system optimization scheduling model, construct a multi-agent reinforcement learning model consisting of a hydrogen-containing integrated energy service provider agent (HIEP-Agent), an electric load aggregator agent (ELA-Agent), a thermal load aggregator agent (TLA-Agent), and a hydrogen load aggregator agent (HLA-Agent);
[0135] Step 2: Initialize parameters and hyperparameters: Initialize the multi-agent policy network, value network parameters, and environment state, and set hyperparameters such as the weight vector and scheduling period of the TOPSIS method.
[0136] Step 3: Generate a scheduling plan: Through the ordered decision generation mechanism, sample the actions of each agent in sequence and build a corresponding scheduling plan;
[0137] Step 4: Run the environment and accumulate information: Input all scheduling plans into the environment, run the complete scheduling cycle, and accumulate their reward values and status information;
[0138] Step 5: TOPSIS evaluation and selection of the optimal compromise solution: Based on the TOPSIS method, the cumulative rewards of all scheduling solutions are evaluated and the optimal compromise solution is selected. The action sequence corresponding to the optimal compromise solution is the final scheduling solution for the current optimization cycle.
[0139] Step 6: MAPPO algorithm updates network parameters: Based on the reward signal of the optimal compromise solution, the multi-agent proximal policy optimization (MAPPO) algorithm is used to update the policy network and value network parameters of the agent;
[0140] Step 7: Determine whether the number of training rounds has been reached: If the number of training rounds has not been reached, proceed to the next round and return to step 3;
[0141] Step 8: Determine whether convergence has occurred and output the final scheduling plan: When the maximum number of training rounds or the convergence condition of the multi-agent proximal strategy optimization algorithm is reached, the training is terminated and the final scheduling plan is output.
[0142] The above description merely represents preferred embodiments of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above disclosure to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.
Claims
1. A multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system, characterized by: The following steps are involved: Step 1: For the multi-agent hydrogen-containing integrated energy system optimization scheduling model, a multi-agent reinforcement learning model consisting of a hydrogen-containing integrated energy service provider agent, an electric load aggregator agent, a thermal load aggregator agent, and a hydrogen load aggregator agent is constructed; Step 2: Initialize the multi-agent strategy network, value network parameters and environment state, and set the weight vector and scheduling period of the ideal solution sorting method; Step 3: Through the ordered decision generation mechanism, the action generation order of each agent is sampled in sequence, and the corresponding scheduling plan is constructed; Step 4: Input all scheduling plans into the environment, run the complete scheduling cycle, and accumulate their reward values and status information; Step 5: Evaluate the cumulative rewards of all scheduling solutions based on the approximate ideal solution ranking method and select the optimal compromise solution. The action sequence corresponding to the optimal compromise solution is the final scheduling solution for the current optimization cycle; Step 6: Based on the reward signal of the optimal compromise solution, use the multi-agent proximal policy optimization algorithm to update the policy network and value network parameters of the agent; Step 7: If the set number of training rounds has not been reached, proceed to the next round and return to step 3; Step 8: When the maximum number of training rounds is reached or the multi-agent proximal policy optimization algorithm converges, the training is terminated and the final scheduling plan is output.
2. The multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system according to claim 1 is characterized in that: The multi-agent hydrogen-containing integrated energy system optimization scheduling model is as follows: ; ; Where: is the system exergy efficiency; t is time, T is the total time; 、 、 are the energy quality coefficients of electrical energy, thermal energy and hydrogen energy respectively; 、 and are the electrical load, thermal load and hydrogen load at time t respectively; Indicates the power of interaction with the upper grid; is the natural gas input during period t; and They are Photovoltaic and wind turbine power output during the time period; is the calorific value of natural gas; 、 are the power generation power and power generation efficiency of the cogeneration unit in period t respectively; 、 are the heating power and heating efficiency of the cogeneration unit during period t respectively; 、 and They are The total amount of electricity load reduced, the total amount of heat load reduced, and the total amount of hydrogen load reduced during the period; Compensate for heat load reduction costs; is the heat load compensation cost coefficient; Compensating for hydrogen load reduction costs; is the hydrogen load compensation cost coefficient; Compensation costs for electrical load reduction; is the electricity load compensation cost coefficient.
3. The multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system according to claim 2 is characterized in that: The constraints of the multi-agent hydrogen-containing integrated energy system optimization scheduling model include cogeneration unit constraints, electric boiler equipment constraints, electric hydrogen production equipment constraints, fuel cell equipment constraints, reducible load constraints, electric power balance constraints, thermal power balance constraints, hydrogen power balance constraints, and energy storage equipment constraints.
4. The multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system according to claim 3 is characterized in that: The state space of each agent is defined as follows: ; ; ; ; Where: 、 、 and They represent the state space collections of the hydrogen-containing integrated energy service provider agent, the electric load aggregator agent, the thermal load aggregator agent, and the hydrogen load aggregator agent respectively; 、 and They represent the electricity price, heat price and hydrogen price of each entity in the internal transaction of hydrogen-containing integrated energy service providers; Indicates the electricity purchase price of the upper power grid, Indicates the electricity selling price of the upper-level power grid; express Total electric load power during the period; express Total heat load power during the period; express Total hydrogen load power during the period; for Thermal output of fuel cell equipment during the period; The heat output of the combined heat and power unit; for The electric power consumed by the hydrogen production equipment during the time period; for The electric power consumed by the electric boiler equipment during the period; for The power output of fuel cell equipment during the period; for Thermal output of electric boiler equipment during the time period; for The hydrogen power consumed by the fuel cell equipment during the period; for Hydrogen output of time-slot electric hydrogen production equipment; express Time period energy storage charge state; express Thermal energy storage capacity during the period; express Hydrogen storage capacity during the period.
5. The multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system according to claim 4 is characterized in that: The action space representation of each agent is as follows: ; Where: 、 、 and denote the action spaces of the hydrogen-containing integrated energy service provider agent, the electric load aggregator agent, the thermal load aggregator agent, and the hydrogen load aggregator agent, respectively; express Time period energy storage power output; express Thermal output of thermal energy storage during the period; express Hydrogen storage output during the period.
6. The multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system according to claim 5 is characterized in that: The reward function of the hydrogen-containing integrated energy service provider agent is: ; ; ; ; Where: 、 、 、 These are the rewards for the hydrogen-containing integrated energy service provider agent, the electric load aggregator agent, the thermal load aggregator agent, and the hydrogen load aggregator agent; is the cost scaling ratio; Auxiliary value to make the reward function return to positive value; Reduce the unsatisfactory cost of electricity load, The operation and maintenance costs of electric energy storage; Reduce unsatisfactory costs for heat load, The operation and maintenance costs of thermal energy storage; The cost of purchasing heat; Reducing unsatisfactory costs for hydrogen load, The operation and maintenance costs of hydrogen energy storage; The cost of purchasing hydrogen.
7. The multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system according to claim 6 is characterized in that: The network update of the multi-agent proximal policy optimization algorithm is based on the Actor-Critic network. The update objective function of the Actor network is: ; Where: It is The total loss of the Actor network of agents; is the shear loss; is the entropy regularization term; is the probability ratio of the new and old strategies; is the generalized advantage estimate; is the shear range; is the entropy regularization coefficient; It is The policy function of an agent, For the Actor network parameters of each agent; is the strategy entropy; It means to find the expectation of the empirical trajectory, and clip is the clipping function; The critic network updates its parameters by minimizing the mean square error between the value function's prediction and the actual return: ; Where: is the loss function of the Critic network, The global state of the Critic network Value prediction; is the target value calculated from the trajectory data.
8. The multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system according to claim 7 is characterized in that: Using the ordered decision generation mechanism, the agent generates actions in sequence according to the set processing order. The action generation of the current agent depends on its own state and the action of the previous agent. The previous action is dynamically added to the observation data of the current agent to form an enhanced observation vector, which is mathematically expressed as follows: ; Where: For the The enhanced observation vector of each agent, Represents a splicing operation, is the state of the i-th agent at time t, is the action of the previous agent at time t.
9. The multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system according to claim 8 is characterized in that: The electric load aggregator agent takes the enhanced observation vector as input and uses the Actor network to generate actions. The thermal load aggregator agent and the hydrogen load aggregator agent also generate actions according to the ordered decision-making mechanism. The four agents are arranged and combined, and there are 24 action generation sequences, that is, 24 scheduling schemes. Each scheduling scheme inputs the environment independently and interacts with the environment to obtain the cumulative reward and state transition sequence within the scheduling cycle. All scheduling schemes are executed in parallel without interfering with each other.
10. The multi-agent optimal compromise scheduling method for a multi-agent hydrogen-containing integrated energy system according to claim 9 is characterized in that, for Given a scheduling period, the cumulative reward of a scheduling scheme is calculated as follows: ; Where: is the cumulative reward of the k-th scheduling solution, is the reward of the kth scheduling plan at time t.
Citation Information
Patent Citations
Multi-agent power generation optimal scheduling method based on reinforcement learning
CN110728406A
Multi-park integrated energy system collaborative optimization scheduling method and system
CN117852710A