Multi-objective adaptive optimization control method and system for proton exchange membrane fuel cell

By constructing a reward function and a health decay model using reinforcement learning algorithms, multi-objective adaptive optimization control of the proton exchange membrane fuel cell system is achieved, solving the stack polarization loss problem caused by load fluctuations, extending system life and improving energy efficiency.

CN122494716APending Publication Date: 2026-07-31SHANDONG ELECTRIC POWER ENG CONSULTING INST CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG ELECTRIC POWER ENG CONSULTING INST CORP
Filing Date
2026-05-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing proton exchange membrane fuel cell systems struggle to achieve globally optimal control under load fluctuations, leading to drastic changes in internal polarization losses and shortening system lifespan. Traditional control strategies prioritize performance over lifespan, making it difficult to meet real-time control requirements.

Method used

A multi-objective adaptive optimization control method is adopted, and a reward function is constructed through reinforcement learning algorithm, including load tracking term, hydrogen consumption penalty term, output fluctuation suppression term and lifetime constraint term, forming a power allocation strategy that prioritizes healthy stacks and limits the load of aging stacks. Combined with fuel cell mechanism model and health degradation model, coordinated control of stack and energy storage unit is achieved.

Benefits of technology

While meeting power requirements, it reduces the rate of fuel cell degradation, extends system life, and improves system energy efficiency optimization and lifespan protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122494716A_ABST
    Figure CN122494716A_ABST
Patent Text Reader

Abstract

This invention relates to the field of battery control technology, and provides a multi-objective adaptive optimization control method and system for proton exchange membrane fuel cells (PEMFCs). The method includes: acquiring the operating state of the PEMFC system; generating control actions through a policy network; sending the control actions to the execution layer, where the fuel cell stack and energy storage unit jointly perform load power tracking; updating the operating state based on the execution results; calculating an immediate reward function; calculating the time-series difference error using a value network; and updating the reinforcement learning policy network and value network. The immediate reward function includes a load tracking term, a hydrogen consumption penalty term, an output fluctuation suppression term, and a lifetime constraint term. This reduces the fuel cell stack degradation rate and extends the overall system lifetime.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of battery control technology, and particularly relates to a multi-objective adaptive optimization control method and system for proton exchange membrane fuel cells. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Proton exchange membrane fuel cells (PEMFCs), as a highly efficient and clean energy conversion device, have shown great potential in the fields of distributed energy and mobile power. However, a PEMFC system is a typical multivariable coupled and strongly nonlinear dynamic system. In actual operation, frequent load fluctuations can lead to drastic changes in polarization losses within the stack, and long-term variable load operation can accelerate the irreversible degradation of the membrane electrode structure, shortening the system lifespan.

[0004] Existing control methods mostly rely on fixed rules or sophisticated physical mechanism models. The former is difficult to achieve global optimization under complex operating conditions, while the latter has an exponential increase in computational load when facing complex systems such as multiple stacks in parallel, making it difficult to meet the stringent requirements of real-time control. More importantly, traditional strategies often prioritize performance over lifespan, resulting in short-term power increases at the expense of long-term reliability. Summary of the Invention

[0005] To address the technical problems mentioned above, this invention provides a multi-objective adaptive optimization control method and system for proton exchange membrane fuel cells. The reward function is set to consist of a load tracking term, a hydrogen consumption penalty term, an output fluctuation suppression term, and a lifetime constraint term. The reinforcement learning algorithm can automatically learn to form a power allocation strategy of "prioritizing healthy stacks and limiting load on aging stacks" while meeting power requirements, thereby reducing the stack degradation rate and extending the overall system lifetime.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: The first aspect of the present invention provides a multi-objective adaptive optimization control method for proton exchange membrane fuel cells, comprising: The operating status of the proton exchange membrane fuel cell system is acquired, and control actions are generated through a strategy network. After the control action is sent to the execution layer, the stack and energy storage unit jointly complete the load power tracking, update the operating status according to the execution result, calculate the instantaneous reward function, calculate the time difference error by the value network, and update the reinforcement learning policy network and the value network. The immediate reward function includes a load tracking term, a hydrogen consumption penalty term, an output fluctuation suppression term, and a lifetime constraint term.

[0007] Furthermore, the control actions include: power allocation or current density allocation of each fuel cell stack; and charging and discharging power of the energy storage unit.

[0008] Furthermore, the operating status includes the output power of each fuel cell stack: ;in, This refers to the number of series-connected battery cells contained in the fuel cell stack. Indicates the voltage of a single cell. This indicates the current of a single cell.

[0009] Furthermore, the voltage of the single cell is: ;in, Current density; This is the equivalent ohmic internal resistance; and These are empirical parameters; Reversible voltage Furthermore, the operating state includes the water tank temperature, and the heat balance model for the water tank temperature is: ; in, Let t be the water tank temperature. To generate heat for fuel cells, For user heat load; The equivalent heat capacity of the water tank is calculated using the following formula: ; The specific heat capacity of water at constant pressure. The density of water, The effective water storage capacity of the water tank. This refers to the equivalent heat capacity of the water tank's insulation layer and shell.

[0010] Furthermore, the operating state includes the health state of each fuel cell stack, and the voltage decay model for the health state of each fuel cell stack is as follows: ;in, For cumulative voltage decay, This represents the maximum voltage decay value at the end of the fuel cell's lifespan, while SOH represents the fuel cell's health status.

[0011] Furthermore, the operating state includes the state of charge of the energy storage unit, and the evolution model of the state of charge of the energy storage unit is as follows: ; in, Let t be the state of charge of the energy storage unit. For battery charging and discharging power, Let be the time interval between time t and time t+1. This refers to the battery capacity.

[0012] A second aspect of the present invention provides a multi-objective adaptive optimization control system for a proton exchange membrane fuel cell, comprising: The action generation module is configured to: acquire the operating status of the proton exchange membrane fuel cell system and generate control actions through a strategy network; The control optimization module is configured to: send control actions to the execution layer, and then the stack and energy storage unit jointly complete load power tracking, update the operating status according to the execution results, calculate the instantaneous reward function, calculate the time difference error by the value network, and update the reinforcement learning policy network and the value network. The immediate reward function includes a load tracking term, a hydrogen consumption penalty term, an output fluctuation suppression term, and a lifetime constraint term.

[0013] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described above.

[0014] A fourth aspect of the present invention provides a computer device including a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, wherein the processor executes the program to implement the steps in the multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described above.

[0015] Compared with the prior art, the beneficial effects of the present invention are: The present invention sets the reward function to consist of a load tracking term, a hydrogen consumption penalty term, an output fluctuation suppression term, and a lifetime constraint term. The reinforcement learning algorithm can automatically learn to form a power allocation strategy of "prioritizing healthy stacks and limiting load on aging stacks" while meeting power requirements, thereby reducing the stack degradation rate and extending the overall system lifetime.

[0016] This invention constructs a closed-loop control framework consisting of a fuel cell mechanism model, a health degradation model, a reinforcement learning decision model, and an electrothermal storage collaborative execution system. This framework enables the fuel cell system to meet electrothermal load requirements while simultaneously optimizing energy efficiency and protecting its lifespan. Attached Figure Description

[0017] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0018] Figure 1 This is a flowchart of the multi-objective adaptive optimization control method for a proton exchange membrane fuel cell according to Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the reinforcement learning agent according to Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the PEMFC system configuration according to Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to Embodiment 4 of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0020] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0021] Example 1 This embodiment provides a multi-objective adaptive optimization control method for proton exchange membrane fuel cells.

[0022] The multi-objective adaptive optimization control method for proton exchange membrane fuel cells provided in this embodiment can integrate the advantages of physical mechanisms and data-driven approaches, while taking into account both efficiency and lifespan.

[0023] To address the common problems in the operation and control of existing fuel cell integrated energy systems, such as the difficulty in incorporating lifespan degradation into real-time scheduling decisions, the difficulty in balancing system energy efficiency and lifespan optimization, and the high complexity of multi-stack collaborative control, this embodiment provides a multi-objective adaptive optimization control method for proton exchange membrane fuel cells. By constructing a closed-loop control framework consisting of a fuel cell mechanism model, a health degradation model, a reinforcement learning decision model, and an electrothermal storage collaborative execution system, this method enables the fuel cell system to meet electrothermal load requirements while simultaneously optimizing energy efficiency and protecting lifespan.

[0024] The multi-objective adaptive optimization control method for proton exchange membrane fuel cells provided in this embodiment, such as Figure 1 As shown, it includes the following steps: Step 1: Construct the PEMFC multi-stack mechanism model and the SOH evolution model.

[0025] Step 101: To describe the physical operating mechanism of the fuel cell system, a PEMFC stack array mechanism model is first established, including a voltage model, a power model, a hydrogen consumption model, and a heat generation model, such as... Figure 3 As shown.

[0026] (1) PEMFC single-cell voltage model.

[0027] The output voltage of a fuel cell unit is obtained by subtracting various polarization losses from the ideal reversible voltage: ; in, Individual unit output voltage; Reversible voltage; Activation polarization loss; Ohmic loss; Concentration polarization loss.

[0028] In the low to medium current density range, concentration polarization losses can be ignored, therefore the voltage expression can be simplified to: ; The activation polarization loss is represented by an empirical function: ; Ohmic loss is: ; Therefore, the final expression for the single-cell voltage is: ; in, Current density; This is the equivalent ohmic internal resistance; and These are empirical parameters.

[0029] The PEMFC cell voltage model is used to describe the output voltage characteristics of a fuel cell under different load conditions.

[0030] (2) Power model of fuel cell stack.

[0031] The current of a single cell is: ;in, This represents the effective reaction area of ​​the electrode.

[0032] The output power of a single battery cell is: ; If the fuel cell stack contains If there are 100 battery cells connected in series, then the power of the stack is: Unit: kilowatt; If the system contains If there are 1 fuel cell stack, then the total system power is: The unit is kilowatt; among which, Let be the power of the k-th fuel cell stack, which will be used later. express.

[0033] The power model of the fuel cell stack is used to calculate the real-time output power of the system and serves as important state information for the reinforcement learning control strategy.

[0034] (3) Hydrogen consumption model.

[0035] Based on the electrochemical reaction of fuel cells: ; The molar consumption rate of hydrogen is: The unit is moles per second; among which, It is Faraday's constant; The corresponding hydrogen energy input power is: The unit is kilowatt; among which, This is the lower heating value of hydrogen.

[0036] (4) Fuel cell heat generation model.

[0037] During operation, fuel cells convert some chemical energy into electrical energy, while the remainder is released as heat.

[0038] The heat generation capacity of the fuel cell stack is: ; The total heat generated by the system is: ;in, The heat output power of the k-th fuel cell stack is expressed in kilowatts.

[0039] This heat can be recovered through a heat exchange system to meet the user's heat load requirements.

[0040] Step 102: Establish an online SOH decay model.

[0041] To describe the changes in fuel cell lifetime, a voltage decay model was established.

[0042] The voltage decay of fuel cells is mainly caused by the following four degradation mechanisms: start-stop cycle degradation, load change degradation, low power operation degradation, and high power operation degradation.

[0043] The voltage attenuation increment is: ; in, Number of start-stop cycles; : Power variation range; Low-power operation time; High-power operating time; , , , For weights.

[0044] The cumulative voltage decay is: Where t represents the current total running time.

[0045] This model discretizes the long-term operation of a fuel cell into a series of continuous operating stages. Within each stage, it collects operating parameters such as the number of start-stop cycles, power fluctuations, and durations of high and low power operation, and calculates the voltage decay increment for each stage. t, as the upper limit of the summation, represents the total number of stages currently completed, indirectly corresponding to the total operating time of the fuel cell, thus enabling dynamic tracking of the voltage decay process throughout its entire lifecycle. Health status is defined as: ;in, This represents the maximum voltage decay value at the end of the fuel cell's lifespan.

[0046] The SOH index serves as a key input variable for reinforcement learning control strategies.

[0047] This embodiment constructs an online assessment model for the state of health (SOH) of a fuel cell oriented towards operational control. By comprehensively considering various degradation mechanisms such as start-stop cycles, power fluctuations, low-power operation, and high-power operation, a fuel cell voltage decay model is established and transformed into a health status index that can be used for real-time control. This model can assess changes in stack lifetime in real time during system operation, enabling the fuel cell degradation state to be dynamically perceived by the control strategy, and providing lifetime constraint information for subsequent optimized control.

[0048] Compared with traditional control methods that only focus on instantaneous power or efficiency, this method can explicitly introduce the factors affecting the long-term lifespan of fuel cells into the control decision-making process, so that the control strategy can meet the load demand while avoiding excessive acceleration of stack degradation.

[0049] Step 2: Establish a collaborative decision-making model based on reinforcement learning, such as... Figure 2 As shown.

[0050] The operation of the PEMFC system is modeled as a Markov Decision Process (MDP), which can be represented as: ; in, : System state space; Controlling the action space; : State transition probability; : Reward function.

[0051] (1) The state space, i.e., the system state vector, is defined as: ; in, , , , , These are electrical load, thermal load, battery SOC, water tank temperature, and SOH of each fuel cell stack.

[0052] (2) Action space, that is, the control action is defined as the current density of each stack. distribute: ; That is, the system allocates different current densities to each fuel cell stack.

[0053] (3) PPO strategy update.

[0054] The policy function is: ; The PPO algorithm is used for updating strategies, and its objective function is: ; in, , Represents the mathematical expectation of an empirical sample. This represents the ratio of the probability of actions under the new and old strategies. Represents the dominance function. This represents the trimming (truncating) function. Indicates the cutting factor. This represents the updated policy function. This represents the policy function before the update.

[0055] The PPO algorithm ensures training stability by limiting the policy update magnitude.

[0056] This embodiment models the operation scheduling problem of a multi-stack fuel cell system as a Markov decision process (MDP) and constructs a reinforcement learning control strategy with the system operating state as input and the power allocation of each stack as output.

[0057] By introducing the Proximal Policy Optimization (PPO) algorithm for policy training, the agent can autonomously learn the optimal control policy in complex dynamic load environments, thereby achieving adaptive power allocation among multiple stacks.

[0058] During the power allocation process, multi-stack collaborative control is achieved through the following strategies: determining the total output power of the fuel cell based on the system load demand; dynamically allocating the power ratio according to the SOH status of each stack; prioritizing stacks with better health status to undertake higher loads; and implementing penalty control for stacks with large power fluctuations.

[0059] This method does not rely on complex analytical optimization models or accurate prediction models, and can achieve efficient control in uncertain environments, significantly improving the system's operational flexibility and real-time performance.

[0060] Step 3: Execution of coordinated control of electricity, heat and energy storage.

[0061] This embodiment introduces an energy storage system and a thermal storage system on the basis of a multi-stack fuel cell system to achieve coordinated scheduling and control of electrical and thermal energy.

[0062] like Figure 3As shown, the core of the system consists of eight independently adjustable parallel fuel cell stacks, complemented by three core subsystems: hydrogen and air supply, hydrothermal management and waste heat recovery, and electrical energy conversion and storage. This allows the system to simultaneously meet users' dual energy needs for electricity and heat. The flexible architecture of the multi-stack parallel configuration provides the hardware platform for reinforcement learning-based optimization control based on health awareness, enabling synergistic optimization of energy efficiency improvement and equipment lifespan protection.

[0063] Energy storage and thermal storage systems include: PEMFC stack arrays, lithium battery energy storage systems, hot water storage tanks, and electric heating devices.

[0064] The reinforcement learning controller outputs the power allocation strategy for the fuel cell and also coordinates the regulation of the energy storage system and the thermal system.

[0065] The battery SOC evolution model for energy storage systems is as follows: ; in, Let t be the state of charge of the energy storage unit. For battery charging and discharging power, Let be the time interval between time t and time t+1. This refers to the battery capacity.

[0066] The waste heat generated during the power generation process of the fuel cell is recovered through a heat exchanger and enters a hot water storage tank. The heat balance model of the hot water storage tank is as follows: ; in, Let t be the water tank temperature. To generate heat for fuel cells (i.e., ), For user heat load, The equivalent heat capacity of the water tank. The equivalent heat capacity of the water tank is calculated using the following formula: ; The specific heat capacity of water at constant pressure. The density of water, The effective water storage capacity of the water tank. This refers to the equivalent heat capacity of the water tank's insulation layer and shell.

[0067] During system operation, the reinforcement learning controller dynamically adjusts the multi-stack output power, battery energy storage charging and discharging power, and heat recovery and heating strategies based on electrical load demand, thermal load demand, and energy storage status. When fuel cell power generation is insufficient, the energy storage system supplements the electrical load; when fuel cell heat generation is insufficient, the electric heating device supplements the thermal load. Through the coordinated scheduling of multiple energy flows—electricity, heat, and storage—the system can improve overall energy utilization efficiency while meeting comprehensive energy demands.

[0068] The reinforcement learning controller outputs control actions based on real-time status to achieve: power distribution among multiple fuel cells, regulation of energy storage charging and discharging, and coordinated supply of heat load, thereby optimizing system energy efficiency and protecting lifespan under dynamic load conditions.

[0069] Step 4: Implementation of the entire process of collaborative optimization control for multi-stack PEMFC-CHP systems, such as... Figure 1 As shown.

[0070] Step 401: Data collection and input parameter preprocessing.

[0071] This step provides the basic inputs and boundary constraints for the entire system operation, corresponding to... Figure 1 The input layer of Step 1.

[0072] First, we completed the collection and standardized preprocessing of load data from multiple scenarios, collecting the user-side electric and heat load demand sequence. It fully covers the dynamic changes in electrical and thermal loads under different application scenarios, providing core input requirements for subsequent system simulation and control.

[0073] The system initial conditions are initialized synchronously, and the initial health state of each fuel cell stack is set. Initial temperature of hot water storage tank For systems equipped with energy storage units, the initial state of charge of the energy storage units is set synchronously. Simultaneously, it initializes operational statistics variables such as cumulative stack operating time, number of start-stop cycles, and load change records, providing a unified initial state benchmark for subsequent modeling and control processes.

[0074] In addition, pre-setting system operation constraint parameters defines the safe operating boundary of the system, specifically covering core elements such as the safe operating range of the fuel cell stack power, the power change rate limit, the upper and lower limits of the operating temperature, and the current density constraint, thus defining a compliant operating range for the execution of subsequent control strategies.

[0075] Step 402: High-fidelity modeling of the multi-stack PEMFC-CHP system.

[0076] This step constructs a digital twin model of the system, providing accurate state evolution and simulation support for the reinforcement learning environment. Figure 1 The high-fidelity modeling layer in Step 2 consists of three coupled model modules: (1) Electrochemical Model of Multi-Stack PEMFC: Establish a power-current characteristic model for a single stack, a power superposition model for multiple stacks, and a hydrogen molar consumption rate model. If a single stack contains If there are 10 fuel cell cells connected in series, then the stack output power is 1000 kW. If the system contains If there are 1 fuel cell stack, then the total output power of the system is 1. Simultaneously, a model for the molar consumption rate of hydrogen was established. This enables precise quantitative calculation of the stack output power and hydrogen consumption characteristics under different current densities.

[0077] (2) Stack attenuation mechanism and online dynamic evaluation model of SOH A multi-condition online SOH degradation model is constructed to achieve online dynamic assessment of the health status of fuel cell stacks.

[0078] During operation, real-time data such as output power, current, and voltage of each fuel cell stack are collected. The operating conditions are classified according to the power level of the fuel cell stack, and the fuel cell operating conditions are divided into four categories: start-stop cycle operating condition, power variation operating condition, low power operating condition, and high power operating condition. Among them, the low power operating condition is defined as the output power of the fuel cell stack being less than 1 / 5 of the rated power, and the high power operating condition is defined as the output power being more than 4 / 5 of the rated power.

[0079] Subsequently, at each discrete time step, based on the degradation factors corresponding to the above operating conditions, including: the number of start-stop cycles, power variation amplitude, duration of low-power operation, and duration of high-power operation at the current time step; based on these operating characteristics, the cumulative voltage degradation of the fuel cell is calculated using a voltage decay model. This model posits that fuel cell lifespan degradation is primarily composed of the superposition of four degradation mechanisms: voltage decay caused by start-stop cycles, voltage decay caused by load variations, voltage decay caused by low-power operation, and voltage decay caused by high-power operation. The voltage decay increment is: ; Its cumulative voltage decay can be expressed as a weighted sum of the above four types of degradation terms, i.e. .

[0080] Each degradation mechanism corresponds to a different attenuation coefficient. For example, the voltage attenuation coefficient for start-stop cycles is approximately 13.79 μV / cycle, the attenuation coefficient for power variation is approximately 0.04185 μV / kW, the attenuation rate for low-power operation is approximately 8.662 μV / h, and the attenuation rate for high-power operation is approximately 10 μV / h. At each time step, the system calculates the voltage attenuation increment based on the operating data and accumulates it to obtain the current total voltage degradation of the fuel cell stack. After obtaining the accumulated voltage degradation, it is compared with the ultimate voltage attenuation value at the end of the fuel cell's lifespan. Normalization is performed to obtain the current health status of the fuel cell stack. ; Updated Reflecting the remaining performance level of the fuel cell stack, the system then enters the next time step cycle, continuously executing the process of "operational data acquisition - degradation calculation - SOH update" to achieve online dynamic assessment of the health status of the fuel cell stack and provide health awareness information for subsequent power distribution and optimization control.

[0081] (3) Thermal management model: Establish a thermal balance model for the temperature of the hot water storage tank. The thermal balance model for the tank temperature is as follows: ,in Let t be the water tank temperature. To generate heat for fuel cells, For user heat load, The total heat capacity of the water tank is used to achieve dynamic simulation of the entire process of heat generation from the fuel cell, heat extraction from the user's heat load, and heat dissipation from the water tank environment, thus accurately depicting the electro-thermal coupling characteristics of the system.

[0082] Step 403: Construction of the reinforcement learning control environment based on the PPO algorithm This step constructs a reinforcement learning framework for closed-loop interaction between the agent and the environment, corresponding to... Figure 1 The core control layer of Step 3 consists of an outer decision-making module encapsulated by the PPO agent and an inner simulation module executed by the system's high-fidelity model. Continuous optimization of the control strategy is achieved through the collaborative updating of the policy network and the value network. The specific implementation process is as follows: (1) Initialization of outer PPO agent: Construct and initialize the Actor-Critic dual neural network structure, in which the policy network (Actor) is used to output the control actions of the multi-heap system, and the value network (Critic) is used to evaluate the value of the system state; (2) System state vector acquisition: The system collects and constructs state vectors in real time. ,correspond Figure 1 The system status received by the intelligent agent specifically includes: system load power demand (user electrical load). User heat load Current output power of each fuel cell stack or current density Health status of each fuel cell stack State of charge of energy storage unit (If the system includes energy storage), fuel cell stack voltage, water tank temperature Key operating parameters such as hydrogen consumption estimation; these variables together constitute the system state vector. This is used to characterize the current operating and health status of the system; (3) Control action generation: The reinforcement learning agent generates actions based on the current state. Control actions are generated through a policy network. ,correspond Figure 1 The intelligent agent outputs action signals to the environment; among these, the control actions mainly include: power distribution to each fuel cell stack. or current density distribution The charging and discharging power of the energy storage unit (if the system includes an energy storage unit), and the adjustment parameters of the relevant auxiliary execution units, that is, the system allocates different current densities to each fuel cell stack to achieve coordinated power distribution among multiple stacks; (4) Inner system simulation and state update: After the control action is sent to the system execution layer, the stack and energy storage unit jointly complete the load power tracking; the system performs PEMFC-CHP simulation based on the high-fidelity model established in step 402, completes the multi-stack power optimization and thermal coupling process calculation, updates the system operating state, and obtains the system state at the next moment. ; Multi-objective reward function calculation: At each time step, the system calculates the instantaneous reward function based on the current control result. ,correspond Figure 1 The reward signal is fed back to the agent from the environment. To achieve synergistic optimization of efficiency and lifespan, the reward function is improved based on the traditional reinforcement learning control framework. Specifically, the reward function is composed of a weighted average of multiple objective terms, comprehensively evaluating the system's performance after executing the control action at the current time step. The reward function R is defined as follows: ; in, This is the power error penalty coefficient; This is the penalty coefficient for battery capacity error; This is the penalty coefficient for water tank temperature error; This is the penalty coefficient for energy curtailment error; This is a penalty coefficient for battery health. The single power fluctuation penalty coefficient is represented by a, b, and c; these are the correlation index coefficients. The specific errors are for power / battery capacity / water tank temperature; Wasting energy for the system; This represents the current power of each fuel cell stack. This represents the power of the fuel cell stack in the previous iteration. This represents the average health status of each fuel cell stack.

[0083] The physical meaning and design basis of each item are as follows: Load tracking items ( This ensures that the system output power meets the external load requirements; by designing error square terms or positive reward terms, the strategy is guided to respond quickly to load changes; Hydrogen consumption penalty item ( To reduce system hydrogen consumption and improve energy utilization efficiency; to introduce a reward value that is negatively correlated with the hydrogen consumption rate in order to pursue the long-term economic efficiency of the system. Output fluctuation suppression term ( This reduces rapid fluctuations in stack power and improves system stability; it also protects the fuel cell from frequent thermal / electrochemical shocks by penalizing the rate of power change (|ΔP|). Lifetime constraints ( A penalty term based on SOH is introduced to suppress operating conditions that accelerate aging. This term, combined with the degradation model established above, imposes a high penalty on severe conditions such as high load and frequent start-stop to extend the system life.

[0084] The reward function is used to evaluate the system's performance after executing control actions at the current time step and to generate an immediate reward value. This reward value reflects the overall performance of the current control strategy on multiple objectives, such as whether the load power requirements are met, whether hydrogen consumption is low, whether the stack output power is stable, and whether the stack health status (SOH) is maintained well.

[0085] By designing a reward function, multiple optimization objectives, such as system efficiency, stability, and lifetime protection, can be unified into a quantifiable evaluation metric. Reinforcement learning algorithms, by maximizing cumulative rewards, gradually guide the control strategy towards the optimal operating mode that satisfies these objectives.

[0086] The reward function is used to quantitatively evaluate the performance of the system after executing control actions, and generates a reward signal at each time step to feed back to the PPO algorithm, which guides the policy network update, so that the reinforcement learning policy can be gradually optimized in long-term operation and achieve multi-objective control.

[0087] By introducing health-related constraints into the reward function, reinforcement learning algorithms can automatically learn to form a power allocation strategy of "prioritizing healthy stacks and limiting loads on aging stacks" while meeting power requirements, thereby reducing the stack degradation rate and extending the overall system lifespan.

[0088] Finally, after multiple rounds of interactive training, a converged optimal control strategy is obtained. During the online operation of the system, the reinforcement learning agent outputs the optimal control action based on real-time observations, realizing adaptive power allocation and lifetime-aware optimization control of the multi-fuel cell stack system.

[0089] Example 2 The multi-objective adaptive optimization control system for proton exchange membrane fuel cells provided in this embodiment includes: The action generation module is configured to: acquire the operating status of the proton exchange membrane fuel cell system and generate control actions through a strategy network; The control optimization module is configured to: send control actions to the execution layer, and then the stack and energy storage unit jointly complete load power tracking, update the operating status according to the execution results, calculate the instantaneous reward function, calculate the time difference error by the value network, and update the reinforcement learning policy network and the value network. The immediate reward function includes a load tracking term, a hydrogen consumption penalty term, an output fluctuation suppression term, and a lifetime constraint term.

[0090] It should be noted that each module in this embodiment corresponds one-to-one with each step in Embodiment 1, and their specific implementation processes are the same, so they will not be repeated here.

[0091] Example 3 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described in Embodiment 1 above.

[0092] Example 4 This embodiment provides a computer device, such as... Figure 4 As shown, the system includes a computer-readable storage medium 1003, a processor 1001, a communication interface 1002, and a computer program stored on the computer-readable storage medium 1003 and executable on the processor 1001. The processor 1001, communication interface 1002, and computer-readable storage medium 1003 can be connected via a bus or other means. The communication interface 1002 is used to receive and transmit data. When the processor 1001 executes the program, it implements the steps in the multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described in Embodiment 1 above.

[0093] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-objective adaptive optimization control method for proton exchange membrane fuel cells, characterized in that, include: The operating status of the proton exchange membrane fuel cell system is acquired, and control actions are generated through a strategy network. After the control action is sent to the execution layer, the stack and energy storage unit jointly complete the load power tracking, update the operating status according to the execution result, calculate the instantaneous reward function, calculate the time difference error by the value network, and update the reinforcement learning policy network and the value network. The immediate reward function includes a load tracking term, a hydrogen consumption penalty term, an output fluctuation suppression term, and a lifetime constraint term.

2. The multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described in claim 1, characterized in that, The control actions include: power allocation or current density allocation of each fuel cell stack; and charging and discharging power of the energy storage unit.

3. The multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described in claim 1, characterized in that, The operating status includes the output power of each fuel cell stack: ;in, This refers to the number of series-connected battery cells contained in the fuel cell stack. Indicates the voltage of a single cell. This indicates the current of a single cell.

4. The multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described in claim 3, characterized in that, The voltage of the individual battery cell is: ;in, Current density; This is the equivalent ohmic internal resistance; and These are empirical parameters; It is a reversible voltage.

5. The multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described in claim 1, characterized in that, The operating status includes the water tank temperature, and the thermal balance model for the water tank temperature is as follows: ; in, Let t be the water tank temperature. To generate heat for fuel cells, For user heat load, The equivalent heat capacity of the water tank is calculated using the following formula: ; The specific heat capacity of water at constant pressure. The density of water, The effective water storage capacity of the water tank. This refers to the equivalent heat capacity of the water tank's insulation layer and shell.

6. The multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described in claim 1, characterized in that, The operating status includes the health status of each fuel cell stack, and the voltage decay model for the health status of each fuel cell stack is as follows: ;in, For cumulative voltage decay, This represents the maximum voltage decay value at the end of the fuel cell's lifespan, while SOH represents the fuel cell's health status.

7. The multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described in claim 1, characterized in that, The operating state includes the energy storage unit's state of charge, and the evolution model of the energy storage unit's state of charge is as follows: ; in, Let t be the state of charge of the energy storage unit. For battery charging and discharging power, Let be the time interval between time t and time t+1. This refers to the battery capacity.

8. A multi-objective adaptive optimization control system for a proton exchange membrane fuel cell, characterized in that, include: The action generation module is configured to: acquire the operating status of the proton exchange membrane fuel cell system and generate control actions through a strategy network; The control optimization module is configured to: send control actions to the execution layer, and then the stack and energy storage unit jointly complete load power tracking, update the operating status according to the execution results, calculate the instantaneous reward function, calculate the time difference error by the value network, and update the reinforcement learning policy network and the value network. The immediate reward function includes a load tracking term, a hydrogen consumption penalty term, an output fluctuation suppression term, and a lifetime constraint term.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the steps in the multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described in any one of claims 1-7.

10. A computer device comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the multi-objective adaptive optimization control method for proton exchange membrane fuel cells as described in any one of claims 1-7.