A method for controlling a gas supply system of a proton exchange membrane fuel cell based on a MOSAC algorithm
By proton exchange membrane fuel cell gas supply system control method based on MOSAC algorithm, the problems of strong coupling of multiple variables, time-varying nonlinearity and lag control such as cathode and anode inlet gas mass flow rate, pressure and gas excess ratio are solved. Multi-objective collaborative optimization is achieved, the dynamic response and reliability of the system are improved and the equipment life is extended.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN UNIV CHONGQING RES INST
- Filing Date
- 2026-02-14
- Publication Date
- 2026-06-02
Smart Images

Figure CN122136407A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of proton exchange membrane fuel cell gas supply system control technology, specifically relating to a proton exchange membrane fuel cell gas supply system control method based on the MOSAC algorithm. Background Technology
[0002] Currently, existing fuel cell gas supply systems still have many problems, including:
[0003] Existing technology (I): Chinese invention patent "A Real-time Control Method for Fuel Cell Gas Supply System Based on MODDPG Algorithm", application number "202411674593.X". This patent overcomes the problem that traditional discrete motion control strategies are difficult to achieve global optimum in complex multi-objective optimization by using the MODDPG algorithm. Through the multi-objective trade-off mechanism of this algorithm, it can better adapt to the optimization needs of fuel cell gas supply systems in complex dynamic environments, enabling the control strategy to achieve optimal balance among different objectives. However, the following technical limitations remain: Although a multi-objective form is introduced, its integration of Pareto theory is superficial, and it is essentially still a two-objective parallel optimization, ignoring key constraints such as pressure safety and system energy efficiency, resulting in insufficient multi-objective trade-off capability and restricting overall control performance. At the same time, the basic algorithm (DDPG) itself has problems such as low exploration efficiency, sensitivity to hyperparameters, and unstable training.
[0004] Prior Art (II): Chinese Invention Patent "Control Method and System for PEMFC Jet Gas Supply System Based on Model Deep Reinforcement Learning", application number "202311611437.4". This patent utilizes an actor-critic framework to interact with the learned dynamic system model of the PEMFC jet gas supply system and maximize the cumulative reward within the prediction interval, learning a neural network strategy based on model predictive control. Ultimately, by fixing the parameters of the actor network model and deploying the actor network in the controller of the PEMFC jet gas supply system, real-time optimal control of the PEMFC jet gas supply system can be achieved. However, this patent has the following technical defects: First, its multi-objective optimization capability is insufficient; it is essentially still a single-objective weighted optimization and does not achieve true multi-objective trade-offs. Second, it has high model dependence and significant robustness risks. Furthermore, due to algorithmic limitations, the actor-critic framework, which employs a deterministic strategy, typically has lower exploration efficiency and policy robustness than the SAC of the maximum entropy framework.
[0005] Prior Art (III): Chinese Invention Patent "A Collaborative Control Method for Hydrogen Fuel Cell Gas Supply System Based on Neural Network", application number "202410271876.3". This invention provides a collaborative control method for a hydrogen fuel cell gas supply system based on a neural network. By abstracting and simplifying the air management system model, it achieves collaborative tracking control of the oxygen ratio and gas pressure in the cathode channel of the fuel cell system. Compared with hydrogen fuel cell gas supply control strategies based on other control algorithms, this invention can not only more effectively reduce tracking errors, but also reduce the total variation of the air compressor control signal, thereby helping to extend the service life of the air compressor. However, this patent application has the following technical defects: This method is an intelligent improvement of the traditional PID control idea, but it lags behind the deep reinforcement learning strategy based on MOSAC in terms of multi-objective collaborative optimization, autonomous learning decision-making, and comprehensive performance (safety, economy, lifespan) improvement.
[0006] Prior Art (IV): Chinese Invention Patent "A Control Strategy for PEMFC Gas Supply System Based on Improved Gray Wolf Optimization", Patent No. "CN 112886036 B". This invention uses an improved gray wolf optimization algorithm (IGWO) to improve and optimize the fuzzy controller; the technical solution of this invention can achieve effective and accurate control of the proton exchange membrane fuel cell gas supply system, prevent oxygen deficiency and damage to the fuel cell stack, and ensure the performance and safety of the fuel cell system. However, this patent application has the following technical defects: First, the optimization objective is singular, only taking the maximization of net power as the objective; second, the convergence speed and global search capability are limited, and it is easy to get trapped in local optima; at the same time, its adaptability is relatively general. Summary of the Invention
[0007] This invention aims to overcome the control challenges posed by the strong coupling, time-varying nonlinearity, and large lag characteristics of multiple variables such as the mass flow rate, pressure, and excess gas ratio of the cathode and anode in a proton exchange membrane fuel cell (PEMFC) gas supply system. This invention proposes a control strategy for a PEMFC gas supply system based on the MOSAC algorithm. The strategy aims to construct a multi-objective state space and action space, design a hierarchical reward function, and utilize the MOSAC algorithm to achieve coordinated optimization control of the air compressor speed, hydrogen circulation pump speed, and valve opening. The improved method is suitable for real-time optimization and intelligent control of PEMFC gas supply systems, and is of great significance for improving the dynamic responsiveness, operational economy, and reliability of fuel cell systems.
[0008] This invention discloses a control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm. The control method for the proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm includes the following steps:
[0009] Step 1: Establish a model of the proton exchange membrane fuel cell gas supply system;
[0010] The proton exchange membrane fuel cell gas supply system model includes a cathode flow channel model and an anode flow channel model.
[0011] Step 2: Obtain data on the actual state of the fuel cell gas supply system;
[0012] Step 3: Control the proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm.
[0013] Furthermore, in step 1, the mass flow rate equation of the cathode flow channel model is expressed as:
[0014] ;
[0015] in The mass of oxygen at the cathode of a proton exchange membrane fuel cell. The mass of nitrogen at the cathode of a proton exchange membrane fuel cell. , These represent the mass change rates of oxygen and nitrogen at the cathode, respectively; This represents the mass flow rate of air entering the cathode, which is affected by the air compressor speed S. comp control; This means that the oxygen content in the air is approximately 21% by mass. This means that nitrogen in the air accounts for approximately 78% by mass. The mass flow rate at the cathode outlet; The total mass of the cathode outlet. , , , These represent the masses of oxygen, nitrogen, and water vapor at the cathode outlet, respectively. The mass flow rate of oxygen consumed in the reaction inside the cathode, where , , , , These are the number of individual cells in the battery stack, the stack current, the number of electrons transferred per mole of O2, the Faraday constant, and the molar mass of oxygen, respectively.
[0016] The cathode pressure is:
[0017] ;
[0018] In the formula: For cathode pressure, The cathode temperature, This indicates the mass of oxygen at the cathode. Indicates the mass of nitrogen gas at the cathode. The gas constant representing oxygen. Where R is the gas constant. Here is the molar mass of oxygen. The gas constant of nitrogen. Where R is the gas constant. Here is the molar mass of nitrogen. For water vapor in Saturation pressure at temperature This indicates the volume of the cathode in the fuel cell stack;
[0019] The oxygen excess ratio is used to measure the oxygen supply at the cathode, and is defined as the ratio of the oxygen supply to the amount of reactant.
[0020] .
[0021] Furthermore, in step 1, the mass flow rate equation of the anode can be expressed as:
[0022] ;
[0023] in The mass of hydrogen gas at the anode of a proton exchange membrane fuel cell. This indicates the rate of mass change of hydrogen at the anode; This represents the mass flow rate of hydrogen entering the anode, which is affected by the rotational speed S of the hydrogen circulation pump. pump and the opening degree of the hydrogen supply circuit valve control; The mass flow rate at the anode outlet; The total mass of the anode outlet. ; , These represent the masses of hydrogen and water vapor at the anode outlet, respectively. The mass flow rate of hydrogen consumed in the reaction inside the anode. , , , , These are the number of individual cells in the battery stack, the stack current, the number of electrons transferred per mole of H2, the Faraday constant, and the molar mass of hydrogen, respectively.
[0024] The anode pressure is:
[0025] ;
[0026] In the formula: For anode pressure, This is the anode temperature. Indicates the mass of hydrogen gas at the anode. The gas constant of hydrogen. Where R is the gas constant. Here is the molar mass of hydrogen. Water vapor is in Saturation pressure at temperature This indicates the volume of the anode in the fuel cell stack;
[0027] Defined as the ratio of hydrogen supply to reaction amount:
[0028] .
[0029] Furthermore, in step 2, the data on the actual state of the fuel cell gas supply system include the oxygen / hydrogen excess ratio, cathode / anode pressure, oxygen / hydrogen inlet mass flow rate, stack current, and voltage.
[0030] Furthermore, step 3 also includes the following steps:
[0031] Step 31: Construct a state space that includes the inlet flow rate and pressure of the cathode and anode, the excess gas ratio, the load current and the stack temperature, as well as an action space that includes the speed of the cathode air compressor, the speed of the anode circulating pump, the opening degree of the cathode / anode valve and the opening degree of the anode supply valve.
[0032] Initialize the experience replay pool and establish the following core components:
[0033] An Actor network for outputting control actions; multiple Critic evaluation networks and their corresponding target networks independently set for each control objective; and a set of preference weight vectors for weighing the importance of different objectives;
[0034] Step 32, convert the current state space Input an Actor network, output the action space By applying the action space to the fuel cell gas supply system environment, a multi-objective reward function is obtained. With the next state space ; Interaction samples Store the data in the experience replay pool and repeat this step until enough training data has been accumulated.
[0035] Step 33: Randomly select a batch of historical interaction samples from the experience replay pool as training data for this network parameter update;
[0036] Step 34, for each target in the batch sample : State space Inputting the Actor network yields the action space. Enter the first The cumulative reward estimate for each Critic evaluation network is obtained. Calculate the target Q value of the objective; update the parameters of all m Critic evaluation networks by minimizing the loss function of each Critic network using gradient descent.
[0037] Step 35, based on the current preference vector The overall Q value is obtained by weighted summation of the Q values of each objective. By maximizing the aggregated Q value and introducing a policy entropy regularization term, the loss function of the Actor network is constructed, and its parameters are updated using the gradient ascent method, thereby driving the control policy to optimize along the Pareto improvement direction.
[0038] Step 36: Based on the performance of the current strategy on each objective, adaptively adjust the preference vector w through gradient update to explore different regions of the Pareto front; periodically resample the preference vector to ensure full coverage of the trade-off space between multiple objectives;
[0039] Step 37: For each target network corresponding to the Critic evaluation network, the parameters are slowly synchronized to the latest parameters of the corresponding evaluation network using a soft update method to enhance training stability; if the adaptive entropy coefficient is enabled, the temperature parameter is dynamically adjusted according to the actual exploration degree of the strategy.
[0040] Step 38: During training, continuously evaluate the policy performance corresponding to different preference vectors and maintain a set of Pareto optimal policies that are not mutually dominant; during actual deployment, select the most suitable control policy from this set based on the real-time target preferences of the system.
[0041] Step 39: Repeat steps 32 to 38 to iteratively optimize each network parameter using the continuously accumulated interaction data until the control strategy performance converges.
[0042] Furthermore, in step 31, the state space defined based on the dynamic characteristics of the proton exchange membrane fuel cell gas supply system is:
[0043] ;
[0044] in , The excess ratio of oxygen to hydrogen; , These are the pressures at time t at the cathode and anode, respectively; , These are the inlet mass flow rates of air and hydrogen at time t, respectively. , These represent the current and voltage of the proton exchange membrane fuel cell stack, respectively. This refers to the rate of change or the deviation from the set value.
[0045] Furthermore, in step 31, the action space defined based on the dynamic characteristics of the proton exchange membrane fuel cell gas supply system is:
[0046] ;
[0047] ;
[0048] This is the maximum hydrogen circulation pump speed. This is the maximum air compressor speed. and All are determined by the power of the fuel cell stack.
[0049] Furthermore, in step 32, the multi-objective reward function includes:
[0050] Excess ratio tracking reward:
[0051] ;
[0052] in: , Indicates the target excess ratio of oxygen. Indicates the target excess ratio of hydrogen; express ; These represent the weighting coefficient and the sensitivity parameter, respectively.
[0053] Smoothing penalties for each actuator:
[0054] ;
[0055] Indicates the weighting coefficient. This indicates the change in air compressor speed. This indicates the changes in the hydrogen circulation pump. This indicates the change in the opening degree of the cathode valve. This indicates the change in the opening degree of the oxygen valve. This indicates the change in the opening degree of the hydrogen supply valve;
[0056] Stress-related safety penalties:
[0057] ;
[0058] Indicates the weighting coefficient. Indicates the maximum pressure at the cathode. Indicates the maximum pressure at the anode;
[0059] Energy efficiency and emissions penalties:
[0060] ;
[0061] Indicates the weighting coefficient. This indicates the mass flow rate of excess oxygen that is discharged. This indicates the mass flow rate of the excess hydrogen gas that is discharged.
[0062] The multi-objective reward function is as follows:
[0063]
[0064] To over-track rewards, For the smoothing penalty of each actuator, As a form of stress-based safety punishment, Penalties for energy efficiency and emissions.
[0065] The beneficial effects achieved by this invention are:
[0066] This invention, through in-depth analysis of the gas supply system of a proton exchange membrane fuel cell, constructs a dynamic model of the cathode and anode flow channels, achieving accurate description of key state variables such as the mass flow rates of oxygen, nitrogen, and hydrogen, real-time mass changes, anode and cathode pressures, and gas excess ratio. This model directly couples these state variables with control variables such as air compressor speed, hydrogen circulation pump speed, anode and cathode valve openings, and stack current, thus forming an accurate mapping with the state space and action space of the MOSAC algorithm. This modeling method lays an accurate and rapid dynamic foundation for multi-objective real-time optimization control of the gas supply system, significantly improving the physical consistency and convergence reliability of the control strategy.
[0067] This invention integrates Pareto optimization theory with the MOSAC method proposed by the SAC algorithm for multi-objective collaborative optimization control of a proton exchange membrane fuel cell (PEMFC) gas supply system. By constructing a multi-objective state space and action space and designing a hierarchical reward function, it achieves collaborative optimization control of air compressor speed, hydrogen circulation pump speed, and valve opening. This ensures that the oxygen-to-hydrogen excess ratio of the fuel cell stack quickly and stably tracks the set value under dynamic load conditions, maintains the safe range of anode and cathode pressures, and improves system energy efficiency and equipment lifespan. It overcomes the control challenges caused by the strong coupling, time-varying nature, nonlinearity, and hysteresis of multiple variables such as cathode and anode inlet mass flow rate, pressure, and gas excess ratio in the PEMFC gas supply system. Attached Figure Description
[0068] Figure 1 This is a flowchart illustrating a control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm.
[0069] Figure 2 This is a logic block diagram of a control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm. Detailed Implementation
[0070] The present invention will be further described below with reference to specific embodiments, and the advantages and features of the present invention will become clearer as a result. However, these embodiments are merely exemplary and do not constitute any limitation on the scope of the present invention. Those skilled in the art should understand that modifications or substitutions can be made to the details and form of the technical solutions of the present invention without departing from the spirit and scope of the present invention, but all such modifications and substitutions fall within the protection scope of the present invention.
[0071] like Figure 1-2 As shown in this embodiment, a control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm is provided. This control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm includes the following steps:
[0072] Step 1: Establish a model of the proton exchange membrane fuel cell gas supply system;
[0073] The gas supply system model for a proton exchange membrane fuel cell includes a cathode flow channel model and an anode flow channel model. The established flow channel model satisfies the following assumptions: 1. All gases are ideal gases; 2. The various gas components inside the stack and auxiliary equipment are fully mixed and uniformly distributed; 3. The gas flow rate, pressure, temperature, and humidity at the two electrodes of the stack are controlled separately; 4. The temperature and humidity of each cell inside the stack are uniformly distributed and remain constant; 5. The proton exchange membrane is fully humidified, the water vapor inside the stack is saturated, and the liquid volume has no effect on the system.
[0074] Cathode flow path model: The cathode mainly contains oxygen, nitrogen, and water. As assumed in point five above, water exists but does not change dynamically. Therefore, the mass flow rate equation for the cathode can be expressed as:
[0075]
[0076] in The mass of oxygen at the cathode of a proton exchange membrane fuel cell. The mass of nitrogen at the cathode of a proton exchange membrane fuel cell. , These represent the mass change rates of oxygen and nitrogen at the cathode, respectively; This represents the mass flow rate of air entering the cathode, which is affected by the air compressor speed (S). comp )control; This means that the oxygen content in the air is approximately 21% by mass. This means that nitrogen in the air accounts for approximately 78% by mass. The mass flow rate at the cathode outlet; The total mass of the cathode outlet ( , , , (These represent the masses of oxygen, nitrogen, and water vapor at the cathode outlet, respectively). The mass flow rate of oxygen consumed in the reaction inside the cathode, where , , , , These are the number of individual cells in the battery stack, the stack current, the number of electrons transferred per mole of O2, the Faraday constant (approximately 96485 C / mol), and the molar mass of oxygen (approximately 32 g / mol).
[0077] Based on assumption 1 above, the gas in the proton exchange membrane fuel cell stack is an ideal gas, therefore the ideal gas equation can be used; thus, the cathode pressure is:
[0078]
[0079] In the formula: For cathode pressure, The cathode temperature, This indicates the mass of oxygen at the cathode. Indicates the mass of nitrogen gas at the cathode. The gas constant representing oxygen. Where R is the gas constant (approximately 8.314). ), This is the molar mass of oxygen (approximately 32 g / mol). The gas constant of nitrogen. Where R is the gas constant (approximately 8.314). ), This is the molar mass of nitrogen (approximately 28 g / mol). For water vapor in The saturation pressure at a given temperature (its magnitude depends on temperature), as assumed in hypothesis 5 above, is the pressure of water vapor. This indicates the volume of the cathode in the fuel cell stack.
[0080] In this invention, the oxygen excess ratio is used to measure the degree of oxygen supply to the cathode, defined as the ratio of oxygen supply to reaction amount:
[0081]
[0082] Anode flow path model: The anode mainly contains hydrogen and water. As assumed in point five above, water exists but does not change dynamically. Therefore, the mass flow rate equation for the anode can be expressed as:
[0083]
[0084] in The mass of hydrogen gas at the anode of a proton exchange membrane fuel cell. This indicates the rate of mass change of hydrogen at the anode; This represents the mass flow rate of hydrogen entering the anode, which is affected by the speed of the hydrogen circulation pump (S). pump ) and the opening degree of the hydrogen supply circuit valve ( )control; The mass flow rate at the anode outlet; The total mass of the anode outlet ( ; , (These represent the masses of hydrogen and water vapor at the anode outlet, respectively). The mass flow rate of hydrogen consumed in the reaction inside the anode. , , , , These are the number of individual cells in the battery stack, the stack current, the number of electrons transferred per mole of H2 reaction, the Faraday constant (approximately 96485 C / mol), and the molar mass of hydrogen (approximately 2 g / mol).
[0085] Based on assumption 1 above, the gas in the proton exchange membrane fuel cell stack is an ideal gas, therefore the ideal gas equation can be used; thus, the anode pressure is:
[0086]
[0087] In the formula: For anode pressure, This is the anode temperature. Indicates the mass of hydrogen gas at the anode. The gas constant of hydrogen. Where R is the gas constant (approximately 8.314). ), This represents the molar mass of hydrogen gas (approximately 2 g / mol). Water vapor is in The saturation pressure at a given temperature (its magnitude depends on temperature), as assumed in hypothesis 5 above, is the pressure of water vapor. This indicates the volume of the anode in the fuel cell stack.
[0088] In this invention, the excess hydrogen ratio is used to measure the supply level of hydrogen gas at the anode, and is defined as the ratio of the hydrogen supply to the reaction amount:
[0089]
[0090] Step 2: Obtain data on the actual status of the fuel cell gas supply system. This data includes the oxygen / hydrogen excess ratio, cathode / anode pressure, oxygen / hydrogen inlet mass flow rate, stack current, and voltage.
[0091] Step 3: Control the proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm. The basic workflow of this algorithm is as follows:
[0092] Step 31: Construct a state space encompassing cathode and anode inlet flow rate, pressure, gas excess ratio, load current, and stack temperature; and an action space including cathode air compressor speed, anode circulating pump speed, cathode / anode valve opening, and anode supply valve opening. Initialize the experience replay pool and establish the following core components: an Actor network for output control actions; multiple Critic evaluation networks and their corresponding target networks independently set for each control objective (such as oxygen excess ratio tracking, hydrogen excess ratio tracking, pressure safety range, and actuator action smoothness); and a set of preference weight vectors for weighing the importance of different objectives.
[0093] Step 32, convert the current state space Input an Actor network, output the action space By applying the action space to the fuel cell gas supply system environment, a multi-objective reward function is obtained. With the next state space Interaction samples Store the data in the experience replay pool and repeat this step until enough training data has been accumulated.
[0094] Step 33: Randomly select a batch of historical interaction samples from the experience replay pool as training data for this network parameter update.
[0095] Step 34, for each target in the batch sample : State space Inputting the Actor network yields the action space. Enter the first The cumulative reward estimate for each Critic evaluation network is obtained. Calculate the target Q value for the objective; update the parameters of all m Critic evaluation networks by minimizing the loss function of each Critic network using gradient descent.
[0096] Step 35, based on the current preference vector The overall Q value is obtained by weighted summation of the Q values of each objective. By maximizing the aggregated Q-value and introducing a policy entropy regularization term, the loss function of the Actor network is constructed, and its parameters are updated using the gradient ascent method, thereby driving the control policy to optimize along the Pareto improvement direction.
[0097] Step 36: Based on the performance of the current strategy on each objective, adaptively adjust the preference vector w through gradient update to explore different regions of the Pareto front; periodically resample the preference vector to ensure full coverage of the trade-off space between multiple objectives.
[0098] Step 37: For each target network corresponding to the Critic evaluation network, the parameters are slowly synchronized to the latest parameters of the corresponding evaluation network using a soft update method to enhance training stability; if the adaptive entropy coefficient is enabled, the temperature parameter is dynamically adjusted according to the actual exploration degree of the strategy.
[0099] Step 38: During training, continuously evaluate the policy performance corresponding to different preference vectors and maintain a set of Pareto optimal policies that are not mutually dominant; during actual deployment, select the most suitable control policy from this set according to the real-time target preference of the system.
[0100] Step 39: Repeat steps 32 to 38, iteratively optimizing the network parameters using the continuously accumulated interactive data until the control strategy performance converges. Ultimately, an intelligent control strategy is obtained that can adaptively and collaboratively optimize the air compressor, hydrogen circulation pump, and valves under dynamic loads, while ensuring rapid and accurate tracking of the gas excess ratio, stable pressure, and smooth actuator movement.
[0101] In step 31, the MOSAC method, which combines Pareto optimization theory and SAC algorithm, is used to address the state space design, action space design, multi-objective reward function design, and multi-objective optimization in the multi-objective optimization control strategy of the proton exchange membrane fuel cell gas supply system as follows:
[0102] The state space defined based on the dynamic characteristics of the proton exchange membrane fuel cell gas supply system is as follows:
[0103]
[0104] in , The excess ratio of oxygen to hydrogen; , These are the pressures at time t at the cathode and anode, respectively; , These are the inlet mass flow rates of air and hydrogen at time t, respectively. , These represent the current and voltage of the proton exchange membrane fuel cell stack, respectively. This refers to the rate of change or the deviation from the set value.
[0105] The action space defined based on the dynamic characteristics of the proton exchange membrane fuel cell gas supply system is as follows:
[0106]
[0107]
[0108] In step 32, based on the dynamic characteristics of the proton exchange membrane fuel cell gas supply system and Pareto theory, the control objective is decomposed into multiple competing sub-objectives, and the resulting multi-objective reward function is as follows:
[0109] (a) Excessive tracking reward
[0110]
[0111] in: , These represent the weighting coefficient and the sensitivity parameter, respectively.
[0112] (ii) Smoothing penalty for each actuator
[0113]
[0114] (III) Stress Safety Penalties
[0115]
[0116] (iv) Energy efficiency and emission penalties
[0117]
[0118] Based on the dynamic characteristics of the proton exchange membrane fuel cell gas supply system and Pareto theory, the multi-objective reward vector is constructed as follows:
[0119]
[0120] The MOSAC algorithm is used to learn a set of non-dominated policies, forming the Pareto optimal policy set. Each strategy can correspond to a preference vector. ,satisfy Simultaneously optimize the weighted objective: This allows for dynamic adjustment of preference weights based on different operating conditions during the operational phase, including startup, load variation, and steady state, achieving adaptive coupling between multiple objectives. For example, during the rapid load variation phase, improving... (Excess ratio tracking) and The weight of (stress safety) is adjusted to ensure dynamic response. During the stable phase, the weight is increased. (Energy consumption) and Weighting (smoothness) optimizes economy and lifespan.
[0121] In one specific implementation:
[0122] (I) First, construct the state space and action space: The state space S includes real-time operating parameters such as cathode and anode inlet mass flow rate, pressure, gas excess ratio, load current, and stack temperature; the action space A includes air compressor speed, hydrogen circulation pump speed, and valve opening degree. Then, establish an independent Critic network for each control objective. There are m actors in total; an Actor network is established simultaneously. Output control actions; then initialize the experience replay pool D to store state transition samples. Finally, initialize a set of preference vectors. , which represents the weight allocation for different control objectives.
[0123] (ii) First, use the current control strategy. Interact with the proton exchange membrane fuel cell gas supply system to collect the real-time status of the proton exchange membrane fuel cell gas supply system. Then, output the action according to the strategy. Execute and obtain multi-objective rewards and the next state Finally, the sample Transfer to experience pool D.
[0124] (III) Randomly sample a batch of data from the experience pool D, and calculate the target Q value for each target i:
[0125]
[0126] in , Let the grid parameters be the target. Then, calculate the loss function for each Critic grid and update the parameters:
[0127]
[0128]
[0129] (iv) Calculate the weighted Q value:
[0130]
[0131] Then, the Actor network is computed by minimizing the loss:
[0132]
[0133] Finally, the Actor parameters are updated using gradient descent:
[0134]
[0135] (v) Adjust the preference vector based on the current strategy performance. It is used to explore different regions of the Pareto front and periodically resamples the preference vector to ensure comprehensive coverage of the multi-objective trade-off space.
[0136]
[0137] (vi) If an adaptive entropy coefficient is used The update is as follows:
[0138]
[0139]
[0140] Then, the Critic network for each target is soft-updated:
[0141]
[0142] (vii) Maintaining the set of non-dominated strategies during training. Meanwhile, in actual control, based on current operational preferences Choose the best control strategy:
[0143]
[0144] (viii) Repeat steps two through seven above until the control strategy converges or the preset number of training steps are reached, and finally achieve multi-objective collaborative optimization control of the proton exchange membrane fuel cell gas supply system.
[0145] The above are merely preferred embodiments of the present invention and do not constitute any limitation on the scope of protection of the present invention; all technical solutions formed by equivalent transformations or equivalent substitutions fall within the scope of protection of the present invention; the parts of the present invention not described in detail are well known to those skilled in the art.
Claims
1. A control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm, characterized in that, The control method for the proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm includes the following steps: Step 1: Establish a model of the proton exchange membrane fuel cell gas supply system; The proton exchange membrane fuel cell gas supply system model includes a cathode flow channel model and an anode flow channel model. Step 2: Obtain data on the actual state of the fuel cell gas supply system; Step 3: Control the proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm.
2. The control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm according to claim 1, characterized in that, In step 1, the mass flow rate equation of the cathode flow channel model is expressed as: ; in The mass of oxygen at the cathode of a proton exchange membrane fuel cell. The mass of nitrogen at the cathode of a proton exchange membrane fuel cell. , These represent the mass change rates of oxygen and nitrogen at the cathode, respectively; This represents the mass flow rate of air entering the cathode, which is affected by the air compressor speed S. comp control; This means that the oxygen content in the air is approximately 21% by mass. This means that nitrogen in the air accounts for approximately 78% by mass. The mass flow rate at the cathode outlet; The total mass of the cathode outlet. , , , These represent the masses of oxygen, nitrogen, and water vapor at the cathode outlet, respectively. The mass flow rate of oxygen consumed in the reaction inside the cathode, where , , , , These are the number of individual cells in the battery stack, the stack current, the number of electrons transferred per mole of O2, the Faraday constant, and the molar mass of oxygen, respectively. The cathode pressure is: ; In the formula: For cathode pressure, The cathode temperature, This indicates the mass of oxygen at the cathode. Indicates the mass of nitrogen gas at the cathode. The gas constant representing oxygen. Where R is the gas constant. Here is the molar mass of oxygen. The gas constant of nitrogen. Where R is the gas constant. Here is the molar mass of nitrogen. For water vapor in Saturation pressure at temperature This indicates the volume of the cathode in the fuel cell stack; The oxygen excess ratio is used to measure the oxygen supply at the cathode, and is defined as the ratio of the oxygen supply to the amount of reactant. 。 3. The control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm according to claim 1, characterized in that, In step 1, the mass flow rate equation of the anode can be expressed as: ; in The mass of hydrogen gas at the anode of a proton exchange membrane fuel cell. This indicates the rate of mass change of hydrogen at the anode; This represents the mass flow rate of hydrogen entering the anode, which is affected by the rotational speed S of the hydrogen circulation pump. pump and the opening degree of the hydrogen supply circuit valve control; The mass flow rate at the anode outlet; The total mass of the anode outlet. ; , These represent the masses of hydrogen and water vapor at the anode outlet, respectively. The mass flow rate of hydrogen consumed in the reaction inside the anode. , , , , These are the number of individual cells in the battery stack, the stack current, the number of electrons transferred per mole of H2, the Faraday constant, and the molar mass of hydrogen, respectively. The anode pressure is: ; In the formula: This is the anode pressure. This is the anode temperature. Indicates the mass of hydrogen gas at the anode. The gas constant of hydrogen. Where R is the gas constant. Here is the molar mass of hydrogen. Water vapor is in Saturation pressure at temperature This indicates the volume of the anode in the fuel cell stack; Defined as the ratio of hydrogen supply to reaction amount: 。 4. The control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm according to claim 1, characterized in that, In step 2, the actual state data of the fuel cell gas supply system includes the oxygen / hydrogen excess ratio, cathode / anode pressure, oxygen / hydrogen inlet mass flow rate, stack current, and voltage.
5. The control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm according to claim 1, characterized in that, Step 3 also includes the following steps: Step 31: Construct a state space that includes the inlet flow rate and pressure of the cathode and anode, the excess gas ratio, the load current and the stack temperature, as well as an action space that includes the speed of the cathode air compressor, the speed of the anode circulating pump, the opening degree of the cathode / anode valve and the opening degree of the anode supply valve. Initialize the experience replay pool and establish the following core components: An Actor network for outputting control actions; multiple Critic evaluation networks and their corresponding target networks independently set for each control objective; and a set of preference weight vectors for weighing the importance of different objectives; Step 32, convert the current state space Input an Actor network, output the action space By applying the action space to the fuel cell gas supply system environment, a multi-objective reward function is obtained. With the next state space ; Interaction samples Store the data in the experience replay pool and repeat this step until enough training data has been accumulated. Step 33: Randomly select a batch of historical interaction samples from the experience replay pool as training data for this network parameter update; Step 34, for each target in the batch sample : State space Inputting the Actor network yields the action space. Enter the first The cumulative reward estimate for each Critic evaluation network is obtained. Calculate the target Q value of the objective; update the parameters of all m Critic evaluation networks by minimizing the loss function of each Critic network using gradient descent. Step 35, based on the current preference vector The overall Q value is obtained by weighted summation of the Q values of each objective. By maximizing the aggregated Q value and introducing a policy entropy regularization term, the loss function of the Actor network is constructed, and its parameters are updated using the gradient ascent method, thereby driving the control policy to optimize along the Pareto improvement direction. Step 36: Based on the performance of the current strategy on each objective, adaptively adjust the preference vector w through gradient update to explore different regions of the Pareto front; periodically resample the preference vector to ensure full coverage of the trade-off space between multiple objectives; Step 37: For each target network corresponding to the Critic evaluation network, the parameters are slowly synchronized to the latest parameters of the corresponding evaluation network using a soft update method to enhance training stability; if the adaptive entropy coefficient is enabled, the temperature parameter is dynamically adjusted according to the actual exploration degree of the strategy. Step 38: During training, continuously evaluate the policy performance corresponding to different preference vectors and maintain a set of Pareto optimal policies that contain all non-dominant policies; during actual deployment, select the most suitable control policy from this set based on the real-time target preferences of the system. Step 39: Repeat steps 32 to 38 to iteratively optimize each network parameter using the continuously accumulated interaction data until the control strategy performance converges.
6. The control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm according to claim 5, characterized in that, In step 31, the state space defined based on the dynamic characteristics of the proton exchange membrane fuel cell gas supply system is: ; in , The excess ratio of oxygen to hydrogen; , These are the pressures at time t at the cathode and anode, respectively; , These are the inlet mass flow rates of air and hydrogen at time t, respectively. , These represent the current and voltage of the proton exchange membrane fuel cell stack, respectively. This refers to the rate of change or the deviation from the set value.
7. The control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm according to claim 5, characterized in that, In step 31, the action space defined based on the dynamic characteristics of the proton exchange membrane fuel cell gas supply system is: ; ; This is the maximum hydrogen circulation pump speed. This is the maximum air compressor speed. and All are determined by the power of the fuel cell stack.
8. The control method for a proton exchange membrane fuel cell gas supply system based on the MOSAC algorithm according to claim 5, characterized in that, In step 32, the multi-objective reward function includes: Excess ratio tracking reward: ; in: , Indicates the target excess ratio of oxygen. Indicates the target excess ratio of hydrogen; express ; These represent the weighting coefficient and the sensitivity parameter, respectively. Smoothing penalties for each actuator: ; Indicates the weighting coefficient. This indicates the change in air compressor speed. This indicates the changes in the hydrogen circulation pump. This indicates the change in the opening degree of the cathode valve. This indicates the change in the opening degree of the oxygen valve. This indicates the change in the opening degree of the hydrogen supply valve; Stress-related safety penalties: ; Indicates the weighting coefficient. Indicates the maximum pressure at the cathode. Indicates the maximum pressure at the anode; Energy efficiency and emissions penalties: ; Indicates the weighting coefficient. This indicates the mass flow rate of excess oxygen that is discharged. This indicates the mass flow rate of the excess hydrogen gas that is discharged. The multi-objective reward function is as follows: To over-track rewards, For smoothing penalties for each actuator, As a form of stress-based safety punishment, Penalties for energy efficiency and emissions.