A control strategy of fuel cell multi-nozzle ejector based on MOHSAC algorithm

By adopting a control strategy for fuel cell multi-nozzle ejectors based on the MOHSAC algorithm, the problem of coordinating discrete nozzle position switching and continuous main valve flow regulation in multi-nozzle ejector systems is solved, achieving multi-objective adaptive optimization and improving the control effect of proton exchange membrane fuel cells.

CN122455844APending Publication Date: 2026-07-24HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUNAN UNIV
Filing Date
2026-04-30
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing multi-nozzle ejector control schemes face challenges in coordinating discrete nozzle position switching and continuous main valve flow regulation, and the strong coupling of multiple variables makes it difficult to achieve the best balance in multi-objective conflict control.

Method used

A control strategy for a fuel cell multi-nozzle ejector based on the MOHSAC algorithm is adopted. By constructing an actuator characteristic sub-model, a multi-objective state space, and a hybrid action space, a hierarchical multi-objective reward function is designed. The control strategy is trained using a multi-objective hybrid action soft actor-commentator algorithm to establish a Pareto optimal strategy set and achieve coordinated control of nozzle position and main proportional valve opening.

Benefits of technology

It effectively reduces model errors in traditional methods, improves the robustness and reliability of control strategies, and achieves rapid and accurate tracking of hydrogen excess ratio under all operating conditions, stable and controlled anode pressure, efficient ejector operation, and low actuator wear.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122455844A_ABST
    Figure CN122455844A_ABST
Patent Text Reader

Abstract

The application discloses a fuel cell multi-nozzle ejector control strategy based on a MOHSAC algorithm, relates to the technical field of proton exchange membrane fuel cell anode gas supply system control, and establishes an actuator characteristic submodel, obtains a mapping relationship from a discrete nozzle gear to a continuous valve opening mixed action to a total flow of primary flow through offline calibration; obtains actual state data of the ejector, constructs a multi-target state space and a mixed action space; constructs a layered multi-target reward function covering excess ratio tracking, pressure safety, ejector ratio efficiency and actuator life; establishes independent double Critic networks and mixed Actor networks based on the MOHSAC algorithm to perform iterative updating, and maintains a Pareto optimal strategy set; and in the deployment stage, calls a matching strategy according to a running condition to output a cooperative control instruction. The application realizes global cooperative optimization of the mixed action under multi-target conflicts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of control technology for anode gas supply systems in proton exchange membrane fuel cells, and more specifically, to a control strategy for multi-nozzle ejectors in fuel cells based on the MOHSAC algorithm. Background Technology

[0002] In proton exchange membrane fuel cell anode gas supply systems, multi-nozzle ejectors significantly expand the high-efficiency operating range and improve hydrogen recirculation capability under all operating conditions compared to single-nozzle ejectors. However, existing multi-nozzle ejector control schemes mostly employ single-closed-loop PID regulation or segmented control based on a fixed power threshold, which suffers from the following insurmountable technical limitations:

[0003] First, multi-nozzle ejector systems face the hybrid characteristics of discrete nozzle switching and continuous main valve flow regulation. Existing methods typically employ manual segmented control, failing to achieve dynamic coordination and matching between discrete and continuous actions. This results in severe pressure shocks and flow fluctuations that are highly likely to occur during operating condition switching or transitions.

[0004] Secondly, the anode gas supply system is characterized by strong coupling of multiple variables such as anode pressure, hydrogen excess ratio, and ejector ratio, dynamic and time-varying load, and strong system nonlinearity. Traditional single-target trackers or control logic relying on offline fixed parameters lack global optimization and adaptive capabilities, cannot handle conflicting control objectives (such as tracking accuracy, system safety, operating efficiency, and actuator lifespan), and find it difficult to find the optimal balance point among multiple objectives. Summary of the Invention

[0005] The purpose of this invention is to provide a control strategy for a fuel cell multi-nozzle ejector based on the MOHSAC algorithm, so as to solve the problem of difficulty in coordinating discrete nozzle position switching and continuous main valve flow regulation in the prior art, as well as the multi-objective conflict control problem caused by strong coupling of multiple variables.

[0006] The technical solution of this invention is: to provide a control strategy for a fuel cell multi-nozzle ejector based on the MOHSAC algorithm, comprising the following steps:

[0007] A characteristic sub-model of the ejector actuator in the anode system of a proton exchange membrane fuel cell was established. The mapping relationship between the mixed action inputs of discrete nozzle position and continuous valve opening and the total primary flow rate was obtained through offline calibration.

[0008] The actual state data of the multi-nozzle ejector in the anode gas supply system of a proton exchange membrane fuel cell are obtained to construct a multi-objective state space and a hybrid action space. The hybrid action space includes the discrete action of the multi-nozzle ejector opening position and the continuous action of the main proportional valve opening command.

[0009] A hierarchical multi-objective reward function is constructed, which includes a reward term for hydrogen excess ratio tracking, a penalty term for anode pressure safety, a penalty term for minimum ejection ratio efficiency guarantee, and a penalty term for actuator lifetime protection.

[0010] Control policy training is based on the multi-objective hybrid action soft actor-critic MOHSAC algorithm: an independent dual-critic evaluation network is established for each control objective, and an Actor network is established to output discrete gear distribution probabilities and continuous valve opening Gaussian distribution parameters in parallel; the network is iteratively updated using the multi-objective state space, hybrid action space and hierarchical multi-objective reward function, and the preference weight vector is periodically resampled during training to maintain the Pareto optimal policy set;

[0011] During the actual deployment phase, based on the current system operating conditions, a pre-trained strategy matching the current target preference is called from the Pareto optimal strategy set, and a coordinated control command for the nozzle position and the opening of the main proportional valve is output.

[0012] In any of the above technical solutions, the actuator characteristic sub-model is further specified as follows:

[0013] ;

[0014] in, This refers to the primary flow rate of the anode ejector. The number of nozzle solenoid valves to open; The single-nozzle flow coefficient of the main proportional valve characterizes the nonlinear mapping relationship between the effective opening of the main proportional valve and the corresponding primary flow reference flow when a single nozzle is open. Represents the temperature correction factor; effective opening degree of the main proportional valve. The conversion formula is:

[0015] ;

[0016] in, The initial opening of the main proportional valve. The dead zone of the proportional valve is obtained through calibration; the opening degree used in subsequent calculations in the actuator characteristic sub-model refers to the effective opening degree. .

[0017] In any of the above technical solutions, further, the state vector in the multi-objective state space... Defined as:

[0018] ;

[0019] in, The secondary flow rate of the ejector. This is the anode inlet pressure. For the current of the proton exchange membrane fuel cell stack, The current nozzle setting. The hydrogen excess ratio, The hydrogen temperature at the anode inlet. The rate of change of current, The effective opening degree of the main proportional valve.

[0020] In any of the above technical solutions, the specific calculations of each item in the hierarchical multi-objective reward function are as follows:

[0021] Rewards for tracking hydrogen excess ratio for:

[0022] ;

[0023] in To track the amplitude coefficient, For sensitivity parameters, This represents the actual excess hydrogen ratio. The target excess ratio;

[0024] Penalties for anode pressure safety For: when the anode inlet pressure hour, ;when hour, ;when hour, ;in This is the upper limit of the safe anode pressure. This is the lower limit of the safe anode pressure. This is the pressure penalty coefficient;

[0025] Penalty for guaranteeing minimum efficiency against ejection ratio For: when the actual ejection ratio hour, ;when hour, ;in The minimum permissible ejection ratio threshold, This is the penalty coefficient;

[0026] Penalty items for actuator life protection For: when hour, ;when hour, ;in To activate the multi-nozzle ejector at the current moment. The gear position that was engaged at the previous moment. This is the penalty coefficient.

[0027] In any of the above technical solutions, the network is further iteratively updated using a multi-objective state space, a hybrid action space, and a hierarchical multi-objective reward function, including updating the Critic evaluation network:

[0028] Sample historical interaction samples from the experience replay pool, and then set the next state. Input the current Actor network to obtain the prediction of the next action ;

[0029] Input it together with the state into the first A dual-crit target network with multiple targets is used to calculate the target Q-value by incorporating an entropy regularization term. :

[0030] ;

[0031] in, For the first The reward value for each objective. For the formula coefficients, For the first The first goal Each Critic target network parameter, For adaptive entropy coefficient, The logarithmic probability of the mixed action;

[0032] Calculate the loss function for each Critic evaluation network and update the network parameters using gradient descent.

[0033] In any of the above technical solutions, further, the iterative update of the network using the multi-objective state space, the hybrid action space, and the hierarchical multi-objective reward function also includes updating the Actor network:

[0034] Based on the current preference weight vector The combined Q-value is obtained by taking the minimum value of the two Critic evaluation network outputs for each target and then performing a linear weighted sum. ;

[0035] Construct an Actor network loss function that includes both discrete and continuous policy log probabilities. :

[0036] ;

[0037] in, These are historical interaction samples sampled from the experience replay pool. Let be the probability distribution of the discrete gear position. This represents the probability distribution of continuous valve opening.

[0038] The loss function is minimized using gradient descent. To update the Actor network parameters .

[0039] In any of the above technical solutions, furthermore, during the network iterative update process, the adaptive entropy coefficient is also included. Update:

[0040] Calculate the average entropy of the current policy ;

[0041] Constructing an adaptive entropy coefficient update loss function ,in The target entropy is preset;

[0042] The adaptive entropy coefficient is updated using gradient descent. .

[0043] In any of the above technical solutions, further, during the actual deployment phase, based on the current system operating conditions, a pre-trained policy matching the current target preference is invoked from the Pareto optimal policy set, specifically as follows:

[0044]

[0045] in, For the target strategy to be invoked, This is the Pareto optimal policy set. To control the total number of targets, To dynamically adjust the system based on its current operating conditions The preference weights of each objective For strategy Next The Q value of each objective.

[0046] The beneficial effects of this invention are:

[0047] To address the challenge of coordinated hybrid actions, this invention constructs a characteristic sub-model of a multi-nozzle ejector actuator using an offline overall calibration method. This establishes a precise mapping relationship between the hybrid action inputs of discrete nozzle positions and continuous valve openings and the total primary flow rate. This method effectively reduces model errors caused by traditional reliance on product manual parameters, provides accurate actuator input-output conversion data for the algorithm, and significantly improves the robustness and convergence reliability of the control strategy when migrating from the simulation environment to the actual physical system.

[0048] The control strategy of this invention relies on a modular, split-type multi-nozzle ejector, which consists of a multi-nozzle injection assembly and a universal mixing chamber assembly, connected by a threaded seal. This design not only facilitates easy disassembly and assembly, and subsequent maintenance and component replacement, but also allows for precise adjustment of the internal flow field of the ejector by changing the relative mixing position of the primary and secondary flows through adjusting the thread fit parameters. This further optimizes the ejection performance and overall operating efficiency of the multi-nozzle ejector.

[0049] To address the bottleneck in multi-objective conflict control, this invention constructs a multi-objective state space and a hybrid action space adapted to the hybrid action characteristics of multi-nozzle ejectors. It also innovatively designs a hierarchical multi-objective reward function encompassing hydrogen excess ratio tracking accuracy, anode pressure safety margin, minimum ejection ratio efficiency guarantee, and actuator switching lifespan protection. This design scientifically quantifies competing physical indicators, forcing the agent to balance system safety and lifespan while ensuring tracking accuracy, effectively preventing a single objective from dominating the optimization process due to magnitude differences.

[0050] This invention deeply integrates Pareto multi-objective optimization theory, hybrid action space processing strategy, and soft actor-critic (SAC) algorithm to propose the MOHSAC algorithm. By establishing an independent dual-critic evaluation network for each control objective and periodically resampling the preference weight vector during training, a Pareto optimal policy set is maintained. In actual deployment, the system can adaptively call the control policy with the best matching preference weight according to the dynamic changes in load, realizing global collaborative optimization of nozzle gear switching timing and continuous opening of the main proportional valve. Under all operating conditions, it simultaneously ensures rapid and accurate tracking of hydrogen excess ratio, stable and controlled anode pressure, efficient ejector operation, and low actuator wear. Attached Figure Description

[0051] The advantages of the above and additional aspects of the present invention will become apparent and readily understood in the description of the embodiments in conjunction with the following drawings, wherein:

[0052] Figure 1 This is a flowchart of the algorithm for a fuel cell multi-nozzle ejector control strategy based on the MOHSAC algorithm according to an embodiment of the present invention.

[0053] Figure 2 This is a schematic diagram of a multi-nozzle component of a fuel cell multi-nozzle ejector based on a MOHSAC algorithm control strategy according to an embodiment of the present invention, including a three-dimensional view and a side view.

[0054] Figure 3 This is a schematic diagram of a general cavity assembly for a multi-nozzle ejector based on a MOHSAC algorithm-based control strategy for a fuel cell multi-nozzle ejector according to an embodiment of the present invention.

[0055] Figure 4 This is a schematic diagram of a multi-nozzle ejector assembly based on a MOHSAC algorithm-based control strategy for a fuel cell multi-nozzle ejector according to an embodiment of the present invention.

[0056] Figure 5 This is a schematic diagram of the application of a fuel cell multi-nozzle ejector control strategy based on the MOHSAC algorithm according to an embodiment of the present invention in a proton exchange membrane fuel cell system.

[0057] Figure 6 This is a flowchart illustrating the MOHSAC algorithm, a control strategy for a fuel cell multi-nozzle ejector based on the MOHSAC algorithm according to an embodiment of the present invention, in a PEMFC anode multi-nozzle ejector. Detailed Implementation

[0058] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.

[0059] In the following description, many specific details are set forth in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0060] This embodiment provides a control strategy for a fuel cell multi-nozzle ejector based on the MOHSAC algorithm. This strategy aims to solve the problem of coordinating the mixed actions of discrete nozzle position switching and continuous main valve flow regulation in a multi-nozzle ejector, as well as the multi-objective conflict control bottleneck caused by the strong coupling of multiple variables such as anode pressure, hydrogen excess ratio, and ejection ratio.

[0061] The control strategy described above in this invention relies on the anode gas supply system of a proton exchange membrane fuel cell, which adopts a modular, split-type multi-nozzle ejector.

[0062] like Figures 2 to 4 As shown, the multi-nozzle ejector is divided into two parts: a multi-nozzle injection assembly and a universal mixing chamber assembly. The two assemblies are connected by a threaded seal, which makes disassembly and assembly convenient and facilitates later maintenance and component replacement. At the same time, the relative mixing position of the primary flow and the secondary flow can be changed by adjusting the thread fit parameters, so as to achieve fine adjustment of the internal flow field of the ejector.

[0063] like Figure 5As shown, based on the above hardware architecture, the workflow of the anode gas supply system is as follows: The high-pressure working gas output from the hydrogen storage tank enters the main proportional valve through the inlet valve. The primary working gas flow rate is precisely controlled by continuously adjusting the opening of the main proportional valve. The primary flow rate at the outlet of the main proportional valve is collected in real time by a flow sensor. Subsequently, the hydrogen enters each nozzle branch of the multi-nozzle ejector. The controller determines the number of nozzle solenoid valves to be opened according to the real-time power requirements of the fuel cell stack, thereby adjusting the primary flow rate. The high-pressure hydrogen accelerated by the nozzles forms a low-pressure zone inside the ejector, drawing in unreacted hydrogen to form a mixed gas. The ejector outlet pressure is monitored in real time by a pressure sensor. The mixed gas finally enters the proton exchange membrane fuel cell stack to participate in the electrochemical reaction. Unconsumed anode tail gas after the fuel cell reaction is discharged from the hydrogen outlet. After liquid water is separated by a water-gas separator, it is divided into two streams: one stream is used as a secondary flow to circulate hydrogen. Its flow rate is collected in real time by a flow sensor and sent to the secondary flow inlet of the multi-nozzle ejector. After mixing with the primary flow working gas, it re-enters the fuel cell to participate in the reaction, realizing the closed-loop recycling of hydrogen; the other stream of waste gas and excess hydrogen is periodically discharged from the system through a tail valve to maintain the stability of the anode gas composition.

[0064] To achieve high-precision intelligent control, this invention first establishes a characteristic sub-model of the ejector actuator in the anode system of a proton exchange membrane fuel cell, to accurately convert the mixed action input of discrete nozzle positions and continuous valve openings into the total primary flow rate. This model satisfies the following assumptions: 1. All gases are ideal gases; 2. Adiabatic flow; 3. Total mass and total energy are conserved before and after mixing of the primary and secondary flows; 4. Nozzle characteristics are consistent; the geometry and fluid characteristics of the three parallel nozzles are identical; the flow characteristics of a single nozzle when fully open are consistent; the total flow area when n nozzles are open is n times that of a single nozzle; 5. Pressure loss, volumetric effect, and heat loss in connecting pipelines are ignored; the outlet pressure of the main proportional valve equals the inlet pressure of the nozzle; the outlet pressure of the ejector equals the inlet pressure of the anode chamber; 6. The total primary flow rate monotonically increases with the number of open nozzle solenoid valves, meaning the total primary flow rate when n nozzles are open is strictly greater than the total primary flow rate when n-1 nozzles are open (n=2,3), and there is no abnormal operating condition where opening more nozzles does not increase or even decreases the primary flow rate.

[0065] To overcome the problem of significant deviations between the model and the actual system caused by relying on product manual parameters in traditional methods, this invention adopts an offline overall calibration method to directly establish the mapping relationship between the discrete positions of the nozzle solenoid valve, the continuous opening degree of the main proportional valve, and the primary flow rate of the ejector. The actuator characteristic sub-model is as follows:

[0066] ;

[0067] in, The primary flow rate of the anode ejector in a proton exchange membrane fuel cell; The number of nozzle solenoid valves to open; The single-nozzle flow coefficient of the main proportional valve, obtained through offline calibration, characterizes the effective opening degree of the main proportional valve when a single nozzle is open. The nonlinear mapping relationship between the primary flow reference flow rate and the corresponding primary flow reference flow rate; This represents the temperature correction factor.

[0068] Because dead zones (where there is opening but no flow) often exist in main proportional valves, it is necessary to convert the original opening to an effective opening.

[0069] ;

[0070] in, The dead zone of the proportional valve is obtained from calibration, and all subsequent formulas refer to it. All represent effective opening degree .

[0071] The temperature correction factor can usually be obtained using the following formula:

[0072] ;

[0073] In the formula For standard calibration temperature, This represents the actual hydrogen inlet temperature.

[0074] The main calibration process is as follows: First, the dead zone of the main proportional valve is calibrated under standard calibration conditions (temperature 25℃) to obtain... Then, the flow coefficient of a single nozzle was analyzed. Calibration: Open any nozzle solenoid valve, based on the calibrated dead zone. Increase the effective opening degree of the main proportional valve The calibration data is divided into 50 uniformly spaced nodes from 0 to 1 with a step size of 0.02. This process iterates through all 50 nodes to obtain a set of discrete data. After calibration, a polynomial is used to fit the discrete data points to obtain a continuous single-nozzle flow coefficient function. Then, the temperature coefficient and the consistency of flow rate across multiple nozzles are verified to complete the calibration.

[0075] Based on this, a hydrodynamic model of the ejector is further established to obtain quantitative calculations of the ejector ratio and outlet flow rate under different operating conditions. The relationship between the primary flow and the total ejector flow rate is established, as shown in the following model:

[0076] ;

[0077] ;

[0078] In the formula For secondary flow rate, For a single flow, This refers to the ejector outlet flow rate. This is the ejection ratio.

[0079] Simultaneously, a model for the excess hydrogen ratio in the anode system of a proton exchange membrane fuel cell was established:

[0080] ;

[0081] in, The value is the ejector outlet flow rate; the denominator is the flow rate of hydrogen consumed in the reaction inside the anode, where... , , , , These are the number of individual cells in the battery stack, the stack current, the number of electrons transferred per mole of H2, the Faraday constant, and the molar mass of hydrogen.

[0082] After completing the construction of the above one-dimensional simulation model, this invention obtains the actual state data of the multi-nozzle ejector in the anode gas supply system of a proton exchange membrane fuel cell, and controls the ejector based on the multi-objective hybrid motion soft actor-commentator (MOHSAC) algorithm.

[0083] like Figure 1 The diagram shows the overall workflow of the algorithm. The core of the algorithm lies in constructing a multi-objective state space and a hybrid action space that adapt to the characteristics of mixed actions, and designing a hierarchical multi-objective reward function.

[0084] Specifically, the state vector is defined based on the dynamic characteristics of the ejector. as follows:

[0085] ;

[0086] The physical meaning and function of each state variable are as follows:

[0087] The excess ratio of hydrogen is the direct feedback quantity of the primary control objective. The agent must observe this value in real time to evaluate the merits of the current action and form a closed-loop control capability.

[0088] The secondary flow rate of the ejector is calculated in real time by comparing the secondary flow with the primary flow. The agent needs to know whether the current ejection capacity is below the safety threshold so that it can increase the reflux by switching nozzles or increasing valve opening when necessary to prevent hydrogen waste.

[0089] It is the pressure at the anode inlet. The anode pressure directly affects the safety of the pressure difference across the membrane electrode of the fuel cell stack, and is strongly nonlinearly coupled with the primary flow rate and ejection ratio of the ejector. Providing this state enables the agent to sense the current system pressure and avoid pressure overshoot caused by valve adjustment or nozzle switching.

[0090] This is the hydrogen temperature at the anode inlet. The hydrogen supply temperature will fluctuate with the operating conditions. Providing this state allows the agent to adapt to the effect of temperature on the flow gain and improves the model transfer robustness.

[0091] The current of the proton exchange membrane fuel cell stack is the load current, which is the reference for flow demand. The agent must know the current value in order to estimate the required target total flow and make corresponding valve and nozzle decisions.

[0092] The current change rate provides feedforward prediction capability. When the load changes rapidly, the introduction of the current change rate can enable the agent to sense the upcoming surge in flow demand in advance, thereby increasing the valve opening and reducing tracking lag and overshoot.

[0093] The current nozzle setting is used as the state input, while the previous setting is used as the state input. This satisfies the Markov property and enables the agent to learn the switching cost, avoiding high-frequency oscillation between two settings and improving valve life.

[0094] This is the effective opening degree of the main proportional valve. The valve opening degree at the previous moment affects the response delay and flow starting point of the current actuator. Inputting this state helps the agent output a smooth and continuous adjustment amount, avoiding severe valve jitter.

[0095] Defined action vector The action constraints are as follows:

[0096] ;

[0097] in The multi-nozzle ejector has three opening positions (discrete action), corresponding to three nozzle combination modes. The higher the position, the more nozzles are opened and the larger the throat flow area. This is the normalized opening command for the main proportional valve (continuous action). 0 indicates that the valve is completely closed, and 1 indicates that the valve is completely open.

[0098] Based on Pareto theory, this invention decomposes the control objective into four competing sub-objectives, and constructs a hierarchical multi-objective reward function as follows:

[0099] (a) Excessive tracking rewards :

[0100] ;

[0101] in Represents the tracking amplitude coefficient; This represents a sensitivity parameter that controls how steeply the reward decays with tracking error; This represents the actual excess hydrogen ratio; This represents the target excess ratio; the reward is maximized at the target value. As the error increases, the reward decreases smoothly, guiding the agent to prioritize tracking accuracy.

[0102] (ii) Anode pressure hazard penalty :

[0103] ;

[0104] in Represents the anode inlet pressure; This represents the upper limit of safe anode pressure; This represents the lower limit of the safe anode pressure. This represents the pressure penalty coefficient; when the anode pressure exceeds the upper safety limit or falls below the lower safety limit, a linear penalty proportional to the extent of the exceedance is applied, forcing the agent to avoid dangerous conditions.

[0105] (iii) Minimum efficiency penalty for ejection ratio :

[0106] ;

[0107] in Represents the actual ejection ratio; This represents the minimum permissible ejection ratio threshold; This represents the penalty coefficient; when the ejection ratio is below the minimum efficiency threshold, a linear negative penalty is applied, and the penalty amount increases as the ejection ratio decreases, thereby ensuring the ejector's minimum ejection capacity for circulating hydrogen and preventing hydrogen waste.

[0108] (iv) Actuator life penalty :

[0109] ;

[0110] in This indicates the current position of the multi-nozzle ejector. This indicates the gear position at the previous moment; This represents the penalty coefficient; this penalty directly suppresses frequent switching of discrete gears and extends the mechanical life of the solenoid valve.

[0111] The resulting multi-objective reward vector is: All reward items are normalized to the [0,1] interval to ensure the fairness of the weighted summation of multiple objectives and to prevent differences in magnitude from causing one objective to dominate the optimization process.

[0112] Using a multi-objective hybrid action soft actor-critic algorithm, a set of non-dominated policies is learned, forming the Pareto optimal policy set. Each strategy can correspond to a preference vector. ,satisfy Simultaneously optimize the weighted objective:

[0113] ;

[0114] In this way, the preference weights are dynamically adjusted according to different operating conditions during the operation phase (including startup, load change, steady state, etc.) to achieve adaptive coupling between multiple objectives.

[0115] like Figure 6 As shown, the specific execution flow of the MOHSAC algorithm during the algorithm training and deployment phases is as follows:

[0116] First, an independent Critic evaluation network is established for each control objective. Where i=1,2,3,4 correspond to the four sub-objectives mentioned above, and each objective employs a dual Critic network structure to suppress overestimation; simultaneously, an Actor network is established. The network outputs the probability distribution of discrete gear positions in parallel. And the mean and standard deviation parameters of the continuous valve opening corresponding to each possible gear position. Initialize the experience replay pool D, and initialize a set of weight preference vectors. ,satisfy , which represents the weight allocation for different control objectives.

[0117] Use the current control policy Interact with the system to obtain the normalized real-time state vector. Specific actions are obtained by sampling from the mixed action distribution. After motion smoothing filtering, it is applied to the anode ejector system, and after execution, a multi-target reward vector is obtained. and the next state Complete interaction sample Store in experience replay pool D.

[0118] A batch of historical interaction samples were randomly sampled from the experience replay pool D. For each optimization objective i, the next state is... Input the current Actor network to obtain the prediction of the next action The state and the target's state are input into the dual-Critic target network of the i-th target, and the target's Q-value is calculated using the entropy regularization term:

[0119] ;

[0120] in For the network parameters of the j-th Critic target of the i-th target; The adaptive entropy coefficient; Let be the logarithmic probability of the mixed action.

[0121] Calculate the loss function for each Critic evaluation network and update the parameters using gradient descent:

[0122] ;

[0123] ;

[0124] Based on the current preference weight vector For each target, the minimum value of the outputs from the two Critic evaluation networks is taken, and then a linear weighted sum is performed to obtain the comprehensive Q value:

[0125] ;

[0126] To handle the mixed action space, an Actor network loss function is constructed that includes both discrete policy log probabilities and continuous policy log probabilities:

[0127] ;

[0128] Update the Actor network parameters using gradient descent: This drives the control strategy to optimize along the Pareto improvement direction specified by the current preference vector.

[0129] At the same time, calculate the average entropy of the current strategy. The adaptive entropy coefficients are updated through gradient descent. :

[0130] ;

[0131] ;

[0132] If the average entropy is lower than the preset target entropy Then increase To encourage exploration; conversely, to reduce it. To suppress randomness.

[0133] Subsequently, a soft update method is applied to the two Critic target networks corresponding to each target i, updating the coefficients. (Much less than 1) Slowly synchronize its parameters to the latest parameters of the corresponding evaluation network:

[0134] ;

[0135] During the training loop, after a fixed number of rounds, a new set of preference weight vectors is resampled using a uniform distribution. This ensures the algorithm can comprehensively explore the entire Pareto front region, from "aggressive tracking priority" to "conservative lifetime priority." During training, the policy performance corresponding to different preference vectors is continuously evaluated, and a Pareto-optimal policy set is maintained based on the Pareto non-dominance relationship of the actual cumulative rewards for each objective. .

[0136] During actual deployment, the fuel cell controller directly calls the pre-trained strategy from the set that best matches the current target preference based on the current system operating conditions (such as rapid load change phase or steady-state operation phase), without the need for online computation.

[0137] ;

[0138] For example, during the rapid load change phase, improve (Excess ratio tracking) and The weight of (stress safety) is adjusted to ensure dynamic response; during the stable phase, the weight is increased. (Ejection efficiency) and The weighting of actuator lifespan optimizes hydrogen utilization and equipment lifespan. Ultimately, it enables adaptive collaborative decision-making between the timing of multi-nozzle ejector gear switching and the continuous opening degree of the main proportional valve.

[0139] In summary, this invention proposes a control strategy for a fuel cell multi-nozzle ejector based on the MOHSAC algorithm, comprising the following steps:

[0140] A characteristic sub-model of the ejector actuator in the anode system of a proton exchange membrane fuel cell is established. The mapping relationship between the mixed action inputs of discrete nozzle position and continuous valve opening and the total primary flow rate is obtained through offline calibration.

[0141] The actual state data of the multi-nozzle ejector in the anode gas supply system of a proton exchange membrane fuel cell are obtained to construct a multi-objective state space and a hybrid action space. The hybrid action space includes the discrete actions of the multi-nozzle ejector opening position and the continuous actions of the main proportional valve opening command.

[0142] A hierarchical multi-objective reward function is constructed, which includes a reward term for hydrogen excess ratio tracking, a penalty term for anode pressure safety, a penalty term for minimum efficiency guarantee of ejector ratio, and a penalty term for actuator lifetime protection.

[0143] Control policy training is based on the multi-objective hybrid action soft actor-critic MOHSAC algorithm: an independent dual-critic evaluation network is established for each control objective, and an Actor network is established to output discrete gear distribution probabilities and continuous valve opening Gaussian distribution parameters in parallel; the network is iteratively updated using the multi-objective state space, hybrid action space and hierarchical multi-objective reward function, and the preference weight vector is periodically resampled during training to maintain the Pareto optimal policy set.

[0144] During the actual deployment phase, based on the current system operating conditions, a pre-trained strategy matching the current target preference is called from the Pareto optimal strategy set, and a coordinated control command for the nozzle position and the opening of the main proportional valve is output.

[0145] The steps in this invention can be adjusted, combined, or deleted according to actual needs.

[0146] The units in the device of the present invention can be merged, divided, or reduced according to actual needs.

[0147] In this invention, the terms "installation," "connection," "linking," and "fixing" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; "linking" can be a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of these terms in this invention according to the specific circumstances.

[0148] The shapes of the components in the accompanying drawings are schematic and may differ from their actual shapes. The drawings are only used to illustrate the principles of the present invention and are not intended to limit the present invention.

[0149] Although the invention has been disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely exemplary and not intended to limit the application of the invention. The scope of protection of the invention is defined by the appended claims and may include various modifications, alterations, and equivalents made to the invention without departing from the scope and spirit of the invention.

Claims

1. A control strategy for a fuel cell multi-nozzle ejector based on the MOHSAC algorithm, characterized in that, Includes the following steps: A characteristic sub-model of the ejector actuator in the anode system of a proton exchange membrane fuel cell was established. The mapping relationship between the mixed action inputs of discrete nozzle position and continuous valve opening and the total primary flow rate was obtained through offline calibration. The actual state data of the multi-nozzle ejector in the anode gas supply system of a proton exchange membrane fuel cell are obtained to construct a multi-objective state space and a hybrid action space. The hybrid action space includes the discrete action of the multi-nozzle ejector opening position and the continuous action of the main proportional valve opening command. A hierarchical multi-objective reward function is constructed, which includes a reward term for hydrogen excess ratio tracking, a penalty term for anode pressure safety, a penalty term for minimum ejection ratio efficiency guarantee, and a penalty term for actuator lifetime protection. Control policy training is based on the multi-objective hybrid action soft actor-critic MOHSAC algorithm: an independent dual-critic evaluation network is established for each control objective, and an Actor network is established to output discrete gear distribution probabilities and continuous valve opening Gaussian distribution parameters in parallel; the network is iteratively updated using the multi-objective state space, hybrid action space and hierarchical multi-objective reward function, and the preference weight vector is periodically resampled during training to maintain the Pareto optimal policy set; During the actual deployment phase, based on the current system operating conditions, a pre-trained strategy matching the current target preference is called from the Pareto optimal strategy set, and a coordinated control command for the nozzle position and the opening of the main proportional valve is output.

2. The fuel cell multi-nozzle ejector control strategy based on the MOHSAC algorithm as described in claim 1, characterized in that, The actuator characteristic sub-model is as follows: ; in, This refers to the primary flow rate of the anode ejector. The number of nozzle solenoid valves to open; The single-nozzle flow coefficient of the main proportional valve characterizes the nonlinear mapping relationship between the effective opening of the main proportional valve and the corresponding primary flow reference flow when a single nozzle is open. Represents the temperature correction factor; effective opening degree of the main proportional valve. The conversion formula is: ; in, The initial opening of the main proportional valve. The dead zone of the proportional valve is obtained through calibration; the opening degree used in subsequent calculations in the actuator characteristic sub-model refers to the effective opening degree. .

3. The fuel cell multi-nozzle ejector control strategy based on the MOHSAC algorithm as described in claim 1, characterized in that, State vector in multi-objective state space Defined as: ; in, The secondary flow rate of the ejector. This is the anode inlet pressure. For the current of the proton exchange membrane fuel cell stack, The current nozzle setting. The hydrogen excess ratio, The hydrogen temperature at the anode inlet. The rate of change of current, The effective opening degree of the main proportional valve.

4. The fuel cell multi-nozzle ejector control strategy based on the MOHSAC algorithm as described in claim 1, characterized in that, The specific calculations for each item in the hierarchical multi-objective reward function are as follows: Rewards for tracking hydrogen excess ratio for: ; in To track the amplitude coefficient, For sensitivity parameters, This represents the actual excess hydrogen ratio. The target excess ratio; Penalties for anode pressure safety For: when the anode inlet pressure hour, ;when hour, ;when hour, ;in This is the upper limit of the safe anode pressure. This is the lower limit of the safe anode pressure. This is the pressure penalty coefficient; Penalty for guaranteeing minimum efficiency against ejection ratio For: when the actual ejection ratio hour, ;when hour, ;in The minimum permissible ejection ratio threshold, This is the penalty coefficient; Penalty items for actuator life protection For: when hour, ;when hour, ;in To activate the multi-nozzle ejector at the current moment. The gear position that was engaged at the previous moment. This is the penalty coefficient.

5. The fuel cell multi-nozzle ejector control strategy based on the MOHSAC algorithm as described in claim 1, characterized in that, The network is iteratively updated using a multi-objective state space, a hybrid action space, and a hierarchical multi-objective reward function, including updates to the Critic evaluation network: Sample historical interaction samples from the experience replay pool, and then set the next state. Input the current Actor network to predict the next action ; Input it together with the state into the first A dual-crit target network with multiple targets is used to calculate the target Q-value by incorporating an entropy regularization term. : ; in, For the first The reward value for each objective. For the formula coefficients, For the first The first goal Each Critic target network parameter, For adaptive entropy coefficient, The logarithmic probability of the mixed action; Calculate the loss function for each Critic evaluation network and update the network parameters using gradient descent.

6. The fuel cell multi-nozzle ejector control strategy based on the MOHSAC algorithm as described in claim 5, characterized in that, The network is iteratively updated using the multi-objective state space, hybrid action space, and hierarchical multi-objective reward function, and the update of the Actor network is also included: Based on the current preference weight vector The combined Q-value is obtained by taking the minimum value of the two Critic evaluation network outputs for each target and then performing a linear weighted sum. ; Construct an Actor network loss function that includes both discrete and continuous policy log probabilities. : ; in, These are historical interaction samples sampled from the experience replay pool. Let be the probability distribution of the discrete gear position. This represents the probability distribution of continuous valve opening. The loss function is minimized using gradient descent. To update the Actor network parameters .

7. The fuel cell multi-nozzle ejector control strategy based on the MOHSAC algorithm as described in claim 6, characterized in that, The network iterative update process also includes adjusting the adaptive entropy coefficient. Update: Calculate the average entropy of the current policy ; Constructing an adaptive entropy coefficient update loss function ,in The target entropy is preset; The adaptive entropy coefficient is updated using gradient descent. .

8. The fuel cell multi-nozzle ejector control strategy based on the MOHSAC algorithm as described in claim 1, characterized in that, During the actual deployment phase, based on the current system operating conditions, a pre-trained policy matching the current target preference is invoked from the Pareto optimal policy set. The specific formula is as follows: in, For the target strategy to be invoked, This is the Pareto optimal policy set. To control the total number of targets, To dynamically adjust the system based on its current operating conditions The preference weights of each objective For strategy Next The Q value of each objective.