Fuel cell system control method, device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610801828.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-08
AI Technical Summary
[0002]针对燃料电池系统的现有控制策略通常将性能目标与寿命约束通过加权求和的方式耦合在一个目标函数中,然而性能与寿命之间往往存在固有的冲突关系,固定权重难以在全工况范围内实现二者的动态平衡,并且燃料电池系统在长期运行过程中会发生性能衰退,初始设定的权重系数随着燃料电池系统的老化逐渐失效,导致控制策略偏离最优折中
[0014] This application includes at least the following beneficial effects: By introducing a reinforcement learning strategy that considers both performance and lifespan as a dual reward mechanism, the operating parameter values of the fuel cell system are first obtained and input into the policy network of the reinforcement learning agent for analysis. Then, the fuel cell system is regulated according to the control command information of the fuel cell system obtained from the analysis. Next, new operating parameter values, efficiency impact parameter values, and damage impact parameter values of the fuel cell system are obtained. Subsequently, the performance reward is determined based on the efficiency impact parameter value of the fuel cell system, and the lifespan reward is determined based on the damage impact parameter value of the fuel cell system. Finally, the performance reward, lifespan reward, and the operating parameter values, new operating parameter values, and control command information of the fuel cell system are combined to form a training sample and stored in an experience pool used to assist in training the reinforcement learning agent. This can improve the control robustness of the fuel cell system.
Smart Images

Figure CN122716367A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fuel cell technology, and in particular to fuel cell system control methods, devices, electronic equipment and storage media. Background Technology
[0002] Existing control strategies for fuel cell systems typically couple performance targets and lifetime constraints into a single objective function through weighted summation. However, there is often an inherent conflict between performance and lifetime, and fixed weights make it difficult to achieve a dynamic balance between the two across the entire operating range. Furthermore, fuel cell systems experience performance degradation during long-term operation, and the initially set weight coefficients gradually become ineffective as the fuel cell system ages, causing the control strategy to deviate from the optimal compromise. Summary of the Invention
[0003] The main objective of this application is to propose a control method, device, electronic equipment, and storage medium for a fuel cell system, aiming to improve the control robustness of the fuel cell system by introducing a reinforcement learning strategy that takes into account both performance and lifespan.
[0004] To achieve the above objectives, one aspect of this application proposes a fuel cell system control method, the method comprising: The operating parameter values of the fuel cell system are obtained and input into the policy network in the reinforcement learning agent for analysis to obtain the control command information of the fuel cell system. Based on the control command information of the fuel cell system, the fuel cell system is regulated, and then new operating parameter values, efficiency impact parameter values, and damage impact parameter values of the fuel cell system are obtained. A performance bonus is determined based on the efficiency impact parameter value of the fuel cell system; and a lifetime bonus is determined based on the damage impact parameter value of the fuel cell system. Based on the performance reward, the lifetime reward, and the operating parameter values, new operating parameter values, and control command information of the fuel cell system, a training sample is generated and stored in the experience pool, which is used to assist in training the reinforcement learning agent.
[0005] Furthermore, the efficiency impact parameters of the fuel cell system include the net output power and hydrogen consumption of the fuel cell system; determining the performance bonus based on the efficiency impact parameters of the fuel cell system includes: The efficiency of the fuel cell system is determined based on the preset low calorific value of hydrogen, the net output power of the fuel cell system, and the hydrogen consumption. The performance bonus is determined based on the net output power, hydrogen consumption, and efficiency of the fuel cell system.
[0006] Furthermore, the damage impact parameters of the fuel cell system include the stack temperature, stack membrane water content, and stack anode-cathode voltage difference of the fuel cell system; determining the life bonus based on the damage impact parameters of the fuel cell system includes: The thermal damage increment is determined based on the preset frequency factor, preset activation energy, preset gas constant, preset control step size, and stack temperature of the fuel cell system. The humidity damage increment is determined based on the preset humidity damage coefficient, the preset control step size, and the water content of the fuel cell stack membrane. The differential pressure damage increment is determined based on the preset differential pressure damage coefficient, the preset control step size, and the differential pressure between the anode and cathode of the fuel cell system stack. The lifetime bonus is determined based on the thermal damage increment, the humidity damage increment, and the differential pressure damage increment.
[0007] Furthermore, the method also includes: A batch of training samples is extracted from the experience pool and the data is split to form a first batch of training sub-samples and a second batch of training sub-samples. Each training sub-sample in the first batch of training sub-samples contains a performance reward as well as the operating parameter values, new operating parameter values and control command information of the fuel cell system. Each training sub-sample in the second batch of training sub-samples contains a lifetime reward as well as the operating parameter values, new operating parameter values and control command information of the fuel cell system. The first value network in the reinforcement learning agent is trained based on the first batch of training sub-samples; and the second value network in the reinforcement learning agent is trained based on the second batch of training sub-samples. The policy network is trained based on the batch of training samples, combined with the first value network and the second value network that have completed training.
[0008] Further, training the policy network based on the batch training samples, combined with the first value network and the second value network that have completed training, includes: All operating parameter values of the fuel cell system are extracted from the batch training samples, and then all operating parameter values of the fuel cell system are input into the strategy network for analysis to obtain all new control command information of the fuel cell system and determine the gradient of the action with respect to the parameters of the strategy network. All operating parameter values of the fuel cell system and their corresponding new control command information are input into the first value network after training to determine the gradient of the first value network for the action. All operating parameter values of the fuel cell system and their corresponding new control command information are input into the trained second value network to determine the gradient of the second value network with respect to the action. The gradient of the policy network with respect to the parameters is determined based on the gradient of the first value network with respect to the action, the gradient of the second value network with respect to the action, and the gradient of the action with respect to the parameters of the policy network. The policy network is updated with parameters based on the gradient of the parameters.
[0009] Further, determining the gradient of the policy network with respect to the parameters based on the gradient of the first value network with respect to the action, the gradient of the second value network with respect to the action, and the gradient of the action with respect to the parameters of the policy network includes: The gradient of the action from the first value network and the gradient of the action from the second value network are added together to obtain the total gradient. The gradient of the policy network with respect to the parameters is determined based on the gradient of the action with respect to the parameters of the policy network and the total gradient.
[0010] Furthermore, the method also includes: When the lifespan reward exceeds a preset safety threshold, the stack membrane impedance and stack hydrogen flow rate of the fuel cell system are obtained, and then the control command information of the fuel cell system is adjusted in combination with the stack temperature and stack anode-cathode pressure difference.
[0011] To achieve the above objectives, another aspect of this application proposes a fuel cell system control device, the device comprising: The first module is used to acquire the operating parameter values of the fuel cell system and input them into the policy network in the reinforcement learning agent for analysis, so as to obtain the control command information of the fuel cell system. The second module is used to regulate the fuel cell system according to the control command information of the fuel cell system, and then obtain the new operating parameter values, efficiency impact parameter values and damage impact parameter values of the fuel cell system. The third module is used to determine the performance bonus based on the efficiency impact parameter value of the fuel cell system; and to determine the life bonus based on the damage impact parameter value of the fuel cell system. The fourth module is used to generate a training sample and store it in an experience pool based on the performance reward, the life reward, the operating parameter values of the fuel cell system, the new operating parameter values, and the control command information. The experience pool is used to assist in training the reinforcement learning agent.
[0012] To achieve the above objectives, another aspect of this application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described fuel cell system control method.
[0013] To achieve the above objectives, another aspect of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described fuel cell system control method.
[0014] This application includes at least the following beneficial effects: By introducing a reinforcement learning strategy that considers both performance and lifespan as a dual reward mechanism, the operating parameter values of the fuel cell system are first obtained and input into the policy network of the reinforcement learning agent for analysis. Then, the fuel cell system is regulated according to the control command information of the fuel cell system obtained from the analysis. Next, new operating parameter values, efficiency impact parameter values, and damage impact parameter values of the fuel cell system are obtained. Subsequently, the performance reward is determined based on the efficiency impact parameter value of the fuel cell system, and the lifespan reward is determined based on the damage impact parameter value of the fuel cell system. Finally, the performance reward, lifespan reward, and the operating parameter values, new operating parameter values, and control command information of the fuel cell system are combined to form a training sample and stored in an experience pool used to assist in training the reinforcement learning agent. This can improve the control robustness of the fuel cell system. Attached Figure Description
[0015] Figure 1 This is a schematic flowchart of a fuel cell system control method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the composition of a fuel cell system control device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0017] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0018] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0020] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0021] Reinforcement learning (RL) is a machine learning method. Its fundamental framework is the Markov decision process, which allows an agent to learn optimal policies through trial and error in its interactions with the environment. The agent performs actions in the environment and receives feedback (rewards) based on the results of those actions. These reward signals guide the agent to adjust its policy to maximize long-term cumulative rewards.
[0022] The commercial application of fuel cell systems places stringent demands on their efficiency, power density, and durability. However, these three objectives are conflicting in terms of physical mechanisms, as explained below: (1) Conflict between efficiency and power: Pursuing high efficiency usually means low current density and high air excess coefficient, but this will limit the net power output of the fuel cell system; pursuing high power often requires increasing reactant pressure and flow rate, which will increase the parasitic power consumption of auxiliary equipment such as air compressor, and thus reduce the net efficiency of the fuel cell system. (2) Conflict between power and lifespan: Rapidly increasing power output will cause the internal temperature of the fuel cell stack to rise sharply and the water content of the membrane to fluctuate drastically, which will accelerate the sintering of the catalyst and the degradation of the electrolyte membrane, and significantly shorten the lifespan of the fuel cell stack.
[0023] Existing control strategies for fuel cell systems typically couple performance targets and lifetime constraints into a single objective function using a weighted summation. However, there is often an inherent conflict between performance and lifetime, and fixed weights make it difficult to achieve a dynamic balance between the two across all operating conditions. The ratio of performance weights to lifetime penalty coefficients usually requires extensive manual adjustment and is not universally applicable under different operating conditions. Furthermore, fuel cell systems experience performance degradation during long-term operation, and the initially set weight coefficients gradually become ineffective as the fuel cell system ages, causing the control strategy to deviate from the optimal compromise.
[0024] Furthermore, although some scholars have proposed that fuel cell systems can be controlled within a reinforcement learning framework, performance rewards and lifetime rewards are often weighted and fused to obtain a comprehensive reward to guide the training of the agent's value network. Sparse and negative lifetime constraint signals are easily overwhelmed by dense performance reward signals, causing the agent to sacrifice long-term lifetime for short-term performance. Moreover, lifetime decay is a long-term cumulative process, and existing agents lack the ability to explicitly predict the future lifetime risk of current actions.
[0025] In view of this, embodiments of this application provide a fuel cell system control method, apparatus, electronic device, and storage medium. This scheme introduces a reinforcement learning strategy that considers both performance and lifespan rewards. First, the operating parameter values of the fuel cell system are acquired and input into the policy network of the reinforcement learning agent for analysis. Then, the fuel cell system is regulated according to the control command information of the fuel cell system obtained from the analysis. Next, new operating parameter values, efficiency impact parameter values, and damage impact parameter values of the fuel cell system are acquired. Subsequently, performance rewards are determined based on the efficiency impact parameter values of the fuel cell system, and lifespan rewards are determined based on the damage impact parameter values of the fuel cell system. Finally, the performance rewards, lifespan rewards, and the operating parameter values, new operating parameter values, and control command information of the fuel cell system are combined to form a training sample, which is stored in an experience pool used to assist in training the reinforcement learning agent. This can improve the control robustness of the fuel cell system.
[0026] This application provides a fuel cell system control method, relating to the field of fuel cell technology. It can be applied to a terminal, a server, or software running on either a terminal or a server. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster consisting of multiple physical servers, or a distributed system. It can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the above method, but is not limited to these forms.
[0027] Please see Figure 1 , Figure 1 This is an optional flowchart illustrating a fuel cell system control method provided in an embodiment of this application. The method may include, but is not limited to, the following steps S101 to S104: Step S101: Obtain the operating parameter values of the fuel cell system and input them into the policy network in the reinforcement learning agent for analysis to obtain the control command information of the fuel cell system. Step S102: Adjust the fuel cell system according to the control command information of the fuel cell system, and then obtain the new operating parameter values, efficiency impact parameter values and damage impact parameter values of the fuel cell system. Step S103: Determine the performance bonus based on the efficiency impact parameter value of the fuel cell system; and determine the lifespan bonus based on the damage impact parameter value of the fuel cell system. Step S104: Based on the performance reward, life reward, and the operating parameter values, new operating parameter values, and control command information of the fuel cell system, a training sample is generated and stored in the experience pool. This experience pool is used to assist in training the reinforcement learning agent.
[0028] Steps S101 to S104, as shown in the embodiments of this application, improve the control robustness of the fuel cell system by introducing a reinforcement learning strategy that takes into account both performance and lifespan.
[0029] In some embodiments, a fuel cell system includes at least a fuel cell stack and connected to it a hydrogen supply subsystem, an air supply subsystem, and a thermal management subsystem. The air supply subsystem includes at least an air compressor and a back pressure valve. The air compressor compresses ambient air to a set pressure to provide sufficient high-pressure oxygen to the fuel cell stack cathode. As a core power-consuming component of the air supply subsystem, it directly affects reaction efficiency and system output power. The back pressure valve is typically installed at the fuel cell stack cathode outlet and is used to control the gas back pressure on the fuel cell stack cathode side, thereby regulating the internal pressure level of the fuel cell stack. The thermal management subsystem includes at least a cooling water pump. The cooling water pump drives the coolant (such as deionized water or a dedicated coolant) to circulate between the fuel cell stack and the radiator, carrying away the heat generated during fuel cell stack operation and dissipating it through the radiator, ensuring the fuel cell stack operates within a suitable temperature window. The hydrogen supply subsystem includes at least a hydrogen circulation pump. The hydrogen circulation pump re-extracts and pressurizes unreacted hydrogen from the fuel cell stack anode outlet and returns it to the fuel cell stack anode inlet for reuse. The specific composition of the fuel cell system is prior art and will not be described further here.
[0030] In step S101 of some embodiments, the policy network in the reinforcement learning agent is configured with a state space and an action space. The operating parameters of the fuel cell system input to the state space include stack current density, stack voltage, stack cathode inlet pressure, stack cathode outlet pressure, stack temperature, stack membrane impedance, and stack cumulative damage variables. The control commands output from the action space include air compressor speed commands, back pressure valve opening commands, cooling water pump speed commands, and hydrogen circulation pump speed commands. It is understood that the operating parameter values and new operating parameter values of the fuel cell system refer to the specific values of the operating parameters of the fuel cell system at different times, and the control command information and new control command information of the fuel cell system refer to the specific manifestations of the control commands of the fuel cell system at different times.
[0031] In step S103 of some embodiments, the efficiency impact parameter value of the fuel cell system may include the net output power and hydrogen consumption of the fuel cell system; regarding the determination of performance bonus based on the efficiency impact parameter value of the fuel cell system, the corresponding implementation may include, but is not limited to, the following: First, based on the preset lower calorific value of hydrogen and the net output power and hydrogen consumption of the fuel cell system, the efficiency of the fuel cell system is determined, which can be calculated using the following expression: ; In the formula, For the efficiency of fuel cell systems, The net output power of a fuel cell system can be obtained by subtracting the power of the fuel cell stack from the power consumption of other auxiliary components. For the hydrogen consumption of fuel cell systems, The preset lower calorific value of hydrogen is typically taken as 33.33 kWh / kg. Refers to the current time; The performance bonus is then determined based on the net output power, hydrogen consumption, and efficiency of the fuel cell system, and can be calculated using the following expression: ; In the formula, For performance rewards, This is the preset maximum net output power limit for the fuel cell system. These are the preset maximum hydrogen consumption limits for the fuel cell system, and these two maximum limits can be determined by technicians based on the current operating conditions of the fuel cell system. , and All are normalized weighting coefficients.
[0032] By incorporating maximizing performance rewards as one of the objectives during the training of reinforcement learning agents, the agents are guided towards optimization towards high net power, low hydrogen consumption, and high efficiency.
[0033] In step S103 of some embodiments, the damage impact parameter values of the fuel cell system may include the stack temperature, stack membrane water content, and stack anode-cathode voltage difference of the fuel cell system. The stack anode-cathode voltage difference can be obtained by subtracting the stack cathode voltage from the stack anode voltage. Regarding the determination of lifespan bonus based on the damage impact parameter values of the fuel cell system, corresponding implementation methods may include, but are not limited to, the following: First, based on the preset frequency factor, preset activation energy, preset gas constant, preset control step size, and stack temperature of the fuel cell system, the thermal damage increment is determined, which can be calculated using the following expression: ; In the formula, The thermal damage increment can be determined based on the existing Arrhenius equation, which states that the stack decay rate doubles for every 10°C increase in stack temperature. The preset frequency factor is related to the stack aging mechanism (such as catalyst roughening, membrane degradation, etc.), and its preferred value range is
[10] . 5 s -1 10 7 s -1 The specific possible value is 10. 6 s -1 , The activation energy is preset, and its preferred range is [50 kJ / mol, 100 kJ / mol], specifically 70 kJ / mol. The gas constant is a preset value, typically taken as 8314 J / (mol·K). For the stack temperature of the fuel cell system, The preset control step size can be understood as the basic control interval of the fuel cell system through a reinforcement learning agent; Based on the preset humidity damage coefficient, preset control step size, and water content of the fuel cell stack membrane, the humidity damage increment is determined and can be calculated using the following expression: , ; In the formula, The humidity-induced damage increment refers to mechanical stress damage caused by excessively low (flooded) or excessively high (dry) impedance of the fuel cell film. The preset humidity damage coefficient is preferably set within the range of
[10] . -6 s -1 10 -4 s -1 The specific possible value is 10. -5 s -1 , The water content of the fuel cell stack membrane in the fuel cell system. This refers to the instantaneous damage caused by the moisture content of the fuel cell stack membrane deviating from the optimal humidity range. The preset membrane dry damage coefficient is preferably set within the range of [0.5, 2], and can specifically be set to 1. To preset the flood damage coefficient, its preferred value range is [0.5, 1.5], and a specific value of 0.8 is acceptable. This represents the optimal lower limit for the water content of the fuel cell stack membrane, preferably within the range of [12, 15], and specifically, a value of 14. The optimal upper limit for the water content of the fuel cell stack membrane is defined as (15, 18), with a preferred value of 17. The saturation value of water content in the stack membrane of the fuel cell system is preferably [20, 24], and can be specifically taken as 22; Based on the preset differential pressure damage coefficient, preset control step size, and the differential pressure between the anode and cathode of the fuel cell system, the differential pressure damage increment is determined and can be calculated using the following expression: ; In the formula, This refers to the incremental pressure difference damage, which is caused by membrane mechanical fatigue due to excessive pressure difference between the cathode and anode of the fuel cell stack. The preset differential pressure damage coefficient is preferably taken in the range of
[10] . -6 s -1 10 -4 s -1 The specific possible value is 5 × 10. -6 s -1 , For the anode-cathode voltage difference of the fuel cell stack, To preset the differential pressure safety threshold, its preferred value range is [30 kPa, 80 kPa], and a specific value of 50 kPa is acceptable. For reference pressure difference, its preferred range is [50 kPa, 100 kPa], and a specific value of 80 kPa is acceptable. The damage index is preferably in the range of [1, 2.5], and can specifically be 1.5. This indicates the positive part operation, i.e. This is used to guide damage when the pressure difference exceeds a safety threshold; The life bonus is then determined based on the increments of thermal damage, humidity damage, and differential pressure damage, and can be calculated using the following expression: ; In the formula, As a lifespan reward, , and All are normalized weighting coefficients.
[0034] By considering maximizing lifetime reward as another objective during the training of the reinforcement learning agent, the reinforcement learning agent is guided to optimize in the direction of slowing down the rate of fuel cell lifetime decay.
[0035] In some embodiments, the reinforcement learning agent can be constructed based on the existing DDPG (Deep Deterministic Policy Gradient) algorithm, the existing TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm, or other existing algorithms. Based on this, the aforementioned fuel cell system control method may further include a training process for the reinforcement learning agent, and the corresponding implementation may include, but is not limited to, the following: First, a batch of training samples is extracted from the experience pool and the data is split to form a first batch of training sub-samples and a second batch of training sub-samples. Each training sub-sample in the first batch of training sub-samples contains a performance reward as well as the operating parameter values, new operating parameter values, and control command information of the fuel cell system. Each training sub-sample in the second batch of training sub-samples contains a lifetime reward as well as the operating parameter values, new operating parameter values, and control command information of the fuel cell system. It can be understood that the first batch of training sub-samples and the second batch of training sub-samples contain the same number of samples, and there are every two training sub-samples in the first batch of training sub-samples and the second batch of training sub-samples that differ only in the type of reward. Secondly, the first value network in the reinforcement learning agent is trained based on the first batch of training sub-samples; and the second value network in the reinforcement learning agent is trained based on the second batch of training sub-samples. The training process of these two value networks is existing technology, and the existing TD (Temporal Difference) algorithm can be introduced to assist in the implementation, which will not be elaborated here. Finally, based on the batch training samples, the policy network is trained by combining the first value network and the second value network that have been trained.
[0036] By constructing a dual-objective decoupled deep reinforcement learning control architecture, a first value network is set up to learn and evaluate based on the performance objective, and a second value network is set up to learn and evaluate based on the lifetime risk avoidance objective. This can solve the problem of mutual interference between performance objectives and lifetime objectives in traditional methods. Subsequently, co-optimization is achieved through decoupled gradient guidance, thereby improving the control robustness of the fuel cell system.
[0037] Furthermore, regarding the training of the policy network based on a batch of training samples, combining the first value network and the second value network that have completed training, the corresponding implementation methods may include, but are not limited to, the following: First, all operating parameter values of the fuel cell system are extracted from the batch training samples. Then, the extracted operating parameter values of the fuel cell system are input into the policy network for analysis to obtain all new control command information of the corresponding fuel cell system. The gradient of the action with respect to the parameters of the policy network is determined, which can be understood as the derivative of the action output by the policy network with respect to its own parameters (such as weights). Specifically, it is obtained by backpropagation in the policy network based on all new control command information of the fuel cell system determined by the forward propagation method. Secondly, all extracted operating parameter values of the fuel cell system and their corresponding new control command information are input into the first value network after training to determine the gradient of the first value network with respect to the action. This can be understood as the derivative of the Q value output by the first value network with respect to the action. Specifically, this is obtained by first performing forward propagation based on all extracted operating parameter values of the fuel cell system and their corresponding new control command information in the first value network, which is in a parameter (such as weights) frozen state, to obtain all Q values, and then performing back propagation based on all Q values. Similarly, all extracted operating parameter values of the fuel cell system and their corresponding new control command information are input into the second value network after training to determine the gradient of the second value network with respect to the action. This can be understood as the derivative of the Q value output by the second value network with respect to the action. Specifically, this is obtained by first performing forward propagation based on all extracted operating parameter values of the fuel cell system and their corresponding new control command information in the second value network, which is in a parameter (such as weights) frozen state, to obtain all Q values, and then performing back propagation based on all Q values. Next, based on the gradients of the first value network with respect to the action, the second value network with respect to the action, and the gradients of the action with respect to the parameters of the policy network, the gradient of the policy network with respect to the parameters is determined. Specifically, the gradients of the first value network with respect to the action and the second value network with respect to the action are first added together to obtain the total gradient. Then, based on this total gradient and the gradient of the action with respect to the parameters of the policy network, the gradient of the policy network with respect to the parameters is determined, which can be calculated using the following expression: ; In the formula, The gradient of the policy network with respect to the parameters can be understood as being calculated based on the chain rule. The gradient of the first value network with respect to the action. The gradient of the second value network with respect to the action. The gradient of the action with respect to the parameters of the policy network. It can represent the operation of calculating the expected value; Finally, the policy network is updated based on the gradient of the policy network with respect to the parameters. In general, the gradient of the policy network with respect to the parameters is multiplied by the preset learning rate, and then the result of the multiplication is added to the current parameters of the policy network to obtain the updated result of the current parameters of the policy network.
[0038] When updating the parameters of the policy network, the gradient function of the policy network can be used to indicate in which direction the control parameters should be adjusted to maximize the objective function. In this application, in order to enable the policy network to learn a strategy that can both improve the performance of the fuel cell system and avoid excessive lifetime damage, the performance gradient and lifetime gradient are introduced to jointly adjust the policy gradient, thereby driving the parameters of the policy network to be updated once in the direction that simultaneously improves performance and reduces lifetime damage. If the performance gradient and lifetime gradient are in the same direction, the parameter update direction is strengthened; if the performance gradient and lifetime gradient are in opposite directions, a trade-off is made based on their magnitudes. Ultimately, the policy network learns to make dynamic decisions between the dual objectives of pursuing high performance and avoiding lifetime risks without the need for manual weight setting. This simplifies the algorithm debugging process and improves the ability of reinforcement learning strategies to be applied across different types of fuel cell stacks.
[0039] In some embodiments, a related dynamic model of the fuel cell system can be constructed, which includes at least an electrochemical model of the fuel cell stack, a fluid dynamics model, a thermodynamic model, and a semi-empirical lifetime decay model. By realizing the interaction between the dynamic model of the fuel cell system and the reinforcement learning agent in the simulation environment, more training samples can be obtained and stored in the experience pool, thereby expanding the number of training samples in the experience pool, which is beneficial to improving the training effect of the reinforcement learning agent.
[0040] In some embodiments, the above-described fuel cell system control method may further include: when the lifespan bonus exceeds a preset safety threshold, acquiring the stack membrane impedance and stack hydrogen flow rate of the fuel cell system, and then adjusting the control command information of the fuel cell system in conjunction with the stack temperature and stack anode-cathode voltage difference; wherein, the preset safety threshold is preferably set to -0.8, which can be obtained by calibration when controlling the fuel cell system to operate normally under rated conditions. By triggering this emergency parameter adjustment mechanism, the fuel cell system can be pulled back to a safe operating range.
[0041] Specifically, if the stack temperature of the fuel cell system exceeds a preset upper temperature limit (preferably set to 80°C), the speed of the cooling water pump should be increased. If the stack membrane impedance exceeds a preset upper impedance limit (which can be obtained by calibrating the stack after a membrane dry test), the speed of the air compressor should be decreased, and the speed of the hydrogen circulation pump should be increased. If the stack membrane impedance is lower than a preset lower impedance limit (which can be obtained by calibrating the stack after a water flooding test), the speed of the air compressor should be increased to strengthen purging. If the anode-cathode pressure difference of the fuel cell system exceeds a preset pressure limit... The differential pressure safety threshold is preferably set to 50 kPa. If the back pressure valve opening is increased, the pressure is balanced. If the hydrogen flow rate of the fuel cell system stack is lower than the preset hydrogen flow rate threshold, it indicates that the hydrogen circulation volume of the stack is insufficient. In this case, the speed of the hydrogen circulation pump is increased. The preset hydrogen flow rate threshold can be determined by: synchronously acquiring the stack current of the fuel cell system, then determining the corresponding rated hydrogen flow rate of the stack under the stack current by looking up a table, and then multiplying the rated hydrogen flow rate of the stack by the preset hydrogen excess coefficient to obtain the preset hydrogen flow rate threshold. The preset hydrogen excess coefficient is preferably set to 1.3.
[0042] In some embodiments, the above-described fuel cell system control method may further include: inputting the operating parameter values of the fuel cell system and their corresponding control command information into a second value network in a reinforcement learning agent; determining the target Q value through forward propagation; and then obtaining the target gradient of the action by the second value network after backpropagation based on the target Q value; acquiring application scenario information of the fuel cell system, such as the fuel cell system being used in buses pursuing long lifespan or in passenger vehicles pursuing high performance; determining the corresponding adaptive adjustment coefficient through a lookup table; and then correcting the control command information of the fuel cell system based on the adaptive adjustment coefficient and the target gradient of the action by the second value network, which can be calculated using the following expression: ; In the formula, The results of the correction regarding the control command information of the fuel cell system. This refers to the control command information for the fuel cell system. The target gradient of the second value network for the action. Used to determine the gradient sign, when This indicates that increasing the action value in the current state will increase the expected cumulative lifespan reward, meaning less long-term damage. When this occurs, it indicates that the action value needs to be reduced to decrease damage in the current state. When this occurs, it indicates that no action value needs to be modified. For adaptive adjustment coefficients, when When this occurs, it indicates that the instruction correction direction is more biased towards lifespan protection, and the action will be adjusted in the direction of reducing long-term damage. This indicates that the instruction correction direction is more performance-oriented, and the action will be adjusted in the opposite direction of the gradient sign, sacrificing some lifetime for higher performance. The preferred value range is [0, 0.2], and when The larger the value, the stronger the correction to the action value.
[0043] By appropriately modifying the control commands according to the application scenario and lifetime gradient of the fuel cell system without compromising the original performance strategy, decoupled guidance between performance and lifetime can be achieved, enabling the fuel cell system to operate in the desired direction.
[0044] Please see Figure 2 , Figure 2 This is a schematic diagram of an optional structural composition of a fuel cell system control device provided in an embodiment of this application, which can implement the above-described fuel cell system control method. The device may include, but is not limited to, the following: The first module 201 is used to acquire the operating parameter values of the fuel cell system and input them into the policy network in the reinforcement learning agent for analysis, so as to obtain the control command information of the fuel cell system. The second module 202 is used to regulate the fuel cell system according to the control command information of the fuel cell system, and then obtain the new operating parameter values, efficiency impact parameter values and damage impact parameter values of the fuel cell system. The third module 203 is used to determine the performance bonus based on the efficiency impact parameter value of the fuel cell system; and to determine the life bonus based on the damage impact parameter value of the fuel cell system. The fourth module 204 is used to generate a training sample and store it in the experience pool based on performance rewards, life rewards, and the operating parameter values, new operating parameter values and control command information of the fuel cell system. This experience pool is used to assist in training the reinforcement learning agent.
[0045] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those specifically implemented by the above method embodiments, and the beneficial effects achieved by the present device embodiments are also the same as those achieved by the above method embodiments.
[0046] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned fuel cell system control method. This electronic device can include any smart terminal such as a tablet computer or an in-vehicle computer.
[0047] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those implemented by the above method embodiments, and the beneficial effects achieved by the present device embodiments are also the same as those achieved by the above method embodiments.
[0048] Please see Figure 3 , Figure 3 This is a schematic diagram illustrating the hardware structure of an electronic device according to another embodiment. The electronic device includes: The processor 301 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 302 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 302 can store the operating system and other applications. When the technical solutions provided in the embodiments of this application are implemented through software or firmware, the relevant program code is stored in the memory 302 and is called and executed by the processor 301. Input / output interface 303 is used to implement information input and output; The communication interface 304 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 305 transmits information between various components of the device (e.g., processor 301, memory 302, input / output interface 303, and communication interface 304); The processor 301, memory 302, input / output interface 303 and communication interface 304 are connected to each other within the device via bus 305.
[0049] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described fuel cell system control method.
[0050] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented by this storage medium embodiment are the same as those implemented by the above method embodiments, and the beneficial effects achieved by this storage medium embodiment are also the same as those achieved by the above method embodiments.
[0051] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0052] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0053] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0054] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0055] Those skilled in the art will understand that all or some of the steps, apparatuses, or functional modules / units in the methods disclosed above can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0056] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0057] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0058] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed between the devices or units may be through some interfaces, and the indirect coupling or communication connection may be electrical, mechanical, or other forms.
[0059] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0060] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0061] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0062] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A control method for a fuel cell system, characterized in that, The method includes: The operating parameter values of the fuel cell system are obtained and input into the policy network in the reinforcement learning agent for analysis to obtain the control command information of the fuel cell system. Based on the control command information of the fuel cell system, the fuel cell system is regulated, and then new operating parameter values, efficiency impact parameter values, and damage impact parameter values of the fuel cell system are obtained. A performance bonus is determined based on the efficiency impact parameter value of the fuel cell system; and a lifetime bonus is determined based on the damage impact parameter value of the fuel cell system. Based on the performance reward, the lifetime reward, and the operating parameter values, new operating parameter values, and control command information of the fuel cell system, a training sample is generated and stored in the experience pool, which is used to assist in training the reinforcement learning agent.
2. The fuel cell system control method according to claim 1, characterized in that, The efficiency impact parameters of the fuel cell system include the net output power and hydrogen consumption of the fuel cell system; determining the performance bonus based on the efficiency impact parameters of the fuel cell system includes: The efficiency of the fuel cell system is determined based on the preset low calorific value of hydrogen, the net output power of the fuel cell system, and the hydrogen consumption. The performance bonus is determined based on the net output power, hydrogen consumption, and efficiency of the fuel cell system.
3. The fuel cell system control method according to claim 1, characterized in that, The damage impact parameters of the fuel cell system include the stack temperature, stack membrane water content, and stack anode-cathode voltage difference; determining the lifespan bonus based on the damage impact parameters of the fuel cell system includes: The thermal damage increment is determined based on the preset frequency factor, preset activation energy, preset gas constant, preset control step size, and stack temperature of the fuel cell system. The humidity damage increment is determined based on the preset humidity damage coefficient, the preset control step size, and the water content of the fuel cell stack membrane. The differential pressure damage increment is determined based on the preset differential pressure damage coefficient, the preset control step size, and the differential pressure between the anode and cathode of the fuel cell system stack. The lifetime bonus is determined based on the thermal damage increment, the humidity damage increment, and the differential pressure damage increment.
4. The fuel cell system control method according to claim 1, characterized in that, The method further includes: A batch of training samples is extracted from the experience pool and the data is split to form a first batch of training sub-samples and a second batch of training sub-samples. Each training sub-sample in the first batch of training sub-samples contains a performance reward as well as the operating parameter values, new operating parameter values and control command information of the fuel cell system. Each training sub-sample in the second batch of training sub-samples contains a lifetime reward as well as the operating parameter values, new operating parameter values and control command information of the fuel cell system. The first value network in the reinforcement learning agent is trained based on the first batch of training sub-samples; and the second value network in the reinforcement learning agent is trained based on the second batch of training sub-samples. The policy network is trained based on the batch of training samples, combined with the first value network and the second value network that have completed training.
5. The fuel cell system control method according to claim 4, characterized in that, The step of training the policy network based on the batch training samples, combined with the first value network and the second value network that have completed training, includes: All operating parameter values of the fuel cell system are extracted from the batch training samples, and then all operating parameter values of the fuel cell system are input into the strategy network for analysis to obtain all new control command information of the fuel cell system and determine the gradient of the action with respect to the parameters of the strategy network. All operating parameter values of the fuel cell system and their corresponding new control command information are input into the first value network after training to determine the gradient of the first value network for the action. All operating parameter values of the fuel cell system and their corresponding new control command information are input into the trained second value network to determine the gradient of the second value network for the action. The gradient of the policy network with respect to the parameters is determined based on the gradient of the first value network with respect to the action, the gradient of the second value network with respect to the action, and the gradient of the action with respect to the parameters of the policy network. The policy network is updated with parameters based on the gradient of the parameters.
6. The fuel cell system control method according to claim 5, characterized in that, The step of determining the gradient of the policy network with respect to the parameters based on the gradient of the first value network with respect to the action, the gradient of the second value network with respect to the action, and the gradient of the action with respect to the parameters of the policy network includes: The gradient of the action from the first value network and the gradient of the action from the second value network are added together to obtain the total gradient. The gradient of the policy network with respect to the parameters is determined based on the gradient of the action with respect to the parameters of the policy network and the total gradient.
7. The fuel cell system control method according to claim 3, characterized in that, The method further includes: When the lifespan reward exceeds a preset safety threshold, the stack membrane impedance and stack hydrogen flow rate of the fuel cell system are obtained, and then the control command information of the fuel cell system is adjusted in combination with the stack temperature and stack anode-cathode pressure difference.
8. A control device for a fuel cell system, characterized in that, The device includes: The first module is used to acquire the operating parameter values of the fuel cell system and input them into the policy network in the reinforcement learning agent for analysis, so as to obtain the control command information of the fuel cell system. The second module is used to regulate the fuel cell system according to the control command information of the fuel cell system, and then obtain the new operating parameter values, efficiency impact parameter values and damage impact parameter values of the fuel cell system. The third module is used to determine the performance bonus based on the efficiency impact parameter value of the fuel cell system; and to determine the lifespan bonus based on the damage impact parameter value of the fuel cell system. The fourth module is used to generate a training sample and store it in an experience pool based on the performance reward, the life reward, the operating parameter values of the fuel cell system, the new operating parameter values, and the control command information. The experience pool is used to assist in training the reinforcement learning agent.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the fuel cell system control method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the fuel cell system control method according to any one of claims 1 to 7.