A method for predicting voltage decay in a variable load fuel cell
By employing a distributed independent Q-learning method, combined with reinforcement learning environment and information interaction between agents, the accuracy problem of voltage decay prediction for fuel cells under varying load conditions is solved, achieving more efficient fuel cell voltage decay prediction, adapting to the complex changes of fuel cells under different operating conditions, and improving the accuracy and efficiency of fuel cell lifespan prediction.
Patent Information
- Application Number
- CN202511358202.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-23
AI Technical Summary
Existing fuel cell voltage decay prediction models cannot effectively address the issue of fuel cell durability and commercialization under varying load conditions.
This paper employs a distributed independent Q-learning method to extract feature values of operating parameters, establish a voltage model, use a greedy algorithm to select actions, and use a double Q-learning algorithm to update network parameters. It combines this with reinforcement learning to construct a reinforcement learning environment, setting reward formulas, and facilitating information interaction between agents. This method indirectly achieves information interaction through the environment, avoiding computational overhead and time delays between agents, thus significantly improving the model's computational efficiency.
The computational efficiency of the fuel cell voltage decay prediction model has been improved, greatly enhancing the model's computational efficiency and enabling it to better cope with various complex situations during fuel cell operation. This results in greater flexibility and adaptability.
Smart Images

Figure CN120850822B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fuel cell vehicle technology, and specifically to a method for predicting voltage decay in variable load fuel cells. Background Technology
[0002] A proton exchange membrane fuel cell (PEMFC) is an electrochemical energy conversion device that efficiently converts chemical energy into electrical energy. With its significant advantages such as high efficiency, high energy density, and low pollution, it has shown broad application prospects in the power supply field of medium and large vehicle power systems and is expected to become a key energy solution in the future transportation sector, driving the automotive industry towards a cleaner and more sustainable direction.
[0003] However, the commercialization of PEMFCs currently faces severe challenges, with limited durability due to degradation being one of the main limiting factors. In the actual operation of fuel cell power systems for vehicles, fuel cell stacks experience degradation due to various factors. Specifically, high potential accelerates the oxidation and corrosion of electrode materials, catalyst deactivation reduces the activity of electrochemical reactions, and membrane electrode assembly (MEA) degradation affects proton conduction and gas diffusion, leading to a gradual decline in the performance of the fuel cell stack. Macroscopically, this performance degradation manifests as a gradual increase in the difference between the measured voltage or power of the fuel cell stack and its initial value under constant current conditions, eventually deviating from design requirements and reaching the end of its lifespan, rendering it unable to continue operating normally.
[0004] To accurately predict the degradation of PEMFCs and assess their lifespan in advance, thereby providing a scientific basis for the design, optimization, and maintenance of fuel cells, researchers have conducted extensive research on degradation prediction methods. Currently, hybrid-driven degradation prediction methods have become one of the research focuses. This method combines the advantages of model-driven and data-driven approaches. It possesses the in-depth understanding and mechanistic explanation of model-driven methods, which can reveal the intrinsic causes of fuel cell degradation; and it also has the flexibility of data-driven methods, which can learn and adjust based on actual operating data to adapt to different operating conditions and environmental conditions.
[0005] The general process of hybrid-driven approaches involves extracting aging parameters from physical, empirical, or semi-empirical models. These parameters reflect the degree of degradation of the fuel cell in different aspects. Based on their sensitivity to current density, aging parameters can be categorized into current-density-sensitive and current-density-insensitive aging parameters. Current-density-sensitive aging parameters change significantly with variations in current density; for example, the loss of catalyst activity may be closely related to the intensity of electrochemical reactions at different current densities. In contrast, current-density-insensitive aging parameters are relatively stable and less affected by current density, such as the physical aging process of certain materials in membrane electrode assemblies.
[0006] However, traditional degradation prediction models have revealed significant shortcomings when dealing with variable load conditions. In actual operation of automotive fuel cell power systems, the load current changes frequently to adapt to different driving conditions, such as acceleration, deceleration, and hill climbing. Traditional degradation prediction models generally cannot simultaneously account for the variation patterns of both current density-sensitive and current density-insensitive aging parameters under variable load conditions. Under variable load conditions, the dynamic characteristics of these two aging parameters are more complex, and their interaction significantly impacts the fuel cell degradation process. Because traditional models lack the ability to effectively handle this complex relationship, their long-term lifespan prediction accuracy is significantly limited, failing to accurately predict the degradation and lifespan of fuel cells during actual use, thus restricting the reliability and commercialization of automotive fuel cell power systems. Summary of the Invention
[0007] The purpose of this invention is to propose a method for predicting voltage decay in variable load fuel cells, which can improve the accuracy and efficiency of voltage decay prediction in variable load fuel cells.
[0008] To achieve the above objectives, the present invention provides a method for predicting voltage decay in a variable load fuel cell, comprising,
[0009] Extract the operating parameters and calculate their characteristic values;
[0010] A voltage model is established, with the characteristic values of operating parameters and aging parameters as inputs and the fuel cell stack voltage as the output.
[0011] The aging parameters are optimized for the first data sample under each load current to obtain the initial values of each aging parameter.
[0012] Agents are assigned based on aging parameter characteristics, where:
[0013] Agents are assigned based on current density-insensitive aging parameters and current density-sensitive aging parameters;
[0014] Construct networks for each agent, including a DQN evaluation network and a target network, use a greedy algorithm to select actions, and employ a double-Q learning algorithm to update network parameters;
[0015] A reinforcement learning environment is constructed, the voltage model is used as the voltage calculation formula and a reward formula is set, and the discrete actions selected by the agent are output as operations on various aging parameters of the fuel cell through base conversion;
[0016] Each agent is assigned a dedicated experience pool to perform model training and voltage prediction.
[0017] The fundamental advantage of this approach lies in its ability to address the shortcomings of traditional methods, which typically train and predict voltage decay separately for each operating condition without considering the differences and temporal variations in voltage decay across operating conditions. This new method, however, considers both the differences in voltage decay across operating conditions and performs temporal decay prediction. Instead of processing data from each operating condition in isolation, it comprehensively considers data from different operating conditions, achieving cross-condition temporal prediction. This allows the model to better capture the voltage decay patterns of fuel cells during transitions between different operating conditions, improving the model's predictive ability under complex operating conditions and better reflecting real-world operating realities.
[0018] Multi-current operating condition data exhibits asynchronous characteristics, with aging parameters varying under different current densities, posing a challenge to differentiated aging parameter modeling. To address this issue, this paper proposes a voltage decay prediction method based on distributed independent Q-learning. An agent is assigned based on whether the aging parameters are insensitive or sensitive to current density, establishing specialized models for different types of aging parameters. This solves the challenge of differentiated aging parameter modeling under multi-current operating conditions, thereby improving decay prediction accuracy.
[0019] This proposed distributed independent Q-learning voltage prediction method employs discretization and number system transformation in its action space. The discrete actions selected by the agents are converted into operations on various aging parameters of the fuel cell through number system transformation, resulting in a more concise and efficient representation of the action space. Furthermore, the agents operate independently without communication, avoiding the computational overhead and time delays associated with inter-agent communication. This significantly improves the model's computational efficiency, enabling faster training and prediction, and making it suitable for scenarios with high real-time requirements.
[0020] By constructing a reinforcement learning environment, using the voltage model as the voltage calculation formula and setting a reward formula, the agent can continuously adjust its behavior strategy based on environmental feedback. This enables the model to adapt to different operating conditions and changes, making it more flexible and adaptable, and better able to cope with various complex situations that may occur during fuel cell operation.
[0021] As a feasible and preferred solution, data preprocessing is performed before extracting operating parameters, including removing non-steady-state data that deviates from the preset load current ±1A range, verifying data integrity, identifying and deleting outliers using the box plot method, and downsampling to reduce data frequency;
[0022] The extracted operating parameters include stack voltage, anode inlet and outlet gas pressure, cathode inlet and outlet gas pressure, air flow rate, and cooling water inlet and outlet temperatures.
[0023] The moving average algorithm is used to smooth and correct the time series data.
[0024] As a feasible and preferred approach, the characteristic values of the operating parameters are calculated, including the following:
[0025] The average voltage across all sections is calculated as the voltage characteristic value, using the following formula:
[0026]
[0027] Where n is the number of fuel cell cells in the stack. Let be the voltage value of the fuel cell in section i;
[0028] The average gas pressure at the anode inlet and outlet is used as the characteristic value of hydrogen pressure. The formula is as follows:
[0029]
[0030] in, The anode inlet gas pressure, This refers to the anode outlet gas pressure.
[0031] The average gas pressure at the cathode inlet and outlet is used as a characteristic value of air pressure, calculated using the following formula:
[0032]
[0033] in, The anode inlet gas pressure, This represents the pressure of the gas at the anode outlet.
[0034] As a feasible and preferred option, the voltage model is shown in equation (1):
[0035]
[0036] In the formula, T is the fuel cell temperature; R is the ideal gas constant; and These are hydrogen pressure and air pressure, respectively. The transport coefficient is used to indicate how changes in potential at the reaction interface alter the magnitude of the forward and reverse activation barriers; n is the number of electrons transferred in the electrochemical reaction; F is the Faraday constant; j is the current density; jleak is the leakage current density; j0 is the exchange current density; ASR is the area-normalized fuel cell resistance. j is an empirical constant; jL is the limiting current density;
[0037] Current density-insensitive aging parameters include exchange current density j0, effective reaction area A, and limiting current density jL. Current density-sensitive aging parameters include area-normalized fuel cell resistance ASRi under various load currents.
[0038] As a feasible and preferred approach, a network of agents is constructed, a greedy algorithm is used to select actions, and a double-Q learning algorithm is employed to update network parameters, including the following:
[0039] The network adopts a two-layer fully connected structure, which can be mathematically expressed as:
[0040]
[0041] In the formula, and These are the trainable parameter matrices for the hidden layer and the output layer, respectively. To observe the spatial dimension, The number of hidden layer neurons. For the state dimension; Represents the ReLU activation function; the input layer receives data in dimension 1. Observation state vector After regularization Linear transformation projection to The final output after nonlinear activation processing of the 3D hidden layer space is... State value estimates under different actions;
[0042] Using a greedy algorithm to select the action, the expression is (3):
[0043]
[0044] In the formula, It is a greedy probability, ranging from (0, 1);
[0045] The maximum value is estimated by forward propagation of the agent DQN. And choose the action that maximizes value. Use actions Target network in intelligent agent Calculate the value to obtain the estimated maximum value .
[0046] As a feasible and preferred solution, the network update process is as follows:
[0047] The formula for calculating the TD target is:
[0048]
[0049] In the formula, For TD objectives, As a reward, To balance the discount factor between current rewards and future rewards;
[0050] The formula for calculating TD error is:
[0051]
[0052] In the formula, For TD error;
[0053] The DQN network parameters are updated using gradient descent, and the formula is:
[0054]
[0055] In the formula, For the updated network parameters, For the network parameters before the update, The learning rate;
[0056] The target network parameters are updated using a weighted average soft update, as shown in the formula:
[0057]
[0058] In the formula, This is the soft update coefficient.
[0059] As a feasible and preferred approach, a reinforcement learning environment is constructed, wherein:
[0060] The voltage model described above is used as the voltage calculation method;
[0061] A reward function is set, which is based on the error between the predicted voltage value and the measured voltage value, and the formula is:
[0062]
[0063] In the formula, R is the reward of the reinforcement learning algorithm; This is the predicted voltage value; This is a voltage measurement value; It is a positive constant; all agents share the reward R.
[0064] As a feasible and preferred approach, a dedicated experience pool is set up for each agent, including the following:
[0065] Sample When storing data into the experience pool, set the sampling probability using the following formula:
[0066]
[0067] In the formula, For the j-th experience of agent i; Let be the gradient of the DQN network for agent i; It is a constant used to ensure that all samples are drawn with a non-zero probability;
[0068] The learning rate for the samples is set using the following formula:
[0069]
[0070] In the formula, Let be the learning rate for sample j; b is the general learning rate set; b is the total number of samples in the priority experience replay array. It's a hyperparameter.
[0071] As a feasible and preferred approach, model training includes:
[0072] Use the initial values of the aging parameters as the initial state of the reinforcement learning model;
[0073] The corresponding agent is activated based on the load current, and each agent selects an action based on the local observation state using a greedy strategy.
[0074] The aging parameters are updated based on the actions, the predicted voltage and reward are calculated, and the experience is stored in the corresponding experience pool.
[0075] The network parameters of each agent are sampled from the experience pool and updated using the double-Q learning algorithm.
[0076] As a feasible and preferred approach, voltage prediction includes:
[0077] The sliding window mechanism is used to select sample data for the time to be predicted;
[0078] The corresponding intelligent agent is activated to update the aging parameters, and then the predicted voltage is calculated through the voltage model;
[0079] At the end of the sliding window, the predicted results are compared with the actual values to update the experience pool, and the next round of prediction continues. Attached Figure Description
[0080] Figure 1This is a logic block diagram of a voltage attenuation prediction method based on distributed independent Q-learning.
[0081] Figure 2 This is a schematic diagram of the distributed independent Q-learning algorithm.
[0082] Figure 3 This is a schematic diagram of the network update method based on the double-Q learning algorithm.
[0083] Figure 4 The flowchart shows the training process for a voltage decay prediction model based on distributed independent Q-learning.
[0084] Figure 5 This is a flowchart of the voltage attenuation prediction model based on distributed independent Q-learning.
[0085] Figure 6 This is the voltage prediction result.
[0086] Figure 7 This is a schematic diagram of the electronic device structure according to an embodiment of the present invention.
[0087] The attached figures indicate the following: electronic device 500, processor 501, communication interface 502, memory 503, and bus 504. Detailed Implementation
[0088] To make the technical solution and advantages of this application clearer, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only some embodiments of the present invention, and are only used to explain this application, not to limit it. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered isolated; they can be combined with each other to achieve better technical effects. The same reference numerals appearing in the accompanying drawings of the following embodiments represent the same features or components, and can be applied to different embodiments.
[0089] Furthermore, unless otherwise defined, the technical or scientific terms used in this invention description shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains.
[0090] The present invention will now be described in further detail with reference to the accompanying drawings.
[0091] Reference Figure 1 This invention provides a method for predicting voltage decay in a variable load fuel cell, taking into account both current density-sensitive and non-sensitive aging parameters, and includes the following steps.
[0092] Step S1, data preprocessing, specifically includes the following steps:
[0093] Step S1.1: In the system's historical operating data, filter the data sample set that is in normal operating condition. The specific steps are as follows:
[0094] Unsteady-state data that deviates from the preset load current ±1 A range are removed.
[0095] Verify data integrity by checking for missing values. For data records with missing values, if the number of missing values is small and does not affect the overall data characteristics, interpolation can be used to fill them; if the number of missing values is large, the data record should be deleted directly.
[0096] Outliers are identified and deleted using box plots. Box plots use quartiles to determine the distribution range of data. Data exceeding a certain range above or below the upper or lower quartile (usually 1.5 times the interquartile range) are identified as outliers and deleted.
[0097] Downsampling reduces the data frequency by filtering data that characterizes decay at 1-hour intervals. For example, if the original data collection frequency is once per minute, after downsampling, only one data point is retained per hour to reduce the amount of data and improve the efficiency of subsequent processing.
[0098] Step S1.2: Extract operating parameters from the preprocessed data, including stack voltage, anode inlet and outlet gas pressure, cathode inlet and outlet gas pressure, air flow rate, and cooling water inlet and outlet temperatures.
[0099] Step S1.3: The time-series data is smoothed and corrected using a moving average algorithm to suppress interference from high-frequency noise and measurement errors. The formula for the moving average algorithm is as follows:
[0100]
[0101] in, These are the data values after smoothing correction. The original data values, The window size for the moving average. For example, setting the window size. =5, then for the current time The data value is then adjusted by taking the average of the five data values from the previous four times and the current time.
[0102] Step S2, eigenvalue calculation, specifically includes the following steps:
[0103] Calculate the average voltage of all sections as the voltage characteristic value.
[0104]
[0105] Where n is the number of fuel cell cells in the stack. Let be the voltage value of the fuel cell in section i.
[0106] The average gas pressure at the anode inlet and outlet is calculated as the characteristic value of hydrogen pressure.
[0107]
[0108] in, The anode inlet gas pressure, This represents the pressure of the gas at the anode outlet.
[0109] The average gas pressure at the cathode inlet and outlet is calculated as a characteristic value of air pressure.
[0110]
[0111] in, The anode inlet gas pressure, This represents the pressure of the gas at the anode outlet.
[0112] The average temperature of cooling water entering and leaving the stack is calculated as the characteristic value of the stack temperature.
[0113] Step S3: Establish a semi-empirical mechanism voltage model for the fuel cell. The voltage model is shown in equation (1):
[0114]
[0115] In the formula, T is the fuel cell temperature in K; R is the ideal gas constant. and These are hydrogen pressure and air pressure, respectively, in atm; The transport coefficient, representing how the potential change at the reaction interface alters the magnitude of the forward and reverse activation barriers, is set to 0.5; n is the number of electrons transferred in the electrochemical reaction, which is 2 in this reaction; F is the Faraday constant; j is the current density in A / cm²; jleak is the leakage current density in A / cm²; j0 is the exchange current density in A / cm²; ASR is the area-normalized fuel cell resistance in A / cm². ; is an empirical constant; jL is the limiting current density, in A / cm2.
[0116] Step S4: The aging parameters in the voltage model, including the exchange current density j0, the effective reaction area A, the limiting current density jL, and the area resistance ASRi under each load current, are used to reflect the performance degradation of the fuel cell during operation.
[0117] Step S5: Aging parameter optimization. The Cuckoo algorithm is used on the first data point under each load current to optimize the aging parameters and unknown parameters in the voltage model. The Cuckoo algorithm is used here as the initial value acquisition method for the aging parameters. Based on the initial conditions of each current condition, parameter optimization is performed to obtain a set of possible initial values for the aging parameters that make the calculated initial voltage value close to the true value.
[0118] In this embodiment, the relevant parameters of the Cuckoo algorithm are set as follows:
[0119] Population size: Generally set between 20 and 50. A population size that is too small may result in insufficient search space coverage, making it difficult to find the global optimum; a population size that is too large will increase computational cost and reduce algorithm efficiency. For example, setting the population size to 30 achieves a good balance between computational efficiency and search capability.
[0120] Maximum number of iterations: Typically set to 100-500. Too few iterations may prevent the algorithm from converging to the optimal solution; too many iterations will increase computation time and may have reached the algorithm's search limit.
[0121] For example, setting the maximum number of iterations to 200 can meet the optimization needs in most cases.
[0122] Step size factor: Generally, the value is between 0.1 and 0.5. A step size factor that is too large may cause the search process to skip the optimal solution; a step size factor that is too small will slow down the search. For example, setting the step size factor to 0.2 allows for a reasonable trade-off between search accuracy and speed.
[0123] Step S6: Assign agents based on aging parameter characteristics. Exchange current density j0, effective reaction area A, and limiting current density jL do not change significantly under different current densities, and are therefore current density insensitive parameters, assigned to agent 0. Area resistivity ASR differs significantly under different current densities, and is therefore current density sensitive, assigned according to load current conditions. If the fuel cell operates under load currents I1, I2, I3, ..., In, then its corresponding area resistivity ASR1, ASR2, ASR3, ..., ASRn are assigned to n different agents: agent 1, agent 2, agent 3, ..., agent n.
[0124] Step S7: Construct the network of each agent. The specific steps are as follows.
[0125] Step S7.1: Establish the network structure. Each agent contains two neural networks—a DQN evaluation network and a target network. The two networks have identical structures but different parameters. The network adopts a two-layer fully connected structure, mathematically expressed as:
[0126]
[0127] In the formula, and These are the trainable parameter matrices for the hidden layer and the output layer, respectively. To observe the spatial dimension, The number of hidden layer neurons. For the state dimension; This represents the ReLU activation function. The input layer receives data in dimension 1. Observation state vector After regularization Linear transformation projection to The final output after nonlinear activation processing of the 3D hidden layer space is... State value estimates under different actions.
[0128] The DQN evaluation network for agent i is a neural network that approximates the optimal action-value function. Its input is the agent's local observations. and the chosen action The output is a scalar representing a score for selecting the action given the observation. The target network of agent i is a copy of the DQN evaluation network. A schematic diagram of the distributed independent Q-learning algorithm is shown below. Figure 2 As shown.
[0129] Step S7.2, use a greedy algorithm to select an action, expressed as (3):
[0130]
[0131] In the formula, It is a greedy probability, ranging from (0, 1).
[0132] Step S7.3: Set the double-Q learning algorithm as the network update method. The algorithm diagram is shown below. Figure 3 As shown, the estimated maximum value is obtained by forward propagation of the agent's DQN. And choose the action that maximizes value. Use actions Target network in intelligent agent Calculate the value to obtain the estimated maximum value .
[0133] The formula for calculating the TD target is:
[0134]
[0135] In the formula, For TD objectives, As a reward, To balance the discount factor between current rewards and future rewards.
[0136] The formula for calculating TD error is:
[0137]
[0138] In the formula, This is the TD error.
[0139] The DQN network parameters are updated using gradient descent, and the formula is:
[0140]
[0141] In the formula, For the updated network parameters, For the network parameters before the update, This is the learning rate.
[0142] The target network parameters are updated using a weighted average soft update, as shown in the formula:
[0143]
[0144] In the formula, This is the soft update coefficient.
[0145] Step S8, construct the reinforcement learning environment, the specific steps are as follows:
[0146] Step S8.1: Use the voltage model as the voltage calculation method and set the reward formula as follows:
[0147]
[0148] In the formula, R is the reward of the reinforcement learning algorithm; This is the predicted voltage value; This is a voltage measurement value; It is a very small positive constant, the purpose of which is to ensure that the denominator of equation (8) is not 0, and that each agent shares the reward R.
[0149] Step S8.2 involves converting the selected action into a number base and outputting it as a specific operation on each aging parameter of the fuel cell. In this embodiment, five types of actions are defined, represented by [0, 1, 2, 3, 4], as shown in Table 1.
[0150] Table 1. Fuzzy rule base for the first-level classifier
[0151]
[0152] Agent 0 needs to operate on all current density-insensitive aging parameters (3 in total), and each parameter has 5 possible actions, so Agent 1 has a total of 53 (i.e. 125) action types. Agents 2 to n only need to operate on their respective aging parameter ASR, so each agent has 5 action types.
[0153] Since the actions are discrete, the Discrete class from the Open AI Gym library is used to describe the discrete action space. The number of action types is assigned to the Discrete class's attribute k. When the function sample is called, it returns an integer m between [0, k-1]. For agent 0, the integer m is in the range [0, 124], which can be converted to a three-digit pentatonic number, where each digit represents the operation type of the parameter. For example, if the returned integer is 107, it is converted to the pentatonic number "412", representing the first aging parameter, the exchange current density. Significantly increased, the second aging parameter, leakage current density. The third aging parameter, effective reaction area, decreased slightly. Unchanged. For agents 2 to 5, the integer m is in the range [0, 4], and the integer m can directly represent the operation type of the aging parameter ASR.
[0154] There is no direct information exchange between the agents; instead, information is exchanged indirectly through the environment. When an agent manipulates aging parameters, this information is transmitted to the Env environment. Then, when another agent needs environmental information at the next moment, it can obtain the aforementioned aging parameter information to complete the information exchange. Therefore, the frequency and triggering conditions of the interaction depend on the load current conditions. For example, under conditions where the load current changes frequently, the information exchange between agents will be more frequent.
[0155] Step S9: Set up a dedicated experience pool for each agent. (The sample...) When storing data into the experience pool, set the sampling probability using the following formula:
[0156]
[0157] In the formula, For the j-th experience of agent i; Let be the gradient of the DQN network for agent i; It is a very small constant to prevent the sampling probability from approaching zero, and is used to ensure that all samples are drawn with a non-zero probability.
[0158] The learning rate for the samples is set using the following formula:
[0159]
[0160] In the formula, Let be the learning rate for sample j; b is the general learning rate set; b is the total number of samples in the priority experience replay array. It's a hyperparameter.
[0161] Step S10, train the model, the training process is as follows: Figure 4 As shown. The specific steps are as follows:
[0162] Step S10.1: Set network parameters. Set parameters such as the number of layers, number of neurons, learning rate, and discount factor of the neural network according to the actual situation.
[0163] Step S10.2: Use the aging parameter optimization result from step S7 as the initial state of the reinforcement learning model, and set the time interval. .
[0164] Step S10.3: Set time t+1, select the sample corresponding to time t, and determine the load current i.
[0165] Step S10.4, agents 0 and i enter the active state, and the local observation state is changed. and Input their respective DQN networks and select actions using a greedy strategy. and .
[0166] Step S10.5: Enter the Env environment and proceed according to the action. and By manipulating the aging parameters, the local observation state at the next moment can be obtained. and Combining the operating parameters at time t, the predicted voltage is calculated using the voltage model in step S5. The current reward is obtained through formula (8). .
[0167] Step S10.5, apply experience and The experience replays are added to the respective experience replay pools of agent 0 and agent i, and it is determined whether the batch size has been reached. If it has, proceed to step S10.6; otherwise, return to step S10.3. Additionally, it is determined whether the number of experience entries in the experience pool has reached the experience pool size; if it has, the oldest experience is deleted.
[0168] Step S10.6: Sample from the experience replay pools of agent 0 and agent i respectively to obtain their respective batches.
[0169] In step S10.7, each batch is input into its DQN network and the target network, and the network is updated using the double Q learning algorithm.
[0170] Step S10.8: Determine if the training set has ended. If it has ended, proceed to the prediction phase; otherwise, return to step t+1.
[0171] Step S11, predict voltage, the prediction process is as follows: Figure 5 As shown, the specific steps are as follows:
[0172] Step S11.1: Set the sliding window size L and the time t.
[0173] Step S11.2: Set time t+1, select the sample corresponding to time t, and determine the load current i.
[0174] Step S11.3, agents 0 and i enter the active state, and the local observation state is changed. and Each input DQN network selects an action using the greedy strategy in step S9.2. and .
[0175] Step S11.4: Enter the Env environment and proceed according to the action. and By manipulating the aging parameters, the local observation state at the next moment can be obtained. and Combining the operating parameters at time t, the predicted voltage is calculated using the voltage model in step S5. .
[0176] Step S11.5: Determine whether the current time has reached the end of the sliding window. If so, retrieve the true voltage values of all data samples within the sliding window, obtain all rewards through equation (8), and put all experiences into their respective experience replay pools. Determine whether the number of experience entries in the experience pool has reached the size of the experience pool. If it has exceeded the size, delete the earliest experience and proceed to the next step; otherwise, return to the step at time t+1.
[0177] Step S11.6: Sample from the experience replay pools of agent 0 and agent i respectively to obtain their respective batches.
[0178] In step S11.7, each batch is input into its DQN network and the target network, and the network is updated using the double Q learning algorithm.
[0179] Step S11.8: Determine if the training set has ended. If it has, the prediction is complete, and the voltage prediction result is obtained. Refer to... Figure 6Otherwise, move the sliding window to the next cycle and set the step at time t+1.
[0180] This application embodiment also provides an electronic device 500 that utilizes the aforementioned variable load fuel cell voltage decay prediction method. The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the aforementioned variable load fuel cell voltage decay prediction method. In this application embodiment, the processor is the control center of the computer method and can be a physical machine processor or a virtual machine processor.
[0181] Reference Figure 7 The electronic device 500 includes at least one processor 501, at least one communication interface 502, at least one memory 503, and at least one bus 504. The bus 504 is used for communication between these components, the communication interface 502 is used for signaling or data communication with other node devices, and the memory 503 stores machine-readable instructions executable by the processor 501. When the electronic device 500 is running, the processor 501 communicates with the memory 503 via the bus 504. When the machine-readable instructions are invoked by the processor 501, they execute the steps of the variable load fuel cell voltage decay prediction method described above.
[0182] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor of an electronic device, can implement the steps of the variable load fuel cell voltage decay prediction method described above.
[0183] Those skilled in the art will understand that implementing all or part of the process in a variable load fuel cell voltage decay prediction method can be accomplished by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium. When executed, the program can include the processes of various embodiments of the variable load fuel cell voltage decay prediction method. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0184] The above content is merely an embodiment of the present invention. Commonly known structures and characteristics of the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all prior art in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can improve and implement this solution based on the guidance provided in this application and their own capabilities. Some typical well-known structures or systems should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. A method for predicting voltage decay in a variable load fuel cell, characterized in that, include, Extract the operating parameters and calculate their characteristic values; A voltage model is established, with the characteristic values of operating parameters and aging parameters as inputs and the fuel cell stack voltage as the output. The aging parameters of the first data sample under each load current are optimized using the Cuckoo Optimization Algorithm to obtain the initial values of each aging parameter. Agents are assigned based on aging parameter characteristics, where: Agents are assigned based on current density-insensitive aging parameters and current density-sensitive aging parameters; Construct networks for each agent, including a DQN evaluation network and a target network. Use a greedy algorithm to select actions and a double-Q learning algorithm to update network parameters. This includes the following: The network adopts a two-layer fully connected structure, which can be mathematically expressed as: In the formula, and These are the trainable parameter matrices for the hidden layer and the output layer, respectively. To observe the spatial dimension, The number of hidden layer neurons. For the state dimension; Represents the ReLU activation function; the input layer receives data in dimension 1. Observation state vector After regularization Linear transformation projection to The final output after nonlinear activation processing of the 3D hidden layer space is... State value estimates under different actions; The action is selected using a greedy algorithm, expressed as: In the formula, It is a greedy probability, ranging from (0, 1); The maximum value is estimated by forward propagation of the agent DQN. And choose the action that maximizes value. Use actions Target network in intelligent agent Calculate the value to obtain the estimated maximum value ; A reinforcement learning environment is constructed, using the voltage model as the voltage calculation formula and setting a reward formula. The discrete actions selected by the agent are converted into operations on various aging parameters of the fuel cell through base conversion, where: The reward function is based on the error between the predicted voltage value and the measured voltage value, and the formula is: In the formula, R is the reward of the reinforcement learning algorithm; This is the predicted voltage value; This is a voltage measurement value; It is a positive constant; all agents share the reward R; Each agent is assigned a dedicated experience pool to perform model training and voltage prediction.
2. The method for predicting voltage decay in a variable load fuel cell according to claim 1, characterized in that, Before extracting operating parameters, data preprocessing is performed, including removing non-steady-state data that deviates from the preset load current ±1A range, verifying data integrity, identifying and deleting outliers using the box plot method, and downsampling to reduce data frequency; The extracted operating parameters include stack voltage, anode inlet and outlet gas pressure, cathode inlet and outlet gas pressure, air flow rate, and cooling water inlet and outlet temperatures. The moving average algorithm is used to smooth and correct the time series data.
3. The method for predicting voltage decay in a variable load fuel cell according to claim 1, characterized in that: Calculate the characteristic values of the operating parameters, including the following: Calculate the average voltage of all sections as the voltage characteristic value; The average gas pressure at the anode inlet and outlet is calculated as the characteristic value of hydrogen pressure. The average gas pressure at the cathode inlet and outlet is calculated as the characteristic value of air pressure; the average temperature of cooling water entering and exiting the stack is calculated as the characteristic value of stack temperature.
4. The method for predicting voltage decay in a variable load fuel cell according to claim 1, characterized in that, The voltage model is shown in equation (1): In the formula, T is the fuel cell temperature; R is the ideal gas constant; and These are hydrogen pressure and air pressure, respectively. The transport coefficient is used to indicate how changes in potential at the reaction interface alter the magnitude of the forward and reverse activation barriers; n is the number of electrons transferred in the electrochemical reaction; F is the Faraday constant; j is the current density; jleak is the leakage current density; j0 is the exchange current density; ASR is the area-normalized fuel cell resistance. j is an empirical constant; jL is the limiting current density; Current density-insensitive aging parameters include exchange current density j0, effective reaction area A, and limiting current density jL. Current density-sensitive aging parameters include area-normalized fuel cell resistance ASRi under various load currents.
5. The method for predicting voltage decay in a variable load fuel cell according to claim 1, characterized in that, The specific process of network update is as follows: The formula for calculating the TD target is: In the formula, For TD objectives, As a reward, To balance the discount factor between current rewards and future rewards; The formula for calculating TD error is: In the formula, For TD error; The DQN network parameters are updated using gradient descent, and the formula is: In the formula, For the updated network parameters, For the network parameters before the update, The learning rate; The target network parameters are updated using a weighted average soft update, as shown in the formula: In the formula, This is the soft update coefficient.
6. The method for predicting voltage decay in a variable load fuel cell according to claim 1, characterized in that, Set up a dedicated experience pool for each agent, including the following: Sample When storing data into the experience pool, set the sampling probability using the following formula: In the formula, For the j-th experience of agent i; Let be the gradient of the DQN network for agent i; It is a constant used to ensure that all samples are drawn with a non-zero probability; The learning rate for the samples is set using the following formula: In the formula, Let be the learning rate for sample j; The general learning rate is set; b is the total number of samples in the priority experience replay array; It's a hyperparameter.
7. The method for predicting voltage decay in a variable load fuel cell according to claim 1, characterized in that, Performing model training includes: Use the initial values of the aging parameters as the initial state of the reinforcement learning model; The corresponding agent is activated based on the load current, and each agent selects an action based on the local observation state using a greedy strategy. The aging parameters are updated based on the actions, the predicted voltage and reward are calculated, and the experience is stored in the corresponding experience pool. The network parameters of each agent are sampled from the experience pool and updated using the double-Q learning algorithm.
8. The method for predicting voltage decay in a variable load fuel cell according to claim 1, characterized in that, Perform voltage prediction, including: The sliding window mechanism is used to select sample data for the time to be predicted; The corresponding intelligent agent is activated to update the aging parameters, and then the predicted voltage is calculated through the voltage model; At the end of the sliding window, the predicted results are compared with the actual values to update the experience pool, and the next round of prediction continues.
Citation Information
Patent Citations
Fuel cell aging prediction method and system, medium and product
CN118393389A
Multi-stack fuel cell life convergence control method and system based on cloud model
CN118782829A