Fuel cell performance degradation long-term prediction method based on deep reinforcement learning
Through a hybrid model based on deep reinforcement learning, combined with semi-empirical and semi-mechanical models and data-driven models, the problem of insufficient prediction accuracy of fuel cell performance decay is solved, and long-term prediction with high accuracy and high robustness is achieved, which enhances the explanatory and generalization capabilities of the model.
Patent Information
- Application Number
- CN202510371182.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-01
AI Technical Summary
The existing fuel cell performance decay prediction methods have problems such as insufficient accuracy, poor versatility and the inability to achieve quantitative accuracy indicators for long-term predictions.
A hybrid model based on deep reinforcement learning is adopted, combined with a semi-empirical semi-mechanical model and a data-driven model, input features are screened through the Shapley additive interpretation method, parameters are calibrated using genetic algorithms, long and short-term memory neural network model is constructed, and dynamic weighted prediction is achieved through the actor-critician algorithm.
It improves the accuracy and robustness of long-term prediction of fuel cell performance decay, realizes accurate evaluation of future information of fuel cell system, and enhances the interpretability and generalization capabilities of the model.
Smart Images

Figure CN120233263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fuel cells, and particularly to a long-term prediction method for fuel cell performance degradation based on a hybrid model. Background Art
[0002] Under the macro background of global climate change and energy structure transformation, the strategic position of renewable energy has become increasingly prominent. Compared with energy forms such as hydropower, wind power, and solar energy, which are restricted by geography, time, and volatility, hydrogen energy stands out due to its high stability. Among them, fuel cells, especially proton exchange membrane fuel cells (PEMFCs), have become an important force in promoting the realization of the "dual-carbon" goal due to their advantages such as high energy conversion efficiency, fast response ability, zero pollution emissions, and low-noise operation. However, the insufficient durability of fuel cells is one of the current bottlenecks restricting their full commercialization. Degradation prediction technology can prospectively provide the health status information of fuel cell systems, providing a solid theoretical basis and data support for optimizing operating conditions, formulating maintenance strategies, and making scientific decisions. Therefore, the prediction research on fuel cell performance degradation has significant practical value for improving its reliability and durability.
[0003] Currently, there are three main types of methods for fuel cell degradation prediction: prediction based on mechanism / empirical models, prediction based on data-driven models, and prediction based on hybrid models. Mechanism-based methods are limited by the complex degradation mechanisms inside fuel cells and cannot establish accurate prediction models. Empirical-based methods rely on experimental data and have poor generality. Data-driven methods usually use machine learning methods, which do not require knowledge of the degradation mechanism but need a large amount of data to discover internal non-linear characteristics and have a large error in the later stage of prediction. Hybrid model-based methods can integrate the advantages of both, capture non-linear characteristics, and have a certain interpretability of the results. However, in the weighted summation process, it is impossible to change dynamically without prior knowledge. In addition, there are no relevant quantitative accuracy indicators for long-term degradation prediction (hundreds of hours) at present.
[0004] Therefore, if a long-term prediction method for fuel cell performance degradation with high precision and high robustness can be developed to accurately evaluate the future information of fuel cell systems, it will naturally become a great blessing for adjusting and optimizing fuel cell operation strategies and improving their durability in the development of fuel cell technology. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a long-term prediction method for fuel cell performance degradation based on deep reinforcement learning. To achieve the above object, the technical solution adopted by the present invention is:
[0006] Long-term prediction method for fuel cell performance degradation based on deep reinforcement learning, the method comprising the following steps:
[0007] S1. Obtain experimental data during the fuel cell degradation process. First, filter the original voltage data to remove spikes and noise, and smooth the curve to facilitate subsequent training and prediction.
[0008] S2. Use the Shapley Additive Explanation method to screen the data, calculate the marginal contribution of each input feature to the model output, and select several factors that have the greatest impact on the voltage as inputs.
[0009] S3. Based on the Nernst voltage and voltage loss calculation expressions, construct a semi-empirical and semi-mechanistic degradation model. This model describes the degradation process by designing degradation factors, including leakage current density, electrochemically active area, proton resistance, electron resistance, and gas transport coefficient, and uses the genetic algorithm to calibrate and fit the polarization curve and degradation curve respectively to obtain the initial parameters and degradation parameters.
[0010] S4. Use the long short-term memory neural network method to construct a data-driven model, divide the training set and test set, and perform long-term prediction on the fuel cell performance.
[0011] S5. Construct a hybrid model, use the actor-critic algorithm to dynamically allocate weights to the prediction results of the semi-empirical and semi-mechanistic degradation model and the data-driven model, use the weighted value as the final model output result, and calculate its correlation coefficient.
[0012] First, construct a semi-empirical and semi-mechanistic (SEMI) model and a data-driven model to predict the battery voltage respectively, and then use the deep reinforcement learning method to dynamically weight the results of the two models to achieve long-term prediction of fuel cell performance degradation.
[0013] Then, each of the above steps is further defined.
[0014] The characteristics and beneficial effects of the present invention are:
[0015] (1) In the present invention, the SHAP method is used to screen the original degradation data, which can greatly reduce the number of input features. This process can significantly reduce the complexity of the model, reduce the dependence on the data set, improve the calculation efficiency without sacrificing accuracy, and broaden its application scenarios under limited data.
[0016] (2) The hybrid model constructed in the present invention first proposes a quantization index in the long-term prediction process. The hybrid model can not only capture the advantages of local non-linear characteristics in the voltage decay process, but also endow the model with a certain interpretability. First, two models are used for prediction respectively, and then the dynamic weighting process is realized through the deep reinforcement learning algorithm (Actor-Critic algorithm). The model can adaptively adjust the weights according to different input features and states, so as to better adapt to the complex and changeable voltage decay process. This dynamic adjustment mechanism not only improves the prediction ability of the model, but also enhances its generalization ability on different data sets, making it have a wide application prospect in the prediction of fuel cell performance decay. Description of the Drawings
[0017] Figure 1 It is the overall method flow chart of the present invention.
[0018] Figure 2 It is the original data and preprocessing result diagram of the fuel cell decay experiment in the embodiment of the present invention.
[0019] Figure 3 It is the feature contribution value result diagram based on the SHAP method in the embodiment of the present invention.
[0020] Figure 4 It is the initial polarization curve fitting result diagram of the semi-empirical and semi-mechanistic model in the embodiment of the present invention.
[0021] Figure 5 It is the decay curve fitting result diagram of the semi-empirical and semi-mechanistic model in the embodiment of the present invention.
[0022] Figure 6 It is the prediction result diagram of the data-driven model in the embodiment of the present invention.
[0023] Figure 7 It is the prediction result diagram of the fuel cell performance decay hybrid model based on reinforcement learning in the embodiment of the present invention. Detailed Embodiments
[0024] The following further illustrates the method steps of the present invention through specific calculation examples. It should be noted that this embodiment is narrative rather than restrictive, and does not limit the protection scope of the present invention.
[0025] A long-term prediction method for fuel cell performance decay based on deep reinforcement learning, which specifically includes:
[0026] S1, obtaining the experimental data in the fuel cell decay process. First, filter the original voltage data to remove spikes and noise, and smooth the curve to facilitate subsequent training and prediction.
[0027] S2. Use the SHapley Additive exPlanations (SHAP) method to screen the data, calculate the marginal contribution of each input feature to the model output, and select several factors that have the greatest impact on the voltage as the inputs.
[0028] S3. Based on the Nernst voltage and the voltage loss calculation expressions, construct a semi-empirical and semi-mechanistic degradation model. This model describes the degradation process by designing a degradation factor, including leakage current density, electrochemically active area, proton resistance, electron resistance, and gas transport coefficient, and uses the genetic algorithm to calibrate and fit the polarization curve and the degradation curve respectively to obtain the initial parameters and degradation parameters.
[0029] S4. Adopt the method of Long short-term memory network (LSTM) to construct a data-driven model, divide the training set and the test set, select appropriate hyperparameters, and make long-term predictions on the performance of the fuel cell.
[0030] S5. Construct a hybrid model, use the actor-critic algorithm to dynamically assign weights to the prediction results of the semi-empirical and semi-mechanistic degradation model and the data-driven model, take the weighted value as the final model output result, and calculate its correlation coefficient R 2 。
[0031] In step S1, the filtering method used is Savitzky-Golay, which is a digital filtering technique based on local polynomial least squares fitting. It realizes smoothing and noise reduction by performing polynomial fitting on data points through a sliding window, while retaining the shape and key features of the signal (such as peaks and trends). Its core advantage lies in flexible parameters (adjustable window size and polynomial order), and it is applicable to fields such as spectral analysis and signal processing.
[0032] In step S2, SHAP (SHapley Additive exPlanations) is a machine learning model interpretation method based on the Shapley value in game theory. By calculating the mean of the marginal contributions of each feature under different feature combinations, it quantifies its impact on the model prediction, thereby screening out the features that have the greatest impact on the predicted value. The DeepExplainer interpreter is applied to solve the Shapley value of each feature.
[0033] In step S3, the construction and parameter calibration of the semi-empirical and semi-mechanistic model include the following steps:
[0034] S31: The components involved in the degradation of a fuel cell include: the membrane, the catalyst layer, and the gas diffusion layer. Among them, voltage is used as an indicator of degradation. First, an electrochemical model is established based on the Nernst voltage equation. Considering the processes of activation loss, ohmic loss, and concentration loss decaying over time, corresponding degradation factors are specifically designed. The coupled formula is as follows:
[0035]
[0036] Where, V ner (V) is the Nernst voltage, I (A) represents the current, d is the degradation factor, which is used to characterize the degradation of the component over time. The subscript leak represents the leakage current density, ECSA represents the effective reaction area, ion represents the proton resistance, ele represents the electron resistance, B represents the concentration loss correction coefficient, and D represents the gas diffusion coefficient.
[0037] It can be seen from the formula that the degradation rates of each component are different, divided into two categories: linear and exponential decay. A initial is the initial effective reaction area, R ion,initial is the initial proton resistance (Ωm 2 ), R ele,initial (Ωm 2 ) is the electron resistance, K c,initial is the initial concentration loss coefficient.
[0038] S32: According to the polarization curve of the fuel cell in the initial state, the genetic algorithm is used to calibrate the initial parameters, and the root mean square error (RMSE) between the simulated value and the true value is used as the fitness function. The initial parameters i 0,a is the cathode reference exchange current density; i 0,c is the anode reference exchange current density, α a is the anode reaction transfer coefficient; α c is the cathode reaction transfer coefficient, i leak,initial is the initial leakage current density, R ion,initial is the initial proton resistance, R ele,initial is the initial electron resistance, K c,initial is the initial concentration loss coefficient.
[0039] S33: According to the experimental data of the fuel cell degradation process, the genetic algorithm is used to calibrate the degradation factors, and the root mean square error (RMSE) between the simulated value and the true value is used as the fitness function. The degradation factor d mainly includes: the leakage current density d leak , the effective reaction area d ECSA , the proton resistance d ion , the electron resistance d ele , the concentration loss correction coefficient d B and the gas diffusion coefficient d D .
[0040] In step S4, a data-driven model is constructed using the long short-term memory neural network (LSTM) method. The training set and the test set are divided, appropriate hyperparameters are selected, and long-term prediction of the fuel cell performance is carried out. Dropout is adopted in the LSTM model to prevent overfitting of the model and improve the generalization ability of the model.
[0041] In step S5, the construction of the hybrid model includes the following steps:
[0042] S51: Collect the voltage prediction values of the data-driven model and the mechanism model. The data-driven model is denoted as V1; the mechanism model is denoted as V2.
[0043] S52: Configure the environment of the deep reinforcement learning algorithm. Among them, the state is set to the predicted voltage values of the data-driven model V1 and the mechanism model V2, the action is set to the weight w of the data-driven model, and the reward is set to the error between the hybrid model value and the true value. When the error is within a certain range (which can be set by oneself), the reward is positive, otherwise it is negative.
[0044] S53: Divide the data-driven model V1 and the mechanism model V2 into the training set and the test set according to the same ratio. The purpose of the training set is to enable the neural network to learn the optimal weight allocation strategy, while the test set is to obtain high-precision voltage prediction values.
[0045] S54: The actor-critic algorithm adopted by the hybrid model includes an actor neural network and a critic neural network. In the training stage, the role of the actor is to determine the optimal weight strategy π θ , in order to obtain the maximum cumulative reward, and the output of each step is the weighted sum of the models:
[0046] V hybrid = wV1+(1 - w)V2 (2)
[0047] Where: V hybrid is the output voltage value of the hybrid model.
[0048] The critic is responsible for evaluating the quality of the actions selected by the actor through the value function, so as to optimize the action strategy of the actor.
[0049] In the test stage, the optimal weight strategy learned in the training stage is used to dynamically sum the mechanism model and the data-driven model to obtain the fuel cell voltage value in the prediction stage.
[0050] S55: The calculation formula of the correlation coefficient R 2 is as follows:
[0051]
[0052] Among them, y i is the true value of the voltage, is the predicted value of the voltage, is the average value of the predicted voltage values. Specific embodiments
[0054] S1 (such as Figure 1 ), obtain the experimental data of the 2014 IEEE PHM Data Challenge. Since the original dataset is too large, select the first 960h of data and perform interval sampling on it. Then use the Savizky-Golay filter to process the original data, remove abnormal data such as noise and spikes, smooth the curve, and obtain the experimental dataset for neural network training.
[0055] S2, the original dataset contains 19 features, namely voltage, time, air-side inlet pressure, air-side outlet pressure, hydrogen inlet temperature, current, cooling water outlet temperature, hydrogen inlet pressure, air outlet temperature, hydrogen outlet flow rate, air-side relative humidity, hydrogen outlet temperature, air outlet flow rate, cooling water inlet temperature, air inlet temperature, hydrogen outlet pressure, cooling water flow rate, air inlet flow rate, hydrogen inlet flow rate. The Deep Explainer is used, which can efficiently calculate the contribution of features to the model. After calculation, select the first 7 features as the input.
[0056] Furthermore, step S3 includes:
[0057] In S31, the meanings of the parameters in the formula will not be elaborated. Based on the initial polarization curve data of the data challenge, use the genetic algorithm to calibrate the initial parameters (i 0,a , i 0,c , R ion,initial etc., a total of ten) in the semi-empirical semi-mechanistic model. The fitness function selects the reciprocal of the RMSE of the simulated voltage and the experimental voltage value. Figure 3 Shows the comparison of the data of the polarization curve simulated by the semi-empirical semi-mechanistic model and the real polarization curve. It can be seen that the results of the two models are in good agreement at different current densities, and the simulation results are highly consistent with the experimental results.
[0058] S32: Next, calibrate the decay factors (d leak , d ECSA , etc.) based on the 880h experimental dataset. Still use the genetic algorithm, and the fitness function selects the reciprocal of the RMSE of the simulated voltage and the experimental voltage value, Figure 4Shows the entire decline process (0 - 880h). The decline curves of the voltage values obtained from model simulation and the smoothed voltage values of the real data over time are presented. It can be seen that both show a linear downward trend, and the overall trends match well. However, during the local decline process, the experimental data shows a non-linear downward trend, while the semi-empirical and semi-mechanistic model has a poor fitting effect and cannot capture these characteristics.
[0059] Furthermore, step S4 includes:
[0060] S41: First, divide the training set and the test set. Since the LSTM network is generally used for time series prediction, the first 30% of the data (t < 228h) is divided as the training set, and the remaining 70% of the data (t > 228h) is used as the test set.
[0061] S42: Set the hyperparameters of the LSTM network. Configure 192 hidden neurons, and the learning rate is 1×10 -5 , and it is trained through 2000 iterations. In addition, dropout is set to 0.15.
[0062] Step S5 includes:
[0063] S51: Collect the voltage prediction values of the mechanistic model and the data-driven model, denoted as V1 (LSTM model) and V2 (semi-empirical and semi-mechanistic model) respectively.
[0064] S52: Configure the environment of the deep reinforcement learning algorithm. Among them, the state is set as the predicted voltage values of the two models (V1, V2), the action is set as the weight w of the data-driven model, and the reward is set as the error between the hybrid model value and the real value. When the error is within 10mV, the reward is positive, otherwise it is negative.
[0065] S53: Divide V1 and V2 into the training set and the test set according to the ratio of 3:7. The purpose of the training set is to enable the neural network to learn the optimal action strategy, while the test set is for verifying the action strategy.
[0066] S54: The model uses the AC algorithm to dynamically optimize the weight w in the training set, seek the best weighted strategy for the two models, and conduct a generalization test on the weighted strategy in the test set.
[0067] S55: Calculate the R 2 of the individual semi-empirical and semi-mechanistic model, data-driven model, and hybrid model respectively. The calculation formula is as follows:
[0068]
[0069] Among them, y i is the real voltage value, is the predicted voltage value, is the average value of the voltage prediction.
[0070] Figure 6 Shows the comparison of three models (semi-empirical semi-mechanistic model, data-driven model, hybrid model) with the experimental values. The correlation coefficient R of the semi-empirical semi-mechanistic model is calculated 2 to be 0.821; the correlation coefficient R of the data-driven model 2 is 0.803; the correlation coefficient R of the hybrid model 2 is 0.865. The R of the hybrid model 2 is the highest and the fitting effect is the best. It can not only capture the non-linear characteristics in the local degradation process of the fuel cell, but also ensure the linear decline trend of the overall degradation. It not only ensures the interpretability of the model, but also reduces the iterative divergence phenomenon of the data-driven model in the later stage of long-term prediction. Compared with the semi-empirical semi-mechanistic model and the data-driven model, the accuracy of the hybrid model is improved by 5.36% and 7.72% respectively, verifying the accuracy of this method.
Claims
1. A long-term prediction method for fuel cell performance degradation based on deep reinforcement learning, characterized by: The method comprises the following steps: S1, obtain the experimental data of the fuel cell decay process, first filter the original voltage data to remove the peak and noise, smooth the curve, and facilitate subsequent training and prediction; S2, the Shapley additive interpretation method is used to screen the data, calculate the marginal contribution of each input feature to the model output, and select the factors that have the greatest impact on voltage as input; S3, based on the Nernst voltage and voltage loss calculation expressions, a semi-empirical and semi-mechanistic decay model is constructed. The decay process is described by designing decay factors, including leakage density, electrochemical active area, proton resistance, electronic resistance, and gas transfer coefficient. Genetic algorithms are used to calibrate and fit the polarization curve and decay curve, respectively, to obtain initial parameters and decay parameters. S4, using the long short-term memory neural network method, constructing a data-driven model, dividing the training set and the test set, and making long-term predictions on fuel cell performance; S5, construct a hybrid model, use the actor-critic algorithm to dynamically assign weights to the prediction results of the semi-empirical and semi-mechanistic decay model and the data-driven model, use the weighted value as the final model output result, and calculate its correlation coefficient.
2. The method for predicting fuel cell performance degradation based on deep reinforcement learning according to claim 1 is characterized in that: In step S1, the filtering method used is Savitzky-Golay, which is used in the fields of signal processing and spectral analysis. It smoothes noise or calculates derivatives by fitting polynomials to data points within a sliding window while retaining key features of the signal, such as peaks and trends.
3. The method for predicting fuel cell performance degradation based on deep reinforcement learning according to claim 1 is characterized in that: In step S2, the SHAP value of the input feature is calculated, and the deep learning interpreter is applied to solve the importance of each feature, and the factor with the greatest impact on the output is selected according to the importance.
4. The method for predicting fuel cell performance degradation based on deep reinforcement learning according to claim 1 is characterized in that: In step S3, the construction and parameter calibration of the semi-empirical and semi-mechanistic model include the following steps: S31: The components involved in the degradation of fuel cells include: membrane, catalyst layer and gas diffusion layer. Among them, voltage is used as a degradation indicator. First, an electrochemical model is established based on the Nernst voltage equation. Considering the process of activation loss, ohmic loss and concentration loss decaying over time, a corresponding degradation factor is specially designed. The coupled formula is as follows: Among them, V ner is the Nernst voltage, I is the current, d is the decay factor, which is used to characterize the degradation of the component over time. The subscript leak represents the leakage density, ECSA represents the effective reaction area, ion represents the proton resistance, ele represents the electronic resistance, B represents the concentration loss correction factor, and D represents the gas diffusion coefficient. It can be seen from the formula that the degradation rate of each component is different, which can be divided into two categories: linear and exponential decay. A initial is the initial effective reaction area, R ion,initial is the initial proton resistance, R ele,initia is the electronic resistance, K c,initial is the initial concentration loss coefficient; S32: According to the polarization curve of the fuel cell in the initial state, the initial parameters are calibrated using a genetic algorithm, and the root mean square error between the simulated value and the true value is used as the fitness function. The initial parameter i 0,a is the cathode reference exchange density; i 0,c is the anode reference exchange density, α a is the anode reaction transport coefficient; α c is the cathode reaction transport coefficient, i leak,initial is the initial leakage current, R ion,initial is the initial proton resistance, R ele,initial is the initial electronic resistance, K c,initial is the initial concentration loss coefficient; S33: Based on the experimental data of the fuel cell decay process, the decay factor is calibrated using a genetic algorithm, and the root mean square error between the simulated value and the true value is used as the fitness function. The decay factor d mainly includes: the leakage current d leak , effective reaction area d ECSA , proton resistance d ion , electronic resistance d ele , concentration loss correction factor d B and the gas diffusion coefficient d D .
5. The method for predicting fuel cell performance degradation based on deep reinforcement learning according to claim 1 is characterized in that: In step S4, random inactivation is used in the long short-term memory neural network model to prevent overfitting of the model and improve the generalization ability of the model.
6. The method for predicting fuel cell performance degradation based on deep reinforcement learning according to claim 1 is characterized in that: In step S5, the construction of the hybrid model includes the following steps: S51: collecting voltage prediction values of the data-driven model and the mechanism model, the data-driven model is recorded as V1; the mechanism model is recorded as V2; S52: Configure the environment of the deep reinforcement learning algorithm, where the state is set to the predicted voltage value of the data-driven model V1 and the mechanism model V2, the action is set to the weight of the data-driven model, and the reward is set to the error between the mixed model value and the true value. The reward is positive, otherwise it is negative; S53: Divide the data-driven model and the mechanism model into a training set and a test set in the same proportion. The purpose of the training set is to enable the neural network to learn the optimal weight allocation strategy, and the purpose of the test set is to obtain a high-precision voltage prediction value. S54: The actor-critic algorithm used in the hybrid model contains an actor neural network and a critic neural network. During the training phase, the role of the actor is to determine the optimal weight strategy to obtain the maximum cumulative reward. The output of each step is the weighted sum of the model: <h2 style=";text-align:left;direction:ltr">V<h2 style=";text-align:left;direction:ltr"> hybrid <h2 style=";text-align:left;direction:ltr"> =wV1+(1-w)V2 (2) Where: V hybrid is the output voltage value of the hybrid model, The evaluator is responsible for evaluating the quality of the actions selected by the actor through the value function, thereby optimizing the actor's action strategy. In the testing phase, the optimal weight strategy learned in the training phase is used to dynamically sum the mechanism model and the data-driven model to obtain the fuel cell voltage value in the prediction phase. S55: The calculation formula of the correlation coefficient is as follows: Among them, R 2 is the correlation coefficient, i is the sample index, N is the number of samples, and y i is the true value of voltage, is the voltage prediction value, is the average value of the voltage prediction value.