A PEMFC degradation prediction method based on TD-Informer network integrating dual attention and temporal features

By constructing a TD-Informer network with a dual sparse self-attention mechanism and a temporal feature decoder, the problem of feature interaction redundancy and dilution effect of the traditional Informer model in PEMFC degradation prediction is solved, and high-precision PEMFC degradation state prediction is achieved to meet industry needs.

CN120413713BActive Publication Date: 2025-09-26GUANGXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510890919.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-26
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The traditional Informer model has the problem of feature interaction redundancy caused by the single-dimensional sparse self-attention mechanism and spectral domain feature dilution effect caused by multiple attention stacking in PEMFC degradation prediction, which leads to the expansion of prediction errors in the key degradation stage.

Method used

The TD-Informer network integrates a dual-dimensional sparse self-attention mechanism and a temporal feature decoder. By constructing a dual-dimensional sparse self-attention mechanism and a lightweight temporal feature decoder, it reduces feature interaction redundancy, improves the spectral domain feature retention rate, and achieves high-precision prediction.

Benefits of technology

The prediction error of key degradation stages is significantly reduced, the accuracy and engineering applicability of PEMFC degradation prediction are improved, the structural limitations of traditional models are overcome, and the requirements of industrial development are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120413713B_ABST
    Figure CN120413713B_ABST
Patent Text Reader

Abstract

The present invention discloses a TD-Informer network PEMFC degradation prediction method that integrates dual attention and temporal features, belonging to the technical field of PEMFC degradation prediction. The method comprises: conducting a PEMFC performance degradation experiment to obtain multi-dimensional actual operating data; extracting key characteristic factors using the SHAP value analysis method; then designing and constructing a TD-Informer prediction model that integrates a multi-head dual-sparse self-attention mechanism and a temporal feature decoder; and finally, by integrating the structural advantages of the dual-sparse mechanism and the temporal feature decoder, achieving high-precision prediction of the PEMFC degradation state. This method has strong adaptability to operating conditions, improves the accuracy of PEMFC degradation prediction, and can effectively support PEMFC life prediction and maintenance decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of PEMFC degradation prediction, and in particular relates to a TD-Informer network PEMFC degradation prediction method integrating dual attention and temporal features. Background Art

[0002] Proton exchange membrane fuel cells (PEMFCs), a key device for hydrogen energy conversion and utilization, face performance degradation, a key technical bottleneck restricting their development. After extended operation, PEMFCs typically experience varying degrees of degradation. This irreversible degradation directly impacts the lifespan and economic viability of PEMFCs. Therefore, accurately predicting their degradation state is crucial for lifespan prediction and health management of PEMFCs.

[0003] Data-driven methods have been widely used in PEMFC degradation prediction due to their powerful adaptive learning capabilities and prediction accuracy. By mining the influencing characteristics of multi-dimensional parameters such as voltage and current, these methods can accurately capture early signs of performance degradation without relying on complex mechanism models, and their engineering practicality is significant. These advantages make them an effective technical means for PEMFC health management. However, the single-dimensional sparse self-attention mechanism used in the traditional Informer model suffers from feature interaction redundancy when processing degradation data of PEMFC multi-parameter coupling. At the same time, the stacking of multiple attentions will also cause a spectral domain feature dilution effect, which leads to an expansion of the prediction error band in the key degradation stage.

[0004] To further improve the accuracy of PEMFC degradation prediction and meet engineering adaptability requirements, this paper proposes a TD-Informer network-based PEMFC degradation prediction method that integrates dual-attention and temporal features. By constructing a dual-dimensional sparse self-attention mechanism and a temporal feature decoder module, this approach improves the retention of spectral domain features during key degradation stages, ultimately achieving high-precision PEMFC degradation prediction that meets engineering applicability requirements. Summary of the Invention

[0005] In order to solve the above problems, the present invention proposes a PEMFC degradation prediction method based on a TD-Informer network that integrates dual attention and temporal features.

[0006] The above-mentioned purpose of the present invention is achieved through the following technical solutions:

[0007] A PEMFC degradation prediction method based on a TD-Informer network that integrates dual attention and temporal features includes the following steps:

[0008] S1: Data Collection

[0009] S101: Conduct performance degradation experiments on the PEMFC test bench, record PEMFC state parameters in real time through the data acquisition module, and construct a multi-dimensional parameter degradation data set;

[0010] S102: Selecting the output voltage as a degradation indicator and taking the other recorded state parameters except the voltage as influencing factors;

[0011] S2: Feature Engineering and Sample Construction

[0012] S201: Sampling the multi-dimensional parameter degradation data set;

[0013] S202: Using the SHAP value method, quantify the contribution of each influencing factor to PEMFC voltage degradation, and select the top five influencing factors as the main influencing features;

[0014] S203: extracting time features, voltage, and main impact feature data from the sampled multi-dimensional parameter degradation data set as degradation data, and performing smoothing preprocessing on the degradation data to eliminate instantaneous noise interference;

[0015] S204: Mapping the degradation data to the interval [0, 1] by the maximum-minimum normalization method, and dividing the training set, validation set, and test set in proportion;

[0016] S3: TD-Informer prediction model construction

[0017] S301: Construct a multi-head dual sparse self-attention mechanism;

[0018] S302: Constructing a temporal feature decoder;

[0019] S303: Integrate the multi-head dual sparse self-attention mechanism and the temporal feature decoder to build the TD-Informer prediction model;

[0020] S304: Using the optimizer and loss function, iteratively optimize the TD-Informer prediction model parameters;

[0021] S305: Input the degradation data into TD-Informer for iterative training and verify whether the model has converged. If it meets the expected requirements, the trained TD-Informer prediction model is obtained after the test passes.

[0022] S4: Online prediction of degradation status

[0023] S401: collecting the time characteristics, voltage and main influencing characteristics of the target PEMFC in real time, and performing standardization processing in accordance with S201, S202, S203 and S204;

[0024] S402: Input the processed data into the trained TD-Informer prediction model to output a voltage degradation prediction sequence for a future preset period;

[0025] S403: Comprehensively evaluate the prediction accuracy of the model, generate an evaluation report, and guide maintenance strategies.

[0026] Furthermore, in the S301 process, the multi-head dual sparse self-attention mechanism is constructed including: three learnable matrices W Q 、W K and W V Used to map the input data X into feature matrices Q, K, and V:

[0027]

[0028]

[0029]

[0030] Where Q is the query matrix, K is the key matrix, and V is the value matrix. By jointly optimizing the sparsity patterns of the query matrix and the key matrix, a multi-head dual-sparse self-attention mechanism is established. That is, all query matrices Q and key matrices K are screened by Kullback-Leibler divergence, and query subsets and key subsets are extracted from Q and K. That is, u query matrices or key matrices with the largest Kullback-Leibler divergence are selected:

[0031]

[0032] Where c is an adjustable hyperparameter, which is used to dynamically adjust the number of query subsets and key subsets, L represents the length of the input sequence; the query subset is defined as Q P ; The unselected key matrix is ​​processed by global averaging and combined with the key subset to generate K P , and then perform attention mechanism calculation:

[0033]

[0034] Where, X DP represents the intermediate matrix calculated by the attention mechanism, d represents the dimension of Q or K; X is calculated by the distillation mechanism. DP Perform sparsity calculation and obtain the output X of the multi-head dual sparse self-attention mechanism D :

[0035]

[0036] Furthermore, in the S302 process, the construction of the temporal feature decoder includes: constructing a linear layer for converting the temporal feature map in the S203 degraded data into a matrix consistent with the decoder output dimension, then using the ReLU activation function to balance the nonlinear expression ability and numerical stability of the temporal feature, and then constructing the Dropout layer and the residual connection layer, and finally constructing the LayerNorm layer to generate the output of the temporal feature decoder.

[0037] Furthermore, in the S303 process, the integration of the multi-head dual sparse self-attention mechanism and the temporal feature decoder to construct the TD-Informer prediction model includes: e The encoder of TD-Informer is composed of a multi-head double sparse self-attention mechanism layer in series; n d The decoder of TD-Informer is constructed by a multi-head dual sparse self-attention mechanism layer and a single-layer multi-head full self-attention mechanism layer. The multi-head full self-attention mechanism is expressed as:

[0038]

[0039] The output of the TD-Informer encoder is combined with the output of the multi-head dual sparse self-attention mechanism layer in the TD-Informer decoder, and input into the multi-head full self-attention mechanism layer in the TD-Informer decoder for processing to obtain a preliminary prediction. The preliminary prediction data is then combined with the output of the temporal feature decoder, and a fully connected layer is added to form the final prediction.

[0040] Furthermore, in the process of S305, the degradation data is input into TD-Informer for iterative training, and the model is verified to determine whether it converges. If it meets the expected requirements, the trained TD-Informer prediction model is obtained after the test passes, including: constructing the input of TD-Informer, assuming T in the jth prediction i and Z i Represent the time characteristics and the observation values ​​of the main influencing characteristics at the i-th moment, Y i represents the output voltage observation value at time i; the sequence consisting of the characteristic observation values ​​and output voltage observation values ​​between historical time 0 and time t can be expressed as:

[0041]

[0042] Represents the input data to the encoder in TD-Informer, where the total sequence length Seq_len=t+1 is a hyperparameter that needs to be set by the user; After time, location and feature encoding, it is input into the encoder of TD-Informer for prediction training;

[0043]

[0044] Represents the input data to the decoder in the TD-Informer network, After encoding time, location and features, they are input into the TD-Informer decoder for prediction training. The total length of the sequence is l token +l pred +1, Tok_len=l token +1 is the hyperparameter that needs to be set, l token +1 represents the length of the historical sequence; in addition, the hyperparameter Pred_len=l needs to be set pred Represents the length of the preset future prediction sequence. The preset future prediction sequence is filled with 0 placeholders in TD-Informer;

[0045] After the TD-Informer prediction model is built, the training set is used to train the TD-Informer prediction model. The validation set is used to determine whether the model has converged. Finally, the test set is used to evaluate the prediction effect of the model.

[0046] Furthermore, in the S403 process, the prediction accuracy of the comprehensive evaluation model is generated, and an evaluation report is generated to guide the maintenance strategy, including: using the root mean square error RMSE, mean absolute percentage error MAPE and mean absolute error MAE as evaluation indicators; if RMSE, MAPE and MAE all meet the expected requirements, the test passes; otherwise, continue training to achieve the target accuracy; the calculation formulas of RMSE, MAPE and MAE are as follows:

[0047]

[0048]

[0049]

[0050] Where, is the predicted value, y i is the actual value, n is the number of samples; the smaller the RMSE, MAPE and MAE values ​​are, the higher the prediction accuracy is.

[0051] The present invention has the following beneficial effects:

[0052] To address the structural limitations of traditional informer networks in PEMFC performance degradation prediction, namely the excessive feature interaction redundancy caused by a single-dimensional sparse self-attention mechanism and the spectral domain feature dilution effect caused by multi-layer attention stacking, this paper innovatively proposes a TD-Informer network prediction method that integrates a dual-sparse self-attention mechanism with a temporal feature decoder. This method significantly reduces feature interaction redundancy by constructing a dual-dimensional constrained sparse attention strategy. It also designs a lightweight MLP temporal feature decoder to efficiently preserve spectral domain features, reducing prediction errors during key degradation stages and significantly improving the accuracy and engineering applicability of PEMFC degradation prediction. The method comprises: conducting PEMFC performance degradation experiments to obtain multi-dimensional actual operating parameters; extracting key characteristic factors using SHAP value analysis; then designing and constructing a TD-Informer prediction model that integrates a multi-head dual-sparse self-attention mechanism with a temporal feature decoder. Finally, by integrating the structural advantages of the dual-sparse strategy and the temporal feature decoder, high-precision prediction of PEMFC degradation states is achieved. This method overcomes the structural limitations of the original Informer model's single-dimensional sparsification strategy when processing PEMFC degradation data, significantly reducing the model's training parameters; the lightweight MLP time series feature decoder completes the spectral domain decomposition and residual fusion of PEMFC global time series features, solving the problem of multi-scale feature dilution in time series data under the Informer multi-layer attention stacking architecture, significantly reducing the prediction error in the key degradation stage, thereby adapting to the development requirements of the PEMFC technology industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 This is a flow chart of a TD-Informer network PEMFC degradation prediction method that integrates dual attention and temporal features of the present invention;

[0054] Figure 2 This is a schematic diagram of the dual sparse self-attention mechanism network of the present invention;

[0055] Figure 3 is a schematic diagram of a timing feature decoder network of the present invention;

[0056] Figure 4 TD-Informer network diagram of the present invention that integrates dual sparse self-attention mechanism and temporal feature decoder;

[0057] Figure 5 This is a schematic diagram of the online prediction process of the PEMFC degradation state of the present invention;

[0058] Figure 6 This is a diagram showing the prediction effect of the Informer prediction model provided by an embodiment of the present invention on a fuel cell experimental dataset;

[0059] Figure 7 This is a diagram showing the prediction effect of the TD-Informer prediction model provided by an embodiment of the present invention on a fuel cell experimental dataset. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0061] Example

[0062] Please refer to Figure 1 The present invention provides a PEMFC degradation prediction method based on a TD-Informer network that integrates dual attention and temporal features. The steps are as follows:

[0063] S1: Data Collection

[0064] S101: Conduct performance degradation experiments on the PEMFC test bench. Record PEMFC state parameters (including voltage, current, current density, air inlet pressure, hydrogen inlet pressure, air inlet temperature, and other state parameters) in real time through the data acquisition module to construct a multi-dimensional parameter degradation dataset.

[0065] S102: Selecting the output voltage as a degradation indicator and taking the other recorded state parameters except the voltage as influencing factors;

[0066] S2: Feature Engineering and Sample Construction

[0067] S201: Sampling the multi-dimensional parameter degradation data set and setting the time sampling interval within a reasonable range;

[0068] S202: Using the SHAP value method, quantify the contribution of each influencing factor to PEMFC voltage degradation, and select the top five influencing factors as the main influencing features;

[0069] S203: extracting time features, voltage, and main impact feature data from the sampled multi-dimensional parameter degradation data set as degradation data, and performing smoothing preprocessing on the degradation data to eliminate instantaneous noise interference;

[0070] S204: Mapping the degradation data to the interval [0, 1] by the maximum-minimum normalization method, and dividing the training set, validation set, and test set in proportion;

[0071] S3: TD-Informer prediction model construction

[0072] S301: Construct a multi-head dual sparse self-attention mechanism;

[0073] S302: Constructing a temporal feature decoder;

[0074] S303: Integrate the multi-head dual sparse self-attention mechanism and the temporal feature decoder to build the TD-Informer prediction model;

[0075] S304: Use the optimizer and loss function to prevent overfitting with early stopping and iteratively optimize the TD-Informer prediction model parameters.

[0076] S305: Input the degradation data as input data into TD-Informer for iterative training, and verify whether the model converges. If it meets the expected requirements, the test passes and the trained TD-Informer prediction model is obtained;

[0077] S4: Online prediction of degradation status

[0078] S401: collecting the time characteristics, voltage and main influencing characteristics of the target PEMFC in real time, and performing standardization processing in accordance with S201, S202, S203 and S204;

[0079] S402: Input the processed data into the trained TD-Informer prediction model to output a voltage degradation prediction sequence for a future preset period;

[0080] S403: Comprehensively evaluate the prediction accuracy of the model, generate an evaluation report, and guide the formulation of maintenance strategies.

[0081] refer to Figure 2 In the S301, the construction of the multi-head dual sparse self-attention mechanism includes: three learnable matrices W Q 、W K and W V Used to map the input data X into feature matrix Q (query matrix), K (key matrix) and V (value matrix):

[0082]

[0083]

[0084]

[0085] Where Q is the query matrix, K is the key matrix, and V is the value matrix. By jointly optimizing the sparsity patterns of the query matrix and the key matrix, a multi-head dual-sparse self-attention mechanism is established. That is, all query matrices Q (number L_Q) and key matrices K (number L_K) are screened using the Kullback-Leibler divergence (KL divergence). The query subset (number Q_top) and key subset (number K_top) are extracted from Q and K. That is, the u query matrices or key matrices with the largest KL divergence are selected (the queries whose attention distribution deviates most from the uniform distribution):

[0086]

[0087] Where c is a constant (i.e., the adjustable hyperparameters Factor1 and Factor2, which are used to dynamically adjust the sparsity of Q_top and K_top, respectively), L represents the length of the input sequence; the query subset is defined as Q P ; The unselected key matrix is ​​processed by global mean pooling and combined with the key subset to generate K P , and then perform attention mechanism calculation:

[0088]

[0089] Where, X DP represents the intermediate matrix calculated by the attention mechanism, d represents the dimension of Q or K; X is calculated by the distillation mechanism. DP Perform sparsity calculation and obtain the output X of the multi-head dual sparse self-attention mechanism D :

[0090]

[0091] refer to Figure 3 In the S302 process, the construction of the temporal feature decoder includes: constructing a linear layer for converting the temporal feature map in the S203 degraded data into a matrix consistent with the decoder output dimension, then using the ReLU activation function to balance the nonlinear expression ability and numerical stability of the temporal feature, and then constructing the Dropout layer and the residual connection layer, and finally constructing the LayerNorm layer to generate the output of the temporal feature decoder.

[0092] refer to Figure 4 In the S303 process, the integration of multi-head dual sparse self-attention mechanism and temporal feature decoder to construct TD-Informer prediction model includes: e The encoder of TD-Informer is composed of a multi-head double sparse self-attention mechanism layer in series; n dThe decoder of TD-Informer is constructed by a multi-head dual sparse self-attention mechanism layer and a single-layer multi-head full self-attention mechanism layer. The multi-head full self-attention mechanism is expressed as:

[0093]

[0094] The output of the TD-Informer encoder is combined with the output of the multi-head dual sparse self-attention mechanism layer in the TD-Informer decoder, and input into the multi-head full self-attention mechanism layer in the TD-Informer decoder for processing to obtain a preliminary prediction, which is then combined with the output of the temporal feature decoder and a fully connected layer is added to form the final prediction.

[0095] In the process of S305, the degradation data is input as input data into the TD-Informer network for iterative training, and the model is verified to determine whether it converges. If it meets the expected requirements, the test is passed and the trained TD-Informer prediction model is obtained, including: constructing the input of TD-Informer, assuming T in the jth prediction i and Z i Represent the time characteristics and the observation values ​​of the main influencing characteristics at the i-th moment, Y i represents the output voltage observation value at time i; the sequence consisting of the characteristic observation values ​​and output voltage observation values ​​between historical time 0 and time t can be expressed as:

[0096]

[0097] Represents the input data to the encoder in the TD-Informer network, where the total sequence length Seq_len=t+1 is a hyperparameter that needs to be set by the user; After time, location and feature encoding, it is input into the encoder of TD-Informer for prediction training;

[0098]

[0099] Represents the input data to the decoder in the TD-Informer network, After encoding time, location and features, they are input into the TD-Informer decoder for prediction training. The total length of the sequence is l token +l pred +1, Tok_len=l token +1 is the hyperparameter that needs to be set, l token +1 represents the length of the historical sequence; in addition, the hyperparameter Pred_len=l needs to be setpred represents the preset future prediction sequence length (filled with placeholder 0 in the Informer network to represent the data to be predicted), After time, location and feature encoding, the data is input into the decoder of TD-Informer for prediction training;

[0100] After the TD-Informer prediction model is built, the training set is used to train the TD-Informer prediction model. The validation set is used to determine whether the model has converged. Finally, the test set is used to evaluate the prediction effect of the model.

[0101] In the S403 process, the prediction accuracy of the comprehensive evaluation model is evaluated, and an evaluation report is generated to guide the formulation of the maintenance strategy, including: using the root mean square error (RMSE), mean absolute percentage error (MAPE), and mean absolute error (MAE) as evaluation indicators; if RMSE, MAPE, and MAE all meet the expected requirements, the test passes; otherwise, training continues to achieve the target accuracy; the calculation formulas for RMSE, MAPE, and MAE are as follows:

[0102]

[0103]

[0104]

[0105] Where, is the predicted value, y i is the actual value, n is the number of samples; the smaller the RMSE, MAPE and MAE values ​​are, the higher the prediction accuracy is.

[0106] refer to Figure 5 , an online prediction flow chart of the PEMFC degradation state of S4 is presented.

[0107] This embodiment of the present invention provides a specific example of a PEMFC degradation prediction method using a TD-Informer network that integrates dual attention and temporal features. This example uses a PEMFC produced by Yuchai Xinlan Technology Co., Ltd. as the experimental platform, and follows step S101 to establish a corresponding experimental platform. This example selects the fuel cell stack and the battery cells at both ends and in the middle of the stack (F1, F82, and F185, respectively) as the research objects. 533 hours of dynamic operating experimental data are collected to train and test the PEMFC degradation dynamic prediction model.

[0108] In order to verify the accuracy of the prediction model in this embodiment, after the model training is completed, the test model sample set is brought into different prediction models for testing and comparative experiments. Figure 6 and Figure 7 The results demonstrate the localized amplification prediction performance of the Informer prediction model and the TD-Informer prediction model proposed in this paper on a stack degradation dataset. Table 1 lists the RMSE, MAE, and MAPE of the prediction results for the Informer, RNN, GRU, LSTM, and TD-Informer prediction models proposed in this paper on different experimental datasets. Analysis shows that the TD-Informer prediction model proposed in this paper has higher prediction accuracy than other models.

[0109] Table 1 RMSE, MAE and MAPE of prediction results of different models

[0110]

[0111] Those skilled in the art will recognize that the technical solutions described in the above embodiments can drive the corresponding hardware devices to execute all or part of the steps through computer programs, and the relevant programs can be stored in non-temporary computer-readable storage media, such as but not limited to disks, optical storage media, read-only memories, or random access memories.

[0112] The above description is only a preferred embodiment of the present invention and is not intended to limit the present method. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A TD-Informer network PEMFC degradation prediction method integrating dual attention and temporal features, characterized by: The following steps are involved: S1: Data Collection S101: Conduct performance degradation experiments on the PEMFC test bench, record PEMFC state parameters in real time through the data acquisition module, and construct a multi-dimensional parameter degradation data set; S102: Selecting the output voltage as a degradation indicator and taking the other recorded state parameters except the voltage as influencing factors; S2: Feature Engineering and Sample Construction S201: Sampling the multi-dimensional parameter degradation data set; S202: Using the SHAP value method, quantify the contribution of each influencing factor to PEMFC voltage degradation, and select the top five influencing factors as the main influencing features; S203: extracting time features, voltage, and main impact feature data from the sampled multi-dimensional parameter degradation data set as degradation data, and performing smoothing preprocessing on the degradation data to eliminate instantaneous noise interference; S204: Mapping the degradation data to the interval [0, 1] by the maximum-minimum normalization method, and dividing the training set, validation set, and test set in proportion; S3: TD-Informer prediction model construction S301: Construct a multi-head dual sparse self-attention mechanism; S302: Constructing a temporal feature decoder; S303: Integrate the multi-head dual sparse self-attention mechanism and the temporal feature decoder to build the TD-Informer prediction model; S304: Using the optimizer and loss function, iteratively optimize the TD-Informer prediction model parameters; S305: Input the degradation data into TD-Informer for iterative training and verify whether the model has converged. If it meets the expected requirements, the trained TD-Informer prediction model is obtained after the test passes. S4: Online prediction of degradation status S401: collecting the time characteristics, voltage and main influencing characteristics of the target PEMFC in real time, and performing standardization processing in accordance with S201, S202, S203 and S204; S402: Input the processed data into the trained TD-Informer prediction model to output a voltage degradation prediction sequence for a future preset period; S403: Comprehensively evaluate the prediction accuracy of the model, generate an evaluation report, and guide maintenance strategies; In the S301 process, the multi-head dual sparse self-attention mechanism is constructed including: three learnable matrices W Q 、W K and W V Used to map the input data X into feature matrices Q, K, and V: , , , Where Q is the query matrix, K is the key matrix, and V is the value matrix. By jointly optimizing the sparsity patterns of the query matrix and the key matrix, a multi-head dual-sparse self-attention mechanism is established. That is, all query matrices Q and key matrices K are screened by Kullback-Leibler divergence, and query subsets and key subsets are extracted from Q and K. That is, u query matrices or key matrices with the largest Kullback-Leibler divergence are selected: , Where c is an adjustable hyperparameter, which is used to dynamically adjust the number of query subsets and key subsets, L represents the length of the input sequence; the query subset is defined as Q P ; The unselected key matrix is ​​processed by global averaging and combined with the key subset to generate K P , and then perform attention mechanism calculation: , Where, X DP represents the intermediate matrix calculated by the attention mechanism, d represents the dimension of Q or K; X is calculated by the distillation mechanism. DP Perform sparsity calculation and obtain the output X of the multi-head dual sparse self-attention mechanism D : 。 2. According to claim 1, a TD-Informer network PEMFC degradation prediction method integrating dual attention and temporal features is characterized in that: In the S302 process, the construction of the temporal feature decoder includes: constructing a linear layer for converting the temporal feature map in the S203 degraded data into a matrix consistent with the decoder output dimension, then using the ReLU activation function to balance the nonlinear expression ability and numerical stability of the temporal feature, then constructing the Dropout layer and the residual connection layer, and finally constructing the LayerNorm layer to generate the output of the temporal feature decoder.

3. According to the TD-Informer network PEMFC degradation prediction method integrating dual attention and temporal features as described in claim 2, it is characterized in that: In the S303 process, the integration of multi-head dual sparse self-attention mechanism and temporal feature decoder to construct TD-Informer prediction model includes: e The encoder of TD-Informer is composed of a multi-head double sparse self-attention mechanism layer in series; n d The decoder of TD-Informer is constructed by a multi-head dual sparse self-attention mechanism layer and a single-layer multi-head full self-attention mechanism layer. The multi-head full self-attention mechanism is expressed as: , The output of the TD-Informer encoder is combined with the output of the multi-head dual sparse self-attention mechanism layer in the TD-Informer decoder, and input into the multi-head full self-attention mechanism layer in the TD-Informer decoder for processing to obtain a preliminary prediction. The preliminary prediction data is then combined with the output of the temporal feature decoder, and a fully connected layer is added to form the final prediction.

4. The TD-Informer network PEMFC degradation prediction method integrating dual attention and temporal features according to claim 1 is characterized in that: In the process of S305, the degradation data is input into TD-Informer for iterative training, and the model is verified to determine whether it converges. If it meets the expected requirements, the trained TD-Informer prediction model is obtained after the test passes, including: constructing the input of TD-Informer, assuming T in the jth prediction i and Z i Represent the time characteristics and the observation values ​​of the main influencing characteristics at the i-th moment, Y i represents the output voltage observation value at time i; the sequence consisting of the characteristic observation values ​​and output voltage observation values ​​between historical time 0 and time t can be expressed as: , Represents the input data to the encoder in TD-Informer, where the total sequence length Seq_len=t+1 is a hyperparameter that needs to be set by the user After time, location and feature encoding, it is input into the encoder of TD-Informer for prediction training; , Represents the input data to the decoder in the TD-Informer network, After encoding time, location and features, they are input into the TD-Informer decoder for prediction training. The total length of the sequence is l token +l pred +1, Tok_len=l token +1 is the hyperparameter that needs to be set, l token +1 represents the length of the historical sequence; in addition, the hyperparameter Pred_len=l needs to be set pred Represents the length of the preset future prediction sequence. The preset future prediction sequence is filled with 0 placeholders in TD-Informer; After the TD-Informer prediction model is built, the training set is used to train the TD-Informer prediction model. The validation set is used to determine whether the model has converged. Finally, the test set is used to evaluate the prediction effect of the model.

5. The TD-Informer network PEMFC degradation prediction method integrating dual attention and temporal features according to claim 1 is characterized in that: In the S403 process, the prediction accuracy of the comprehensive evaluation model is generated, and an evaluation report is generated to guide the maintenance strategy, including: using the root mean square error (RMSE), mean absolute percentage error (MAPE), and mean absolute error (MAE) as evaluation indicators; if RMSE, MAPE, and MAE all meet the expected requirements, the test passes; otherwise, continue training to achieve the target accuracy; the calculation formulas for RMSE, MAPE, and MAE are as follows: , , , Where, is the predicted value, y i is the actual value, n is the number of samples; the smaller the RMSE, MAPE and MAE values ​​are, the higher the prediction accuracy is.

Citation Information

Patent Citations

  • Long time series data prediction method based on bidirectional sparse mechanism Transform

    CN116541435A

  • PEMFC degradation prediction method and system based on spatial dynamic non-stationary reconstruction attention

    CN120048947A