A method for predicting the remaining useful life of aero-engines based on diffusion models and spatio-temporal attention mechanisms

By combining the diffusion model and the time attention mechanism, the problems of data scarcity and noise impact of aircraft engines are solved, the accuracy and generalization capabilities of RUL prediction are improved, and the health management and maintenance decisions of aircraft engines are supported.

CN119782867BActive Publication Date: 2025-07-29DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411984723.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-07-29
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The existing methods for the remaining service life prediction of aircraft engines are limited by data scarcity and noise, making it difficult to accurately capture time characteristics, resulting in insufficient prediction accuracy and generalization capabilities.

Method used

Combining the diffusion model and the temporal attention mechanism, data diversity is enhanced through the diffusion process, and time attention units and multi-headed attention layers are introduced into the neural network to improve the recognition ability of time features and build an end-to-end RUL prediction model.

Benefits of technology

Improve the accuracy of aircraft engine residual service life prediction and generalization capabilities, enhance the ability to capture performance degradation processes, and provide reliable support for health management and maintenance decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782867B_ABST
    Figure CN119782867B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of aero-engine health management, and discloses a method for predicting the remaining useful life of an aero-engine based on a diffusion model and a spatio-temporal attention mechanism. The present invention proposes a novel learning framework for RUL prediction of aero-engines. This framework integrates a diffusion process and an attention-based time module to achieve an end-to-end accurate estimation from sensor signals to RUL. Aiming at the problem that the scarcity of aero-engine data limits the application of deep learning networks in improving the accuracy of RUL prediction, this method uses a diffusion process to enrich data diversity and improve data quality. In order to improve the accuracy of RUL prediction, an attention mechanism is used to extract intra-sample and inter-sample temporal features, effectively utilizing temporal information, thereby enhancing the ability to capture the performance degradation process of aero-engines. The RUL prediction method of the present invention provides more reliable support for the health management and maintenance decision-making of aero-engines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of aero-engine health management, and is a method for predicting the remaining useful life (RUL) of an aero-engine, involving data augmentation and time series processing, and can be used in technical fields such as aerospace and transportation. Background Art

[0002] In modern industry, the Prognostics and Health Management (PHM) system is crucial for ensuring the reliability of industrial activities, and it includes key links such as anomaly detection, fault diagnosis, and RUL estimation. As a key component of an aircraft, the health state of an aero-engine is directly related to flight safety and reliability. The RUL prediction technology can evaluate the operating condition of an aero-engine and effectively prevent faults. For airlines and maintenance teams, it is an important tool that can help them formulate scientific maintenance plans, timely detect and solve potential fault problems, thereby ensuring the safety and availability of the engine.

[0003] The faults of an aero-engine usually do not occur suddenly, but are caused by the gradual wear and performance degradation of internal components during long-term operation. Key components inside the engine, such as turbine blades, bearings, and seals, operate under extreme conditions such as high temperature, high pressure, and high speed, and are prone to wear and loss, thereby affecting the performance and reliability of the engine. In order to effectively monitor and manage the degradation process of the engine, airlines usually adopt various means for regular inspections and maintenance. Currently, the aero-engine RUL prediction methods are mainly divided into model-based methods and data-driven methods. The model-based methods rely on a deep understanding of the aero-engine to establish a degradation mechanism model, such as the Kalman filter and the particle filter. However, due to the complexity of the aero-engine, it is difficult to establish an accurate physical model. With the progress of modern instruments and measurement technologies, a large amount of monitoring data can be obtained from the sensors of the aero-engine, making the data-driven methods become a research hotspot. These methods attempt to directly establish a non-linear mapping relationship between historical monitoring data and the engine health state. In recent years, artificial intelligence technologies, especially deep learning technologies, have achieved remarkable results in multiple fields. Neural networks can directly model the sensor monitoring data and perform RUL prediction in a data-driven manner without relying on complex physical models and expert knowledge.

[0004] Deep learning technology has shown great potential in the field of aero-engine RUL prediction, mainly including methods based on RNN, CNN, and Transformer, etc. For example, RNN and its variants LSTM (Ma M, Mao Z. Deep Convolution-based LSTM Network for Remaining Useful Life Prediction[J], IEEE Transactions on Industrial Informatics, 2020, PP(99):1-1. DOI:10.1109 / TII.2020.2991796.) and GRU (R. Gong, J. Li, and C. Wang, “Remaining useful life prediction based on multisensor fusion and attention TCN-BiGRU model,” IEEE Sensors J., vol. 22, no. 21, pp. 21101–21110, Nov. 2022.) perform excellently in capturing the complex non-linear relationship between input and output and short-term correlations in time series, becoming the main framework for predicting engine RUL. In addition, based on the fact that RNN is good at capturing time correlations for sequence learning while CNN is good at extracting local features, combining the two is also a common solution for RUL prediction. In addition, improved CNN and multi-scale CNN can enhance the model's learning ability for complex features (Ren L, Sun Y, Wang H, et al. Prediction of Bearing Remaining Useful Life With Deep Convolution Neural Network[J]. IEEE Access, 2018:13041-13049. DOI:10.1109 / ACCESS.2018.2804930.). Methods based on Transformer can also achieve more accurate RUL prediction by focusing on more critical information in the monitoring data (Zhang Z, Song W, Li Q. Dual Aspect Self-Attention based on Transformer for Remaining Useful Life Prediction[J]. 2021. DOI:10.48550 / arXiv.2106.15842.).

[0005] Although existing methods have achieved certain results in mining sequence data features for predicting the remaining useful life (RUL) of aero-engines, there are still challenges that need to be further improved. For example, existing methods are often affected by the time-dependence of sequences, which are easily masked by redundant data or noise, making it difficult to effectively extract the components that truly play an important role in prediction. This patent discovers a time attention mechanism that can specifically focus on the time features between samples to obtain the time correlation in the data, and can suppress the influence of redundant data or noise. In addition, due to limitations in equipment and testing costs, it is very difficult to obtain experimental samples throughout the entire life cycle of an aero-engine, which directly limits the number of training samples that can be collected. For deep learning technologies that rely on a large amount of data support, this limitation seriously affects the prediction performance of the RUL model. To alleviate this problem, the present invention enhances data diversity through a symmetric diffusion model, thereby increasing the number of training samples, improving the generalization ability of the model to prevent overfitting, and thus enhancing the prediction performance of the model. In recent years, diffusion generative models have shone brightly in the field of vision. They are the most cutting-edge neural network model architectures in the field of generating images. DDIM (Song J, Meng C, Ermon S. Denoising Diffusion Implicit Models [C] / / International Conference on Learning Representations. 2021.) and DDPM (Nichol A, Dhariwal P. Improved Denoising Diffusion Probabilistic Models [J]. 2021. DOI: 10.48550 / arXiv.2102.09672.), which are improvements based on diffusion models, have gradually been applied to research directions such as time series (Rasul K, Seward C, Schuster I, et al. Autoregressive Denoising Diffusion Models for Multivariate Probabilistic Time Series Forecasting [J]. 2021. DOI: 10.48550 / arXiv.2101.12072.). However, there has been no research on applying diffusion generative models to the RUL prediction task of aero-engines. While introducing a diffusion generative model, the present invention improves the inverse process of the model, only uses the forward diffusion process to augment data, and directly replaces the inverse process of the diffusion model with an RUL prediction network, reducing the model training cost and obtaining a new prediction method for aero-engines. This model can be extended and applied to related technical fields such as aerospace and transportation.

[0006] In summary, the present invention proposes a method for predicting the remaining useful life (RUL) of an aero-engine by combining a symmetric diffusion process and an attention mechanism. This method adopts the integration of a temporal attention unit, enhancing its ability to identify features at different time scales, enabling the model to accurately identify the process of aero-engine performance degradation, which is crucial for predicting the complex degradation process of an aero-engine. Additionally, by applying the diffusion process to the training stage of model prediction, noise is added to the input sequence and the target sequence to achieve sample amplification, increasing the diversity of the dataset, reducing the empirical risk in model training, improving the generalization ability, and preventing overfitting. The combination of these techniques not only improves the accuracy of RUL prediction but also enhances the generalization and robustness of the model, providing effective decision support for the health management and maintenance of aero-engines. Summary of the Invention

[0007] The present invention proposes a method for predicting the remaining useful life (RUL) of an aero-engine, which combines a diffusion process and a temporal attention mechanism. By augmenting the original data through the diffusion process, the present invention not only maintains the randomness of the data but also reduces the training burden of the inference network. In addition, the introduction of a temporal attention unit further enhances the network's comprehensive ability to capture temporal features and dependencies between samples, thereby improving the accuracy of prediction.

[0008] Technical Solution of the Present Invention:

[0009] A method for predicting the remaining useful life of an aero-engine based on a diffusion model and a spatio-temporal attention mechanism, comprising the following steps:

[0010] Step 1: Process the data acquired by sensors to construct training samples, where the training samples include a training set and a test set;

[0011] 1.1) Use the K-means clustering algorithm to perform conditional classification on the data acquired by sensors, and perform z-score standardization processing on the data of each category. The calculation formula is:

[0012]

[0013] where i and j respectively represent the i-th sensor and the j-th conditional classification, and ν j and τ j respectively represent the average value and standard deviation of the i-th sensor under the j-th conditional classification, and x ij and α ij respectively represent the data before and after standardization processing;

[0014] 1.2) Write a function to change the shape of the input data, which includes the training set and the data for training in the test set, and generate the input data with the desired shape according to the window length and offset; use a linear degradation model or a piecewise degradation model to predict the RUL, and set the parameter early_rul; obtain the shape of the specified array according to different requirements;

[0015] Step 2: Perform data augmentation according to the diffusion model. The data includes the training data set and the label data set in the training set;

[0016] 2.1) Use the forward process of the diffusion model to perform synchronous data augmentation on the training data set and the label data set; for the given sequence distribution x0 ∼ q(x0), the approximate posterior q(x1,…,x T |x0) is expressed as:

[0017]

[0018] where the transition kernel q(x t |x t-1 ) is defined as:

[0019]

[0020] where t and T represent time steps, and q(x t |x t-1 ) represents the conditional probability distribution of x t-1 given x t ; represents the normal distribution; x t and x t-1 are the states at consecutive time steps in the diffusion process; is the scaling factor of the mean; β t is the noise variance at time step t, which is an increasing variance schedule used to control the level of noise added; I is the identity matrix;

[0021] 2.2) By sampling the Gaussian vector, we get:

[0022]

[0023] where α t = 1 - β t , ∈ x is the sampled Gaussian noise of the training data set;

[0024] 2.3) To reduce the uncertainty brought by the generative model, the diffusion process adds the same type of noise to both the training dataset and the label dataset simultaneously, so that the sequence of the training dataset remains consistent and intact with the sequence of the label dataset, generating the label formula:

[0025]

[0026] Among them, β′ t = ωβ t , where ω is a fixed proportional parameter, ∈ y is the Gaussian noise sampled from the label dataset;

[0027] Step 3: Based on the training samples constructed in Steps 1 and 2, construct a neural network model as the remaining useful life prediction model of the engine; the neural network model structure includes a time series information encoding layer, a spatio-temporal convolutional layer, a multi-head attention layer, and an output layer;

[0028] 3.1) To solve the problem that the Transformer self-attention mechanism cannot directly capture the temporal sequence of the input data, after the input data enters the time series information encoding layer, sine and cosine position encodings are added through the position encoding layer to provide the relative position information of the input data, thereby enhancing the network structure model's ability to capture temporal features. The corresponding calculation formula is:

[0029]

[0030] Among them, P(t, 2j) represents the value of the t-th time step and the 2j-th sensor in the position encoding matrix; d is the dimension of the position encoding, that is, the number of elements in each position encoding vector;

[0031] 3.2) To obtain the spatio-temporal feature information of the input data, a spatio-temporal convolutional layer is added:

[0032] F st = σ(W st ·(X * T s ))

[0033] Among them, F st represents the spatio-temporal feature, W st is the weight matrix, X is the sensor data matrix, T s is the time feature matrix, σ is the activation function, and * represents the spatio-temporal convolution operation;

[0034] 3.3) The attention mechanism is the core component of the Transformer model. The upper-layer input sensor data matrix X is combined with three different weight matrices W q , W k and W vMultiply to obtain the query vector Q, the key vector K, and the content vector V. The corresponding formulas are as follows:

[0035] Q = XW Q

[0036] K = XW K

[0037] V = XW V

[0038] 3.4) By calculating the dot product of the query vector Q and the key vector K, an association matrix is obtained. After activation by the Softmax function, the weights corresponding to each sensor data in the upper layer are obtained; multiply the weights obtained by the Softmax function by the spatio-temporal feature information F st , and the content vector V to obtain the fused spatio-temporal attention feature output. The calculation formula is as follows:

[0039]

[0040] where F att is the fused spatio-temporal attention feature, d k is the dimension of the key vector K, is used as a scaling factor to alleviate the vanishing gradient problem, and ⊙ represents element-wise multiplication;

[0041] 3.5) The multi-head attention layer adopts the multi-head attention mechanism, that is, multiple groups of Q, K, and V are calculated, and then the outputs of multiple groups of attention are concatenated as the final output to balance the possible biases generated by the same attention mechanism, thereby improving the effect of the neural network model. The corresponding calculation formula is as follows:

[0042] F fused = Concat(A1, A2,..., A H )W O

[0043] where F fused represents the fused multi-head spatio-temporal attention feature; H is the number of heads, that is, the number of parallel self-attention layers in the neural network model; A i is the self-attention output of the i-th head; A i = Attention(Q i , K i , V i ) is the self-attention calculation of the i-th head, which is the same as the calculation of a single self-attention layer; W O is the weight matrix used to combine the outputs of all heads;

[0044] The calculation process of multi-head attention is as follows: (1) For each head i, calculate the query Q i , the key K i , and the value Vi Obtained by multiplying the input sensor data matrix X by the corresponding weight matrix; (2) Calculate the self-attention A for each head i ; (3) Concatenate the outputs A1, A2, …, A of all heads H and then multiply by the weight matrix W O to obtain the final multi-head attention output;

[0045] 3.6) The predicted RUL is obtained through the output layer. The output layer processes the normalized attention weight information of the attention module, maps the input to the hidden vector space, and predicts the final RUL. The expression is:

[0046]

[0047] where is the predicted remaining useful life, W out and b out are the weight and bias of the output layer respectively;

[0048] 3.7) The definition of the loss function is as follows:

[0049]

[0050] where L is the loss function, y i and are the true value and the predicted value respectively, λ and α are regularization parameters, and ∥W st ∥ 2 is the L2 regularization term of the weight matrix W st ;

[0051] Step Four: According to the dataset obtained in Step One and Step Two and the neural network model obtained in Step Three, conduct training and testing to obtain the predicted rul value. Use the root mean square error loss function, the aviation fault diagnosis scoring function, and the early stopping method to evaluate the obtained predicted rul value and evaluate the life prediction effect. The calculation formulas are as follows:

[0052]

[0053] where y i is the true value, is the model predicted value.

[0054] Advantages of the present invention: The present invention proposes a novel learning framework specifically for the remaining useful life (RUL) prediction of aero-engines. This framework integrates a diffusion process and an attention-based temporal module to achieve an accurate end-to-end estimation from sensor signals to RUL. Aiming at the scarcity problem of aero-engine data, which often limits the application of deep learning networks in improving the accuracy of RUL prediction, this method uses a diffusion process to enrich data diversity and improve data quality. To improve the accuracy of RUL prediction, an attention mechanism is used to extract intra-sample and inter-sample temporal features, effectively utilizing temporal information, thereby enhancing the ability to capture the performance degradation process of aero-engines. Through the application of these technologies, the RUL prediction method of the present invention can provide more reliable support for the health management and maintenance decision-making of aero-engines. Description of the Drawings

[0055] Figure 1 It is the overall flowchart of the remaining useful life prediction algorithm of the present invention.

[0056] Figure 2 It is the framework diagram of the neural network model.

[0057] Figure 3 It is the graph of RUL prediction results. Detailed Embodiments

[0058] The following further illustrates the detailed embodiments of the present invention in conjunction with the drawings and technical solutions.

[0059] Experimental verification is carried out by using the aero-engine dataset C-MAPSS provided by NASA.

[0060] Step 1: Process the data acquired by the sensors to construct training samples, where the training samples include a training set and a test set;

[0061] 1.1) Use the K-means clustering algorithm to perform conditional classification on the data acquired by the sensors, and perform z-score normalization processing on the data of each category. The calculation formula is:

[0062]

[0063] where i and j respectively represent the i-th sensor and the j-th conditional classification, v j and τ j respectively represent the average value and standard deviation of the i-th sensor under the j-th conditional classification, x ij and α ij respectively represent the data before and after normalization processing;

[0064] 1.2) Write a function to change the shape of the input data, which includes the training set and the data for training in the test set. Generate the input data with the desired shape according to the window length and offset; use a linear degradation model or a piecewise degradation model to predict the RUL, and set the parameter early_rul; obtain the shape of a specified array according to different requirements. In the experiment, the shape of the initial training data is (20631×26), and after data preprocessing, the data shape becomes (20631×14).

[0065] Step 2: Perform data augmentation according to the diffusion model. The data includes the training data set and the label data set in the training set;

[0066] 2.1) Use the forward process of the diffusion model to perform synchronous data augmentation on the training data set and the label data set; for a given sequence distribution \(x_0 \sim q(x_0)\), the approximate posterior \(q(x_1,\ldots,x_T|x_0)\) of the diffusion process is expressed as:

[0067]

[0068] where the transition kernel \(q(x_t|x_{t - 1})\) of the Markov process is defined as: t \(|x_{t - 1})\) t-1 is defined as:

[0069]

[0070] where \(t\) and \(T\) represent time steps, \(q(x_t|x_{t - 1})\) represents the conditional probability distribution of \(x_t\) given \(x_{t - 1}\); t \(|x_{t - 1})\) t-1 ; represents the normal distribution; \(x_t\) and \(x_{t - 1}\) are the states at consecutive time steps in the diffusion process; t-1 \(x_{t - 1}\) t ; is the scaling factor of the mean; \(\beta_t\) is the noise variance at time step \(t\), which is an increasing variance schedule used to control the level of noise added; \(I\) is the identity matrix; t \(x_t\) t-1 and \(x_{t - 1}\) are the states at consecutive time steps in the diffusion process; is the scaling factor of the mean; \(\beta_t\) t is the noise variance at time step \(t\), which is an increasing variance schedule used to control the level of noise added; \(I\) is the identity matrix;

[0071] 2.2) By sampling Gaussian vectors, we get:

[0072]

[0073] where \(\alpha_t = 1 - \beta_t\), t \(= 1 - \beta_t\) t , \(\epsilon\) x \(\in N(0, I)\) is the sampled Gaussian noise of the training data set;

[0074] 2.3) To reduce the uncertainty brought by the generative model, the diffusion process adds the same type of noise to both the training dataset and the label dataset simultaneously, so that the sequence of the training dataset remains consistent and intact with that of the label dataset, generating the label formula:

[0075]

[0076] Among them, β′ t = ωβ t , where ω is a fixed proportional parameter, ∈ y is the Gaussian noise sampled from the label dataset;

[0077] Step 3: Based on the training samples constructed in Steps 1 and 2, construct a neural network model as the engine remaining life prediction model; the structure of the neural network model includes a time series information encoding layer, a spatio-temporal convolutional layer, a multi-head attention layer, and an output layer;

[0078] 3.1) To solve the problem that the Transformer self-attention mechanism cannot directly capture the temporal characteristics of the input data, after the input data enters the time series information encoding layer, sine and cosine position encodings are added through the position encoding layer to provide the relative position information of the input data, thereby enhancing the network structure model's ability to capture temporal features. The corresponding calculation formula is:

[0079]

[0080] Among them, P(t, 2j) represents the value of the t-th time step and the 2j-th sensor in the position encoding matrix; d is the dimension of the position encoding, that is, the number of elements in each position encoding vector;

[0081] 3.2) To obtain the spatio-temporal feature information of the input data, add a spatio-temporal convolutional layer:

[0082] F st = σ(W st ·(X * T s ))

[0083] Among them, F st represents the spatio-temporal feature, W st is the weight matrix, X is the sensor data matrix, T s is the time feature matrix, σ is the activation function, and * represents the spatio-temporal convolution operation;

[0084] 3.3) The attention mechanism is the core component of the Transformer model. The upper-layer input sensor data matrix X is combined with three different weight matrices W q , W k and W vMultiply to obtain the query vector Q, the key vector K, and the content vector V. The corresponding formulas are as follows:

[0085] Q = XW Q

[0086] K = XW K

[0087] V = XW V

[0088] 3.4) By calculating the dot product of the query vector Q and the key vector K, an association matrix is obtained. After activation by the Softmax function, the weights corresponding to each sensor data in the upper layer are obtained; Multiply the weights obtained by the Softmax function with the spatio-temporal feature information F st , and the content vector V to obtain the fused spatio-temporal attention feature output. The calculation formula is as follows:

[0089]

[0090] where F att is the fused spatio-temporal attention feature, d k is the dimension of the key vector K, is used as a scaling factor to alleviate the vanishing gradient problem, and ⊙ represents element-wise multiplication;

[0091] 3.5) The multi-head attention layer adopts the multi-head attention mechanism, that is, multiple groups of Q, K, and V are calculated, and then the outputs of multiple groups of attention are concatenated as the final output to balance the possible biases generated by the same attention mechanism, thereby improving the performance of the neural network model. The corresponding calculation formula is as follows:

[0092] F fused = Concat(A1, A2, …, A H )W O

[0093] where F fused represents the fused multi-head spatio-temporal attention feature; H is the number of heads, that is, the number of parallel self-attention layers in the neural network model; A i is the self-attention output of the i-th head; A i = Attention(Q i , K i , V i ) is the self-attention calculation of the i-th head, which is the same as the calculation of a single self-attention layer; W O is the weight matrix used to combine the outputs of all heads;

[0094] The calculation process of multi-head attention is as follows: (1) For each head i, calculate the query Q i , the key K i , and the value Vi Obtained by multiplying the input sensor data matrix X by the corresponding weight matrix; (2) Calculate the self-attention A for each head i ; (3) Concatenate the outputs A1, A2, …, A H of all heads, and then multiply by the weight matrix W O to obtain the final multi-head attention output;

[0095] 3.6) The final result is the RUL prediction obtained through the output layer. The output layer processes the normalized attention weight information of the attention module, maps the input to the hidden vector space, and predicts the final RUL. The expression is:

[0096]

[0097] where, is the predicted remaining useful life, W out and b out are the weight and bias of the output layer respectively;

[0098] 3.7) The definition of the loss function is as follows:

[0099]

[0100] where L is the loss function, y i and are the true value and the predicted value respectively, λ and α are regularization parameters, and ∥W st ∥ 2 is the L2 regularization term of the weight matrix W st ;

[0101] Step Four: According to the dataset obtained in Step One and Step Two and the neural network model obtained in Step Three, train and test to obtain the predicted rul value. Use the root mean square error loss function, the aviation fault diagnosis scoring function, and the early stopping method to evaluate the obtained predicted rul value and evaluate the life prediction effect. The calculation formulas are as follows respectively:

[0102]

[0103] where, y i is the true value, is the model predicted value.

[0104] Finally, the prediction result is as Figure 3 , where True RUL represents the true RUL, and Pred RUL is the prediction result graph of the model.

Claims

1. A method for predicting the remaining useful life of an aeroengine based on a diffusion model and a spatio-temporal attention mechanism, characterized in that, It includes the following steps: Step 1: Process the data obtained by the sensor to construct training samples, where the training samples include a training set and a test set; 1.1) Use the K-means clustering algorithm to perform conditional classification on the data obtained by the sensor, and perform z-score normalization processing on the data of each category. The calculation formula is: wherein, i and j respectively represent the i-th sensor and the j-th condition classification, v j and τ j respectively represent the average value and the standard deviation of the i-th sensor under the j-th condition classification, x ij and α ij respectively represent the data before and after the standardization process; 1.2) Write a function to change the shape of the input data. The input data includes the training data of the training set and the test set. Generate input data with the desired shape according to the window length and offset; use a linear degradation model or a piecewise degradation model to predict the RUL, and set the parameter early_rul; obtain the shape of the specified array according to different requirements; Step 2: Perform data augmentation according to the diffusion model. The data includes the training data set and the label data set in the training set; 2.1) Synchronous data augmentation of the training dataset and the label dataset using the forward process of the diffusion model; for a given sequence distribution x0 ∼ q(x0), the approximate posterior q(x1, …, x T |x0) is expressed as: where the transition kernel \(q(x t |x t-1 ) is defined as: where t and T represent time steps, q(x t ∣x t-1 ) represents the conditional probability distribution of x given x t-1 ; t represents the normal distribution; x and x t and x t-1 are the states at consecutive time steps in the diffusion process; is the scaling factor of the mean; β t is the noise variance at time step t, which is an increasing variance schedule used to control the level of noise added; I is the identity matrix; 2.2) By sampling Gaussian vectors, we get: where α t = 1 - β t , ∈ x is the sampled Gaussian noise of the training dataset; 2.3) To reduce the uncertainty brought by the generative model, the same type of noise is added to both the training data set and the label data set during the diffusion process, so that the sequence of the training data set is consistent and complete with the sequence of the label data set, and the label formula is generated: Among them, β′ t = ωβ t , where ω is a fixed proportional parameter, ∈ y is the Gaussian noise sampled from the labeled data set; Step 3: Based on the training samples constructed in Steps 1 and 2, construct a neural network model as the engine remaining useful life prediction model; the neural network model structure includes a time series information encoding layer, a spatio-temporal convolutional layer, a multi-head attention layer, and an output layer; 3.1) To solve the problem that the Transformer self-attention mechanism cannot directly capture the temporal characteristics of the input data, after the input data enters the time series information encoding layer, sine position encoding and cosine position encoding are added through the position encoding layer to provide the relative position information of the input data, thereby enhancing the network structure model's ability to capture temporal features. The corresponding calculation formula is: Among them, P(t, 2j) represents the value of the t-th time step and the 2j-th sensor in the position encoding matrix; d is the dimension of the position encoding, that is, the number of elements in each position encoding vector; 3.2) To obtain the spatio-temporal feature information of the input data, add a spatio-temporal convolutional layer: F st = σ(W st ·(X * T s )) Among them, F st represents spatio-temporal features, W st is the weight matrix, X is the sensor data matrix, T s is the time feature matrix, σ is the activation function, and * represents the spatio-temporal convolution operation; 3.3) The attention mechanism is the core component of the Transformer model, multiplying the upper-layer input sensor data matrix X by three different weight matrices W q , W k and W v to obtain the query vector Q, the key vector K, and the content vector V. The corresponding formula is as follows: Q = XW Q K = XW K V = XW V 3.4) By calculating the dot product of the query vector Q and the key vector K, an association matrix is obtained. After activation by the Softmax function, the weights corresponding to each sensor data in the upper layer are obtained; Multiply the weights obtained by the Softmax function by the spatio-temporal feature information F st and the content vector V to obtain the fused spatio-temporal attention feature output. The calculation formula is: Among them, F att is the fused spatio-temporal attention feature, d k is the dimension of the key vector K, serves as a scaling factor to alleviate the problem of vanishing gradients, and ⊙ represents element-wise multiplication; 3.5) The multi-head attention layer adopts the multi-head attention mechanism, that is, multiple groups of Q, K, and V are calculated and then the outputs of multiple groups of attention are concatenated as the final output to balance the possible biases generated by the same attention mechanism, thereby improving the effect of the neural network model. The corresponding calculation formula is: F fused = Concat(A1, A2, …, A H )W O Among them, F fused represents the fused multi-head spatio-temporal attention feature; H is the number of heads, that is, the number of self-attention layers in parallel in the neural network model; A i is the output of the fused spatio-temporal attention feature calculated in step 3.4); W O is the weight matrix used to combine the outputs of all heads; The calculation process of multi-head attention is as follows: (1) For each head i, calculate the query Q i , the key K i , and the value V i by multiplying the input sensor data matrix X with the corresponding weight matrix; (2) Calculate the self-attention A i for each head; (3) Concatenate the outputs A1, A2, …, A H of all heads, and then multiply by the weight matrix W O to obtain the final multi-head attention output; 3.6) The final result is the RUL prediction obtained through the output layer. The output layer processes the normalized attention weight information of the attention module, maps the input to the hidden vector space, and predicts the final RUL. The expression is: Among them, is the predicted remaining useful life, W out and b out are the weight and bias of the output layer, respectively; 3.7) The definition of the loss function is as follows: Among them, L is the loss function, y i and are the true value and the predicted value respectively, λ and α are regularization parameters, ||W st || 2 is the L2 regularization term of the weight matrix W st ; Step 4: According to the data set obtained in Steps 1 and 2 and the neural network model obtained in Step 3, perform training and testing to obtain the predicted rul value. Use the root mean square error loss function, the aviation fault diagnosis scoring function, and the early stopping method to evaluate the obtained predicted rul value and evaluate the life prediction effect. The calculation formulas are as follows: Among them, y i is the true value, is the model predicted value.

Citation Information

Patent Citations

  • Aero-engine rolling bearing fault diagnosis method based on self-supervised learning

    CN117592543A

  • Multi-scale hybrid attention mechanism modeling method for predicting remaining useful life of aero engine

    WO2024087128A1