A Method for Predicting the Remaining Useful Life of Aero-Engines Based on Data Augmentation
Through data augmentation and multi-path feature fusion prediction model, the problems of insufficient data and poor generalization performance in the aircraft engine residual life prediction are solved, and high-precision life prediction is achieved.
Patent Information
- Application Number
- CN202210188342.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-02-28
AI Technical Summary
Among the existing aircraft engine residual life prediction methods, it is difficult to establish an accurate degradation model based on prediction methods based on physical models. The data-driven method has problems such as insufficient data and poor generalization performance, resulting in low prediction accuracy.
The data set is expanded by using data augmentation algorithm to build a multi-path feature fusion prediction model, and the spatial and timing features are extracted using CNN, GRU and LSTM networks, and fusion prediction is performed through the full connection layer.
It improves the accuracy and accuracy of aircraft engine residual life prediction, meets the demand for data volume of deep learning, reduces the computational complexity and overfitting phenomenon, and improves the prediction performance of the model.
Smart Images

Figure CN114547986B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the remaining useful life of an aero-engine, and in particular to a method for predicting the remaining useful life of an aero-engine based on data augmentation and a multi-path feature fusion model. Background Art
[0002] The turbofan engine, as the "heart" of various aircraft systems in the aviation field, its performance changes can directly affect the safe operation of the aircraft. However, due to the daily operation of the aero-turbofan engine in an environment of high temperature, high pressure and high vibration speed, it is very easy to have accidents and cause failures as the working time increases. Therefore, predicting the remaining useful life (RUL) of the aero-turbofan engine and changing regular maintenance to proactive maintenance are of great significance for reducing flight safety accidents and saving equipment maintenance costs.
[0003] Currently, the methods for predicting the remaining useful life can be roughly divided into two categories, namely physical model-based prediction and data-driven prediction. Due to the increasing complexity of mechanical systems and the high coupling between components, it is difficult to establish an accurate degradation model based on physical model prediction methods, and their flexibility and portability are poor. The data-driven prediction method uses technologies such as signal processing to analyze the collected equipment degradation data and fault data, and extracts the features reflecting system degradation and faults to perform RUL prediction. In the data-driven method, machine learning and deep learning have been widely used. Machine learning models have simple structures and poor ability to extract deep features of complex non-linear multi-dimensional samples, resulting in low prediction accuracy. In recent years, deep learning methods have gradually emerged in the prediction field. The convolutional neural network (CNN) has strong learning ability and can well extract the spatial features of data, but convolution and pooling operations will cause data loss. The recurrent neural network (RNN) is a network with unique advantages in processing time series data, but RNN will have problems of gradient explosion or gradient disappearance in long-term prediction. The long short-term memory network (LSTM) alleviates the extreme gradient problem of RNN through gate structures. The gated recurrent unit (GRU) reduces the gate structure based on LSTM and improves the calculation efficiency. Most of the existing deep learning models are constructed based on a single model, unable to mine and comprehensively consider multi-faceted features, and have poor generalization performance. In addition, in terms of degradation data, the existing performance degradation data sets are few and cannot meet the needs of deep learning models for a large amount of training data, resulting in low accuracy of the existing deep learning life prediction models. Summary of the Invention
[0004] In view of the deficiencies in the background art, the present invention provides a method for predicting the remaining useful life of an aero-engine based on data augmentation.
[0005] The present invention includes the following steps:
[0006] Step 1: Use a data augmentation algorithm to perform data augmentation and normalization on the acquired data.
[0007] Step 2: Construct a multi-path feature fusion prediction model, and input the data processed in Step 1 into the model to extract different features in the data;
[0008] Step 3: Fuse the features extracted in Step 2 and input them into the fully connected layer for final remaining useful life prediction;
[0009] Step 4: Use two evaluation methods to evaluate the model, verify the effectiveness of data augmentation, and make comparisons with other models.
[0010] Further, in Step 1:
[0011] Step 1.1: Use a data augmentation algorithm to perform data augmentation on the acquired data to meet the requirements of deep learning for a large amount of data. This data augmentation algorithm uses the complete training trajectory of a certain engine to generate partial training trajectories. Each partial training trajectory is generated by generating a random point during the linear degradation process of the complete training trajectory and truncating the complete training trajectory at this point. The obtained partial training trajectories are spliced after the complete training trajectory to achieve data augmentation. The partial training trajectories generated by this method represent the performance trajectories of the device before reaching the failure point, and this trajectory is relatively close to the test data, which can not only increase the number of training data but also better simulate the test data.
[0012] Step 1.2: After Step 1.1, use the z-score method to perform dimensionless processing on the augmented data, and limit the sizes of various parameters in the data within the same interval. The z-score method is defined as follows:
[0013]
[0014] where μ i and σ i represent the mean and standard deviation of the data of the i-th sensor respectively.
[0015] Further, in Step 2:
[0016] Step 2.1: Select three network structures, CNN, GRU, and LSTM, to construct a multi-path feature fusion prediction model, and use their different network advantages to extract features from the performance degradation data.
[0017] Step 2.2: The first path of the model consists of a CNN and a GRU. The CNN is used to extract the spatial features of the degraded data, and then the features are input into the GRU to extract the temporal features. Since data loss occurs during the convolution and pooling operations in the process of feature extraction by the CNN, an LSTM is used as the second path of the model to separately extract the temporal features of the performance degradation data. Since the data is enhanced in Step 1, a GRU with a fast calculation speed is selected to extract the temporal features of the first path. In the model, the Dropout mechanism is used to reduce overfitting and improve the fitting effect of the model.
[0018] Further, in Step 3:
[0019] Step 3.1: The features separately extracted on the two paths in the multi-path feature fusion prediction model in Step 2 are fused using the concat method in TensorFlow into spatio-temporal features.
[0020] Step 3.2: The fused spatio-temporal features in Step 3.1 are input into two fully connected layers for calculation. The last fully connected layer has only one output value, which is the predicted value of the remaining useful life of the engine.
[0021] Further, in Step 4:
[0022] The root mean square error (RMSE) and the scoring function (Score) are used to evaluate the predicted value of the remaining useful life in Step 3 to evaluate the prediction ability of the model. RMSE evaluates the unbiased estimation ability of the model, and Score increases the penalty weight for lag prediction. The expressions of RMSE and Score are as follows:
[0023]
[0024] In the formula, E i represents the error of the i-th prediction, E i <0 represents a leading prediction, and E i >0 represents a lag prediction.
[0025] Finally, the two evaluation methods of RMSE and Sorce are used to evaluate the data with and without data augmentation to verify the effectiveness of data augmentation, and the multi-path feature fusion prediction model is compared with other prediction models in terms of RMSE and Sorce.
[0026] Compared with the existing methods, the present invention has the following beneficial effects:
[0027] The number of model training data is increased through a data augmentation algorithm, meeting the requirements of deep learning for a large amount of data; some of the training trajectory data generated by the data augmentation algorithm is similar to the test data in form and can enhance the prediction performance of the model during the model training process, improving the prediction accuracy.
[0028] The multi-path feature fusion prediction model, through a parallel structure, concentrates the advantages of multiple algorithms in one model, simultaneously learns local spatial features and temporal features, realizes the full mining of input data by the network, and thus achieves the purpose of high-precision prediction; by combining the different advantages of CNN and RNN, the number of parameters is reduced, the calculation speed is further improved, the calculation complexity is reduced, and the overfitting phenomenon in training is reduced. Brief Description of the Drawings
[0029] Figure 1 For data augmentation; (a) Without data augmentation; (b) Data augmentation.
[0030] Figure 2 It is the structure of the multi-path feature fusion prediction model.
[0031] Figure 3 It is the flow chart of the multi-path feature fusion prediction model.
[0032] Figure 4 It is the RUL prediction result. Detailed Implementation Manner
[0033] The following further describes a method for predicting the remaining useful life of an aero-engine based on data augmentation of the present invention with reference to the accompanying drawings.
[0034] The present invention uses the turbine fan engine degradation simulation data set publicly available from NASA and invents a method for predicting the remaining useful life of an aero-engine based on a data augmentation method and a multi-path feature fusion prediction model.
[0035] Step 1: Data prediction processing
[0036] Perform data augmentation and data normalization processing on the adopted CMAPSS data set.
[0037] The specific operation of data augmentation is as Figure 1 shown, Figure 1 (a) in is the complete training trajectory of the RUL of a certain engine, and each moment in the trajectory contains multi-dimensional data. In the augmentation algorithm, partial training trajectories are generated using the complete training trajectory. As Figure 1 (b) in shows the use of Figure 1The complete training trajectory in (a) generates three partial training trajectories. Each partial training trajectory is generated by truncating the complete training trajectory at a random point during the linear degradation process, and the obtained partial training trajectory is placed after the complete training trajectory for data augmentation purposes. The generated partial training trajectories represent the performance trajectories of the device before reaching the failure point, and this trajectory is relatively close to the data in the test set. The data augmentation method in the present invention can not only increase the amount of training set data but also better simulate the data in the test set, enabling the model to better learn the data features and improve the prediction accuracy of the model.
[0038] After data augmentation, the z-score method is used to dimensionless process the augmented data, limiting the magnitudes of various parameters in the data within the same interval to prevent affecting the prediction results. The z-score method is defined as follows:
[0039]
[0040] In the formula, μ i and σ i respectively represent the mean and standard deviation of the data of the i-th sensor.
[0041] Step two: Construct a multi-path feature fusion prediction model to extract the hidden features in the data.
[0042] CNN, GRU, and LSTM are all typical models for solving time series problems currently. The multi-path feature fusion prediction model in the present invention integrates the unique advantages of multiple models, and its structure is as Figure 2 shown, consisting of two parallel paths and a fully connected layer. The first path consists of CNN and GRU. After the data is preprocessed, it is input into CNN to extract spatial features and then into GRU to extract temporal features. Since CNN uses convolution and pooling operations for feature extraction, some information will be lost during this process, resulting in incomplete extraction of the temporal features of the original data by GRU. Therefore, a second parallel path consisting of LSTM is added to extract the complete temporal features. Finally, the extracted data features are fused into spatio-temporal features for the final RUL prediction. For ease of understanding, the proposed flow chart is as Figure 3 shown.
[0043] Step three: Feature fusion
[0044] For the features extracted on the two paths in the multi-path feature fusion prediction model in step two, the concat() method in TensorFlow is used for fusion, fusing them into spatio-temporal features. The fused spatio-temporal features are input into two layers of fully connected layers for calculation. The last fully connected layer has only one neuron, and the output result is the predicted value of the remaining life of the engine.
[0045] Step 4: Evaluate the model, verify the effectiveness of data augmentation, and make comparisons with other models.
[0046] The present invention uses two evaluation criteria, Root Mean Square Error (RMSE) and Score. The nth prediction error E n , and the expressions of RMSE and Score are as follows:
[0047] E n = RUL Est - RUL True
[0048] RUL Est and RUL True represent the predicted value and the true value respectively. E i <0 indicates an over-prediction, and E i >0 indicates a lag-prediction. In practice, lag-predictions may pose safety hazards. Therefore, a scoring function that imposes a greater penalty on lag-predictions is adopted:
[0049]
[0050] The scoring function is sensitive to outliers. Since the prediction errors are not normalized, a single outlier can significantly change the value of the scoring function. Therefore, RMSE is also used to evaluate the unbiased estimation ability of the algorithm:
[0051]
[0052] The data with data augmentation and the data without data augmentation are normalized and then input into the multi-path feature fusion prediction model to predict the remaining useful life of the engine. Two evaluation metrics, RMSE and Score, are used to verify the performance. The results are shown in Table 1. It can be seen from the table that the evaluation metrics of the prediction after data augmentation are significantly better numerically than the prediction results without data augmentation, improving the prediction performance. It can be concluded that the data augmentation method in the present invention has a positive impact on life prediction, making the model fit better and the prediction results more accurate to a certain extent.
[0053] Table 1 Comparison of data augmentation results
[0054]
[0055] The present invention uses the test data in the dataset to verify the model performance. Figure 4For the RUL prediction results of engine No. 68 in the test set, as can be seen from the figure, engine No. 68 showed a relatively stable state throughout the test, and the predicted value curve closely fit the actual value curve. Although there may be fluctuations in the early stage of the test, as the operating cycle of the engine continues to increase, the predicted curve gradually becomes stable.
[0056] The multi-path feature fusion prediction model of the present invention is compared with other RUL prediction models. Due to the certain randomness of model prediction, the present invention conducts multiple experiments to obtain the average results of RMSE and Sorce. As shown in Table 2, compared with other methods, the method used in this article has obtained better overall results.
[0057] As can be seen from Table 2, the model structure of the present invention has obtained the lowest values in both RMSE and Score, indicating that the model structure of this article has higher prediction accuracy and has obtained better performance in the prediction of turbofan engines. Compared with a single LSTM, the RMSE values of the model of the present invention in the 4 sub-datasets are reduced by 26.19%, 25.19%, 19.78%, and 32.02% respectively. Compared with CNN-GRU, the RMSE of the model of the present invention is reduced by 35.4%, 32.54%, 30.35%, and 29.36% respectively. The model of the present invention also obtains a lower Score value in the prediction, showing the advancement of the prediction results, which has higher safety and effectiveness in the application of actual prediction and health management.
[0058] Table 2 Comparison results of different prediction methods
[0059]
[0060] The above introduces a method for predicting the remaining useful life of an aero-engine based on data augmentation proposed by the present invention, and elaborates on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; for those skilled in the art, without departing from the idea of the present invention, there will be changes in the specific implementation manner and application. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for predicting the remaining useful life of an aero-engine based on data augmentation, characterized in that It includes the following steps: Step 1: Use a data augmentation algorithm to augment the training data, and normalize it and the test data using the z-score method; The data augmentation method is as follows: Generate partial training trajectories using the complete training trajectories of the engine. Each partial training trajectory is generated by generating a random point during the linear degradation process of the complete training trajectory and truncating the complete training trajectory at this point. The obtained partial training trajectories are spliced after the complete training trajectory to achieve data augmentation; the generated partial training trajectories represent the performance trajectories of the device before reaching the fault point; Step 2: Construct a multi-path feature fusion prediction model, and input the data processed in Step 1 into the model to extract different features; Use CNN, GRU, and LSTM to construct a multi-path feature fusion prediction model to extract features from the input data; CNN and GRU are the first path of the model. After CNN extracts the spatial features of the input data, it inputs them into GRU to extract the temporal features of the data. LSTM is the second path of the model, which extracts the temporal features of the input data; Step 3: Fuse the features extracted in Step 2 and input them into the fully connected layer for the final remaining useful life prediction; Specifically, use the concat() method in TensorFlow to fuse the features on the two paths separately extracted in the multi-path feature fusion prediction model in Step 2. After fusing them into spatio-temporal features, input them into two fully connected layers for the remaining useful life prediction; Step 4: Use two evaluation methods to evaluate the model, verify the effectiveness of data augmentation, and make comparisons with other models; Use the root mean square error RMSE and the scoring function Score to evaluate the remaining useful life values predicted in Step 3. RMSE evaluates the unbiased estimation ability of the model, and Score increases the penalty weight for lagged prediction. The expressions of RMSE and Score are as follows: where E i represents the error of the i-th prediction, E i < 0 indicates a lead prediction, and E i > 0 indicates a lag prediction.
2. The method for predicting the remaining useful life of an aero-engine based on data augmentation according to claim 1, wherein, In Step 4, use the two evaluation methods of RMSE and Score to evaluate the data with and without data augmentation, verify the effectiveness of data augmentation, and make comparisons of RMSE and Score between the multi-path feature fusion prediction model and other prediction models.
Citation Information
Patent Citations
Turbofan engine remaining service life prediction method based on spatiotemporal feature fusion
CN112580263A