Aero-engine missing data filling method based on trajectory similarity
By using variable weight parameter alignment and variable window filtering algorithms, combined with trajectory similarity measurement, the problem of insufficient accuracy in missing data imputation for aero-engines is solved, achieving efficient data imputation and effective prediction tasks in practical applications.
Patent Information
- Application Number
- CN202510912088.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-28
AI Technical Summary
Existing methods for filling in missing data for aero-engines suffer from insufficient accuracy when considering data dispersion and practical applications. In particular, data-driven methods fail to effectively integrate the dispersion of engine cluster parameters and neglect the functionality of data in practical applications.
A new trajectory similarity measurement algorithm is proposed, which optimizes cross-section alignment by employing a variable weight parameter alignment algorithm and combines it with a variable window filtering algorithm to consider the individual differences of aero-engines. This algorithm comprehensively measures the similarity of data segments from multiple dimensions and verifies the accuracy and functionality of the data.
It significantly improves the accuracy of missing data filling for aero-engines, ensuring that the filled data can play the same role as the real data in predicting remaining service life, thus enhancing the accuracy and functionality of data filling.
Smart Images

Figure CN120849801A_ABST
Abstract
Description
Technical Field
[0001] This invention proposes a method for filling in missing data of aero-engines based on trajectory similarity, which belongs to the field of intelligent aero-engine health management. Background Technology
[0002] Predictive and health management technologies for aero-engines are crucial for assessing engine operating condition, lifespan, and aircraft safety. NASA, in discussions with industry experts, identified five key challenges, including data availability, quality, and ownership. The success of predictive maintenance technologies relies on data, which is essential for training and improving models. However, aero-engines, operating in harsh environments, often suffer from data gaps due to sensor failures, data storage issues, and data privacy requirements. Therefore, accurately filling in these missing aero-engine data has become an important research topic.
[0003] Centralized analysis of field flight data for aircraft of the same engine type revealed significant differences in identical air path parameters within a certain range, exhibiting a highly dispersed cluster. The main reasons for these differences include: 1) tolerances during production and assembly; 2) variations in operating time, status, and mission on different aircraft; and 3) varying degrees of noise contamination in sensor data. These factors cause the same mechanistic characteristics to exhibit different external behaviors in different individuals or stages, increasing the difficulty of filling in missing data.
[0004] Currently, mainstream aero-engine missing data imputation algorithms can be broadly categorized into three types: traditional methods based on historical information, physical model methods, and data-driven methods. Data-driven methods have garnered significant attention due to their avoidance of complex mathematical models. Patent CN116562141A discloses a method for predicting the remaining life of multi-dimensional degraded equipment, which constructs high-dimensional feature information by extracting local features to address data missingness. However, this method treats aero-engine missing data imputation as a data preprocessing step in the prediction of remaining life or important parameters, lacking a detailed description of the imputation process and verification of the accuracy of the imputed data. Patent CN118296758A discloses an imputation method based on bidirectional dynamics and attribute association, utilizing a long short-term memory neural network to extract temporal dependency information. While it can integrate attribute association and temporal dependency, it does not consider the dispersion of engine cluster parameters and only evaluates the similarity between the imputed data and the real data from a statistical perspective. Although it can provide some quantitative results, it often neglects the functionality of the data in practical applications. This invention proposes an aero-engine missing data imputation method based on trajectory similarity, which can effectively solve these problems. Summary of the Invention
[0005] To address the problems of existing data-driven methods, this invention proposes a method for filling in missing aero-engine data based on trajectory similarity, which achieves accurate filling of missing aero-engine data.
[0006] To achieve the above objectives, the concept and technical solution of the present invention are as follows:
[0007] The basic concept of this invention is to propose a variable-weight parameter alignment algorithm to optimize cross-sectional alignment; to propose a variable window filtering algorithm to consider the parameter dispersion caused by the individual differences of aero-engines, and to combine it with the variable-weight parameter alignment algorithm to find the segment most similar to the missing data segment on each engine in the sample library; to propose a new trajectory similarity measurement algorithm to comprehensively measure the similarity of data segments from multiple dimensions, including the value of the variable-weight parameter alignment formula, window length, and starting position; and to verify the accuracy of the data and its functionality in predicting remaining service life. This method pays particular attention to the performance of the data in actual prediction tasks, significantly improving the accuracy of data imputation.
[0008] Based on the above basic concept, the technical solution proposed in this invention is a method for filling in missing aero-engine data based on trajectory similarity, comprising the following steps:
[0009] Step 1: Propose a variable weight parameter alignment algorithm to optimize cross-section alignment;
[0010] Step 2: Propose a variable window filtering algorithm to consider the parameter dispersion caused by the individual differences of aero-engines;
[0011] Step 3: Propose a new trajectory similarity measurement algorithm to comprehensively measure the similarity of data segments from multiple dimensions, including the value of the variable weight parameter alignment formula, window length, and starting position;
[0012] Step 4: Verify the accuracy of the data and its functionality in predicting remaining useful life.
[0013] Furthermore, step 1 proposes a variable weight parameter alignment algorithm to optimize cross-sectional alignment, achieving cross-sectional alignment from the perspective of remaining useful life prediction. Unlike ordinary cross-sectional data alignment methods, the variable weight parameter alignment algorithm not only pursues the alignment of the data itself, but also considers the performance of the incomplete data in actual prediction tasks such as remaining useful life prediction. Using correlation analysis, each parameter is assigned a corresponding weight, as shown in the formula:
[0014]
[0015] Where h is in the formula kThe Pearson correlation coefficient between a certain parameter and remaining useful life, where p is the parameter, n is the number of parameters, and i and j represent two different cross-sections. The smaller the VPA value, the higher the similarity.
[0016] The VPA formula differs from ordinary cross-sectional data alignment methods. It not only pursues data alignment itself but also considers the specific tasks the imputed data will be used for, such as remaining service life prediction. It uses correlation analysis to assign appropriate weights to each parameter. If a parameter is strongly correlated with remaining service life, its weight is high; otherwise, its weight is low. The weight of a parameter is the Pearson correlation coefficient between that parameter and remaining service life. This allows the imputed data to achieve better results in remaining service life prediction. The parameters of each engine in the sample database are weighted according to their correlation with its remaining service life. When comparing missing data segments with data from each engine in the sample database, the weight of that engine is used. The similarity between two cross-sections is obtained according to the VPA formula. Similarly, if the imputed data is used for other data analysis, remaining service life can be replaced with other important parameters.
[0017] Furthermore, step 2 proposes a variable window filtering algorithm to consider the parameter dispersion caused by the individual differences of aero-engines. Besides sharing many basic characteristics, a group of engines of the same model will also exhibit some individual differences. These differences are visually reflected in sensor data as variations such as data stretching and distortion. These variations can be abstractly represented using simple data structures.
[0018] When data stretching occurs, the expression is:
[0019]
[0020] Within an engine group, the same failure mode may exhibit different degradation rates on different individuals, and the performance degradation rate of each individual is not constant at different stages. Therefore, different degradation rates can be characterized by randomly stretching the gas path parameter data along the time axis.
[0021] When data shift occurs, the transformation expression is:
[0022]
[0023] Due to differences in manufacturing and assembly processes and variations in operating environments, the onset time of engine performance degradation varies. By shifting random data segments along the timeline, certain degradation characteristics can be advanced or delayed, thus characterizing different degradation onset times.
[0024] For the cross-sectional data at both ends of the missing data segment, the aforementioned data distortion can be addressed by changing the window size. Based on the length of the missing data, the parameter dispersion of the engine group in the sample library is analyzed to determine the window length range. For each reference engine trajectory, windows of different sizes within the range are slid sequentially, and similarity comparison is performed using VPA. The most similar data is selected and recorded to form a new similar trajectory library. The new similar trajectory library contains a segment of data for each engine, rather than the entire trajectory.
[0025] Furthermore, step 3 proposes a novel trajectory similarity measurement algorithm that comprehensively measures the similarity of data segments from multiple dimensions, including the value of the variable weight parameter alignment formula, window length, and starting position. Window length may reflect the deformation of certain features in sensor data, while position reflects the time when the feature occurred. The length and starting position of the window should be analyzed in conjunction. For example, if the length of a window is 1.5 times that of the missing data segment, the difference may seem large at first glance, but if the starting position of the window is also far away, the window may still serve as a reference for filling in the missing data. This is because in an engine cluster, the same degradation pattern may exhibit different degradation rates on different individuals, and even the degradation rate of each individual at different stages is not constant. If, in the time series of a certain engine, the degradation pattern is stretched overall or partially, it can still serve as a reference even if the window length is longer than the missing data segment. When the VPA value is small and the difference is not significant, windows with both length and position close to the missing data segment are preferred; if none are available, all indicators of the window are considered.
[0026] The formula for the new trajectory similarity measurement algorithm is defined as follows:
[0027]
[0028] In the formula, P represents the starting position of the window, and L represents the window length. The subscript t represents the cross-sectional characteristics of the engine under test. c represents a constant, and the setting of the constant coefficients needs to be considered in conjunction with the length and position of the window. Adjusting these constant coefficients according to the actual situation can reasonably achieve data filling.
[0029] Furthermore, step 4 verifies the accuracy of the data and its functionality in predicting remaining useful life. The specific steps are as follows:
[0030] Step 4.1: After obtaining the sensor data, preprocess the data;
[0031] Step 4.2: Verify that the filled data can numerically follow the curve of the true value.
[0032] Step 4.3: Verify whether the supplemented data can achieve the same effect as the real data in the important work of engine health management. Remaining service life prediction is one of the core tasks of engine health management. The effectiveness of the supplementation method is judged by comparing the effect of supplemented data and real data in predicting remaining service life.
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0034] This invention addresses the issue of missing data in aero-engine data acquisition and storage caused by sensor malfunctions, data storage server problems, and human factors. It proposes a trajectory similarity-based method for filling in missing aero-engine data, accurately filling in the gaps. A variable-weight parameter alignment algorithm is proposed to optimize cross-sectional alignment, and a variable-window filtering algorithm is proposed to consider parameter dispersion caused by individual differences in aero-engines, finding the segment most similar to the missing data segment on each engine in the sample library. A novel trajectory similarity measurement algorithm is proposed to comprehensively measure the similarity of data segments from multiple dimensions, including the value of the variable-weight parameter alignment formula, window length, and starting position, to achieve missing data filling. The accuracy of the data and its functionality in remaining service life prediction are verified. This invention focuses on the performance of the data in actual prediction tasks. Attached Figure Description
[0035] Figure 1 This is a flowchart of the missing data filling method for aero-engines based on trajectory similarity provided by the present invention;
[0036] Figure 2 This is a schematic diagram of the variable weight parameter alignment algorithm provided by the present invention;
[0037] Figure 3 This is a schematic diagram of the variable window filtering algorithm provided by the present invention;
[0038] Figure 4 This is a schematic diagram of the data filling and verification method provided by the present invention;
[0039] Figure 5 This is a schematic diagram showing the results of filling in the missing data of the engine AD provided by the present invention;
[0040] Figure 6 This is a schematic diagram showing the results of filling in the missing data of the engine eh provided by the present invention. Detailed Implementation
[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] To verify the effectiveness of the proposed method, this example uses the publicly available NASA C-MAPSS simulation dataset. This example uses the FD001 dataset for model training and testing; the training set contains data from 100 engines, and the test set contains data from 100 engines.
[0043] Please see Figure 1 The flowchart shows a method for filling missing data in aero-engines based on trajectory similarity provided by the present invention, which specifically includes the following steps:
[0044] Step 1: Please refer to Figure 2 This diagram illustrates a variable-weight parameter alignment algorithm that achieves more accurate alignment from the perspective of remaining useful life prediction. Unlike traditional cross-sectional data alignment methods, the variable-weight parameter alignment algorithm not only focuses on the alignment of the data itself but also considers the performance of the imputed data, such as in remaining useful life prediction. It utilizes correlation analysis to assign corresponding weights to each parameter, as shown in the formula:
[0045]
[0046] Where h is in the formula k The Pearson correlation coefficient between a certain parameter and remaining useful life, where p is the parameter, n is the number of parameters, and i and j represent two different cross-sections. The smaller the VPA value, the higher the similarity.
[0047] The VPA formula not only focuses on the alignment of the data itself but also considers the application of the imputed data in a specific task. Through correlation analysis, each parameter is assigned a corresponding weight. If a parameter has a strong correlation with remaining useful life, it is given a larger weight; conversely, if the correlation is weak, the weight is smaller. The weight of a parameter is determined by its Pearson correlation coefficient with remaining useful life. In this way, the imputed data can perform better in predicting remaining useful life. The parameters of each engine in the sample database are assigned weights based on their correlation with its remaining useful life. When comparing the missing data segment with the data of each engine in the sample database, the weight of that engine is used. Finally, the similarity between the two segments is calculated using the VPA formula.
[0048] Step 2: Please refer to Figure 3The diagram illustrates a variable window filtering algorithm. Engine groups of the same model not only possess numerous basic characteristics but also exhibit some individual differences. These differences are visually represented in sensor data as stretching, deformation, and other variations. These changes can be abstractly represented using simple data structures.
[0049] When data stretching occurs, the expression is:
[0050]
[0051] Within an engine group, the same failure mode may exhibit different degradation rates on different individuals, and the performance degradation rate of each individual is not constant at different stages. Therefore, different degradation rates can be characterized by randomly stretching the gas path parameter data along the time axis.
[0052] When data shift occurs, the transformation expression is:
[0053]
[0054] Due to differences in manufacturing and assembly processes and variations in operating environments, the onset time of engine performance degradation varies. By shifting random data segments along the timeline, certain degradation characteristics can be advanced or delayed, thus characterizing different degradation onset times.
[0055] After performing similarity matching between cross-sectional data using a variable-weight parameter alignment algorithm, a filtering algorithm with a variable window is used to find the most similar data segment on each engine trajectory in the sample library, forming a new similar trajectory library. Considering the evolution of engine characteristics across individuals, reflected in data stretching, translation, and other transformations, windows of different sizes are used to slide sequentially on each engine trajectory. The range of window length is determined based on the characteristics of the training set samples. Emphasis is placed on the lifetime distribution characteristics of the training samples. The upper and lower limits of the window length are set to half and twice the length of the missing data, respectively.
[0056] Step 3: Propose a novel trajectory similarity measurement algorithm that comprehensively measures the similarity of data segments from multiple dimensions, including the value of the variable weight parameter alignment formula, window length, and starting position. The window length may reflect the deformation of certain features in the sensor data, while the position reflects the time when the feature occurred. The length and starting position of the analysis window should be considered together.
[0057] The formula for the new trajectory similarity measurement algorithm is defined as follows:
[0058]
[0059] In the formula, P represents the starting position of the window, L represents the window length, the subscript t represents the cross-sectional characteristics of the engine under test, and c represents a constant.
[0060] The setting of constant coefficients needs to be considered in conjunction with the length and position of the window. If the position is far back and the length is short, increase the weight of c2 to emphasize the location of the missing data and facilitate the learning of features near the same position. If the position is near the front and the length is long, increase the weight of c3 to emphasize the length of the missing data and facilitate the learning of features contained in long data segments. If the position and length are both moderate, increase the weight of c1 to balance the position and the features of the data.
[0061] Step 4: Please refer to Figure 4 The diagram illustrates a data validation method to verify the accuracy and functionality of the data in predicting remaining useful life. The specific steps are as follows:
[0062] Step 4.1: After obtaining the sensor data, preprocess the data.
[0063] To eliminate the dimensional influence between sensor data features, a maximum-minimum standard normalization method is selected. The normalization calculation formula is:
[0064]
[0065] In the formula, X represents sensor data, X min and X max These represent the minimum and maximum values of the corresponding sensor data.
[0066] This invention uses aircraft engine exhaust temperature as an example to verify the missing data imputation method. The exhaust temperatures and threshold values of individual engines are distributed within a certain range. For the same exhaust temperature value, some engines can still operate normally for a long time, while others are already approaching the threshold. Therefore, accurately imputing the missing data for each engine and the deeper information contained within that missing data is crucial.
[0067] Step 4.2: Verify that the filled data can numerically follow the curve of the true value.
[0068] Considering the specific reasons for lost data from aero-engines, if it's due to sensor damage, it typically manifests as the loss of partial data for a certain parameter, and the lost data won't be too long in the time series, as sensor anomalies will be detected in real-time monitoring and analysis. However, when all sensor data for an engine is lost over a period of time, it's highly likely due to human error, such as data storage problems, or the engine being transferred from one management system to another due to a change in service location, resulting in a temporary inaccessibility of a segment of data. In this case, the missing data segment is likely to be relatively long, so the chosen length should be randomized within a certain range. The average lifespan of the engines in the training set is around 200 cycles, with a minimum lifespan of around 130 cycles and a maximum lifespan of around 350 cycles. Therefore, the length of the missing data is set between 20 and 80 cycles. The location also considers the early, middle, and late stages of the lifespan.
[0069] To visually demonstrate the data imputation effect, we selected some representative engines with missing data. Four engines had lifespans near the average, with missing data in the early-mid, mid-late, and late stages, respectively; two engines had longer lifespans, with missing data in the early and mid-late stages; the remaining engines had shorter lifespans, with missing data in the mid-late stages. For a detailed description of the imputation effect, please refer to [link to documentation / reference]. Figure 5 and Figure 6 The results show that the new method can follow the real data curve well and capture some seemingly inconspicuous features in the real data that may be of great help to subsequent health management work;
[0070] Step 4.3: Verify whether the supplemented data can achieve the same effect as the actual data in predicting the remaining service life of the engine.
[0071] The Long Short-Term Memory (LSTM) network used for lifetime prediction uses 100 units. The learning rate is set to 0.001, and the Adam optimizer is used. The Xavier initializer is used to initialize the weights of all networks. The epochs are 2000. The batch size is set to 450. L1 regularization is used to better distribute the weights of specific hidden layers. The sliding window size is 15×30, with a stride of 1. The model automatically records the best training result.
[0072] The last 30 flight cycles of each engine in the test set were used as the data segment to be filled. The data for the last 30 flight cycles of all engines were estimated using a new method, and the remaining service life was predicted using the same Long Short-Term Memory (LSTM) network. Since the sliding window size used was 15×30, the data used to predict the remaining service life of the test set samples came entirely from the estimation of the new method, which strongly demonstrates the effectiveness of the data filling.
[0073] Using both the actual and predicted values as inputs to the model, a Long Short-Term Memory (LSTM) network was used to predict the remaining useful life. The root mean square error (RMSE) and score of the results are shown in Table 1. The scores for both are almost identical, and the RSME results are also very similar. The results indicate that data imputation not only numerically follows the curve of the actual values but also achieves the same effect as the actual data in predicting remaining useful life.
[0074] Table 1 Comparison of Results
[0075] Metric Algorithm Actual data Fill in the data RMSE 12.76 12.71 Score 324 325
[0076] To demonstrate the superiority of the proposed method, it is compared with missing data estimation methods such as no missing data filling, mean imputation, nearest neighbor imputation, linear regression, and Long Short-Term Memory Neural Network (ConvLSTM) with convolutional kernels based on classical machine learning methods. The results are shown in Table 2.
[0077] Table 2 Comparison of Methods
[0078] method RMSE Score Real data 12.76 324 No filling 27.91 12303 Mean interpolation 15.57 537 Nearest neighbor interpolation method 15.15 458 Linear regression method 13.74 382 Based on ConvLSTM 15.52 526 This invention 12.71 325
[0079] This invention is not limited to the above embodiments. Based on the technical solutions disclosed in this invention, those skilled in the art can make some simple modifications, equivalent changes and alterations to some of the technical features without creative effort, all of which fall within the scope of the technical solutions of this invention.
Claims
1. A method for filling in missing data of aero-engines based on trajectory similarity, characterized in that, Specifically, it includes the following steps: Step 1: Propose a variable weight parameter alignment algorithm to optimize cross-section alignment; Step 2: Propose a variable window filtering algorithm to consider the parameter dispersion caused by the individual differences of aero-engines; Step 3: Propose a new trajectory similarity measurement algorithm to comprehensively measure the similarity of data segments from multiple dimensions, including the value of the variable weight parameter alignment formula, window length, and starting position, to achieve missing data filling; Step 4: Verify the accuracy of the data and its functionality in predicting remaining useful life.
2. The method for filling in missing aero-engine data based on trajectory similarity according to claim 1, characterized in that, Step 1 proposes a variable-weight parameter alignment algorithm to optimize cross-sectional alignment. Unlike ordinary cross-sectional data alignment methods, the variable-weight parameter alignment algorithm not only pursues the alignment of the data itself, but also considers the performance of the imputed data in actual prediction tasks such as remaining useful life prediction. Using correlation analysis, each parameter is assigned a corresponding weight, as shown in the formula: Where h is in the formula k The Pearson correlation coefficient between a certain parameter and remaining useful life, where p is the parameter, n is the number of parameters, and i and j represent two different cross-sections. The smaller the VPA value, the higher the similarity.
3. The method for filling in missing aero-engine data based on trajectory similarity according to claim 1, characterized in that, Step 2 proposes a variable window filtering algorithm that considers the parameter dispersion caused by individual differences in aero-engines. Based on the length of missing data, the range of window length is determined by analyzing the parameter dispersion of the aircraft group in the sample database. Linear transformations, translation transformations, and combinations thereof for two-dimensional data can be expressed by a single formula: In the formula, the leftmost side of the equation represents the transformed data, the rightmost side represents the original data, and the 3×3 matrix is the transformation matrix.
4. The method for filling in missing aero-engine data based on trajectory similarity according to claim 1, characterized in that, Step 3 proposes a novel trajectory similarity measurement algorithm that comprehensively measures the similarity of data segments from multiple dimensions, including the value of the variable weight parameter alignment formula, window length, and starting position, to achieve missing data imputation. The formula for the new trajectory similarity measurement algorithm is defined as follows: In the formula, P represents the starting position of the window, and L represents the window length. The subscript t represents the cross-sectional characteristics of the engine under test. c represents a constant, and the setting of the constant coefficients needs to be considered in conjunction with the length and position of the window. Adjusting these constant coefficients according to the actual situation can reasonably achieve data filling.
5. The method for filling in missing aero-engine data based on trajectory similarity according to claim 1, characterized in that, Step 4 verifies the accuracy of the data and its functionality in predicting remaining service life. The supplementary data should not only numerically follow the curve of the actual values, but also achieve the same effect as the actual data in the crucial work of engine health management.
Citation Information
Patent Citations
Method for predicting residual life of multivariate degradation equipment in consideration of data missing
CN116562141A
Aero-engine parameter-oriented missing value filling method based on bidirectional dynamic and attribute association
CN118296758A