A method and apparatus for predicting the RUL of aero-engines based on MSFormer and domain adaptation
By using MSFormer and domain-adaptive methods, data features at multiple time scales are extracted and the loss function is improved. This solves the problem that local details and global correlations of data are not comprehensively considered in the prediction of the remaining service life of aero-engines, and achieves efficient and accurate prediction of remaining service life.
Patent Information
- Application Number
- CN202411705528.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-11-26
AI Technical Summary
Existing technologies fail to comprehensively consider both local details and global correlations in predicting the remaining service life of aero-engines, and domain-adaptive methods fail to effectively align the remaining service life relationships between samples, resulting in unsatisfactory prediction results.
We employ an MSFormer-based and domain-adaptive approach, extracting data features through a multi-scale transformer, comprehensively considering information using Transformer, and improving the loss function to align the feature distributions of the source and target domains. We also use selectable source domain labels to assist in correcting the alignment process and restore the target domain features to the target domain data to preserve domain-specific information.
It improves the accuracy and processing efficiency of predicting the remaining service life of aero-engines, ensures the feature alignment effect of the model across different domains, and enhances the accuracy of prediction.
Smart Images

Figure CN119669678B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of aero-engine health management technology, and in particular to an aero-engine RUL prediction method and device based on MSFormer and domain-adaptive technology. Background Technology
[0002] As the primary power source for modern aircraft, the health of aircraft engines plays a crucial role. A failure in an aircraft engine during operation can jeopardize aircraft safety and result in significant losses of human and financial resources. Remaining Useful Life (RUL) refers to the time an aircraft engine is expected to operate normally under current conditions until it fails. Predicting the RUL of an aircraft engine before operation can help prevent engine failures, avoid major accidents, and ensure the reliability of aircraft operation.
[0003] In recent years, with the continuous development of deep learning, a number of methods using deep learning to predict the RUL of aero-engines have emerged in the field of RUL prediction. One method utilizes a multi-cell long short-term memory network model to process input data according to its importance, achieving better predictive performance compared to traditional machine learning methods. Another method employs a multi-encoder layer Transformer network architecture to achieve parallel processing of data of different time lengths, enhancing the model's ability to handle short-term time-series features while preserving long-term dependencies. A third method, based on operating condition clustering and residual self-attention, eliminates the impact of complex engine operating conditions on data analysis through data preprocessing methods such as clustering and normalization of the original aero-engine data, and uses the proposed model to predict the RUL of aero-engines.
[0004] With the continuous development of computer technology, deep learning techniques, such as Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Temporal Convolutional Networks (TCN), and Transformers, have been applied to RUL (Remaining Lifetime) prediction for aero-engines, achieving certain predictive results. However, algorithms using traditional deep learning methods for aero-engine RUL prediction require a large amount of labeled data, which is difficult to obtain in practice due to the scarcity of labeled data throughout the entire lifecycle of aero-engines. To address the problem of limited labeled data, domain adaptation techniques have been applied to RUL prediction. Existing domain-adaptive aero-engine RUL prediction methods do not comprehensively consider the local details and global correlations of the data when extracting features, and only perform overall alignment between domains in terms of domain alignment, without considering the relationship between remaining lifetimes of samples. Furthermore, the alignment process also loses a significant amount of target domain-specific information, all of which contribute to unsatisfactory prediction results.
[0005] In the existing technology, there is a lack of a high-efficiency and accurate RUL prediction method for aero-engines based on multi-scale transformers and domain adaptation. Summary of the Invention
[0006] To address the technical problems of existing technologies that fail to comprehensively consider local details and global correlations in feature extraction, and fail to take into account the relationship between remaining lifetimes of samples in domain alignment, this invention provides a method and apparatus for predicting the Remaining Lifetime (RUL) of aero-engines based on MSFormer and domain adaptation. The technical solution is as follows:
[0007] On the one hand, a method for predicting the RUL of aero-engines based on MSFormer and domain adaptation is provided. This method is implemented by an aero-engine RUL prediction device and includes:
[0008] Acquire data from the first aero-engine; based on the operating conditions of the aero-engine, preprocess the data from the first aero-engine to obtain source domain data and target domain data;
[0009] Based on domain adaptation technology, a domain adaptation model is constructed according to the MSFormer model structure and the NCE model structure;
[0010] The source domain data and the target domain data are used to train the domain adaptive model to obtain the trained domain adaptive model.
[0011] Acquire the aero-engine data to be predicted; input the aero-engine data to be predicted into the trained domain adaptive model to perform aero-engine RUL prediction and obtain the remaining service life of the aero-engine.
[0012] On the other hand, a field-adaptive RUL prediction device for aero-engines is provided, which is applied to the field-adaptive RUL prediction method for aero-engines. The device includes:
[0013] The data acquisition module is used to acquire data from the first aero-engine; based on the operating conditions of the aero-engine, the data from the first aero-engine is preprocessed to obtain source domain data and target domain data;
[0014] The model building module is used to build a domain-adaptive model based on the MSFormer model structure and the NCE model structure, using domain adaptation technology.
[0015] The model training module is used to train the domain adaptive model using the source domain data and the target domain data to obtain the trained domain adaptive model.
[0016] The model application module is used to acquire the aero-engine data to be predicted; the aero-engine data to be predicted is input into the trained domain adaptive model to predict the aero-engine's RUL (Remaining Service Life) and obtain the remaining service life of the aero-engine.
[0017] On the other hand, an aircraft engine RUL prediction device is provided, the aircraft engine RUL prediction device comprising: a processor; a memory storing computer-readable instructions, wherein when the computer-readable instructions are executed by the processor, any one of the above-described aircraft engine RUL prediction methods based on MSFormer and domain adaptation is implemented.
[0018] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement any of the above-described methods for predicting the RUL of aero-engines based on MSFormer and domain adaptation.
[0019] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0020] This invention proposes a method for predicting the RUL (Range Limiting) of aero-engines based on MSFormer and domain adaptation.
[0021] This invention addresses the shortcomings of existing research techniques by proposing a RUL prediction method for aero-engines based on multi-scale transformers and domain adaptation. This method extracts data features across multiple time scales and utilizes a Transformer to comprehensively consider information, improving the domain-adaptive loss function. It aligns the feature distributions of the source and target domains, ensuring that internal data are aligned with similar labels as much as possible, and uses selectable source domain labels to assist in correcting the alignment process. Furthermore, this invention ensures that target domain features retain as much domain-specific information as possible by restoring them to target domain data, thus improving the model's RUL prediction performance. This invention is a highly efficient and accurate RUL prediction method for aero-engines based on multi-scale transformers and domain adaptation. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of an aero-engine RUL prediction method based on MSFormer and domain adaptation provided in an embodiment of the present invention;
[0024] Figure 2 This is a block diagram of an aero-engine RUL prediction device based on MSFormer and domain adaptation provided in an embodiment of the present invention;
[0025] Figure 3 This is a schematic diagram of the structure of an aircraft engine RUL prediction device provided in an embodiment of the present invention. Detailed Implementation
[0026] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0027] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0028] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0029] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0030] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0031] This invention provides a method for predicting the RUL (Responsive Limitation) of aero-engines based on MSFormer and domain adaptation. This method can be implemented by an aero-engine RUL prediction device, which can be a terminal or a server. Figure 1 The flowchart shown is for aero-engine RUL prediction method based on MSFormer and domain adaptation. The processing flow of this method may include the following steps:
[0032] S1. Obtain the data of the first aero-engine; based on the operating conditions of the aero-engine, preprocess the data of the first aero-engine to obtain source domain data and target domain data.
[0033] Optionally, based on the operating conditions of the aero-engine, the data of the first aero-engine is preprocessed to obtain source domain data and target domain data, including:
[0034] Calculate the RUL tag based on the first aero-engine data; add the RUL tag to the first aero-engine data to obtain the second aero-engine data;
[0035] Based on the operating conditions of the aero-engine, the data of the second aero-engine is divided to obtain the first source domain data and the first target domain data.
[0036] The first source domain data and the first target domain data are normalized to obtain the second source domain data and the second target domain data.
[0037] Based on preset sliding window parameters, the sliding window method is used to divide the second source domain data and the second target domain data into samples to obtain source domain data and target domain data.
[0038] In one feasible implementation, this invention addresses the problem of limited labeled data by applying domain adaptation technology to RUL prediction. This technology requires two parts of data: one part consists of data whose RULs need to be predicted but lack labels, referred to as the target domain; the other part consists of labeled data with a distribution similar to the target domain, referred to as the source domain. Domain adaptation technology primarily uses labeled source domain data and unlabeled target domain data to train the network model, ensuring that the features extracted from the two domains are as similar as possible, thereby enabling the network model to predict RULs on the target domain data.
[0039] The collected raw data of the first aero-engine is preprocessed, and the source domain and target domain are divided according to the operating conditions.
[0040] The RUL label is obtained by subtracting the current cycle number from the maximum cycle number for each engine. The second aero-engine data was obtained. Engine features were extracted and normalized to eliminate the influence of different feature scales. This process is as follows (1):
[0041] (1);
[0042] in, For the normalized data, For the maximum value of the characteristic, The minimum feature value is used. The samples are divided using a sliding window method, and the sliding window length is set. and sliding window step length The data required for the sliding window is time-series data from sensors that record aero-engine data, including both time and feature dimensions.
[0043] S2. Based on domain adaptation technology, construct a domain adaptation model according to the MSFormer model structure and the NCE model structure.
[0044] The domain-adaptive model includes an extractor, a predictor, a discriminator, and a reconstructor.
[0045] The extractor is built based on the MSFormer model structure; the predictor is built using a single fully connected layer; the discriminator is built using a three-layer fully connected layer; and the reconstructor is built based on the NCE model structure.
[0046] The extractor includes a source domain extractor and a target domain extractor; the structure of the extractor includes a patching module, a sample mapping module, an LSTM model, a Transformer encoder, and a target domain mapping module.
[0047] In one feasible implementation, the domain adaptive model constructed by the present invention mainly includes four modules: extractor, predictor, discriminator, and reconstructor.
[0048] The extractor extracts high-dimensional features from the sample input; the predictor obtains the RUL prediction result based on the high-dimensional features; the discriminator determines whether the sample comes from the source domain or the target domain based on the high-dimensional features; and the reconstructor restores the target domain features to the data to retain as much domain-specific information as possible in the features. The extractor is implemented using a Multi-Scale Transformer (MSFormer) model, the predictor uses a single fully connected layer, the discriminator uses a three-layer fully connected layer, and the reconstructor uses a Noise Contrastive Estimation (NCE) model.
[0049] The extractor includes a source domain extractor. and target domain extractor The process of extracting features from a sample using an extractor can be represented by the following equation (2):
[0050] (2);
[0051] in, This represents the samples segmented by the sliding window method, which will be fed into the MSFormer model to extract high-dimensional features. and and the target of feature restoration .sample The shape is , The length of the sliding window represents The time dimension; for The number of features.
[0052] Internally, the MSFormer model contains samples. The size will be set in advance. and step length The Patching module divides data into groups of different lengths. The process is as follows: (3) and (4):
[0053] (3);
[0054] (4);
[0055] For a shape of The sample, after being processed by a size of Step size is After the Patching module, the resulting sub-data shape is , This indicates the number of sub-data segments obtained after one operation.
[0056] According to the analysis, it can be seen that The varying shapes of the samples, to some extent, reflect the information contained within the samples at different time intervals. This is achieved through a sample mapping module composed of fully connected layers. Can To transform it into the same shape, the process is as follows (5):
[0057] (5);
[0058] The shape of the input sample mapping module is The sub-data, after being processed by the fully connected layer, can ultimately yield a shape of... The mapping data is given by R, where R represents the dimensions to which the sample and target domains are mapped. During the mapping process, intermediate dimensions are used to ensure the data dimensions remain consistent. After being mapped and stitched together, the final shape obtained is as follows: Data This data will be used in the reconstructor during the transfer training phase.
[0059] After the data is segmented using the Patching module, each segment contains more detailed information within a fixed time period. This information is extracted using the LSTM model, as shown in equation (6):
[0060] (6);
[0061] Input shape is The data ultimately yields a shape of The output data will be After inputting the model, the output data is obtained and concatenated to finally obtain the shape as follows: of It contains a wealth of time-related details.
[0062] exist At the start of the time dimension, a classification token (CLS) is added, and the result is passed to a single fully connected layer for embedding, resulting in a shape of... Data The data is then fed into the Transformer encoder to extract the integrated information between different time scales, as shown in equation (7):
[0063] (7);
[0064] in, For the encoder's output characteristics, the shape should be similar to... Consistent. Divide it into two parts, as shown in equation (8):
[0065] (8);
[0066] in, It refers to the features extracted from the model at the beginning of the time dimension after adding the CLS marker, which are used for training transfer. concat means concatenation.
[0067] In the target domain mapping module, it is necessary to use To reconstruct It needs to go through the feature mapping module first. Will Mapping to and With the same shape, the target domain mapping module is composed of fully connected layers as shown in equation (9):
[0068] (9);
[0069] in, These are the processing characteristics passed to the refactorer.
[0070] The predictor is a single-layer fully connected layer structure, with input from the target domain extractor. Features obtained The RUL prediction result is obtained as shown in equation (10):
[0071] (10);
[0072] Here, FC represents a single fully connected layer. In the following text, FC also represents a single fully connected layer, but with different parameters.
[0073] The discriminator used in this invention has a three-layer fully connected structure, and the input features... The domain label judgment results of the sample are obtained, and the process is as follows: (11), (12), (13):
[0074] (11);
[0075] (12);
[0076] (13);
[0077] The expression for the refactorer is as follows (14):
[0078] (14);
[0079] in, This represents an NCE model consisting of multiple fully connected layers. For target domain features, The target domain data reconstructed for the model.
[0080] In order for the reconstructor to pass the target domain features Restore target domain data ,but It retains a relatively large amount of domain-specific information. This invention uses fully connected layers... Perform restoration. The number of fully connected layers in the model is represented by The time dimension is the same, and each fully connected layer receives... Input and output a time step And ultimately, in the time dimension splicing into and Vectors with exactly the same dimensions.
[0081] Information Noise Contrastive Estimation (InfoNCE) is chosen as the fourth loss function to ensure that as much useful information as possible is retained at each time step during reconstruction. The calculation formula is shown in Equation (15) below:
[0082] (15);
[0083] in, This indicates that the mean is calculated.
[0084] S3. Using source domain data and target domain data, train the domain adaptation model to obtain the trained domain adaptation model.
[0085] Optionally, the domain adaptation model is trained using source domain data and target domain data to obtain a trained domain adaptation model, including:
[0086] Based on the extractor, the predictor is pre-trained using source domain data to obtain an optimized predictor;
[0087] Based on the extractor and the optimized predictor, the discriminator is trained by transfer learning using source domain data and target domain data to obtain an optimized discriminator;
[0088] Based on the optimized discriminator and reconstructor, the extractor is trained by transfer learning according to the target domain data to obtain the optimized extractor.
[0089] In one feasible implementation, the training process of the present invention is divided into two stages: pre-training and transfer training.
[0090] The transfer training process includes training the discriminator. and training feature extractor Two parts. Primarily through an adversarial approach, using a continuously trained discriminator. and feature extractor The confrontation between them, that is During training, it tends to favor domain labels that can recognize features, while It tends to confuse This process prevents the recognition of domain labels, thereby continuously improving the capabilities of both modules and ultimately achieving the effect of bringing the target domain features closer to the source domain features. The two parts of transfer training are performed alternately.
[0091] Optionally, based on the extractor, the predictor is pre-trained using source domain data to obtain an optimized predictor, including:
[0092] The source domain data is input into the extractor for feature extraction to obtain the source domain data features;
[0093] The source domain data features are input into the predictor to predict the remaining useful life, and the first useful life prediction value is obtained.
[0094] Based on the source domain data, the loss function is calculated according to the first lifetime prediction value to obtain the mean squared error loss;
[0095] Based on the mean squared error loss, the predictor parameters are optimized to obtain an optimized predictor.
[0096] In one feasible implementation, during the pre-training process, the source domain training data is fed into the extractor. and predictor After obtaining the RUL prediction results of the samples, the mean squared error loss (MSELoss) is calculated and the network parameters are trained. The loss function is calculated as follows (16):
[0097] (16);
[0098] in, Represents true RUL, Indicates the prediction result. Indicates the number of samples.
[0099] Optionally, based on the extractor and the optimized predictor, the discriminator is transferred-trained using source domain data and target domain data to obtain an optimized discriminator, including:
[0100] The source domain data and target domain data are input into the extractor for feature extraction and CLS labeling is added to obtain CLS-labeled source domain data features and CLS-labeled target domain data features;
[0101] The source domain data features and target domain data features of CLS-tagged data are input into the discriminator for domain discrimination to obtain the first domain discrimination result and the second domain discrimination result.
[0102] Based on the CLS-labeled source domain data features, the CLS-labeled target domain data features, and the preset domain label values, a first loss function is calculated according to the first domain discrimination result and the second domain discrimination result.
[0103] The CLS-marked target domain data features are input into the optimization predictor to predict the lifetime, and a second lifetime prediction value is obtained.
[0104] In the source domain data, the five data sets with the smallest difference from the second lifetime prediction value are selected to obtain the sample source domain data.
[0105] Weights are calculated based on the sample source domain data and the second lifetime prediction value.
[0106] Construct a second loss function based on the weights and the first loss function;
[0107] Based on the first loss function and the second loss function, the discriminator parameters are optimized to obtain an optimized discriminator.
[0108] In one feasible implementation, the source domain extractor is obtained by replicating the pre-trained extractor. and target domain extractor Input the training data from the source domain and the target domain respectively. and Obtain CLS-labeled source domain data features CLS marks target domain data features The process is as follows: (17) and (18):
[0109] (17);
[0110] (18);
[0111] Will and enter The domain discrimination result for each sample is obtained, and the process is as follows: (19) and (20):
[0112] (19);
[0113] (20);
[0114] The discriminator is trained using Binary Cross Entropy With Logits Loss (BCEWithLogitsLoss). The mathematical expression for BCEWithLogitsLoss is as follows (21):
[0115] (twenty one)
[0116] in, Indicates the real domain tag, This indicates the prediction result.
[0117] In this step, the domain label of the source domain data is set to 1, and the domain label of the target domain data is set to 0. Then the first loss function can be expressed as the following equation (22):
[0118] (twenty two);
[0119] in, and Representing the number of samples in the source domain and the number of samples in the target domain, respectively, the first domain discrimination result. Second domain discrimination result These represent the domain discrimination results for the source and target domain samples, respectively.
[0120] To further achieve relative alignment of features with similar RULs when aligning features between two domains, it is necessary to use the true RUL values of the source domain. RUL prediction pseudo-labels for the target domain Calculate the weights. Input predictor Chinese calculation It can be expressed as the following formula (23):
[0121] (twenty three);
[0122] Assigning the weights to the loss function shown in equation (22), we obtain the second loss function as shown in equation (24):
[0123] (twenty four)
[0124] From equation (24), we can see that, and The closer the samples are, the larger their weights will be, and the more the model will tend to train on these samples; that is, the model will tend to group these samples closer together in the feature space. and The larger the difference, the smaller the weight value, and the smaller the impact of the loss function of this group of samples on the training result. The magnitude of adjusting the model parameters according to this group of samples during training is also smaller. That is, the model is more likely to ignore this group of samples, and finally achieve the effect of aligning samples from different domains with similar RUL in the feature space.
[0125] To improve training efficiency, this invention employs a method for selecting source domain samples used in weight calculation: all samples in the training set are placed into a sample pool, and for each target domain sample, weights are calculated based on... Find the 5 source domain samples in the sample pool with the smallest difference between the label and the value, randomly select one to calculate the weighted loss function, and do not replace the label after calculation. When the number of samples in the pool is less than 5, refill it.
[0126] Because if you had chosen it before and If samples with large differences are fed into the model, their weights will become very small, meaning that the samples will not play a significant role in training the model. After screening, the weights of each group of samples will be relatively large, which can improve training efficiency.
[0127] Based on the above steps, it is finally used to train the discriminator. The loss function is as follows (25):
[0128] (25);
[0129] in, and This represents the balancing hyperparameter of the loss function, where epoch represents the current training epoch, and EPOCH represents the total number of training epochs. The hyperparameter for the number of starting rounds that control the effect of the weighted loss function.
[0130] Optionally, based on the optimized discriminator and reconstructor, the extractor is transferred to the target domain data for training to obtain an optimized extractor, including:
[0131] The target domain data is input into the extractor for feature extraction and CLS labels are added to obtain the labeled target domain data features and the unlabeled target domain data features.
[0132] The labeled target domain data features are input into the optimized discriminator for domain discrimination to obtain a third discrimination result;
[0133] Reverse the field labels of the third discrimination result to obtain the fourth discrimination result;
[0134] Based on the target domain data and the preset domain label values, the third loss function is calculated according to the fourth discrimination result;
[0135] Input the unlabeled target domain data features into the reconstructor to reconstruct the data and obtain the reconstructed target domain data.
[0136] Calculate the fourth loss function based on the target domain data and the reconstructed target domain data;
[0137] Based on the third and fourth loss functions, the parameters of the extractor are optimized to obtain an optimized extractor.
[0138] In one feasible implementation, a target domain feature extractor is trained. ,let Distribution as close as possible .
[0139] like If it is difficult to distinguish the domain label of a sample that actually comes from the target domain, then it is considered... The distribution is relatively close At this point, a pre-trained algorithm capable of predicting RUL based on source domain features is used. This will yield better RUL prediction results for the target domain data.
[0140] Input the target domain training data Obtain the features of the labeled target domain data Unlabeled target domain data features and target domain data As shown in equation (26):
[0141] (26);
[0142] Features enter The domain discrimination result is obtained as shown in equation (20).
[0143] Flipping the domain labels to achieve obfuscation The purpose is that the domain label of the target domain data is 1 at this time. Then, according to equation (21), the third loss function at this time is as follows (27):
[0144] (27);
[0145] To preserve the domain-specific information of the target domain as much as possible, a target domain mapping module is designed and a fourth loss function is calculated according to equations (14) and (15). Assisted model training. The loss function for the overall training process can be listed as shown in equation (28):
[0146] (28);
[0147] in, and This represents the balancing hyperparameter of the loss function.
[0148] S4. Obtain the data of the aero-engine to be predicted; input the data of the aero-engine to be predicted into the trained domain adaptive model to predict the RUL of the aero-engine and obtain the remaining service life of the aero-engine.
[0149] In one feasible implementation, this invention utilizes the NASA-C-MAPSS aero-engine simulation dataset. This dataset comprises four subsets, each with a varying number of operating conditions and fault states. Each subset includes a training set and a test set. The training set contains cyclic data spanning the entire lifecycle of the aero-engine, while the test set contains cyclic data spanning an incomplete lifecycle of the aero-engine. Data at each time step includes the engine number, cycle number, three operating parameters, and data from 21 sensors.
[0150] In this invention, 14 features from the original data were selected for verification. The verification results show that the invention can successfully align two different domains, and the labels predicted after domain alignment using this invention are mostly distributed within an error range of 30 from the true labels.
[0151] This invention addresses the shortcomings of existing research techniques by proposing a RUL prediction method for aero-engines based on multi-scale transformers and domain adaptation. This method extracts data features across multiple time scales and utilizes a Transformer to comprehensively consider information, improving the domain-adaptive loss function. It aligns the feature distributions of the source and target domains, ensuring that internal data are aligned with similar labels as much as possible, and uses selectable source domain labels to assist in correcting the alignment process. Furthermore, this invention ensures that target domain features retain as much domain-specific information as possible by restoring them to target domain data, thus improving the model's RUL prediction performance. This invention is a highly efficient and accurate RUL prediction method for aero-engines based on multi-scale transformers and domain adaptation.
[0152] Figure 2 This is a block diagram illustrating an aircraft engine RUL prediction device based on MSFormer and domain adaptation, according to an exemplary embodiment. The device is used in an aircraft engine RUL prediction method based on MSFormer and domain adaptation. (Refer to...) Figure 2 The device includes a data acquisition module 210, a model building module 220, a model training module 230, and a model application module 240. Wherein:
[0153] The data acquisition module 210 is used to acquire data from the first aero-engine; based on the operating conditions of the aero-engine, it preprocesses the data from the first aero-engine to obtain source domain data and target domain data;
[0154] Model building module 220 is used to build a domain adaptive model based on the MSFormer model structure and the NCE model structure, using domain adaptive technology.
[0155] The model training module 230 is used to train the domain adaptive model using source domain data and target domain data to obtain the trained domain adaptive model.
[0156] The model application module 240 is used to acquire the data of the aero-engine to be predicted; the data of the aero-engine to be predicted is input into the trained domain adaptive model to predict the RUL of the aero-engine and obtain the remaining service life of the aero-engine.
[0157] Optionally, the data acquisition module 210 is further used for:
[0158] Calculate the RUL tag based on the first aero-engine data; add the RUL tag to the first aero-engine data to obtain the second aero-engine data;
[0159] Based on the operating conditions of the aero-engine, the data of the second aero-engine is divided to obtain the first source domain data and the first target domain data.
[0160] The first source domain data and the first target domain data are normalized to obtain the second source domain data and the second target domain data.
[0161] Based on preset sliding window parameters, the sliding window method is used to divide the second source domain data and the second target domain data into samples to obtain source domain data and target domain data.
[0162] The domain-adaptive model includes an extractor, a predictor, a discriminator, and a reconstructor.
[0163] The extractor is built based on the MSFormer model structure; the predictor is built using a single fully connected layer; the discriminator is built using a three-layer fully connected layer; and the reconstructor is built based on the NCE model structure.
[0164] The extractor includes a source domain extractor and a target domain extractor; the structure of the extractor includes a patching module, a sample mapping module, an LSTM model, a Transformer encoder, and a target domain mapping module.
[0165] Optionally, the model training module 230 is further used for:
[0166] Based on the extractor, the predictor is pre-trained using source domain data to obtain an optimized predictor;
[0167] Based on the extractor and the optimized predictor, the discriminator is trained by transfer learning using source domain data and target domain data to obtain an optimized discriminator;
[0168] Based on the optimized discriminator and reconstructor, the extractor is trained by transfer learning according to the target domain data to obtain the optimized extractor.
[0169] Optionally, the model training module 230 is further used for:
[0170] The source domain data is input into the extractor for feature extraction to obtain the source domain data features;
[0171] The source domain data features are input into the predictor to predict the remaining useful life, and the first useful life prediction value is obtained.
[0172] Based on the source domain data, the loss function is calculated according to the first lifetime prediction value to obtain the mean squared error loss;
[0173] Based on the mean squared error loss, the predictor parameters are optimized to obtain an optimized predictor.
[0174] Optionally, the model training module 230 is further used for:
[0175] The source domain data and target domain data are input into the extractor for feature extraction and CLS labeling is added to obtain CLS-labeled source domain data features and CLS-labeled target domain data features;
[0176] The source domain data features and target domain data features of CLS-tagged data are input into the discriminator for domain discrimination to obtain the first domain discrimination result and the second domain discrimination result.
[0177] Based on the CLS-labeled source domain data features, the CLS-labeled target domain data features, and the preset domain label values, a first loss function is calculated according to the first domain discrimination result and the second domain discrimination result.
[0178] The CLS-marked target domain data features are input into the optimization predictor to predict the lifetime, and a second lifetime prediction value is obtained.
[0179] In the source domain data, the five data sets with the smallest difference from the second lifetime prediction value are selected to obtain the sample source domain data.
[0180] Weights are calculated based on the sample source domain data and the second lifetime prediction value.
[0181] Construct a second loss function based on the weights and the first loss function;
[0182] Based on the first loss function and the second loss function, the discriminator parameters are optimized to obtain an optimized discriminator.
[0183] Optionally, the model training module 230 is further used for:
[0184] The target domain data is input into the extractor for feature extraction and CLS labels are added to obtain the labeled target domain data features and the unlabeled target domain data features.
[0185] The labeled target domain data features are input into the optimized discriminator for domain discrimination to obtain a third discrimination result;
[0186] Reverse the field labels of the third discrimination result to obtain the fourth discrimination result;
[0187] Based on the target domain data and the preset domain label values, the third loss function is calculated according to the fourth discrimination result;
[0188] Input the unlabeled target domain data features into the reconstructor to reconstruct the data and obtain the reconstructed target domain data.
[0189] Calculate the fourth loss function based on the target domain data and the reconstructed target domain data;
[0190] Based on the third and fourth loss functions, the parameters of the extractor are optimized to obtain an optimized extractor.
[0191] This invention addresses the shortcomings of existing research techniques by proposing a RUL prediction method for aero-engines based on multi-scale transformers and domain adaptation. This method extracts data features across multiple time scales and utilizes a Transformer to comprehensively consider information, improving the domain-adaptive loss function. It aligns the feature distributions of the source and target domains, ensuring that internal data are aligned with similar labels as much as possible, and uses selectable source domain labels to assist in correcting the alignment process. Furthermore, this invention ensures that target domain features retain as much domain-specific information as possible by restoring them to target domain data, thus improving the model's RUL prediction performance. This invention is a highly efficient and accurate RUL prediction method for aero-engines based on multi-scale transformers and domain adaptation.
[0192] Figure 3 This is a schematic diagram of the structure of an aircraft engine RUL prediction device provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the aircraft engine RUL prediction device may include the above-mentioned Figure 2 The illustrated aero-engine RUL prediction device based on MSFormer and domain adaptation. Optionally, the aero-engine RUL prediction device 310 may include a first processor 2001.
[0193] Optionally, the aircraft engine RUL prediction device 310 may also include a memory 2002 and a transceiver 2003.
[0194] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.
[0195] The following is combined with Figure 3 A detailed introduction to each component of the aircraft engine RUL prediction device 310:
[0196] The first processor 2001 is the control center of the aero-engine RUL prediction device 310. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0197] Optionally, the first processor 2001 can perform various functions of the aircraft engine RUL prediction device 310 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0198] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 3 CPU0 and CPU1 are shown in the diagram.
[0199] In a specific implementation, as one example, the aircraft engine RUL prediction device 310 may also include multiple processors, for example... Figure 3 The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0200] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0201] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be connected via the interface circuit of the aircraft engine RUL prediction device 310. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0202] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0203] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 3 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0204] Alternatively, the transceiver 2003 can be integrated with the first processor 2001, or it can exist independently and be connected to the interface circuit of the aircraft engine RUL prediction device 310. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0205] It should be noted that, Figure 3 The structure of the aircraft engine RUL prediction device 310 shown does not constitute a limitation on the router. Actual knowledge structure identification devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0206] Furthermore, the technical effectiveness of the aero-engine RUL prediction device 310 can be referenced from the technical effectiveness of the aero-engine RUL prediction method based on MSFormer and domain adaptation described in the above method embodiments, and will not be repeated here.
[0207] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0208] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0209] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0210] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0211] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0212] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0213] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0214] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0215] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0216] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0217] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0218] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0219] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for predicting the RUL (Range Limit Indicator) of aero-engines based on MSFormer and domain adaptation, characterized in that, The method includes: Acquire data from the first aero-engine; based on the operating conditions of the aero-engine, preprocess the data from the first aero-engine to obtain source domain data and target domain data; Based on domain adaptation technology, a domain adaptation model is constructed according to the MSFormer model structure and the NCE model structure; The domain adaptive model includes an extractor, a predictor, a discriminator, and a reconstructor. The extractor is constructed based on the MSFormer model structure; the predictor is constructed using a single fully connected layer; the discriminator is constructed using a three-layer fully connected layer; and the reconstructor is constructed based on the NCE model structure. The extractor includes a source domain extractor and a target domain extractor; the structure of the extractor includes a Patching module, a sample mapping module, an LSTM model, a Transformer encoder, and a target domain mapping module. The source domain data and the target domain data are used to train the domain adaptive model to obtain the trained domain adaptive model. The step of training the domain adaptation model using the source domain data and the target domain data to obtain the trained domain adaptation model includes: Based on the extractor, the predictor is pre-trained according to the source domain data to obtain an optimized predictor; Based on the extractor and the optimized predictor, the discriminator is transferred to the source domain data and the target domain data to obtain an optimized discriminator. Based on the optimized discriminator and the reconstructor, the extractor is subjected to transfer training according to the target domain data to obtain an optimized extractor; Acquire the aero-engine data to be predicted; input the aero-engine data to be predicted into the trained domain adaptive model to perform aero-engine RUL prediction and obtain the remaining service life of the aero-engine.
2. The aero-engine RUL prediction method based on MSFormer and domain adaptation according to claim 1, characterized in that, The operating conditions based on the aero-engine involve preprocessing the first aero-engine data to obtain source domain data and target domain data, including: Calculate the RUL tag based on the first aero-engine data; add the RUL tag to the first aero-engine data to obtain the second aero-engine data; Based on the operating conditions of the aero-engine, the second aero-engine data is divided to obtain the first source domain data and the first target domain data; The first source domain data and the first target domain data are normalized to obtain the second source domain data and the second target domain data. Based on preset sliding window parameters, the sliding window method is used to divide the second source domain data and the second target domain data into samples to obtain source domain data and target domain data.
3. The aero-engine RUL prediction method based on MSFormer and domain adaptation according to claim 1, characterized in that, The step of pre-training the predictor based on the extractor and the source domain data to obtain an optimized predictor includes: The source domain data is input into the extractor for feature extraction to obtain the source domain data features; The source domain data features are input into the predictor to predict the remaining useful life, and a first useful life prediction value is obtained. Based on the source domain data, the loss function is calculated according to the first predicted lifetime value to obtain the mean squared error loss; Based on the mean squared error loss, the predictor parameters are optimized to obtain an optimized predictor.
4. The aero-engine RUL prediction method based on MSFormer and domain adaptation according to claim 1, characterized in that, The step of performing transfer learning on the discriminator based on the extractor and the optimized predictor, according to the source domain data and the target domain data, to obtain an optimized discriminator includes: The source domain data and the target domain data are input into the extractor for feature extraction and CLS tagging is added to obtain CLS-tagged source domain data features and CLS-tagged target domain data features. The source domain data features and target domain data features of the CLS marker are input into the discriminator for domain discrimination to obtain the first domain discrimination result and the second domain discrimination result; Based on the CLS-labeled source domain data features, the CLS-labeled target domain data features, and the preset domain label values, a first loss function is calculated according to the first domain discrimination result and the second domain discrimination result; The CLS-marked target domain data features are input into the optimization predictor to predict the lifetime, thereby obtaining a second lifetime prediction value. From the source domain data, the five sets of data with the smallest difference from the second lifespan prediction value are selected to obtain sample source domain data; Based on the sample source domain data, the weights are calculated according to the second predicted lifetime value; Construct a second loss function based on the weights and the first loss function; Based on the first loss function and the second loss function, the discriminator is optimized to obtain an optimized discriminator.
5. The aero-engine RUL prediction method based on MSFormer and domain adaptation according to claim 1, characterized in that, The step of performing transfer learning on the extractor based on the optimized discriminator and the reconstructor, according to the target domain data, to obtain an optimized extractor includes: The target domain data is input into the extractor for feature extraction and CLS labeling is added to obtain labeled target domain data features and unlabeled target domain data features. The labeled target domain data features are input into the optimized discriminator for domain discrimination to obtain a third discrimination result. The field labels of the third discrimination result are flipped to obtain the fourth discrimination result; Based on the target domain data and the preset domain label values, a third loss function is calculated according to the fourth discrimination result; The unlabeled target domain data features are input into the reconstructor to reconstruct the data and obtain the reconstructed target domain data. Calculate the fourth loss function based on the target domain data and the reconstructed target domain data; Based on the third loss function and the fourth loss function, the parameters of the extractor are optimized to obtain an optimized extractor.
6. A region-adaptive RUL prediction device for aero-engines, wherein the region-adaptive RUL prediction device for aero-engines is used to implement the region-adaptive RUL prediction method for aero-engines as described in any one of claims 1-5, characterized in that, The device includes: The data acquisition module is used to acquire data from the first aero-engine; based on the operating conditions of the aero-engine, the data from the first aero-engine is preprocessed to obtain source domain data and target domain data; The model building module is used to build a domain-adaptive model based on the MSFormer model structure and the NCE model structure, using domain adaptation technology. The model training module is used to train the domain adaptive model using the source domain data and the target domain data to obtain the trained domain adaptive model. The model application module is used to acquire the aero-engine data to be predicted; the aero-engine data to be predicted is input into the trained domain adaptive model to predict the aero-engine's RUL (Remaining Service Life) and obtain the remaining service life of the aero-engine.
7. An aircraft engine RUL prediction device, characterized in that, The aircraft engine RUL prediction device includes: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for predicting residual service life of aero-engine through time-frequency domain analysis
CN114492184A
Cross-subject electroencephalogram signal emotion recognition method
CN118633938A