Aero-engine gas path system fault diagnosis method based on cross-domain small sample learning
Through the cross-domain small sample learning method, combined with dynamic adaptive weight optimization and multi-level feature enhancement modules, the problems of data scarcity and imbalance in aircraft engine gas path system fault diagnosis are solved, and efficient fault diagnosis and good generalization performance are achieved.
Patent Information
- Application Number
- CN202510880487.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-23
AI Technical Summary
Existing small-sample learning methods have problems of algorithm overfitting and low fault diagnosis accuracy in aircraft engine gas path system fault diagnosis, and there are label offset and feature offset in transfer learning, especially poor performance under unbalanced datasets.
A cross-domain small sample learning method is adopted to obtain the gas path system data of aircraft engines under different operating conditions, perform balanced sampling and probabilistic data enhancement, and construct a multi-level feature enhancement module and diagnosis module. The dynamic adaptive weight optimization coefficient, LDAM loss function, contrast projection module and domain-aware contrast loss function are combined to optimize the feature extraction and diagnosis process.
It improves the accuracy and generalization ability of fault diagnosis, reduces the feature distribution differences and label distribution differences between the source domain and the target domain, enhances the model's adaptability to unbalanced samples, and improves the feature discrimination ability in complex noise environments.
Smart Images

Figure CN120687877A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aero-engine automatic control, and in particular to a method for diagnosing faults in an aero-engine gas path system based on cross-domain small sample learning. Background Art
[0002] With the advancement of aviation technology, the gas flow systems of aircraft engines, the core propulsion systems of aircraft, have become increasingly complex, and the types and modes of failures have become more diverse. The stability and reliability of aircraft engine gas flow systems are directly related to the safety and economic efficiency of aircraft. However, in practical applications, aircraft engine failures not only occur infrequently, making data difficult to collect, but also often have an unbalanced distribution, resulting in a limited number of failure samples that can be used for model training.
[0003] Intelligent fault diagnosis technology has developed rapidly in recent years because it does not require precise physical models. Among these methods, deep learning has demonstrated promising performance, as it not only automatically extracts discriminative fault feature representations without prior knowledge but also simultaneously optimizes both the fault feature extractor and the fault classifier. Many deep learning-based methods, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), have been successfully applied in fault diagnosis. However, they struggle with small sample sizes and imbalanced datasets. To address this problem, data augmentation, transfer learning, meta-learning, and multimodal learning are commonly used. Jiang et al. proposed a novel data augmentation classifier (DAC) to address the data imbalance problem in industrial fault classification. By combining a generative adversarial network (GAN) with a data selection strategy, they effectively improved fault classification performance. Ding et al. proposed a novel deep imbalanced domain adaptation (DIDA) framework to address feature and label shift in bearing fault diagnosis under multiple operating conditions. By optimizing classification boundaries through static weight learning and class alignment, the accuracy and generalization of fault diagnosis were significantly improved. Wang proposed a few-sample mechanical fault diagnosis framework based on self-supervised signal representation learning and improved twin networks. By utilizing a large amount of unlabeled data to learn the intrinsic characteristics of the signal and using a variety of data enhancement methods to expand the data, the fault diagnosis performance with a small number of labeled samples was significantly improved. The above scheme shows that expanding data and simplifying the model have a good effect in solving the small sample problem.
[0004] Small-sample learning, a method that can efficiently solve fault diagnosis problems using only a small number of samples, is gaining increasing attention. However, this method faces huge challenges in actual operation. For example, when the number of training samples is too small, the model may overfit the features in these small samples, resulting in a significant decrease in its generalization ability. This means that although the model can recognize patterns in the training data well, it performs poorly when faced with new, unseen data. In addition, how to effectively extract sufficient effective semantic information from limited data and build a robust model that can adapt to the needs of different tasks is also one of the key challenges in current FSL research. In existing small-sample learning, the problem of sample scarcity is often solved by using data augmentation methods to expand samples or migrating source domain data, but few methods combine the two. Large amounts of data augmentation will reduce the generalization ability of the model. Using a single data augmentation method will lead to low sample quality. Although generative adversarial networks and diffusion models can produce high-quality data, they have relatively high requirements for the number of original samples and model construction, and the model is also prone to collapse. Migrating source domain data will not only cause label offset and sample offset problems, but also be too idealistic in some cases. For example, in actual applications, the normal operation data of aircraft engines will be significantly more than the fault data. Therefore, the method of combining data augmentation with the migration of unbalanced data has greater practical significance. Summary of the Invention
[0005] The present invention provides an aero-engine gas path system fault diagnosis method based on cross-domain small sample learning, so as to overcome the problems of algorithm overfitting and low fault diagnosis accuracy caused by data scarcity in existing small sample learning, as well as the technical problems of label offset and feature offset caused by unbalanced source domain data in existing small sample transfer learning.
[0006] In order to achieve the above object, the technical solution of the present invention is:
[0007] A method for fault diagnosis of an aircraft engine gas path system based on cross-domain small sample learning includes:
[0008] S1: Acquire gas path system data under different operating states of an aircraft engine, wherein the gas path system data includes a source domain dataset and a target domain dataset, perform balanced sampling on the source domain dataset, perform probabilistic data enhancement on the target domain dataset, expand the number of samples, and obtain processed source domain data and target domain data;
[0009] S2: Construct a fault diagnosis module, which includes a multi-level feature enhancement module and a diagnosis module connected in sequence. The processed source domain data and target domain data are respectively input into the multi-level feature enhancement module for feature extraction to obtain source domain features and target domain features. The source domain features and target domain features are input into the diagnosis module to obtain a fault diagnosis result.
[0010] S3: construct a dynamic adaptive weight optimization coefficient, and construct an LDAM loss function based on the dynamic adaptive weight optimization coefficient. The LDAM loss function is used to calculate the loss between the source domain features and the true source domain labels.
[0011] S4: Construct a contrast projection module, input the source domain features and the target domain features into the contrast projection module for contrast projection, and obtain source domain projection features and target domain projection features;
[0012] S5: constructing an intra-domain contrast loss function and a cross-domain contrast loss function, and constructing a domain-aware contrast loss function based on the intra-domain contrast loss function and the cross-domain contrast loss function; the domain-aware contrast loss function is used to calculate the loss between the source domain projection feature and the target domain projection feature;
[0013] S6: Use the backpropagation algorithm to update the multi-level feature enhancement module based on the LDAM loss function and the domain awareness comparison loss function to obtain an updated multi-level feature enhancement module. Input the aircraft engine data to be diagnosed into the updated multi-level feature enhancement module and the diagnosis module to obtain the final fault diagnosis result.
[0014] Furthermore, the multi-level feature enhancement module includes a feature extraction module, a bottleneck layer, a lightweight attention mechanism, a feature enhancement module and an output layer connected in sequence;
[0015] The feature extraction module includes a convolutional neural network;
[0016] The lightweight attention mechanism includes a first linear layer, a first RELU activation function layer, a second linear layer and a Sigmoid activation function layer connected in sequence;
[0017] The feature enhancement module comprises a third linear layer, a first batch normalization layer and a second RELU activation function layer connected in sequence;
[0018] The feature extraction module is used to perform local feature extraction and pooling processing on the input data through a convolutional neural network to obtain high-dimensional local features;
[0019] The bottleneck layer is used to reduce or compress the high-dimensional local features output by the feature extraction module to obtain compressed features;
[0020] The lightweight attention mechanism is used to dynamically adjust the response weight of the feature dimension, weight the compressed features by channel importance, and obtain attention-weighted features;
[0021] The feature enhancement module is used to perform nonlinear transformation and normalization on the attention weighted features to obtain enhanced high-order features.
[0022] Furthermore, the diagnosis module includes a fourth linear layer, which is used to calculate the scores of different fault categories of the source domain features and the target domain features to obtain the fault classification result of each feature, that is, the fault diagnosis result.
[0023] Furthermore, the contrast projection module includes a fifth linear layer, a third RELU activation function layer, a sixth linear layer, and a second batch normalization layer connected in sequence;
[0024] The fifth linear layer is used to perform linear transformation on the input source domain features and target domain features respectively, and output high-dimensional features of the source domain features and the target domain features;
[0025] The third RELU activation function layer is used to activate the high-dimensional features of the input source domain features and the target domain features respectively, and output enhanced features of the source domain features and the target domain features;
[0026] The sixth linear layer is used to perform linear transformation on the input source domain features and the enhanced features of the target domain features, compress the feature dimensions, and output low-dimensional embedding representations of the source domain features and the target domain features;
[0027] The second batch normalization layer is used to normalize the low-dimensional embedding representations of the input source domain features and target domain features, respectively, and output source domain projection features and target domain projection features.
[0028] Furthermore, a dynamic adaptive weight optimization coefficient is constructed, including:
[0029] S31. Design a performance adjustment factor based on the modified Sigmoid function, as shown in formula (1).
[0030]
[0031] Among them, α1 is the weight coefficient of the performance adjustment factor, k1, T perf is a hyperparameter, P gap Represents the performance gap relative to the global performance, as shown in formula (2),
[0032]
[0033] Among them, μ Acc Represents the historical accuracy mean, μ Acc_globalrepresents the global performance benchmark;
[0034] S32, design stability factor, as shown in formula (3),
[0035]
[0036] Among them, α2 represents the weight coefficient of the stability factor, k2 represents the hyperparameter, σ Acc represents the standard deviation of category accuracy;
[0037] S33, constructing a dynamic adaptive weight optimization coefficient based on the performance adjustment factor and the stability factor, as shown in formula (4),
[0038] F adj =α2+(F perf -α2)·F stab (4)
[0039] Among them F adj Limited to preset range Inside;
[0040] S34, constructing the dynamic weight of the fault category based on the dynamic adaptive weight optimization coefficient, as shown in formula (5),
[0041] w dyn =w base ·F′ adj (5)
[0042] Among them, w dyn represents the dynamic weight, w base Represents the base weight.
[0043] Furthermore, the LDAM loss function is constructed based on the dynamic adaptive weight optimization coefficient, including:
[0044] S34, constructing the LDAM loss function based on the dynamic adaptive weight optimization coefficient, as shown in formula (6),
[0045]
[0046] Where B represents the batch size, C represents the total number of categories, Indicates that the source domain feature corresponds to the real source domain category y l The original logit value calculated above is the logit value after LDAM boundary adjustment and scaling, z i,a Represents the final logit value of the source domain feature on the ath fault category.
[0047] Furthermore, we construct an intra-domain contrast loss function and a cross-domain contrast loss function, and then construct a domain-aware contrast loss function based on the intra-domain contrast loss function and the cross-domain contrast loss function, including:
[0048] S51. Construct the intra-domain contrast loss function, as shown in formula (7):
[0049]
[0050] Where N is the number of samples in each batch, f i is the feature vector of the sample in the corresponding data domain, f j is the feature vector of the positive sample in the sample, |P(i)| represents the total number of samples in the set, f k is the feature vector of any sample k in a batch, τ is the temperature coefficient;
[0051] S52. Construct a cross-domain contrast loss function, as shown in formula (8):
[0052]
[0053] Among them, P ST (n) = {m | l S,n =l T,m} is the index set of target domain data samples m with the same label as source domain data sample n, |P ST (n)| represents the set P ST The total number of samples in (n), f S,n Represents the feature vector of source domain data sample n, f T,m Represents the feature vector of the target domain data sample m, f T,r Represents the feature vector of the target domain data sample r;
[0054] S53. Based on the intra-domain contrast loss function and the cross-domain contrast loss function, a domain-aware contrast loss function is constructed, as shown in formula (9):
[0055]
[0056] Among them, γ represents the weight coefficient, represents the intra-domain contrast loss function of the source domain, represents the intra-domain contrastive loss function of the target domain.
[0057] Furthermore, the back-propagation algorithm is used to update the multi-level feature enhancement module based on the LDAM loss function and the domain awareness contrast loss function to obtain an updated multi-level feature enhancement module. The aircraft engine data to be diagnosed is input into the updated multi-level feature enhancement module and the diagnosis module to obtain the final fault diagnosis results, including:
[0058] S61. Construct a total loss function based on the LDAM loss function and the domain awareness contrast loss function, as shown in formula (10).
[0059]
[0060] Among them, λ is the weight coefficient;
[0061] S62. Calculate the total loss function value of the source domain features and the target domain features, calculate the gradient of the total loss function based on the back propagation algorithm, and use the Adam optimizer to update and optimize the parameters in the multi-level feature enhancement module and the contrast projection module based on the gradient to obtain an optimized multi-level feature enhancement module and an optimized contrast projection module;
[0062] S63. Use the optimized contrast projection module to perform contrast projection on the source domain features and the target domain features to obtain new source domain projection features and new target domain projection features. During the training process, the multi-level feature enhancement module and the contrast projection module are continuously iterated to update and optimize until the total loss function value converges, thereby obtaining the final multi-level feature enhancement module.
[0063] S64: Input the aircraft engine data to be diagnosed into the optimized multi-level feature enhancement module and the diagnosis module to obtain the final fault diagnosis result.
[0064] Beneficial effects: The present invention provides an aero-engine gas path system fault diagnosis method based on cross-domain small sample learning, which has the following advantages:
[0065] 1. A dynamic adaptive weight optimization mechanism is proposed. With the help of adaptive adjustment factors and stability factors, the weight updates are stable and timely. It can reweight the imbalanced samples and pay more attention to the sample data with difficult-to-distinguish sample features and unclear boundaries during prediction.
[0066] 2. Constructing a domain-aware comparative loss function based on intra-domain loss function and cross-domain loss function to alleviate the problem of feature distribution differences and label distribution differences between the source domain and the target domain;
[0067] 3. A multi-level feature enhancement module is proposed, which combines the improved convolutional neural network, bottleneck layer and lightweight attention mechanism to better mine the potential features of the data and has good generalization performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0069] Figure 1 A flow chart of the method for diagnosing faults in an aero-engine gas path system based on cross-domain small sample learning provided by the present invention;
[0070] Figure 2 This is a structural diagram of the aero-engine gas path system fault diagnosis method based on cross-domain small sample learning provided by the present invention;
[0071] Figure 3 This is a comparison chart before and after data enhancement for the probabilistic data enhancement method;
[0072] Figure 4 A structural diagram of the multi-level feature enhancement module provided by the present invention;
[0073] Figure 5 A structural diagram of the contrast projection module provided by the present invention;
[0074] Figure 6 This is a comparison chart of the probabilistic data enhancement results of an embodiment of the present invention;
[0075] Figure 7 Graph showing experimental results of the provided embodiment on CD-CNN;
[0076] Figure 8 Graph showing experimental results of the provided embodiment on LANMSFF;
[0077] Figure 9 Graph showing experimental results of the provided embodiment on MLFEM;
[0078] Figure 10 The figure is a diagram of experimental results of applying the method provided by the present invention to the provided embodiment. DETAILED DESCRIPTION
[0079] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0080] This embodiment provides an aero-engine gas path system fault diagnosis method based on cross-domain small sample learning, such as Figure 1 As shown, including:
[0081] S1: Acquire gas path system data under different operating states of an aircraft engine, wherein the gas path system data includes a source domain dataset and a target domain dataset, perform balanced sampling on the source domain dataset, perform probabilistic data enhancement on the target domain dataset, expand the number of samples, and obtain processed source domain data and target domain data;
[0082] S2: Construct a fault diagnosis module, which includes a multi-level feature enhancement module and a diagnosis module connected in sequence. The processed source domain data and target domain data are respectively input into the multi-level feature enhancement module for feature extraction to obtain source domain features and target domain features. The source domain features and target domain features are input into the diagnosis module to obtain a fault diagnosis result.
[0083] S3: construct a dynamic adaptive weight optimization coefficient, and construct an LDAM loss function based on the dynamic adaptive weight optimization coefficient. The LDAM loss function is used to calculate the loss between the source domain features and the true source domain labels.
[0084] S4: Construct a contrast projection module, input the source domain features and the target domain features into the contrast projection module for contrast projection, and obtain source domain projection features and target domain projection features;
[0085] S5: constructing an intra-domain contrast loss function and a cross-domain contrast loss function, and constructing a domain-aware contrast loss function based on the intra-domain contrast loss function and the cross-domain contrast loss function; the domain-aware contrast loss function is used to calculate the loss between the source domain projection feature and the target domain projection feature;
[0086] S6: Use the backpropagation algorithm to optimize the multi-level feature enhancement module based on the LDAM loss function and the domain awareness contrast loss function to obtain the optimized multi-level feature enhancement module. Input the aircraft engine data to be diagnosed into the optimized multi-level feature enhancement module and the diagnosis module to obtain the final fault diagnosis result.
[0087] Specifically, if Figure 2 As shown, first, the air path system data of the aircraft engine under different operating states is obtained. The air path system data includes a source domain data set and a target domain data set. Balanced sampling is performed on the source domain data set, and probabilistic data enhancement is performed on the target domain data set to expand the number of samples to obtain processed source domain data and target domain data. The probabilistic data enhancement method is used to expand the number of samples while ensuring data diversity and improving the generalization ability of the model.
[0088] Secondly, a fault diagnosis module is constructed, which includes a multi-level feature enhancement module and a diagnosis module connected in sequence. The processed source domain data and target domain data are respectively input into the multi-level feature enhancement module for feature extraction to obtain source domain features and target domain features. The source domain features and target domain features are input into the diagnosis module to obtain fault diagnosis results. The improved convolutional neural network, bottleneck layer and attention mechanism are combined. The convolutional neural network is used to capture local features of the data, and convolution kernels at different levels capture features from fine-grained to coarse-grained; the bottleneck layer is mainly used to remove redundant information, retain key features, and improve the generalization ability of the model; the lightweight attention mechanism enhances the model's sensitivity to key faults and improves the model's feature discrimination ability in complex noise environments. This module can better mine the potential features of the data and has good generalization performance;
[0089] Thirdly, a dynamic adaptive weight optimization coefficient is constructed. Based on the dynamic adaptive weight optimization coefficient, an LDAM loss function is constructed. Performance adjustment factors and stability factors are designed to form a dynamic adaptive weight optimization coefficient to ensure stable and timely weight updates. Unbalanced samples are reweighted, so that the fault diagnosis process pays more attention to samples with difficult-to-distinguish sample features and unclear boundaries, thereby improving the accuracy of diagnosis. The LDAM loss function constructed based on the dynamic adaptive weight optimization coefficient can make the fault diagnosis process pay more attention to the tail data in the long-tail source domain data set, so that the network learns features that are fairer and more balanced for all categories.
[0090] Construct a contrast projection module, input the source domain features and the target domain features into the contrast projection module for contrast projection, obtain the source domain projection features and the target domain projection features, and construct an intra-domain contrast loss function and a cross-domain contrast loss function. Based on the intra-domain contrast loss function and the cross-domain contrast loss function, a domain-aware contrast loss function is constructed to calculate the loss function of the source domain projection features and the target domain projection features; construct a domain-aware contrast loss function based on the intra-domain loss function and the cross-domain loss function, which alleviates the problem of feature distribution differences and label distribution differences between the source domain and the target domain;
[0091] Finally, the back-propagation algorithm is used to optimize the multi-level feature enhancement module based on the LDAM loss function and the domain awareness contrast loss function to obtain the optimized multi-level feature enhancement module. The aircraft engine data to be diagnosed is input into the optimized multi-level feature enhancement module and the diagnosis module to obtain the final fault diagnosis result.
[0092] In a specific embodiment, a scheme for obtaining gas path system data of an aircraft engine under different operating states, wherein the gas path system data includes a source domain data set and a target domain data set, performing balanced sampling on the source domain data set, performing probabilistic data enhancement on the target domain data set, and expanding the number of samples to obtain processed source domain data and target domain data is as follows:
[0093] S11. Obtain the air path system data of the aircraft engine under different operating conditions to form a source domain data set and target domain dataset
[0094] S12. Perform balanced sampling on the source domain data set to obtain sampled source domain data;
[0095] S13, perform probabilistic data enhancement on the single data sample in the target domain data set to expand the number of samples. In this embodiment, four transformation methods are selected to transform the data: random amplitude scaling, adding Gaussian noise, time jittering, and frequency masking. Each transformation method T g The data D after the previous transformation is transformed with a preset probability. (g-1) Perform another transformation, as shown in formula (11),
[0096]
[0097] D (g) Represents the probability enhanced data, P aug Represents the preset probability, P aug ∈[0.5,1];
[0098] The sampled source domain data is used to train the multi-level feature enhancement module and the contrast projection module. Five data are selected from each category in the target domain dataset, and the selected data are probability enhanced to obtain the probability enhanced data. This data is combined with the sampled source domain data to form a dataset for training the multi-level feature enhancement module and the contrast projection module. Each category in the target domain includes 400 sample data, and all the data in the target domain dataset are used to test the multi-level feature enhancement module and the contrast projection module.
[0099] In this solution, the probabilistic data enhancement method is used to expand the number of samples while ensuring the diversity of the data and improving the generalization ability of the model. The results of data enhancement using four transformation methods: random amplitude scaling, adding Gaussian noise, time jittering and frequency masking are shown. Figure 3 As shown in the figure, it can be seen that the enhanced signal graph has a wider range than the original signal graph, and the number of samples has been expanded.
[0100] In a specific embodiment, a fault diagnosis module is constructed, which includes a multi-level feature enhancement module and a diagnosis module connected in sequence. The processed source domain data and target domain data are respectively input into the multi-level feature enhancement module for feature extraction to obtain source domain features and target domain features. The source domain features and target domain features are input into the diagnosis module to obtain the fault diagnosis result scheme:
[0101] like Figure 4As shown, the multi-level feature enhancement module includes a feature extraction module, a bottleneck layer, a lightweight attention mechanism, a feature enhancement module and an output layer connected in sequence;
[0102] The feature extraction module includes a convolutional neural network;
[0103] The feature extraction module in this solution uses a multi-layer convolutional neural network structure, consisting of several sequentially connected convolutional units. Each convolutional unit includes a convolution layer, a normalization layer, a nonlinear activation function layer, and a pooling operation. The convolution layer uses a rectangular convolution kernel structure to enhance the model's feature extraction capabilities in specific directions. The module input is a single-channel signal and outputs a fixed-length high-dimensional feature vector.
[0104] In this scheme, the convolutional neural network includes the first, second and third convolution modules and the maximum pooling layer connected in sequence;
[0105] The first convolution module and the second convolution module each include a convolution layer, a batch normalization layer, a RELU activation function layer, and a maximum pooling layer connected in sequence;
[0106] The third convolution module contains a sequentially connected convolution layer, batch normalization layer, RELU activation function layer, adaptive global average pooling layer, flattening layer, dropout layer and fully connected layer;
[0107] The convolution layer of the first convolution module has 1 input channel and 32 output channels, and uses a rectangular convolution kernel. The size of the rectangular convolution kernel can be m×n, where m and n are unequal positive integers and can be set according to the input data structure to achieve feature perception capabilities in different directions.
[0108] The convolution layer of the second convolution module has 32 input channels and 64 output channels;
[0109] The convolution layer of the third convolution module has 64 input channels and 128 output channels, the Dropout layer has an inactivation probability of 0.5, and the fully connected layer has an input dimension of 128 and an output dimension of 4;
[0110] The lightweight attention mechanism includes a first linear layer, a first RELU activation function layer, a second linear layer and a Sigmoid activation function layer connected in sequence;
[0111] In this scheme, the dimension of the output feature of the first linear layer is 512, and the dimension of the output feature of the second linear layer is 256;
[0112] The bottleneck layer in this scheme includes a linear layer, a batch normalization layer, and a Dropout layer connected in sequence;
[0113] The linear layer input dimension is 4, the output dimension is set according to the mode, and the inactivation probability of the Dropout layer is 0.4;
[0114] The feature enhancement module includes a third linear layer, a batch normalization layer and a second RELU activation function layer connected in sequence;
[0115] The dimension of the output feature of the third linear layer is 512;
[0116] The feature extraction module is used to perform local feature extraction and pooling processing on the input data through a convolutional neural network to obtain high-dimensional local features;
[0117] The bottleneck layer is used to reduce or compress the high-dimensional local features output by the feature extraction module to obtain compressed features;
[0118] Specifically, the bottleneck layer is used to reduce or compress the high-dimensional vector output by the feature extraction module to improve the efficiency and generalization ability of the model. The bottleneck layer includes a linear transformation layer and normalization processing, combined with a regularization mechanism to avoid overfitting;
[0119] The lightweight attention mechanism is used to dynamically adjust the response weight of the feature dimension, weight the compressed features by channel importance, and obtain attention-weighted features;
[0120] Specifically, this module is used to dynamically adjust the response weights of feature dimensions, highlight key information areas, and improve the model's ability to model the distribution of semantic importance;
[0121] The feature enhancement module is used to perform nonlinear transformation and normalization on the attention weighted features to obtain enhanced high-order features;
[0122] Specifically, the module includes at least one set of linear transformation and nonlinear activation structures, with consistent input and output dimensions, to further enhance the expressiveness of bottleneck features and improve classification performance and stability.
[0123] The diagnosis module includes a fourth linear layer, which is used to calculate the scores of different fault categories of the source domain features and the target domain features to obtain the fault classification result of each feature, that is, the fault diagnosis result;
[0124] The dimension of the input feature of the fourth linear layer is 512, and the dimension of the output result is 4, representing the four types of faults.
[0125] This solution proposes a multi-level feature enhancement module that combines an improved convolutional neural network, a bottleneck layer, and an attention mechanism. The convolutional neural network is used to capture local features of the data, and convolution kernels at different levels capture features from fine-grained to coarse-grained. The bottleneck layer is mainly used to remove redundant information, retain key features, and improve the generalization ability of the model. The lightweight attention mechanism enhances the model's sensitivity to key faults and improves the model's feature discrimination ability in complex noisy environments. This module can better mine the potential features of the data and has good generalization performance.
[0126] In a specific embodiment, a dynamic adaptive weight optimization coefficient is constructed, and an LDAM loss function is constructed based on the dynamic adaptive weight optimization coefficient. The LDAM loss function is used to calculate the loss between the source domain features and the true source domain labels. The scheme is:
[0127] S31. In order to smoothly adjust the weight according to the performance gap, a performance adjustment factor is designed based on the modified Sigmoid function, as shown in formula (12).
[0128]
[0129] Among them, α1 is the weight coefficient of the performance adjustment factor, which is 2.0 in this embodiment, k1, T perf is a hyperparameter, k1=5, T perf =0.85, this factor makes the performance below the threshold T perf The category with performance below the threshold can obtain a gain greater than 1, and the category with performance below the threshold can obtain a gain less than 1. gap Represents the performance gap relative to the global performance, as shown in formula (13),
[0130]
[0131] Among them, μ Acc Represents the historical accuracy mean, μ Acc_global represents the global performance benchmark;
[0132] S32. In order to avoid over-adjustment of categories with drastic performance fluctuations, a stability factor is designed, as shown in formula (14):
[0133]
[0134] Among them, α2 is the weight coefficient of the stability factor, which is set to 1.0 in this embodiment, k2 is a hyperparameter, and σ Acc represents the standard deviation of category accuracy;
[0135] S33, constructing a dynamic adaptive weight optimization coefficient based on the performance adjustment factor and the stability factor, as shown in formula (15),
[0136] F adj =α2+(F perf -α2)·F stab (15)
[0137] Among them F adj Limited to preset range Inside;
[0138] S34, constructing the dynamic weight of the fault category based on the dynamic adaptive weight optimization coefficient, as shown in formula (16),
[0139] w dyn =w base ·F′ adj (16)
[0140] Among them, w dyn represents the dynamic weight, w base represents the basic weight;
[0141] S34, constructing the LDAM loss function based on the dynamic adaptive weight optimization coefficient, as shown in formula (17),
[0142]
[0143] Where B represents the batch size, C represents the total number of categories, Indicates that the source domain feature corresponds to the real source domain category y l The original logit value calculated above is the logit value after LDAM boundary adjustment and scaling, z i,a Represents the final logit value of the source domain feature on the ath fault category.
[0144] In this scheme, performance adjustment factors and stability factors are designed, and dynamic adaptive weight optimization coefficients are formed to ensure stable and timely weight updates. Unbalanced samples are reweighted, so that the fault diagnosis process pays more attention to samples with difficult-to-distinguish sample features and unclear boundaries, thereby improving the accuracy of diagnosis. The LDAM loss function is constructed based on the dynamic adaptive weight optimization coefficient, which can make the fault diagnosis process pay more attention to the tail data in the long-tail source domain dataset, so that the model can learn features that are fairer and more balanced for all categories.
[0145] In a specific embodiment, a comparative projection module is constructed, and source domain features and target domain features are input into the comparative projection module for comparative projection. The scheme for obtaining source domain projection features and target domain projection features is:
[0146] like Figure 5 As shown, the contrast projection module includes a fifth linear layer, a third RELU activation function layer, a sixth linear layer and a second batch normalization layer connected in sequence;
[0147] The fifth linear layer is used to perform linear transformation on the input source domain features and target domain features respectively, and output high-dimensional features of the source domain features and the target domain features;
[0148] The third RELU activation function layer is used to activate the high-dimensional features of the input source domain features and the target domain features respectively, and output enhanced features of the source domain features and the target domain features;
[0149] The sixth linear layer is used to perform linear transformation on the input source domain features and the enhanced features of the target domain features, compress the feature dimensions, and output low-dimensional embedding representations of the source domain features and the target domain features;
[0150] The second batch normalization layer is used to normalize the low-dimensional embedding representations of the input source domain features and target domain features, respectively, and output source domain projection features and target domain projection features.
[0151] In this scheme, the feature representation is mapped to a latent space dedicated to calculating the contrast loss to obtain the projection feature. The contrast projection module includes a multi-layer nonlinear mapping structure, which is used to project the intermediate feature vector into a low-dimensional embedding space. It is suitable for contrast loss optimization to improve the discrimination and consistency between different samples. This module can achieve a dual goal: it can effectively narrow the distance between the source domain and the target domain by comparing in the projection space, and it can promote the aggregation of similar categories and the separation of heterogeneous categories by retaining rich classification information in the feature representation space.
[0152] In a specific embodiment, an intra-domain contrast loss function and a cross-domain contrast loss function are constructed, and a domain-aware contrast loss function is constructed based on the intra-domain contrast loss function and the cross-domain contrast loss function; the domain-aware contrast loss function is used to calculate the loss between the source domain projection feature and the target domain projection feature. The scheme is:
[0153] S51. Construct the intra-domain contrast loss function, as shown in formula (18):
[0154]
[0155] Where N is the number of samples in each batch, f i is the feature vector of the sample in the corresponding data domain, f j is the feature vector of the positive sample in the sample, |P(i)| represents the total number of samples in the set, f k is the feature vector of any sample k in a batch, τ is the temperature coefficient; the positive sample in the sample refers to the sample of the same category as the anchor sample;
[0156] S52. Construct a cross-domain contrast loss function, as shown in formula (19):
[0157]
[0158] Among them, P ST (n) = {m | l S,n =l T,m} is the index set of target domain data samples m with the same label as source domain data sample n, |P ST (n)| represents the set P ST The total number of samples in (n), f S,n Represents the feature vector of source domain data sample n, f T,m Represents the feature vector of the target domain data sample m, f T,r Represents the feature vector of the target domain data sample r;
[0159] S53. Based on the intra-domain contrast loss function and the cross-domain contrast loss function, a domain-aware contrast loss function is constructed, as shown in formula (20).
[0160]
[0161] Among them, γ represents the weight coefficient, represents the intra-domain contrast loss function of the source domain, represents the contrast loss function within the target domain.
[0162] In this scheme, a domain-aware comparative loss function based on the intra-domain loss function and the cross-domain loss function is constructed to alleviate the problem of feature distribution differences and label distribution differences between the source domain and the target domain.
[0163] In a specific embodiment, a back-propagation algorithm is used to optimize the multi-level feature enhancement module based on the LDAM loss function and the domain awareness contrast loss function to obtain an optimized multi-level feature enhancement module. The aircraft engine data to be diagnosed is input into the optimized multi-level feature enhancement module and the diagnosis module to obtain the final fault diagnosis result. The scheme is:
[0164] S62. Calculate the total loss function value of the source domain features and the target domain features, calculate the gradient of the total loss function based on the back propagation algorithm, and use the Adam optimizer to update and optimize the parameters in the multi-level feature enhancement module and the contrast projection module based on the gradient to obtain an optimized multi-level feature enhancement module and an optimized contrast projection module;
[0165] S63. Use the optimized contrast projection module to perform contrast projection on the source domain features and the target domain features to obtain new source domain projection features and new target domain projection features. During the training process, the multi-level feature enhancement module and the contrast projection module are continuously iterated to update and optimize until the total loss function value converges, thereby obtaining the final multi-level feature enhancement module.
[0166] S64: Input the aircraft engine data to be diagnosed into the optimized multi-level feature enhancement module and the diagnosis module to obtain the final fault diagnosis result.
[0167] Example 1:
[0168] The actual operating data of aircraft engines usually contains a large amount of sensitive information, such as engine performance parameters, failure modes, maintenance records, etc. This information has important commercial value for aircraft engine manufacturers and operators and is difficult to obtain. Therefore, a laboratory-based aircraft engine digital health management and fault diagnosis twin platform is used to simulate and output data of different components of the JT9D aircraft engine air path system under various operating conditions. Data of three fault states, namely normal operating state and fan and low-pressure compressor, are selected. The JT9D air path system data output by the simulation platform is divided into two different data sets. The training set samples are shown in Table 1. The number of samples of each of the four labels in the target domain is 5, the number of samples of source domain label 1 is 240, and the number of samples of labels 1-3 is 36 each:
[0169] Table 1 Training set samples
[0170]
[0171]
[0172] The specific situation of the test set is shown in Table 2, with 400 samples for each label:
[0173] Table 2 Samples of the target domain test set
[0174] Failure Mode Label Sample size Efficiency decline Normal (healthy) 0 400 0 Fan02 1 400 0.2 Fan03 2 400 0.3 LPC02 3 400 0.2
[0175] Experimental results:
[0176] Probabilistic data enhancement results are as follows Figure 6 As shown, there are some obvious differences between different samples; in the time domain, the enhanced signal roughly follows the original signal with slight adjustments. In the frequency domain, the signal spectra before and after enhancement are also very similar. By comparing the distributions, the two distributions are almost identical and both are close to the normal distribution. The box plot shows that the enhanced data is very similar to the original data box plot, the medians are almost at the same level, the interquartile ranges are of similar size, and the statistical indicators of the data before and after enhancement, the mean, standard deviation, maximum and minimum values are very close. Therefore, the data generated by the probabilistic data enhancement technology proposed in the present invention is diversified and retains the original characteristics, which can help the model better generalize to unseen data.
[0177] The proposed method is compared with four other methods, namely, cross-domain convolutional neural network (CD-CNN), lightweight attention network (LANMSFF), and multi-level feature module (MLFEM):
[0178] CD-CNN uses the basic convolutional neural network module to extract features, and the results are as follows Figure 7 As shown, the overall accuracy is 84.5%, of which the accuracy of difficult sample label 2 (Fan03) is 45.0%;
[0179] The LANMSFF method mainly changes the attention mechanism, and the results are as follows Figure 8 As shown, the overall accuracy is 85.25%, and the accuracy of difficult sample label 3 is 51%;
[0180] MLFEM uses the same feature extraction module as the present invention, and the results are as follows: Figure 9 As shown in the figure, using the basic LDAM rebalancing module and data enhancement module, the overall accuracy is 88.75%, of which the accuracy for difficult sample label 2 is 58.2%;
[0181] Figure 10 The experimental results of the method proposed in this invention show an overall accuracy of 95%, and the accuracy of difficult sample label 3 is 84.8%, which is significantly higher than the diagnostic results of several other methods. This shows that the dynamic optimization mechanism of the present invention can increase the focus on difficult data while paying attention to a smaller number of samples. Probabilistic data enhancement can better increase the generalization ability of the model.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for fault diagnosis of an aircraft engine gas path system based on cross-domain small sample learning, characterized in that: include: S1: Acquire gas path system data under different operating states of an aircraft engine, wherein the gas path system data includes a source domain dataset and a target domain dataset, perform balanced sampling on the source domain dataset, perform probabilistic data enhancement on the target domain dataset, expand the number of samples, and obtain processed source domain data and target domain data; S2: Construct a fault diagnosis module, which includes a multi-level feature enhancement module and a diagnosis module connected in sequence. The processed source domain data and target domain data are respectively input into the multi-level feature enhancement module for feature extraction to obtain source domain features and target domain features. The source domain features and target domain features are input into the diagnosis module to obtain a fault diagnosis result. S3: construct a dynamic adaptive weight optimization coefficient, and construct an LDAM loss function based on the dynamic adaptive weight optimization coefficient. The LDAM loss function is used to calculate the loss between the source domain features and the true source domain labels. S4: Construct a contrast projection module, input the source domain features and the target domain features into the contrast projection module for contrast projection, and obtain source domain projection features and target domain projection features; S5: constructing an intra-domain contrast loss function and a cross-domain contrast loss function, and constructing a domain-aware contrast loss function based on the intra-domain contrast loss function and the cross-domain contrast loss function; the domain-aware contrast loss function is used to calculate the loss between the source domain projection feature and the target domain projection feature; S6: Use the backpropagation algorithm to optimize the multi-level feature enhancement module based on the LDAM loss function and the domain awareness contrast loss function to obtain the optimized multi-level feature enhancement module. Input the aircraft engine data to be diagnosed into the optimized multi-level feature enhancement module and the diagnosis module to obtain the final fault diagnosis result.
2. The method for diagnosing faults in an aero-engine gas path system based on cross-domain small sample learning according to claim 1, characterized in that: The multi-level feature enhancement module includes a feature extraction module, a bottleneck layer, a lightweight attention mechanism, a feature enhancement module and an output layer connected in sequence; The feature extraction module includes a convolutional neural network; The lightweight attention mechanism includes a first linear layer, a first RELU activation function layer, a second linear layer and a Sigmoid activation function layer connected in sequence; The feature enhancement module comprises a third linear layer, a first batch normalization layer and a second RELU activation function layer connected in sequence; The feature extraction module is used to perform local feature extraction and pooling processing on the input data through a convolutional neural network to obtain high-dimensional local features; The bottleneck layer is used to reduce or compress the high-dimensional local features output by the feature extraction module to obtain compressed features; The lightweight attention mechanism is used to dynamically adjust the response weight of the feature dimension, weight the compressed features by channel importance, and obtain attention-weighted features; The feature enhancement module is used to perform nonlinear transformation and normalization on the attention weighted features to obtain enhanced high-order features.
3. The method for diagnosing faults in an aero-engine gas path system based on cross-domain small sample learning according to claim 2, characterized in that: The diagnosis module includes a fourth linear layer, which is used to calculate the scores of different fault categories of the source domain features and the target domain features to obtain the fault classification result of each feature, that is, the fault diagnosis result.
4. The method for diagnosing faults in an aero-engine gas path system based on cross-domain small sample learning according to claim 3, characterized in that: The contrast projection module includes a fifth linear layer, a third RELU activation function layer, a sixth linear layer and a second batch normalization layer connected in sequence; The fifth linear layer is used to perform linear transformation on the input source domain features and target domain features respectively, and output high-dimensional features of the source domain features and the target domain features; The third RELU activation function layer is used to activate the high-dimensional features of the input source domain features and the target domain features respectively, and output enhanced features of the source domain features and the target domain features; The sixth linear layer is used to perform linear transformation on the input source domain features and the enhanced features of the target domain features, compress the feature dimensions, and output low-dimensional embedding representations of the source domain features and the target domain features; The second batch normalization layer is used to normalize the low-dimensional embedding representations of the input source domain features and target domain features, respectively, and output source domain projection features and target domain projection features.
5. The method for diagnosing faults in an aero-engine gas path system based on cross-domain small sample learning according to claim 3, characterized in that: Construct dynamic adaptive weight optimization coefficients, including: S31. Design a performance adjustment factor based on the modified Sigmoid function, as shown in formula (1). Among them, α1 is the weight coefficient of the performance adjustment factor, k1, T perf is a hyperparameter, P gap Represents the performance gap relative to the global performance, as shown in formula (2), Among them, μ Acc Represents the historical accuracy mean, μ Acc_global represents the global performance benchmark; S32, design stability factor, as shown in formula (3), Among them, α2 represents the weight coefficient of the stability factor, k2 represents the hyperparameter, σ Acc represents the standard deviation of category accuracy; S33, constructing a dynamic adaptive weight optimization coefficient based on the performance adjustment factor and the stability factor, as shown in formula (4), F adj =α2+(F perf -α2)·F stab (4) Among them F adj Limited to preset range Inside; S34, constructing the dynamic weight of the fault category based on the dynamic adaptive weight optimization coefficient, as shown in formula (5), w dyn =w base ·F′ adj (5) Among them, w dyn represents the dynamic weight, w base Represents the base weight.
6. The method for diagnosing faults in an aero-engine gas path system based on cross-domain small sample learning according to claim 5, characterized in that: The LDAM loss function is constructed based on the dynamic adaptive weight optimization coefficient, including: S34, constructing the LDAM loss function based on the dynamic adaptive weight optimization coefficient, as shown in formula (6), Where B represents the batch size, C represents the total number of categories, Indicates that the source domain feature corresponds to the real source domain category y l The original logit value calculated above is the logit value after LDAM boundary adjustment and scaling, z i,a Represents the final logit value of the source domain feature on the ath fault category.
7. The method for diagnosing faults in an aero-engine gas path system based on cross-domain small sample learning according to claim 1, characterized in that: Construct the intra-domain contrast loss function and the cross-domain contrast loss function, and construct the domain-aware contrast loss function based on the intra-domain contrast loss function and the cross-domain contrast loss function, including: S51. Construct the intra-domain contrast loss function, as shown in formula (7): Where N is the number of samples in each batch, f i is the feature vector of the sample in the corresponding data domain, f j is the feature vector of the positive sample in the sample, |P(i)| represents the total number of samples in the set, f k is the feature vector of any sample k in a batch, τ is the temperature coefficient; S52. Construct a cross-domain contrast loss function, as shown in formula (8): Among them, P ST (n) = {m | l S,n =l T,m } is the index set of target domain data samples m with the same label as source domain data sample n, |P ST (n)| represents the set P ST The total number of samples in (n), f S,n Represents the feature vector of source domain data sample n, f T,m Represents the feature vector of the target domain data sample m, f T,r Represents the feature vector of the target domain data sample r; S53. Based on the intra-domain contrast loss function and the cross-domain contrast loss function, a domain-aware contrast loss function is constructed, as shown in formula (9): Among them, γ represents the weight coefficient, represents the intra-domain contrast loss function of the source domain, represents the intra-domain contrastive loss function of the target domain.
8. The method for diagnosing faults in an aero-engine gas path system based on cross-domain small sample learning according to claim 7, characterized in that: The backpropagation algorithm is used to optimize the multi-level feature enhancement module based on the LDAM loss function and the domain awareness contrast loss function to obtain the optimized multi-level feature enhancement module. The aircraft engine data to be diagnosed is input into the optimized multi-level feature enhancement module and the diagnosis module to obtain the final fault diagnosis results, including: S61. Construct a total loss function based on the LDAM loss function and the domain awareness contrast loss function, as shown in formula (10). Among them, λ is the weight coefficient; S62. Calculate the total loss function value of the source domain features and the target domain features, calculate the gradient of the total loss function based on the back propagation algorithm, and use the Adam optimizer to update and optimize the parameters in the multi-level feature enhancement module and the contrast projection module based on the gradient to obtain an optimized multi-level feature enhancement module and an optimized contrast projection module; S63. Use the optimized contrast projection module to perform contrast projection on the source domain features and the target domain features to obtain new source domain projection features and new target domain projection features. During the training process, the multi-level feature enhancement module and the contrast projection module are continuously iterated to update and optimize until the total loss function value converges, thereby obtaining the final multi-level feature enhancement module. S64: Input the aircraft engine data to be diagnosed into the optimized multi-level feature enhancement module and the diagnosis module to obtain the final fault diagnosis result.