Aeroengine Fault Identification Method Based on Multi-Source Domain Adaptive Network
By introducing a multi-source domain adaptive network with dynamic migration modules and attention mechanisms in aero engine fault identification, the problem of different distributions of training and test data is solved, and a more efficient fault recognition effect is achieved.
Patent Information
- Application Number
- CN202211109458.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-13
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-09-13
AI Technical Summary
In the identification of rolling bearing faults of aircraft engines, existing deep learning models cannot effectively adapt to the data distribution differences between training data and test data, resulting in degradation of recognition capabilities, especially in multi-source domain scenarios, which is difficult to fully utilize the knowledge of each source domain.
A multi-source domain adaptive network that adopts dynamic migration module and attention mechanism, transforms multi-source domain problems into single-source domain problems by designing dynamic migration modules in feature extractors, and embeds attention mechanisms in the classifiers, and dynamically adjusts model parameters to adapt to the data distribution differences of different source domains.
It effectively reduces the data distribution differences between multi-source domains, improves the accuracy and robustness of fault identification, and can better utilize the knowledge of multiple source domains to complete the target identification task.
Smart Images

Figure CN115587289B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of aero-engines, and particularly relates to a method for identifying aero-engine faults. Background Art
[0002] Rolling bearings are important components of aero-engines and are widely used in various scenarios. Rolling bearings work in environments with variable rotational speeds, temperatures, loads, strengths, etc. for a long time, and are vulnerable components in aero-engines. When rolling bearings are in a harsh environment for a long time, it will affect the working efficiency of aero-engines and cause economic losses at least, and pose a potential threat to personal safety at worst. Therefore, the research on the faults of rolling bearings has always attracted much attention.
[0003] With the increasing digitization and intelligence of modern industry, some traditional models such as artificial neural networks, backpropagation, support vector machines, etc. can no longer meet the development needs due to their dependence on feature extraction techniques. The emergence of deep learning overcomes the deficiencies of these traditional models and has been continuously proven to be an effective data-driven bearing fault diagnosis method. Many scholars have applied deep learning to the identification of rolling bearing faults and verified the identification effect through a large number of experiments. Some scholars use long short-term memory networks to design identification models for bearing fault diagnosis, and some combine long short-term memory networks and residual learning to diagnose gearbox faults. Some use variational autoencoders to solve the problem of bearing fault identification under data imbalance, and some use deep regularized variational autoencoders to solve fault diagnosis problems. Some combine convolutional neural networks with infrared thermal imaging technology to achieve fault diagnosis, and some conduct a series of bearing fault diagnosis experiments based on convolutional neural networks with self-attention modules. For these deep diagnosis models mentioned above, when the test data and the training data satisfy the same data distribution, a satisfactory identification level can be achieved. However, in actual scenarios, due to the complex and variable operating conditions of aero-engines, there will be certain differences in the data distributions of the collected training data and test data, resulting in a significant degradation of the model's identification ability obtained from the training data on the test data. Since the training data and the test data do not satisfy the same distribution, a domain adaptation model is trained, which aims to map different data distributions into a high-dimensional space to reduce the data distribution difference, thereby improving the model's identification ability and generalization ability.
[0004] Typical deep learning models that do not consider data distribution differences are not suitable for these practical scenarios. As a popular transfer learning method, domain adaptation aims to mine shared knowledge from the source domain and the target domain to solve target tasks such as recognition and detection. Current research on domain adaptation mainly includes two categories: single-source domain adaptation and multi-source domain adaptation. Single-source domain adaptation methods aim to apply the model trained on the source domain to the target domain. A widely used strategy is to minimize the distribution distance between the source domain and the target domain. Some scholars have applied the joint maximum mean discrepancy to design an adversarial adaptation framework to solve the recognition problem under variable rotational speeds. Others have applied multiple dense blocks with central moment differences to construct a bearing fault diagnosis model. When the data distribution differences between different domains are large, these single-source domain methods are not competitive because they only transfer the relevant knowledge of a single source domain to solve the target diagnosis task and usually have difficulty obtaining satisfactory recognition results. In some practical scenarios, multi-source domains with different data distributions can be obtained from different but related working conditions to make up for the lack of recognition knowledge in the single-source domain. Multi-source domain adaptation methods solve the domain adaptation problem by using a series of source domains with various data distributions. Some use the maximum mean discrepancy to align the distributions between each source domain and the target domain to construct a multi-source domain adaptation model. Some have designed a matrix distance metric strategy to construct a domain adaptation framework based on multi-source domains. Others have applied multiple sub-networks to align the data distributions of each source domain and the target domain to identify unknown faults. The above research all uses a "static" transfer model, that is, the parameters do not change according to the different data distributions of each source domain. However, when there are large distribution differences not only between each source domain but also between each source domain and a single target domain, these "static" transfer models are difficult to fully utilize the knowledge of each source domain to complete the target recognition task, usually affecting the adaptation performance and leading to optimization difficulties. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the present invention provides a method for aero-engine fault recognition based on a multi-source domain adaptation network. This method transforms the multi-source domain fault recognition problem into a single-source domain fault recognition problem by designing a dynamic transfer module in the feature extractor, thereby promoting the alignment of data distributions between different domains. Then, an attention mechanism is embedded into two different classifiers to enhance the influence of relevant source domains, and thus fully utilize the knowledge contained in multiple source domains to complete the target recognition task. The method proposed by the present invention is easy to operate and has obvious fault recognition effects, providing a feasible solution for aero-engine fault recognition under obvious data distribution differences.
[0006] The technical solutions adopted by the present invention to solve its technical problems include the following steps:
[0007] Step 1: Collect the vibration signals of aero-engine components using sensors. According to the different operating conditions of aero-engine components, obtain multiple different source domain data and a single target domain data;
[0008] Step 2: Perform Fourier transform on the collected multiple different source domain data and a single target domain data to obtain the spectral samples corresponding to each source domain and the single target domain;
[0009] Step 3: Construct a feature extractor with a dynamic migration module; the dynamic migration module is divided into a static convolution module and a dynamic convolution module;
[0010] Step 3-1: Define the dynamic migration module as:
[0011] H(x) = H c (x) + △H(x) (1)
[0012] where H c (x) is the output of the static convolution module, △H(x) represents the output of the dynamic convolution module depending on the input x; H(x) represents the output of the dynamic migration module; the input x is the spectral sample obtained in Step 2;
[0013] The static convolution module consists of one convolutional layer, and the dynamic convolution module consists of W identical but independent convolutional layers, which is described by formula (2). The inputs of the static convolution module and the dynamic convolution module are the same;
[0014]
[0015] where represents the i th th convolutional layer, is the dynamic coefficient of, regarded as the projection of △H(x) in the i th th convolutional layer;
[0016] The dynamic coefficient is obtained by a coefficient branch, which consists of an average pooling layer and two fully connected layers; in addition, the Softmax activation function is used to normalize , and the acquisition process of is as follows:
[0017]
[0018] In the formula, avg(·), fc1 ReLU (·) and fc2 Softmax (·) respectively represent the average pooling layer, the first fully connected layer with the ReLU activation function, and the second fully connected layer with the Softmax activation function;
[0019] Step 3-2: The feature extractor G is serially composed of a first convolutional layer, a first dynamic migration module, a first pooling layer, a second convolutional layer, a second dynamic migration module, and a second pooling layer in sequence; the structures of the first dynamic migration module and the second dynamic migration module are the same;
[0020] Step 4: Construct a first classifier and a second classifier, both classifiers are composed of three fully connected layers and an output layer;
[0021] Step 5: Input each source domain data into its respective feature extractor with a dynamic migration module first, and then input the outputs of the feature extractor into the two classifiers respectively to obtain the corresponding classification losses; use loss s1 to represent the average loss of the first classifier on all source domain data, and use loss s2 to represent the average loss of the second classifier on all source domain data;
[0022] The attention mechanism automatically reduces the negative information caused by different domains by assigning coefficients to loss s1 and loss s2 respectively, that is, when the value of loss s1 is greater than the value of loss s2 , the smaller coefficient is assigned to loss s1 , otherwise the larger coefficient is assigned; the acquisition processes of the coefficients of loss s2 and loss s2 are as follows formulas respectively:
[0023]
[0024] In the formula, att1 and att2 are the coefficients assigned to loss s2 and loss s2 respectively, concat(loss s1 , loss s2 ) represents connecting two matrices according to the standard of horizontal connection, d k represents the number of rows of the matrix, and Softmax(·) represents the Softmax activation function;
[0025] According to the coefficients att1 and att2, the optimization objectives of the two classifiers with the attention mechanism are summarized as
[0026]
[0027] Step 6: Use multiple source domain data to train the feature extractor and the two classifiers, and after training, obtain a multi-source domain adaptive network; then perform fault identification on single target domain data.
[0028] Preferably, for the classification losses of the two classifiers, when there are two sets of source domain data, the losses of the first source domain data and the second source domain data on the two classifiers are respectively as follows:
[0029]
[0030] where x s1 and x s2 respectively represent all input samples of the first source domain and the second source domain, y s1 and y s2 respectively represent the labels corresponding to x s1 and x s2 , G(x s1 ) and G(x s2 ) respectively represent the outputs of the feature extractor on x s1 and x s2 , and are respectively the cross-entropy losses of the first classifier and the second classifier on x s1 ; and are respectively the cross-entropy losses of the first classifier and the second classifier on x s2 , loss s1 and loss s2 are respectively the average losses of the two classifiers on x s1 and x s2 .
[0031] The beneficial effects of the present invention are as follows:
[0032] By introducing a dynamic transfer module, the present invention enables the model parameters to adapt to different data distributions in different source domains. The dynamic transfer module breaks the barriers of multiple source domains, converts the multi-source domain recognition problem into a single-source domain recognition problem, simplifies the alignment of data distributions between each source domain and a single target domain, and thus more effectively reduces the distribution differences between data. Then, an attention mechanism is embedded in the two classifiers, which can fully utilize the knowledge in different source domains to align the data distributions by assigning different weights to different source domains, thereby obtaining better fault recognition results. To illustrate the effectiveness of the proposed method as much as possible, the present invention constructs a series of target recognition tasks using two sets of aero-engine fault data to verify the effectiveness and feasibility of the proposed method. A large number of experimental results prove that the method proposed by the present invention is more effective and has better robustness in identifying aero-engine faults than other methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart of aero-engine key component fault recognition based on a dynamic transfer multi-source domain adaptive network according to the present invention.
[0034] Figure 2 This is the structural diagram of the dynamic migration module of the present invention;
[0035] Figure 3 This is the structural diagram of the feature extractor with a dynamic migration module constructed by the present invention.
[0036] Figure 4 This is the overall framework diagram of the adaptive network of the present invention.
[0037] Figure 5 This is the schematic diagram of the influence of different numbers of convolutional layers in Embodiment 1 of the present invention on the recognition result. Detailed implementation manners
[0038] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0039] The present invention first designs a dynamic migration module in the feature extractor to convert the multi-source domain fault recognition problem into a single-source domain fault recognition problem, thereby promoting the alignment of data distributions between different domains. Then, an attention mechanism is embedded into two different classifiers to enhance the influence of relevant source domains, and thus make full use of the diagnostic knowledge contained in multiple source domains to complete the target recognition task.
[0040] Aero-engine fault recognition method based on a multi-source domain adaptive network, comprising the following steps:
[0041] Step 1: Use sensors to collect vibration signals of aero-engine components, and obtain multiple different source domain data and a single target domain data according to different operating conditions of aero-engine components;
[0042] Step 2: Perform Fourier transform on the collected multiple different source domain data and the single target domain data to obtain the spectral samples corresponding to each source domain and the single target domain;
[0043] Step 3: Construct a feature extractor with a dynamic migration module, where the structure of the dynamic migration module is as Figure 2 shown; the dynamic migration module is divided into a static convolution module and a dynamic convolution module;
[0044] Step 3-1: Define the dynamic migration module as:
[0045] H(x) = H c (x) + △H(x) (1)
[0046] wherein, H c (x) is the output of the static convolution module, and △H(x) represents the output of the dynamic convolution module depending on the input x; H(x) represents the output of the dynamic migration module; the input x is the spectral sample obtained in Step 2;
[0047] The static convolution module consists of a single convolutional layer, and the dynamic convolution module consists of W identical but independent convolutional layers, as described by Equation (2). The inputs to the static convolution module and the dynamic convolution module are the same;
[0048]
[0049] where, represents the i th th convolutional layer, is 's dynamic coefficient, regarded as the projection of △H(x) in the i th th convolutional layer; the dynamic convolution module selects these projections in a way that depends on the input, thus choosing different feature subspaces to learn different inputs;
[0050] The dynamic coefficient is obtained from a coefficient branch, which consists of an average pooling layer and two fully connected layers; in addition, the Softmax activation function is used to normalize, 's acquisition process is as follows:
[0051]
[0052] In the formula, avg(·), fc1 ReLU (·) and fc2 Softmax {·} represent the average pooling layer, the first fully connected layer with ReLU activation function, and the second fully connected layer with Softmax activation function, respectively;
[0053] Step 3-2: Embed the constructed dynamic transfer module into the feature extractor to align the data distributions between each source domain and a single target domain; the structure of the feature extractor with the dynamic transfer module is as Figure 3 shown. The feature extractor G is serially composed of a first convolutional layer, a first dynamic transfer module, a first pooling layer, a second convolutional layer, a second dynamic transfer module, and a second pooling layer in sequence; the structures of the first dynamic transfer module and the second dynamic transfer module are the same; the feature extractor with the dynamic transfer module makes the proposed method more flexible and has a strong ability to align the large data distributions between each source domain and a single target domain;
[0054] Step 4: Construct a first classifier and a second classifier, both of which consist of three fully connected layers and an output layer;
[0055] Step 5: Input each source domain data into its respective feature extractor with a dynamic migration module first, and then input the outputs of the feature extractor into two classifiers respectively to obtain the corresponding classification losses; use loss s1 to represent the average loss of the first classifier on all source domain data, and use loss s2 to represent the average loss of the second classifier on all source domain data;
[0056] When there are two sets of source domain data, the losses of the first source domain data and the second source domain data on the two classifiers are shown as follows respectively:
[0057]
[0058] where, x s1 and x s2 represent all input samples of the first source domain and the second source domain respectively, y s1 and y s2 represent the corresponding labels of x s1 and x s2 respectively, G(x s1 ) and G(x s2 ) represent the outputs of the feature extractor on x s1 and x s2 respectively, and are the cross-entropy losses of the first classifier and the second classifier on x s1 respectively; and are the cross-entropy losses of the first classifier and the second classifier on x s2 respectively, loss s1 and loss s2 are the average losses of the two classifiers on x s1 and x s2 respectively;
[0059] The attention mechanism automatically reduces the negative information caused by different domains by assigning coefficients to loss s1 and loss s2 respectively, that is, when the value of loss s1 is greater than the value of loss s2 , the smaller coefficient is assigned to loss s1 , otherwise the larger coefficient is assigned; by using this attention mechanism, the relevant knowledge from different source domains can be fully utilized to achieve multi-source domain adaptation; the processes of obtaining the coefficients of loss s2 and loss s2 are as follows in the respective formulas:
[0060]
[0061] In the formula, att1 and att2 are the coefficients assigned to loss s2 and loss s2 respectively, and concat(loss s1 , loss s2 ) represents connecting two matrices according to the standard of horizontal connection. d k represents the number of rows of the matrix, and Softmax(·) represents the Softmax activation function;
[0062] According to the coefficients att1 and att2, the optimization objectives of two classifiers with attention mechanisms are summarized as
[0063]
[0064] Step 6: Use multiple source domain data to train the feature extractor and two classifiers. After training, a multi-source domain adaptive network is obtained; then, fault identification is performed on single target domain data. Specific embodiments:
[0066] The content of the present invention will be further described in detail below with reference to the accompanying drawings: Refer to Figure 1 As shown, the content of the present invention can be mainly divided into two parts: The first part is to construct a feature extractor with a dynamic migration module; the second part is to design a dual classifier with an attention mechanism.
[0067] Refer to Figure 2 As shown, the dynamic migration module mines the features of the input data through a static convolution module and a dynamic convolution module;
[0068] Refer to Figure 3 As shown, the feature extractor with a dynamic migration module designed by the present invention can adapt to significantly different data distributions;
[0069] Refer to Figure 4 As shown, the method proposed by the present invention can utilize the dynamic migration module and the attention mechanism to fully transfer the knowledge in multiple source domains, and thus more effectively complete the target recognition task.
[0070] Refer to Figure 5 As shown, when the dynamic migration module is embedded in the feature extractor of the proposed method, different numbers of the same but independent convolution layers have a certain impact on the recognition result. The abscissa in the figure represents different target recognition tasks; the ordinate represents the recognition accuracy rate, and the unit is %.
[0071] Example 1: In this example, a related rolling bearing dataset is used to carry out a series of target recognition tasks. The dataset includes four different operating condition types: normal type, inner race fault, outer race fault, and roller fault. The data sampling frequency is 50 kHz. According to three different rotational speeds (600 rpm, 800 rpm, and 1000 rpm), the dataset is divided into three different domains, namely T 600 , T 800 , and T 1000 , where T 600 represents the domain obtained at a motor speed of 600 rpm, and the others are similar. Each domain contains four operating conditions. After processing the original data through fast Fourier transform, each health type contains 250 samples, and each sample has 784 data points. According to the principle of randomly selecting 2 domains from 3 domains as the source domains and the remaining 1 domain as the target domain, 3 different target recognition tasks can be constructed, namely (T 800 , T 1000 → T 600 ), (T 600 , T 1000 → T 800 ), and (T 600 , T 800 → T 1000 ). In the (T 800 , T 1000 → T 600 ) target recognition task, the proposed method is compared with four representative domain adaptation methods (Deep Domain Adaptation Related Alignment, Adversarial Domain Adaptation Network, Joint Distribution Adaptation, Principal Component Analysis) and a deep learning method (Convolutional Neural Network). To avoid the contingency of experimental results, each method runs independently 10 times in their respective same environment. As can be seen from Table 1, the proposed Dynamic Transfer Multi-Source Domain Adaptation Network of the present invention can make full use of more effective information from multiple source domains to solve the target recognition task. In addition, the recognition results of the method of the present invention with different numbers of independent convolutional layers in this example are shown in Table 2.
[0072] Table 1 Experimental comparison results
[0073]
[0074] Table 2 Recognition results of the method of the present invention based on different numbers of independent convolutional layers
[0075]
[0076]
[0077] Example 2: In this example, a series of target recognition tasks are carried out using a relevant rolling bearing dataset to demonstrate the performance of the proposed method. According to different load torques, radial forces, and rotational speeds, the dataset is divided into three different sub-datasets (M1, M2, and M3). Each sub-dataset includes seven health types: normal type, mild inner race fault type, severe inner race fault-I type, severe inner race fault-II type, mild outer race fault type, severe outer race fault type, and compound fault type. The sampling frequency for each health type in each sub-dataset is 64 kHz. 313,600 data points of each health type are converted into 200 samples through fast Fourier transform. Therefore, each sample contains 784 data points. Based on the three sub-datasets, three different domain adaptation tasks can be established, namely (M2, M3→M1), (M1, M3→M2), and (M1, M2→M3). Taking (M2, M3→M1) as an example, M2 and M3 are used as the source domains, and the remaining M1 is used as the target domain. In the (M2, M3→M1) target recognition task, the proposed method is compared with deep domain adaptation related alignment, adversarial domain adaptive network, joint distribution adaptation, principal component analysis, and convolutional neural network. To avoid the contingency of experimental results, each method runs independently 10 times in their respective identical environments. As can be seen from Table 3, the dynamic transfer multi-source domain adaptive network proposed in the present invention can make full use of information from multiple source domains to more effectively solve the target recognition task. In addition, the recognition results of the method of the present invention based on different numbers of independent convolutional layers are shown in Table 4.
[0078] Table 3 Experimental comparison results
[0079]
[0080] Table 4 Recognition results of the method of the present invention based on different numbers of independent convolutional layers
[0081]
[0082]
Claims
1. A method for identifying aero-engine faults based on a multi-source domain adaptive network, characterized in that, It includes the following steps: Step 1: Use sensors to collect vibration signals of aero-engine components. According to different operating conditions of aero-engine components, multiple different source domain data and a single target domain data are obtained; Step 2: Perform Fourier transform on the collected multiple different source domain data and a single target domain data to obtain the spectral samples corresponding to each source domain and the single target domain; Step 3: Construct a feature extractor with a dynamic transfer module; the dynamic transfer module is divided into a static convolution module and a dynamic convolution module; Step 3-1: Define the dynamic transfer module as: H(x) = H c (x) + △H(x) (1) Among them, H c (x) is the output of the static convolution module, and △H(x) represents the output of the dynamic convolution module that depends on the input x; H(x) represents the output of the dynamic migration module; the input x is the spectral sample obtained in step 2; The static convolution module consists of one convolutional layer, and the dynamic convolution module consists of W identical but independent convolutional layers, which is described by formula (2). The inputs of the static convolution module and the dynamic convolution module are the same; Among them, represents the i-th th convolutional layer, is the dynamic coefficient of, regarded as the projection of △H(x) in the i-th th convolutional layer; Dynamic coefficient Obtained from a coefficient branch, which consists of an average pooling layer and two fully connected layers; in addition, the Softmax activation function is used to normalize it, The acquisition process of is as follows: where avg(·), fc1 ReLU (·) and fc2 Softmax {·} represent the average pooling layer, the first fully connected layer with ReLU activation function, and the second fully connected layer with Softmax activation function, respectively; Step 3-2: The feature extractor G is sequentially and serially composed of a first convolutional layer, a first dynamic transfer module, a first pooling layer, a second convolutional layer, a second dynamic transfer module, and a second pooling layer; the structures of the first dynamic transfer module and the second dynamic transfer module are the same; Step 4: Construct a first classifier and a second classifier, and both classifiers are composed of three fully connected layers and an output layer; Step 5: Input each source domain data into its respective feature extractor with a dynamic migration module first, and then input the outputs of the feature extractors into two classifiers respectively to obtain the corresponding classification losses; use loss s1 to represent the average loss of the first classifier on all source domain data, and use loss s2 to represent the average loss of the second classifier on all source domain data; The attention mechanism automatically reduces the negative information caused by different domains by assigning coefficients to loss s1 and loss s2 respectively, that is, when the value of loss s1 is greater than the value of loss s2 , a smaller coefficient is assigned to loss s1 , otherwise a larger coefficient is assigned; the acquisition processes of the coefficients of loss s2 and loss s2 are as follows in the formulas respectively: where att1 and att2 are the coefficients assigned to loss s2 and loss s2 respectively, concat(loss s1 , loss s2 ) represents connecting two matrices according to the standard of horizontal connection, d k represents the number of rows of the matrix, and Softmax(·) represents the Softmax activation function; According to coefficients att1 and att2, the optimization objectives of two classifiers with attention mechanisms are summarized as Step 6: Use multiple source domain data to train the feature extractor and the two classifiers. After training, a multi-source domain adaptive network is obtained; then, fault identification is performed on the single target domain data.
2. The aero-engine fault identification method based on a multi-source domain adaptive network according to claim 1, wherein, Regarding the classification losses of the two classifiers, when there are two sets of source domain data, the losses of the first source domain data and the second source domain data in the two classifiers are respectively as follows: Among them, x s1 and x s2 respectively represent all input samples of the first source domain and the second source domain, y s1 and y s2 respectively represent the labels corresponding to x s1 and x s2 G(x s1 ) and G(x s2 ) respectively represent the outputs of the feature extractor on x s1 and x s2 . and are respectively the cross-entropy losses of the first classifier and the second classifier on x s1 ; and are respectively the cross-entropy losses of the first classifier and the second classifier on x s2 , loss s1 and loss s2 are respectively the average losses of the two classifiers on x s1 and x s2 .
Citation Information
Patent Citations
Deep transfer learning intelligent fault diagnosis method and device, storage medium and equipment
CN111898095A
Model and method for multi-source domain adaptation by aligning partial features
US20220138495A1