Rotating machine transfer learning fault diagnosis method fusing semi-supervised contrast learning in field

The fuzzy classification boundary problem in rotating machinery fault diagnosis under cross-work conditions is solved through the semi-supervised contrast transfer learning network (SSCTLN). The performance of the diagnostic model is improved by using SSCL and LMMD strategies, achieving higher fault identification accuracy and stability.

CN120257041APending Publication Date: 2025-07-04BEIJING UNIV OF CHEM TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510237311.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing deep learning-based fault diagnosis methods are difficult to effectively deal with domain offset problems under cross-work conditions, resulting in fuzzy classification boundaries and affecting diagnostic performance, especially in unsupervised target domains.

Method used

Semi-supervised contrast transfer learning network (SSCTLN) is adopted to improve cross-working condition diagnostic performance and eliminate fuzzy classification boundaries through the in-field semi-supervised contrast learning (SSCL) strategy, combining local maximum mean difference (LMMD) and dynamic limitation related contrast learning losses.

Benefits of technology

It effectively eliminates the fuzzy classification boundaries, improves the accuracy and stability of fault diagnosis of cross-work conditions, and improves the generalization ability of the diagnostic model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257041A_ABST
    Figure CN120257041A_ABST
Patent Text Reader

Abstract

The invention provides a rotating machine transfer learning fault diagnosis method fusing semi-supervised contrast learning in the field. In the method, an intra-domain semi-supervised contrast learning (SSCL) algorithm is designed, the SSCL takes category information as supervision, discriminative learning of different categories in each domain is guided to eliminate fuzzy classification boundaries, convenience is provided for domain adaptation of cross-domain features carried out by adopting local maximum mean difference (LMMD), and then cross-working-condition diagnosis performance is improved. Meanwhile, domain confrontation is introduced in a mode of dynamically limiting related contrast learning loss and transfer learning loss gains, so that negative effects of target domain pseudo labels with poor quality on SSCL and feature transfer learning are reduced, and the stability of a diagnosis model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent migration fault diagnosis, and relates to a method for intelligent migration diagnosis of rotating machinery component faults, in particular to a rotating machinery transfer learning fault diagnosis method based on semi-supervised contrast learning within a domain. Background Art

[0002] As one of the most widely used mechanical devices, the health status of internal related components of rotating machinery directly affects the safe and reliable operation of the mechanical system, such as bearings and gears. Moreover, in the rotating machinery system, these components are often in complex and changeable operating environments, making them vulnerable components. Effective condition monitoring and fault diagnosis of rotating machinery components are of great significance for ensuring the safe and reliable operation of the mechanical system and economic benefits. Therefore, it is necessary to carry out research on fault diagnosis of rotating machinery components. The existing fault diagnosis methods are mainly divided into traditional signal analysis-based and model-based methods, and data-driven intelligent methods.

[0003] Traditional methods extract fault features from the operating data of rotating machinery according to the fault mechanism, which requires relying on professional prior knowledge and is difficult to process a large amount of data. In recent years, with the development of artificial intelligence technology, data-driven intelligent fault diagnosis (IFD) methods have become a research hotspot. Among them, the IFD method based on deep learning (DL) has received extensive attention because it can automatically extract features from massive data and achieve end-to-end fault diagnosis.

[0004] However, the existing IFD methods need to meet the following conditions: the training and test data follow the same distribution, and there is abundant labeled data for model training. In actual industrial scenarios, rotating machinery components are often affected by different loads and speeds and are in cross-condition conditions, resulting in distribution differences (i.e., domain shift) between the collected data, and collecting sufficient labeled data will consume a huge labor cost. It is difficult to meet the above conditions, and at the same time, it also limits the generalization performance of the IFD method. Therefore, there is an urgent need for a new algorithm to construct a diagnostic model to effectively diagnose the faults of unlabeled target rotating machinery components in cross-condition scenarios.

[0005] As a way in transfer learning (TL), feature transfer learning (FTL) provides a solution idea for fault diagnosis in such cross-condition scenarios. FTL reduces the distribution difference between the training data (source domain) and the test data (target domain) to fit these two data domains and learn domain-invariant features for fault diagnosis in the target domain. Due to such characteristics, the research on fault diagnosis based on FTL has become a hotspot.

[0006] Although the FTL method can help achieve cross-condition fault diagnosis, the problem of fuzzy classification boundaries in the model training process cannot be ignored. Specifically, during the training of some FTL-based models, due to only focusing on the global alignment of cross-domain features and lacking discriminative and targeted learning between different fault categories, there are fuzzy classification boundaries, especially in the unsupervised target domain, which will limit the performance of the diagnostic model.

[0007] However, although relevant research and methods have been proposed, they solve the problem of fuzzy classification boundaries by achieving class-level alignment of cross-domain features. For example, Deep Subdomain Adaptation (DSAN), which uses Local Maximum Mean Discrepancy (LMMD) as the FTL means, and Conditional Domain Adversarial Network (CDAN), etc. However, these methods are difficult to handle fault samples located at the classification boundary and still result in limited performance. At the same time, there is little research on solving the problem of fuzzy classification boundaries by achieving class-level alignment of features within each domain.

[0008] The present invention introduces contrastive learning (CL) into the feature transfer learning of features and proposes a semi-supervised contrastive transfer learning network (SSCTLN) to solve this problem and improve the performance of the diagnostic model. Among them, the semi-supervised contrastive learning (SSCL) strategy designed by constructing supervised contrastive learning (SCL) within the source domain and self-supervised contrastive learning (Self-SCL) within the target domain enables the diagnostic model to specifically learn the discriminative features between different fault categories within each domain to eliminate fuzzy classification boundaries, while facilitating the transfer learning of cross-domain features and improving the cross-condition diagnosis performance. Summary of the Invention

[0009] To solve the technical problem of cross-condition rotating machinery component fault diagnosis, the present invention provides a cross-condition rotating machinery fault diagnosis method based on a semi-supervised contrastive transfer learning network.

[0010] The purpose of the present invention is to propose a semi-supervised contrastive transfer learning network (SSCTLN) for realizing the fault diagnosis of cross-condition rotating machinery. Specifically, a semi-supervised contrastive learning (SSCL) strategy is designed. SSCL uses class information as supervision to guide the discriminative learning of different classes within each domain to eliminate fuzzy classification boundaries and facilitates the transfer learning of cross-domain features using Local Maximum Mean Discrepancy (LMMD). At the same time, in a dynamic way of restricting the relevant contrastive learning loss and transfer learning loss gain, and introducing domain adversarial, to reduce the negative impact of poor-quality target domain pseudo-labels on SSCL and feature transfer learning and better improve the cross-condition diagnosis performance.

[0011] S1 Construct a feature encoder

[0012] To improve the fault feature extraction ability of the diagnostic model, the proposed SSCTLN adopts an enhanced residual convolution module (ERCM), and improves it according to the characteristics of the fault diagnosis task to construct the backbone network, and finally constructs the required one-dimensional convolutional neural network (1D-CNN).

[0013] When constructing the 1D-CNN, we choose the stacking method of one-dimensional convolutional layers (Conv1d-BN, Conv1d is one-dimensional convolution, BN is batch normalization) consistent with the ERCM. By stacking two convolutional layers in series without an activation function in the middle, we can efficiently simulate a convolutional layer with a large number of channels. At the same time, by stacking convolutional layers, the receptive field of the network is gradually increased, enabling it to have better feature representation ability than the traditional structure CNN. Different from the ERCM, the constructed 1D-CNN abandons the residual structure because the fault diagnosis task does not need to pay extra attention to the temporal sequence between features, but needs to focus on the typicality of fault features. At the same time, the 1D-CNN replaces the ReLU activation function in the ERCM with the Swish activation function, which is shown in formula (1), to help the diagnostic model capture and distinguish fault features more accurately. Then, a convolutional block with the structure of Conv1d-BN-Conv1d-BN-Swish_act-MaxPool is obtained, and the 1D-CNN is constructed by stacking a convolutional block of Conv1d-BN-Conv1d-BN-Swish_act-AdaMaxPool and four of the above convolutional blocks.

[0014] Swish_act(x)*Sigmoid(x) (1)

[0015]

[0016] Among them, x represents the features of the layer before the activation function extracted in the convolutional block, and Sigmoid is also an activation function, as shown in formula (2). Then, to further enhance the fault feature representation ability of the diagnostic model, a convolutional block attention mechanism (CBAM) with feature adaptive correction ability is used to enhance the key information in the fault features. By combining channel attention and spatial attention, it provides more comprehensive and effective feature representation ability for the diagnostic model. Finally, the constructed 1D-CNN is concatenated with the CBAM to form the used feature encoder G E (·).

[0017] S2 Construct a semi-supervised contrastive learning algorithm

[0018] The designed semi-supervised contrastive learning (SSCL) algorithm within the domain is respectively composed of supervised contrastive learning (SCL) within the source domain and self-supervised contrastive learning (Self-SCL) within the target domain. The within-domain SSCL enables the diagnostic model to learn the discriminability between samples of different fault categories by maximizing the similarity between sample representations belonging to the same fault category while minimizing the similarity between sample representations of different fault categories. Compared with unsupervised contrastive learning (UCL), SCL uses the labels corresponding to the samples as supervision to participate in model training, enabling the diagnostic model to better learn discriminative features. However, the data in the target domain has no labels, but the predicted values y' y generated by the classifier G t can be obtained, and this paper uses it as pseudo-labels to construct the supervision medium and constructs the target domain Self-SCL.

[0019] The specific implementation of the within-domain SSCL strategy is as follows: The source domain data and the target domain data (X s and X t respectively represent the data sample sequences of the source domain and the target domain in a mini-batch, which contain B source domain and target domain data samples x s and x t ) are input into the shared feature encoder G E (·) to obtain the source domain feature F s = G E (X s ) and the target domain feature F t = G E (X t ), where and respectively represent the source domain and target domain feature sequences containing B source domain and target domain features f s and f t . Then they are input into the projector to obtain the source domain and target domain feature representations and where Z s and Z t respectively represent the source domain and target domain feature representation sequences containing B source domain and target domain feature representations z s and z t . And the corresponding fault category labels are given, that is, and where y' t = G y (G E (x t)) represents inputting the target domain feature f t = G E (x t ) into the classifier G y (·) and obtaining the predicted value (i.e., the pseudo label). Then Z s and Z t are respectively used to calculate the SCL loss of the source domain and the Self-SCL loss of the target domain:

[0020]

[0021] where k ∈ {1, 2,..., B}, k ≠ i ≠ j, sim(·, ·) represents the similarity calculation between feature vectors, ||·||2 represents the 2-norm representation of feature vectors, and τ is the temperature coefficient, which is set to 0.07 according to the empirical principle. or represents, in the feature representation sequence of a mini-batch in the source domain or the target domain, the feature representation that belongs to the same fault category as the feature representation z s k or z t k , that is or their labels are the same. On the contrary, or represents, in the feature representation sequence of a mini-batch, the feature representation that belongs to different fault categories from the feature representation z s k or z t k , that is or their labels are different. Minimizing the SCL loss L SCL and the Self-SCL loss L Self-SCL can enable the diagnostic model to better define the classification boundary, which is beneficial for learning discriminative features for the classification task and also provides convenience for the transfer learning of cross-domain features. Specifically, the existence of a fuzzy classification boundary increases the risk of misalignment during cross-domain feature alignment. After the in-domain SSCL solves the problem of the fuzzy classification boundary, it reduces the probability of misalignment of features to a certain extent.

[0022] S3 sets the feature transfer learning algorithm

[0023] While performing within-domain SSCL, this paper uses Local Maximum Mean Discrepancy (LMMD) to calculate the similarity between the source domain and the target domain, which is used to guide the transfer learning of cross-domain features. LMMD aims to calculate the feature distribution corresponding to each fault category to make the distributions of each category similar, rather than the overall distribution. The within-domain SSCL strategy in S2 clarifies the classification boundaries between categories within each domain and enhances the distinguishability between categories, providing a more accurate basis for measuring distribution similarity in LMMD calculation, which facilitates the transfer learning of cross-domain features.

[0024] Suppose the feature representation sequences of a mini-batch within the source domain and the target domain are respectively and where Z s or Z t is a feature representation sequence that contains B feature representations z s or z t from the source domain or the target domain and the corresponding fault labels y s or y′ t . y′ t = Gy(GE(xt)) represents the predicted value obtained by inputting the target domain feature f t = G E (x t ) into the classifier G y (·), that is, the pseudo-label. Then, the LMMD loss calculation between Z s and Z t is as follows:

[0025]

[0026] where C represents the total number of fault categories (depending on the diagnostic task), c represents the fault category to which the feature representation z s or z t belongs, represents the Reproducing Kernel Hilbert Space (RKHS), φ(·) represents the non-linear mapping from the original feature space to the RKHS, represents the feature kernel. and represent the weights of the source domain feature representation z s i and the target domain feature representation z t j belonging to category c respectively. It should be noted that the sum of the weights is 1, that is, and and is the weighted sum over category c. The weight i of the feature representation z is calculated as follows:

[0027]

[0028] Among them, represents the c-th entry of the vector obtained after one-hot encoding the class label y i For the source domain, the true label y s i is transformed into a one-hot vector for calculating the weight of each sample For the case where the true label in the unlabeled target domain is not available, after one-hot encoding the pseudo-label obtained by the classifier the weight of each sample is calculated

[0029] S4 Construct a semi-supervised contrastive transfer learning algorithm

[0030] When calculating the Self-SCL loss and LMMD loss of the target domain, the pseudo-labels of the target domain are both used, which results in the performance improvement effect of the above algorithm on the diagnostic model depending on the quality of the pseudo-labels. Therefore, from the perspective of reducing the impact of poor-quality pseudo-labels on model training and directly improving the quality of pseudo-labels, the following improvements are made to the above algorithm:

[0031] Improvement 1: Considering that the quality of the pseudo-labels output by the model is unstable in the early stage of training, to reduce the adverse impact of poor-quality pseudo-labels on model training, an increasing penalty coefficient λ is assigned to the Self-SCL loss and LMMD loss respectively increase , and the stability of the model is ensured by restricting their early gains.

[0032] Improvement 2: Considering that after reducing the early LMMD loss gain, the model training will focus more on the source domain rather than the target domain, which will lead to limited performance on the target domain. To improve the attention of the model training to the target domain and thus directly improve the quality of the target domain pseudo-labels, source domain features and target domain features are introduced for domain adversarial training to additionally guide the transfer learning of cross-domain features. The domain adversarial loss is calculated as follows:

[0033]

[0034] Among them, L CE represents the cross-entropy loss, G Adv (·) represents the domain adversarial module, that is, the discriminator, and d k is the domain label corresponding to the feature f k .

[0035] By using the supervised representative learning loss L of the source domainSCL , the self-supervised contrastive learning loss L in the target domain Self-SCL , the Local Maximum Mean Discrepancy (LMMD) loss L LMMD and the domain adversarial loss L Adv are combined to obtain the proposed semi-supervised contrastive transfer learning (SSCTL) algorithm. The objective loss function L corresponding to the SSCTL algorithm SSCTL is as follows:

[0036]

[0037] where μ is a fixed penalty coefficient that restricts the gain of the domain adversarial loss, and it is set to 0.5 according to the empirical principle. epoch is the number of iterative training times. Based on the empirical principle and multiple experiments, the maximum number of iterative training times is set to 200, so epoch ∈ {0, 1, 2,..., 199}. λ increase is an exponential increasing penalty coefficient, and its value increases from 0 to 1 as epoch increases.

[0038] S5 Construct a cross-condition diagnosis model based on the semi-supervised contrastive transfer learning network

[0039] Combined with the feature encoder G in S1 E (·) and the semi-supervised contrastive transfer learning (SSCTL) algorithm in S4 to construct a cross-condition diagnosis model based on the semi-supervised contrastive transfer learning network (SSCTLN). This diagnosis model has two optimization objectives during training: 1) Minimize the various contrastive learning and transfer learning losses shown in formula (8), 2) Minimize the supervised classification loss L y on the labeled source domain. For a mini-batch of source domain data samples The supervised classification loss of the source domain is defined as follows:

[0040]

[0041] In the invention, the SSCTL loss and the supervised classification loss L y of the source domain are jointly optimized in an end-to-end manner. The joint objective loss function is shown in formula (11), and the Adam optimizer is used to update the network parameters.

[0042]

[0043] The hyperparameters during the training of the diagnostic model based on SSCTLN are set as follows according to the empirical principle: the batch size batch_size is set to 128, i.e., B = 128; the learning rate learning_rate is set to 0.001, and a stepwise decay strategy is adopted, and learning_rate is decayed at a decay rate of 0.1 at epoch = 119 and epoch = 159; the weight decay rate weight_decay is set to 3e-6. Description of the Drawings

[0044] Figure 1 is the network model framework diagram of the invention;

[0045] Figure 2 is the 1D-CNN in the feature encoder of the invention;

[0046] Figure 3 is the F1-score index of each fault category of the invention on the PU dataset;

[0047] Figure 4 is the F1-score index of each fault category of the invention on the WTG dataset;

[0048] Figure 5 is the confusion matrix diagram of the invention on the PU dataset;

[0049] Figure 6 is the confusion matrix diagram of the invention on the WTG dataset;

[0050] Figure 7 is the feature visualization diagram of the invention on the PU dataset;

[0051] Figure 8 is the feature visualization diagram of the invention on the WTG dataset. Detailed Description of the Invention

[0052] The present invention will be further described below in conjunction with the drawings and the detailed description of the invention.

[0053] Figure 1 is the framework flowchart of the cross-condition diagnostic model of the semi-supervised contrast transfer learning network (SSCTLN) of the present invention, which includes: a shared feature encoder composed of a one-dimensional convolutional neural network (1D-CNN) and a convolutional block attention module (CBAM) connected in series; a fault label classifier composed of a fully connected layer and a Softmax activation function; and a semi-supervised contrast transfer learning (SSCTL) module proposed for feature transfer learning.

[0054] In order to describe the technical solution of the above invention in more detail, the following lists the specific algorithm construction process.

[0055] First, the acquisition of rotating machinery data and the encapsulation of the datasets required for model training and testing are introduced in detail.

[0056] For various operating states of rotating machinery under different working conditions (different working condition parameters, such as rotational speed, load, etc.), the present invention samples the fault vibration signals of rotating machinery components through acceleration sensors installed near the faulty components, and constructs a labeled source domain dataset and an unlabeled target domain dataset. For the cross-condition transfer fault diagnosis scenario, the data collected under the working condition of the labeled data is defined as a source domain with n s source domain with n labeled samples The data collected under the working condition of the unlabeled data is defined as a target domain with n t target domain with n unlabeled samples where x and y represent the fault samples and fault class labels respectively. Since the data is collected from different working conditions, there will be differences between the cross-domain data and their distributions, that is, X s ≠X t and Finally, the entire source domain is divided into a training set, and the target domain is divided into a training set and a testing set to complete the data preparation.

[0057] Next, the feature encoder designed in the present invention is introduced in detail.

[0058] To improve the fault feature extraction ability of the diagnostic model, the proposed SSCTLN adopts an enhanced residual convolution module (ERCM), and improves it according to the characteristics of the fault diagnosis task to construct the backbone network, and finally constructs a 1D-CNN as Figure 2 shown.

[0059] When constructing the 1D-CNN, we selected the stacking method of one-dimensional convolutional layers (Conv1d-BN, where Conv1d is one-dimensional convolution and BN is batch normalization) consistent with ERCM. By cascading and stacking two convolutional layers without an activation function in the middle, we efficiently simulated a convolutional layer with a large number of channels. At the same time, by stacking convolutional layers, we gradually increased the receptive field of the network, enabling it to have better feature representation ability than traditional CNN structures. Different from ERCM, the constructed 1D-CNN abandoned the residual structure because the fault diagnosis task does not require additional attention to the temporal sequence between features but rather focuses on the typicality of fault features. At the same time, 1D-CNN replaced the ReLU activation function in ERCM with the Swish activation function, which is shown in formula (1), to help the diagnostic model more accurately capture and distinguish fault features. Then, a convolutional block with the structure of Conv1d-BN-Conv1d-BN-Swish_act-MaxPool was obtained, and a 1D-CNN was constructed by stacking a convolutional block of Conv1d-BN-Conv1d-BN-Swish_act-AdaMaxPool and four of the above convolutional blocks.

[0060] Swish_act(x) = x * Sigmoid(x) (1)

[0061]

[0062] Among them, x represents the features of the layer before the activation function extracted within the convolutional block, and Sigmoid is also an activation function, as shown in formula (2). Then, to further enhance the fault feature representation ability of the diagnostic model, a convolutional block attention mechanism (CBAM) with feature adaptive correction ability was used to enhance the key information in the fault features. By combining channel attention and spatial attention, it provides more comprehensive and effective feature representation ability for the diagnostic model. Finally, the constructed 1D-CNN was concatenated with CBAM to form the feature encoder G used. E (·).

[0063] Then, the semi-supervised contrast transfer learning strategy algorithm designed in the present invention is introduced in detail.

[0064] (1) Semi-supervised contrast learning strategy

[0065] The designed semi-supervised contrastive learning (SSCL) algorithm within the domain is respectively composed of supervised contrastive learning (SCL) in the source domain and self-supervised contrastive learning (Self-SCL) in the target domain. The within-domain SSCL enables the diagnostic model to learn the discriminability between samples of different fault categories by maximizing the similarity between sample representations belonging to the same fault category while minimizing the similarity between sample representations of different fault categories. Compared with unsupervised contrastive learning (UCL), SCL uses the labels corresponding to the samples as supervision to participate in model training, enabling the diagnostic model to better learn discriminative features. However, the data in the target domain is unlabeled, but the predicted values y' of the target domain generated by the classifier G y (·) can be obtained t . In this paper, it is used as a pseudo-label to construct a supervision medium and a target domain Self-SCL is constructed.

[0066] The specific implementation of the within-domain SSCL strategy is as follows: The source domain data and the target domain data (X s and X t respectively represent the data sample sequences of the source domain and the target domain in a mini-batch, which contain B source domain and target domain data samples x s and x t ) in each, are input into the shared feature encoder G E (·), and the source domain features F s =G E (X s ) and the target domain features F t =G E (X t ) are obtained, where and respectively represent the source domain and target domain feature sequences containing B source domain and target domain features f s and f t . Then they are input into the projector to obtain the feature representations and of the source domain and the target domain respectively, where Z s and Z t respectively represent the source domain and target domain feature representation sequences containing B source domain and target domain feature representations z s and z t . And the corresponding fault category labels are given, that is and , where y' t =G y (G E (x t)) represents inputting the target domain feature f t = G E (x t ) into the classifier G y (·) and obtaining the predicted value (i.e., pseudo-label). Then Z s and z t are respectively used to calculate the SCL loss of the source domain and the Self-SCL loss of the target domain:

[0067]

[0068] where k ∈ {1, 2,..., B}, k ≠ i ≠ j, SIM(·, ·) represents the similarity calculation between feature vectors, ||·||2 represents the 2-norm representation of feature vectors, τ is the temperature coefficient, and according to the empirical principle, it is set to 0.07. or represents, in the feature representation sequence of a mini-batch in the source domain or the target domain, the feature representation that belongs to the same fault category as the feature representation z s k or z t k , that is or their labels are the same. On the contrary, or represents, in the feature representation sequence of a mini-batch, the feature representation that belongs to different fault categories from the feature representation z s k or z t k , that is or their labels are different. Minimizing the SCL loss L SCL and the Self-SCL loss L Self-SCL can enable the diagnostic model to better define the classification boundary, which is beneficial for learning discriminative features for the classification task and also provides convenience for the transfer learning of cross-domain features. Specifically, the existence of a fuzzy classification boundary increases the risk of misalignment during cross-domain feature alignment. After the in-domain SSCL solves the problem of the fuzzy classification boundary, it reduces the probability of feature misalignment to a certain extent.

[0069] (2) Transfer learning strategy

[0070] While performing SSCL within the domain, this paper adopts the Local Maximum Mean Discrepancy (LMMD) to calculate the similarity between the source domain and the target domain, so as to guide the transfer learning of cross-domain features. LMMD aims to calculate the feature distribution corresponding to each fault category to make the distributions of each category similar, rather than the overall distribution similar. The designed SSCL strategy within the domain clarifies the classification boundaries between categories within each domain and enhances the distinguishability between categories, making the LMMD calculation have a more accurate basis for measuring distribution similarity, which provides convenience for the transfer learning of cross-domain features.

[0071] Suppose the feature representation sequences of a mini-batch within the source domain and the target domain are respectively and where Z s or Z t is a feature representation sequence containing B source domain or target domain feature representations Z s or z t and the corresponding fault labels y s or y′ t The feature representation sequence, y′ t = G y (G E (x t )) represents the predicted value obtained by inputting the target domain feature f t = G E (x t ) into the classifier G y (·), that is, the pseudo-label. Then, the LMMD loss calculation between Z s and Z t is as follows:

[0072]

[0073] where C represents the total number of fault categories (depending on the diagnostic task), c represents the fault category to which the feature representation Z s or z t belongs, represents the Reproducing Kernel Hilbert Space (RKHS), φ(·) represents the non-linear mapping from the original feature space to the RKHS, represents the feature kernel. and respectively represent the weights of the source domain feature representation z s i and the target domain feature representation z t j belonging to category c. It should be noted that the sum of the weights is 1, that is, and and is the weighted sum on category c. The feature representation zi The weight of is calculated as follows:

[0074]

[0075] where, represents the c-th entry of the vector obtained after one-hot encoding the class label y i For the source domain, the true label y s i is transformed into a one-hot vector for calculating the weight of each sample For the case where the true label in the unlabeled target domain is not available, the pseudo-label obtained by the classifier is one-hot encoded, and the weight of each sample is calculated

[0076] (3) Algorithm details

[0077] When calculating the Self-SCL loss and LMMD loss of the target domain, the pseudo-labels of the target domain are both used, which leads to the performance improvement effect of the above algorithm on the diagnostic model depending on the quality of the pseudo-labels. Therefore, from the perspective of reducing the impact of poor-quality pseudo-labels on model training and directly improving the quality of pseudo-labels, the following improvements are made to the above algorithm:

[0078] Improvement 1: Considering that the quality of the pseudo-labels output by the model in the early stage of training is unstable, to reduce the adverse impact of poor-quality pseudo-labels on model training, an increasing penalty coefficient λ is assigned to the Self-SCL loss and LMMD loss respectively increase to ensure the stability of the model by restricting their early gains.

[0079] Improvement 2: Considering that after reducing the early LMMD loss gain, the model training will focus more on the source domain rather than the target domain, which will lead to limited performance on the target domain. To improve the attention of the model training to the target domain and thus directly improve the quality of the target domain pseudo-labels, source domain features and target domain features are introduced for domain adversarial training to additionally guide the transfer learning of cross-domain features. The domain adversarial loss is calculated as follows:

[0080]

[0081] where, L CE represents the cross-entropy loss, G Adv (·) represents the domain adversarial module, i.e., the discriminator, and d k is the domain label corresponding to the feature f k ​

[0082] By combining the supervised representative learning loss \(L\) in the source domain SCL , the self-supervised contrastive learning loss \(L\) in the target domain Self-SCL , the Local Maximum Mean Discrepancy (LMMD) loss \(L\) LMMD and the domain adversarial loss \(L\) Adv , the proposed semi-supervised contrastive transfer learning (SSCTL) algorithm is obtained. The objective loss function \(L\) corresponding to the SSCTL algorithm SSCTL is as follows:

[0083]

[0084] where \(\mu\) is a fixed penalty coefficient that restricts the gain of the domain adversarial loss, and it is set to 0.5 according to empirical principles. epoch is the number of iterative training times. Based on empirical principles and through multiple experiments, the maximum number of iterative training times is set to 200, so epoch ∈ {0, 1, …, 199}. \(\lambda\) increase is an exponential increasing penalty coefficient, and its value increases from 0 to 1 as epoch increases.

[0085] Finally, a cross-condition diagnosis model based on the semi-supervised contrastive transfer learning network (SSCTLN) designed in the present invention is introduced in detail.

[0086] A cross-condition diagnosis model based on the semi-supervised contrastive transfer learning network (SSCTLN) is constructed by combining the constructed feature encoder and the designed semi-supervised contrastive transfer learning (SSCTL) algorithm. The specific structural parameters of SSCTLN are shown in Table 1.

[0087] Table 1 Structural parameters of SSCTLN

[0088]

[0089] The SSCTLN model has two optimization objectives during training: 1) Minimize the various contrastive learning and transfer learning losses shown in formula (8), 2) Minimize the supervised classification loss \(L\) on the labeled source domain y . For a mini-batch of source domain data samples The supervised classification loss of the source domain is defined as follows:

[0090]

[0091] In the present invention, the SSCTL loss and the supervised classification loss \(L\) of the source domain y are jointly optimized in an end-to-end manner. The joint objective loss function is shown in formula (11), and the Adam optimizer is used to update the network parameters.

[0092]

[0093] The hyperparameters during the training of the diagnostic model based on SSCTLN are set as follows according to the empirical principle: the batch size batch_size is set to 128, i.e., B = 128; the learning rate learning_rate is set to 0.001, and a staged decay strategy is adopted. The learning_rate is decayed at a decay efficiency of 0.1 when epoch = 119 and epoch = 159; the weight decay rate weight_decay is set to 3e-6.

[0094] Algorithm 1 Training Process of SSCTLN

[0095]

[0096]

[0097] The training process of the fault diagnosis model based on SSCTLN is shown in Algorithm 1, and the process of realizing cross-condition fault diagnosis is as follows:

[0098] (1) Sample the vibration signals of rotating machinery components under different working conditions, construct a labeled source domain and an unlabeled target domain, and divide the target domain into a training set and a test set.

[0099] (2) Set the network structure and parameters of the diagnostic model, as well as the hyperparameters related to model training, and initialize the network parameters.

[0100] (3) According to the model training process shown in Algorithm 1, use the labeled training set in the source domain and the unlabeled training set in the target domain to iteratively train the diagnostic model, minimize the target loss until the maximum number of iterative training times is reached, and then obtain the trained diagnostic model.

[0101] (4) Input the target domain test set into the trained diagnostic model, perform the cross-condition fault diagnosis task, and output the diagnostic results. At the same time, use relevant evaluation indicators and visualization techniques to evaluate the diagnostic performance of SSCTLN.

[0102] Aiming at the problem of limited performance caused by fuzzy classification boundaries when training a cross-operating rotating machinery fault diagnosis model, this paper proposes a diagnosis method based on SSCTLN. In SSCTLN, first, cross-operating fault diagnosis is achieved by adopting LMMD for feature transfer learning. Then, a domain SSCL strategy is designed by constructing SCL in the source domain and Self-SCL in the target domain to solve the fuzzy classification boundary problem existing in the training of the diagnosis model based on feature transfer learning, and at the same time facilitate feature transfer learning to improve the performance of the diagnosis model. Finally, considering that Self-SCL and LMMD in the target domain need to use target domain pseudo labels, their learning effects depend on the quality of pseudo labels. To this end, this paper limits their gains by assigning dynamic penalty terms to reduce the negative impact of poor-quality pseudo labels, and introduces domain adversarial to improve the quality of pseudo labels to improve the stability of the diagnosis model. In order to verify the rationality and effectiveness of the proposed SSCTLN, ablation experiments and comparative experiments were carried out on the cross-operating diagnosis task on the Paderborn University (PU) and wind turbine gearbox (WTG) datasets. The experimental results show that SSCTLN can effectively realize the fault diagnosis of rotating machinery under cross-operating scenarios. The cross-operating diagnosis accuracy on the PU and WTG datasets reaches 84.19% and 96.52%, respectively, as shown in Tables 2 and 3, and its diagnostic performance is better than that of popular feature transfer learning methods such as Figure 3 and 4 F1-score indicator comparison chart, Figure 5 and 6 The confusion matrix comparison chart of Figure 7 and 8 The feature visualization comparison is shown in the figure.

[0103] Table 2 Average accuracy on PU dataset (%)

[0104]

[0105] Table 3 Average accuracy (%) on the WTG dataset.

[0106]

Claims

1. A fault diagnosis method for rotating machinery based on transfer learning that integrates semi-supervised contrastive learning in the fusion domain, characterized in that: The method includes: S1 designing an improved one-dimensional convolutional neural network (1D-CNN) to construct a feature encoder; S2 designing a semi-supervised contrastive learning (SSCL) algorithm in the domain, which includes supervised contrastive learning (SCL) in the source domain and self-supervised contrastive learning (Self-SCL) in the target domain; S3 using local maximum mean discrepancy (LMMD) for cross-domain feature transfer learning. Aiming at the fuzzy classification boundary problem existing in the feature transfer learning process, S2 is adopted to enable the transfer diagnosis model to specifically learn the discriminative features between different fault categories to eliminate the fuzzy classification boundary, and at the same time facilitate the transfer learning of cross-domain features. Finally, a semi-supervised contrastive transfer learning (SSCTL) module is constructed; S4 optimizing the SSCTL module constructed in S3, and finally constructing a semi-supervised contrastive transfer learning network (SSCTLN) to realize the fault diagnosis of rotating machinery components under different working conditions.

2. A fault diagnosis method for rotating machinery transfer learning integrating semi-supervised contrastive learning in the fusion field, characterized in that: S1 designing an improved one-dimensional convolutional neural network (1D-CNN) to construct a feature encoder, and the steps are as follows: Step 1: When constructing the 1D-CNN, a one-dimensional convolutional layer consistent with ERCM was selected, and two convolutional layers were stacked in series without an activation function in the middle; different from ERCM, the constructed 1D-CNN discarded the residual structure; Step 2: Replace the ReLU activation function in ERCM with the Swish activation function, and this activation function is shown in formula (1); then, a convolutional block with the structure of Conv1d-BN-Conv1d-BN-Swish_act-MaxPool was obtained, and a 1D-CNN was constructed by stacking a convolutional block of Conv1d-BN-Conv1d-BN-Swish_act-AdaMaxPool and four of the above convolutional blocks; where Conv1d is one-dimensional convolution and BN is batch normalization; Swish_act(x) = x * Sigmoid(x) (1) Among them, x represents the features of the layer before the activation function extracted in the convolutional block, and Sigmoid is also an activation function, as shown in formula (2); Step 3: Adopt the Convolutional Block Attention Module (CBAM) with feature adaptive correction ability to enhance the key information in the fault features; it combines channel attention and spatial attention; finally, the constructed 1D-CNN is concatenated with CBAM to form the feature encoder G E (·).

3. A fault diagnosis method for rotating machinery transfer learning integrating semi-supervised contrastive learning in the fusion field, characterized in that: S2 designing a semi-supervised contrastive learning (SSCL) algorithm in the domain, and the steps are as follows: First, for the cross-condition migration fault diagnosis scenario, the data collected under the condition of labeled data is defined as the source domain with n s labeled samples The data collected under the condition of unlabeled data is defined as the target domain with n t unlabeled samples where x and y represent samples and labels respectively; since the data is collected from different conditions, there will be differences between cross-domain data and its distribution, that is, X s ≠ X t and Then, construct the SSCL strategy within the domain, and its specific implementation is as follows: The source domain data in a mini-batch and the target domain data X s and X t respectively represent the data sample sequences of the source domain and the target domain in a mini-batch, which respectively contain B source domain and target domain data samples x s and x t ; Input them into the shared feature encoder G E (·), and obtain the source domain feature F s = G E (X s ) and the target domain feature F t = G E (X t ), where and respectively represent the source domain and target domain feature sequences containing B source domain and target domain features f s and f t ; Then input them into the projector to obtain the feature representations of the source domain and the target domain respectively and where Z s and Z t respectively represent the source domain and target domain feature representation sequences containing B source domain and target domain feature representations z s and z t ; And give the corresponding labels, that is, and where y' t = G y (G E (x t )) represents the predicted value (i.e., pseudo-label) obtained after inputting the target domain feature f t = G E (x t ) into the classifier G y (·); Then Z s and Z t are respectively used to calculate the SCL loss of the source domain and the Self-SCL loss of the target domain: where \(k\in\{1,2,\ldots,B\}\), \(k\neq i\neq j\), \(sim(\cdot,\cdot)\) represents the similarity calculation between feature vectors, \(\|\cdot\|_2\) represents the 2-norm representation of feature vectors, and \(\tau\) is the temperature coefficient, set to 0.07; or represents that in the feature representation sequence of a mini-batch in the source domain or target domain, the feature representation \(z\) s k or \(z\) t k belong to the feature representations of the same fault class, that is or their labels are the same; on the contrary, or represents that in the feature representation sequence of a mini-batch, the feature representation \(z\) s k or \(z\) t k belong to the feature representations of different fault classes, that is or their labels are different; minimizing the SCL loss \(L\) SCL and the Self-SCL loss \(L\) Self-SCL can make the diagnostic model clarify the classification boundary.

4. A fault diagnosis method for rotating machinery transfer learning integrating semi-supervised contrastive learning in the fusion field, characterized in that: S3 is specifically as follows: While performing SSCL within the domain, Local Maximum Mean Discrepancy (LMMD) is adopted to calculate the similarity between the source domain and the target domain, so as to guide the transfer learning of cross-domain features; the feature representation sequences of a mini-batch within the source domain and the target domain are respectively and where Z s or Z t is a feature representation sequence that contains B source domain or target domain feature representations z s or z t and the corresponding fault labels y s or y′ t The feature representation sequence of y′ t = G y (G E (x t )) represents the predicted value obtained after inputting the target domain feature f t = G E (x t ) into the classifier G y (·), that is, the pseudo label; then, the LMMD loss between Z s and Z t is calculated as follows: Where C represents the total number of fault categories, and c represents the fault category to which the feature representation z s or z t belongs; represents the reproducing kernel Hilbert space (RKHS), φ(·) represents the non-linear mapping from the original feature space to the RKHS, represents the feature kernel; and represent the source domain feature representation z s i and the target domain feature representation z t j are the weights belonging to category c; the sum of the weights is 1, i.e., and and is the weighted sum over category c; the weight of the feature representation z i is calculated as follows: Among them, represents the c-th entry of the vector obtained after one-hot encoding the class label y i ; for the source domain, the true label y s i is converted into a one-hot vector for calculating the weight of each sample For the case where the true label of the unlabeled target domain is not available, the pseudo-label obtained by the classifier is one-hot encoded and then the weight of each sample is calculated 5. A fault diagnosis method for rotating machinery transfer learning that integrates semi-supervised contrastive learning in the fusion field, characterized in that: S4 is specifically as follows: Step 1: Considering that the pseudo-labels of the target domain are used when calculating the Self-SCL loss and LMMD loss of the target domain, the following improvements were made to the above algorithm: Improvement 1: An increasing penalty coefficient λ is assigned to the Self-SCL loss and the LMMD loss respectively increase to ensure the stability of the model by restricting their early gains; Improvement 2: Introduce source domain features and target domain features for domain adversarial training to additionally guide the transfer learning of cross-domain features; the domain adversarial loss is calculated as follows: Among them, L CE represents the cross-entropy loss, and G Adv (·) represents the domain adversarial module, i.e., the discriminator, and d k is the domain label corresponding to the feature f k ; by combining the supervised representative learning loss L SCL of the source domain, the self-supervised contrastive learning loss L Self-SCL of the target domain, the local maximum mean discrepancy (LMMD) loss L LMMD and the domain adversarial loss L Adv , the proposed semi-supervised contrastive transfer learning (SSCTL) algorithm is obtained. Combining the feature encoder G E (·) and the SSCTL algorithm, a cross-condition diagnosis model based on the semi-supervised contrastive transfer learning network (SSCTLN) is constructed; the objective loss function corresponding to the SSCTL algorithm is as follows: Among them, μ is a fixed penalty coefficient that restricts the gain of the domain adversarial loss, and is set to 0.5; epoch is the number of iterative training times, and the maximum number of iterative training times is set to 200, so epoch ∈ {0, 1, …, 199}; λ increase is an exponential increasing penalty coefficient; Step 2: The SSCTLN model has two optimization objectives during training: 1) Minimize the contrastive learning and transfer learning losses shown in Equation (8), and 2) Minimize the supervised classification loss \(L\) on the labeled source domain. y For a mini-batch of source domain data samples The supervised classification loss on the source domain is defined as follows: The SSCTL loss and the supervised classification loss of the source domain are jointly optimized in an end-to-end manner. The joint objective loss function is shown in formula (11), and the Adam optimizer is used to update the network parameters; The hyperparameters during the training of the diagnostic model based on SSCTLN are set as follows: the batch size batch_size is set to 128, i.e., B = 128; the learning rate learning_rate is set to 0.001, and a stepwise decay strategy is adopted. The learning_rate is decayed at a decay efficiency of 0.1 at epoch = 119 and epoch = 159; the weight decay rate weight_decay is set to 3e-6.

Citation Information

Cited By

  • Boundary-guided rolling bearing semi-supervised fault diagnosis method and system

    CN121051573A

  • A boundary-guided semi-supervised fault diagnosis method and system for rolling bearings

    CN121051573B

  • Rotating machinery cross-domain fault diagnosis method based on dynamic evolution and wavelet two-way structure

    CN121256489A

  • Rotating machine fault classification method based on semi-supervised transfer learning

    CN121542878A