Cross-Domain Fault Diagnosis Method for Rolling Bearings Based on Self-Supervised Learning
Through the self-supervised learning dual-classifier domain adaptation method, the cross-domain fault diagnosis problem of rolling bearings under variable operating conditions is solved, the accuracy of feature extraction and classification is improved, and the diagnostic accuracy in the target domain is enhanced.
Patent Information
- Application Number
- CN202310963119.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-02
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-08-02
AI Technical Summary
The existing deep learning-based cross-domain fault diagnosis method of rolling bearings is difficult to effectively extract the feature distribution of the target domain under variable working conditions, resulting in low diagnostic accuracy, and the existing domain adaptation methods have deviations and inaccuracies in feature extraction and classifier output.
Using a dual classifier domain adaptation method based on self-supervised learning, by constructing a tag source domain and a tagless target domain, combining cross entropy loss and target clustering loss, dual classifier combination and individual classification accuracy loss are introduced, feature extractor and classifier parameters are optimized, and end-to-end fault diagnosis is achieved.
The fault diagnosis accuracy of rolling bearings under variable working conditions is improved, the domain adaptation model pays attention to the target domain data, the accuracy of feature extraction and classification is improved, and the diagnostic performance is achieved.
Smart Images

Figure CN116975718B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of rolling bearing fault diagnosis, and particularly relates to a cross-domain fault diagnosis method for rolling bearings based on self-supervised learning. Background Art
[0002] Rolling bearings are one of the indispensable basic transmission components in modern mechanical equipment, and their health status will directly affect the normal operation of mechanical equipment. In complex mechanical systems, the condition monitoring and fault diagnosis of rolling bearings are of great significance for improving industrial production efficiency and reducing the accident rate. Traditional fault diagnosis methods can be roughly divided into two categories: model-based methods and signal processing-based methods. Due to the random factors and noises in the actual equipment working environment being difficult to estimate in advance, it is often difficult to construct an effective mathematical model for model-based methods. And signal processing-based methods rely on certain expert prior knowledge and are difficult to process a large amount of data.
[0003] In recent years, due to the rapid development of artificial intelligence technology and the continuous accumulation of a large amount of monitoring data, data-driven intelligent fault diagnosis technology has become a research hotspot. Rolling bearing fault diagnosis methods represented by deep learning (DL) can automatically extract features from the collected data and achieve an end-to-end diagnosis mode, and have now achieved remarkable results. However, the premise for the good performance of DL-based fault diagnosis methods is to meet two key conditions: rich labeled data and independent and identically distributed between training and test data. Due to the distribution differences between data caused by changes in working conditions and the influence of the working environment, it is difficult for DL-based methods to train a robust fault diagnosis model with high generalization performance.
[0004] Transfer learning attempts to transfer the trained knowledge from the labeled domain (i.e., the source domain) to a similar domain (i.e., the target domain) without labels, but sharing semantic information to assist deep learning in overcoming the limitations mentioned above. As one of the representative methods of transfer learning, the unsupervised domain adaptation (UDA) method bridges the domain gap between the source domain and the target domain by mining domain-invariant feature representations. To achieve this task, existing domain adaptation methods are divided into two main categories. One is to use certain specific metrics to align the data feature distributions; the other is inspired by the generative adversarial network (GAN) and uses the idea of adversarial learning to enhance the feature extraction ability of the network model, thereby obtaining domain-invariant feature representations. Existing UDA methods based on adversarial learning are mainly implemented through two paradigms: introducing an additional domain discriminator or using two independent classifiers for learning. The domain discriminator can improve the transferability of features by confusing the feature extractor and the domain discriminator, but it ignores the class information of the target samples, resulting in the deterioration of feature discriminability. By using the difference metric between the two classifiers and through the min-max game of the output difference between the feature extractor and the classifier, the decision boundary of the model is extended, achieving feature alignment while maintaining sample discriminability.
[0005] It is worth noting that in the field of bearing fault diagnosis, there are still certain deficiencies in the existing dual-classifier adversarial domain adaptation methods. For example, the network trained with source domain data often tends to be more biased towards the source domain, resulting in certain deviations in the obtained target features. In this state, it is difficult to extract better adaptive features from the feature distributions of rolling bearing fault data adapted to the two domains. In addition, this type of method assumes that consistency is equivalent to discriminability in all cases. In fact, the two classifiers can consistently make ambiguous distinctions about the samples, resulting in the network model misclassifying the target samples.
[0006] Currently, UDA methods used in the field of intelligent fault diagnosis are roughly divided into two main categories. The difference-based method is one of the more popular ones. It reduces the distribution difference between source domain and target domain samples through distance metrics at the feature layer of the model, thereby alleviating the problem of low cross-domain fault diagnosis accuracy. Its basic network framework is as Figure 1 (a) shown. For example, the maximum mean discrepancy (MMD), as a classic distance metric in the field of domain adaptation, has been widely used in the intelligent fault diagnosis of bearings. Since then, various variants of MMD have emerged in an endless stream. The JMMD metric was proposed to measure the joint distribution relationship between sample data to achieve more comprehensive domain adaptation. To achieve the transformation from global alignment to local alignment, the LMMD metric was proposed to measure the distribution difference of the subdomains related to the transferable features.
[0007] Another type of domain adaptation method has a similar working idea to generative adversarial networks (GANs), mainly focusing on learning domain-invariant feature representations of the source domain and the target domain in an adversarial manner. The domain adversarial neural network (DANN) was the first to introduce the adversarial idea into the field of domain adaptation, and its basic network framework is as shown in Figure 1 (b). The key to this type of method lies in whether the constructed domain discriminator can identify whether the input samples come from the source domain or the target domain through a special gradient reversal layer (GRL) structure. By training the feature extractor of the network to confuse the domain discriminator, domain-invariant features can be extracted during the adversarial process. Due to its strong flexibility and robustness, it has been successfully applied to mechanical fault diagnosis.
[0008] The above introduction of the domain discriminator is one of the domain adaptation methods based on adversarial learning, while the other is the dual-classifier adversarial paradigm. By introducing an additional classifier to compare the output differences for the same sample, it guides the optimization of the feature extractor and realizes the expansion of the decision boundary of the target model. Its basic network framework is as shown in Figure 1 (c). MCD was the first method to propose the dual-classifier paradigm, using the L1 distance as the measurement standard for the difference between classifiers. Later, SWD used the sliced Wasserstein distance to obtain the output consistency from a geometric perspective, improving the computational efficiency of the network through the optimal transport path. Recently, BCDM designed a new classifier determinacy difference (CDD) metric to focus on the discriminability of the outputs between classifiers. In the field of bearing fault diagnosis, scholars have also introduced domain adaptation methods based on the dual-classifier paradigm.
[0009] Although the two types of domain adaptation methods based on adversarial learning have been applied to the field of intelligent fault diagnosis, they still have corresponding defects. The domain adaptation method based on the domain adversarial paradigm has a certain degree of deterioration in the discriminative characteristics of the extracted features because the classifier is not involved in the adversarial process. The domain adaptation method based on the dual-classifier adversarial paradigm often uses the difference metric between classifiers to obtain the output consistency of samples, but the output consistency does not represent the accuracy of the output. The same output results of the two classifiers may both be incorrect.
[0010] In the bearing fault diagnosis based on the dual-classifier paradigm, the feature extractor F is used to extract the deep features of the input bearing samples, and two classifiers C1 and C2 with the same structure are designed to distinguish the extracted features and confuse the feature extractor F. The training process of a typical dual-classifier domain adaptation model consists of the following three steps.
[0011] S1: Initialization training. Using the labeled source domain data, the model is trained by minimizing the cross-entropy loss to optimize the parameters of F, C1, and C2, so that they have initial generalization and discriminative capabilities respectively. The target optimization function is obtained through Equation a:
[0012]
[0013] Among them, N s represents the total number of source domain samples, represents the softmax output of the classifier for the i-th sample in the source domain, c represents the bearing samples with a total of c categories, represents the indicator vector of the source domain label, represents the standard cross-entropy loss function.
[0014] S2: Narrow the decision boundary between classifiers. Fix the feature extractor F, and update the parameters of classifiers C1 and C2 to narrow the decision boundary between classifiers. At the same time, in order to maintain the accuracy of the model on the source domain, the classification loss of the source domain is still added in this step. This step is usually achieved by designing specific metrics, such as L1 distance or sliced Wasserstein distance, to maximize the difference between the outputs of the two classifiers. The target optimization function is obtained by Equation b:
[0015]
[0016] Among them, represents the metric for measuring the difference in the outputs of the two classifiers. and represent the output vectors of classifiers C1 and C2 for the same target domain sample.
[0017] S3: Adjust the data distribution. Fix classifiers C1 and C2, and update the feature extractor F so that the feature of the target domain sample is concentrated in the spatial distribution within the decision boundary of the classifier. The classification of the data is achieved by minimizing the output difference between classifiers C1 and C2, and the target optimization function is obtained by Equation c:
[0018]
[0019] Combined with Figure 2 it can be seen that in the three-step training strategy based on the dual-classifier paradigm, in the first step, the network model is trained only by the cross-entropy loss of the source domain data. When the data distribution difference between the source domain and the target domain is large in the rolling bearing fault diagnosis, its generalization ability on the target domain cannot be guaranteed. Moreover, in the second step, the model parameters are directly updated on the target domain by the output difference between classifiers, which largely depends on the initial discrimination ability of the model trained by the source domain on the target domain.
[0020] Due to the lack of true target domain sample labels in unsupervised domain adaptation learning, it is difficult for the decision boundaries of the two classifiers to distinguish the low data density within the target domain, which may lead to the deterioration of the discriminability of the features between samples.
[0021] In summary, there is no suitable method for cross - domain fault diagnosis of rolling bearings in the existing technology. Summary of the Invention
[0022] The present invention provides a cross - domain fault diagnosis method for rolling bearings based on self - supervised learning to solve the problem of cross - domain fault diagnosis of rolling bearings under variable working conditions.
[0023] In order to solve the above - mentioned technical problems, the technical solution of the present invention is as follows: The cross - domain fault diagnosis method for rolling bearings based on self - supervised learning includes the following steps:
[0024] S1: Construct a labeled source domain and an unlabeled target domain from the original vibration signals collected under different working conditions, and randomly divide them into a training set and a test set in a ratio of 8:2.
[0025] S2: Input the training set data into the network model, calculate the overall loss function of the network three times in sequence, and update the overall parameters of the network, the parameters of the classifier, and the parameters of the feature extractor through the backpropagation algorithm in sequence.
[0026] S3: Repeat step S2 until the maximum number of iterations, stop training, and generate a domain - adaptation network with good generalization performance.
[0027] S4: Input the test set data in step S1 into the trained network model and output the diagnosis results under different working conditions.
[0028] In a preferred embodiment of the present invention, the domain - adaptation network structure includes: a feature extractor based on a five - layer convolutional block and two classifiers with the same structure. The feature extractor extracts transferable features from the source domain and target domain data, and the two classifiers can obtain the deep features generated by the feature extractor and output the corresponding prediction probabilities.
[0029] In a preferred embodiment of the present invention, the method for constructing the domain - adaptation network model in step S1 is as follows:
[0030] First, introduce the original vibration signals sampled from the source domain and the target domain into the feature extractor.
[0031] Then, two binary classifiers use the cross - entropy loss to classify the source samples, and at the same time embed an SSL loss based on target clustering.
[0032] Finally, construct a min - max game, and introduce the joint and individual classification accuracy losses of the binary classifiers to guide the output of the binary classifiers.
[0033] In a preferred embodiment of the present invention, step S1 specifically includes:
[0034] S11: Collect the vibration data of rolling bearings under different working conditions, construct the labeled source domain and the unlabeled target domain, randomly divide 80% of the target domain data into the training set, and the remaining 20% into the test set;
[0035] S12: Set the maximum number of iterations of the network, the size and stride of the convolutional kernels in the convolutional layer, set the size and stride of the pooling kernels in the pooling layer, and initialize the weight parameters of the convolutional layer and the fully connected layer.
[0036] In a preferred embodiment of the present invention, step S2 specifically includes:
[0037] S21: Input the source domain training set data, the corresponding true labels, and the target domain training set data, calculate the pseudo-labels of the target domain samples, and calculate the total loss function of the network at this time Update the overall parameters of the network using the backpropagation algorithm;
[0038] S22: Input the source domain training set data, the corresponding true labels, and the target domain training set data, and calculate the total loss function of the network at this time Update the classifier parameters using the backpropagation algorithm;
[0039] S23: Input the source domain training set data, the corresponding true labels, and the target domain training set data, and calculate the total loss function of the network at this time Update the feature extractor parameters using the backpropagation algorithm.
[0040] In a preferred embodiment of the present invention, in step S21, the total loss function of the network is obtained through Equation 1:
[0041]
[0042] Where, N t represents the total number of target domain samples, represents the j-th sample in the target domain, represents the indicator vector of the target domain pseudo-label; c represents the bearing samples of a total of c categories, represents the softmax output of the classifier for the j-th sample in the target domain; is the pseudo-label, which is obtained through Equation 2:
[0043]
[0044] Where, represents the j-th sample in the target domain; F(·) represents the feature vector of the sample after passing through the feature extractor, cen c represents the centroid of the target domain data of the c-th category.
[0045] In a preferred embodiment of the present invention, the centroid of the target domain data of category c is obtained by Equation 3:
[0046]
[0047] where δ c represents the prediction for the c-th category after the softmax operation, represents the j-th sample of the target domain; F(·) represents the feature vector of the sample after passing through the feature extractor, and C(·) represents the output vector after passing through the classifier; C k (k = 1, 2) are two task classifiers.
[0048] In a preferred embodiment of the present invention, the loss functions in steps S22 and S23 are calculated by Equation 4:
[0049]
[0050] where μ and ω are the balance parameters between the joint classification certainty and the individual classification certainty respectively; is a metric for the joint classification accuracy of the dual classifiers; is the individual classification accuracy index output by the classifier.
[0051] In a preferred embodiment of the present invention, the individual classification accuracy index output by the classifier in Equation 4
[0052] is calculated by Equation 5:
[0053]
[0054] where δ c represents the prediction for the c-th category after the softmax operation, F(·) represents the feature vector of the sample after passing through the feature extractor, and C(·) represents the output vector after passing through the classifier; C k (k = 1, 2) are two task classifiers.
[0055] In a preferred embodiment of the present invention, the metric for the joint classification accuracy of the dual classifiers in Equation 4 is calculated by Equation 6:
[0056]
[0057] where δ c represents the prediction for the c-th category after the softmax operation, F(·) represents the feature vector of the sample after passing through the feature extractor, and C(·) represents the output vector after passing through the classifier; C k (k = 1, 2) are two task classifiers.
[0058] The technical solution provided by the present invention has the following advantages compared with the prior art:
[0059] The present invention proposes a metric for measuring the classification accuracy of a dual classifier. This metric combines the individual classification accuracies of the classifiers and the joint classification accuracy between the classifiers, further enhancing the discriminability of samples. It enhances the attention of the domain adaptation model to the target domain data and performs self-supervised learning to guide the optimization of the target domain, so as to improve the fault diagnosis accuracy of rolling bearings under variable working conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0061] Figure 1 is a network framework diagram of a domain adaptation method in the prior art;
[0062] Figure 2 is a training flow chart based on the dual classifier paradigm in the prior art;
[0063] Figure 3 is a flow chart of a rolling bearing cross-domain fault diagnosis method based on self-supervised learning in an embodiment of the present invention;
[0064] Figure 4 is a structural diagram of a dual classifier domain adaptation network in a rolling bearing cross-domain fault diagnosis method based on self-supervised learning in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] For ease of understanding, the following describes the rolling bearing cross-domain fault diagnosis method based on self-supervised dual classifier domain adaptation in combination with embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0066] For ease of understanding of the present invention, the present invention will be described more comprehensively with reference to the relevant drawings. The preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0067] Aiming at the problem that traditional fault diagnosis methods are difficult to improve the diagnosis performance under variable working conditions, a new fault diagnosis method based on self-supervised learning dual-classifier domain adaptation is proposed. The collected original vibration signals are directly used as the input of the diagnosis model to realize end-to-end fault diagnosis from signal input to diagnosis output. By simultaneously considering the joint classification accuracy between classifiers and the individual classification accuracy of each classifier, the performance of fault diagnosis is improved. In addition, self-supervised training is introduced into the domain adaptation fault diagnosis method based on dual-classifiers to enhance the domain adaptation learning ability.
[0068] Compared with the conventional dual-classifier domain adaptation method, the dual-classifier classification accuracy metric proposed in this invention effectively solves the problem of the consistent ambiguous discrimination of samples by the two classifiers. At the same time, the introduction of self-supervised learning can make full and effective use of the pseudo-label information of the target domain and can learn the data features of the target domain more efficiently. Applying the dual-classifier domain adaptation method based on self-supervised learning to the fault diagnosis of rolling bearings under variable working conditions can improve the ability to extract bearing fault features and further enhance the fault diagnosis performance.
[0069] Specifically, referring to Figure 3 As shown, a cross-domain fault diagnosis method for rolling bearings based on self-supervised dual-classifier domain adaptation proposed in this invention includes the following steps:
[0070] S1: Construct a labeled source domain and an unlabeled target domain from the original vibration signals collected under different working conditions, and randomly divide them into a training set and a test set in a ratio of 8:2.
[0071] S11: Collect the vibration data of rolling bearings under different working conditions, construct a labeled source domain and an unlabeled target domain, and randomly divide 80% of the target domain data into the training set and the remaining 20% into the test set.
[0072] S12: Set the maximum number of iterations of the network, the size and stride of the convolutional kernels in the convolutional layer, set the size and stride of the pooling kernels in the pooling layer, and initialize the weight parameters of the convolutional layer and the fully connected layer.
[0073] The overall structure of the constructed domain adaptation network is as Figure 4 shown. The input of this network is composed of the labeled source domain vibration signals and the unlabeled target vibration signals. The network structure mainly includes: 1) A feature extractor F based on a five-layer convolutional block to extract transferable features from the source domain and target domain data; 2) Two classifiers C1 and C2 with the same structure to obtain the depth features generated by F and output the corresponding prediction probabilities P1 and P2 to realize the classification of bearing faults.
[0074] Specifically, the method for constructing the domain adaptation network model is:
[0075] First, the original vibration signals sampled from the source domain and the target domain are introduced into the feature extractor F to learn high-level feature representations;
[0076] Then, the two classifiers C1 and C2 use the cross-entropy loss to classify the source samples, and at the same time embed an SSL loss based on target clustering to learn more semantic information and improve the bearing fault diagnosis performance;
[0077] Finally, a min-max game is constructed, and the joint and individual classification accuracy losses of the two classifiers are introduced to guide the output of the two classifiers, thus correctly guiding the model optimization.
[0078] S2: Input the training set data into the network model, calculate the overall loss function of the network three times in sequence, and update the overall network parameters, classifier parameters, and feature extractor parameters through the backpropagation algorithm in sequence.
[0079] S21: Input the source domain training set data, the corresponding true labels, and the target domain training set data, calculate the pseudo-labels of the target domain samples corresponding thereto, and calculate the overall loss function of the network at this time based on this Update the overall network parameters using the backpropagation algorithm.
[0080] The present invention proposes a strategy based on data weighted clustering to obtain more reliable target pseudo-labels. The samples are weighted by the softmax outputs of the two classifiers, and the centroid of the target domain data of the c-th category is obtained through Equation 3:
[0081]
[0082] where δ c represents the prediction for the c-th category after the softmax operation, represents the j-th sample in the target domain. F(·) represents the feature vector of the sample after passing through the feature extractor, and C(·) represents the output vector after passing through the classifier. C k (k = 1, 2) are the two task classifiers. Then the pseudo-label is calculated using the nearest centroid strategy through Equation 2:
[0083]
[0084] In order to make full use of the unlabeled label information of the target domain samples, on the basis of the original source domain supervised learning, a self-supervised learning mechanism is adopted to strengthen it, so that the network model has a certain discrimination ability on the target domain at the initial stage of learning. The loss of self-supervised learning is obtained through Equation 1:
[0085]
[0086] where, N t represents the total number of samples in the target domain, and
[0087] represents the indication vector of the target domain pseudo - labels. By combining the self - supervised process of the target domain data with the supervised process of the source domain data, it essentially promotes better alignment of samples between the two domains at the class level.
[0087] S22: Input the source domain training set data, the corresponding true labels, and the target domain training set data, and calculate the total loss function of the network at this time Update the classifier parameters using the backpropagation algorithm.
[0088] The output difference metric is a key feature of the dual - classifier paradigm. A good difference metric can better assist in optimizing the network model through the outputs between classifiers. The key lies in that during the min - max adversarial process of the network, the generalization performance of the feature extractor is greatly enhanced.
[0089] Ideally, when the outputs between the two classifiers are consistent, it can be considered that the model can correctly classify the target samples. However, this situation ignores the certainty of the output, which may lead to the problem of ambiguous output, that is, the probabilities of each class are equal. Obviously, this is not helpful for classification. If the outputs of classifiers C1 and C2 are consistent and are certain for the same sample, then the model can correctly identify the sample. Based on this, a metric for the accuracy of the dual - classifier is proposed, which is obtained through Equation 6:
[0090]
[0091] On the one hand, in terms of consistency, when the difference is large, the predictions for the same class must be relative (i.e., one large and one small), so the joint classification accuracy is small; on the other hand, in terms of accuracy, when the difference is small but the classification is ambiguous, a small joint classification certainty will also be generated. Therefore, this metric can consider both the consistency and accuracy between the outputs of the two classifiers simultaneously. The square - root constraint is used to accelerate the network convergence speed and prevent overfitting.
[0092] In addition, when there are large differences between the two classifiers, there may still be the problem of ambiguous output, especially for the predicted classes with the highest confidence. For example, if the predictions of the two classifiers are [0.8, 0.1, 0.1] and [0.2, 0.5, 0.3], their outputs can be optimized in the direction of [1 / 3, 1 / 3, 1 / 3] simultaneously to increase the joint classification accuracy, but this leads to a decrease in the individual discrimination ability of each classifier. Therefore, the individual accuracy of each classifier's output must be considered, which is obtained through Equation 5:
[0093]
[0094] Maximizing the joint classification accuracy between classifiers and the individual classification accuracy of each classifier can better ensure the output of the classifier, guide the optimization of the model, and ultimately achieve accurate recognition of the model in the target domain. The final classification certainty loss can be calculated by Equation 4:
[0095]
[0096] where μ and ω are the balance parameters between the joint classification certainty and the local classification certainty, respectively.
[0097] S23: Input the source domain training set data, the corresponding true labels, and the target domain training set data, and calculate the total loss function of the network at this time. Update the parameters of the feature extractor using the backpropagation algorithm.
[0098] S3: Repeat step S2 until the maximum number of iterations, stop training, and generate a domain adaptation network with good generalization performance.
[0099] S4: Input the test set data in step S1 into the trained network model and output the diagnostic results under different working conditions.
[0100] In summary, supervised adversarial training combines self-supervised learning optimization to achieve high-performance unsupervised domain adaptation of a dual classifier.
[0101] The present invention verifies the effectiveness of the proposed unsupervised domain adaptation method through two rolling bearing data sets. By comparing with other different classical transfer methods, its good superiority is further demonstrated. The following is the experimental verification of the present invention.
[0102] All experimental methods are written in the Pytorch environment and run on a computer configured with an NVIDIA GeForce RTX3060 GPU and an Intel i5-12500H CPU.
[0103] A well-known publicly available dataset provided by the University of Paderborn is widely used in the research of rolling bearing fault diagnosis. This dataset uses the 6203 type ball bearing as the test object and includes two major categories: artificial damage and real damage. In order to be closer to the application in actual industrial scenarios, the present invention selects real bearing damage samples generated by accelerated life tests. The test bench mainly consists of a motor, a torque measurement shaft, a rolling bearing test module, a flywheel, and a load motor. Table 1 describes the detailed parameters of the specific faulty bearings, which are divided into 13 categories of faults in total. According to different operating conditions of drive speed, radial force, and load torque, the PU dataset is set to four working conditions, and the sampling frequency is 64 kHz for all. Each experiment selects two of the four working conditions as the source domain and the target domain. For example, task 0-1 means collecting source domain and target domain data from working conditions 0 and 1 respectively. Obviously, there are distribution differences in the samples collected under different working conditions, but they are all in the same label space. Therefore, the diagnostic knowledge learned from the data under one working condition can be transferred to another working condition. A total of 6 transfer tasks are set in this experiment as shown in Table 2.
[0104] Table 1: Detailed Parameters of Faulty Bearings in the PU Dataset
[0105]
[0106]
[0107] Table 2: 6 Transfer Tasks under Different Working Conditions of PU
[0108]
[0109] In this experiment, the samples truncated by the sliding window sampling method do not overlap, and the length of each sample is 1024. According to the total length of each type of original vibration signal, 250 samples of each type in each task are collected. Therefore, the total number of samples in each task is N = 3250. In each experiment, 80% of the sampled data in the target domain is randomly divided into the training set, and the remaining 20% is divided into the test set. All samples are normalized using Z-Score before being input into the network to unify the data size and reduce the differences. The calculation process is shown in Equations 7 - 9:
[0110]
[0111]
[0112]
[0113] In all experiments, the number of iterations and the batch size were set to 15000 and 32, respectively. During network training, the SGD optimizer was used to update the model parameters. The initial learning rate was set to 0.01, and a dynamic learning rate update technique was adopted. The update algorithm is shown in Equation 10. The momentum was set to 0.9, and the weight decay was set to 5e-4.
[0114]
[0115] where lr0 is the initial learning rate, i is the real-time iteration number, and it is the total number of iterations.
[0116] To further verify the superiority of the proposed method under variable working conditions, eight other classic transfer methods were used for comparative experiments. In particular, these comparative methods are mainly the three most common paradigms of UDA. As a general model for supervised learning, using 1D-CNN as the baseline method can fully demonstrate the advantages of UDA algorithms. For methods based on metric differences, the present invention selected the DAN method based on the MMD metric and the DSAN method based on the LMMD metric. For methods based on domain adversarial training, the present invention selected the DANN and DCTLN methods; and for methods based on dual-classifier adversarial training, the present invention selected the MCD, SWD, and BCDM methods. Table 3 summarizes the details of the network architectures, transfer types, and key points of all comparative methods. For a fair comparison, the feature extractors and classifiers of all methods shared the same network structure. In addition, the domain discriminators of DANN and DCTLN are binary classifiers.
[0117] Table 3: Structural description of comparative methods
[0118]
[0119] To reduce the influence of experimental randomness, each type of task was repeated 10 times to obtain the average diagnostic accuracy. The experimental results of the proposed domain adaptation method and other 8 comparison methods are shown in Table 4. It can be seen from Table 4 that the proposed method achieved the highest diagnostic accuracy in all 6 tasks, and the average accuracy rate of the 6 types of tasks reached 84.77%, fully demonstrating the superior robustness and reliability of this method when dealing with different transmission tasks. Due to the lack of a domain adaptation process, the traditional 1D-CNN method was unable to extract domain-invariant features of each domain and reduce the differences between the two domains, and its performance in various tasks was the worst, with an average diagnostic accuracy of only 53.96%. This indicates that it is difficult to solve the fault classification problem of rolling bearings under variable working conditions only through deep learning models. It can be seen from the table that it is obvious that the UDA methods based on metric differences and the UDA methods based on domain adversarial training perform significantly lower than the UDA method based on dual-classifier adversarial training. Among various methods based on dual-classifier adversarial training, in tasks 0→3 and 1→3, the diagnostic accuracy of the proposed method far exceeded that of the same type of methods. In addition, it can be found that tasks 0→1 and 1→3 are relatively difficult tasks in the PU dataset, but the proposed method still had the best performance.
[0120] Table 4: Average Diagnostic Accuracy of the PU Dataset (%)
[0121] Tasks 0-1 0-2 0-3 1-2 1-3 2-3 Average 1D-CNN 48.38±1.74 83.98±1.64 48.11±3.60 54.26±2.16 33.20±3.26 55.82±2.32 53.96% DAN 52.43±2.53 84.46±2.66 58.55±3.74 56.45±3.23 36.71±2.69 65.28±3.00 58.98% DSAN 56.42±1.18 88.78±1.36 72.59±5.25 60.31±3.22 38.35±2.61 76.60±2.93 65.51% DANN 52.85±2.27 84.17±1.04 60.00±2.99 55.94±2.54 35.82±2.77 66.06±1.81 59.14% DCTLN 53.92±1.68 84.02±2.17 57.92±2.24 57.00±2.57 35.54±2.22 66.17±2.55 59.09% MCD 64.95±4.28 93.72±1.10 67.54±5.35 78.00±2.83 48.31±2.19 71.85±4.03 70.73% SWD 62.31±2.88 91.95±1.29 57.98±3.51 71.48±2.01 42.65±4.35 67.45±4.58 65.64% BCDM 71.60±1.59 97.60±0.91 77.37±8.15 88.51±1.61 48.17±4.36 82.81±4.31 77.68% Proposed 76.17±1.51 97.74±0.46 87.20±6.41 92.57±1.58 66.05±6.61 88.89±4.78 84.77%
[0122] The framework of the proposed method consists of a self-supervised learning module based on data clustering, a joint classification certainty module, and an individual classification certainty module. To test the effectiveness of each module, an ablation study was conducted on task 1→2 of the PU dataset. It should be noted that the dual-classifier paradigm with two classifiers and one feature extractor was the backbone of the ablation analysis. The average accuracy of the cross-domain diagnostic performance of 10 trials of all models is shown in Table 5. It can be seen from the table that the proposed method achieved the highest average accuracy, and each module contributed to improving the performance of the model in the target domain. In particular, the degree of reduction was the most significant without maximizing joint classification certainty, with a reduction of 8.57%. Since joint classification certainty is the basis of the dual-classifier paradigm, in the final test, the two classifiers were used to jointly determine the target domain samples. In addition, the self-supervised learning module and the local certainty maximization module increased by 1.88% and 6.22% respectively. This is because the former aligns the source and target domain distributions at the feature level to reduce the domain gap, while the latter avoids the ambiguous output of the classifier, and thus can further optimize the generalization ability of the feature extractor compared with the previous dual-classifier paradigm methods.
[0123] Table 5: Average Accuracy of All Models in the Ablation Experiment of the PU Dataset
[0124] Without Ljcc Without Llcc Without Lssl All Average accuracy 84.00% 86.35% 90.69% 92.57%
[0125] To further verify the effectiveness of the proposed method, a fault diagnosis experiment was conducted on a rotating machinery fault diagnosis test platform (H0205). The simulation platform mainly consists of a driving motor, a gearbox, a test bearing, an acceleration sensor, and a data acquisition system, etc. The experiment selected a UPH205 type rolling bearing with a sampling frequency of 48 kHz. A wire electrical discharge linear cutting machine was used to set damage sizes of 0.4 mm and 0.65 mm on the inner ring, outer ring, and rolling elements of the bearing respectively. A total of six bearing health state categories were simulated in the experiment, as shown in Table 6, including cage fault, inner and outer ring compound fault, spherical fault, inner ring fault, outer ring fault, and normal condition. Signal acquisition was carried out under three different working conditions (working condition C0 = 1800 rpm / min, working condition C1 = 2100 rpm / min, working condition C2 = 2400 rpm / min), and 750 samples were set under each working condition. A total of 6 migration tasks were designed.
[0126] Table 6: Detailed parameters of faulty bearings in the H0205 dataset
[0127] Category label 0 1 2 3 4 5 Fault category CF IOF BF IF OF N Fault scale(mm) 0.4 0.65 0.4 0.65 0.65 0
[0128] Similarly, 8 UDA comparison models were adopted for experimental comparison, and the average diagnostic accuracy of 10 trials is shown in Table 7. It can be seen from Table 7 that this method has good fault recognition performance on the H0205 bearing dataset. The average fault recognition accuracy reached 98.88%, which is 9% higher than the 1D-CNN method. The proposed method has an obvious advantage in diagnostic accuracy compared with the method based on metric difference and the method based on domain adversarial; compared with the method based on dual-classifier adversarial, it has better comprehensive performance.
[0129] Table 7: Average diagnostic accuracy (%) of the H0205 dataset
[0130] Tasks 0-1 0-2 1-0 1-2 2-0 2-1 Average 1D-CNN 93.80±1.54 87.33±2.33 93.40±1.84 95.07±1.84 77.87±7.72 91.80±3.81 89.88% DAN 97.94±1.46 97.87±1.40 96.40±1.34 98.27±1.22 91.33±2.08 98.27±1.34 96.88% DSAN 99.47±0.61 97.73±1.10 95.73±1.89 98.93±1.00 94.20±3.05 97.13±4.31 97.20% DANN 97.47±1.21 95.67±1.92 94.13±1.36 97.20±1.69 92.60±4.28 96.13±1.36 95.53% DCTLN 98.27±0.78 95.60±1.73 95.27±1.55 98.27±1.27 92.73±2.02 97.00±1.52 96.19% MCD 99.87±0.42 98.67±0.94 96.53±0.88 99.40±0.49 95.47±1.17 98.47±0.77 98.07% SWD 99.33±0.89 97.33±1.51 96.93±1.61 99.40±0.49 95.47±1.74 98.07±1.02 97.76% BCDM 99.87±0.28 98.54±0.88 97.00±1.01 99.53±0.45 97.14±1.00 99.27±0.66 98.56% Proposed 100.00±0.00 99.73±0.47 97.40±0.49 99.93±0.21 97.20±0.52 99.00±1.30 98.88%
[0131] To solve the problem of bearing fault diagnosis under variable working conditions, the present invention proposes a new self-supervised learning dual-classifier domain adaptation method. Aiming at the limitations of the original dual-classifier domain adaptation method, a measure of maximizing the accuracy of the dual-classifier is proposed, and the domain adaptation model is optimized by simultaneously considering the joint certainty between classifiers and the individual certainty of each classifier. To enhance the attention of the network model to the learning of target domain data, a self-supervised learning method based on data clustering is proposed, achieving better alignment of feature distributions. Experimental verification was carried out on two rolling bearing datasets, PU and H0205. In the 6 migration task test experiments on the PU dataset, compared with the basic dual-classifier MCD method, the diagnostic accuracy of the proposed method was increased by 14% on average. The experimental results show the effectiveness of the present invention.
[0132] The future focus of work will be to explore the transfer of diagnostic knowledge between different machines through adversarial training and self-supervised learning, so as to achieve a wider range of cross-domain fault diagnosis.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A cross-domain fault diagnosis method for rolling bearings based on self-supervised learning, characterized in that, It includes the following steps: S1: Construct a labeled source domain and an unlabeled target domain from the original vibration signals collected under different working conditions, and randomly divide them into a training set and a test set in a ratio of 8:
2. S2: Input the training set data into the network model, calculate the overall loss function of the network three times in sequence, and update the overall parameters of the network, the parameters of the classifier, and the parameters of the feature extractor through the backpropagation algorithm in sequence. S3: Repeat step S2 until the maximum number of iterations, stop training, and generate a domain adaptation network with good generalization performance. S4: Input the test set data in step S1 into the trained network model and output the diagnostic results under different working conditions. The specific content of step S2 includes: S21: Input the source domain training set data, the corresponding true labels, and the target domain training set data, calculate the pseudo-labels of the target domain samples corresponding thereto, and calculate the total loss function of the network at this time based on this. Use the backpropagation algorithm to update the overall parameters of the network; S22: Input the source domain training set data, the corresponding true labels, and the target domain training set data, and calculate the total loss function of the network at this time. Update the classifier parameters using the backpropagation algorithm; S23: Input the source domain training set data, the corresponding true labels, and the target domain training set data, and calculate the total loss function of the network at this time Update the parameters of the feature extractor using the backpropagation algorithm; In step S21, the overall loss function of the network is obtained through Equation 1: Among them, N t represents the total number of target domain samples, represents the j-th sample in the target domain, represents the indication vector of the target domain pseudo-label; c represents c categories of bearing samples This represents the softmax output of the classifier for the j-th sample in the target domain; is the pseudo-label, which is obtained by Equation 2: Among them, represents the j-th sample in the target domain; F(·) represents the feature vector of the sample after passing through the feature extractor, and cen c represents the centroid of the target domain data of the c-th category; The centroid of the target domain data of the c-th category is obtained through Equation 3: Among them, δ c represents the prediction for the c-th class after the softmax operation, represents the j-th sample in the target domain; F(·) represents the feature vector of the sample after passing through the feature extractor, and C(·) represents the output vector after passing through the classifier; C k (k = 1, 2) are two task classifiers; The loss functions in the step S22 and the step S23 are obtained by calculating according to Equation 4: Among them, μ and ω are the balance parameters between the joint classification certainty and the individual classification certainty, respectively; It is a metric for the joint classification accuracy of the dual classifier; The individual classification accuracy index output by the classifier; The individual classification accuracy metric output by the classifier in Equation 4 Obtained by calculation using Equation 5: Among them, δ c represents the prediction for the c-th class after the softmax operation, F(·) represents the feature vector of the sample after passing through the feature extractor, and C(·) represents the output vector after passing through the classifier; C k (k = 1, 2) are two task classifiers; Metric for the combined classification accuracy of the two classifiers in Equation 4 Obtained by calculation using Equation 6: Among them, δ c represents the prediction for the c-th class after the softmax operation, f(·) represents the feature vector of the sample after passing through the feature extractor, and C(·) represents the output vector after passing through the classifier; C k (k = 1, 2) are two task classifiers.
2. The cross-domain fault diagnosis method for rolling bearings based on self-supervised learning according to claim 1, wherein The domain adaptation network structure includes: a feature extractor based on a five-layer convolutional block and two classifiers with the same structure. The feature extractor extracts transferable features from the source domain and target domain data, and the two classifiers can obtain the deep features generated by the feature extractor and output the corresponding prediction probabilities.
3. A cross-domain fault diagnosis method for rolling bearings based on self-supervised learning according to claim 2, characterized in that, The method for constructing the domain adaptation network model in step S1 is: First, introduce the original vibration signals sampled from the source domain and the target domain into the feature extractor. Then, the two binary classifiers use the cross-entropy loss to classify the source samples, and at the same time embed a self-supervised learning loss based on target clustering. Finally, construct a minimax game and introduce the joint and individual classification accuracy losses of the binary classifiers to guide the output of the binary classifiers.
4. A cross-domain fault diagnosis method for rolling bearings based on self-supervised learning according to claim 2, characterized in that, Step S1 specifically includes: S11: Collect the vibration data of rolling bearings under different working conditions, construct a labeled source domain and an unlabeled target domain, randomly divide 80% of the target domain data into the training set, and the remaining 20% into the test set. S12: Set the maximum number of iterations of the network, the size and stride of the convolutional kernels in the convolutional layer, set the size and stride of the pooling kernels in the pooling layer, and initialize the weight parameters of the convolutional layer and the fully connected layer.
Citation Information
Patent Citations
Rolling bearing fault diagnosis method based on self-supervised learning and clustering
CN113792758A
Self-learning-based unsupervised cross-working-condition bearing fault diagnosis method
CN115358259A