Step-by-step guided adaptive transfer learning rolling bearing fault diagnosis method
By constructing the 1D-LKMS-ECAResNet model and step-by-step guide domain adaptive strategy, the problems of noise interference and feature distribution differences are solved, and the high accuracy and stability of bearing fault diagnosis are achieved, and the performance of the model under different operating conditions is improved.
Patent Information
- Application Number
- CN202510543774.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
AI Technical Summary
The existing bearing fault diagnosis method based on domain adaptive transfer learning fails to effectively consider the influence of noise, ignores the local differences in feature distribution, and the model is difficult to converge stably during training, resulting in insufficient diagnostic accuracy under different operating conditions.
The 1D-LKMS-ECAResNet model is constructed, and a large-core multi-scale structure and an efficient channel attention residual structure are embedded. Combined with step-by-step guide domain adaptive strategy, the loss function is calculated through DJP-MMD and LMMD to realize the global and local alignment of the source domain and target domain features, and stabilize the model training process.
It improves the accuracy of the fault diagnosis of the model under different operating conditions, reduces noise interference, ensures the stable convergence of the model in high-dimensional space, and improves the domain adaptation capability.
Smart Images

Figure CN120448785A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bearing fault diagnosis, and in particular to a rolling bearing fault diagnosis method based on step-by-step guided adaptive transfer learning. Background Art
[0002] With the development of science and technology, traditional bearing fault diagnosis technology has gradually been replaced by deep learning-based technology due to its shortcomings such as excessive reliance on manual experience and complex diagnostic processes. Intelligent diagnostic methods based on deep learning can deeply explore the fault characteristics contained in bearing vibration signals by learning from a large number of data samples, and then identify the type of bearing fault. However, bearing fault diagnosis driven by deep learning is highly dependent on large-scale labeled samples, and sample labeling requires a lot of labor costs. At the same time, equipment operating conditions change dynamically with production demand, resulting in significant distribution differences in bearing fault samples under different operating conditions. As a result, models trained based on data from a single operating condition have poor model generalization capabilities when applied to other operating conditions.
[0003] Transfer learning can achieve bearing fault diagnosis under different operating conditions (target domains) by learning transferable fault features from samples of a single operating condition (i.e., the source domain). Unsupervised domain adaptation (UDA), one of the main transfer learning methods, can narrow the feature distribution differences between the source and target domains through specific algorithmic techniques, enabling the model to adaptively learn target domain features, achieve knowledge transfer, and improve the model's diagnostic performance on target domain tasks. However, current domain adaptive transfer learning bearing fault diagnosis methods have the following shortcomings: First, current domain adaptive transfer learning bearing fault diagnosis methods do not consider the impact of noise on fault diagnosis; second, current domain adaptive transfer learning bearing fault diagnosis methods only perform global alignment of feature distributions between the source and target domains, ignoring the alignment of related subdomains and the differences between different categories; third, when combining different domain adaptive methods, the problem that different domain adaptive losses when adjusting network parameters may make it difficult for the model to effectively converge towards a stable direction during training is not considered. These three issues affect the accuracy of transfer learning fault diagnosis models for bearing fault identification to a certain extent. Therefore, a step-by-step guided adaptive transfer learning method for rolling bearing fault diagnosis is proposed. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems in the prior art and to propose a rolling bearing fault diagnosis method based on step-by-step guided adaptive transfer learning.
[0005] A rolling bearing fault diagnosis method using step-by-step guided adaptive transfer learning includes the following steps:
[0006] S1. Process the vibration signal of the collected bearing data by segmenting the vibration signal to form samples containing different fault types and normal states, and divide the samples into a training set and a test set according to an 8:2 ratio;
[0007] S2. Construct a 1D-LKMS-ECAResNet model to fully extract the fault features contained in the source and target domains and reduce the interference of noise on the domain adaptation process;
[0008] S3, construct a step-by-step guided domain adaptation strategy, and constrain the 1D-LKMS-ECAResNet model through the step-by-step guided domain adaptation strategy;
[0009] S4. Input labeled training samples of the source domain working condition and unlabeled training samples of the target domain working condition into the 1D-LKMS-ECAResNet model to extract features. Calculate the classification loss of the source domain through the classifier of the 1D-LKMS-ECAResNet model. Calculate the distribution difference loss of the two domain features extracted by the 1D-LKMS-ECAResNet model under the step-by-step guided domain adaptation strategy. Backpropagate the two losses to adjust the parameters of the 1D-LKMS-ECAResNet model. Save the model after multiple iterations.
[0010] S5. Input the test set samples of the target domain into the trained 1D-LKMS-ECAResNetz model for testing.
[0011] Preferably, in step S2, a large-core multi-scale structure and an efficient channel attention residual structure are embedded in the 1D-LKMS-ECAResNet model.
[0012] Preferably, in step S3, the step-by-step guided domain adaptation strategy consists of DJP-MMD and LMMD, DJP-MMD is used for domain adaptation in the first 50 iterations of training, LMMD is used for domain adaptation in 50 to 100 iterations of training, and after 100 iterations of training, both LMMD and DJP-MMD are used for domain adaptation;
[0013] The calculation formula of DJP-MMD is as follows:
[0014]
[0015] Where d represents the Euclidean distance, M T represents the joint probability difference between the source domain and the target domain on the same category, M D represents the joint probability difference between the source domain and the target domain in different categories, μ represents the balance parameter, where C represents the number of categories, n and m represent the number of samples in the source domain and target domain datasets respectively, It is represented by the number of samples predicted as category c in the target domain dataset, n c represents the number of samples belonging to category c in the source domain dataset, φ represents the mapping, represents the i-th sample belonging to category c in the source domain, represents the jth sample belonging to category c in the target domain;
[0016] The calculation formula of LMMD is as follows:
[0017]
[0018] Among them, D s and D t Represents the source domain and target domain data, x s and x t Respectively represent D s and D t In order to express LMMD more clearly, the concept of sample weight is introduced. Assuming that both the source domain and the target domain have C categories, the unbiased estimator of LMMD can be expressed as follows:
[0019]
[0020] In the above formula and Represented as source domain samples and target domain samples The weights belonging to category C, where can be expressed as:
[0021]
[0022] Where y ic Represents the cth element of the label vector, representing sample x i The probability of belonging to category c, when domain D is the source domain, the weight From the source domain samples The true label Calculated; when domain D is the target domain, since the target domain has no label, the label of the target domain is the probability prediction value output by the label classifier as the pseudo label of the target domain sample;
[0023] The reproducing kernel Hilbert space H has the inner product completeness property <φ(x s ),φ(x t )>=K(x s ,x t ), where <·,·> represents the inner product of the vector, and K represents the kernel function. Through kernel embedding, the sample can be mapped into a high-order moment. Assume that the samples in the source domain and the target domain are output after feature extraction in a certain network layer L. and Then the distance between the source domain and the target domain distribution can be converted into a kernel function form as follows:
[0024]
[0025] The calculation formula for the step-by-step guided domain adaptation loss is as follows:
[0026]
[0027] Among them, L DJP-MMD represents the global domain adaptation DJP-MMD distribution difference loss, LMMD (D s ,D t ) represents the local domain adaptation LMMD distribution difference loss, μ represents the adjustment parameter, and epoch represents the number of iterations of the model.
[0028] Preferably, in step S4, the cross entropy loss of the classifier is:
[0029]
[0030] in, represents the source domain sample, represents the source domain features, The predicted labels of the source domain, represents the true label of the source domain, J(·,·) represents the cross entropy loss;
[0031] Total loss function Loss Total for:
[0032] Loss Total =Loss y +αLoss d .
[0033] Compared with the existing technology, the advantages of the present invention are:
[0034] 1. The present invention constructs a 1D-LKMS-ECAResNet model to fully extract source and target domain features. The model embeds a large-kernel multi-scale structure and an efficient channel attention residual structure, which can perform smoother feature extraction of the input signal, reduce the interference of high-frequency noise, and effectively improve the noise resistance of the network through feature weighting. It can also successfully avoid the gradient explosion problem caused by too many network layers by leveraging the advantages of the residual structure.
[0035] 2. The present invention constructs a step-by-step guided domain adaptation strategy, which can achieve global and local alignment of source and target domain features in high-dimensional space and enhance the distribution distance of different categories by step-by-step guided network iteration. In addition, the step-by-step guided strategy can enable the model to converge stably, further improving the model's performance and domain adaptation capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Schematic diagram of the structure of the 1D-LKMS-ECAResNetz model in the present invention.
[0037] Figure 2 This is a schematic diagram of the rolling bearing fault diagnosis method using step-by-step guided adaptive transfer learning in the present invention.
[0038] Figure 3 This is the T-SNE visualization effect diagram of the test set features of the JNU bearing dataset A→B migration task in this invention.
[0039] Figure 4 This is the test set confusion matrix of the JNU bearing dataset A→B migration task in this invention. DETAILED DESCRIPTION
[0040] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below in conjunction with specific implementation methods.
[0041] Reference Figure 1-2 As shown, a step-by-step guided adaptive transfer learning rolling bearing fault diagnosis method includes the following steps:
[0042] S1. Process the vibration signal of the collected bearing data by segmenting the vibration signal to form samples containing different fault types and normal states, and divide the samples into a training set and a test set according to an 8:2 ratio;
[0043] S2. Construct a 1D-LKMS-ECAResNet model to fully extract the fault features contained in the source and target domains and reduce the interference of noise on the domain adaptation process;
[0044] S3, construct a step-by-step guided domain adaptation strategy, and constrain the 1D-LKMS-ECAResNet model through the step-by-step guided domain adaptation strategy;
[0045] S4. Input labeled training samples of the source domain working condition and unlabeled training samples of the target domain working condition into the 1D-LKMS-ECAResNet model to extract features. Calculate the classification loss of the source domain through the classifier of the 1D-LKMS-ECAResNet model. Calculate the distribution difference loss of the two domain features extracted by the 1D-LKMS-ECAResNet model under the step-by-step guided domain adaptation strategy. Backpropagate the two losses to adjust the parameters of the 1D-LKMS-ECAResNet model. Save the model after multiple iterations.
[0046] S5. Input the test set samples of the target domain into the trained 1D-LKMS-ECAResNetz model for testing.
[0047] In step S2, a large-kernel multi-scale structure and an efficient channel attention residual structure are embedded in the 1D-LKMS-ECAResNet model. The large-kernel multi-scale structure uses large-kernel one-dimensional convolutions such as 31, 15, and 11. These large-kernel convolutions can cover a longer time window of the input signal, perform smoother feature extraction on the input signal, and reduce the interference of high-frequency noise. The efficient channel attention residual structure can effectively improve the noise resistance of the network through feature weighting, and can also successfully avoid the problem of gradient explosion caused by too many network layers by leveraging the advantages of the residual structure.
[0048] In step S3, the step-by-step guided domain adaptation strategy consists of DJP-MMD and LMMD. DJP-MMD is used for domain adaptation in the first 50 iterations of training. In this stage, the model can quickly narrow the overall distribution difference between the source domain and the target domain from a macro perspective, learn the global features shared by the source domain and the target domain, ensure the model's grasp of the overall distribution, and lay the foundation for subsequent local adaptation; LMMD is used for domain adaptation in 50 to 100 iterations of training, which can make more detailed adjustments to the local structure of the data; after 100 iterations of training, LMMD and DJP-MMD are used simultaneously for domain adaptation, which can integrate global and local information. The combination of the two can further improve the performance and domain adaptation ability of the model;
[0049] The calculation formula of DJP-MMD is as follows:
[0050]
[0051] Where d represents the Euclidean distance, M T represents the joint probability difference between the source domain and the target domain on the same category, M D represents the joint probability difference between the source domain and the target domain in different categories, μ represents the balance parameter, where C represents the number of categories, n and m represent the number of samples in the source domain and target domain datasets respectively, It is represented by the number of samples predicted as category c in the target domain dataset, n c represents the number of samples belonging to category c in the source domain dataset, φ represents the mapping, represents the i-th sample belonging to category c in the source domain, represents the jth sample belonging to category c in the target domain;
[0052] The calculation formula of LMMD is as follows:
[0053]
[0054] Among them, D s and D t Represents the source domain and target domain data, x s and x t Respectively represent D s and D t In order to express LMMD more clearly, the concept of sample weight is introduced. Assuming that both the source domain and the target domain have C categories, the unbiased estimator of LMMD can be expressed as follows:
[0055]
[0056] In the above formula and Represented as source domain samples and target domain samples The weights belonging to category C, where can be expressed as:
[0057]
[0058] Where y ic Represents the cth element of the label vector, representing sample x i The probability of belonging to category c, when domain D is the source domain, the weight From the source domain samples The true label Calculated; when domain D is the target domain, since the target domain has no label, the label of the target domain is the probability prediction value output by the label classifier as the pseudo label of the target domain sample;
[0059] The reproducing kernel Hilbert space H has the inner product completeness property <φ(x s ),φ(x t )>=K(x s ,x t ), where <·,·> represents the inner product of the vector, and K represents the kernel function. Through kernel embedding, the sample can be mapped into a high-order moment. Assume that the samples in the source domain and the target domain are output after feature extraction in a certain network layer L. and Then the distance between the source domain and the target domain distribution can be converted into a kernel function form as follows:
[0060]
[0061] The calculation formula for the step-by-step guided domain adaptation loss is as follows:
[0062]
[0063] Among them, L DJP-MMD represents the global domain adaptation DJP-MMD distribution difference loss, LMMD (D s ,D t ) represents the local domain adaptation LMMD distribution difference loss, μ represents the adjustment parameter, and epoch represents the number of iterations of the model.
[0064] In step S4, the cross entropy loss of the classifier is:
[0065]
[0066] in, represents the source domain sample, represents the source domain features, The predicted labels of the source domain, represents the true label of the source domain, J(·,·) represents the cross entropy loss;
[0067] Total loss function Loss Total for:
[0068] Loss Total =Loss y +αLoss d .
[0069] Example
[0070] This example uses the Jiangnan University (JNU) bearing dataset, which includes four bearing states: normal, inner ring fault, outer ring fault, and rolling element fault. The bearings operate under three different operating conditions (the motor speed in each operating condition is 600 r / min, 800 r / min, and 1000 r / min, respectively). Vibration signals of the four bearing states are collected under these three operating conditions.
[0071] S1. The vibration signal of the bearing data of Jiangnan University is processed. The vibration signal is segmented and intercepted to form samples containing different fault types and normal states. The segmentation length of the vibration signal is 3072, that is, 3072 data points are one sample. Table 1 details the division of samples under different working conditions, where the three working conditions are represented by A, B, and C (A: 600r / min, B: 800r / min, C: 1000r / min). In addition, in order to facilitate the output of the model, the four types of normal, inner ring fault, outer ring fault, and rolling element fault are set as labels 0, 1, 2, and 3 respectively. The training set and test set of the model are divided into 8:2.
[0072] Table 1. Sample table of three working conditions of JNU bearing data set
[0073]
[0074] S2. Construct a 1D-LKMS-ECAResNet model. The structure of the model is as follows Figure 1 As shown in the figure, the model embeds a large kernel multi-scale structure and an efficient channel attention residual structure.
[0075] S3. Construct a step-by-step guided domain adaptation strategy. By constraining 1D-LKMS-ECAResNet, the extracted source domain features and target domain features are better aligned in high-dimensional space.
[0076] S4, such as Figure 2 As shown in the figure, the training samples of the source domain working conditions (with labels) and the training samples of the target domain working conditions (without labels) of the JNU bearing dataset are input into 1D-LKMS-ECAResNet for feature extraction. The extracted source domain features are calculated through the classifier to calculate the cross entropy loss. The distribution difference loss of the target domain features and the source domain features extracted by the 1D-LKMS-ECAResNet model under the step-by-step guided domain adaptation strategy is calculated. The two losses are back-propagated to adjust the 1D-LKMS-ECAResNet parameters. Through multiple iterations, the classifier of the model can also well identify the fault type of the target domain samples.
[0077] S5. Input the test set samples of the target domain of the JNU bearing dataset into the trained 1D-LKMS-ECAResNetz for testing. To verify the practicality of the present invention, it is experimentally compared with other methods. The verification results of cross-speed transfer learning on the Jiangnan University bearing dataset are shown in Table 2:
[0078] Table 2. Accuracy of cross-condition transfer learning of different methods on the Jiangnan University bearing dataset
[0079]
[0080] The following conclusions can be drawn from the above experiments: In the six cross-working condition migration experiments of the Jiangnan University bearing dataset, the method of the present invention showed excellent performance. Compared with the maximum mean difference (MMD) method, the average accuracy of the method of the present invention increased by 19.37%; compared with the deep adversarial network (DANN), the average accuracy increased by 13.12%; compared with the correlation alignment (CORAL) method, the average accuracy increased by 13.88%; compared with the kernelized maximum mean difference (KM-MMD) method, the average accuracy was also 13.37% higher. In order to further demonstrate the test effect of the present invention, the test set of the A→B migration task of the JNU University bearing dataset was selected to display the T-SNE visualization effect diagram to highlight the aggregation effect of its different category features, such as Figure 3 shown.
[0081] It is understood from common technical knowledge that the present invention may be implemented by other embodiments that do not depart from its spirit or essential features. Therefore, the embodiments disclosed above are, in all respects, merely illustrative and not exclusive. All modifications within the scope of the present invention or equivalent to the scope of the present invention are intended to be encompassed by the present invention.
Claims
1. A rolling bearing fault diagnosis method using step-by-step guided adaptive transfer learning, characterized by: The following steps are involved: S1. Process the vibration signal of the collected bearing data by segmenting the vibration signal to form samples containing different fault types and normal states, and divide the samples into a training set and a test set according to an 8:2 ratio; S2. Construct a 1D-LKMS-ECAResNet model to fully extract the fault features contained in the source and target domains and reduce the interference of noise on the domain adaptation process; S3, construct a step-by-step guided domain adaptation strategy, and constrain the 1D-LKMS-ECAResNet model through the step-by-step guided domain adaptation strategy; S4. Input the labeled training samples of the source domain working condition and the unlabeled training samples of the target domain working condition into the 1D-LKMS-ECAResNet model to extract features. The classification loss of the source domain is calculated through the classifier of the 1D-LKMS-ECAResNet model. The distribution difference loss of the two domain features extracted by the 1D-LKMS-ECAResNet model under the step-by-step guided domain adaptation strategy is calculated. The two losses are backpropagated to adjust the parameters of the 1D-LKMS-ECAR esNet model. The model is saved after multiple iterations. S5. Input the test set samples of the target domain into the trained 1D-LKMS-ECAResNetz model for testing.
2. The rolling bearing fault diagnosis method based on step-by-step guided adaptive transfer learning according to claim 1 is characterized in that: In step S2, a large-core multi-scale structure and an efficient channel attention residual structure are embedded in the 1D-LKMS-ECAResNet model.
3. The rolling bearing fault diagnosis method based on step-by-step guided adaptive transfer learning according to claim 1, characterized in that: In step S3, the step-by-step guided domain adaptation strategy consists of DJP-MMD and LMMD, using DJP-MMD for domain adaptation in the first 50 iterations of training, using LMMD for domain adaptation in 50-100 iterations of training, and using both LMMD and DJP-MMD for domain adaptation after 100 iterations of training; The calculation formula of DJP-MMD is as follows: Where d represents the Euclidean distance, M T represents the joint probability difference between the source domain and the target domain on the same category, M D represents the joint probability difference between the source domain and the target domain in different categories, μ represents the balance parameter, where C represents the number of categories, n and m represent the number of samples in the source domain and target domain datasets respectively, It is represented by the number of samples predicted as category c in the target domain dataset, n c represents the number of samples belonging to category c in the source domain dataset, φ represents the mapping, represents the i-th sample belonging to category c in the source domain, represents the jth sample belonging to category c in the target domain; The calculation formula of LMMD is as follows: Among them, D s and D t Represents the source domain and target domain data, x s and x t Respectively represent D s and D t In order to express LMMD more clearly, the concept of sample weight is introduced. Assuming that both the source domain and the target domain have C categories, the unbiased estimator of LMMD can be expressed as follows: In the above formula and Represented as source domain samples and target domain samples The weight of category C, where can be expressed as: Where y ic Represents the cth element of the label vector, representing sample x i The probability of belonging to category c, when domain D is the source domain, the weight From the source domain samples The true label Calculated; when domain D is the target domain, since the target domain has no label, the label of the target domain is the probability prediction value output by the label classifier as the pseudo label of the target domain sample; The reproducing kernel Hilbert space H has the inner product completeness property <φ(x s ),φ(x t )>=K(x s ,x t ), where <·,·> represents the inner product of the vector, and K represents the kernel function. Through kernel embedding, the sample can be mapped into a high-order moment. Assume that the samples in the source domain and the target domain are output after feature extraction in a certain network layer L. and Then the distance between the source domain and the target domain distribution can be converted into a kernel function form as follows: The calculation formula for the step-by-step guided domain adaptation loss is as follows: Among them, L DJP-MMD represents the global domain adaptation DJP-MMD distribution difference loss, LMMD (D s ,D t ) represents the local domain adaptation LMMD distribution difference loss, μ represents the adjustment parameter, and epoch represents the number of iterations of the model.
4. The rolling bearing fault diagnosis method based on step-by-step guided adaptive transfer learning according to claim 1, characterized in that: In step S4, the cross entropy loss of the classifier is: in, represents the source domain sample, represents the source domain features, The predicted labels of the source domain, represents the true label of the source domain, J(·,·) represents the cross entropy loss; Total loss function Loss Total for: Loss Total =Loss y +αLoss d 。
Citation Information
Cited By
Transmission chain cross-domain diagnosis method for optimal transmission through fusion of zero-sequence current and vibration signals
CN120873986A