Bearing fault diagnosis method based on multi-core weighted joint domain adaptive network
By introducing a multi-core weighted joint domain adaptation method in the domain adaptive network, the weights of edge and conditional distributions are dynamically adjusted, and the pseudo-labels are corrected, the problem of large differences in fault diagnosis accuracy in the prior art is solved, and higher fault diagnosis accuracy and stability are achieved.
Patent Information
- Application Number
- CN202411238561.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-05
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-09-05
AI Technical Summary
The existing domain adaptive methods cannot dynamically adjust the weight ratio of edge distribution and conditional distribution according to the actual data situation during joint distribution alignment, resulting in a large difference in the accuracy of fault diagnosis; at the same time, the pseudo-label confidence generated by neural networks trained with the source domain is low, resulting in low fault diagnosis accuracy.
A multi-core weighted joint domain adaptive network is adopted to adjust the weight ratio of edge distribution and conditional distribution according to actual data through dynamic weight adaptation, and correct the pseudo-label to improve the confidence of the pseudo-label.
It improves the accuracy and stability of bearing fault diagnosis, enhances the adaptability of the neural network model, and can better adapt to fault diagnosis tasks under different data conditions.
Smart Images

Figure CN119202717B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a bearing fault diagnosis method based on a multi-core weighted joint domain adaptive network, and belongs to the field of intelligent diagnosis of fault data. Background Art
[0002] With the rapid development of deep learning, intelligent fault diagnosis technology has been used more and more widely. However, ideal fault diagnosis technology often needs to satisfy the independent and identical distribution between domains and requires a large number of labeled samples, which to a certain extent limits its application in engineering practice. Domain adaptation can solve these problems well. It alleviates the domain differences of data distribution through joint learning of source domain samples and target domain samples. The main idea of deep domain adaptation is to use deep networks for domain-invariant feature learning. For example, some scholars in the industry have proposed a weighted maximum mean difference (WMMD) model to align the marginal distribution between source domain data and target domain data; some scholars have used a generator model based on multi-kernel maximum mean difference (MKMMD) to align the distribution of labeled data and unlabeled industrial data; and some scholars have proposed a new method based on pseudo-labels and clustering assumptions to achieve class-conditional distribution alignment.
[0003] Based on these people's work, more and more researchers have extended these unilateral distributions to joint distributions. For example, some scholars use the Joint Domain Adaptation Network (JAN) to learn domain transfer knowledge by aligning the joint distribution. Some scholars have developed a joint adversarial domain adaptation (JADA) method to improve the generalization of the model. Other scholars have improved diagnostic accuracy by extending marginal distribution adaptation (MDA) to joint distribution adaptation (JDA).
[0004] However, the above-mentioned existing domain adaptation methods still have the following two problems: First, in the process of joint distribution alignment, the marginal distribution and conditional distribution are usually only quantitatively assigned weight ratios, which makes it impossible to adapt and adjust according to the actual data situation, resulting in large differences in the fault diagnosis accuracy under different data conditions; second, pseudo-labels of the target domain are required in the conditional distribution. The existing method is to use the neural network trained in the source domain to pseudo-label the target domain samples. This method will produce pseudo-labels with low confidence, which will ultimately lead to low fault diagnosis accuracy. Summary of the invention
[0005] In view of the problems existing in the above-mentioned prior art, the present invention provides a bearing fault diagnosis method based on a multi-core weighted joint domain adaptive network, which adopts a dynamic weight adaptive method to adjust the weight ratio of the marginal distribution and the conditional distribution according to the actual data situation, and adopts a correction value method to correct the pseudo-label, thereby improving the confidence of the pseudo-label, and ultimately effectively improving the accuracy of bearing fault diagnosis.
[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is: a bearing fault diagnosis method based on a multi-core weighted joint domain adaptive network, the specific steps are:
[0007] S1. Obtain sample data of the source domain and the target domain respectively, divide the sample data of the target domain into training sample data and test sample data, and use the sample data of the source domain and the training sample data of the target domain as training data;
[0008] S2, inputting the sample data of the source domain and the training sample data of the target domain into the neural network, thereby obtaining the sample features of the source domain and the target domain extracted by the deep part of the neural network, the difference in the output features of the target domain training samples in the shallow part and the deep part of the neural network, and the pseudo labels of the target domain training samples;
[0009] S3, after processing the data obtained in step S2, a source domain classification loss function and a pseudo label correction loss function are obtained, and a joint domain adaptation function of the source domain and the target domain is obtained by using the maximum multi-core mean difference calculation process;
[0010] S4, combining the source domain classification loss function, the joint domain adaptation function of the source domain and the target domain, and the pseudo-label correction loss function obtained in step S3 as the overall network loss function;
[0011] S5, training the neural network using the training data in step S1 and the overall network loss function obtained in step S4, thereby updating the network parameters;
[0012] S6. Input the test sample data of the target domain in step S1 into the neural network trained in step S5. Finally, the neural network performs bearing fault diagnosis on the input data to obtain the accuracy of the bearing fault diagnosis of the neural network.
[0013] Furthermore, the sample data of each of the source domain and the target domain include data of faulty bearings and data of normal bearings.
[0014] Furthermore, in step S1, the sample data of the source domain is labeled, and the sample data of the target domain is unlabeled; the sample data categories of the source domain and the target domain are the same (that is, the fault type in the source domain is consistent with the fault type in the target domain, that is, the labels of the same fault are consistent); the sample data of the source domain and the sample data of the target domain are respectively expressed as:
[0015]
[0016] in is the labeled source domain sample data, n s is the number of source domain samples, Represents the i-th source domain sample and the corresponding label; is the unlabeled target domain sample data, n t is the number of samples in the target domain, Indicates the jth target domain sample, the source domain and the target domain have a common category C n .
[0017] Further, after the data in step S2 is input into the neural network:
[0018] S2.1. Use a neural network to extract features from the sample data of the source domain and the target domain to obtain the corresponding sample features:
[0019]
[0020] Among them, F D is the deep network part, θ s ,θ t are network parameters;
[0021] S2.2, Jensen-Shannon divergence is used to measure the difference between the sample data of the target domain in the shallow part and the deep part of the neural network. The specific calculation is as follows:
[0022] Calculate the target domain sample data input shallow part and deep part to get the average distribution of feature distribution
[0023]
[0024] Where F S It is the shallow network part;
[0025] Calculate the Jensen-Shannon divergence after obtaining the mean distribution:
[0026]
[0027] D JS The value is used as the pseudo label correction value;
[0028] S2.3. Obtain the one-hot encoding of the pseudo-label of the target domain sample data through the trained neural network.
[0029] Further, the step S3 is specifically as follows:
[0030] S3.1. Use multi-kernel maximum mean difference (MKMMD) to calculate the marginal distribution difference between the source domain and the target domain, where the marginal distribution adaptive loss L MDA The calculation is as follows:
[0031]
[0032]
[0033] In the formula, G(·) is the network feature extraction part, is the training sample data of the target domain;
[0034] S3.2. Use multi-kernel maximum mean difference (MKMMD) to calculate the conditional distribution difference between the source domain and the target domain, where the conditional distribution adaptation loss L CDA The calculation is as follows:
[0035]
[0036]
[0037] S3.3. Calculation of adaptive factor u:
[0038]
[0039] S3.4, the weighted joint domain adaptation function is calculated as follows:
[0040] L WJDA =(1-μ)·L MDA +μ·L CDA
[0041] S3.5. Calculation of source domain classification loss:
[0042]
[0043] in is the one-hot encoding of the true label of the i-th source domain sample data;
[0044] S3.6, combined with step S2 to obtain the pseudo-label correction value D JS , the corrected pseudo-label loss function is calculated as follows:
[0045]
[0046] L rect =E[exp{-D JS}·L CE +D JS ]
[0047] in is the one-hot encoding of the pseudo-label of the j-th target domain sample data.
[0048] Furthermore, the overall network loss function in step S4 is specifically:
[0049] L total =L cls +λ·L WJDA +0.1 L rect
[0050] Where λ is the balance factor between losses.
[0051] Further, in step S5, the network parameters are updated using an Adamax optimizer;
[0052]
[0053] Where θ represents the network parameters and α represents the learning rate.
[0054] Compared with the prior art, the present invention has the following advantages:
[0055] 1. The present invention first inputs sample data of the source domain and training sample data of the target domain into a neural network, and then obtains sample features of the source domain and the target domain, the difference in output features of the target domain training samples in the shallow part and the deep part of the neural network, and the pseudo-label of the target domain training samples. Then, the above data are processed to obtain a source domain classification loss function and a pseudo-label correction loss function, and the maximum multi-core mean difference is used to calculate and process to obtain a joint domain adaptation function of the source domain and the target domain; the above functions are combined as the overall network loss function, and the weight ratio of the marginal distribution and the conditional distribution is adjusted according to the actual data situation in a dynamic weight adaptation manner through the function to obtain the required neural network model, which has better adaptability when facing different fault diagnosis tasks and meets different engineering application scenarios.
[0056] 2. In the process of training and generating the corresponding neural network model, the present invention adopts the calculated JS divergence as the correction value, and uses the correction value to correct the pseudo-label with low confidence, thereby improving the confidence of the pseudo-label, thereby combining with the above-mentioned dynamic weight adaptive method, and finally effectively improving the accuracy and stability of bearing fault diagnosis under different data conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a MKWJDAN network structure diagram used for bearing fault diagnosis in the present invention;
[0058] Figure 2 It is a network structure diagram for pseudo-label correction in the present invention;
[0059] Figure 3 It is a comparison result diagram of the method of the present invention and other mainstream methods on task J2-J1 in the experimental proof;
[0060] Figure 4 This is a visualization result diagram using t-SNE of the method of the present invention and other mainstream methods on tasks D1-D2 in experimental proof. DETAILED DESCRIPTION
[0061] The present invention will be further described below.
[0062] like Figure 1 and 2 As shown, the specific steps of the present invention are:
[0063] S1. Obtain sample data of the source domain and the target domain respectively, divide the sample data of the target domain into training sample data and test sample data, and use the sample data of the source domain and the training sample data of the target domain as training data; the sample data of the source domain and the target domain respectively include data of faulty bearings and data of normal bearings.
[0064] The sample data in the source domain is labeled, and the sample data in the target domain is unlabeled; the sample data categories of the source domain and the target domain are the same (that is, the fault type in the source domain is consistent with the fault type in the target domain, that is, the labels of the same fault are consistent); the sample data in the source domain and the sample data in the target domain are respectively expressed as:
[0065]
[0066] in is the labeled source domain sample data, n s is the number of source domain samples, Represents the i-th source domain sample and the corresponding label; is the unlabeled target domain sample data, n t is the number of samples in the target domain, Indicates the jth target domain sample, the source domain and the target domain have a common category C n .
[0067] S2, inputting the sample data of the source domain and the training sample data of the target domain into the neural network, thereby obtaining the sample features of the source domain and the target domain extracted by the deep part of the neural network, the difference in the output features of the target domain training samples in the shallow part and the deep part of the neural network, and the pseudo labels of the target domain training samples;
[0068] After the data is input into the neural network:
[0069] S2.1. Use a neural network to extract features from the sample data of the source domain and the target domain to obtain the corresponding sample features:
[0070]
[0071] Among them, F D is the deep network part, θ s ,θ t are network parameters;
[0072] S2.2, Jensen-Shannon divergence is used to measure the difference between the sample data of the target domain in the shallow part and the deep part of the neural network. The specific calculation is as follows:
[0073] Calculate the target domain sample data input shallow part and deep part to get the average distribution of feature distribution
[0074]
[0075] Where F S It is the shallow network part;
[0076] Calculate the Jensen-Shannon divergence after obtaining the mean distribution:
[0077]
[0078] D JS The value is used as the pseudo label correction value;
[0079] S2.3. Obtain the one-hot encoding of the pseudo-label of the target domain sample data through the trained neural network.
[0080] S3, after processing the data obtained in step S2, the source domain classification loss function and the pseudo label correction loss function are obtained, and the joint domain adaptation function of the source domain and the target domain is obtained by using the maximum multi-core mean difference calculation process, which is specifically:
[0081] S3.1. Use multi-kernel maximum mean difference (MKMMD) to calculate the marginal distribution difference between the source domain and the target domain, where the marginal distribution adaptive loss L MDA The calculation is as follows:
[0082]
[0083]
[0084] In the formula, G(·) is the network feature extraction part, is the training sample data of the target domain;
[0085] S3.2. Use multi-kernel maximum mean difference (MKMMD) to calculate the conditional distribution difference between the source domain and the target domain, where the conditional distribution adaptation loss L CDA The calculation is as follows:
[0086]
[0087]
[0088] S3.3. Calculation of adaptive factor u:
[0089]
[0090] S3.4, the weighted joint domain adaptation function is calculated as follows:
[0091] L WJDA =(1-μ)·L MDA +μ·L CDA
[0092] S3.5. Calculation of source domain classification loss:
[0093]
[0094] in is the one-hot encoding of the true label of the i-th source domain sample data;
[0095] S3.6, combined with step S2 to obtain the pseudo-label correction value D JS , the corrected pseudo-label loss function is calculated as follows:
[0096]
[0097] L rect =E[exp{-D JS}·L CE +D JS ]
[0098] in is the one-hot encoding of the pseudo-label of the j-th target domain sample data.
[0099] S4, combining the source domain classification loss function, the joint domain adaptation function of the source domain and the target domain, and the pseudo-label correction loss function obtained in step S3 as the overall network loss function; the overall network loss function is specifically:
[0100] L total =L cls +λ·L WJDA +0.1 L rect
[0101] Where λ is the balance factor between losses.
[0102] S5, using the training data in step S1 and the overall network loss function obtained in step S4 to train the neural network, and then using Adamax optimizer to update the network parameters;
[0103]
[0104] Where θ represents the network parameters and α represents the learning rate.
[0105] S6. Input the test sample data of the target domain in step S1 into the neural network trained in step S5. Finally, the neural network performs bearing fault diagnosis on the input data to obtain the accuracy of the bearing fault diagnosis of the neural network.
[0106] Tests prove:
[0107] In order to verify that the bearing fault diagnosis method of the present invention has better accuracy than the existing methods, the bearing fault diagnosis data sets selected by the present invention are respectively from the bearing data of Polytechnic University of Turin (DIRG), the bearing data set of Jiangnan University (JNU), and the bearing data set of Huazhong University of Science and Technology (HUST).
[0108] Step 1. Dataset division: The bearing dataset of Politecnico di Torino (DIRG) and the bearing dataset of Huazhong University of Science and Technology (HUST) are divided into D1, D2 and H1, H2 according to different loads; the bearing dataset of Jiangnan University (JNU) is divided into J1, J2 according to different speeds. The division details are shown in Table 1.
[0109] Table 1 Information on different working conditions for each dataset
[0110]
[0111] Step 2. Experimental data division: We use two datasets of different working conditions of the same test bench as the source domain and target domain, such as D1-D2. It can be interpreted that D1 is a labeled source domain and D2 is an unlabeled target domain. Therefore, the three datasets can be divided into six tasks, namely D1-D2, D2-D1, J1-J2, J2-J1, H1-H2, and H2-H1. The specific division of the experimental data is shown in Table 2, including normal state (Normal), inner race fault (IF), outer race fault (OF), rolling element fault (BF) Table 2. Specific details of the experimental data
[0112]
[0113] Step 3: Input the data into the method of the present invention for training and testing, and set the learning rate to 0.01.
[0114] Step 4: In order to evaluate the performance of fault diagnosis, the measurement of diagnostic accuracy is introduced. Diagnostic accuracy is defined as the ratio of the number of correctly classified test samples to the total number of test samples.
[0115] Step 5. In order to evaluate the superiority of the present invention, we compared the method proposed by the present invention (MKWJDAN) with other current mainstream bearing fault diagnosis methods, and measured them using the F1score index. The comparison results are shown in Table 3. Figure 3 and Figure 4 shown.
[0116] Table 3 F1score results of each model (%)
[0117]
[0118] From Table 3, and Figure 3 and 4 It can be seen that the present invention has more outstanding accuracy and robustness in the unsupervised bearing fault diagnosis task on different data sets.
[0119] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A bearing fault diagnosis method based on a multi-core weighted joint domain adaptive network, characterized in that: The specific steps are: S1. Obtain sample data of the source domain and the target domain respectively, divide the sample data of the target domain into training sample data and test sample data, and use the sample data of the source domain and the training sample data of the target domain as training data; S2. Input the sample data of the source domain and the training sample data of the target domain into the neural network, so as to obtain the sample features of the source domain and the target domain extracted by the deep part of the neural network, the difference in the output features of the target domain training samples in the shallow part and the deep part of the neural network, and the pseudo labels of the target domain training samples; after the data is input into the neural network: S2.
1. Use a neural network to extract features from the sample data of the source domain and the target domain to obtain the corresponding sample features: where F D is the deep network part, θ s ,θ t are network parameters; represents the i-th source domain sample; represents the jth target domain sample; S2.2, Jensen-Shannon divergence is used to measure the difference between the sample data of the target domain in the shallow part and the deep part of the neural network. The specific calculation is as follows: Calculate the target domain sample data input shallow part and deep part to get the average distribution of feature distribution Where F S It is the shallow network part; Calculate the Jensen-Shannon divergence after obtaining the mean distribution: D JS The value is used as the pseudo label correction value; S2.3, obtaining the one-hot encoding of the pseudo-label of the target domain sample data through the trained neural network; S3, after processing the data obtained in step S2, a source domain classification loss function and a pseudo label correction loss function are obtained, and a joint domain adaptation function of the source domain and the target domain is obtained by using the maximum multi-core mean difference calculation process; S4, combining the source domain classification loss function, the joint domain adaptation function of the source domain and the target domain, and the pseudo-label correction loss function obtained in step S3 as the overall network loss function; S5, training the neural network using the training data in step S1 and the overall network loss function obtained in step S4, thereby updating the network parameters; S6. Input the test sample data of the target domain in step S1 into the neural network trained in step S5, and finally the neural network performs bearing fault diagnosis on the input data.
2. The bearing fault diagnosis method based on multi-core weighted joint domain adaptive network according to claim 1 is characterized in that: The sample data of each of the source domain and the target domain include data of faulty bearings and data of normal bearings.
3. The bearing fault diagnosis method based on multi-core weighted joint domain adaptive network according to claim 1 is characterized in that: In step S1, the sample data of the source domain is labeled, and the sample data of the target domain is unlabeled; the sample data categories of the source domain and the target domain are the same; the sample data of the source domain and the sample data of the target domain are respectively expressed as: where θ S is the labeled source domain sample data, n s is the number of source domain samples, Represents the i-th source domain sample and the corresponding label; θ T is the unlabeled target domain sample data, n t is the number of samples in the target domain, Indicates the jth target domain sample, the source domain and the target domain have a common category C n .
4. The bearing fault diagnosis method based on multi-core weighted joint domain adaptive network according to claim 1 is characterized in that: The step S3 is specifically as follows: S3.
1. Use the multi-core maximum mean difference to calculate the marginal distribution difference between the source domain and the target domain, where the marginal distribution adaptive loss L MDA The calculation is as follows: L MDA =MKMMD(θ S ,θ T1 ) In the formula, G(·) is the network feature extraction part, θ T1 is the training sample data of the target domain; S3.
2. Use the multi-core maximum mean difference to calculate the conditional distribution difference between the source domain and the target domain, where the conditional distribution adaptive loss L CDA The calculation is as follows: L CDA =MKMMD conditional (θ S ,θ T1 ) S3.
3. Calculation of adaptive factor μ: S3.
4. The joint domain adaptation function of the source domain and the target domain is calculated as follows: L WJDA =(1-μ)·L MDA +μ·L CDA S3.
5. Calculation of source domain classification loss function: in is the one-hot encoding of the true label of the i-th source domain sample data; S3.6, combined with step S2 to obtain the pseudo-label correction value D JS , the corrected pseudo-label loss function is calculated as follows: L rect =E[exp(-D JS )·L CE +D JS ] in is the one-hot encoding of the pseudo-label of the j-th target domain sample data.
5. The bearing fault diagnosis method based on multi-core weighted joint domain adaptive network according to claim 4 is characterized in that: The overall network loss function in step S4 is specifically: L total =L cls +λ·L WJDA +0.1 L rect Where λ is the balance factor between losses.
6. The bearing fault diagnosis method based on multi-core weighted joint domain adaptive network according to claim 1 is characterized in that: In step S5, the Adamax optimizer is used to update the network parameters; Where θ represents the network parameters and α represents the learning rate.
Citation Information
Patent Citations
Rolling bearing fault cross-equipment depth domain adaptive migration diagnosis method
CN118410306A