An Unsupervised Bearing Fault Diagnosis Method Based on Adaptive Residual Adversarial Network
By adopting an adaptive residual adversarial network in bearing fault diagnosis, depth features are extracted and probability distributions are aligned, the problem of cross-domain data distribution differences is solved, and bearing fault diagnosis with high accuracy and good generalization performance is achieved.
Patent Information
- Application Number
- CN202210863740.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-07-20
AI Technical Summary
The existing bearing fault diagnosis methods are difficult to meet the distribution differences of cross-domain data in actual engineering, resulting in limited generalization capabilities of the model and difficult to effectively apply in actual industrial environments.
Unsupervised method based on adaptive residual adversarial network is adopted to extract the deep features of the source domain and the target domain through the deep residual network, and use adversarial learning and multi-core maximum mean difference to accurately align the edge probability distribution and conditional probability distribution, thereby realizing cross-domain bearing fault diagnosis.
It realizes high identification accuracy and good generalization performance for cross-domain bearing fault diagnosis, and can be effectively applied in actual industrial environments.
Smart Images

Figure CN115127814B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of rolling bearing fault detection in rotating machinery, and in particular to an unsupervised bearing fault diagnosis method based on an adaptive residual adversarial network. Background Art
[0002] With the development and progress of industrial production and scientific and technological levels, rotating machinery is also continuously developing towards high speed, continuity, and automation, effectively improving production efficiency, ensuring product quality, and saving energy and manpower. As a key component of rotating machinery, the health status of rolling bearings is related to the operation safety of the entire equipment. However, bearings work in harsh environments, and it is inevitable that abnormalities will occur during high-load operation. Once a failure occurs, it may cause economic losses or even major safety accidents. To ensure the stable operation of rotating machinery, it is necessary to conduct early fault diagnosis on rolling bearings.
[0003] In recent years, bearing fault diagnosis methods based on deep learning have been widely used, mainly because deep learning has powerful data processing capabilities and feature learning capabilities, and accurate and efficient fault diagnosis can be achieved without the need for human labor and prior knowledge. Common deep learning-based diagnosis methods include convolutional neural networks and autoencoders, and both have been successfully applied to bearing fault diagnosis. However, most existing studies achieve ideal results based on a premise that the training data and test data have the same distribution. In actual engineering, due to factors such as changes in operating conditions, wear of mechanical equipment, and environmental noise, it is difficult for the training data and test data of the same equipment to meet the above premise. Therefore, for the actual use scenarios of bearings, it is an urgent need in current bearing fault diagnosis research to study unsupervised fault diagnosis methods with strong generalization ability and high accuracy.
[0004] Transfer learning can transfer the powerful skills learned to related problems and has now been widely applied to handle bearing fault diagnosis tasks in complex environments. As a popular branch of transfer learning, unsupervised domain adaptation methods have the ability to bridge the distribution differences between domains and explore domain-invariant features, and have been widely used in the field of image recognition and have been introduced into the field of cross-domain bearing fault diagnosis. Although unsupervised domain adaptation methods have been applied to the field of fault diagnosis, most current methods focus on considering the marginal probability distribution differences between the source domain and the target domain, while ignoring the conditional probability distribution differences. Even when some researchers consider the cross-domain overall distribution differences, they simply add the cross-domain marginal probability distribution and the conditional probability distribution directly, without considering the relationship between the two, resulting in limited model generalization ability and difficulty in applying to actual industrial environments to solve cross-domain bearing fault diagnosis problems.
[0005] From the above analysis, it can be seen that neither the difference in marginal probability distributions between a single measurement source domain and the target domain nor simply adding cross-domain marginal probability distributions and conditional probability distributions can achieve good model generalization ability, let alone achieve good bearing cross-domain fault diagnosis effect. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide an unsupervised bearing fault diagnosis method based on an adaptive residual adversarial network, which uses a deep residual network to extract deep features of the original vibration data of the source domain and the target domain, and uses adversarial learning and multi-kernel maximum mean discrepancy to simultaneously and accurately align the marginal probability distributions and conditional probability distributions of the source domain and the target domain, and realizes cross-domain bearing fault diagnosis with high recognition accuracy and generalization performance.
[0007] To solve the above technical problem, the technical solution adopted by the present invention is: an unsupervised bearing fault diagnosis method based on an adaptive residual adversarial network, including the following steps:
[0008] Step S1, generate a source domain classification network model, and obtain a feature extractor and a source domain classifier through supervised training according to the bearing historical vibration signals;
[0009] Step S2, generate an unsupervised bearing fault diagnosis model based on an adaptive residual adversarial network, and optimize the parameters of the unsupervised model based on the adaptive residual adversarial network according to the bearing historical vibration signals under different working conditions;
[0010] Step S3, perform bearing fault diagnosis based on the measured values of the target bearing vibration.
[0011] A further improvement of the technical solution of the present invention lies in: the specific process of the step S1 is as follows:
[0012] Step S11, construct a one-dimensional residual network as a signal feature extractor: an improved residual block constructed based on a convolutional layer with residual connection is used as the basic structure of the feature extractor network. Structurally, multiple improved residual blocks are connected, and a global average pooling layer is applied after multiple improved residual blocks to further reduce the dimension of the features and compress the network parameters;
[0013] Step S12, construct the network structure of the source domain classifier: use three fully connected layers to learn the data features extracted in the step S11, connect a Softmax layer after the three fully connected layers, and classify the fault categories through the Softmax layer;
[0014] Step S13, obtain the bearing vibration signals with labels to construct a source domain data set, and use the normalized data as the input data for training the model;
[0015] Step S14: Use the source domain data obtained in step S13 as the input data of the model, and input it into the feature extractor and the source domain classifier network generated in steps S11 and S12. The source domain data is used as the input of the feature extractor, and the output of the feature extractor is used as the input of the source domain classifier. Continuously adjust the model parameters of the feature extractor and the source domain classifier through backpropagation, and stop training when the maximum number of training times is reached or when the loss function of the classifier reaches a preset value within the range of the number of training times, so as to obtain a pre-trained feature extractor and source domain classifier.
[0016] A further improvement of the technical solution of the present invention lies in that: the one-dimensional residual network in step S11 includes two initial convolutional blocks, four improved residual blocks and a global average pooling layer. The four improved residual blocks are each composed of two convolutional layers, and each convolutional layer uses the SELU activation function. The SELU activation function is as follows:
[0017]
[0018] where α and λ are constants.
[0019] A further improvement of the technical solution of the present invention lies in that: the specific process of step S2 is as follows:
[0020] Step S21: Use a three-layer fully connected layer followed by a Softmax layer to construct the network structure of the target domain classifier;
[0021] Step S22: Construct the network structure of the domain discriminator according to the idea of adversarial learning. The domain discriminator is used to distinguish whether the features processed by the feature extractor come from the source domain or the target domain. The domain discriminator is composed of three layers of fully connected layers and one layer of Softmax layer;
[0022] Step S23: Assign weights to the domain discriminator and the target domain classifier by randomly initializing the weights, and normalize the weights to meet the requirements of the SELU activation function;
[0023] Step S24: Introduce the feature extractor network and the source domain classifier network obtained in step S14, and use the network parameters obtained in S14 as the initialization parameters;
[0024] Step S25: Obtain bearing vibration data under two different working conditions, and use them as the source domain dataset and the target domain dataset respectively. The data after normalization of the two is used as the input data for training the model;
[0025] Step S26: Use the source domain and target domain data obtained in step S25 as the input of the feature extractor, and extract the data features of the source domain and the target domain respectively;
[0026] Step S27: Use the source domain data features obtained in step S26 as the input of the source domain classifier, and the target domain data features as the input of the target domain classifier. Calculate the gap between the outputs of the source domain classifier and the target domain classifier using the multi-kernel maximum mean discrepancy. Based on the label classification results, generate the first loss function using the cross-entropy loss; based on the obtained domain classification results, generate the second loss function using the cross-entropy loss. Use the source domain data features and target domain data features obtained in step S26 as the input of the domain discriminator, and generate the third loss function according to the MK-MMD loss in the fully connected layer;
[0027] Step S28: Construct the overall loss function of the unsupervised model based on the residual joint adaptive network using the first loss function, the second loss function, and the third loss function obtained in step S27 according to the corresponding weights. Calculate a dynamic weighting coefficient μ by calculating the MK-MMD distance between the domain and the category. The dynamic weighting coefficient μ quantitatively and qualitatively combines the MK-MMD distance with multiple adversarial domain losses to form a joint distribution distance to dynamically update the weight coefficient;
[0028] Step S29: Repeat steps S25 - S28 and continuously adjust the model parameters of the feature extractor, classifier, and domain discriminator through backpropagation according to the corresponding loss functions. Stop training when the maximum number of training times is reached or when the loss function reaches a preset value within the range of training times, and obtain the optimized unsupervised model based on the adaptive residual adversarial network.
[0029] A further improvement of the technical solution of the present invention lies in that: based on the label classification results in step S27, the formula for using the cross-entropy loss as the first loss function is as follows:
[0030]
[0031] where, is the cross-entropy classification loss function, G f is the feature extractor representation function, G y is the classifier representation function, x i is the i-th sample, y i is the category label corresponding to x o n s is the total number of source domain samples;
[0032] Based on the domain classification results, the formula for using the cross-entropy loss as the second loss function is as follows:
[0033]
[0034] where, is the cross-entropy loss function of the domain classifier, D is the domain discriminator representation function, n tis the total number of samples in the target domain, d i is x i corresponding domain label and d i = 0 or 1;
[0035] Aligning the conditional probability distribution of MK-MMD in the fully connected layer as the third loss function:
[0036]
[0037] Among them, is the high-level feature of the i-th sample in the source domain, is the high-level feature of the j-th sample in the source domain, is the high-level feature of the i-th sample in the target domain, is the high-level feature of the j-th sample in the target domain, and K(.,.) is the Gaussian kernel function.
[0038] A further improvement of the technical solution of the present invention is that the overall loss function in step S28 is as follows:
[0039]
[0040] Among them, θ f , θ y , θ d and respectively represent the parameters of the feature extractor, classifier, domain discriminator, and MK-MMD, and the coefficient μ is a dynamic weighting coefficient.
[0041] A further improvement of the technical solution of the present invention is that the specific process of step S3 is as follows:
[0042] Step S31, obtain the bearing vibration data generated under the same working conditions as the target domain in step S25 and perform normalization processing. The normalized data is used as the input of the unsupervised model based on the adaptive residual adversarial network. According to the output structure of the network, judge the type of fault, and compare with the true fault label to judge whether the transfer diagnosis is successful;
[0043] Step S32, repeat step S31 and record whether the fault diagnosed each time is correct. After repeating a sufficient number of times, calculate the accuracy of the transfer diagnosis.
[0044] Due to the adoption of the above technical solution, the technical progress achieved by the present invention is:
[0045] The unsupervised bearing fault diagnosis method based on the adaptive residual adversarial network provided by the present invention improves the basic residual structure, uses a deep residual network to extract the deep features of the original vibration data in the source domain and the target domain, and uses adversarial learning and multi-kernel maximum mean discrepancy to accurately align the marginal probability distribution and conditional probability distribution of the source domain and the target domain at the same time. In addition, in order to accelerate the optimization iteration speed and improve the diagnosis accuracy, a dynamic weighting coefficient is designed to dynamically measure the relative importance of the differences between the two distributions, realizing cross-domain bearing fault diagnosis and having a high recognition accuracy and generalization performance. Description of the Drawings
[0046] Figure 1 It is a flowchart of the implementation method of the present invention;
[0047] Figure 2 It is a corresponding table of fault categories and labels of the data used in the present invention;
[0048] Figure 3 It is a table of network structure parameters of the feature extractor of the present invention;
[0049] Figure 4 It is a structural diagram of the improved residual block;
[0050] Figure 5 It is a structural diagram of the adaptive residual adversarial network based on the present invention;
[0051] Figure 6 It is a confusion matrix diagram of the 0Hp→2Hp migration task of the adaptive residual adversarial network based on the present invention. Detailed Embodiment
[0052] The following further describes the present invention in detail with reference to the embodiments:
[0053] As Figure 1 shown, an unsupervised bearing fault diagnosis method based on an adaptive residual adversarial network includes the following steps:
[0054] Step S1, generate a source domain classification network model, and obtain a feature extractor and a source domain classifier through supervised training according to the bearing historical vibration signals;
[0055] Step S2, generate an unsupervised bearing fault diagnosis model based on the adaptive residual adversarial network, and optimize the parameters of the unsupervised model based on the adaptive residual adversarial network according to the bearing historical vibration signals under different working conditions;
[0056] Step S3, perform bearing fault diagnosis based on the measured values of the target bearing vibration.
[0057] In this embodiment, the Case Western Reserve University (CWRU) bearing dataset is selected. The CWRU bearing dataset is provided by Case Western Reserve University in the United States. This dataset consists of vibration signals of a motor bearing under different operating conditions. The operating conditions include two variable parameters, load and rotational speed. The states of the motor bearing include normal, ball fault, inner race fault, and outer race fault. In this embodiment, ten types of fault vibration data are used. The processing results of the data set labels are shown in Figure 2 。
[0058] The specific steps of step S1 in this embodiment are as follows:
[0059] Step S11: Construct an improved one-dimensional residual network as a signal feature extractor. The one-dimensional residual network includes two initial convolutional blocks, four improved residual blocks, and a global average pooling layer. The four improved residual blocks are each composed of two convolutional layers, and each convolutional layer uses the SELU activation function. The SELU activation function is as follows:
[0060]
[0061] where α and λ are constants. The fixed value of α is 1.6732632423543772848170429916717, and the fixed value of λ is 01.0507009873554804934193349852946.
[0062] The parameter design of the one-dimensional residual network is shown in Figure 3 。Based on the convolutional layer with residual connection, the improved residual block is constructed as the basic structure of the feature extraction network. The design of the improved residual block is shown in Figure 4 ,and multiple improved residual blocks are connected in structure. The activation function is an important source of the non-linear representation ability of the neural network, and the normalization operation of the neural network can improve the convergence speed of the neural network and reduce the internal covariate shift during the training process of the neural network. In the improved one-dimensional residual network, the SELU activation function combines the two operations. The SELU activation function not only retains the function of the activation function but also can perform self-normalization processing on the input data, which can replace a large number of traditional normalization functions such as BN.
[0063] The improved residual block has a skip connection that directly connects the input and output. The improved residual block is more conducive to the backpropagation of errors inside the neural network, so its parameters are easier to train. In addition, the zero-padding strategy is adopted. The padding size of the initial convolutional layer is 5, and the padding size of the rest is 1. Global average pooling is applied after multiple improved residual blocks to further reduce the dimension of the features and compress the number of network parameters;
[0064] Step S12, construct the network structure of the source domain classifier: Use three fully connected layers to learn the data features extracted in step S11. The first fully connected layer uses 50 neurons, the second fully connected layer uses 20 neurons, and the third fully connected layer uses the same number as the number of fault categories, which is 10 neurons. Using three fully connected layers is more conducive to comprehensively extracting the extracted features and strengthening the classification effect. After the three fully connected layers, connect a Softmax layer, and classify the fault categories through the Softmax layer;
[0065] Step S13, obtain the bearing vibration signals with labels to construct the source domain dataset, and use the normalized data as the input data for training the model:
[0066] In this embodiment, the bearing dataset of Case Western Reserve University is selected, and the data collected under different loads (0Hp, 1Hp, 2Hp, 3Hp) is selected for training. Taking the transfer task 0Hp→2Hp as an example, first use the 0Hp data after normalization as the source domain labeled dataset. Use the drive end vibration data with a sampling frequency of 12KHz, set the length of a single sample to 1024, and take 1000 samples for each category. During the experiment, the samples of each category are divided into a training set and a test set in a ratio of 8:2 to establish a labeled source domain dataset
[0067] Step S14, use the source domain data obtained in step S13 as the input data of the model, and input it into the feature extractor and the source domain classifier network generated in steps S11 and S12. Set the network learning rate to 0.0001, the batch size to 128, and the number of training times to 50. The source domain data is used as the input of the feature extractor, and the output of the feature extractor is used as the input of the source domain classifier. Continuously adjust the model parameters of the feature extractor and the source domain classifier through backpropagation, and stop training when the maximum number of training times is reached or when the loss function of the classifier reaches the preset value within the number of training times, and obtain the pre-trained feature extractor and source domain classifier.
[0068] Step 2 of this embodiment specifically includes:
[0069] Step S21, construct the network structure of the target domain classifier by using three fully connected layers followed by a Softmax layer. The first fully connected layer uses 50 neurons, the second fully connected layer uses 20 neurons, and the third fully connected layer uses the same number as the number of fault categories, which is 10 neurons. Using three fully connected layers is more conducive to comprehensively extracting the extracted features and strengthening the classification effect. Finally, classify the fault categories through Softmax;
[0070] Step S22: Construct the network structure of the domain discriminator according to the idea of adversarial learning. The domain discriminator is used to distinguish whether the features processed by the feature extractor come from the source domain or the target domain. The domain discriminator consists of three fully connected layers and one Softmax layer. Among them, the first fully connected layer uses 50 neurons, the second fully connected layer uses 20 neurons, and the third fully connected layer uses 2 neurons. Using three fully connected layers is more conducive to comprehensively extracting the extracted features and strengthening the classification effect. Finally, the fault categories are classified through Softmax;
[0071] Step S23: Assign weights to the domain discriminator and the target domain classifier by randomly initializing the weights, and normalize the weights to meet the requirements of the SELU activation function;
[0072] Step S24: As the transfer task 0Hp→2Hp, introduce the feature extractor network and the source domain classifier network obtained in step S14, and use the network parameters obtained in S14 as the initialization parameters;
[0073] Step S25: Obtain the bearing vibration data under two different working conditions, which are used as the source domain dataset and the target domain dataset respectively. The normalized data of the two are used as the input data for training the model.
[0074] In this embodiment, the bearing dataset of Case Western Reserve University is selected, and the data collected under different loads (0Hp, 1Hp, 2Hp, 3Hp) is used for training. Taking the transfer task 0Hp→2Hp as an example, the normalized 0Hp data is used as the source domain data, and the normalized 2Hp data is used as the target domain data. The driving end vibration data with a sampling frequency of 12KHz is used, the length of a single sample is set to 1024, and 1000 samples are taken for each category. During the experiment, the samples of each category are divided into a training set and a test set in a ratio of 8:2, and an unlabeled target domain training set is established
[0075] Step S26: Use the source domain and target domain data obtained in step S25 as the input of the feature extractor to extract the data features of the source domain and the target domain respectively;
[0076] Step S27: Use the source domain data features obtained in step S26 as the input of the source domain classifier, and the target domain data features as the input of the target domain classifier. Calculate the gap between the outputs of the source domain classifier and the target domain classifier using the multi-kernel maximum mean discrepancy. Based on the label classification results, generate the first loss function using the cross-entropy loss; based on the obtained domain classification results, generate the second loss function using the cross-entropy loss. Use the source domain data features and target domain data features obtained in step S26 as the input of the domain discriminator, and generate the third loss function according to the MK-MMD loss in the fully connected layer;
[0077] Adversarial learning generally consists of three parts: a feature extractor, a classifier, and a domain discriminator. The optimization objective of the feature extractor is to confuse the domain discriminator so that it cannot distinguish whether the data comes from the source domain or the target domain; while the optimization objective of the domain discriminator is to distinguish whether the features processed by the feature extractor come from the source domain or the target domain, and a loss function is designed based on this.
[0078] Based on the label classification results, cross-entropy loss is used as the first loss function:
[0079]
[0080] Among them, is the cross-entropy classification loss function, G f is the feature extractor representation function, G y is the classifier representation function, x i is the i-th sample, y i is the corresponding one of x i category label, n s is the total number of samples in the source domain;
[0081] Based on the domain classification results, cross-entropy loss is used as the second loss function:
[0082]
[0083] Among them, is the cross-entropy loss function of the domain classifier, D is the domain discriminator representation function, n t is the total number of samples in the target domain, d i is the corresponding one of x i domain label and d i = 0 or 1;
[0084] Based on the alignment of the conditional probability distributions of MK-MMD in the fully connected layer as the third loss function:
[0085]
[0086] Among them, is the high-level feature of the i-th sample in the source domain, is the high-level feature of the j-th sample in the source domain, is the high-level feature of the i-th sample in the target domain, is the high-level feature of the j-th sample in the target domain, K(.,.) is the Gaussian kernel function.
[0087] Step S28: Use a dynamic weighting coefficient μ to dynamically measure the relative importance of the marginal probability distribution and the conditional probability distribution. Construct the overall loss function of the unsupervised model based on the adaptive residual adversarial network for the first loss function, the second loss function, and the third loss function obtained in Step S27 according to the corresponding weights, and use the update of the overall loss function to dynamically update the weights.
[0088] Most researchers use the methods of average search and random guessing to select appropriate weights, but these methods are time-consuming and prone to missing the optimal value. In the present invention, a dynamic weighting coefficient μ is obtained by calculating the MK-MMD distance between the domain and the category. Specifically, MK-MMD is added to the second FC layer of the domain discriminator to calculate the difference in the distribution of fault features within the same category, and the weight size is determined by the ratio of the MK-MMD between domains. When the coefficient is initialized, the dynamic weighting coefficient is set to 0.5. During subsequent training, the dynamic weighting coefficient is calculated in real time according to the actual loss. During the training process, the dynamic weighting coefficient μ is used to dynamically measure the relative importance of the marginal probability distribution and the conditional probability distribution. If the marginal loss is large, the dynamic weighting factor will increase, prompting the network to pay more attention to the alignment of the marginal probability distribution, and vice versa.
[0089] The overall loss function is as follows:
[0090]
[0091] where θ f , θ y , θ d and represent the parameters of the feature extractor, the classifier, the domain discriminator, and MK-MMD respectively, and the coefficient μ is the dynamic weighting coefficient.
[0092] Step S29: Set the network learning rate to 0.0001, the batch size to 128, and the number of training times to 50. Repeat Step S25 - S28 and continuously adjust the model parameters of the feature extractor, the classifier, and the domain discriminator through backpropagation according to the corresponding loss functions. Stop training when the maximum number of training times is reached or when the loss function reaches the preset value within the range of the number of training times, and obtain the optimized unsupervised model based on the adaptive residual adversarial network. The overall structure of the model is shown in Figure 5 .
[0093] Step 3 of this embodiment specifically includes the following steps:
[0094] Step S31: Obtain the bearing vibration data generated under the same working conditions as the target domain in Step S25 and perform normalization processing. The normalized data is used as the input of the unsupervised model based on the residual joint adaptive network. Determine the type of fault according to the output structure of the network, and compare with the true fault label to determine whether the transfer diagnosis is successful.
[0095] In this embodiment, the Case Western Reserve University bearing dataset is selected, and the data collected under different loads (0 Hp, 1 Hp, 2 Hp, 3 Hp) is selected for training. Taking the transfer task 0 Hp→2 Hp as an example, the data after normalization of 2 Hp is used as the target domain data. The vibration data at the driving end with a sampling frequency of 12 KHz is used, the length of a single sample is set to 1024, and 1000 samples are taken for each category. During the experiment, the samples of each category are divided into a training set and a test set at a ratio of 8:2 to construct a labeled target domain test set The normalized data is used as the input of the unsupervised model based on the residual joint adaptive network. The type of fault is judged according to the output structure of the network, and whether the transfer diagnosis is successful is judged by comparing with the true fault label.
[0096] Step S32: Repeat step S31 and record whether the fault diagnosed each time is correct. After repeating a sufficient number of times, calculate the accuracy of the transfer diagnosis and draw a confusion matrix as shown in Figure 6 .
Claims
1. An unsupervised bearing fault diagnosis method based on an adaptive residual adversarial network, characterized in that: It includes the following steps: Step S1: Generate a source domain classification network model, and obtain a feature extractor and a source domain classifier through supervised training based on bearing historical vibration signals; The specific process of step S1 includes: Step S11: Construct a one-dimensional residual network as a signal feature extractor: Use an improved residual block constructed based on convolutional layers with residual connections as the basic structure of the feature extractor network. Structurally, multiple improved residual blocks are connected, and a global average pooling layer is applied after multiple improved residual blocks to further reduce the dimension of the features and compress the network parameters. The one-dimensional residual network includes two initial convolutional blocks, four improved residual blocks, and a global average pooling layer. Each of the four improved residual blocks consists of two convolutional layers, and each convolutional layer uses the SELU activation function; Step S12: Construct the network structure of the source domain classifier: Use three fully connected layers to learn the data features extracted in step S11. A Softmax layer is connected after the three fully connected layers to classify the fault categories through the Softmax layer; Step S2: Generate an unsupervised bearing fault diagnosis model based on an adaptive residual adversarial network, and optimize the parameters of the unsupervised model based on the adaptive residual adversarial network according to bearing historical vibration signals under different working conditions; The specific process of step S2 is as follows: Step S21: Construct the network structure of the target domain classifier by connecting a Softmax layer after three fully connected layers; Step S22: Construct the network structure of the domain discriminator according to the adversarial learning idea. The domain discriminator is used to distinguish whether the features processed by the feature extractor come from the source domain or the target domain. The domain discriminator consists of three fully connected layers and one Softmax layer; Step S23: Assign weights to the domain discriminator and the target domain classifier by randomly initializing the weights, and normalize the weights to meet the requirements of the SELU activation function; Step S24: Introduce the feature extractor network and the source domain classifier network obtained in step S14, and use the network parameters obtained in S14 as the initial parameters; Step S25: Obtain bearing vibration data under two different working conditions, and use them as the source domain dataset and the target domain dataset respectively. The normalized data of the two are used as the input data for training the model; Step S26: Use the source domain and target domain data obtained in step S25 as the input of the feature extractor to extract the respective data features of the source domain and the target domain; Step S27: Use the source domain data features obtained in step S26 as the input of the source domain classifier, and the target domain data features as the input of the target domain classifier. Calculate the gap between the outputs of the source domain classifier and the target domain classifier using the multi-kernel maximum mean discrepancy. Based on the label classification results, generate a first loss function using the cross-entropy loss. The formula of the first loss function is as follows: Among them, is the cross-entropy classification loss function, G f is the feature extractor representation function, G y is the classifier representation function, x i is the i-th sample, y i is x i corresponding class label, n s is the total number of samples in the source domain; Based on the obtained domain classification results, generate a second loss function using the cross-entropy loss. The formula of the second loss function is as follows: Among them, is the cross-entropy loss function of the domain classifier, D is the domain discriminator representation function, and n t is the total number of samples in the target domain, and d i is the domain label corresponding to x i and d i = 0 or 1; Use the source domain data features and target domain data features obtained in step S26 as the input of the domain discriminator, and generate a third loss function according to the MK-MMD loss in the fully connected layer. The formula of the third loss function is as follows: Among them, is the high-level feature of the i-th sample in the source domain, is the high-level feature of the j-th sample in the source domain, is the high-level feature of the i-th sample in the target domain, is the high-level feature of the j-th sample in the target domain, and K(.,.) is the Gaussian kernel function; Step S28: Construct the overall loss function of the unsupervised model based on the residual joint adaptive network according to the corresponding weights for the first loss function, the second loss function, and the third loss function obtained in Step S27. The overall loss function is as follows: Among them, θ f , θ y , θ d and respectively represent the parameters of the feature extractor, classifier, domain discriminator, and MK-MMD. The coefficient μ is a dynamic weighting coefficient. By calculating the MK-MMD distance between the domain and the category, a dynamic weighting coefficient μ is obtained. The dynamic weighting coefficient μ quantitatively and qualitatively combines the MK-MMD distance with multiple adversarial domain losses to form a joint distribution distance to dynamically update the weight coefficient; Step S29: Repeat Steps S25 - S28 and continuously adjust the model parameters of the feature extractor, the classifier, and the domain discriminator through backpropagation according to the corresponding loss functions. Stop training when the maximum number of training times is reached or when the loss function reaches a preset value within the range of the number of training times, and obtain the optimized unsupervised model based on the adaptive residual adversarial network. Step S3: Perform bearing fault diagnosis based on the measured vibration values of the target bearing.
2. The unsupervised bearing fault diagnosis method based on an adaptive residual adversarial network according to claim 1, wherein: The specific process of Step S1 further includes the following: Step S13: Obtain the bearing vibration signals with labels to construct the source domain dataset, and use the normalized data as the input data for training the model. Step S14: Use the source domain data obtained in Step S13 as the input data of the model, and input it into the feature extractor and the source domain classifier network generated in Steps S11 and S12. The source domain data is used as the input of the feature extractor, and the output of the feature extractor is used as the input of the source domain classifier. Continuously adjust the model parameters of the feature extractor and the source domain classifier through backpropagation. Stop training when the maximum number of training times is reached or when the classifier loss function reaches a preset value within the range of the number of training times, and obtain the pre-trained feature extractor and source domain classifier.
3. The unsupervised bearing fault diagnosis method based on an adaptive residual adversarial network according to claim 1, characterized in that: The SELU activation function in Step S11 is as follows: where α and λ are constants.
4. The unsupervised bearing fault diagnosis method based on an adaptive residual adversarial network according to claim 1, wherein: The specific process of Step S3 is as follows: Step S31: Obtain the bearing vibration data generated under the same working conditions as the target domain in Step S25 and perform normalization processing. Use the normalized data as the input of the unsupervised model based on the adaptive residual adversarial network. Judge the type of the fault according to the output structure of the network, and compare with the true fault label to judge whether the transfer diagnosis is successful. Step S32: Repeat Step S31 and record whether the fault diagnosed each time is correct. Calculate the accuracy of the transfer diagnosis after repeating a sufficient number of times.
Citation Information
Patent Citations
Rolling bearing fault diagnosis method based on dynamic index antagonism self-adaption
CN114429152A
Vibration signal diagnostic analysis method based on EEMD and deep domain adversarial network
CN114648044A