Fault Diagnosis Method Based on Semi-Supervised Adversarial Domain Generalization Intelligent Model
By adopting a semi-supervised adversarial domain generalization intelligent model in the mechanical fault diagnosis model, combined with domain fuzzy strategy and measurement learning, the existing models have solved the problems of insufficient generalization ability and high complexity when dealing with unknown working conditions and inconsistent distribution problems, and achieved high precision and generalization ability mechanical fault diagnosis.
Patent Information
- Application Number
- CN202211558882.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-12-06
AI Technical Summary
When the existing domain generalization fault diagnosis model deals with the problems of missing labels and inconsistent distribution of some source domains, the generalization capability is insufficient and the model is highly complex, making it difficult to be applicable to mechanical fault diagnosis in unknown working conditions.
The fault diagnosis method based on the semi-supervised adversarial domain generalization intelligent model is adopted. By constructing feature extractors, label classifiers, domain classifiers and class-level optimization modules, combining domain fuzzy strategies and metric learning, domain invariance and discriminant features are extracted, and adversarial training is realized to improve the generalization ability of the model.
It significantly improves the generalization ability and recognition accuracy of the mechanical fault diagnosis model, and can realize intelligent identification of semi-supervised mechanical health status under unknown working conditions.
Smart Images

Figure CN116089812B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mechanical fault identification, and in particular to a fault diagnosis method based on a semi-supervised adversarial domain generalization intelligent model. Background Art
[0002] The intelligent identification of the health state of a mechanical system can diagnose mechanical faults in a timely manner, which is beneficial to the safe and stable operation of the mechanical system. Most of the currently developed mechanical intelligent fault diagnosis technologies are data-driven methods. Inspired by the application of deep neural networks in the field of computer pattern recognition, researchers have been keen on exploring intelligent fault diagnosis methods based on deep learning in recent years. The deep neural network has strong fault feature extraction ability and is convenient for directly constructing an end-to-end fault diagnosis scheme. Since mechanical key components are often in an environment of variable speed or variable load, there are distribution differences in data under different working conditions. However, deep learning fault diagnosis methods usually assume that the test data and the training data are of the same distribution, which hinders the application of deep learning-based fault diagnosis methods in actual diagnosis tasks. Facing the problem of fault diagnosis under variable working conditions, some scholars introduced transfer learning. By combining the domain adaptation method of transfer learning with deep neural networks, a deep domain adaptation method was proposed. Further, the adversarial idea was introduced from the generative adversarial network into the deep neural network, forming a domain adversarial neural network, which can dynamically select transferable features between different domains. These domain adaptation methods have good effects on cross-domain fault diagnosis problems and can better handle the problem of inconsistent distributions of training data and test data.
[0003] A prerequisite for the domain adaptation method to achieve excellent performance is the prior distribution of the target data. However, in actual industrial scenarios, the fault data of the target working condition is usually invisible during the training stage. When the target data cannot be obtained in advance to participate in model training, if the trained model is applied to a new working condition, catastrophic failures often occur. Therefore, it is necessary to explore a more realistic fault diagnosis model that can be generalized to invisible working conditions. During the training stage, only source domain data is required, without accessing the target data, and the trained model can be applied to the bearing fault diagnosis of the target task. A model suitable for this kind of fault diagnosis task is called a domain generalization model.
[0004] The existing domain generalization fault diagnosis model is a fully supervised generalization network model, which generally consists of three parts: feature extractor, domain discriminator and label classifier. The feature extractor takes labeled multi-source domain data as input and outputs implicit features. Both the domain discriminator and the label classifier take the extracted implicit feature vector as input and output the source and type of the feature vector respectively. Adversarial training is performed between multi-source data to extract domain invariant features to improve the generalization and robustness of the network. When using a fully supervised generalization network for intelligent fault diagnosis, it is first necessary to input the labeled multi-source domain dataset into the network for model parameter training. After the training is completed, the target domain test data is input into the trained feature extractor and label classifier for fault diagnosis to obtain its fault category.
[0005] The fully supervised generalization network model requires all source domains to have category labels in the intelligent fault diagnosis of machinery under unknown working conditions. However, labeling engineering data requires domain expert knowledge and the workload is huge. Therefore, the actual multi-source domain data set may only have some source domains with category labels while the other source domains do not have category labels. At this time, the fully supervised generalization network model is no longer applicable. In addition, the fully supervised generalization network model inputs multi-source domain data into the traditional domain adversarial neural network for adversarial training. The model either does not consider the fine-grained alignment between multiple source domain data, resulting in insufficient domain generalization ability, or requires the use of multiple domain discriminators and classifiers to improve the domain generalization ability, thereby significantly increasing the complexity of the model. Therefore, the existing domain generalization model has the following disadvantages: 1) It is not applicable when some source domain labels are missing; 2) The domain generalization ability is weak; 3) The model complexity is high. Summary of the invention
[0006] To this end, the technical problem to be solved by the present invention is to provide a fault diagnosis method based on a semi-supervised adversarial domain generalized intelligent model with strong generalization ability and high recognition accuracy.
[0007] In order to solve the above technical problems, the present invention provides a fault diagnosis method based on a semi-supervised adversarial domain generalized intelligent model, which comprises the following steps:
[0008] S1, the collected mechanical vibration time domain signal is cut into data samples, the sample length is unified, and the sample amplitude is normalized, and the data set is divided into a multi-source domain data set and a target domain data set;
[0009] S2. Construct a feature extractor, wherein the feature extractor is used to map the preprocessed data sample to the target feature space, wherein the feature extractor takes the data sample as input and takes the high-level implicit features of the data sample as output;
[0010] S3. Construct a label classifier which is used to predict the class labels of high-level implicit features. The label classifier takes the output of the feature extractor as the input and the predicted sample class labels as the output.
[0011] S4. Construct a domain classifier which is used to predict the domain labels of high-level implicit features. The domain classifier takes the output of the feature extractor as the input and the predicted domain labels as the output.
[0012] S5. Construct a class-level optimization module which is used to perform metric learning on the extracted high-level implicit features to optimize the class boundaries.
[0013] S6. Combine the feature extractor, the label classifier, the domain classifier and the class-level optimization module to construct a fault diagnosis training model.
[0014] S7. Input the multi-source domain training data set into the constructed fault diagnosis training model, extract domain-invariant data features through the domain blurring strategy, and perform model training according to the given loss function and optimization algorithm.
[0015] S8. Input the target domain test data set into the trained feature extractor and label classifier to online identify the health state categories of the samples.
[0016] In one embodiment of the present invention, the domain blurring strategy adopts the multi-class probability output of the domain classifier, and the output probability dimension is the same as the number of source domains. Through the continuous adversarial training of the feature extractor and the domain classifier, the feature extractor can blur the domain classifier, that is, it cannot determine which domain the extracted high-level implicit features come from, and it is considered that the probabilities of coming from each domain are the same.
[0017] In one embodiment of the present invention, step S7 includes:
[0018] S71. Extract high-level implicit features from the multi-source domain samples through the feature extractor, and input the features into the label classifier, the domain classifier and the class-level optimization module.
[0019] S72. For all the data from multiple source domains, minimize the cross-entropy loss of the domain classifier to optimize the domain classifier.
[0020] S73. For the labeled data from the source domain, minimize the cross-entropy loss of the label classifier to optimize the label classifier.
[0021] S74. Minimize the cross-entropy loss of the label classifier and the center loss of the class-level optimization module, while maximizing the cross-entropy loss of the domain classifier, the Shannon entropy loss of the sample domain label prediction probability, and the Shannon entropy loss of the mean of the mini-batch sample domain label prediction probabilities to optimize the feature extractor; the training objective of optimizing the domain classifier is to try to identify the input features as the correct domain labels, while the training objective of optimizing the feature extractor is to obfuscate the domain classifier so that it cannot correctly determine which domain the features come from and output an uncertain domain label, thus forming a min-max adversarial relationship;
[0022] S75. Stop training when the adversarial training makes the model reach the Nash equilibrium.
[0023] In an embodiment of the present invention, in step S1, the basis for dividing the data set is the working condition of the machine. Data samples under the same working condition are placed in the data set of the same domain, and the categories of mechanical health states included in different data sets are the same; the multi-source domain data set includes a labeled source domain and multiple unlabeled source domains for training the model; the target domain data set does not participate in the model training and is only used to test the accuracy of the model prediction results.
[0024] In an embodiment of the present invention, the optimization algorithm is the root mean square propagation algorithm, the stochastic gradient descent method, or the adaptive moment estimation algorithm.
[0025] In an embodiment of the present invention, the class-level optimization module adopts a metric learning method. By learning a class center, it constrains the features belonging to this class to approach the class center as much as possible. For unlabeled source domain samples, the class center is updated through the predicted labels given by the label classifier.
[0026] In an embodiment of the present invention, the domain classifier consists of a fully connected layer and a Softmax classifier, and the domain classification loss of the domain classifier is the cross-entropy loss of the sample prediction domain labels in the multi-source domain.
[0027] In an embodiment of the present invention, the label classifier consists of a fully connected layer and a Softmax classifier, and the label classification loss of the label classifier is the cross-entropy loss of the sample prediction class labels in the labeled source domain.
[0028] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method described in any one of the above.
[0029] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in any one of the above.
[0030] The above technical solution of the present invention has the following advantages compared with the prior art:
[0031] Based on the domain adversarial neural network, the fault diagnosis method of the present invention based on the semi-supervised adversarial domain generalization intelligent model adopts the semi-supervised learning method. By proposing a domain fuzzy strategy to eliminate the data distribution differences between multiple source domains, it learns the distribution-independent representations in the labeled source domain and the unlabeled source domain, extracts domain-invariant features, and introduces metric learning to promote the intra-class aggregation and inter-class separability of the features learned by the model, improving the discriminability of the features. Thus, it greatly improves the generalization ability of the mechanical fault diagnosis model, realizes the semi-supervised intelligent identification of mechanical health states under unknown working conditions, and has high identification accuracy.
[0032] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the drawings, details are described as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to make the content of the present invention easier to be clearly understood, the present invention will be further described in detail below according to the specific embodiments of the present invention in conjunction with the drawings, where
[0034] Figure 1 is a flowchart of the fault diagnosis method based on the semi-supervised adversarial domain generalization intelligent model in the embodiment of the present invention;
[0035] Figure 2 is a schematic diagram of the fault diagnosis training model in the embodiment of the present invention;
[0036] Figure 3 is a schematic diagram of the bearing fault diagnosis model applied to unknown working conditions in the embodiment of the present invention;
[0037] Figure 4 is a schematic diagram of the visualization clustering result of the high-level hidden features of the source domain and the target domain extracted by the fault diagnosis model in the embodiment of the present invention;
[0038] Figure 5 is a schematic diagram of the confusion matrix of the target domain bearing health state result predicted by the fault diagnosis model in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The present invention will be further described below in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments given are not intended to limit the present invention.
[0040] Embodiment 1
[0041] Refer toFigure 1 As shown, this embodiment discloses a fault diagnosis method based on a semi-supervised adversarial domain generalized intelligent model, which includes the following steps:
[0042] S1, the collected mechanical vibration time domain signal is cut into data samples, the sample length is unified, and the sample amplitude is normalized to the range of [0,1], and the data set is divided into a multi-source domain data set and a target domain data set;
[0043] The data sets are divided based on the working conditions of the machinery. Data samples under the same working condition are placed in the same domain data set. Different data sets contain the same categories of machinery health status. The multi-source domain data set contains a labeled source domain and multiple unlabeled source domains for model training. The target domain data set does not participate in model training and is only used to test the accuracy of the model prediction results. Among them, machinery includes bearings, etc.
[0044] S2. Construct a feature extractor F, wherein the feature extractor F is used to map the preprocessed data sample to the target feature space, wherein the feature extractor F takes the data sample as input and takes the high-level implicit features of the data sample as output;
[0045] Optionally, the feature extractor F includes but is not limited to being constructed by one of a fully connected network, a deep convolutional network, a deep belief network, and a deep residual network.
[0046] S3, constructing a label classifier C, wherein the label classifier C is used to predict the category label of the high-level implicit feature, and the label classifier C takes the output of the feature extractor F as input and takes the predicted sample category label as output;
[0047] Optionally, the label classifier C is composed of a fully connected layer and a Softmax classifier, and the label classification loss of the label classifier C is the cross entropy loss L of the predicted category label of the sample in the labeled source domain C .
[0048] S4, constructing a domain classifier D, wherein the domain classifier D is used to predict the domain label of the high-level implicit feature, and the domain classifier D takes the output of the feature extractor F as input and takes the predicted domain label as output;
[0049] Optionally, the domain classifier D is composed of a fully connected layer and a Softmax classifier, and the domain classification loss of the domain classifier D is the cross entropy loss L of the domain labels predicted by samples in the multi-source domains. D .
[0050] S5, constructing a class-level optimization module M, wherein the class-level optimization module M is used to perform metric learning on the extracted high-level implicit features to optimize class boundaries;
[0051] The class-level optimization module M adopts a metric learning method. By learning a class center, it constrains the features belonging to this class to approach the class center as much as possible. For unlabeled source domain samples, the class center is updated using the predicted labels given by the label classifier C. The distance metric loss of the class-level optimization module M is the center loss L of the features extracted by the feature extractor F. O 。
[0052] S6. Combine the feature extractor F, the label classifier C, the domain classifier D, and the class-level optimization module M to construct a fault diagnosis training model; refer to Figure 2 Figure 7, which is a schematic diagram of the fault diagnosis training model (semi-supervised adversarial domain generalization intelligent model).
[0053] The feature extractor F and the label classifier C form a feed-forward neural network, which is used to predict the class labels of input data samples.
[0054] The feature extractor F and the domain classifier D form a feed-forward neural network, which is used to predict the domain labels of input data samples.
[0055] S7. Input the multi-source domain training data set into the constructed fault diagnosis training model, extract domain-invariant data features through the domain blurring strategy, and perform model training according to the given loss function and optimization algorithm.
[0056] The domain blurring strategy uses the multi-class probability output of the domain classifier. The output probability dimension is the same as the number of source domains. Through the continuous adversarial training of the feature extractor F and the domain classifier D, the feature extractor F can blur the domain classifier D, that is, it cannot determine which domain the extracted high-level hidden features come from, and it is considered that the probability of coming from each domain is the same.
[0057] Specifically, step S7 includes:
[0058] S71. Extract high-level hidden features from multi-source domain samples through the feature extractor F, and input the features into the label classifier C, the domain classifier D, and the class-level optimization module M.
[0059] S72. For all data from multiple source domains, minimize the cross-entropy loss L D of the domain classifier M to optimize the domain classifier D.
[0060] S73. For the labeled data from the source domain, minimize the cross-entropy loss L D of the label classifier C to optimize the label classifier C.
[0061] S74. Minimize the cross-entropy loss L C of the label classifier C and the center loss L O of the class-level optimization module M, and at the same time maximize the cross-entropy loss L D, the Shannon entropy loss L of the predicted probability of the sample domain label S , the Shannon entropy loss L of the mean of the predicted probabilities of the mini-batch sample domain labels MS to optimize the feature extractor F; the training objective of optimizing the domain classifier D is to try to identify the input features as the correct domain labels, while the training objective of optimizing the feature extractor F is to blur the domain classifier D so that it cannot correctly determine which domain the features come from and output an uncertain domain label, thus forming a min-max adversarial relationship;
[0062] S75. Stop training when the adversarial training makes the model reach the Nash equilibrium.
[0063] Optionally, the optimization algorithm is the root mean square propagation algorithm, the stochastic gradient descent method, the adaptive moment estimation algorithm, etc.
[0064] S8. Input the target domain test data set into the trained feature extractor F and label classifier C to online identify the health state category of the sample. Implement fault diagnosis. Among them, the trained feature extractor F and label classifier C form a fault diagnosis model, refer to Figure 3 .
[0065] To verify the effectiveness of the present invention, in a specific embodiment, taking the bearing data set provided by the KAt data center of the University of Paderborn as an example, the data set contains four types of health state data: normal state (N), inner race fault (I), outer race fault (O), and inner and outer race compound fault (IO), and the fault category labels are represented by 0, 1, 2, and 3 respectively. The data collected under four working conditions are included in the experiment, and the data situation adopted in this experiment is described in Table 1. The data under each working condition is a domain, and the number of samples in each health state in each domain is 100. Use one labeled domain and two unlabeled domains in Table 1 as the source domain training data set, and the remaining one domain as the target domain test data set.
[0066] Table 1 Description of the experimental bearing data set
[0067]
[0068] In step S1, data preprocessing. Intercept the bearing time-domain signal data into samples with a length of 4096 points, and then perform [0,1] normalization processing on the sample amplitudes, and use the processed samples as the model input samples.
[0069] In step S2, construct the feature extractor F. The feature extractor F adopts the classic MK-ResCNN network, uses the preprocessed time-domain signal data samples as the model input, and outputs a high-level hidden feature vector with a length of 768.
[0070] In step S3, a label classifier C is constructed. The label classifier C uses a fully connected network, with a total of two layers designed. The dimensions of the hidden layers are 768 and 4 respectively. After the two fully connected layers, ReLU and Softmax activation functions are connected respectively. Finally, the model outputs a four-dimensional vector to represent the health status category of the input data.
[0071] In step S4, a domain classifier D is constructed. The domain classifier D uses a fully connected network, with a total of two layers designed. The dimensions of the hidden layers are 768 and 3 respectively. After the two fully connected layers, ReLU and Softmax activation functions are connected respectively. Finally, the model outputs a three-dimensional vector to represent the domain label of the input data.
[0072] In step S5, a class-level optimization module M is constructed. The class-level optimization module M is established, and the center loss of the extracted features is calculated by using the true class labels in the labeled source domain and the predicted class labels assigned by the label classifier C to the unlabeled source domain during the training process.
[0073] In step S6, a fault diagnosis training model is constructed. The feature extractor F, the label classifier C, the domain classifier D, and the class-level optimization module M are combined to construct a complete fault diagnosis training model. Refer to Figure 2 。
[0074] In step S7, model training is performed. The multi-source domain training data set is input into the constructed fault diagnosis training model. Domain-invariant data features are extracted through the domain blur strategy, and the model is trained according to the given loss function and optimization algorithm.
[0075] The target loss function of the model includes the alternating optimization training target losses for the feature extractor F, the label classifier C, and the domain classifier D.
[0076] The optimization algorithm uses the Stochastic Gradient Descent (SGD) algorithm, with a learning rate of 0.001 and a momentum of 0.9. After 200 iterations, the loss of the model's target function tends to balance, and the model training ends.
[0077] In step S8, online fault diagnosis is performed. The target domain test data set is input into the trained feature extractor F and label classifier C. As Figure 3 shown, the health status category of the online identified samples is determined.
[0078] Figure 4It is the visual clustering result of the high-level implicit features of the source domain and the target domain extracted by the fault diagnosis model. Among them, the source domain B has class labels, while the source domains C and D do not have class labels. The data under the working condition A is the target domain. It can be seen that the method of the present invention can effectively align the samples of the same class in multiple source domains, and make the boundaries between the features of different class samples more obvious, proving that the method of the present invention can learn distribution-independent representations, improve the intra-class aggregation and inter-class separability of features, so as to extract domain-invariant and discriminative features. The confusion matrix of the prediction results of the method of the present invention for the target domain is as Figure 5 shown. It can be seen that the diagnostic accuracy rate of the method of the present invention is very high, reaching 100%, and the samples of the four health states have not been misidentified, demonstrating the excellent generalization performance of the method of the present invention.
[0079] In summary, by establishing a domain fuzziness strategy to eliminate the distribution differences between multiple source domains, learning distribution-independent representations, and optimizing the class boundaries by designing a class-level optimization module, the model can extract domain-invariant and discriminative features with generalization ability, realizing accurate diagnosis of bearing faults under invisible working conditions.
[0080] Embodiment 2
[0081] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method in Embodiment 1 are implemented.
[0082] Embodiment 3
[0083] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method in Embodiment 1 are implemented.
[0084] Embodiment 4
[0085] The present invention also provides a processor, which is used to run a program. When the program runs, the method described in Embodiment 1 is executed.
[0086] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0087] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0088] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0089] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0090] Obviously, the above embodiments are only examples for clear illustration and are not limitations on the implementation. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. A fault diagnosis method based on a semi-supervised adversarial domain generalization intelligent model, characterized in that, It includes the following steps: S1. Intercept the collected mechanical vibration time-domain signals into data samples, unify the sample length, normalize the sample amplitudes, and divide the data set into a multi-source domain data set and a target domain data set; S2. Construct a feature extractor, which is used to map the preprocessed data samples to a target feature space. The feature extractor takes the data samples as inputs and the high-level hidden features of the data samples as outputs; S3. Construct a label classifier, which is used to predict the class labels of the high-level hidden features. The label classifier takes the output of the feature extractor as an input and the predicted sample class labels as outputs; S4. Construct a domain classifier, which is used to predict the domain labels of the high-level hidden features. The domain classifier takes the output of the feature extractor as an input and the predicted domain labels as outputs; S5. Construct a class-level optimization module, which is used to perform metric learning on the extracted high-level hidden features to optimize the class boundaries; S6. Combine the feature extractor, the label classifier, the domain classifier, and the class-level optimization module to construct a fault diagnosis training model; S7. Input the multi-source domain training data set into the constructed fault diagnosis training model, extract domain-invariant data features through a domain blur strategy, and perform model training according to a given loss function and optimization algorithm; S7 includes: S71. Extract high-level hidden features from the multi-source domain samples through the feature extractor, and input the features into the label classifier, the domain classifier, and the class-level optimization module; S72. For all the data from multiple source domains, minimize the cross-entropy loss of the domain classifier to optimize the domain classifier; S73. For the labeled data from the source domain, minimize the cross-entropy loss of the label classifier to optimize the label classifier; S74. Minimize the cross-entropy loss of the label classifier and the center loss of the class-level optimization module, and at the same time maximize the cross-entropy loss of the domain classifier, the Shannon entropy loss of the sample domain label prediction probability, and the Shannon entropy loss of the mean of the mini-batch sample domain label prediction probabilities to optimize the feature extractor; The training objective of optimizing the domain classifier is to try to identify the input features as the correct domain labels, while the training objective of optimizing the feature extractor is to blur the domain classifier so that it cannot correctly determine which domain the features come from and outputs an uncertain domain label, thus forming a min-max adversarial relationship; S75. Stop training when the adversarial training makes the model reach the Nash equilibrium; S8. Input the target domain test data set into the trained feature extractor and label classifier to online identify the health state categories of the samples.
2. The fault diagnosis method based on a semi-supervised adversarial domain generalization intelligent model according to claim 1, characterized in that, The domain blur strategy adopts the multi-class probability output of the domain classifier, and the output probability dimension is the same as the number of source domains. Through the continuous adversarial training of the feature extractor and the domain classifier, the feature extractor can blur the domain classifier, that is, it cannot determine which domain the extracted high-level hidden features come from and considers the probabilities of coming from each domain to be the same.
3. The fault diagnosis method based on a semi-supervised adversarial domain generalization intelligent model according to claim 1, characterized in that, In step S1, the dataset is divided according to the working conditions of the machine. Data samples under the same working condition are placed in the datasets of the same domain, and different datasets contain the same categories of mechanical health states. The multi-source domain dataset contains one labeled source domain and multiple unlabeled source domains, which are used for model training. The target domain dataset does not participate in model training and is only used to test the accuracy of the model prediction results.
4. The fault diagnosis method based on a semi-supervised adversarial domain generalization intelligent model according to claim 1, characterized in that, The optimization algorithm is the root mean square propagation algorithm, the stochastic gradient descent method, or the adaptive moment estimation algorithm.
5. The fault diagnosis method based on a semi-supervised adversarial domain generalization intelligent model according to claim 1, characterized in that, The class-level optimization module adopts a metric learning method. By learning a class center, it constrains the features belonging to this class to approach the class center as much as possible. For unlabeled source domain samples, the class center is updated through the predicted labels given by the label classifier.
6. The fault diagnosis method based on a semi-supervised adversarial domain generalization intelligent model according to claim 1, characterized in that, The domain classifier consists of a fully connected layer and a Softmax classifier. The domain classification loss of the domain classifier is the cross-entropy loss of the predicted domain labels of the samples in the multi-source domain.
7. The fault diagnosis method based on a semi-supervised adversarial domain generalization intelligent model according to claim 1, characterized in that, The label classifier consists of a fully connected layer and a Softmax classifier. The label classification loss of the label classifier is the cross-entropy loss of the predicted class labels of the samples in the labeled source domain.
8. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Deep transfer learning intelligent fault diagnosis method and device, storage medium and equipment
CN111898095A
Intelligent fault diagnosis method based on deep adversarial domain self-adaption
CN111898634A
Cited By
A quality detection method based on a shared query attention mechanism
CN122509754A