A Multi-Classifier Equipment Fault Diagnosis Method Based on DMSTFA

By using the multi-classifier method of DMSTFA in equipment fault diagnosis, dynamically aligning the characteristics of the source domain and the target domain, the problem of insufficient generalization ability of equipment fault diagnosis in different scenarios is solved, and high-accuracy fault diagnosis under different operating conditions is achieved.

CN114970702BActive Publication Date: 2025-06-24HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210540631.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2025-06-24
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

Existing equipment fault diagnosis methods have weak generalization capabilities under different scenarios or conditions, especially when the data distribution of the source and target domains is large, it is difficult to achieve accurate fault diagnosis.

Method used

The fault diagnosis method of multi-classifier equipment based on DMSTFA is adopted, and the domain invariant features and class distinction characteristics are extracted through dynamic multi-subspace transferable feature alignment, and the alignment adaptive factor and similarity weight are used to optimize the feature extraction and classifier combination to improve the diagnostic accuracy of the model on the target domain.

Benefits of technology

It significantly improves the diagnostic accuracy of the equipment fault diagnosis model in the target domain, enhances the generalization ability of the model, and can realize intelligent fault diagnosis under different working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114970702B_ABST
    Figure CN114970702B_ABST
Patent Text Reader

Abstract

The present invention proposes a multi-classifier device fault diagnosis method based on DMSTFA. First, the data is divided into source domain and target domain data according to different working conditions; then a device fault diagnosis neural network model is constructed and its parameters are initialized; the source domain and target domain data are input into the neural network, and the domain adaptation loss of each feature subspace is calculated according to DMSTFA; the feature similarity weights of each subspace are calculated; according to the similarity weights, the outputs of multiple classifiers are combined and the classification loss is calculated; the domain adaptation loss and the classification loss on each feature space are added together to obtain the total loss function value, and then the model parameters are iteratively trained and updated to obtain the final model; when diagnosing device faults, the target domain data is input into the final model to obtain the device fault diagnosis result. The present invention enhances the generalization ability of the model on the target domain with large distribution differences and significantly improves the diagnosis accuracy of the device fault diagnosis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for diagnosing equipment faults, in particular to a method for diagnosing equipment faults based on a multi-classifier using DMSTFA. Background Art

[0002] Fault diagnosis plays an important role in studying the relationship between monitoring data and the health state of machines. Accurate fault diagnosis has always been at the forefront of research for engineers and scientists. Traditional model-based methods rely only on thresholds of different signals at fault frequencies to determine the presence of faults. These models can only describe some fault types with clear signal characteristics. In reality, faults are often more complex. For example, in the early stage of a fault, the characteristics may not be very clear; multiple faults may occur simultaneously, which may modify the fault characteristics and create new characteristics due to coupling effects. Therefore, the data itself may have many unique features or patterns, and it is almost impossible for humans to identify these complex features through manual observation or interpretation. Machine learning-based algorithms, including artificial neural networks (ANN), principal component analysis (PCA), support vector machines (SVM), etc., can analyze data more comprehensively and make intelligent decisions about the presence of faults by applying the learned knowledge. To obtain better performance under various operating conditions and noisy environments, deep learning-based methods are becoming increasingly popular, such as autoencoders (AE), deep belief networks (DBN), convolutional neural networks (CNN), etc. Deep learning helps to automatically learn fault characteristics from the collected data. A large number of experiments have shown that deep learning-based methods have higher generalization performance than machine learning-based methods. However, both machine learning- and deep learning-based equipment fault diagnosis methods are based on a necessary assumption: the samples of the training data set from the laboratory (source domain) should have the same distribution as the samples of the test data set from the engineering scenario (target domain). However, machine equipment in different scenarios usually operates under different working conditions, such as load, speed, and temperature. This means that under different working conditions, the distribution of data may be significantly different. Then, when the learning model based on the training data set is deployed to the test data set in practical applications, its generalization ability is weak, which further affects the decision-making of the model on the test set.

[0003] This problem can be solved by the idea of reusing diagnostic knowledge in different scenarios or conditions. For example, the diagnostic knowledge of a laboratory equipment dataset may help identify the health status of equipment in an engineering scenario, which is exactly what can be achieved by the domain adaptation problem of transfer learning. As a subfield of transfer learning, unsupervised domain adaptation can alleviate the inconsistent distribution between labeled training samples and unlabeled test samples by learning the shared features of two domains. Its core idea is that since the data in the source domain and the target domain are different, then map the data into a feature space and find a metric criterion in the feature space to make the feature distributions of the source domain data and the target domain data close. Then the classifier trained based on the source domain features can be used on the target domain. For fault diagnosis, due to the inherent similarity of fault features under different application scenarios or working conditions, the shared features existing in these two domains allow transfer learning to have a certain degree of feasibility and practical significance.

[0004] In summary, finding an effective domain adaptation-based device fault diagnosis method to enhance the generalization ability of the training model in the source domain on the target domain, especially when there is a large distribution difference between the two, has become an urgent problem to be solved currently. Summary of the Invention

[0005] To solve the above technical problems, the present invention proposes a multi-classifier device fault diagnosis method based on DMSTFA, which significantly improves the diagnostic accuracy of the device fault diagnosis model on the target domain, and can achieve intelligent diagnosis of device faults even when the working conditions are quite different.

[0006] The technical solution of the present invention includes the following steps:

[0007] Step 1: Divide the data into source domain and target domain data according to different working conditions;

[0008] Step 2: Construct a device fault diagnosis neural network model and initialize its parameters;

[0009] The described equipment fault diagnosis neural network consists of three parts: a feature extractor, three classifiers, and a domain discriminator; specifically, the feature extractor contains a backbone network composed of multiple convolutional layers, and three feature subspace networks composed of single convolutional layers and pooling layers; the three classifiers have the same structure, and each classifier consists of a fully connected hidden layer and a Softmax layer; the domain discriminator consists of a fully connected hidden layer and a binary classifier with logistic regression; the backbone network is used to extract low-dimensional signal features of equipment data, the feature subspace network is used to map the low-dimensional features to different spaces and extract high-dimensional features from multiple perspectives, and there are batch normalization and Relu layers after each convolutional layer; the pooling layer selects the average pooling strategy for feature downsampling, is placed after the convolutional layer of each feature subspace, and is connected to the fully connected hidden layer, and finally outputs the probability value on each category through the Softmax layer;

[0010] Step 3: Input the source domain and target domain data into the neural network, and calculate the domain adaptation loss of each feature subspace according to DMSTFA;

[0011] The described DMSTFA is fully named Dynamic Multiple Sub-structure Transferable Features Alignment, that is, dynamic multi-subspace transferable feature alignment. This method designs a domain adaptation loss to extract transferable features on multiple subspaces; L transfer represents the alignment loss of transferable features between the source domain and the target domain on the subspace. The total transferable feature alignment loss consists of the inter-domain feature difference loss and the class discrimination loss between the source domain and the target domain on each subspace. This part is called the domain adaptation part, L transfer The calculation method is as follows:

[0012]

[0013] where n represents the number of feature subspaces, and j represents the j-th subspace, represents the transferable feature alignment loss of the source domain and the target domain on the jth subspace, where the transferable features are divided into two parts, one is the domain invariant features of the source domain and the target domain, and the other is the category discrimination features of the source domain and the target domain. Therefore, the transferable feature alignment loss is composed of two parts accordingly. One is the inter-domain feature difference loss, which is used to measure the inter-domain feature difference between the source domain and the target domain, thereby guiding the feature extractor to extract the domain invariant features of the source domain and the target domain, and explicitly aligning the transferable features of the source domain and the target domain; the other is the category discrimination loss, which is used to measure whether the features of the two domains are discriminative in a certain category. In other words, the loss is used to measure whether the extracted features help the classifier to make accurate predictions, thereby guiding the feature extractor to extract the category discrimination features of the two domains. In particular, since the categories of the source domain and the target domain are consistent, extracting the category discrimination features of the two helps to implicitly align the transferable features of the source domain and the target domain. The formula is as follows:

[0014]

[0015] In the formula, represents the inter-domain feature difference loss between the source domain and the target domain on the j-th subspace, represents the category discrimination loss of the source domain and the target domain in the j-th subspace, μ represents the alignment adaptation factor, which dynamically adjusts the importance of the alignment domain invariant features and the category discrimination features in the feature alignment process; the three important components in the formula are further explained below.

[0016] The specific calculation is as follows:

[0017]

[0018] Where g(·) represents a convolutional neural network, which is used to extract low-dimensional features of the original input data, h(·) maps the low-dimensional features to different feature subspaces, and d j Represents the MK-MMD distance between the source domain and the target domain in the jth subspace; previous methods based on feature difference loss often map the source domain and the target domain to a low-dimensional feature space for feature alignment. This single feature extraction method may miss some important information. The DMSTFA method proposed in the present invention can align the distribution of the source domain and the target domain in multiple feature subspaces, and learn multiple domain-invariant features by minimizing the distribution differences of multiple feature subspaces.

[0019] The specific calculation is as follows:

[0020]

[0021] In the formula Denotes the probability of the classifier's predicted output for the source domain or target domain data, C s,t Denotes the number of classes in two domains, and c denotes a predicted output class.

[0022] μ = 1 - 2|d(h(g(X s,t )) - 0.5|(5)

[0023] In the formula, d(·) denotes the discrimination result of the domain discriminator for the high-dimensional features of the source domain or target domain subspace, which discriminates whether the input data belongs to the source domain or the target domain. In the last layer of the discriminator in the network, the sigmoid activation function is used to make the discrimination result a floating-point number between [0, 1]. When the discrimination result is close to 1, it is considered that the input data belongs to the source domain. When the discrimination result is close to 0, it is considered that the input data belongs to the target domain. When the discrimination result is close to 0.5, it is considered that the input features represent the domain-invariant features of the source domain and the target domain, because at this time the domain discriminator cannot distinguish whether the input data belongs to the source domain or the target domain. Therefore, μ ∈ [0, 1]. When μ → 0, it means that no domain-invariant features are extracted, and at this time, more attention should be paid to the domain difference loss. When μ → 1, it means that domain-invariant features have been extracted, and at this time, more attention should be paid to the class discrimination loss. Therefore, the alignment adaptation factor μ dynamically adjusts the importance of the two losses according to the actual extracted feature situation, guides the feature extractor to extract more transferable features, and achieves a good feature alignment effect.

[0024] Step Four: Calculate the feature similarity weights for each subspace;

[0025] Step Five: Combine the outputs of multiple classifiers according to the similarity weights and calculate the classification loss;

[0026] Step Six: Add the domain adaptation loss and the classification loss on each feature space to obtain the total loss function value, and then perform iterative training to update the model parameters to obtain the final model;

[0027] Step Seven: When diagnosing equipment faults, input the target domain data into the final model to obtain the equipment fault diagnosis result.

[0028] Preferably, the specific content of Step One is: Divide the data collected by the sensor into labeled source domain data and unlabeled target domain data where S represents the source domain, T represents the target domain, and i represents the i-th sample in the source domain or target domain. NS Denotes the number of source domain samples, N T Denotes the number of target domain samples; the working conditions of the source domain and target domain data are different, that is, the source domain and target domain data have different distributions.

[0029] Preferably, in order to utilize the information on the feature differences between the source domain and the target domain, the domain adaptation part provides a similarity weight reflecting the similarity between the source domain and the target domain on multiple feature subspaces to re-weight multiple classifiers. The similarity weight is calculated as follows:

[0030]

[0031] where w j represents the similarity weight between the source domain and the target domain on the j-th subspace, represents the loss of the inter-domain feature difference between the source domain and the target domain on the j-th subspace, n represents the number of subspaces. If the two domains on a certain subspace are similar to each other, it is reflected in a lower loss of the inter-domain feature difference and a higher weight w j .

[0032] Preferably, the parameters of the device fault diagnosis neural network in step two are initialized by a normal distribution random initialization method, and its parameters are updated by the Adam algorithm.

[0033] Preferably, the output of the combined multi-classifier is specifically as follows: The multi-classifier strategy separately establishes source classifiers on each feature subspace, uses the weights of the domain adaptation part to evaluate the similarity between the source and the target on each subspace, and weights each classifier to learn the target fault classifier. The output of the c-th class of the classifier is as follows:

[0034]

[0035] where v i is the i-th weight connecting the input to the i-th output neuron, C s represents the number of source domain categories, w j represents the similarity weight between the source domain and the target domain on the j-th subspace.

[0036] Preferably, the MK-MMD distance d j between the source domain and the target domain on the j-th subspace is calculated as follows:

[0037]

[0038] Preferably, in step five, the classification loss is calculated as follows:

[0039]

[0040] where S represents the source domain, i represents the i-th sample in the source domain, N S represents the number of source domain samples, J uses the cross-entropy loss function, represents the data of the i-th sample in the source domain, The data label of the i-th sample in the source domain, represents the predicted output of the combined classifier for the i-th sample.

[0041] Preferably, adding the domain adaptation loss and the classification loss on each feature space to obtain the total loss function value, specifically:

[0042]

[0043] In the formula, S represents the source domain, i represents the i-th sample in the source domain, N S represents the number of samples in the source domain, J adopts the cross-entropy loss function, represents the data of the i-th sample in the source domain, The data label of the i-th sample in the source domain, represents the predicted output of the combined classifier for the i-th sample.

[0044] The beneficial effects of the present invention are as follows: The present invention proposes a multi-classifier method for DMSTFA under working conditions with large differences. The transferable features are divided into domain-invariant features and class discrimination features. The designed domain adaptation loss is used to measure the deviation of the inter-domain feature distribution between the source domain and the target domain on multiple feature spaces and the deviation of the inter-class feature distribution between the source domain and the target domain on multiple classifiers. The alignment adaptation factor is used to dynamically guide the feature extraction network to allocate the importance of extracting domain-invariant features and class discrimination features, so that the features of the source domain and the target domain have a better alignment effect; the similarity between the source domain and the target domain is evaluated from multiple angles, and the designed similarity weight is used to re-weight the results of the multi-classifier to learn a more accurate target fault classifier, making full use of the information learned on each feature subspace and focusing on the importance of the features extracted from the space corresponding to the maximum similarity; the domain adaptation losses on multiple feature subspaces are statistically calculated, and the classification loss is calculated according to the similarity weight by combining the outputs of the classifiers corresponding to each subspace. The two are added to obtain the total loss function value, and the model parameters are updated and optimized through backpropagation iteration. Based on the powerful learning ability of the deep learning model, the feature representations of the source domain and the target domain are learned. Further, the multi-classifier method of DMSTFA measures the deviation of the transferable feature distribution between the source domain and the target domain from multiple angles and combines the results of the multi-classifier, enhancing the generalization ability of the model on the target domain with large distribution differences, significantly improving the diagnostic accuracy of the equipment fault diagnosis model, and solving the problem that the domain adaptation algorithm on a single space has poor generalization ability for the target domain with large differences, thereby realizing the intelligent diagnosis of equipment faults. Description of the Drawings

[0045] Figure 1 is the flow chart of the method of the present invention;

[0046] Figure 2Overall structure diagram of the neural network of the present invention;

[0047] Figure 3 Overall flow chart of the DMSTFA multi-classifier algorithm of the present invention; Specific implementation manners

[0048] The present invention will be further described in detail below with reference to the accompanying drawings.

[0049] As Figure 1 shown, the present invention includes the following steps.

[0050] 1) Divide the data into source domain and target domain data according to different working conditions;

[0051] 2) Construct a device fault diagnosis neural network model and randomly initialize its parameters;

[0052] 3) Clear the parameter gradients of the neural network model and perform one round of iterative training;

[0053] 4) Input the source domain and target domain data into the neural network, and calculate the domain adaptation loss of each feature subspace according to DMSTFA;

[0054] 5) Calculate the feature similarity weights of each subspace;

[0055] 6) Combine the outputs of multiple classifiers according to the similarity weights and calculate the classification loss;

[0056] 7) Add the domain adaptation loss and classification loss on each feature space to obtain the total loss function value;

[0057] 8) Perform backpropagation according to the total loss value, and the optimization method adopts the Adam algorithm;

[0058] 9) Iteratively train and update the model parameters, judge whether the iteration times are reached, and save the final model;

[0059] 10) When diagnosing the device fault, input the target domain data into the final model to obtain the device fault diagnosis result.

[0060] In step 1), the data collected by the sensor is divided into labeled source domain data and unlabeled target domain data where S represents the source domain, T represents the target domain, i represents the i-th sample in the source domain or target domain, N S represents the number of source domain samples, N T represents the number of target domain samples; the working condition conditions of the source domain and target domain data are different, that is, the source domain and target domain data have different distributions.

[0061] The overall structure schematic diagram of the device fault diagnosis model in step 2) is asFigure 2 As shown in the figure, the specific construction steps are as follows:

[0062] The device fault diagnosis neural network consists of three parts: a feature extractor, three classifiers, and a domain discriminator. Specifically, the feature extractor contains a backbone network composed of multiple convolutional layers, and three feature subspace networks composed of single convolutional layers and pooling layers; the three classifiers have the same structure, and each classifier consists of a fully connected hidden layer and a Softmax layer; the domain discriminator consists of a fully connected hidden layer and a binary classifier with logistic regression. The backbone network is used to extract the low-dimensional signal features of the device data, and the feature subspace network is used to map the low-dimensional features to different spaces and extract high-dimensional features from multiple perspectives. After each convolutional layer, there are batch normalization and Relu layers; the pooling layer selects the average pooling strategy for feature downsampling, which is placed after the convolutional layer of each feature subspace and connected to the fully connected hidden layer, and finally the probability value of each category is output through the Softmax layer. The parameters of the device fault diagnosis neural network are initialized by the normal distribution random initialization method, and its parameters are updated by the Adam algorithm.

[0063] The domain adaptation loss of each feature subspace in step 4) consists of two parts, the inter-domain feature difference loss and the class discrimination loss, and the importance of the two parts of the loss is dynamically adjusted by an alignment adaptation factor. The inter-domain feature difference loss is calculated after the three pooling layers. After the backbone network extracts the low-dimensional features of the source domain and target domain data, they are input into the three subspaces to further extract high-dimensional feature representations from multiple perspectives. The smaller the inter-domain feature difference loss, the higher the inter-domain similarity features between the source domain and the target domain in this space; the class discrimination loss is calculated after the three Softmax layers. The Softmax layer outputs the predicted probability value of each category. The smaller the class discrimination loss, the higher the class discrimination of the features extracted in this space; the alignment adaptation factor is calculated after the output of the domain discriminator, and 0.5 is used as the critical value to measure whether the extracted features already have the optimal domain-invariant features, so as to guide the feature extraction network to allocate the importance of extracting domain-invariant features and class discrimination features. On the one hand, the network optimizes the domain adaptation loss to align the features of the source domain and the target domain in multiple subspaces; on the other hand, the domain adaptation loss provides the similarity weights of the two domains for the subsequent multi-classifiers. The feature difference loss is measured by the MMD (Maximum Mean Discrepancy) distance. MMD can weight the distance between two distributions in the reproducing kernel Hilbert space (RKHS). The MMD distance calculation formula is as follows:

[0064]

[0065] Among which H k The RKHS is represented by using the kernel k. Generally, the Gaussian kernel is used as the kernel. φ(*) represents the mapping to the RKHS, tr represents the trace of the matrix, M is the coefficient matrix, and M is calculated as follows:

[0066]

[0067] The idea of MMD is based on the samples in the source domain and the target domain. By finding the mapping function φ in the sample space, and then calculating the mean of the function values of the samples with different distributions on φ, the mean distance between the data distributions of the source domain and the target domain corresponding to φ can be obtained by taking the difference between the two means. As a test statistic, MMD can judge whether two distributions are the same. If the MMD distance is small enough, it can be considered that the two distributions are the same, otherwise it is considered that they have a large deviation. However, for practical applications, the selection of parameters for each kernel is crucial for the final performance of the feature mapping. To better select the Gaussian kernel parameters, MK-MMD (Multi-Kernel Maximum Mean Discrepancy) is used instead of MMD. MK-MMD is an effective estimator that provides a mapping using a convex combination of m kernels. The calculation method for kernel selection is as follows:

[0068]

[0069] In the formula, u represents the u-th kernel, β u represents the weighted parameter of the u-th kernel, and d represents the number of kernels. Using MK-MMD, the mean distance value of the feature mapping between the source domain and a single target domain can be calculated, and this distance is used as the value of the loss function of the target domain in a single space. The calculation method is as follows:

[0070]

[0071] The model is iteratively optimized through MK-MMD, so that the model has the domain-invariant representation ability of the source domain and the target domain.

[0072] In step 4), DMSTFA is the abbreviation of Dynamic Multiple Sub-structure Transferable Features Alignment, that is, Dynamic Multiple Subspace Transferable Feature Alignment. This method designs a domain adaptation loss to extract transferable features on multiple subspaces; in addition, a weight is designed according to the feature difference loss of each subspace to represent the similarity between the source domain and the target domain, so as to adaptively combine the outputs of multiple classifiers; device

[0073] The total loss function value of the fault diagnosis neural network is calculated as follows: L = L C + λL transfer (5)

[0074] Among them, L C represents the classification loss of the output of the source multi-classifier combination. L transfer represents the loss of aligning the transferable features of the source domain and the target domain on the subspace. λ represents the weight parameter for weighing the two parts. The total transferable feature alignment loss is composed of the transferable feature alignment losses of the source domain and the target domain on each subspace. L transfer The calculation method is as follows:

[0075]

[0076] Among them, n represents the number of feature subspaces, j represents the j-th subspace, and L j transfer represents the transferable feature alignment loss of the source domain and the target domain on the j-th subspace. This method believes that the transferable features are divided into two parts. One is the domain-invariant features of the source domain and the target domain, and the other is the class discriminative features of the two domains. Therefore, the transferable feature alignment loss consists of two parts correspondingly. One is the inter-domain feature difference loss, which is used to measure the feature difference between the source domain and the target domain, so as to guide the feature extractor to extract the domain-invariant features of the source domain and the target domain. The other is the class discriminative loss, which is used to measure whether the features of the two domains are discriminative in a certain class, so as to guide the feature extractor to extract the class discriminative features of the two domains. The formula is as follows:

[0077]

[0078] In the formula, L j mmd represents the feature difference loss of the source domain and the target domain on the j-th subspace, and L j dis represents the class discriminative loss of the source domain and the target domain on the j-th subspace. μ represents the alignment adaptation factor, which dynamically adjusts the importance of aligning the domain-invariant features and the class discriminative features in the feature alignment process. The following further explains the three important components in the formula.

[0079] i. L j mmd The specific calculation is as follows:

[0080]

[0081] In the formula, g(·) represents the convolutional neural network, which is used to extract the low-dimensional features of the original input data. h(·) maps the low-dimensional features to different feature subspaces, and d jDenote the MK-MMD distance between the source domain and the target domain on the j-th subspace. Previous methods based on feature difference loss often map the source domain and the target domain to a low-dimensional feature space for feature alignment. This single feature extraction method may miss some important information. The DMSTFA method proposed by the present invention can align the distributions of the source domain and the target domain on multiple feature subspaces, and learn multiple domain-invariant features by minimizing the distribution differences of multiple feature subspaces.

[0082] ii.L j dis The specific calculation is as follows:

[0083] In the formula Denote the probability of the classifier's prediction output for the source domain or the target domain data, which is provided by the Softmax layer in the network, and C s,t Denote the number of classes of the two domains, and c denotes a certain predicted output class.

[0084] iii.μ = 1 - 2|d(h(g(X s,t )) - 0.5|(10)

[0085] In the formula, d(·) denotes the discrimination result of the domain discriminator for the high-dimensional features of the source domain or the target domain subspace, which discriminates whether the input belongs to the source domain or the target domain. In the last layer of the discriminator in the network, a sigmoid activation function is used to make the discrimination result a floating-point number between [0, 1]. When the discrimination result is close to 1, it is considered that the input data belongs to the source domain. When the discrimination result is close to 0, it is considered that the input data belongs to the target domain. When the discrimination result is close to 0.5, it is considered that the input features can represent the domain-invariant features of the source domain and the target domain. Therefore, μ ∈ [0, 1]. When μ → 0, it means that no domain-invariant features have been extracted, and at this time, more attention should be paid to the domain-invariant feature loss. When μ → 1, it means that domain-invariant features have been extracted, and at this time, more attention should be paid to the class discrimination loss. The adaptive factor μ can dynamically adjust the importance of the two losses according to the actual extracted feature situation and achieve a good feature alignment effect.

[0086] Furthermore, in order to utilize the information of the feature differences between the source domain and the target domain, the adaptive part can also feedback a similarity weight reflecting the similarity between the source domain and the target domain on multiple feature subspaces to re-weight the learning classifier. The similarity weight is calculated as follows:

[0087]

[0088] In the formula, w j Denote the similarity weight between the source domain and the target domain on the j-th subspace, and L j mmdIt represents the feature difference loss between the source domain and the target domain on the j-th subspace. n represents the number of subspaces. If the two domains on a certain subspace are similar to each other, it is reflected in a lower feature difference loss L j mmd and a higher weight w j .

[0089] Furthermore, the multi-classifier strategy separately establishes source classifiers on each feature subspace, uses the weights of the adaptive part to evaluate the similarity between the source and the target on each space, and weights each classifier to learn the target fault classifier. The output of the c-th class of the classifier is as follows:

[0090]

[0091] where v i is the i-th weight connecting the input to the i-th output neuron, and C s represents the number of source domain categories. Generally speaking, this method measures on multiple feature spaces, focuses on improving the learning ability of the space with the smallest feature difference, and at the same time, the multi-source fault classifier is re-weighted by the similarity scores provided by the improved adaptive part to obtain more accurate classification results. The total loss function value is calculated as follows:

[0092]

[0093] To more specifically and intuitively illustrate the multi-classifier strategy of DMSTFA, the overall calculation process is as Figure 3 shown, and the calculation steps are as follows:

[0094] Step1: Calculate the feature difference loss L between the source domain and the target domain on each feature subspace mmd ;

[0095] Step2: Calculate the class discrimination loss L of the classifier between the source domain and the target domain on each feature subspace dis ;

[0096] Step3: Calculate the alignment adaptive factor μ of the domain discriminator for the source domain features and the target domain features on each feature subspace;

[0097] Step4: Calculate the similarity weight w between the source domain and the target domain. This weight evaluates the similarity between the two domains on each feature space, so as to make full use of the information on each subspace to learn the classifier;

[0098] Step5: Re-weight the outputs of each classifier according to the similarity weight to obtain the output of the combined classifier, and calculate the classification loss L of the source domain C ;

[0099] Step 6: According to the alignment adaptation factor, add the cross-domain feature difference loss and the class discrimination loss of each feature subspace to obtain the total domain adaptation loss, and finally add the value of the classification loss function to obtain the total loss function value of the neural network;

[0100] Step 7: Iteratively optimize the parameters of the equipment fault diagnosis model according to the total loss function value, and finally obtain the final model.

[0101] Step 4) of the present invention proposes a multi-classifier method for DMSTFA under working conditions with large differences, divides the transferable features into domain-invariant features and class discrimination features, uses the designed domain adaptation loss to measure the cross-domain feature distribution deviation between the source domain and the target domain on multiple feature spaces and the inter-class feature distribution deviation between the source domain and the target domain on multiple classifiers, and dynamically guides the feature extraction network to allocate importance to the extraction of domain-invariant features and class discrimination features through the alignment adaptation factor, so that the features of the source domain and the target domain have a better alignment effect; evaluates the similarity between the source domain and the target domain from multiple perspectives, re-weights the results of the multi-classifier using the designed similarity weights to learn a more accurate target fault classifier, fully utilizes the information learned on each feature subspace and focuses on the importance of the features extracted from the space corresponding to the maximum similarity; statistically calculates the domain adaptation loss on multiple feature subspaces, and calculates the classification loss according to the similarity weights by combining the outputs of the classifiers corresponding to each subspace, and the sum of the two obtains the total loss function value, and iteratively updates and optimizes the model parameters through backpropagation. Based on the powerful learning ability of the deep learning model, the feature representations of the source domain and the target domain are learned. Further, the multi-classifier method of DMSTFA measures the transferable feature distribution deviation between the source domain and the target domain from multiple perspectives and combines the results of the multi-classifier, enhancing the generalization ability of the model on the target domain with large distribution differences, significantly improving the diagnostic accuracy of the equipment fault diagnosis model, and solving the problem that the domain adaptation algorithm on a single space has poor generalization ability for the target domain with large differences, thereby realizing the intelligent diagnosis of equipment faults.

Claims

1. A multi-classifier device fault diagnosis method based on DMSTFA, characterized in that: Step 1: Divide the data into source domain and target domain data according to different working conditions; Step 2: Construct a device fault diagnosis neural network model and initialize its parameters; The device fault diagnosis neural network consists of three parts: a feature extractor, three classifiers, and a domain discriminator; specifically, the feature extractor contains a backbone network composed of multiple layers of convolution, and three feature subspace networks composed of a single layer of convolution and a pooling layer; the three classifiers have the same structure, and each classifier consists of a fully connected hidden layer and a Softmax layer; the domain discriminator consists of a fully connected hidden layer and a binary classifier with logistic regression; the backbone network is used to extract the low-dimensional signal features of the device data, and the feature subspace network is used to map the low-dimensional features to different spaces and extract high-dimensional features from multiple angles. After each convolution layer, there are batch normalization and Relu layers; the pooling layer all selects the average pooling strategy for feature downsampling, which is placed after the convolution layer of each feature subspace and connected to the fully connected hidden layer, and finally the probability value on each category is output through the Softmax layer; Step 3: Input the source domain and target domain data into the neural network, and calculate the domain adaptation loss of each feature subspace according to DMSTFA, that is, Dynamic Multi-Subspace Transferable Feature Alignment; The DMSTFA designs a domain adaptation loss to extract transferable features on multiple subspaces; L transfer denotes the alignment loss of transferable features between the source domain and the target domain on the subspace. The total alignment loss of transferable features consists of the inter-domain feature difference loss and the class discrimination loss between the source domain and the target domain on each subspace. This part is called the domain adaptation part, L transfer The calculation method is as follows: where n represents the number of feature subspaces, and j represents the j-th subspace. denotes the transferable feature alignment loss between the source domain and the target domain on the j-th subspace. The transferable features are divided into two parts. One is the domain-invariant features of the source domain and the target domain, and the other is the class discriminative features of the source domain and the target domain. Therefore, the transferable feature alignment loss consists of two parts correspondingly. One is the inter-domain feature difference loss, which is used to measure the inter-domain feature difference between the source domain and the target domain, so as to guide the feature extractor to extract the domain-invariant features of the source domain and the target domain and explicitly align the transferable features of the source domain and the target domain. The other is the class discriminative loss, which is used to measure whether the features of the two domains are discriminative in a certain class. Since the classes of the source domain and the target domain are the same, extracting the class discriminative features of the two helps to implicitly align the transferable features of the source domain and the target domain. The formula is as follows: In the formula, represents the cross-domain feature difference loss between the source domain and the target domain on the j-th subspace, represents the class discrimination loss between the source domain and the target domain on the j-th subspace, μ represents the alignment adaptation factor, and dynamically adjusts the importance of the alignment domain-invariant features and the class discrimination features in the feature alignment process; the following further explains the three important components in the formula; The specific calculation is as follows: where \(g(\cdot)\) represents a convolutional neural network for extracting low-dimensional features of the original input data, and \(h\) j (\cdot)\) maps the low-dimensional features to the \(j\)-th feature subspace, and \(d\) j represents the MK-MMD distance between the source domain and the target domain on the \(j\)-th subspace, that is, the multi-kernel maximum mean discrepancy; The specific calculation is as follows: where represents the probability of the classifier's predicted output for the source domain or target domain data, and C s,t represents the number of classes in the two domains, and c represents a predicted output class; μ=1 - 2|d(h(g(X s,t )) - 0.5| (5) In the formula, h(·) maps the low-dimensional features to the feature subspace, and d(·) represents the domain discrimination result of the domain discriminator on the high-dimensional features of the source domain or target domain subspace, discriminating whether the input data belongs to the source domain or the target domain. In the last layer of the discriminator in the network, the sigmoid activation function is used to make the discrimination result a floating point number between [0,1]. When the discrimination result is close to 1, it is considered that the input data belongs to the source domain. When the discrimination result is close to 0, it is considered that the input data belongs to the target domain. When the discrimination result is close to 0.5, it is considered that the input feature represents the domain-invariant features of the source domain and the target domain. Because at this time, the domain discriminator can no longer distinguish whether the input data belongs to the source domain or the target domain, so μ∈[0,1]. When μ→0, it means that no domain-invariant features are extracted, and at this time, more attention should be paid to the domain difference loss between domains. When μ→1, it means that domain-invariant features have been extracted, and at this time, more attention should be paid to the category discrimination loss; therefore, the alignment adaptation factor μ dynamically adjusts the importance of the two losses according to the actual extracted feature situation, guides the feature extractor to extract more transferable features, and achieves a good feature alignment effect; Step 4: Calculate the feature similarity weight of each subspace; Step 5: Combine the outputs of multiple classifiers according to the similarity weight and calculate the classification loss; Step 6: Add the domain adaptation loss and the classification loss on each feature space to obtain the total loss function value, and then perform iterative training to update the model parameters to obtain the final model; Step 7: When diagnosing the device fault, input the target domain data into the final model to obtain the device fault diagnosis result.

2. A multi-classifier device fault diagnosis method based on DMSTFA according to claim 1; characterized in that: The specific content of the first step is as follows: according to the operating speed of the device, the data collected by the sensor is divided into labeled source domain data and unlabeled target domain data where S represents the source domain, T represents the target domain, i represents the i-th sample in the source domain or the target domain, and N S represents the number of source domain samples, and N T represents the number of target domain samples; the working conditions of the source domain and target domain data are different, that is, the source domain and target domain data have different distributions.

3. A multi-classifier device fault diagnosis method based on DMSTFA according to claim 1; characterized in that: The calculation of the similarity weight is as follows: where w j represents the similarity weight between the source domain and the target domain on the j-th subspace, represents the inter-domain feature difference loss between the source domain and the target domain on the j-th subspace, n represents the number of subspaces. If two domains on a certain subspace are similar to each other, it is reflected in a lower inter-domain feature difference loss and a higher weight w j .

4. A multi-classifier device fault diagnosis method based on DMSTFA according to claim 1, characterized in that: The neural network parameter initialization for equipment fault diagnosis in the second step adopts the normal distribution random initialization method, and its parameters are updated by the Adam algorithm.

5. A multi-classifier device fault diagnosis method based on DMSTFA according to claim 1, characterized in that: The output of the combined multi-classifier is specifically as follows: The multi-classifier strategy separately establishes source classifiers on each feature subspace, uses the weights of the domain adaptation part to evaluate the similarity between the source and the target on each subspace, weights each classifier to learn the target fault classifier, and the output of the c-th class of the classifier is as follows: where v i is the i-th weight connecting the input to the i-th output neuron, C s represents the number of source domain classes, w j represents the similarity weight between the source domain and the target domain on the j-th subspace.

6. A multi-classifier device fault diagnosis method based on DMSTFA according to claim 1, characterized in that: The MK-MMD distance d between the source domain and the target domain on the j-th subspace j is calculated as follows:

7. A multi-classifier device fault diagnosis method based on DMSTFA according to claim 1, characterized in that: In the fifth step, the classification loss is calculated as follows: where \(S\) represents the source domain, \(i\) represents the \(i\)-th sample in the source domain, and \(N\) S represents the number of samples in the source domain. \(J\) adopts the cross-entropy loss function, represents the data of the \(i\)-th sample in the source domain, represents the data label of the \(i\)-th sample in the source domain, represents the predicted output of the combined classifier for the \(i\)-th sample.

8. A multi-classifier device fault diagnosis method based on DMSTFA according to claim 1, characterized in that: The domain adaptation loss and classification loss L on each feature space C are added to obtain the total loss function value, specifically: Where S represents the source domain, i represents the i-th sample in the source domain, and N S represents the number of samples in the source domain. J uses the cross-entropy loss function, represents the data of the i-th sample in the source domain, represents the data label of the i-th sample in the source domain, represents the predicted output of the combined classifier for the i-th sample.

Citation Information

Patent Citations

  • Fault diagnosis method based on adaptive manifold embedding dynamic distribution alignment

    CN111829782A

  • Deep transfer learning intelligent fault diagnosis method and device, storage medium and equipment

    CN111898095A