Intelligent diagnosis method and system for mechanical equipment fault under open set
Through the Fourier and convolution embedded FCformer encoder and the open set dynamic subdomain adaptive module, combined with adaptive unknown class detection threshold learning and minimum-maximum entropy game training, unknown class interference problems in mechanical equipment fault diagnosis under the open set are solved, and accurate fault identification and detection are achieved.
Patent Information
- Application Number
- CN202510380073.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-28
AI Technical Summary
In the mechanical equipment fault diagnosis under open set, the prior art failed to effectively overcome the interference of unknown classes on feature alignment, and the unknown class detection threshold of traditional methods relies on prior knowledge, resulting in poor generalization and stability, making it difficult to apply to complex and changeable industrial scenarios.
The FCformer encoder with Fourier and convolutional embedding is used for time-frequency feature extraction, and an open set dynamic subdomain adaptive module and adaptive unknown class detection threshold learning method are designed, combined with the minimum-maximum entropy game training strategy to realize dynamic feature alignment and intelligent detection.
Accurate fault diagnosis under open sets is realized, the generalization ability and stability of the model is improved, and known fault modes can be identified and unknown fault modes can be intelligently detected.
Smart Images

Figure CN120336780A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mechanical equipment fault diagnosis, and particularly relates to an intelligent diagnosis method and system for mechanical equipment faults under an open set. Background Art
[0002] With the development of modern industrial technology, mechanical equipment is gradually developing towards precision, complexity, and intelligence, which also puts forward more stringent requirements for the safety and reliability of mechanical equipment. Rolling bearings in mechanical equipment may experience faults such as pitting and fracture during operation, resulting in serious economic losses and safety accidents. Therefore, carrying out health monitoring of mechanical equipment is of great significance for ensuring mechanical safety, reducing maintenance costs, and protecting personal safety.
[0003] In recent years, deep learning technology has become an effective solution in the field of mechanical equipment fault diagnosis due to its powerful non-linear feature extraction ability. The excellent performance of these methods depends on the independent and identical distribution of training data and test data. However, in actual industrial scenarios, mechanical equipment often operates under complex and variable working conditions, and there are obvious differences in the fault characteristic frequencies and amplitude ranges presented by the collected vibration signals. Directly applying the domain knowledge learned from training data to test data will lead to a decline in diagnostic performance. Therefore, domain adaptation technology, as a branch of transfer learning, can align the features from the source domain and the target domain, thereby transferring the knowledge of the source domain to the target domain. The success of domain adaptation technology is attributed to the same label space of the source domain and the target domain. However, in actual industrial scenarios, the working conditions of mechanical equipment show instability and uncertainty, and new and unknown fault types may occur at any time, resulting in the inapplicability of the fault diagnosis method based on domain adaptation. In this case, the label space of the target domain is larger than that of the source domain, which is the open set fault diagnosis scenario. Although scholars at home and abroad have carried out extensive research on mechanical equipment fault diagnosis under an open set, there are still the following two problems:
[0004] (1) Most open set fault diagnosis methods mainly focus on using domain adversarial training strategies to align the global feature distributions of the source domain and the target domain, while ignoring the conditional distribution deviation and discriminative features unique to the target domain samples. At the same time, due to the existence of unknown classes, traditional domain adaptation methods will misclassify unknown classes as known classes, and simply using LMMD to align the conditional distributions of the two domains will lead to feature negative transfer.
[0005] (2) In terms of unknown class detection, traditional open-set fault diagnosis methods usually adopt fixed thresholds based on prior expert knowledge as evaluation indicators for unknown classes. Their generalization and stability are poor, and it is difficult to apply them to actual industrial scenarios. In addition, some open-set fault diagnosis methods based on adaptive thresholds usually only use indicators such as confidence or distance metrics as detection thresholds, lacking the correlation information between samples. Summary of the Invention
[0006] The purpose of the present invention is to provide an intelligent diagnosis method and system for mechanical equipment faults under an open set, enabling it to overcome the interference of unknown classes to feature alignment and automatically generate more discriminative unknown class detection thresholds, realizing accurate fault diagnosis under different open set scenarios.
[0007] The technical solution to achieve the purpose of the present invention is as follows:
[0008] An intelligent diagnosis method for mechanical equipment faults under an open set, including:
[0009] S1. Use a linear embedding layer to serialize the one-dimensional vibration signal and convert it into a one-dimensional token embedding sequence that meets the requirements of the Transformer encoder;
[0010] S2. Construct an FCformer encoder with Fourier and convolutional embedding layers, and input the one-dimensional token embedding sequence into the FCformer encoder for time-frequency feature extraction;
[0011] S3. Use a global average pooling layer to obtain shared features, and combine the linear embedding layer and T FCformer encoders to form a shared feature extractor;
[0012] S4. Design a fault classifier based on a convolutional neural network to identify fault patterns for the shared features;
[0013] S5. Build an open-set dynamic subdomain adaptation module to capture the fine-grained information of each known class to achieve dynamic subdomain feature alignment under an open set;
[0014] S6. Design an adaptive unknown class detection threshold learning method to automatically update the unknown class detection threshold in each iteration cycle using a sample evaluation metric function. After the iterative training is completed, the final unknown class detection threshold is obtained;
[0015] S7. For the source domain sample data, design a source domain loss function including cross-entropy loss and center loss to discriminate fault features;
[0016] S8. For the target domain sample data, introduce entropy loss to improve the class discrimination ability of the target domain samples, and use the DLMMD loss function to align the dynamic conditional feature distributions;
[0017] S9. Design a min-max entropy game training strategy, and perform game training on the shared feature extractor and the fault classifier through the source domain loss function, the entropy loss function, and the DLMMD loss function;
[0018] S10. Use the evaluation metric function to generate the evaluation metrics of the target domain test samples, and compare them with the unknown class detection threshold to complete the intelligent detection of unknown classes and the cross-domain identification of fault modes in known classes.
[0019] An intelligent diagnosis system for mechanical equipment faults in an open set, comprising:
[0020] Data acquisition and dataset creation module: Collect the vibration signals of mechanical equipment in various health states under different working conditions, and perform sample division and normalization preprocessing on the vibration signal data; Randomly select some source domain data in healthy states as known class samples, while the target domain data includes all health states, that is, the target domain contains both known classes and unknown classes; Select samples from the source domain and the target domain according to a ratio to create a training dataset, and select a quantitative target domain data as a test dataset;
[0021] Model construction and training module: Construct a dynamic sub-domain adaptive network model architecture, including a shared feature extractor, a fault classifier, an open set dynamic sub-domain adaptive module, and an adaptive unknown class detection threshold learning method, and initialize the model structure parameters; Use the min-max entropy game strategy to train the model, update the parameters of the model by optimizing the objective, and complete the training of the dynamic sub-domain adaptive network model;
[0022] Model performance verification module: Use the test dataset to verify the diagnostic performance of the trained dynamic sub-domain adaptive network model, analyze the test results, and output the optimal model and the corresponding unknown class detection threshold;
[0023] Fault diagnosis module: In different open set domain adaptation fault diagnosis scenarios, use the optimal dynamic sub-domain adaptive network model to perform fault diagnosis on the real-time collected vibration signals, so as to achieve cross-domain identification of known fault modes and intelligent detection of unknown fault modes.
[0024] Compared with the prior art, the remarkable advantages of the present invention are:
[0025] 1. The present invention proposes a dynamic sub-domain adaptive network model for the intelligent diagnosis method of mechanical equipment faults in an open set, realizing cross-domain knowledge transfer of known classes and intelligent detection of unknown classes.
[0026] 2. The present invention constructs an improved Transformer encoder based on Fourier and convolutional embeddings, providing a lightweight Transformer encoder structure and a long-distance modeling scheme in the frequency domain.
[0027] 3. Considering the subdomain space differences and subdomain sample imbalances caused by unknown classes, the present invention proposes a novel open-set dynamic subdomain adaptation module based on DLMMD, which can finely align the subdomain features of known classes by dynamically assigning specific class weights.
[0028] 4. The present invention designs an adaptive threshold learning method using the correlation information between samples to automatically generate the unknown class detection threshold, which can reduce the dependence on prior knowledge when setting the threshold and improve the generalization ability of the model for different open-set diagnostic tasks.
[0029] 5. The present invention suppresses the problem that the model is overconfident in unknown classes due to blindly minimizing the entropy value through the min-max entropy game training strategy between the shared feature extractor and the fault classifier. Finally, the fault dataset of cylindrical roller bearings in offshore wind power lifting equipment is used to verify the stable and accurate fault diagnosis performance of the model proposed in the present invention in the open-set domain adaptation scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Specific implementation flowchart of an intelligent diagnosis method and system for mechanical equipment faults under an open set of the present invention.
[0031] Figure 2 Schematic diagram of the structure of the dynamic subdomain adaptation network model of the present invention.
[0032] Figure 3 Schematic diagram of the structure of the FCformer encoder of the present invention.
[0033] Figure 4 Schematic diagram of the structure of the open-set dynamic subdomain adaptation module of the present invention.
[0034] Figure 5 Appearance diagram of the cylindrical roller bearing test bench in this embodiment.
[0035] Figure 6 Comparison diagram of the average H-Score between the model proposed in the present invention and 4 open-set domain adaptation models in this embodiment.
[0036] Figure 7 Confusion matrix diagram of the model proposed in the present invention and different benchmark models in the first test of the T9 task in this embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0037] The specific embodiments of the present invention will be described below in conjunction with the accompanying drawings so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be ignored here.
[0038] Figure 1 It is a specific implementation flowchart of an intelligent diagnosis method and system for mechanical equipment faults under an open set of the present invention. As Figure 2 shown, the dynamic sub-domain self-adaptive network structure of the present invention mainly consists of a shared feature extractor G and a fault classifier C. Among them, the shared feature extractor mainly includes a linear embedding layer, T FCformer encoders, and a global average pooling layer. The fault classifier is used to identify the fault patterns of the shared features. The specific implementation steps of the present invention are as follows:
[0039] Step 1: Use the linear embedding layer to serialize the collected one-dimensional vibration signal so that it is converted into a one-dimensional token embedding sequence z0 that meets the input requirements of the traditional Transformer encoder. The specific steps are as follows:
[0040] Assume that the size of the sequence block is fixed at L, then the collected one-dimensional vibration signal is converted into a new sequence block: where W is the length of the input data, and Q = W / L is the number of sequence blocks. Then, use the linear embedding function ξ(·) to project the sequence block X into a token embedding sequence with a dimension of N, and the token embedding sequence for inputting the Transformer encoder can be obtained as:
[0041]
[0042] In the formula, x q represents the qth sequence block.
[0043] Step 2: Construct a new Transformer framework with a Fourier and Convolutional Embedding (FCE) layer - the FCformer encoder, and input the one-dimensional token embedding sequence z0 obtained in Step 1 into the FCformer encoder. As Figure 3 shown, the FCformer encoder uses a lightweight FCE layer to replace the multi-head self-attention layer of the traditional Transformer. The FCformer encoder includes an FCE layer, layer normalization (LN), a feed-forward layer, and layer normalization. The FCE layer includes a 1D-DFT layer, a one-dimensional convolutional layer, a batch normalization (BN) layer, and a rectified linear unit (ReLU) layer; the feed-forward layer includes a one-dimensional convolutional layer, a BN layer, a ReLU layer, a dropout layer, and a one-dimensional convolutional layer.
[0044] Specifically, for a given vibration signal sequence S n , where n ∈ [0,..., N - 1], the discrete Fourier transform (DFT) is processed as follows:
[0045]
[0046] In the formula, m is the frequency index, and F(m) is the frequency-domain feature representation, is the rotation factor, and j is the imaginary unit.
[0047] Furthermore, through Euler's formula, the discrete Fourier transform is split into a real part term and an imaginary part term, and we can get:
[0048] F(m) = F real (m) + jF imag (m) (3)
[0049] In the formula, F real (m) and F imag (m) are the real part and the imaginary part of F(m), respectively.
[0050] Utilizing the frequency feature learning ability of the 1D-DFT layer, an FCE layer is designed in the FCformer encoder to perform DFT transformation on the length dimension of the token embedding sequence. It should be noted that the FCE layer only retains the real part term after the 1D-DFT layer transformation. Therefore, for the input token embedding sequence z t-1 of the t-th FCformer encoder, the feature map z' t after being processed by the 1D-DFT layer is:
[0051]
[0052] In the formula, is the function that retains the real part after performing DFT transformation along the length dimension.
[0053] Subsequently, a one-dimensional convolutional layer is introduced into the FCformer encoder, and a BN layer and a RELU layer are connected, and the feature map z t ″ is:
[0054]
[0055] In the formula, σ(·) represents the processing function of the BN layer and the RELU layer, and f(·) is the one-dimensional convolutional operation.
[0056] Using residual connection, the feature maps z t ′ and z t-1 are superimposed to obtain the feature map z″′ tis:
[0057] z t ″′=LN(z t ′ t ′)+z t-1 (6)
[0058] where LN(·) is the layer normalization function.
[0059] Finally, a feed-forward layer is connected to provide a non-linear transformation, and the LN layer and residual connection are accessed. At this time, the output token embedding sequence z t of the t-th FCformer encoder can be expressed as:
[0060]
[0061] where FF(·) represents the processing function of the feed-forward layer.
[0062] Step 3: The token embedding sequence z T output after the time-frequency feature extraction by T FCformer encoders is input into the global average pooling layer (GAP), and the shared feature u1 can be obtained as:
[0063]
[0064] where R represents the kernel size of the GAP layer, and GAP(·) represents the GAP operation.
[0065] Combining the linear embedding layer described in Step 1 and the T FCformer encoders described in Step 2 together forms the shared feature extractor G.
[0066] Step 4: Design a fault classifier C based on a convolutional neural network to perform fault mode recognition on the shared feature u1 in Step 3. The fault classifier includes a convolutional extraction layer and two fully connected layers (FC). The convolutional extraction layer includes a one-dimensional convolutional layer, a BN layer, a ReLU layer, and a max pooling layer (MP). Flattening is performed between the convolutional extraction layer and the fully connected layer. Therefore, the output feature u2 of the first FC layer and the final output probability vector y of the second FC layer are respectively expressed as:
[0067]
[0068] where F C (·) represents the operation of the convolutional extraction layer, ψ(·) is the flattening function, f FC1 (·) and f FC2 (·) respectively represent the operations of the first and second FC layers. D is the output dimension of the first FC layer, and K is the number of known classes.
[0069] Step 5: Build an open-set dynamic subdomain adaptation module to achieve dynamic subdomain feature alignment under the open set by capturing the fine-grained information of each known class. The core idea of the open-set dynamic subdomain adaptation module is to dynamically align the subdomain distributions of known classes by masking potential unknown classes, which overcomes the risk of negative feature transfer in traditional LMMD methods due to the presence of unknown classes. Figure 4 The schematic diagram of the open-set dynamic subdomain adaptation module is shown, and the specific implementation steps of this module are described as follows:
[0070] 1. Quantitatively evaluate the probability that each target domain sample belongs to an unknown class to ensure that the model accurately suppresses unknown class alignment during the subdomain adaptation process. Since the probability of an unknown class sample belonging to any known class is low, its confidence sensitivity to uncertain predictions is low. On the contrary, the entropy value representing the uncertainty of the class distribution is high. Therefore, the open-set dynamic subdomain adaptation module in the present invention utilizes the complementary advantages of entropy value and confidence to construct a sample evaluation feature to comprehensively characterize each sample. For a sample with a probability vector y, its evaluation feature function fea(·) can be expressed as:
[0071] fea(y) = [ε1, 1 - ε2] (12)
[0072] In the formula, is the entropy value of the sample probability vector, where y(k) represents the k-th element in the probability vector y; ε2 = max(y(k)), k = 1, 2,..., K is the confidence of the sample probability vector (consistent with the number of known classes).
[0073] 2. Use the K-means algorithm to perform clustering analysis on the evaluation features of the target domain samples, so as to provide pseudo-labels of known classes and unknown classes for the subsequent training of the model. Since the number of clusters k' in this example is 2, that is, k' = 2, therefore, the present invention does not need to consider the influence of the setting of the number of clusters on the sample clustering performance. It should be noted that in the results of this clustering, the clustering samples closer to zero are known classes. Thus, the pseudo-label groups of known classes or unknown classes in the target domain samples can be expressed as:
[0074]
[0075] In the formula, R(·) is the marking operation, Kmeans(·) represents the K-means algorithm operation, n t is the number of target domain samples, represents the probability of the a-th sample in the target domain. represents the pseudo-label value of the a-th target domain sample. When the target domain sample is a known class; the target domain sample is an unknown class.
[0076] 3. According to the pseudo-label h of the target domain samples t , mask the unknown class samples in the probability vector of the a-th target domain sample, and the optimized probability vector obtained is as follows:
[0077]
[0078] 4. Considering that potential unknown class samples do not participate in feature alignment, it will lead to an imbalance in the sub-domain sample distribution between the source domain and the target domain. Traditional LMMD assigns the same weight to different sub-domains, resulting in inaccurate alignment of features within the sub-domain space. Therefore, in order to more precisely align the sub-domain features of known classes, a dynamic local maximum mean discrepancy (DLMMD) method is designed in the open set dynamic sub-domain adaptation module. This method can assign different weights to specific sub-domains according to the number of sub-domain samples, so as to more accurately align the sub-domain features of known classes. Therefore, for the source domain feature distribution P s and the target domain feature distribution P T , the DLMMD calculation is as follows:
[0079]
[0080] In the formula, is the DLMMD function, n s is the number of target domain samples, and respectively represent the weights of the i-th source domain sample and the a-th target domain sample belonging to class k, H is the reproducing kernel Hilbert space (RKHS), φ(·) represents the non-linear mapping function from the feature space to the RKHS, and ||·|| 2 represents the L2 norm. For the source domain sample , convert the true label into a one-hot vector to calculate the weight For the target domain samples without labels , use the optimized prediction vector as a one-hot vector to calculate the weight represents the number of samples belonging to class k in the source domain, represents the number of samples belonging to the unknown class in the target domain, represents the number of target domain samples with known classes and belonging to class k in the target domain, and its specific definition is as follows:
[0081]
[0082] In the formula, G(·) is the indicator function, that is: if the target domain sample If it is a known class and belongs to class k, return 1; otherwise, return 0. Denote the a-th target domain sample with the prediction result of class k.
[0083] Step 6: Design an adaptive unknown class detection threshold learning method to automatically update the unknown class detection threshold in each iteration cycle. Currently, most open-set domain adaptation fault diagnosis methods use a fixed threshold as the indicator for detecting unknown classes, without considering the characteristics of different open-set tasks, resulting in poor robustness of the model. Therefore, the present invention designs an adaptive unknown class detection threshold learning method for this problem. The specific implementation steps of this method are described as follows:
[0084] 1. Use the average value of all elements in the sample evaluation feature as the evaluation index of the sample, and calculate the maximum evaluation index of the source domain samples in each training batch:
[0085]
[0086] In the formula, B is the batch size, P = n s / B is the number of batches in each iteration cycle, is the operation function for taking the average of all elements in the evaluation feature, represents the probability vector of the b-th source domain sample in the p-th batch.
[0087] 2. Combine the maximum evaluation index obtained in each batch in the e-th iteration cycle into a vector:
[0088]
[0089] In the formula, represents the maximum evaluation index of the p-th batch obtained in the e-th iteration cycle, p ∈ [1, P].
[0090] 3. In order to retain the feature information of the previous iteration cycle and reduce the cumulative error, the present invention fuses the vector of the current iteration cycle with the vector of the previous iteration cycle, and thus obtains the unknown class detection threshold as:
[0091]
[0092] In the formula, κ is a balance factor used to balance the relative importance between and . It should be noted that when the iteration cycle E is obtained through training, the final unknown class detection threshold
[0093] 4. As the model training progresses, the feature distributions of the known classes in the source domain and the target domain will gradually converge, and their evaluation metrics will continuously converge. Therefore, to accurately measure the strength of the linear relationship between the evaluation metrics during the convergence process, the Pearson correlation coefficient ρ(·,·) is used to analyze the correlation between the vectors and and then used to adaptively update the balance factor in the unknown class detection threshold function, that is:
[0094]
[0095] When the vectors and are uncorrelated, that is, the model is still in the convergence state, at this time κ is calculated as 0.5, that is: the unknown detection threshold will be and the average of the maximum values.
[0096] 5. After the E-th iteration training is completed, the final unknown class detection threshold
[0097] Step 7: For the source domain sample data, two optimization objectives are designed to prompt the model to learn more discriminative fault features. First, the cross-entropy loss is used to increase the inter-class separability by taking the distance between source samples as a penalty:
[0098]
[0099] In the formula, represents the probability vector of the b-th source domain sample in the current batch, and are the k-th and l-th elements of the Softmax parameter in the source domain samples respectively, and are the k-th and l-th elements of the probability vector of the b-th source domain sample in the current batch respectively. 1{·} is an indicator function, that is: it returns 1 if the condition is true, otherwise it returns 0.
[0100] In addition, the center loss is introduced to improve the intra-class compactness of the source domain samples, that is, by calculating the distance between the feature and its feature center as a penalty to make the features with the same label close to each other, and its loss function is expressed as:
[0101]
[0102] In the formula, c yb represents the feature center of class y b , and u 2,b represents the output feature of the first FC layer of the b-th sample in the current batch.
[0103] Combining formula (22) and formula (23), the loss function of the source domain samples is obtained as follows:
[0104]
[0105] where μ is the penalty factor.
[0106] Step 8: For the target domain sample data, introduce the entropy loss function to improve the class discrimination ability of the target domain samples. As a component of evaluating features, minimizing the entropy value can provide a more reliable detection basis for the DLMMD method and the adaptive threshold. Therefore, the entropy loss function of the target domain samples is expressed as:
[0107]
[0108] where represents the k-th element in the b-th probability vector in the current batch in the target domain.
[0109] Subsequently, use the DLMMD method described in Step 5 to perform conditional feature distribution dynamic alignment on the shared feature u1 described in Step 3 and the first FC layer feature u2 described in Step 4, and its loss function is expressed as:
[0110]
[0111] where λ is the weight factor, and respectively represent the shared features in the source domain and the target domain, and respectively represent the output features of the first FC layer in the source domain and the target domain.
[0112] Step 9: Design a min-max entropy game training strategy. Through the source domain loss function L s described in Step 7 and the entropy loss function DA described in Step 8 and the DLMMD loss function L
[0113] 1. Update the parameters θ of the shared feature extractorG and the fault classifier parameter θ C . Use the training samples of the source domain and the target domain to preliminarily train the structural parameters of the shared feature extractor and the fault classifier. At this time, the loss function of the model is as follows:
[0114]
[0115] In the formula, α and β are trade-off coefficients.
[0116] 2. Update the parameter θ of the shared feature extractor G and fix the parameter θ of the fault classifier C . While minimizing the intra-class and inter-class distances of the source domain samples, perform maximum entropy optimization on the target domain samples. At this time, the loss function of the model is as follows:
[0117]
[0118] 3. Repeat the training steps of the above 1-2, continuously optimize the class boundary and align the cross-domain feature distribution of the known classes, and finally a more accurate and clearer open-set fault decision boundary can be generated.
[0119] Step 10: For the sample of the probability vector y test , use the evaluation feature function described in step 5 to generate the evaluation feature fea(y test ) of the test sample. After averaging all the elements in the evaluation feature, the evaluation index is obtained and compared with the final unknown class detection threshold described in step 6 . Finally, the detection of the unknown class and the identification of the fault modes in the known classes are completed. Specifically: if the evaluation index is less than the threshold , then the sample is a known class; otherwise, the sample is an unknown class. Thus, the final output result of the network model proposed by the present invention is:
[0120]
[0121] In the formula, is the final diagnostic output result of the test sample.
[0122] Based on the above method, this embodiment also proposes an intelligent diagnosis system for mechanical equipment faults under an open set, and the specific content includes:
[0123] Data acquisition and dataset creation module: Collect vibration signals of mechanical equipment in various health states under different working conditions, and perform sample division and normalization preprocessing on the vibration signal data; According to the requirements of the open-set task, randomly select some source domain data in the healthy state as known class samples, while the target domain data includes all health states, that is, the target domain contains both known classes and unknown classes; Select samples from the source domain and the target domain in proportion to create a training dataset, and select a quantitative target domain data as the test dataset.
[0124] Model construction and training module: Build a dynamic sub-domain adaptive network model architecture, including a shared feature extractor, a fault classifier, an open-set dynamic sub-domain adaptive module, and an adaptive unknown class detection threshold learning method, and initialize the model structure parameters; Use the min-max entropy game strategy to train the model, and continuously update the model parameters through the optimization objective to complete the training of the dynamic sub-domain adaptive network model.
[0125] Model performance verification module: Use the test dataset to verify the diagnostic performance of the trained dynamic sub-domain adaptive network model, analyze the test results, and output the optimal model and the corresponding unknown class detection threshold.
[0126] Fault diagnosis module: In different open-set domain adaptation fault diagnosis scenarios, use the optimal dynamic sub-domain adaptive network model to perform fault diagnosis on the vibration signals collected in real time, so as to realize cross-domain recognition of known fault modes and intelligent detection of unknown fault modes.
[0127] 1 Example 1
[0128] To better illustrate the technical effects of the present invention, a specific embodiment is used to experimentally verify the present invention. The specific embodiment is as follows:
[0129] Step 1: Data acquisition and dataset creation. Use the fault dataset of the cylindrical roller bearing test bench in the hoisting drive system of the offshore wind power lifting equipment for experimental verification. As Figure 5 shown, the test bench mainly includes a host computer, a motor, a servo control system, a data acquisition system, a magnetic brake, a tension controller, and an acceleration sensor, etc. In this example, the operating speed of the motor is set to 900 rpm, and the loads are set to two working conditions of 10 N / m and 20 N / m, and the vibration signals of the cylindrical roller bearing are acquired at a sampling frequency of 10240 Hz. In terms of fault setting, in this example, the faulty cylindrical roller bearing is installed in the bearing seat near the magnetic brake, and the test bearing is set to four fault states of rolling element fault (RF), inner race fault (IF), outer race fault (OF), and compound fault (CF, rolling element and outer race fault) (corresponding to labels 0-4 respectively). Among them, the damaged diameter width of the fault is 0.42 mm and the depth is 0.8 mm.
[0130] Table 1 Domain Adaptation Tasks
[0131]
[0132] Based on the above fault data description, in this example, vibration signals of five different health states are used. Each health state contains 300 samples, and each sample consists of 1024 data points. Therefore, according to the number of source domain health states in different domain adaptation tasks, 180 (60%) samples are selected from each source domain health state, and 18×n (where n is the number of source domain health states) samples are selected from each target domain health state to form the training set. Finally, 60 (20%) remaining target domain samples are randomly selected from each health state as the test data set. It should be noted that the labels of the source domain samples are randomly selected, while the target domain samples cover all health states. Thus, 10 domain adaptation tasks of the cylindrical roller bearing fault data set as shown in Table 1 can be obtained.
[0133] Step 2: Model structure and parameter settings. Table 2 shows the specific structure parameter table of the dynamic sub-domain adaptation network model in this case. Among them, zero padding is used in each one-dimensional convolutional layer to ensure the consistency of the feature dimensions.
[0134] Table 2 Structure Parameters of the Dynamic Sub-Domain Adaptation Network Model
[0135]
[0136]
[0137] In addition, the main training hyperparameters in the model training process of this case are shown in Table 3. It should be noted that the learning rate is decreased by 90% after every 10 iteration cycles, and each model training and test runs independently 10 times, and the mean value of the 10 experimental test results is used as the final diagnostic performance of the model to reduce the influence of randomness.
[0138] Table 3 Model Training Hyperparameters
[0139]
[0140] Step 3: Comparison benchmark model settings. To prove the superiority of the model proposed in the present invention, existing advanced and relevant methods are used in this example to compare with the present invention. The specific information of these benchmark models is described as follows:
[0141] (1) DANN: A domain adversarial neural network model consisting of a feature extractor, a task classifier, and a domain classifier (see details in: Ganin Y, Ustinova E, Ajakan H, et al. Domain-Adversarial Training of Neural Networks. J Mach Learn Res. 2016, 17, 2096-30.).
[0142] (2) OSBP: An open-set domain adaptation model based on backpropagation, which identifies unknown classes by adding additional predicted probability elements to the output layer (see details in: Saito K, Yamamoto S, Ushiku Y, Harada T. Open Set Domain Adaptation by Backpropagation. Computer Vision-Eccv 2018, Pt V. 2018, 11209, 156-71.).
[0143] (3) UDA: A classic general domain adaptation method that can achieve feature alignment of known classes through a weighted strategy based on sample domain similarity (see details in: You KC, Long MS, Cao ZJ, Wang JM, Jordan MI. Universal Domain Adaptation. Proc Cvpr Ieee 2019. p. 2715-24.).
[0144] (4) OSDAM: An advanced open-set domain adaptation method for mechanical fault diagnosis (see details in: Zhang W, Li X, Ma H, Luo Z, Li X. Open-Set Domain Adaptation in Machinery Fault Diagnostics Using Instance-Level Weighted Adversarial Learning. Ieee T Ind Inform. 2021, 17, 7445-55.).
[0145] (5) DANet: An open-set fault diagnosis method, whose model structure includes an auxiliary domain discriminator, an extended classifier, and a feature generator, and realizes cross-domain diagnosis of known classes and separation of unknown classes through dual adversarial learning (see details in: Zhao C, Shen WM. Dual adversarial network for cross-domain open set fault diagnosis. Reliab Eng Syst Safe. 2022, 221, 108358.).
[0146] Step 4: Analysis of experimental results. To more precisely verify the effectiveness and superiority of the model proposed in the present invention, four performance indicators are used for quantitative comparison: the accuracy OS of all test samples, the accuracy OS* of known classes, the accuracy UK of unknown classes, and the harmonic mean H-Score of OS and UK. Table 4 lists the open-set diagnosis test results of all comparison models. It can be seen from the table that DANN has the best diagnostic performance for the closed-set domain adaptation tasks T1 and T2, while for all other open-set domain adaptation tasks, the average OS of the model proposed in the present invention is the highest. Overall, the average OS of the model proposed in the present invention in all test tasks is the highest at 98.60%, while the average OS of DANN, OSBP, UDA, OSDAM, and DANet in all test tasks is only 67.84%, 74.07%, 75.11%, 69.75%, and 81.41% respectively, far lower than the model proposed in the present invention. Therefore, the model proposed in the present invention has better diagnostic performance and robustness, and shows great potential in the fault diagnosis of mechanical equipment under open sets.
[0147] Table 4 Open-set diagnosis test results
[0148]
[0149] In addition, Figure 6 A comparison graph of the average H-Score between the model proposed in the present invention and 4 open-set domain adaptation models is given. It can be seen from the graph that the average H-Score of the model proposed in the present invention in each open-set task is better than that of other open-set domain adaptation models, and the smallest average H-Score is also greater than 94%. In addition, the average H-Score of the model proposed in the present invention in 8 open-set domain adaptation tasks is the highest at 98.35%. Generally speaking, the model proposed in the present invention has better accuracy and stability in simultaneously handling cross-domain diagnosis and unknown class detection.
[0150] To intuitively understand the accuracy of the model proposed in the present invention for each health state, a confusion matrix is used to visually analyze the representative DANN, OSBP, and DANet. Figure 7The confusion matrix diagram of the first test of the model proposed in the present invention and different baseline models in the T9 task is shown. It can be seen from the figure that in terms of detecting unknown categories, DANN cannot detect unknown categories, and although OSBP has the highest diagnostic accuracy for unknown categories, 6.67% in the second category is identified as an unknown category. In addition, DANet has the worst diagnostic performance for known categories. On the contrary, the model proposed in the present invention has a diagnostic accuracy of 93.33% for unknown categories and a diagnostic accuracy of 100% for all known categories, indicating that the model proposed in the present invention has better comprehensive diagnostic performance under the open set.
[0151] Although the illustrative specific embodiments of the present invention have been described above for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. An intelligent diagnosis method for mechanical equipment faults under an open set, characterized in that, Including: S1. Use a linear embedding layer to serialize the one-dimensional vibration signal and convert it into a one-dimensional token embedding sequence that meets the requirements of the Transformer encoder. S2. Construct an FCformer encoder with Fourier and convolutional embedding layers, and input the one-dimensional token embedding sequence into the FCformer encoder for time-frequency feature extraction. S3. Use a global average pooling layer to obtain shared features, and combine the linear embedding layer and T FCformer encoders to form a shared feature extractor. S4. Design a fault classifier based on a convolutional neural network to identify fault patterns for the shared features. S5. Build an open-set dynamic subdomain adaptation module to achieve dynamic subdomain feature alignment under the open set by capturing the fine-grained information of each known class. S6. Design an adaptive unknown class detection threshold learning method, use the sample evaluation metric function to automatically update the unknown class detection threshold in each iteration cycle, and obtain the final unknown class detection threshold after iterative training. S7. For the source domain sample data, design a source domain loss function including cross-entropy loss and center loss to discriminate fault features. S8. For the target domain sample data, introduce entropy loss to improve the class discrimination ability of the target domain samples, and use the DLMMD loss function to perform dynamic conditional feature distribution alignment. S9. Design a min-max entropy game training strategy, and perform game training on the shared feature extractor and the fault classifier through the source domain loss function, entropy loss function, and DLMMD loss function. S10. Use the evaluation metric function to generate the evaluation metrics of the target domain test samples, and compare them with the unknown class detection threshold to complete the intelligent detection of unknown classes and cross-domain identification of fault patterns in known classes.
2. The intelligent diagnosis method according to claim 1, wherein The FCformer encoder consists of an FCE layer, a layer normalization, a residual connection, a feed-forward layer, a layer normalization, and a residual connection in sequence. The FCE layer includes a 1D-DFT layer, a one-dimensional convolutional layer, a batch normalization (BN) layer, and a rectified linear unit (ReLU) layer. After the 1D-DFT layer performs DFT transformation, the real part of the feature map is retained, and the residual connection superimposes the feature maps before and after the corresponding layer normalization. The feed-forward layer includes a one-dimensional convolutional layer, a BN layer, a ReLU layer, a dropout layer, and a one-dimensional convolutional layer.
3. The intelligent diagnosis method according to claim 1, wherein The output of the FCformer encoder for time-frequency feature extraction is: where z t is the output token embedding sequence of the FCformer encoder, LN(·) is the layer normalization function, FF(·) represents the processing function of the feed-forward layer, and z t ″′ is the feature map obtained by the second residual connection, Q is the number of sequence blocks, and N is the dimension of the token embedding sequence.
4. The intelligent diagnosis method according to claim 1, characterized in that The fault classifier includes a convolutional extraction layer and two fully connected layers. The convolutional extraction layer includes a one-dimensional convolutional layer, a BN layer, a ReLU layer, and a max pooling layer. Flattening is performed between the convolutional extraction layer and the fully connected layer.
5. The intelligent diagnosis method according to claim 4, wherein The output features u2 of the first FC layer and the final output probability vector y of the second FC layer are respectively expressed as: where, F C (·) represents the operation of the convolutional extraction layer, ψ(·) is the flattening function, f FC1 (·) and f FC2 (·) represent the operations of the first and second FC layers respectively; D is the output dimension of the first FC layer, and K is the number of known classes.
6. The intelligent diagnosis method according to claim 1, wherein Step 5 specifically includes: 5.
1. Construct sample evaluation features using entropy value and confidence. 5.
2. Use the K-means algorithm to perform clustering analysis on the evaluation features of the target domain samples to provide pseudo-labels for known classes and unknown classes for subsequent training. 5.
3. According to the pseudo-label h of the target domain samples t , mask the unknown class samples in the probability vector of the a-th target domain sample to obtain an optimized probability vector; 5.
4. Assign different weights to specific subdomains according to the number of subdomain samples to align the subdomain features of known classes.
7. The intelligent diagnosis method according to claim 1, characterized in that, The calculation formula for the unknown class detection threshold is: Among them, Where κ is a balance factor used to balance the relative importance between the vector of the current iteration cycle and the vector of the previous iteration cycle , and ρ(·,·) is the Pearson correlation coefficient; represents the maximum evaluation index of the p-th batch obtained in the e-th iteration cycle.
8. The intelligent diagnosis method according to claim 1, wherein The source domain sample loss function is: Among them, the cross-entropy loss: Center loss: where μ is the penalty factor, B is the batch size, K is the number of known classes, represents the probability vector of the b-th source domain sample in the current batch, and are the k-th and l-th elements of the Softmax parameter in the source domain sample, respectively, and are the k-th and l-th elements of the probability vector of the b-th source domain sample in the current batch, respectively; 1{·} is an indicator function, that is, it returns 1 if the condition is true and 0 otherwise; c yb represents the feature center of class y b , and u 2,b represents the output feature of the first FC layer of the b-th sample in the current batch.
9. The intelligent diagnosis method according to claim 1, wherein The DLMMD loss function is: Wherein, is the DLMMD function, P s and P T are the feature distributions of the source domain and the target domain respectively, n s is the number of target domain samples, n t is the number of target domain samples, and respectively represent the weights of the i-th source domain sample and the a-th target domain sample belonging to the class k, K is the number of known classes, H is the Reproducing Kernel Hilbert Space (RKHS), φ(·) represents the non-linear mapping function from the feature space to the RKHS, ||·|| 2 represents the L2 norm; represents the number of samples belonging to the class k in the source domain, represents the number of samples belonging to the unknown class in the target domain; For source domain samples Convert the true label into a one-hot vector To calculate the weights For target domain samples without labels Use the optimized prediction vector As a one-hot vector to calculate the weights Denote the number of target domain samples in the target domain that are of known classes and belong to class k. Its specific definition is as follows: where \(G(\cdot)\) is the indicator function, that is: if the target domain sample is a known class and belongs to class \(k\), then return 1, otherwise return 0; represents the \(a\)-th target domain sample with the predicted result of class \(k\); represents the pseudo-label value of the \(a\)-th target domain sample, indicating that the target domain sample is a known class.
10. An intelligent diagnosis system for mechanical equipment faults under an open set, characterized in that, including: Data acquisition and dataset creation module: Collect the vibration signals of mechanical equipment in various health states under different working conditions, and perform sample division and normalization preprocessing on the vibration signal data; Randomly select some source domain data in the healthy state as known class samples, while the target domain data includes all health states, that is, the target domain contains both known classes and unknown classes; Select samples from the source domain and the target domain according to a ratio to create a training dataset, and select a quantitative target domain data as a test dataset; Model construction and training module: Construct a dynamic subdomain adaptive network model architecture, including a shared feature extractor, a fault classifier, an open set dynamic subdomain adaptive module, and an adaptive unknown class detection threshold learning method, and initialize the model structure parameters; Use the min-max entropy game strategy to train the model, update the parameters of the model by optimizing the objective, and complete the training of the dynamic subdomain adaptive network model; Model performance verification module: Use the test dataset to verify the diagnostic performance of the trained dynamic subdomain adaptive network model, analyze the test results, and output the optimal model and the corresponding unknown class detection threshold; Fault diagnosis module: In different open set domain adaptation fault diagnosis scenarios, use the optimal dynamic subdomain adaptive network model to perform fault diagnosis on the real-time collected vibration signals, so as to realize cross-domain recognition of known fault modes and intelligent detection of unknown fault modes.
Citation Information
Patent Citations
Cross-domain fault diagnosis method and system for rolling bearing with unknown inter-domain data label relation
CN117312984A
Bearing fault diagnosis method based on Transform model
CN117387948A
Self-supervised 360-degree depth estimation method, device, equipment and medium
CN117808857A
Dynamic joint distribution alignment network-based bearing fault diagnosis method under variable working conditions
US20230168150A1
Cited By
Open set anomaly detection method for rotating machinery based on wavelet network
CN121615024A