Fault Diagnosis Method for Rotating Machinery under Variable Working Conditions Based on Domain Adaptation Features
By evaluating the time-frequency domain feature data of rotating machinery, combining feature quantization evaluation and joint distribution adaptation, an integrated learning classifier is used to solve the problem of weak generalization ability of fault diagnosis model under variable operating conditions of rotating machinery, and achieve higher fault diagnosis accuracy and model generalization ability.
Patent Information
- Application Number
- CN202311131734.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-04
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-09-04
AI Technical Summary
The prior art lacks sufficient labeled fault samples under rotating mechanical variable conditions, resulting in weak generalization capabilities of deep learning models, and transfer learning methods fail to effectively consider feature discrimination capabilities and distribution adaptation effects, affecting the accuracy of fault diagnosis.
By extracting the time-frequency domain statistical feature data of the rotating machinery, using feature classification accuracy, structural similarity index and Frecher distance evaluation characteristics, feature quantification evaluation indexes are constructed, joint distribution adaptation is performed, and fault diagnosis is used using integrated learning classifiers.
It improves the accuracy of cross-domain fault diagnosis and model generalization capabilities, effectively removes interference and redundant features, and improves the fault identification performance under rotating mechanical variable operating conditions.
Smart Images

Figure CN117093924B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mechanical fault diagnosis methods, and specifically to a variable working condition fault diagnosis method for rotating machinery based on domain adaptation features. Background Technique
[0002] With the rapid development of a series of artificial intelligence methods such as machine learning, in the data-driven rotating machinery fault diagnosis method, the artificial intelligence-based fault diagnosis framework has gradually become a research hotspot. Currently, deep learning methods have attracted the attention and research of many researchers due to their powerful hidden feature mining ability, and many research results have been obtained. However, due to the complex working conditions of rotating machinery in actual industrial scenarios, the fault diagnosis model based on deep learning faces two technical problems: (1) There is a lack of sufficient labeled fault samples. In actual industrial scenarios, under variable and complex working conditions of rotating machinery, the sample data in different fault states is lacking, and the cost of obtaining sufficient labeled samples is very high. (2) Under different working conditions, there are distribution differences in samples of the same fault category, resulting in weak generalization ability of the trained artificial intelligence-based fault diagnosis model and low accuracy in fault diagnosis applied to actual working conditions. (3) Currently, although the deep learning-based fault diagnosis method has received extensive attention and research due to its powerful feature mining ability, it has defects such as hyperparameters, high time consumption, and computational complexity.
[0003] As a promising research direction to solve the above problems, domain adaptation based on transfer learning has gradually been concerned and studied by researchers in recent years. The transfer learning method can mine and learn knowledge from an existing domain (source domain: labeled fault samples under existing working conditions) and train a fault diagnosis model to identify and classify fault samples from different domains (target domain: unlabeled fault samples under variable working conditions).
[0004] In the prior art, in the paper "Joint Distribution Adaptation Transfer Fault Diagnosis of Bearings under Variable Working Conditions" (authors: Liu Yingdong; Liu Tao; Li Hua; Wang Tingxuan) published in the Journal of Electronic Measurement and Instrumentation in May 2021, a bearing fault diagnosis method based on transfer learning and joint distribution was disclosed, which mainly includes four steps:
[0005] (1) Dataset division. The original bearing data is divided into a training set, a test set, and an auxiliary dataset according to different working conditions, where the test set and the auxiliary dataset are of the same working condition.
[0006] (2) Feature extraction. Time domain features of the bearing data are extracted, and weights of each feature are calculated for the extracted time domain features through the FLDA method.
[0007] (3) Joint distribution adaptation. The feature vectors composed of features with larger weight values are respectively subjected to dimensionality reduction learning through PCA and KPCA, and transfer learning through TCA and JDA. On the basis of dataset division, the model is assisted in training by adding an auxiliary dataset with the same working conditions as the test set to the training set, and the test set remains unchanged. The classification accuracies of each method are compared under the condition of adding auxiliary datasets with different proportions.
[0008] (4) Fault identification. Finally, the source domain data after learning is used as the training set, and the target domain data is used as the test set and sent to the KNN classifier for diagnostic classification. The classification accuracies of each method are compared. The main problems in the technical solution of this paper are as follows:
[0009] (A) Using FLDA to calculate the weights of each feature. Since only weights are used to evaluate the importance of features in the process of transferable feature extraction, the influence of feature discriminative ability on cross-domain fault diagnosis is ignored, resulting in incomplete feature extraction results.
[0010] (B) JDA can make up for the limitation of TCA that only considers marginal probability distribution adaptation, comprehensively considers two probability distributions, and thus improves the transfer learning effect. However, in the process of distribution adaptation for source domain and target domain samples, the inter-domain conditional probability distribution and marginal distribution differences are not reasonably considered at the same time, and the feature distortion problem existing in the process of reducing distribution differences in the high-dimensional feature space is not considered either, resulting in poor distribution adaptation effect and affecting the generalization ability performance of the fault diagnosis model.
[0011] (C) Using the KNN classifier for diagnostic classification, the classification effect of a single classifier is not good, and the fault diagnosis ability is not strong. Summary of the Invention
[0012] The present invention provides a method for rotating machinery variable working condition fault diagnosis based on domain adaptation features to solve the problems existing in the prior art, such as ignoring the influence of feature discriminative ability on cross-domain fault diagnosis and poor distribution adaptation effect.
[0013] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0014] A method for rotating machinery variable working condition fault diagnosis based on domain adaptation features, comprising the following steps:
[0015] Step 1: Obtain the rotating machinery vibration signals with labels under existing working conditions and the rotating machinery vibration signals without labels under variable working conditions, extract the time-frequency domain statistical feature data of the rotating machinery vibration signals with labels under existing working conditions as the labeled source domain feature sample set, and extract the time-frequency domain statistical feature data of the rotating machinery vibration signals without labels under variable working conditions as the unlabeled target domain feature sample set;
[0016] Step 2: Based on the statistical feature data in the labeled source domain feature sample set, calculate the feature classification accuracy acc of each statistical feature data in the source domain feature sample set to characterize the discriminant performance of the features; based on the statistical feature data in the normal state of the source domain feature sample set and the statistical feature data in the normal state of the target domain feature sample set, calculate the structural similarity index SSIM and FID score of each statistical feature data to characterize the domain invariance of the features;
[0017] Based on the obtained feature classification accuracy acc, SSIM, and FID, construct the feature quantization evaluation index of each statistical feature data
[0018] Then set a threshold, and select multiple time-frequency domain statistical feature data with a feature quantization evaluation index Z greater than the set threshold from the source domain feature sample set to construct a labeled source domain feature sample subset, and select multiple time-frequency domain statistical feature data with a feature quantization evaluation index Z greater than the set threshold from the target domain feature sample set to construct an unlabeled target domain feature sample subset;
[0019] Step 3: Perform joint distribution adaptation on the time-frequency domain statistical feature data in the source domain feature sample subset and the target domain feature sample subset obtained in Step 2 to obtain the source domain feature sample subset and the target domain feature sample subset after joint distribution adaptation;
[0020] Step 4: Use the data in the source domain feature sample subset after joint distribution adaptation obtained in Step 3 to train the fault diagnosis classifier, and then input the data in the target domain feature sample subset after joint distribution adaptation obtained in Step 3 into the trained fault diagnosis classifier, and obtain the fault diagnosis result of the target domain through the fault diagnosis classifier.
[0021] Furthermore, in Step 1, perform wavelet transform decomposition and reconstruction on the rotating machinery vibration signals with existing working conditions and labels and the rotating machinery vibration signals without labels under variable working conditions respectively to obtain the reconstructed signals, then extract the time-domain statistical features of multiple statistical parameters based on the reconstructed signals respectively, and then extract the frequency-domain statistical features of multiple statistical parameters based on the Hilbert envelope spectrum calculation results of the reconstructed signals respectively, so as to correspondingly obtain the time-frequency domain statistical feature data of the rotating machinery vibration signals with existing working conditions and labels and the time-frequency domain statistical feature data of the rotating machinery vibration signals without labels under variable working conditions.
[0022] Furthermore, the statistical parameters include mean, standard deviation, kurtosis, energy, energy entropy, peak factor, shape factor, skewness, extreme value, range, power spectrum entropy, singular spectrum entropy, approximate entropy, sample entropy, fuzzy entropy, permutation entropy, envelope entropy.
[0023] In the further step 2, the feature classification accuracy acc of each time-frequency domain statistical feature data in the source domain feature sample set is calculated by using the Xgboost classifier.
[0024] Further, when performing joint adaptation distribution in step 3, the minimum of both the maximum mean difference between the marginal probability distributions of the time-frequency domain statistical feature data in the source domain feature sample subset and the target domain feature sample subset, and the maximum mean difference between the conditional probability distributions is used as the total optimization objective for joint adaptation distribution.
[0025] In the further step 3, the stacking ensemble learning model is trained by using the time-frequency domain statistical feature data in the labeled source domain feature sample subset, and then the trained stacking ensemble learning model is used to predict the class labels of the time-frequency domain statistical feature data in the target domain feature sample subset. The obtained class labels are the pseudo-labels of the target domain feature sample subset. Based on the time-frequency domain statistical feature data and the corresponding pseudo-labels in the target domain feature sample subset, the conditional probability distribution of the time-frequency domain statistical feature data in the target domain feature sample subset is calculated.
[0026] In the further step 4, the fault diagnosis classifier is an SVM classifier.
[0027] The present invention effectively quantifies and evaluates the time-frequency statistical feature data of the source domain and the target domain by using the feature classification accuracy, FID, and SSIM. Then, the joint distribution adaptation with the minimum of the differences in marginal probability distribution and conditional probability distribution is used to adapt the distributions of the source domain and target domain feature subsets. Finally, an ensemble learning classifier is used for fault classification diagnosis, improving the cross-domain fault recognition performance. Therefore, compared with the prior art, the beneficial effects of the present invention are as follows:
[0028] (1) The present invention proposes a domain adaptation feature selection method based on feature classification accuracy, FID score, and structural similarity index, which can quantitatively evaluate the domain adaptation ability of statistical features, helps to select feature data more conducive to bearing fault diagnosis across different domains, effectively removes interfering and redundant features, and improves the accuracy of cross-domain fault diagnosis.
[0029] (2) The improved joint distribution adaptation based on strengthening domain generalization ability proposed in the present invention fully considers the differences in conditional probability distribution and marginal probability distribution. At the same time, due to the introduction of ensemble learning, the adaptation ability of distribution differences is strengthened. Compared with classical feature-based transfer learning methods (transfer component analysis, joint distribution adaptation, etc.), its ability to reduce the distribution differences between domains is better, and it can promote the improvement of the generalization ability of the fault diagnosis model. Description of the Drawings
[0030] Figure 1It is the flowchart of the embodiment of the present invention.
[0031] Figure 2 It is the schematic diagram of the stacking ensemble learner in the embodiment of the present invention. Specific implementation manners
[0032] In order to enable those skilled in the art to better understand the solution of the present invention, the following will, in conjunction with the accompanying drawings and embodiments, elaborate in detail on the implementation manners of the present invention, so as to fully understand the implementation process of how the present invention uses technical means to solve technical problems and achieve corresponding technical effects and be implemented accordingly. Each feature in the embodiments of the present invention and in the embodiments can be combined with each other on the premise of not conflicting, and the formed technical solutions are all within the protection scope of the present invention.
[0033] Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] It should be noted that the terms "including" and "having" in the specification and claims of the present invention and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] As Figure 1 shown, this embodiment discloses a rotating machinery variable working condition fault diagnosis method based on domain adaptation features, including the following steps:
[0036] Step 1: Obtain the rotating machinery vibration signals with labels in the existing working conditions (the vibration signals collected under working condition 1 in Figure 1 ), and the rotating machinery vibration signals without labels under variable working conditions (the vibration signals collected under working condition 2 in Figure 1 ), extract the time-frequency domain statistical feature data of the rotating machinery vibration signals with labels in the existing working conditions as the labeled source domain feature sample set, and extract the time-frequency domain statistical feature data of the rotating machinery vibration signals without labels under variable working conditions as the unlabeled target domain feature sample set.
[0037] In this embodiment, the dual-tree complex wavelet packet transform is used to process the vibration signals collected under working condition 1 and working condition 2 of the rotating machinery respectively, so as to realize the decomposition of the corresponding signals. The obtained terminal node signals are respectively reconstructed to obtain reconstructed signals, and then 18 statistical parameters of the reconstructed signals and their Hilbert envelope spectra are calculated respectively. Furthermore, the time-domain and frequency-domain statistical features of the rotating machinery vibration signals with labels under the existing working conditions and the time-domain and frequency-domain statistical features of the rotating machinery vibration signals without labels under variable working conditions are extracted. The time-frequency domain statistical feature data of the rotating machinery vibration signals with labels under the existing working conditions are used as the labeled source domain feature sample set, and the time-frequency domain statistical feature data of the rotating machinery vibration signals without labels under variable working conditions are used as the unlabeled target domain feature sample set.
[0038] Specifically, for the rotating machinery vibration signals with labels under the existing working conditions or the rotating machinery vibration signals without labels under variable working conditions, in order to effectively extract fault features from the original vibration signals, the dual-tree complex wavelet packet transform is used to decompose the corresponding vibration signals by 4 layers. 18 statistical parameters (including mean, standard deviation, kurtosis, energy, energy entropy, peak factor, shape factor, skewness, extreme value, range, power spectrum entropy, singular spectrum entropy, approximate entropy, sample entropy, fuzzy entropy, permutation entropy, envelope entropy) are calculated based on the reconstructed signals of the terminal nodes in the fourth layer, and 288 time-domain statistical features are extracted. Then, the Hilbert envelope spectrum of the reconstructed signal is calculated, and the obtained spectral signal is used to calculate 18 statistical parameters to extract 288 frequency-domain statistical features. Finally, a total of 576 statistical features are obtained, that is, the labeled source domain feature sample set contains a total of 288 time-domain statistical feature data and a total of 288 frequency-domain statistical feature data, and the unlabeled target domain feature sample set also contains a total of 288 time-domain statistical feature data and a total of 288 frequency-domain statistical feature data.
[0039] Step 2: Based on the statistical feature data in the labeled source domain feature sample set, calculate the feature classification accuracy acc of each statistical feature data in the source domain feature sample set to characterize the discriminant performance of the features; based on the statistical feature data in the normal state of the source domain feature sample set and the statistical feature data in the normal state of the target domain feature sample set, calculate the structural similarity index SSIM and FID score of each statistical feature data to characterize the domain invariance of the features.
[0040] Although the time-frequency analysis method based on wavelet analysis can extract fault features from vibration signals with non-stationarity, it will cause a high-dimensional feature set, with interfering and redundant features, thus affecting the accuracy of fault mode recognition and classification. In addition, since rotating machinery often works under complex working conditions, it will lead to distribution differences in vibration signals under different working conditions, and the lack of actual fault signals, thus resulting in poor fault diagnosis effects.
[0041] In this regard, in order to reduce the interference and redundant features in the high-dimensional original feature set and select features that are beneficial to fault mode recognition and classification and have small inter-domain distribution differences, this embodiment proposes a domain adaptation feature selection method DAFS-AFS. The DAFS-AFS method evaluates features from two aspects: the class separability of features and the domain invariance of features. For class separability, the feature classification accuracy acc is used for quantitative evaluation; for domain invariance, one-dimensional feature data is converted into two-dimensional data, and the structural similarity index SSIM (Structural Similarity) and the Frechet Inception Distance score FID are used to calculate the similarity of features between different domains, which is used to characterize its domain invariance.
[0042] The specific calculation process is described as follows:
[0043] (S1) Calculate the feature classification accuracy acc of each time-frequency domain statistical feature data in the source domain feature sample set.
[0044] In this embodiment, the Xgboost classifier is used to calculate the feature classification accuracy acc of each time-frequency domain statistical feature data in the source domain feature sample set to characterize the discriminant performance of the features, that is, the discriminant ability of the statistical feature data in the source domain feature sample set and the target domain feature sample set can be quantified through the feature classification accuracy acc.
[0045] The Xgboost classifier is one of the boosting algorithms. The idea of the Boosting algorithm is to integrate many weak classifiers to form a strong classifier. When training the Xgboost classifier, the time-frequency domain statistical feature data in the source domain feature sample set is used as the data of the training set sample I, denoted as I = {(x1, y1), (x2, y2),...(x m , y m )}, where x m is the mth data in the training set sample I, and y m is the class label of x m . Let the maximum number of iterations be T and the loss function be L t . The regularization coefficients are λ and γ respectively. Then the specific steps of training the Xgboost classifier are as follows:
[0046] (a) Define the loss function L t As shown in formula (1):
[0047]
[0048] In formula (1), x i is the ith data in I; y i is xi Category label; f t-1 (x i ) is the prediction result of the first t - 1 classifiers; ω tj is the value of the leaf node weight; T t (x i ) is the i-th classifier; J is the number of leaf nodes; L(y i , f t-1 (x i )) is a constant.
[0049] Perform a second-order Taylor expansion on the loss function L t to obtain formula (2):
[0050]
[0051] where g ti and h ti are the first-order and second-order gradient statistics of the loss function respectively.
[0052] In formula (2), L(y i , f t-1 (x i )) is a constant and will not affect the derivative. Since the values of the j-th leaf node of each decision tree will ultimately be the same value, it can be further simplified to formula (3):
[0053]
[0054] where R tj represents the instance set of the j-th leaf node.
[0055] Let to obtain formula (4):
[0056]
[0057] (b) Calculate g ti , h ti for the i-th data in the training set sample, and there are formulas (5), (6):
[0058]
[0059]
[0060] (c) Try to split the decision tree based on the current node. By default, the score score = 0, and G and H are the sums of the first-order and second-order derivatives of the current node respectively.
[0061] Let G L = 0, H L= 0, arrange the data in the training set samples in ascending order of the feature values of feature k, sequentially take out the sample data corresponding to the j-th feature value, and calculate the first-order and second-order derivatives of the left and right subtrees after putting the current sample data into the left subtree. As shown in formula (7):
[0062] G L = G L + g ti , G R = G - G L
[0063] H L = H L + h ti , H R = H - H L (7),
[0064] Among them, G L and H L are the sums of the first-order and second-order derivatives of non-sparse value sample data.
[0065] Try to update the maximum score score as shown in formula (8):
[0066]
[0067] Split the subtree based on the splitting feature and feature value corresponding to the maximum score.
[0068] If the maximum score score is 0, the current decision tree is established, calculate ω tj for all leaf regions, and obtain the weak learner T t , update the strong learner f t (x) = f t-1 (x)+εT t (x), where ε is the step size, usually taking the value of 0.1, and then enter the next round of weak learner iteration. If the maximum score score is not 0, continue to try to split the decision tree.
[0069] After the decision tree is established, use the XGBoost classifier after the decision tree is established to calculate the feature classification accuracy acc of each time-frequency domain statistical feature data in the source domain feature sample set. The larger the calculated feature classification accuracy acc value, the higher the class separability of the corresponding time-frequency domain statistical feature data, which is more conducive to classification.
[0070] (S2), calculation of the structural similarity index SSIM of each time-frequency domain statistical feature data.
[0071] The structural similarity index SSIM is a metric for measuring the structural similarity between two given signals or samples, which is determined according to the luminance comparison function, contrast comparison function, and structure comparison function. Among them:
[0072] The luminance value of the feature sample set is taken from the average value of all statistical feature data therein. In this embodiment, it is assumed that the labeled source domain feature sample set is x, and the time-frequency domain statistical feature data in x is x = {x1, x2, …, x N}, then the luminance value μ x of the labeled source domain feature sample set is as shown in formula (9):
[0073]
[0074] In formula (9), x i is the i-th time-frequency domain statistical feature data in the labeled source domain feature sample set, and N is the total number of feature values.
[0075] Similarly, it is assumed that the unlabeled target domain feature sample set is y, and the time-frequency domain statistical feature data in y is y = {y1, y2, …, y N}, y i is the i-th time-frequency domain statistical feature data in the labeled source domain feature sample set, and N is the total number of feature values. Then the luminance value
[0076] The contrast of the feature sample set is taken from the standard deviation (square root of the variance) of all statistical feature data therein. In this embodiment, it is assumed that the contrast of the labeled source domain feature sample set (i.e., the standard deviation of all statistical feature data in the labeled source domain feature sample set) is σ x , then σ x is as shown in formula (10):
[0077]
[0078] Similarly, it is assumed that the contrast of the unlabeled target domain feature sample set (i.e., the standard deviation of all statistical feature data in the unlabeled target domain feature sample set) is σ y , then
[0079] It is assumed that the luminance comparison function of the labeled source domain feature sample set x and the unlabeled target domain feature sample set y is l(x, y), then the luminance comparison function l(x, y) is as shown in formula (11):
[0080]
[0081] In formula (11), C1 is a constant to ensure stability when the denominator is 0.
[0082] Given a contrast comparison function \(c(x, y)\) for the labeled source domain feature sample set \(x\) and the unlabeled target domain feature sample set \(y\), the contrast comparison function \(c(x, y)\) is as shown in formula (12):
[0083]
[0084] In formula (12), \(C2\) is a constant to ensure stability when the denominator is 0.
[0085] Given a structure comparison function \(s(x, y)\) for the labeled source domain feature sample set \(x\) and the unlabeled target domain feature sample set \(y\), the structure comparison function \(s(x, y)\) is as shown in formula (13):
[0086]
[0087] In formula (13), \(C3\) is a constant to ensure stability when the denominator is 0.
[0088] Given that the covariance between the feature samples of the labeled source domain feature sample set \(x\) and the unlabeled target domain feature sample set \(y\) is \(\sigma\) xy , then there is formula (14):
[0089]
[0090] Then the final structural similarity index SSIM is as shown in formula (15):
[0091] SSIM(x,y)=[l(x,y)] α ·[c(x,y)] β ·[s(x,y)] γ (15),
[0092] In formula (15), \(\alpha\), \(\beta\), and \(\gamma\) are weighting coefficients, generally taken as 1.
[0093] Formula (15) can finally be simplified to as shown in formula (16):
[0094]
[0095] (S3) Calculate the FID score FID for each time-frequency domain statistical feature data.
[0096] In this embodiment, the Fréchet distance algorithm is used to calculate the FID score FID for each time-frequency domain statistical feature data in the source domain feature sample set and the target domain feature sample set respectively. The FID index is used to characterize the distance between two distributions. The smaller the distance, the higher the similarity between the two distributions. Therefore, the smaller the FID, the better. The calculation of FID is as shown in formula (17):
[0097]
[0098] After obtaining the structural similarity index SSIM according to formula (16) and the FID score according to formula (17), calculate As a comprehensive index, the higher this index, the higher the class separability of the characterized features, and the more conducive to classification.
[0099] Finally, after obtaining the feature classification accuracy acc, the structural similarity index SSIM, and the FID score FID, based on these three indicators, the time-frequency domain statistical feature data in each of the source domain feature sample set and the target domain feature sample set are respectively subjected to feature quantization evaluation.
[0100] Specifically, construct the feature quantization evaluation index for each statistical feature data The larger the feature quantization evaluation index Z, the stronger the transferability of the corresponding time-frequency domain statistical feature. Then set a threshold, and select multiple time-frequency domain statistical feature data with the feature quantization evaluation index Z greater than the set threshold from the source domain feature sample set to construct a labeled source domain feature sample subset, and select multiple time-frequency domain statistical feature data with the feature quantization evaluation index Z greater than the set threshold from the target domain feature sample set to construct an unlabeled target domain feature sample subset.
[0101] In this embodiment, considering that the source domain feature sample set is a labeled feature set and the label information is available a priori, the data in the target domain feature sample set are all unlabeled data except for the normal state data, the target domain feature sample set is not a priori and the label information is not available, and the statistical feature data categories in the source domain feature sample set are the same as those in the target domain feature sample set, both having a total of 288 time domain statistical feature data and a total of 288 frequency domain statistical feature data. Therefore, in this embodiment, the feature classification accuracy acc of each type of statistical feature data is obtained using the source domain feature sample set, then the structural similarity index SSIM and the FID score of each statistical feature data are calculated, and finally the feature quantization evaluation index Z of each statistical feature data is constructed. After obtaining the Z values of each statistical feature data, they are sorted in descending order, and a threshold for the feature quantization evaluation index is set. Statistical feature data with Z values greater than the set threshold are selected from the source domain feature sample set to construct the corresponding sample subset, and statistical feature data with Z values greater than the set threshold are selected from the target domain feature sample set to construct the corresponding sample subset for subsequent processing steps.
[0102] Step 3: Perform joint distribution adaptation on the time-frequency domain statistical feature data in the source domain feature sample subset and the target domain feature sample subset obtained in Step 2 to obtain the source domain feature sample subset and the target domain feature sample subset after joint distribution adaptation.
[0103] This embodiment proposes an Improved Joint Distribution Adaptation (IJDA) to perform joint distribution adaptation on the source domain feature sample subset and the target domain feature sample subset to reduce the distribution difference. The process of the improved joint distribution adaptation is as follows:
[0104] Let there be a labeled source domain feature sample subset Let there be an unlabeled target domain feature sample subset Where:
[0105] x i is the i-th sample data; y i is the class label of the i-th sample data; n S and n T respectively represent the number of source domain and target domain samples.
[0106] There are differences in both the marginal probability distribution and the conditional probability distribution between the source domain and the target domain, that is, Q S (y S |x S ) ≠ Q T (y T |x T ) and P S (x S ) ≠ P T (x T ). Among them, P S (W T x S ) represents the marginal probability distribution of the time-frequency domain statistical feature data in the source domain feature sample subset, P T (W T x T ) represents the marginal probability distribution of the time-frequency domain statistical feature data in the target domain feature sample subset, Q S (y S |W T x S ) represents the conditional probability distribution of the time-frequency domain statistical feature data in the source domain feature sample subset, Q T (y T| W T x T ) represents the conditional probability distribution of the time-frequency domain statistical feature data in the target domain feature sample subset.
[0107] The goal of the IJDA algorithm is to use D S and D T to learn a feature mapping transformation matrix W, such that after the transformation P S(W T x S ) and P T (WT x T ) distance, Q S (y S |W T x S ) and Q T (y T |W T x T ) distances are minimized as much as possible. Therefore, the IJDA algorithm includes two aspects of optimization objectives:
[0108] (A) Achieve the adaptation of the marginal probability distributions of the time-frequency domain statistical feature data in the source domain feature sample subset and the target domain feature sample subset, that is, P S (W T x S ) and P T (W T x T ) with the minimum maximum mean discrepancy MMD. The optimization objective expression is shown in Equation (18):
[0109]
[0110] In Equation (18), X is the data matrix containing the source domain and target domain feature samples; M0 is the maximum mean discrepancy MMD matrix between the marginal probability distribution P S of the time-frequency domain statistical feature data in the source domain feature sample subset D S (W T x S ) and the marginal probability distribution P T of the time-frequency domain statistical feature data in the target domain feature sample subset D T (W T x T ).
[0111] The calculation of the maximum mean discrepancy MMD matrix M0 between the marginal probability distributions is shown in Equation (19):
[0112]
[0113] (B) Achieve the adaptation of the conditional probability distributions of the time-frequency domain statistical feature data in the source domain feature sample subset and the target domain feature sample subset, that is, Q S (y S |W T x S ) and Q T (y T |W T x T ) with the minimum maximum mean discrepancy MMD. The optimization objective expression is shown in Equation (20):
[0114]
[0115] In formula (20), and are respectively the number of samples of the c-th class in the source domain feature sample subset and the target domain feature sample subset; and are respectively the samples of the c-th class in the source domain feature sample subset and the target domain feature sample subset, and C is the total number of sample classes.
[0116] M c is the maximum mean discrepancy (MMD) matrix between the conditional probability distributions Q S (y S (y S |W T x S ) of the time-frequency domain statistical feature data in the source domain feature sample subset D T and the conditional probability distribution Q T (y T |W T x T ) of the time-frequency domain statistical feature data in the target domain feature sample subset D.
[0117] The maximum mean discrepancy (MMD) matrix M between the conditional probability distributions c is calculated as shown in formula (21):
[0118]
[0119] Based on the above two optimization objectives, the improved overall optimization objective of joint distribution adaptation can be obtained as shown in formula (22):
[0120]
[0121] In formula (22), the unification of the two distances in formulas (18) and (20) is achieved by c = 0, 1, 2, …, C. In formula (22), is the regularization term, λ is a trade-off parameter, and W T XHX T W = I is the constraint condition. Thus, the source domain feature sample subset and the target domain feature sample subset after joint distribution adaptation are obtained.
[0122] Reducing only the difference in marginal distributions does not guarantee a reduction in the overall distribution difference between domains, and the difference in conditional distributions still needs to be further considered. In fact, minimizing the conditional distributions Q S (y S |x S ) and Q T (y T |x T) The difference between them is crucial for achieving robust distribution adaptation. Matching the conditional distribution is very important, even by exploring the sufficient statistics of the distribution, because there is no labeled data in the target domain, that is QT(yT|xT) it cannot be directly modeled and calculated. Most methods require some labeled data in the target domain. Therefore, it is necessary to explore the pseudo-labels of the target data. By applying some basic classifiers trained on the labeled source domain data to the unlabeled target domain data, the pseudo-labels can be easily predicted.
[0123] In this embodiment, according to the above idea, the pseudo-labels of the subset of target domain feature samples in the conditional probability distribution are calculated based on the stacking ensemble learner.
[0124] As Figure 2 shown, the ensemble learning algorithm is to train a series of base models and integrate the output results of each model through a certain ensemble principle, so as to obtain a machine learning method with better performance than a single model. The ensemble principle of Stacking is to hierarchically combine multiple models, iteratively learn the classification bias of the upper-layer model, and improve the overall performance of the model. The Stacking algorithm can integrate different types of models, fuse the classification characteristics of various models, and the integration effect is often better; at the same time, the hierarchical structure of Stacking can further learn on the basis of the first-layer base model, train the meta-model, and finally output the results. The Stacking ensemble learning model generally has a two-layer structure. The first layer combines multiple base models with high classification performance and large differences, trains them on the original data set, and outputs the classification results of each model; the second layer combines the output results of the upper layer into new data features, trains a single meta-model on the newly constructed data set, and outputs the classification results.
[0125] In this embodiment, the stacking ensemble learner is trained using the time-frequency domain statistical feature data in the subset of labeled source domain feature samples, and then the trained stacking ensemble learning model is used to predict the class labels of the time-frequency domain statistical feature data in the subset of target domain feature samples. The obtained class labels are the pseudo-labels of the subset of target domain feature samples. Based on the data of the subset of target domain feature samples and the corresponding pseudo-labels, the conditional probability distribution is calculated.
[0126] Step 4: Use the SVM classifier as the fault diagnosis classifier, train the SVM classifier with the data in the subset of source domain feature samples after joint distribution adaptation obtained in Step 3, and then input the data in the subset of target domain feature samples after joint distribution adaptation obtained in Step 3 into the trained SVM classifier, and obtain the fault diagnosis result of the target domain through the SVM classifier.
[0127] The preferred embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. The embodiments described in the present invention are only descriptions of the preferred embodiments of the present invention, and do not limit the concept and scope of the present invention. Among the various specific technical features described in the above specific embodiments, they can be combined in any suitable manner without contradiction. As long as such a combination does not violate the idea of the present invention, it should also be regarded as the content disclosed in the present disclosure. To avoid unnecessary repetition, the present invention will not separately describe various possible combination methods.
[0128] The present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention and without departing from the design idea of the present invention, various modifications and improvements made by those skilled in the art to the technical solutions of the present invention should fall within the protection scope of the present invention. The technical content claimed by the present invention has been fully recorded in the claims.
Claims
1. A method for rotating machinery variable operating condition fault diagnosis based on domain adaptation features, characterized in that It includes the following steps: Step 1: Obtain the vibration signals of rotating machinery with labels under existing working conditions and the vibration signals of rotating machinery without labels under variable working conditions. Extract the time-frequency domain statistical feature data of the vibration signals of rotating machinery with labels under existing working conditions as the labeled source domain feature sample set, and extract the time-frequency domain statistical feature data of the vibration signals of rotating machinery without labels under variable working conditions as the unlabeled target domain feature sample set; Step 2: Based on the statistical feature data in the labeled source domain feature sample set, calculate the feature classification accuracy acc of each statistical feature data in the source domain feature sample set to characterize the discriminant performance of the features; Based on the statistical feature data in the normal state of the source domain feature sample set and the statistical feature data in the normal state of the target domain feature sample set, calculate the structural similarity index SSIM and FID score of each statistical feature data to characterize the domain invariance of the features; Based on the obtained feature classification accuracy acc, SSIM, and FID above, construct the feature quantization evaluation index for each statistical feature data Then set a threshold, and select multiple time-frequency domain statistical feature data with the feature quantization evaluation index Z greater than the set threshold from the source domain feature sample set to construct a labeled source domain feature sample subset, and select multiple time-frequency domain statistical feature data with the feature quantization evaluation index Z greater than the set threshold from the target domain feature sample set to construct an unlabeled target domain feature sample subset; Step 3: Perform joint distribution adaptation on the time-frequency domain statistical feature data in the source domain feature sample subset and the target domain feature sample subset obtained in Step 2 to obtain the source domain feature sample subset and the target domain feature sample subset after joint distribution adaptation; Step 4: Use the data in the source domain feature sample subset after joint distribution adaptation obtained in Step 3 to train the fault diagnosis classifier, and then input the data in the target domain feature sample subset after joint distribution adaptation obtained in Step 3 into the trained fault diagnosis classifier, and obtain the fault diagnosis result of the target domain through the fault diagnosis classifier.
2. The method for rotating machinery variable working condition fault diagnosis based on domain adaptation features according to claim 1, wherein In Step 1, the vibration signals of rotating machinery with labels under existing working conditions and the vibration signals of rotating machinery without labels under variable working conditions are respectively subjected to wavelet transform decomposition and reconstruction to obtain reconstructed signals, and then the time-domain statistical features of multiple statistical parameters are respectively extracted based on the reconstructed signals. Then, the frequency-domain statistical features of multiple statistical parameters are respectively extracted based on the calculation results of the Hilbert envelope spectrum of the reconstructed signals, so as to correspondingly obtain the time-frequency domain statistical feature data of the vibration signals of rotating machinery with labels under existing working conditions and the time-frequency domain statistical feature data of the vibration signals of rotating machinery without labels under variable working conditions.
3. The method for rotating machinery variable working condition fault diagnosis based on domain adaptation features according to claim 2, characterized in that The statistical parameters include mean, standard deviation, kurtosis, energy, energy entropy, peak factor, shape factor, skewness, extreme value, range, power spectrum entropy, singular spectrum entropy, approximate entropy, sample entropy, fuzzy entropy, permutation entropy, envelope entropy.
4. The method for rotating machinery variable working condition fault diagnosis based on domain adaptation features according to claim 1, characterized in that In Step 2, use the Xgboost classifier to calculate the feature classification accuracy acc of each time-frequency domain statistical feature data in the source domain feature sample set.
5. The method for rotating machinery variable working condition fault diagnosis based on domain adaptation features according to claim 1, characterized in that, When performing joint adaptation distribution in step 3, the maximum mean difference between the marginal probability distributions of the time-frequency domain statistical feature data in the source domain feature sample subset and the time-frequency domain statistical feature data in the target domain feature sample subset, and the maximum mean difference between the conditional probability distributions are both minimized as the total optimization objective of the joint adaptation distribution.
6. The method for rotating machinery variable working condition fault diagnosis based on domain adaptation features according to claim 5, characterized in that, In step 3, the stacking ensemble learning model is trained using the time-frequency domain statistical feature data in the labeled source domain feature sample subset, and then the trained stacking ensemble learning model is used to predict the class labels of the time-frequency domain statistical feature data in the target domain feature sample subset. The obtained class labels are the pseudo-labels of the target domain feature sample subset. Based on the time-frequency domain statistical feature data and the corresponding pseudo-labels in the target domain feature sample subset, the conditional probability distribution of the time-frequency domain statistical feature data in the target domain feature sample subset is calculated.
7. The method for rotating machinery variable working condition fault diagnosis based on domain adaptation features according to claim 1, characterized in that In step 4, the fault diagnosis classifier is an SVM classifier.
Citation Information
Patent Citations
New fault diagnosis method for rotating machinery based on deep confrontation convolutional neural network
CN114358124A
Dynamic joint distribution alignment network-based bearing fault diagnosis method under variable working conditions
US20230168150A1