A semi-supervised bearing compound fault classification method and system based on two-stage SVM and wavelet kernel diffusion
By employing a two-stage SVM and wavelet kernel diffusion method, and utilizing maximal discrete wavelet packet transform and wavelet kernel diffusion techniques, the problem of multi-class classification with few samples in bearing composite fault diagnosis was solved, achieving high-accuracy fault classification.
Patent Information
- Application Number
- CN202310397376.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-13
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2043-04-13
AI Technical Summary
Existing technologies struggle to effectively utilize limited sample data for multi-class classification in bearing composite fault diagnosis, and traditional machine learning methods lack incremental learning capabilities during semi-supervised processes, resulting in low classification accuracy.
A method based on two-stage SVM and wavelet kernel diffusion is adopted. The time-frequency features of bearing signals are extracted by maximal discrete wavelet packet transform, and wavelet kernel is used instead of Euclidean distance for label diffusion. Boundary samples are processed in the coarse and fine classification stages to improve classification accuracy.
It effectively improved the classification accuracy of bearing composite faults, reduced the number of training set samples, avoided insufficient incremental learning in the semi-supervised process, and improved the model's classification ability.
Smart Images

Figure CN116431989B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of fault diagnosis and state monitoring, and particularly relates to a semi-supervised bearing compound fault classification method and system based on two-stage SVM and wavelet kernel diffusion. BACKGROUND
[0002] In recent years, machine learning methods based on big data have made remarkable achievements in the field of mechanical fault diagnosis. Among them, support vector machine (SVM) is a machine learning method based on statistical learning theory, which can realize linear separability in high-dimensional feature space through kernel function transformation, and is widely used in intelligent fault diagnosis of bearings. Traditional machine learning to realize fault classification usually needs a large number of known samples, and the prediction results of the model are greatly affected by the size of the training set data. Supervised learning using only a small amount of sample data is difficult to achieve the expected effect of accurate classification. However, in practical applications, it is often difficult to obtain a large number of known sample data, and it is necessary to classify multiple compound faults. Therefore, it is of great significance to study a high-accuracy classification method suitable for small sample data and multiple bearing compound faults. SUMMARY
[0003] The purpose of the application is to study a high-accuracy classification method suitable for small sample data and multiple bearing compound faults, and to provide a bearing compound fault classification method and system based on two-stage SVM and wavelet kernel diffusion. A time-frequency feature of bearing signal based on maximum overcomplete discrete wavelet packet transform (MODWPT) and root mean square is proposed to replace the traditional time domain and frequency domain to represent the bearing signal. A wavelet kernel is used instead of the original kernel function, and a wavelet kernel diffusion method in high-dimensional space is proposed to replace the original Euclidean distance to label the known data in the test set, thereby improving the accuracy of the diffusion samples and reducing the number of samples in the training set. Finally, two-stage SVM is used to improve the classification accuracy of the model. Specifically, the boundary samples are obtained in the coarse classification stage, and the wavelet kernel diffusion samples and boundary samples are learned in the fine classification stage. This method can avoid the problem of not having incremental learning ability in the semi-supervised process, and has higher accuracy compared with other traditional popular algorithms.
[0004] The application achieves the above technical purpose through the following technical means.
[0005] A semi-supervised bearing compound fault classification method and system based on two-stage SVM and wavelet kernel diffusion includes the following steps:
[0006] Step S1, feature extraction: input the monitored signal, perform time-frequency feature extraction to obtain the bearing signal features;
[0007] Step S2, label propagation: divide the data set, and perform label propagation in the test set by using the known label information through an incremental kernel propagation method;
[0008] Step S3, class coarse division: perform kernel propagation on single-class known under one-to-many coding conditions, and obtain initial predicted labels using a lower penalty parameter. Then, reclassification is performed under a high penalty parameter, and boundary samples are obtained by comparing the prediction scores of the two times;
[0009] Step S4, class fine division: the boundary samples obtained in step S3 are given pseudo-labels using the Euclidean distance under low-dimensional conditions. Under one-to-one coding conditions, the classification of the test data is completed by using the pseudo-labels and the propagation label method.
[0010] In the above scheme, the time series X t is processed by using a time-frequency analysis method of a maximal discrete wavelet packet transform in the step S1
[0011]
[0012] wherein N is a sample capacity, mod represents a remainder of division of two numbers, j is a wavelet decomposition layer number, n=1, 2, …, 2 j indicates a feature point number. and is a scale filter and a wavelet filter of the jth layer, and the width of the filter is L j =(2 j -1)(L1-1)+1, L1 is an initial wavelet length, and the value of n of the jth layer is n=0, …, 2 j -1. When the remainder of n divided by 4 is 0 or 3, when the remainder of n divided by 4 is 1 or 2,
[0013] Further, the root mean square characteristic formula for calculating the time domain characteristics of the bearing signal in the step S1 is as follows:
[0014]
[0015] In the above scheme, the dual quadratic programming problem and the discriminant function constructed in the step S2 are as follows:
[0016]
[0017]
[0018] wherein N is a sample point number, y i indicates a class to which a sample x i belongs, and y j indicates a class to which a sample x jCategory to which it belongs, x n Let b be the support vector, c be the bias vector, C be the penalty factor, and α be the bias vector. i ,α j It is a Lagrange multiplier, K(x) i ·x j ) is the kernel function.
[0019] A kernel function is used to map data from low-dimensional linearly inseparable to high-dimensional separable data, and then the optimal classification surface is solved in the high-dimensional space. The Mexican hat wavelet kernel function, constructed from the Mexican hat wavelet basis functions, is as follows:
[0020]
[0021] Among them, a i Let a be the wavelet scaling parameter, and a i >0, where d is the length of the input vector x.
[0022] Furthermore, in step S2, a high-dimensional incremental nuclear diffusion method is used to perform tag diffusion on the test set. The specific steps of the incremental nuclear diffusion method are as follows:
[0023] Step S21: Apply the Mexican hat wavelet kernel function k(x) to the known label sample x. i ·x j High-dimensional mapping is performed to obtain the high-dimensional kernel matrix of the known samples;
[0024] Step S22: Obtain the kernel mean vector x by taking the mean. m N kernel vectors are obtained by performing a high-dimensional mapping between N test set data and the training set. For the kernel mean vector x m With N kernel vectors The calculation is as follows:
[0025]
[0026] Where m represents the number of iterations and the final number of diffused samples, and N represents the number of unknown samples. The minimum value is taken as the minimum kernel difference, and the corresponding vector is the minimum kernel difference vector;
[0027] Step S23: For the kernel vector with the minimum kernel difference... To expand the kernel matrix, if the original kernel matrix is X... m The new kernel matrix is then:
[0028]
[0029] Step S24: Based on the set number of expanded samples iter, repeat steps S1 to S3 to obtain the diffused sample sequence and the diffused kernel matrix.
[0030] Further, the coarse classification stage of the two-stage TSSVM in step S3 is specifically obtaining initial prediction labels using a lower penalty parameter under one-to-many coding conditions for kernel expansion of single-class known, and using the prediction results to reclassify under a high penalty parameter, and comparing the two prediction scores to obtain boundary samples. The specific steps are as follows:
[0031] Step S31, for common a-class faults, kernel expansion is performed on single-class faults as class one, and the remaining fault classes are taken as class two, and known sample data is used to complete prediction of the test set under a low penalty parameter to obtain prediction labels and the specific score Y1.score of the class is determined:
[0032]
[0033] wherein, α i is a Lagrange multiplier, and b is a bias vector;
[0034] Step S32, using the prediction labels obtained in S31 and the test set as the training set of the SVM, prediction is performed on itself under a high penalty parameter to obtain Y2.score;
[0035] Step S33, comparing the difference between the two prediction results, taking the first n samples with the largest difference as the boundary samples of the first classifier, and finally obtaining n×a boundary sample values under a-class fault conditions.
[0036] Further, the subdivision stage of the TSSVM in step S4 is specifically first assigning pseudo-labels to the obtained boundary samples using the Euclidean distance under low-dimensional conditions, and then performing kernel expansion on double-class known samples under one-to-one coding conditions, and using the pseudo-labels of the samples as the training samples of the SVM, and finally obtaining the prediction results of the test set.
[0037] A system for implementing the semi-supervised bearing composite fault classification method and system based on two-stage SVM and wavelet kernel expansion, comprising a feature extraction module, a label diffusion module, a class coarse classification module and a class subdivision module connected in sequence;
[0038] The feature extraction module is used for inputting the monitored signal in bearing fault diagnosis, performing time-frequency feature extraction to obtain bearing signal features;
[0039] The label diffusion module is used for dividing the data set, and using known label information to perform label diffusion in the test set by the wavelet kernel expansion method;
[0040] The category coarse division module is used for nuclear diffusion of single-class known under one-to-many coding conditions, and an initial prediction label is obtained using a lower penalty parameter.
[0041] The category subdivision module is used for pseudo-label assignment of the obtained boundary samples using the Euclidean distance under low-dimensional conditions. Under one-to-one coding conditions, the classification of the test data is completed using the pseudo-label and the diffusion label method.
[0042] Compared with the prior art, the beneficial effects of the present application are that the present application provides a semi-supervised bearing composite fault classification method and system based on two-stage SVM and wavelet kernel diffusion, and the time-frequency features of the bearing signal using the maximum discrete wavelet packet transform (MODWPT) and the root mean square are used instead of the traditional time domain and frequency domain to represent the bearing signal. The original kernel function is replaced with a wavelet kernel, and on this basis, a wavelet kernel diffusion method in a high-dimensional space is proposed to replace the original Euclidean distance to diffuse the known data in the test set. Finally, the two-stage SVM is used to improve the classification accuracy of the model, specifically, the boundary samples are obtained in the coarse division stage, and the wavelet kernel diffusion samples and the boundary samples are learned in the subdivision stage, and finally the prediction result of the test set is obtained. This method can effectively improve the accuracy of the diffusion samples, reduce the number of samples in the training set, and avoid the problem of not having incremental learning ability in the semi-supervised process.
[0043] To verify the feasibility and method advantages of the present application, the inter-class distance comparison of the proposed time-frequency feature method and other features, the visualization analysis of the incremental kernel diffusion method, and the comparison with the Euclidean distance and geodesic distance diffusion are performed. The self-made test bench data and the public data set are used to verify the semi-supervised bearing composite fault classification method and system based on two-stage SVM and wavelet kernel diffusion, and the proposed method is compared with the SVM and convolutional neural network methods. The experimental results show that the method can improve the linear separability of the composite bearing fault and improve the classification accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 is a general flowchart of an embodiment of the present application;
[0045] Figure 2 is a time-frequency feature extraction flowchart of an embodiment of the present application;
[0046] Figure 3 is an X n The root mean square time-frequency feature map of the wavelet packet coefficients generated by five-layer MODWPT is shown in the following table: Figure 3 (a) is the root mean square time-frequency feature map of the normal bearing, Figure 3 (b) is the root mean square time-frequency feature map of the outer ring fault,Figure 3 (c) is the inner ring fault root mean square time-frequency feature map, Figure 3 (d) is the rolling element fault root mean square time-frequency feature map;
[0047] Figure 4 is the intra-class and inter-class distance map of seven types of faults in the four types of features of five-layer wavelet packet coefficient root mean square, kurtosis, waveform factor and time domain feature, wherein Figure 4 (a) is the intra-class distance map of seven types of faults, Figure 4 (b) is the inter-class distance map of seven types of faults;
[0048] Figure 5 is the incremental kernel diffusion result map of an embodiment of the present application, wherein the single-class training set of faults in group A is 10, and the test set is 100 as shown in Figure 5 A(1), Figure 5 A(2) is the label map after diffusion when iter=20 in the incremental kernel diffusion, Figure 5 A(3) is the label map after diffusion when iter=40, the single-class training set in group B is 5, and the test set is 100 as shown in Figure 5 B(1), Figure 5 B(2) is the label after diffusion when iter=20, Figure 5 B(3) is the label map after diffusion when iter=40, the single-class training set in group C is 10, and the test set is 100 as shown in Figure 5 C(1), Figure 5 C(2) is the label map after diffusion when iter=20, Figure 5 C(3) is the label map after diffusion when iter=40, wherein the gray represents unknown labels;
[0049] Figure 6 is the TSSVM classification result map of an embodiment of the present application, wherein Figure 6 (a) is the TSSVM classification result map of group A with 6 training samples, Figure 6 (b) is the TSSVM classification result map of group A with 10 training samples, Figure 6 (c) is the TSSVM classification result map of group C data under single-class 5 training samples;
[0050] Figure 7 is a schematic diagram of an experimental platform and a sensor installation method of an embodiment of the present application. DETAILED DESCRIPTION
[0051] Embodiments of the present application are described below in detail with reference to the accompanying drawings, wherein the same or similar components have the same or similar designations throughout the various figures. The embodiments described below are exemplary and are intended to be illustrative of the present application, and are not to be construed as limiting the present application.
[0052] In this embodiment, preferably, in order to verify the effectiveness and practicability of the present application, the data of the Fan Wei research group and the operation data of the self-made experimental platform are used to perform experimental verification of the incremental kernel diffusion method, and the bearing vibration data under the conditions of inner ring failure, outer ring failure, roller failure, inner ring roller failure, outer ring roller failure, inner ring and outer ring failure, and normal condition are used.
[0053] Figure 1 The preferred embodiment of the semi-supervised bearing compound fault classification method and system based on two-stage SVM and wavelet kernel diffusion is shown, and the semi-supervised bearing compound fault classification method and system based on two-stage SVM and wavelet kernel diffusion comprises the following steps:
[0054] Step S1, feature extraction: input the monitored signal, perform time-frequency feature extraction to obtain bearing signal features;
[0055] Step S2, label diffusion: divide the data set, and perform label diffusion in the test set by using the known label information through the wavelet kernel diffusion method;
[0056] Step S3, class coarse division: kernel diffusion is performed on single-class known under one-versus-all coding conditions, and an initial prediction label is obtained using a lower penalty parameter. Then, reclassification is performed under a high penalty parameter, and boundary samples are obtained by comparing the prediction scores of the two times;
[0057] Step S4, class subdivision: the boundary samples obtained in step S3 are assigned with pseudo-labels using the Euclidean distance under low-dimensional conditions. Under one-versus-one coding conditions, the classification of the test data is completed using the pseudo-labels and the diffusion label method.
[0058] According to this embodiment, preferably, the data source in step S1 is 200 groups of seven types of fault data collected under the condition of simulated tension of 4.5 KN and rotation speed of 500 r / min from 9 bearing rolling elements of a self-made experimental platform model 6205, and 100 groups are selected from each type of fault data as a test set. The time series X t The jth layer wavelet packet coefficient can be expressed as:
[0059]
[0060] wherein, L jN is the sample size, mod means the remainder of division, j is the number of wavelet decomposition layers, n=1,2…,2 j The number of feature points is represented.
[0061] The root mean square characteristic formula of the time domain characteristics of the bearing signal is calculated as follows:
[0062]
[0063] Let j=5, that is, the root mean square value of the fifth layer wavelet packet coefficient after 5 layers of MODWPT decomposition is used as the time-frequency feature, and the intra-class and inter-class distances of the root mean square value of the 5-layer MODWPT wavelet packet coefficient, kurtosis, waveform factor and time domain characteristics are compared to verify the effect of the method t. The experiment is as shown in Figure 4 , wherein Figure 4 (a) is the intra-class distance diagram of the seven types of faults, Figure 4 (b) is the inter-class distance diagram of the seven types of faults, and Figure 4 It can be observed that the root mean square feature of the 5-layer MODWPT has the highest inter-class distance and smaller intra-class distance, and therefore it can be used as a bearing signal feature for fault classification algorithm and is superior to other features.
[0064] According to the embodiment, preferably, the high-dimensional incremental kernel diffusion method is used in step S2 to perform label diffusion in the test set. The specific steps of the high-dimensional incremental kernel diffusion method are as follows:
[0065] (1) The known label sample x is subjected to high-dimensional mapping by using the Mexican hat wavelet kernel function k(x i ·x j ) to obtain a high-dimensional kernel matrix of the known sample;
[0066] (2) The mean value is obtained to obtain a kernel mean vector x m , and N test set data and the training set are subjected to high-dimensional mapping to obtain N kernel vectors The kernel mean vector x m and the N kernel vectors are calculated as follows:
[0067]
[0068] Wherein, m represents the number of iterations and the final number of diffusion samples, and N represents the number of unknown samples. The minimum value is the minimum kernel difference, and the corresponding vector is the minimum kernel difference vector;
[0069] (3) The kernel matrix of the minimum kernel difference in the kernel vector is expanded, and if the original kernel matrix is X m , then the new kernel matrix is:
[0070]
[0071] (4) According to the set number of extended samples iter, steps (1) to (3) are repeated to obtain a diffusion sample sequence and a diffusion kernel matrix.
[0072] According to the embodiment, preferably, the dual quadratic programming problem constructed in step S2 and the discriminant function are as follows:
[0073]
[0074]
[0075] wherein, wherein N is the number of sample points, y i represents the sample x i belonging to the category, x n is a support vector, b is a bias vector, C is a penalty factor, and α i ,α j is a Lagrange multiplier, and K(x i ·x j ) is a kernel function.
[0076] The data is mapped from low-dimensional linearly inseparable to high-dimensional separable by the kernel function, and then an optimal classification surface is solved in the high-dimensional space. A Mexican hat wavelet kernel function constructed by a Mexican hat wavelet basis function is as follows:
[0077]
[0078] wherein, a i is a wavelet scale parameter, and a i > 0.
[0079] Further verification of the effectiveness of the incremental kernel diffusion method, using a self-made experimental table to obtain seven types of bearing fault vibration data, each of which has 200 groups, and 100 groups of which are test sets. Denote the kurtosis feature group as group A, the time domain feature group as group B, and the root mean square feature group as group C. After dimensionality reduction of the obtained three groups of bearing compound fault features by principal component analysis (PCA), visual analysis is performed. The incremental kernel diffusion result is as shown in Figure 5 , wherein the fault single-class training set in group A is 10, and the test set is 100 as shown in Figure 5 A(1), Figure 5 A(2) is the label graph after diffusion when iter = 20 in the incremental kernel diffusion, Figure 5 A(3) is the label graph after diffusion when iter = 40, the single-class training set in group B is 5, and the test set is 100 as shown in Figure 5 B(1), Figure 5 B(2) is the label after diffusion when iter = 20, Figure 5B(3) is the label map after diffusion when iter = 40, the single-class training set in group C is 10, and the test set is 100 Figure 5 C(1) is shown in Figure 5 C(2) is the label map after diffusion when iter = 20, Figure 5 C(3) is the label map after diffusion when iter = 40, wherein the gray represents unknown labels. The experimental results show that the diffusion accuracy of the root mean square feature group is 97.10% and 95.00%, which is higher than that of other groups.
[0080] According to the embodiment, preferably, the coarse classification stage of the TSSVM (Two-Stage SVM, TSSVM) in step S3 is specifically as follows:
[0081] (1) There are a types of faults, and the kernel diffusion of one type of fault is taken as class one, and the remaining fault classes are taken as class two. The known sample data is used to complete the prediction of the test set under a low penalty parameter to obtain a predicted label and the specific score Y1.score of the class is determined:
[0082]
[0083] In the formula, α i is a Lagrange multiplier, and b is a bias vector;
[0084] (2) The predicted label obtained in (1) is used and the test set is taken as the training set of the SVM to predict itself under a high penalty parameter and obtain Y2.score;
[0085] (3) The difference between the two prediction results is compared, the first n samples with the largest difference value are taken as the boundary samples of the first classifier, and n×a boundary sample values are finally obtained under the condition of a types of faults. In the embodiment, n = 20 and a = 6, that is, 120 boundary sample values are finally obtained.
[0086] According to the embodiment, preferably, the subdivision stage of the TSSVM in the step S4 respectively performs kernel diffusion of the single class and the remaining faults to iter-x1 to ensure the balance of each class of samples when constructing the first classifier. The diffused data is used to predict the test set under the penalty parameter, and 15xN prediction results are obtained under the six-class condition, and the final prediction result is obtained through multi-class voting. The TSSVM, LibSVM and SVM with incremental kernel diffusion, namely the TSSVM subdivision stage, are used to verify three groups of bearing composite fault feature data with different samples and kernel functions. The penalty parameter C is taken as 200, iter is taken as 40, and the number of single-class test set samples is 100, and there are 700 groups in total. The following conclusions can be obtained from the experimental results: 1. Under the condition of fewer training samples, the incremental kernel diffusion classification can reach an accuracy of 98%, which is obviously better than the SVM at the same stage. 2. The results show that the classification effect of the TSSVM is obviously improved compared with the original SVM, and the classification effect is improved by 7.5% compared with the SVM under the same conditions. The TSSVM classification result of the application is shown in Figure 6 Figure 6 (a) is a TSSVM classification result diagram of a group of 5 training samples A, Figure 6 (b) is a TSSVM classification result diagram of a group of 10 training samples A, Figure 7 (c) shows the TSSVM result of group C data under 5 training samples, and the prediction accuracy is 100%.
[0087] According to the embodiment, preferably, the method in the step S4 reaches an accuracy of 100% in the seven-class bearing composite fault classification. Now the method proposed in the application is compared with the convolutional neural network CNN. The 2500 collection points of the signal data in group A are taken as a single sample, and there are 700 groups of data in total. 500 groups are taken as a training set, and 200 groups are taken as a test set. After 20 iterations, the optimal result is reached. The loss rate of the training set is 0.0760, the accuracy is 0.9720, the loss rate of the test set is 0.1359, and the accuracy of the test set is 0.9450, which is lower than the classification result of the TSSVM method proposed in the application.
[0088] Therefore, by the above steps, the semi-supervised bearing composite fault classification method and system based on the two-stage SVM and wavelet kernel diffusion can avoid the problem that the semi-supervised process does not have the incremental learning ability, and can improve the linear separability of the composite bearing fault, thereby improving the classification accuracy, and providing a new method for rolling bearing fault diagnosis. The experimental platform configuration is shown in Table 1.
[0089] Table 1 Hardware platform configuration
[0090]
[0091]
[0092] A system for implementing the semi-supervised bearing compound fault classification method and system based on two-stage SVM and wavelet kernel diffusion, comprising a feature extraction module, a label diffusion module, a class coarse division module and a class subdivision module;
[0093] The feature extraction module is used for inputting the monitored signal in bearing fault diagnosis, performing time-frequency feature extraction to obtain bearing signal features;
[0094] The label diffusion module is used for dividing the data set, and performing label diffusion in the test set by using the known label information through the wavelet kernel diffusion method;
[0095] The class coarse division module is used for kernel diffusion of single-class known under one-to-many coding conditions, and an initial prediction label is obtained by using a lower penalty parameter. Then, reclassification is performed under a high penalty parameter, and boundary samples are obtained by comparing the prediction scores of the two times;
[0096] The class subdivision module is used for pseudo-label assignment of the obtained boundary samples under low-dimensional conditions by using the Euclidean distance. Under one-to-one coding conditions, the classification of the test data is completed by using the pseudo-label and the diffusion label method.
[0097] The above modules are realized by software programming in MATLAB, and the implementation process refers to the step content of the method part.
[0098] In summary, the application provides a semi-supervised bearing composite fault classification method and system based on two-stage SVM and wavelet kernel diffusion, which uses maximum discrete wavelet packet transform (MODWPT) and the time-frequency features of bearing signals with root mean square to replace the traditional time domain and frequency domain for representing bearing signals. The original kernel function is replaced with a wavelet kernel, and on this basis, a wavelet kernel diffusion method in high-dimensional space is proposed to replace the original Euclidean distance to label the known data in the test set. Finally, two-stage SVM is used to improve the classification accuracy of the model, specifically, the boundary samples are obtained in the coarse classification stage, and the wavelet kernel diffusion samples and boundary samples are learned in the fine classification stage, and finally the prediction results of the test set are obtained. To verify the feasibility and method advantage of the application, the inter-class distance comparison of the proposed time-frequency feature method and other features, the visualization analysis of the incremental kernel diffusion method, and the comparison with Euclidean distance and geodesic distance diffusion are performed. The self-made test bench data and public data set are used to verify the effectiveness of the semi-supervised bearing composite fault classification method and system based on two-stage SVM and wavelet kernel diffusion, and the proposed method is compared with SVM, convolutional neural network and other methods. The experimental results show that the method can reduce the number of samples in the training set, avoid the problem of not having incremental learning ability in the semi-supervised process, improve the linear separability of the composite bearing fault, and improve the classification accuracy.
[0099] The semi-supervised bearing composite fault classification method and system based on two-stage SVM and wavelet kernel diffusion provided by the application are described in detail above. In this paper, specific examples are applied to briefly describe the basic principles and implementation methods of the application, but the protection scope of the application is not limited thereto. It can be understood by those skilled in the art that various changes, modifications, replacements and variations of the examples in this paper can be made without departing from the principles and spirits of the application, and the scope of the application is defined by the appended claims and their equivalents.
[0100] It should be understood that although the present specification is described in terms of various embodiments, not every embodiment contains only one independent technical solution, and the description manner of the specification is only for clarity, those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementation manners that can be understood by those skilled in the art.
[0101] The series of detailed descriptions listed above are only specific descriptions of the feasible embodiments of the application, and are not used to limit the protection scope of the application, and equivalent embodiments or changes made without departing from the spirits of the technical art of the application should be included in the protection scope of the application.
Claims
1. A semi-supervised bearing compound fault classification method based on two-stage SVM and wavelet kernel diffusion, characterized in that, The method comprises the following steps: Step S1, feature extraction: input the monitored signal, perform time-frequency feature extraction to obtain bearing signal features; Step S2, label diffusion: divide the data set, and perform label diffusion in the test set by using known label information through an incremental kernel diffusion method; The specific process of step 2 is as follows: Step S21, mapping high dimension for known label samples x by using the Mexican hat wavelet kernel function k x i ▪ x j to obtain a high-dimensional kernel matrix of the known samples Step S22, taking the mean to obtain the kernel mean vector x m , the test set data is mapped to the high dimension space by the kernel function, and the kernel vector of the test set data is obtained N N ; core mean vector x m with N core vector is calculated as follows: (6); wherein, m represents the number of iterations and the final number of diffusion samples, N represents the number of unknown samples; the minimum value is the minimum core difference, and the corresponding vector is the minimum core difference vector; Step S23, the core vector with the minimum core difference The core matrix is expanded, if the original core matrix is X m The new core matrix is: (7); Step S24, according to the set number of extended samples iter , repeating steps S1 to S3, obtaining the diffusion sample sequence and the diffusion matrix. Step S3, class coarse division: perform kernel diffusion on single-class known under one-to-many encoding conditions, and obtain initial predicted labels using a lower penalty parameter; then, reclassification is performed under a high penalty parameter, and boundary samples are obtained by comparing the prediction scores of the two times; Step S4, class subdivision: the boundary samples obtained in step S3 are given pseudo-labels using the Euclidean distance under low-dimensional conditions; under one-to-one encoding conditions, the pseudo-labels and diffusion label method are used to complete the classification of the test data.
2. The method of claim 1, wherein, When the time-frequency feature extraction in the step S1 is performed to obtain the bearing signal feature, a maximum discrete wavelet packet transform in a time-frequency analysis method is used to process a time sequence as X t , the first j layer wavelet packet coefficient can be expressed as: (1); wherein, t =0, …, N -1; is a decomposition coefficient of the discrete wavelet transform, N is a sample size, mod represents a remainder of division of two numbers, j is a wavelet decomposition level, n =1, 2, …, 2 j represents a number of feature points, and the width of the filter is L j =(2 j -1)( L 1-1)+1; L 1 is an initial wavelet length, and the value of the first j layer n is n =0, …, 2 j -1; when the remainder of division of n by 4 is 0 or 3, the filter parameter ; when the remainder of division of n by 4 is 1 or 2, ; and are a scale filter and a wavelet filter of the first j layer; The formula for calculating the wavelet packet root mean square feature of the bearing signal time domain feature in step S1 is as follows: (2) 。 3. The method of claim 1, wherein, The dual quadratic programming problem and the discriminant function are used in the incremental kernel diffusion method in step S2, the data is mapped from low-dimensional linearly inseparable to high-dimensional separable through a kernel function, and then the optimal classification surface is solved in the high-dimensional space through dual quadratic programming, the dual quadratic programming problem and the discriminant function are as follows: (3), (4); wherein, N is the number of sample points, y i represents a sample x i the class to which it belongs, y j represents a sample x i the class to which it belongs, x n is the support vector, b is the bias vector, C is the penalty factor, α i , α j is the Lagrange multiplier, K x i ▪ x j is the kernel function. 4. The method of claim 3, wherein, The Mexican hat wavelet kernel function constructed by the Mexican hat wavelet basis function is as follows: (5); wherein a i is a wavelet scale parameter, and a i > 0, d is the length of the input vector x .
5. The method of claim 1, wherein, The specific steps of step S3 are as follows: Step S31, common a Class fault, a class of faults is nuclear proliferation as class one, and the remaining fault classes are as class two. The prediction label is obtained by completing the prediction of the test set under the low penalty parameter with the known sample data And the specific score of judging its class Y 1. score : (8); wherein α i is a Lagrange multiplier, b is a bias vector; Step S32, using the prediction labels obtained in S31 with the test set as the training set for the SVM, predicting itself at a high penalty parameter and obtaining Y 2. score ; Step S33, comparing the difference of the two prediction results, taking the first sample with the largest difference as the boundary sample of the first classifier, and finally obtaining n boundary sample values under the fault condition a . n × a 6. The method of claim 1, wherein, The class subdivision stage of step S4 is specifically: first, pseudo-labels are given to the obtained boundary samples using the Euclidean distance under low-dimensional conditions, and then kernel diffusion is performed on double-class known samples under one-to-one encoding conditions, and the pseudo-labels of the samples are used as the training samples of the SVM, and finally the prediction results of the test set are obtained.
7. A system for implementing the semi-supervised bearing compound fault classification method based on two-stage SVM and wavelet kernel diffusion according to any one of claims 1-6, characterized in that, The method comprises the following steps: The feature extraction module is used in bearing fault diagnosis to input the monitored signal, perform time-frequency feature extraction to obtain bearing signal features; The label diffusion module is used to divide the data set, and perform label diffusion in the test set by using known label information through a wavelet kernel diffusion method; The class coarse division module is used to perform kernel diffusion on single-class known under one-to-many encoding conditions, and obtain initial predicted labels using a lower penalty parameter; then, reclassification is performed under a high penalty parameter, and boundary samples are obtained by comparing the prediction scores of the two times; The class subdivision module is used to give pseudo-labels to the obtained boundary samples using the Euclidean distance under low-dimensional conditions; under one-to-one encoding conditions, the pseudo-labels and diffusion label method are used to complete the classification of the test data.
Citation Information
Patent Citations
Data annotation processing method and device, storage medium and electronic device
CN114548195A
Data labeling processing method and apparatus, and storage medium and electronic apparatus
WO2022111284A1