A rolling bearing fault diagnosis method based on a semi-supervised learning framework
By introducing 1DCrop data augmentation and the MixMatch semi-supervised framework, combined with consistency regularization and MixUp technology, the problem of insufficient data and labels in rolling bearing fault diagnosis is solved, achieving efficient and low-cost fault diagnosis and improving diagnostic accuracy and model generalization ability.
Patent Information
- Application Number
- CN202310267340.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-17
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-03-17
AI Technical Summary
Existing deep learning models rely on a large amount of labeled data for rolling bearing fault diagnosis. However, due to limitations in the actual engineering environment and the deployment of sensors, they are difficult to fully utilize. Furthermore, existing semi-supervised learning methods do not perform well on one-dimensional data, especially when both data and labels are scarce, resulting in a significant drop in diagnostic accuracy.
The 1DCrop data augmentation algorithm is introduced to process unlabeled data. Combined with the MixMatch semi-supervised framework, the algorithm uses consistency regularization and MixUp technology to reduce entropy using a sharpening function, construct pseudo-labels, form a new loss function to train the model, and make full use of limited data for diagnosis.
With minimal data and labels, it achieves efficient rolling bearing fault diagnosis, improves data utilization and diagnostic accuracy, reduces costs, and has good generalization ability and robustness.
Smart Images

Figure CN116164966B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of rolling bearing fault diagnosis, and relates to a rolling bearing fault diagnosis method based on a semi-supervised learning framework. Background Technology
[0002] Rolling bearings are crucial components of rotating machinery widely used in engineering. However, due to their complex working environment and metal fatigue caused by prolonged operation, bearings are subjected to alternating and impact loads for extended periods, making them prone to wear and fracture, leading to operational failures. Statistics show that approximately 30%-40% of equipment failures are caused by bearing malfunctions. This often results in damage to large machinery and personal injury. Therefore, research on rolling bearing fault diagnosis is of great significance.
[0003] Methods for diagnosing rolling bearing faults, besides traditional palpation and auscultation, include time-domain analysis (Gini index), frequency-domain analysis (Fourier envelope spectrum analysis), and time-frequency domain analysis (short-time Fourier transform) of vibration signals, as well as the widely used traditional machine learning (Support Vector Machine, Naive Bayes classifier, etc.) and deep learning (Convolutional Neural Network, Ladder Network). However, palpation and auscultation rely on rich human experience and their accuracy is not ideal; analyzing vibration signals in the time, frequency, and time-frequency domains requires extracting effective indicators from complex signals, leading to poor results; in traditional machine learning methods, feature extraction requires human intervention and is also limited to some extent by human intervention. Deep learning algorithms, compared to traditional machine learning algorithms, achieve automated feature extraction, improving accuracy and efficiency. Zhang used convolutional neural networks to directly analyze and classify time-domain signals, avoiding information loss and eliminating the need to analyze the internal mechanisms of mechanical devices. Xu Di et al. proposed a quantum genetic algorithm-optimized SVM rolling bearing fault diagnosis method by collecting time-domain and frequency-domain features from vibration signals. Wang Bin used wavelet analysis to decompose the denoised signal into different frequency bands, extracted signal features from each frequency band to form a fault feature vector, and established a rolling bearing fault diagnosis system using a BP neural network as the bearing fault diagnosis network. Yu Zhongqing, after comparing C4.5, Cart, BP, and SVM, concluded that SVM had the best classification effect and further improved the key feature selection of vibration signals.
[0004] The success of deep learning largely depends on large amounts of data. We call data with known fault types labeled data, and fault types labels. However, in practical engineering, due to limitations in the working environment and sensor deployment, acquiring large amounts of labeled data is extremely costly, making it difficult to fully utilize deep learning models. Models that utilize a small portion of labeled data and a large portion of unlabeled data to complete classification tasks are called semi-supervised learning models. Zhang proposed a small-sample fault diagnosis method based on dual-path convolutional attention (DCA) and bidirectional gated recurrent units (DCA-bigru). He used BiGRU to achieve spatiotemporal feature fusion and extracted vibration signal fusion features with attention weights through DCA. Furthermore, applying global average pooling (GAP) to dimensionality reduction and fault diagnosis gives the model good generalization ability and robustness. Ding chose to use a probabilistic mixture model to first fit the bearing signal and then sample to expand the dataset, and used a semi-supervised ladder network to complete the classification, which alleviated the problem of insufficient label information in the bearing dataset to some extent. However, the above methods still have low data utilization, and the diagnostic accuracy will significantly decrease when both the dataset and labeled data are extremely limited. Summary of the Invention
[0005] The technical problem this invention aims to solve is to introduce the MixMatch semi-supervised framework for the first time into the fault diagnosis of rolling bearings and improve its data augmentation methods to overcome the difficulty of applying the MixMatch semi-supervised framework to one-dimensional data, thereby fully utilizing its performance to complete the fault diagnosis task of rolling bearings with very limited data and labels. Currently, semi-supervised learning has been widely applied and developed primarily in the image domain. In recent years, algorithms such as the MixMatch semi-supervised framework, which integrates most mainstream semi-supervised algorithms, have emerged and achieved excellent results on most image datasets. However, its application is still limited to the image domain, and it has not yet been widely applied in one-dimensional domains such as vibration signals. This is mainly because image domain data augmentation methods involving cropping and random augmentation are not suitable for one-dimensional signals, as they destroy the overall temporal information of the signal itself, leading to poor performance and thus hindering the application of MixMatch and its subsequent frameworks in the one-dimensional domain. One-dimensional data augmentation methods have recently been developed, and determining a suitable data augmentation method is key to applying the MixMatch semi-supervised framework to the one-dimensional domain. This invention uses the 1DCrop method for bearing vibration signals as a data augmentation method to complete the one-dimensionalization of the MixMatch semi-supervised framework, enabling it to exhibit good performance in the one-dimensional domain. The efficiency and simplicity of this semi-supervised framework are used to complete classification and diagnosis tasks with very little data and labels.
[0006] The technical solution of this invention:
[0007] Introduce the 1D Crop data augmentation algorithm to augment the unlabeled data while preserving the transient shock characteristics. Then, based on the idea of consistency regularization, encourage the model to produce the same output when the input is perturbed. After the classifier guesses the label, use the sharpening function to minimize the entropy. Then, perform Mixup processing on the obtained unlabeled data and labeled data, and use the weighted sum of the cross-entropy of the new unlabeled data and the L2 loss term of the new labeled data as the loss function. This ensures that entropy is reduced while maintaining consistency and compatibility with traditional regularization techniques, and can efficiently use limited data to complete classification.
[0008] The specific steps are as follows:
[0009] (1) Vibration signal data acquisition: Acquire the vibration signals of all states at the bearing drive end as unlabeled data, and acquire a small amount of labeled vibration signals as labeled data. Use these unlabeled data and labeled data as the dataset of the model.
[0010] (2) Vibration signal data preprocessing: To make full use of the data, when dividing a segment of vibration data into samples, adopt the commonly used enhanced sliding window sampling method for one-dimensional data to preprocess the data, such as Figure 1 .
[0011] The length of the entire segment of data is N, and a series of samples x with a length of n are sampled from it i (i = 1, 2,...), where n should be greater than the number of samples collected in two rotations of the bearing to ensure complete features. The starting point of each sampling is shifted backward by p (0 < p < n). This can make full use of the data and retain features on the premise of ensuring reliable and complete data.
[0012] (3) Send the data into the 1D MixMatch semi-supervised framework for training to obtain the model and save it: As Figure 2 is the improved 1D MixMatch semi-supervised framework of the present invention.
[0013] Among them, U, U′ represent unlabeled data, u, represents the sample point of unlabeled data, X, X′ represent labeled data, x represents the sample point of labeled data, p represents the one-hot label of labeled data, ^q0, q, is the guessed label of unlabeled data, K represents the number of times of data augmentation, T is the hyperparameter of the sharpening function, used to control the low entropy degree of the guessed label, and α is the hyperparameter of Mix up, used for the mixing degree of labeled data and unlabeled data.
[0014] This framework mainly consists of four steps: data augmentation, label guessing, label processing, and MixUp. Labeled data X=(x,p) and unlabeled data U=(u) are transformed into X′ and U′ after these steps. These two data points are then used as new datasets to train the network, with the weighted sum of the cross-entropy of X′ and the L2 loss of U′ as the loss function. The network is trained to achieve optimal performance for rolling bearing fault diagnosis.
[0015] 1) Data Augmentation: A segment of sample points u in the input unlabeled data U is transformed into K new sample points through K random augmentations (specifically including random reversal, 1DCrop, and adding Gaussian noise). In this process, the data undergoes a random reversal process with a certain probability of having their order changed, i.e., u = {u1, u2, ..., u...} n After reversal After passing through the Gaussian noise addition stage, the entire sample point will be superimposed with a segment of Gaussian white noise; after passing through the 1DCrop stage, the data will be randomly split into two segments and their order will be swapped, i.e., the input data u={u1,u2,...,u n}, with a certain data point u m As a boundary, the input sample is divided into two segments, u. f ={u1,...,u m-1} and u b ={u m ,...,u n Then, the order of the samples was rearranged to form a new sample. The choice of m must satisfy equation (1):
[0016] ρn <m<(1-ρ)n (1)
[0017] Where ρ is the piecewise parameter (empirically taken to be below 0.2). To preserve transient characteristics, the selection of ρ should satisfy equations (2) and (3):
[0018]
[0019]
[0020] Among them, f s f i f o f b These are the sampling frequency, the bearing inner ring failure frequency, the bearing outer ring failure frequency, and the rolling element failure frequency, respectively. These frequencies can be calculated using empirical formulas (4), (5), and (6):
[0021] f i =0.6*Z*f r (4)
[0022] f o =0.4*Z*f r (5)
[0023]
[0024] Among them, f r Z represents the bearing frequency, and Z represents the number of rolling elements.
[0025] The purpose of the above data augmentation is to further mine information from limited data and generate new data to improve the generalization of the model. It is also to ensure that the MixMatch semi-supervised framework can complete the core idea of consistency regularization on one-dimensional vibration data, that is, the classifier should output the same class distribution even after the data is expanded.
[0026] 2) Guessing labels: Using K data-enhanced sample points... The guessed label q0 is obtained by inputting each input into the classifier. 1 ,...,q0 K The classifier described above is an arbitrary neural network, and its parameters are consistent with the network used by the framework after data processing, trained on datasets X′ and U′. The purpose of guessing this series of labels is to facilitate MixUp's recombination of labeled data X and unlabeled data U, and to provide a basis for finding a reliable pseudo-label for the unlabeled data.
[0027] 3) Label processing: Obtain the guessed label q0 1 ,...,q0 K , where q0 k =[i1,...,i L [,k∈(1,K), L is the number of categories. To determine the pseudo-labels for unlabeled data, the guessed labels are averaged to obtain...] As shown in equation (7):
[0028]
[0029] This framework tends to make low-entropy predictions to provide a pseudo-label with a clearer category, so it uses the Sharpen function to implicitly... After completing the entropy minimization process, q is obtained, as shown in equations (8) and (9):
[0030]
[0031]
[0032] Where L is the number of categories, i is one of the L possible labels, and T is the temperature parameter. When T approaches 0, q will be close to a one-hot distribution, so reducing the temperature T will encourage the model to produce low-entropy predictions.
[0033] 4) MixUp: After the above processing, we obtain unlabeled data U = (u, q) and labeled data X = (x, p), where x and u are sample points, p is the label of the labeled data, and q is the pseudo-label of the unlabeled data. Then, to make full use of the data, this framework uses MixUp to mix the labeled and unlabeled data. To ensure that the data and labels are compatible with the loss term, modified versions of MixUp are used, as shown in equations (10), (11), (12), and (13):
[0034] λ~Beta(α,α)(10)
[0035] λ′=max(λ,1-λ) (11)
[0036] x′=λ′x1+(1-λ′)x2 (12)
[0037] p′=λ′p1+(1-λ′)p2 (13)
[0038] Where λ is a parameter following a Beat distribution, used to control the degree of mixing, α is a hyperparameter, and x and p represent the sample points and labels, respectively. In the MixUp operation, this framework first combines and shuffles the labeled data X and unlabeled data U to form W, and then lets X' i =MixUp(X i W i ),U′ i =MixUp(U i W i Thus, X′ and U′ are obtained.
[0039] (4) After obtaining X′ and U′ and using them to form a new dataset, input it into the neural network, and use the weighted sum of the cross-entropy of the labeled data and the L2 loss of the unlabeled data as the loss function L, as shown in equations (14), (15), and (16):
[0040]
[0041]
[0042] L = L X +λ u L U (16)
[0043] Among them, p model (y|x;θ) represents the distribution of an input x with respect to the class label y under parameter θ, where λ uThese are the weight parameters of the loss function. The L2 loss is chosen for unlabeled data because it is bounded and insensitive to incorrect predictions. This makes it a common measure of unlabeled data loss and prediction uncertainty in semi-supervised learning, especially since this framework does not propagate gradients by guessing labels. The network is saved after multiple iterations of training and used for fault diagnosis of rolling bearing vibration signals.
[0044] In the method described in this invention, a dataset composed of vibration signals from rolling bearings is used to train the model. Only a very small portion of this dataset is labeled, and the overall data volume requirement is relatively low. The framework described in the method performs data augmentation on the unlabeled data and constructs pseudo-labels, which are then superimposed and mixed with the labeled data to form a new dataset. On this new dataset, the model is trained using a weighted sum of the losses from the labeled and unlabeled data as the loss function. Finally, the optimal model is used to complete the diagnostic task.
[0045] The beneficial effects of the present invention are:
[0046] This invention makes high use of data, fully explores the information in unlabeled data for complete diagnosis, and truly achieves semi-supervised learning of rolling bearing diagnostic models with only a few labels and low cost, resulting in high cost-effectiveness.
[0047] Instruction manual illustrations
[0048] Figure 1 This is a schematic diagram of the sliding window sampling method;
[0049] Figure 2 This is a diagram of the 1DMixMatch semi-supervised framework.
[0050] Figure 3 This is a schematic diagram of the WDCNN network structure. Detailed Implementation
[0051] The specific embodiments of the present invention are described in detail below with reference to the invention content:
[0052] The process of using the 1DMixMatch semi-supervised framework to diagnose rolling bearing faults is described in detail below. The following examples further illustrate the purpose and usage of this invention, but the invention is not limited thereto.
[0053] 1. Experimental equipment and environment configuration
[0054] Experimental equipment: Dell PowerEdge R840 rack server, Intel Xeon processors*4, Nvidia p2200 graphics cards
[0055] Software system: Windows 10
[0056] Programming language: Python 3.6.2
[0057] Deep learning framework: PyTorch 1.0.1
[0058] Dataset: CWRU Bearing Dataset from Case Western Reserve University, USA
[0059] 2. Experimental Methods
[0060] (1) Bearing vibration data acquisition: The dataset used in this invention example comes from the Electrical Engineering Laboratory of Case Western Reserve University in the United States. The bearing at the motor drive end of the test bench is a deep groove ball bearing SKF6205, and the bearing at the fan end is a deep groove ball bearing SKF6203. Vibration signals are collected using a 16-channel data logger with a sampling frequency of 12kHz.
[0061] The experiment sampled vibration signals of different types of rolling bearings under four load conditions: 0HP, 1HP, 2HP, and 3HP. These included normal state (N), inner ring fault (IR), outer ring fault (OR), rolling element fault (B), and various damage states, totaling 40 types. Damage diameters included 0.1778mm, 0.3556mm, and 0.5334mm, represented by 07, 14, and 21, respectively. This example selected 50 sets of samples (only one labeled set) of the SKF6205 bearing at the drive end under 0HP load for ten different states as the original dataset.
[0062] Under Windows 10, the code executes the `prepare.py` file in the `dataset` directory of the main folder to read and label the vibration signals from the `.mat` files in the `org_data` folder and perform sliding window sampling preprocessing with a sliding step size of 28. Since this framework has many hyperparameters, a large proportion of the data is needed as a validation set to adjust the hyperparameters. Therefore, the dataset is divided into a training set, a validation set, and a test set, as shown in Table 1.
[0063] Table 1. Original CWRU dataset
[0064]
[0065] (2) Data Augmentation: To fully extract information from unlabeled data, unlabeled samples need to undergo K data augmentation operations to transform a sample into K segments. This example uses the Compose function to combine several data augmentation methods. The random inversion method generates a random number between 0 and 1; if the number is less than 0.5, the sample is inverted. The Gaussian noise method directly adds a vector that follows a Gaussian distribution at a factor of 0.15. The 1DCrop method selects a clipping point x. m Then, after trimming, the order of the front and back is reversed. The bearing used in this example is the SKF6205 deep groove ball bearing, with a sampling frequency f. s=12kHz, bearing speed r =1797rpm, bearing frequency is 29.95Hz, number of rolling elements Z =9. The inner ring fault frequency f is calculated using equations (4), (5), and (6). i =Hz, outer ring fault frequency f o =Hz, rolling element failure frequency f b =Hz. In order to preserve the transient impact characteristics of the bearing, the selection of ρ should satisfy equations (2) and (3). In this example, ρ = 0.15 is selected. Thus, x can be determined according to equation (1). m Let be a random point between ρN and (1-ρ)N.
[0066] (3) Guessing Labels: After passing the enhanced K groups of samples through a classifier, the guessed labels q0 are obtained. To ensure code simplicity, the classifier should be as simple and reliable as possible. In this example, the classifier used is a fine-tuned WDCNN convolutional neural network. The first layer of WDCNN has a large convolutional kernel. Its purpose is similar to that of the Short-Time Fourier Transform (SFT), which is to extract short-time features and perform the first screening of the sample features. The difference is that the window function of the SFT is a sine function, while the large convolutional kernel of the first layer of WDCNN is obtained through optimization algorithm training. Its advantage is that it can automatically learn diagnostic features and automatically remove features that are not helpful for diagnosis. To highlight the expressive power of WDCNN, the kernel size of the convolutional layers other than the first layer is 3×1. Since the kernel parameters are small, this is conducive to deepening the network and can suppress overfitting to ensure a simple and reliable structure. After each convolutional operation, batch normalization (BN) is performed, followed by 2×1 max pooling. The purpose of Batch Normalization (BN) is to reduce internal covariate shifts, improve network training efficiency, and enhance the network's generalization ability. The WDCNN network structure is shown in Table 2. Figure 3 .
[0067] Table 2. WDCNN Network Structure
[0068]
[0069] (4) Label processing: In order to obtain reliable pseudo-labels, K sets of guessed labels q0 are processed. 1 ,...,q0 K Take the average to get Then, the Sharpen function shown in equations (8) and (9) is used to target... Each category element in the model is sharpened, implicitly minimizing entropy to obtain a pseudo-label q. The purpose of this framework in making low-entropy predictions for pseudo-labels is to obtain pseudo-labels with clearer categories and more defined centers, which is beneficial for achieving better consistency with the labels of labeled data X in the MixUp operation, thus facilitating successful model training.
[0070] (5) MixUp: To make full use of the data, MixUp is used to mix labeled and unlabeled data as shown in equations (10), (11), (12), and (13). First, the labeled data X and the unlabeled data U are combined and shuffled to form W, and then X' is set to... i =MixUp(X i W i ),U′ i =MixUp(U i W i Thus, we obtain X′=(x′,p′) and U′=(u′,q′).
[0071] (6) Network Training and Parameter Tuning: After the above steps, the processed X′ and U′ are obtained and used as the final input of the current batch. These are then input into the WDCNN neural network, the classifier used for guessing labels. It is worth noting that the classifier in both stages is the same, and the parameters are identical in each iteration. The model uses a weighted sum of the cross-entropy of labeled data X′ and the L2 loss of unlabeled data U′ as its loss function. Regularization is a general method of imposing constraints on a model to make it difficult for it to memorize training data, thus aiming to make it better generalize to unseen data and prevent overfitting to the training set. In this algorithm, weight decay is used to penalize the model's L2 norm, with a decay coefficient of e. The MixUp stage also acts as a regularizer to encourage "convex behavior" between samples. Simultaneously, to accelerate convergence, the Adam optimization algorithm is used, which saves training time and effectively speeds up convergence. The best-performing model is ultimately obtained for fault diagnosis of rolling bearings.
[0072] This algorithm has many hyperparameters, but experiments have shown that most of them can be fixed, except for α and λ. u Adjustments are needed; in practice, a starting point α = 0.75, λ u =100, and the specific values in this example are shown in Table 3.
[0073] Table 31 Hyperparameters of the DCrop-MixMatch Algorithm
[0074]
[0075] 3. Experimental Results
[0076] (1) Comparative experiment:
[0077] To present the comparative experimental results more intuitively and comprehensively, the results of various semi-supervised algorithms on the dataset for the ten classes of tasks are compared, as shown in Table 4.
[0078] Table 4 Comparison of Algorithms for the CWRU Bearing Dataset
[0079]
[0080] When the data scale and number of labels are large, comparing 1DCrop+MixMatch with VAE and PCA-SVM, VAE, as a deep generative model, still significantly outperforms the machine learning algorithm PCA-SVM. However, when the data scale and number of labels become extremely small, VAE's performance drops significantly, while 1DCrop+MixMatch maintains good results, indicating that this algorithm has higher efficiency in utilizing data information and is more advantageous when the data volume and number of labels are small. Furthermore, it provides possibilities for the application of subsequent image domain frameworks in the one-dimensional domain. With the standardization and further development of one-dimensional vibration signal data enhancement methods, the application areas of more advanced frameworks such as ReMixMatch and FixMatch will also expand to the fault diagnosis of rolling bearings.
[0081] (2) Ablation experiment:
[0082] To further verify the effectiveness of the 1DCrop method for one-dimensional signal processing using the MixMatch semi-supervised framework, ablation experiments were also conducted, and the results are shown in Table 5.
[0083] Table 5. Ablation experimental results of the 1DCrop enhancement method on the MixMatch semi-supervised framework.
[0084]
[0085] It can be observed that directly applying the MixMatch semi-supervised framework to one-dimensional data does not achieve satisfactory results. However, the 1DCrop data augmentation method can significantly improve MixMatch's performance on one-dimensional data. It performs well in diagnostic tasks with small data sizes and few labels, and maintains high accuracy even when the data size and labels become extremely small. Conversely, the performance of the MixMatch semi-supervised framework shows a significant decline when the data size decreases. Therefore, it can be concluded that the 1DCrop data augmentation method can effectively help the MixMatch semi-supervised framework achieve one-dimensionality, thereby efficiently utilizing data to complete the fault diagnosis task of rolling bearings.
Claims
1. A method for diagnosing rolling bearing faults based on a semi-supervised learning framework, characterized in that, The steps are as follows: The first step: Vibration signal data acquisition: Obtain the vibration signals of all states at the driving end of the bearing as unlabeled data, and obtain the labeled vibration signals as labeled data. Use the obtained unlabeled data and labeled data as the dataset of the model; The second step: Vibration signal data preprocessing: When dividing a segment of vibration data into samples, adopt the commonly used enhanced sliding window sampling method for one-dimensional data to preprocess the data; The third step: Send the data into the 1DMixMatch semi-supervised learning framework for training to obtain a model and save it: The improved 1DMixMatch semi-supervised learning framework includes data augmentation, guessed labels, label processing, and MixUp; 3.1 Data Augmentation: A segment of sample points u in the input unlabeled data U is transformed into K new segments of sample points through K random augmentations. In this process, the data undergoes a random reversal process with a certain probability of having their order changed, i.e., u = {u1, u2, ..., u...} n After reversal After passing through the Gaussian noise addition stage, the entire sample point will be superimposed with a segment of Gaussian white noise; after passing through the 1DCrop stage, the data will be randomly split into two segments and their order will be swapped, i.e., the input data u={u1,u2,...,u n }, with a certain data point u m As a boundary, the input sample is divided into two segments, u. f ={u1,...,u m-1 } and u b ={u m ,...,u n Then, the order of the samples was rearranged to form a new sample. The choice of m must satisfy equation (1): ρn < m < (1 - ρ)n (1) where ρ is the segmentation parameter; Among them, f s f i f o f b These are the sampling frequency, the bearing inner ring fault frequency, the bearing outer ring fault frequency, and the rolling element fault frequency, respectively; these frequencies can be calculated using empirical formulas (4), (5), and (6): f i =0.6*Z*f r (4) f o =0.4*Z*f r (5) Among them, f r Z represents the bearing rotational frequency, and Z represents the number of rolling elements. 3.2 Guessing Labels: Using the data-augmented K-segment sample points... The guessed label q0 is obtained by inputting each input into the classifier. 1 ,...,q0 K The classifier described above is an arbitrary neural network, and its parameters are consistent with those of the network used by the semi-supervised learning framework after data processing, which is trained on the datasets X′ and U′. The purpose of guessing this series of labels is to facilitate MixUp in recombining and processing labeled data X and unlabeled data U, and to provide a basis for finding a reliable pseudo-label for the unlabeled data. 3.3 Tag Processing: Obtaining the guessed tag q0 1 ,...,q0 K , where q0 k =[i1,...,i L ], k∈(1,K), L is the number of categories; in order to determine the pseudo-labels of the unlabeled data, the guessed labels are averaged to obtain As shown in equation (7): Semi-supervised learning frameworks tend to make low-entropy predictions to provide a pseudo-label that is clearer in terms of category, so the Sharpen function is used to implicitly reduce entropy. After completing the entropy minimization process, q is obtained, as shown in equations (8) and (9): where L is the number of categories, i is one of the L possible labels, T is the temperature parameter. When T approaches 0, q will approach the one-hot distribution. Therefore, reducing the temperature T will encourage the model to produce low-entropy predictions; 3.4 MixUp: After the above processing, obtain the unlabeled data U = (u, q) and the labeled data X = (x, p), where x and u are sample points, p is the label of the labeled data, and q is the pseudo-label of the unlabeled data; then, in order to make full use of the data, the semi-supervised learning framework uses MixUp to mix the labeled data and the unlabeled data; in order to make the data and labels compatible with the loss terms respectively, a modified MixUp version is adopted here, such as equations (10), (11), (12), (13): λ ∼ Beta(α, α) (10) where λ is a parameter obeying the Beat distribution, used to control the mixing degree, and α is a hyperparameter. x and p represent the sample point and the label respectively; The fourth step: After obtaining X′ and U′ and forming a new dataset with them, input them into the neural network. Use the weighted sum of the cross-entropy of the labeled data and the L2 loss of the unlabeled data as the loss function L, such as equations (14), (15), (16): Among them, p model (y|x;θ) represents the distribution of an input x with respect to the class label y under parameter θ, where λ u These are the weight parameters of the loss function.
2. The rolling bearing fault diagnosis method based on a semi-supervised learning framework as described in claim 1, characterized in that, The second step involves the following specific operations: The entire data segment has a length of N, and a series of samples x of length n are sampled from it. i (i = 1, 2, ...), where n must be greater than the number of samples collected after the bearing rotates two revolutions to ensure feature integrity; The starting point of each sampling is shifted backward by p (0 < p < n); this can make full use of the data and retain features on the premise of ensuring the reliability and integrity of the data.
3. A rolling bearing fault diagnosis method based on a semi-supervised learning framework as described in claim 1 or 2, characterized in that, In the above-mentioned 3.1 data augmentation, the K - time random augmentation specifically includes: random inversion, 1D Crop, and adding Gaussian noise.
4. A rolling bearing fault diagnosis method based on a semi-supervised learning framework as described in claim 1 or 2, characterized in that, In the above-mentioned 3.1 data augmentation, ρ is the segmentation parameter, taking values below 0.
2. The selection of ρ should satisfy equations (2) and (3):
5. The rolling bearing fault diagnosis method based on a semi-supervised learning framework as described in claim 3, characterized in that, In the above-mentioned 3.1 data augmentation, ρ is the segmentation parameter, taking values below 0.
2. The selection of ρ should satisfy equations (2) and (3):
6. A rolling bearing fault diagnosis method based on a semi-supervised learning framework as described in claim 1, 2, or 5, characterized in that, In the 3.4 MixUp operation, the semi-supervised learning framework first combines and shuffles the labeled data X and unlabeled data U to form W, and then lets X' i =MixUp(X i W i ),U′ i =MixUp(U i W i Thus, X′ and U′ are obtained.
7. The rolling bearing fault diagnosis method based on a semi-supervised learning framework as described in claim 3, characterized in that, In the 3.4 MixUp operation, the semi-supervised learning framework first combines and shuffles the labeled data X and unlabeled data U to form W, and then lets X' i =MixUp(X i W i ),U′ i =MixUp(U i W i Thus, X′ and U′ are obtained.
8. The rolling bearing fault diagnosis method based on a semi-supervised learning framework as described in claim 4, characterized in that, In the 3.4 MixUp operation, the semi-supervised learning framework first combines and shuffles the labeled data X and unlabeled data U to form W, and then lets X' i =MixUp(X i W i ),U′ i =MixUp(U i W i Thus, X′ and U′ are obtained.
Citation Information
Patent Citations
Fault diagnosis method based on semi-supervised learning deep adversarial network
CN110823574A
Bearing fault diagnosis method based on semi-supervised adversarial network
CN113324758A