Self-Supervised Bearing Fault Diagnosis Method Based on Clustering Algorithm
By using a self-supervision method based on clustering algorithm in bearing fault diagnosis, comparing the learning and clustering algorithms to extract vibration signal characteristics, the problem that data in the existing technology cannot reflect actual faults is solved, and higher diagnostic accuracy and lower labor costs are achieved.
Patent Information
- Application Number
- CN202211507266.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-11-29
AI Technical Summary
The existing self-learning method for bearing fault diagnosis has the problem that data cannot respond well to actual faults.
The self-supervised method based on the clustering algorithm is used to pre-train large sample data sets, and vibration signal characteristics are extracted through comparative learning and clustering algorithms, and the model is pre-trained, and the number of classifications of deep neural networks is replaced by the clustering center to realize the model self-classification.
It effectively solves the problem that data cannot reflect actual failures, improves the diagnostic accuracy of the model in a real industrial environment, reduces dependence on data labels, and reduces labor costs.
Smart Images

Figure CN115791179B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a self-supervised bearing fault diagnosis method based on a clustering algorithm. Background Art
[0002] Traditional intelligent fault diagnosis methods usually take a clear fault type as the final result, and usually it is a single fault category. However, in actual industrial scenarios, most of the equipment is operating normally, and faulty equipment will be repaired. And when the equipment fails, multiple faults may occur simultaneously. Therefore, while it is difficult to obtain fault data, it is even more difficult to define what fault category the final result belongs to.
[0003] Based on the idea of self-supervised learning, using the information provided by the data points themselves in the dataset collected from vibration signals, the model is enabled to self-learn. After the model is pre-trained, the model is then fine-tuned according to a high-quality small-sample real industrial environment dataset, so that the final model can have an effective diagnostic ability. Since the model self-learns from the data, the entire model does not require the participation of artificial labels in the pre-training stage, reducing a large amount of labor costs. Currently, the common self-learning solutions for bearing fault diagnosis mainly have the problem that the data cannot well reflect the actual faults. Summary of the Invention
[0004] The purpose of the present invention is to provide a self-supervised bearing fault diagnosis method based on a clustering algorithm to solve the problem that the existing self-learning of bearing fault diagnosis has data that cannot well reflect the actual faults.
[0005] To solve the above technical problems, the present invention provides a self-supervised bearing fault diagnosis method based on a clustering algorithm, including:
[0006] Pre-training a large-sample dataset, using the large-sample dataset to replace prior knowledge; and
[0007] Extracting the features of the information inherent in the vibration signal, and using the features to perform self-learning through contrastive learning and a clustering algorithm to assist in model pre-training.
[0008] Optionally, in the above-mentioned self-supervised bearing fault diagnosis method based on a clustering algorithm, it further includes:
[0009] Using the number of cluster centers calculated by the clustering algorithm to replace the number of classifications finally output by the deep neural network;
[0010] Based on the characteristics of the clustering algorithm, in the input feature space, the clustering centers are initially generated at random positions, and through continuous iteration, all the input data is self-classified. By applying the clustering algorithm, the pre-trained model self-classifies to generate the final clustering centers;
[0011] where the number of clustering centers is several times the number of specific fault categories.
[0012] Optionally, in the self-supervised bearing fault diagnosis method based on the clustering algorithm, it further includes:
[0013] While the output of the autoencoder of the deep model is input to the classifier for classification, the clustering centers are calculated synchronously through the clustering algorithm;
[0014] Taking the clustering centers as pseudo-labels to train the model for model self-learning;
[0015] Combining the clustering algorithm with contrastive learning, and using the clustering algorithm as a proxy task for contrastive learning;
[0016] Generating pseudo-labels from the encoder output of one deep neural network in the siamese network through the clustering algorithm, and calculating the cross-entropy with the output of the classifier of the other deep neural network in the siamese network to train the model.
[0017] Optionally, in the self-supervised bearing fault diagnosis method based on the clustering algorithm, transfer learning is performed by dividing the training of the entire model into a pre-training stage and a fine-tuning stage;
[0018] In the pre-training stage, the ResNet50 model is trained using the information of unlabeled data;
[0019] In the fine-tuning stage, the encoder part of the model is transferred to the new model, and the classifier is redesigned;
[0020] In the fine-tuning stage, the data with specific labels is input into the new model, and classification is performed through the new classifier with the encoder frozen;
[0021] Calculating the cross-entropy loss using the known label information and the model output, and optimizing the parameters of the new model with the cross-entropy loss;
[0022] When the loss of the model converges, the trained model is put into use in the real industrial scenario.
[0023] Optionally, in the self-supervised bearing fault diagnosis method based on the clustering algorithm,
[0024] Using the Sinkhorn-Knopp algorithm as the clustering algorithm;
[0025] When the model learns through the dataset, the dataset is split into multiple small batches;
[0026] After the encoder of the model, a prototype layer is used to continue training the model, and the prototype layer runs through different small batches of the dataset;
[0027] Using the Sinkhorn-Knopp algorithm, the eigenvalue clustering algorithm calculation is converted into an optimal transport problem for processing;
[0028] The small batch is input into the model, and the output of the prototype layer is clustered through Sinkhorn-Knopp to calculate the final loss;
[0029] An online clustering algorithm is used for training to avoid repeated calculations.
[0030] Optionally, in the self-supervised bearing fault diagnosis method based on the clustering algorithm, the pre-training stage includes:
[0031] Collect 500 groups of signals from each state of the collected vibration signals as the source dataset, where the length of each group of signals is 2048 data points;
[0032] The collected data are respectively subjected to random data augmentation to obtain datasets X1 and X2, which are input into the siamese network;
[0033] The data augmentation methods include adding random Gaussian noise, performing random masking, and randomly varying the signal amplitude.
[0034] The overall structure of the ResNet encoder includes:
[0035] The first part trains the input x of 1 * 2048 to 64 * 512, and the subsequent parts continue to extract features from the data;
[0036] The output of the second part is 256 * 512;
[0037] The output of the third part is 512 * 256;
[0038] The output of the fourth part is 1024 * 128; and
[0039] The output of the fifth part is 2048 * 64;
[0040] In the sixth part, the previously obtained output passes through the average pooling layer, and the ResNet encoder finally outputs features of length 2048;
[0041] The features are calculated by the classifier, and the final output of the entire model has a length of 128;
[0042] The respective clustering centers are obtained through the Sinkhorn-Knopp algorithm and used as pseudo-labels label1 and label2;
[0043] The output of the model is interacted with the generated pseudo-labels to calculate the information entropy. The cross-entropy is calculated between y1 and label2, and between y2 and label1 respectively. The parameters of the entire model are optimized in reverse through the calculated loss;
[0044] Repeat the iteration until the model converges.
[0045] Optionally, in the self-supervised bearing fault diagnosis method based on the clustering algorithm, the fine-tuning stage includes:
[0046] Collect 500 groups of signals from each state of the vibration signal as the dataset X'. Each group of signals x in X' has a length of 2048 data points. The collected data is input into the deep neural network after data augmentation;
[0047] Reconstruct the deep neural network, use the encoder of the model in the pre-training stage, and freeze all the weights and bias values therein;
[0048] Redesign the classifier of the new deep neural network according to the number of fault categories included in the dataset X'. The output length of the classifier is the same as the number of fault categories;
[0049] Input the dataset into the model, extract features of length 2048 according to the encoder, and input them into the classifier for classification to obtain the final output Y;
[0050] Calculate the cross-entropy loss between the output Y of the model and the label information Labels, and use this loss to optimize the parameters of the classifier of the model;
[0051] Repeat the iteration until the model converges.
[0052] Optionally, in the self-supervised bearing fault diagnosis method based on the clustering algorithm, it further includes:
[0053] Define positive and negative samples in the proxy task to train the model;
[0054] Introduce the idea of the clustering algorithm, and use the clustering center to replace the negative sample, so that the positive sample directly calculates the difference with the clustering center;
[0055] The clustering algorithm uses several times the number of clustering centers of the number of fault categories as the prototype layer for training to solve the problem of difficultly calculating the number of fault categories in the dataset in the real industrial environment.
[0056] Optionally, in the self-supervised bearing fault diagnosis method based on the clustering algorithm, it further includes:
[0057] Introduce the Sinkhorn-Knopp algorithm, collect all eigenvalues, and transform the K-Means algorithm for clustering according to the data distribution into optimal transport.
[0058] After the output of each small batch of data is obtained through the prototype layer, directly calculate the pseudo-labels through the Sinkhorn-Knopp algorithm, and directly calculate the loss to optimize the model;
[0059] The model synchronously calculates the cluster centers and trains the model when traversing the dataset.
[0060] The inventors of the present invention have studied and summarized that the common self-learning solutions for bearing fault diagnosis mainly have the following problems:
[0061] First of all, the simulated dataset used during model training usually uses the vibration signals in a fixed environment as the data source. However, the factors that can affect vibration signals are extremely extensive, especially the rotational speed of rotating machinery. For the same bearing component, when the rotational speed of the machinery changes, the generated vibration signals are extremely different.
[0062] Secondly, the vibration signals in the commonly recognized datasets in use are all collected by replacing a single faulty component in the machinery. Therefore, the datasets collected in this way are all vibration signals generated due to clear single fault causes. However, the real industrial situation is different. On the one hand, when a fault occurs, such as cracks, outer ring wear, etc., the severity and location distribution of the faults are inconsistent for each device, and it is impossible to directly unify seemingly similar fault causes into the same fault category. On the other hand, there is the situation where multiple fault causes occur simultaneously. When a bearing component fails due to wear, it is very likely that two or more faults will occur simultaneously, and the vibration signals in this regard are not collected in the currently mainstream datasets.
[0063] In addition, during the pre-training stage of the current model, the datasets used are as described above, and the vibration signals are all collected from machinery with clear single fault causes at a fixed rotational speed. Therefore, the datasets used by the current model are clean and the categories are clear. Based on this, during the pre-training stage, most of the existing models output the exact number of fault categories in the classifier. Even when clustering the features output by the encoder using the clustering algorithm, the number of cluster centers is also the exact number of fault categories. However, this situation does not exist in the real industrial environment. The datasets collected in the real industrial environment are basically difficult to fully explore the specific conditions of the machinery and faults represented by the vibration signals used as the data source.
[0064] In the self-supervised bearing fault diagnosis method based on the clustering algorithm provided by the present invention, by utilizing the characteristics of the clustering algorithm, in the input feature space, the clustering centers at random positions are initially generated, and through continuous iteration, all the input data are self-classified. The number of clustering centers calculated by the clustering algorithm is used to replace the number of classifications finally output by the deep neural network, enabling the model to self-classify and generate clustering centers. Moreover, the number of clustering centers can be several times the number of specific fault categories included. Therefore, the problem that the deep neural network needs to clearly know how many fault categories there are in the dataset during pre-training classification is avoided. Brief Description of the Drawings
[0065] Figure 1 FIG. is a schematic diagram of transfer learning of the self-supervised bearing fault diagnosis method based on the clustering algorithm according to an embodiment of the present invention;
[0066] Figure 2 FIG. is a schematic diagram of introducing a residual block into the encoder of the self-supervised bearing fault diagnosis method based on the clustering algorithm according to an embodiment of the present invention. Detailed Embodiments
[0067] The present invention will be further described below in conjunction with the detailed embodiments with reference to the drawings.
[0068] It should be noted that the components in the drawings may be exaggerated for illustration purposes and are not necessarily drawn to scale. In the drawings, the same or functionally identical components are provided with the same reference numerals.
[0069] In the present invention, unless otherwise specified, the expressions "arranged on...", "arranged above...", and "arranged over..." do not exclude the existence of intermediate objects therebetween. In addition, "arranged on or above..." only represents the relative positional relationship between two components, and in certain cases, such as when the product direction is reversed, it can also be converted to "arranged under or below...", and vice versa.
[0070] In the present invention, each embodiment is only intended to illustrate the solution of the present invention and should not be construed as restrictive.
[0071] In the present invention, unless otherwise specified, the quantifier "a" or "one" does not exclude the scenario of multiple elements.
[0072] It should also be noted here that, in the embodiments of the present invention, for the sake of clarity and simplicity, only a part of the components or assemblies may be shown. However, those of ordinary skill in the art can understand that, under the teaching of the present invention, the required components or assemblies can be added according to the specific scenario requirements. In addition, unless otherwise specified, the features in different embodiments of the present invention can be combined with each other. For example, a certain feature in the second embodiment can be used to replace the corresponding or functionally identical or similar feature in the first embodiment, and the obtained embodiment also falls within the scope of the disclosure or the scope of the record of this application.
[0073] It should also be noted here that within the scope of the present invention, terms such as "identical", "equal", "equal to", etc. do not mean that the two values are absolutely equal, but allow a certain reasonable error. That is to say, these terms also cover "substantially identical", "substantially equal", "substantially equal to". By analogy, in the present invention, the directional terms "perpendicular to", "parallel to", etc. also cover the meanings of "substantially perpendicular to" and "substantially parallel to".
[0074] In addition, the numbering of the steps of each method of the present invention does not limit the execution order of the method steps. Unless otherwise specified, the method steps can be executed in different orders.
[0075] The following further elaborates in detail the self-supervised bearing fault diagnosis method based on the clustering algorithm proposed by the present invention in combination with the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise scales, only for the purpose of conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0076] The purpose of the present invention is to provide a self-supervised bearing fault diagnosis method based on the clustering algorithm to solve the problem that the existing self-learning of bearing fault diagnosis has data that cannot well reflect the actual faults.
[0077] To achieve the above object, the present invention provides a self-supervised bearing fault diagnosis method based on the clustering algorithm, including: pre-training a large sample data set and using the large sample data set to replace the prior knowledge; and extracting the features of the information carried by the vibration signal, and using the features of the information carried by the vibration signal for contrast learning and clustering algorithm to assist the pre-training of the model.
[0078] The present invention proposes a self-supervised bearing fault classification method based on contrastive clustering tasks, aiming to solve the problem that in a real industrial scenario, when the dataset collected for pre-training is relatively mixed and it is difficult to distinguish the fault categories represented by vibration signals, a high-quality model can still be pre-trained using this dataset. Then, through fine-tuning and re-training, the model can be put into use in a real industrial environment. The present invention combines the ideas of contrastive learning and clustering algorithms and designs a self-supervised bearing fault classification method based on contrastive clustering tasks.
[0079] The large sample dataset used in the pre-training stage of the present invention does not require prior knowledge. Both contrastive learning and clustering algorithms extract and utilize features from the information inherent in the vibration signals to assist in model training.
[0080] The present invention utilizes the idea of contrastive learning. Instead of simply directly classifying the features extracted from the source data through deep learning, a siamese network is added on this basis, and the features obtained by the siamese network are compared with each other to distinguish similar and different data. This enables the present invention to better extract features from the data itself and better self-distinguish different categories of data.
[0081] The present invention uses the number of cluster centers calculated by the clustering algorithm to replace the number of classifications finally output by the deep neural network. Due to the characteristics of the clustering algorithm, that is, in the input feature space, the cluster centers are generated at random positions initially, and through continuous iteration, all the input data are self-classified. This enables the model to self-classify and generate cluster centers by applying the clustering algorithm. And the number of cluster centers can be several times the number of specific fault categories included, thus avoiding the problem that the deep neural network needs to clearly know how many fault categories there are in the dataset during pre-training classification.
[0082] The present invention calculates the cluster centers through the clustering algorithm while the output of the autoencoder of the deep model is input to the classifier for classification. The cluster centers are used as pseudo-labels and participate in the training of the model again to achieve the purpose of model self-learning.
[0083] The present invention combines the clustering algorithm into contrastive learning and uses the clustering algorithm as a proxy task for contrastive learning, that is, after generating pseudo-labels for the outputs of the two encoders in the siamese network through the clustering algorithm, the cross-entropy is calculated respectively with the output of the classifier of the other network to train the model, effectively improving the overall prediction accuracy of the model.
[0084] The specific process includes: obtaining x1 and x2 respectively by performing different data augmentations on the vibration signal source data x, and inputting them into the two models of the siamese network respectively. The direct outputs of the models are used to calculate the cross-entropy loss with the pseudo-labels calculated by the other model through the objective function, and then the parameters of ResNet50 are optimized by this loss.
[0085] As Figure 2 shown, the present invention also includes transfer learning, which divides the training of the entire model into two parts: the pre-training stage and the fine-tuning stage. In the pre-training stage, a large amount of unlabeled data is used to train the ResNet50 model by utilizing the information of the data itself; after the pre-training is completed, in the fine-tuning stage, the encoder part of the model is migrated into the new model, and the classifier is redesigned; in the fine-tuning stage, a small amount of data with specific labels is input into the new model, and under the condition of freezing the encoder, classification is performed through the new classifier; the cross-entropy loss is calculated by using the known label information and the model output, and the parameters of the new model are optimized by this loss; when the loss of the model converges, the trained model can be put into use in the real industrial scenario. The present invention is divided into two stages in total, namely the pre-training stage and the fine-tuning and retraining stage, and among them, the contrast learning model based on the clustering task is used in the pre-training stage of the present invention.
[0086] Contrast learning is a type of self-supervised learning, and its typical paradigm is: proxy task + objective function. Different from supervised learning, the proxy task uses the true label Label of the dataset and the output y of the dataset after passing through the model, and calculates the loss through the objective function to optimize the model parameters. Since contrast learning uses unlabeled data, that is, there is no true label Label, a proxy task is used to replace it.
[0087] The siamese network is used in the proxy task, that is, two identical networks are trained simultaneously. Its overall structure is: after the source data undergoes different data augmentations, two inconsistent inputs are obtained, and then different outputs are obtained through two identical neural networks (siamese network).
[0088] The main task of the proxy task is to define positive and negative samples, calculate the loss through the objective function according to the differences between the positive and negative samples, and optimize the model parameters by this. In the present invention, the clustering algorithm is used as the proxy task, and the clustering centers generated by the clustering algorithm replace the concept of negative samples in contrast learning; at the same time, according to the idea in contrast learning that although two inputs with different forms are obtained after the same source data undergoes different data augmentations, their semantics should be the same, the output of the model in the siamese network is alternately compared with the pseudo-labels obtained by the other network through the clustering algorithm, so as to calculate the loss of the entire siamese network and optimize the entire model.
[0089] The selection of the deep neural network includes: the deep neural network used in the Siamese network is a 50-layer Residual Network (ResNet50), and its main architecture is divided into an encoder part and a classifier part. Residual blocks are introduced in the encoder. In the final output H(x) of the residual block, in addition to the output F(x) calculated by the convolutional layer for the input x, the input x itself also participates in the operation, effectively alleviating the problems of gradient explosion and gradient disappearance generated during the training of the entire training model, and enabling the model to be deep enough. The residual block retains the information of the original input when outputting, enabling it to effectively retain the original features of the input data in the high-dimensional feature space.
[0090] In the present invention, a Prototype Layer is added after ResNet50, which is used as the clustering center in the clustering algorithm for learning, and the output of the Prototype Layer is used for subsequent clustering algorithms and the calculation of the final model loss. Through experiments, when the number of clustering centers is 3 to 5 times the number of specific fault categories, good training effects can be obtained. Therefore, the number of output nodes of the Prototype Layer in the present invention is 4 times the number of estimated fault categories.
[0091] In the present invention, the Sinkhorn-Knopp algorithm is used as the clustering algorithm. Due to hardware limitations, the video memory capacity of the graphics card is small. Therefore, when the model learns through the dataset, the dataset is usually split into multiple small batches (min-batch). Different from the traditional K-means algorithm, since the traditional K-means algorithm needs to traverse all min-batches first to obtain all feature values, and then calculate the clustering center based on the differences between the feature values, the traditional clustering algorithm needs to collect the features obtained after all data in the dataset are extracted by the model, and then perform clustering calculations to obtain the clustering center as the pseudo-label. Then, all min-batches are traversed again, and the loss is calculated using the output of the model classifier and the pseudo-label. This results in the entire calculation being repeated once.
[0092] In the present invention, after the encoder of the model, a Prototype Layer is used to continue training the model, and the Prototype Layer can penetrate different min-batches of the entire dataset. And the present invention uses the Sinkhorn-Knopp algorithm to convert the problem of clustering that requires all feature values into an optimal transport problem. This enables the present invention to directly input a min-batch into the model, perform clustering calculations on the output of the Prototype Layer using Sinkhorn-Knopp, and thus calculate the final loss. Therefore, the entire model is trained using an online clustering algorithm, without repeated calculations, and the overall computational amount is directly halved compared with the traditional clustering algorithm.
[0093] The specific process of the present invention includes a pre-training stage: Step 1: Collect 500 groups of signals for each state of the collected vibration signals as the source dataset, and the length of each group of signals is 2048 data points. The collected data is respectively subjected to random data augmentation to obtain two datasets X1 and X2, and then input into the siamese network. The methods used for data augmentation include: adding random Gaussian noise (Gaussian Noise), performing random masking (Mask Noise), randomly changing the signal amplitude (Amplitude Scale), etc. Step 2: The overall structure of the ResNet encoder is divided into 5 parts. Part 1 trains the input x of 1 * 2048 to 64 * 512, and the subsequent parts continue to extract features from the data. The output of Part 2 is 256 * 512, the output of Part 3 is 512 * 256, the output of Part 4 is 1024 * 128, and the output of Part 5 is 2048 * 64. Finally, through the average pooling layer (Average Pooling), the encoder of ResNet finally outputs a feature feature with a length of 2048. The feature is further calculated through the classifier, and finally the output length of the entire model is 128. Step 3: Obtain their respective clustering centers through the Sinkhorn-Knopp algorithm and use them as pseudo-labels label1 and label2 for subsequent use. Step 4: Interact and calculate the information entropy between the output of the model and the generated pseudo-labels, and reverse-optimize the parameters of the entire model through the calculated loss. Repeat steps 2 to 4 until the model converges. In the present invention, the number of iterations is set to 400 times.
[0094] The specific process of the present invention includes a fine-tuning stage: after the pre-training stage is completed, the trained model will be used for fine-tuning. The vibration signals used in this stage need to be accurately collected to determine the true fault categories of the collected vibration signals. The amount of data can be small, that is, small samples with high precision. Step 1: Consistent with the pre-training stage, 500 groups of signals are collected from each state of the vibration signals as the data set X'. Each group of signals x in X' has a length of 2048 data points. The collected data is input into the deep neural network after data augmentation. Step 2: Reconstruct the deep neural network. Use the encoder of the model in the pre-training stage and freeze all the weights and bias values therein. Then, according to the number of fault categories included in the data set X', redesign the classifier of the new deep neural network. The output length of the classifier is the same as the number of fault categories. Step 3: Input the data set into the model, extract features with a length of 2048 according to the encoder, and input them into the classifier for classification to obtain the final output Y. Step 4: Calculate the cross-entropy loss between the output Y of the model and the label information Labels, and then use this loss to optimize the parameters of the classifier of the model. Repeat steps 3 and 4 until the model converges. In the present invention, the number of iterations is set to 20 times. The model completed in the fine-tuning stage can be put into use in the real industrial environment. The model trained by the present invention can efficiently diagnose the fault categories included in the data set X' collected in the fine-tuning stage.
[0095] The present invention introduces the idea of contrastive learning. Contrastive learning conducts comparisons by defining examples that are semantically similar to it (positive samples) and examples that are semantically different (negative samples) in the sample. By designing the model structure and contrastive loss, it makes the positive samples closer to each other in the feature space, while the positive samples and negative samples are farther away from each other, so as to achieve an effect similar to clustering. Compared with simply directly extracting features from data through a deep neural network for clustering, contrastive learning can make the differences between different samples larger, thus having a better "classification" performance.
[0096] The present invention is different from traditional contrastive learning. Positive and negative samples are defined in the proxy task to train the model. The idea of the clustering algorithm is introduced in the present invention, and the clustering center replaces the concept of negative samples. First, it enables the positive samples to directly calculate the difference with the clustering center instead of calculating the difference with all negative samples, reducing the overall computational amount of the model. And different from accurately classifying faults with the clustering algorithm, the clustering algorithm of the present invention uses several times the number of clustering centers as the prototype layer for training compared to the number of fault categories to address the problem that it is difficult to calculate the number of fault categories in the data set in the real industrial environment.
[0097] In the present invention, the Sinkhorn-Knopp algorithm is introduced to transform the K-Means algorithm, which requires collecting all eigenvalues first and then clustering according to the data distribution, into an optimal transport problem. This enables, after the output is obtained through the prototype layer for each min-batch of data, the pseudo-labels to be directly calculated by the Sinkhorn-Knopp algorithm and the loss to be directly calculated to optimize the model. Thus, an online clustering algorithm is achieved. This allows the model to not need to first traverse the complete dataset to calculate the clustering centers and then traverse the dataset again to train the model. Instead, the clustering centers and the model are calculated synchronously while traversing the dataset, saving half of the overall training time compared to traditional clustering algorithms; when the dataset is too large, even if the feature set is obtained after the dataset is extracted through the model, but if the number of samples is too large, it may still occupy a large amount of storage space. An overly large feature set is not only difficult to calculate and time-consuming, but also more difficult to calculate if it exceeds the video memory size of the graphics card. However, the online clustering algorithm of the present invention has no such concerns.
[0098] The pre-training of the present invention generally divides the entire model into two major parts, namely the pre-trained model and the fine-tuning model for subsequent training. This allows the models included in the two parts, although the goal is to classify the fault of the vibration signal in the final industrial environment, there can be certain differences in the ways to achieve this goal, the training sets used, and the methods employed, etc. In the present invention, in view of the situation where the number of fault data is small, the cost of manually labeled data is high, and it is difficult to clearly and completely distinguish the fault categories, in the pre-training stage, a large amount of unlabeled data is adopted and self-supervised learning is used to better train the overall model.
[0099] The present invention uses a clustering algorithm because the reasons for generating vibration signals are more complex and diverse, making it an extremely labor-consuming or even impossible task to correctly label all data completely. In the present invention, by extracting the information carried by the input data itself and then using the clustering idea that similar data has similar features and different data has large feature differences, the input data is self-classified. Finally, the generated clustering centers are used as pseudo-labels and applied back to the model to train the classification ability of the model, transforming the unsupervised learning of the clustering algorithm of the entire model into self-supervised learning, effectively improving the accuracy of model prediction.
[0100] The present invention provides a fault diagnosis algorithm with good diagnostic effect for scenarios where there is insufficient fault data with classification labels or it is difficult to distinguish the fault categories of all the collected data.
[0101] The present invention applies clustering algorithm to self-supervised learning, and uses the cluster centers of the clustering algorithm to replace the number of fault categories in traditional classification learning, so that in the model pre-training stage, the strict requirements on the data source are reduced and the amount of data in the unlabeled data set is expanded.
[0102] The present invention uses the Sinkhorn-Knopp algorithm to transform the traditional clustering algorithm into an optimal transmission problem, making it an online clustering algorithm. This greatly improves the operating efficiency of the model and greatly reduces the video memory required for model calculation.
[0103] The present invention separates the data sets and training methods used in the two parts by dividing the entire solution into two parts, namely the pre-trained model and the fine-tuned model to be trained later. This allows the model migrated in the retraining phase to be trained with a wider range of data sets, effectively improving the training accuracy.
[0104] An embodiment of the present invention includes migrating a well-labeled data set under different working conditions. To a certain extent, the bearing vibration signals collected under different working conditions are different, such as the influence of temperature pressure, air humidity, etc. This is an extremely common situation in a real industrial environment, but compared to the vibration signal changes caused by different speeds of machinery, it is still considered to be data with the same label. Then when the data collected in the pre-training model is in the above relationship with the data used in fine-tuning training, the present invention can extract high-dimensional features from the vibration signal through a deep learning model, and this feature will remain basically unchanged under the above circumstances. Therefore, a large amount of existing data can be used for pre-training, and in actual working conditions, a small amount of homologous data sets can be used for fine-tuning to achieve a good diagnostic effect.
[0105] An embodiment of the present invention includes migrating an incompletely labeled data set under different working conditions. When the machine is using different speeds or contains an undetected fault, the vibration signal generated will be very different from the expected type of vibration signal. In this case, the fault categories contained in the collected data set will be more than expected, and the vibration signals that are considered to be classified as the same type of fault are actually different fault categories. If a large amount of data needs to be collected manually and the fault categories of all data need to be accurately distinguished, this will consume a huge amount of human resources and is unreasonable. The present invention uses cluster learning to self-learn and self-classify using the information contained in the data itself. Afterwards, it is trained again with a small amount of carefully sampled data sets to avoid the decrease in model training accuracy due to the inclusion of "wrong" data in the training set.
[0106] In summary, the above embodiments have described in detail different configurations of the self-supervised bearing fault diagnosis method based on the clustering algorithm. Of course, the present invention includes but is not limited to the configurations listed in the above embodiments. Any content obtained by transformation based on the configurations provided in the above embodiments falls within the scope of protection of the present invention. Those skilled in the art can draw inferences by analogy based on the content of the above embodiments.
[0107] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for related parts.
[0108] The above description is only a description of the preferred embodiments of the present invention and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art of the present invention based on the above disclosure fall within the scope of protection of the claims.
Claims
1. A self-supervised bearing fault diagnosis method based on a clustering algorithm, characterized in that, Including: Pre-training a large sample dataset, using the large sample dataset to replace prior knowledge; And Extracting the features of the information carried by the vibration signal, and using the features to perform self-learning through contrast learning and clustering algorithms to assist model pre-training; Combining the clustering algorithm into the contrast learning, and using the clustering algorithm as a proxy task for contrast learning; While the output of the encoder of the deep model is input to the classifier for classification, simultaneously calculating the cluster centers through the clustering algorithm; Using the cluster centers as pseudo-labels to train the model for model self-learning; Obtaining x1 and x2 respectively through different data augmentations of the vibration signal source data x, and inputting them into the two models of the siamese network respectively; Generating pseudo-labels through the clustering algorithm for the outputs of the encoders of the two deep neural networks in the siamese network, and calculating the cross-entropy respectively with the outputs of the classifiers of the other deep neural network in the siamese network to train the model.
2. The self-supervised bearing fault diagnosis method based on a clustering algorithm according to claim 1, characterized in that, Also including: Replacing the number of classifications finally output by the deep neural network with the number of cluster centers calculated by the clustering algorithm; Based on the characteristics of the clustering algorithm, in the input feature space, initially generating cluster centers at random positions, and through continuous iteration enabling all input data to perform self-classification, and through applying the clustering algorithm, enabling the pre-trained model to perform self-classification to generate the final cluster centers; Where the number of cluster centers is several times the number of specific fault categories.
3. The self-supervised bearing fault diagnosis method based on a clustering algorithm according to claim 1, characterized in that, Performing transfer learning by dividing the training of the entire model into a pre-training stage and a fine-tuning stage; In the pre-training stage, training the ResNet50 model using the information of unlabeled data; In the fine-tuning stage, migrating the encoder part of the model into a new model and redesigning the classifier; In the fine-tuning stage, inputting the data with specific labels into the new model, and performing classification through the new classifier with the encoder frozen; Calculating the cross-entropy loss using the known label information and the model output, and optimizing the parameters of the new model with the cross-entropy loss; When the loss of the model converges, putting the trained model into real industrial scenarios for application.
4. The self-supervised bearing fault diagnosis method based on a clustering algorithm according to claim 1, characterized in that, Using the Sinkhorn-Knopp algorithm as the clustering algorithm; When the model learns through the dataset, splitting the dataset into multiple small batches; Using a prototype layer after the encoder of the model to continue training the model, and the prototype layer runs through different small batches of the dataset; Using the Sinkhorn-Knopp algorithm to convert the eigenvalue clustering algorithm calculation into an optimal transport problem for processing; Inputting the small batch into the model, and performing clustering calculation on the output of the prototype layer through Sinkhorn-Knopp to calculate the final loss; Adopting an online clustering algorithm for training to avoid repeated calculations.
5. The self-supervised bearing fault diagnosis method based on a clustering algorithm according to claim 3, characterized in that, The pre-training stage includes: Collecting 500 groups of signals for each state of the collected vibration signals as the source dataset, where the length of each group of signals is 2048 data points; Performing random data augmentations on the collected data respectively to obtain datasets X1 and X2, and inputting them into the siamese network; Where the data augmentation methods include adding random Gaussian noise, performing random masking, and randomly varying the signal amplitude.
6. The self-supervised bearing fault diagnosis method based on a clustering algorithm according to claim 3, characterized in that, The fine-tuning stage includes: Collect 500 groups of signals from each state of the vibration signal as the dataset X'. Each group of signals x in X' has a length of 2048 data points. Input the collected data into the deep neural network after data augmentation; Reconstruct the deep neural network, use the encoder of the model in the pre-training stage, and freeze all the weights and bias values therein; Redesign the classifier of the new deep neural network according to the number of fault categories included in the dataset X'. The output length of the classifier is the same as the number of fault categories; Input the dataset into the model, extract features with a length of 2048 according to the encoder, and input them into the classifier for classification to obtain the final output Y; Calculate the cross-entropy loss between the output Y of the model and the label information Labels, and use this loss to optimize the parameters of the classifier of the model; Repeat the iteration until the model converges.
7. The self-supervised bearing fault diagnosis method based on the clustering algorithm according to claim 1, characterized in that, It also includes: Define positive and negative samples in the proxy task to train the model; Introduce the idea of the clustering algorithm, and use the cluster centers to replace the negative samples, so that the positive samples directly calculate the differences with the cluster centers; The clustering algorithm uses several times the number of cluster centers equal to the number of fault categories as the prototype layer for training to solve the problem that it is difficult to calculate the number of fault categories in the dataset in the real industrial environment.
8. The self-supervised bearing fault diagnosis method based on the clustering algorithm according to claim 1, characterized in that, It also includes: Introduce the Sinkhorn-Knopp algorithm to transform the K-Means algorithm that collects all eigenvalues and then clusters according to the distribution of the data into an optimal transport problem; After each small batch of data obtains the output through the prototype layer, directly calculate the pseudo-labels through the Sinkhorn-Knopp algorithm, and directly calculate the loss to optimize the model; The model synchronously calculates the cluster centers and trains the model when traversing the dataset.
Citation Information
Patent Citations
Self-supervised bearing fault diagnosis method
CN116026590A