A small sample bearing fault diagnosis method based on hierarchical contrast learning optimization
By using a hierarchical comparative learning optimization method, a CMBDC model was constructed, which solved the problem of insufficient accuracy in bearing fault diagnosis under small sample conditions and achieved efficient fault identification and classification with limited data.
Patent Information
- Application Number
- CN202411262953.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-09-10
AI Technical Summary
Under small sample conditions, deep learning models struggle to effectively utilize limited fault data for bearing fault diagnosis, resulting in insufficient diagnostic accuracy.
A few-sample bearing fault diagnosis method based on hierarchical contrastive learning optimization is adopted. By constructing a hierarchical contrastive learning module, a parallel attention fusion module, a stacked activation function and a Brownian distance covariance metric module, the feature extraction and classification process is optimized. The backpropagation algorithm is used to update the network parameters, thereby improving the nonlinear expressive power and classification accuracy of the model.
It effectively reduces the distance between samples of the same type and increases the distance between samples of different types, thereby improving the classification accuracy of bearing fault diagnosis for small samples, especially the fault identification capability under cross-load conditions.
Smart Images

Figure CN119469767B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of bearing fault diagnosis, and particularly relates to a small sample bearing fault diagnosis method based on hierarchical contrast learning optimization. BACKGROUND
[0002] Fault diagnosis as a key technical means plays a crucial role in many fields such as industrial production, transportation, energy and power. With the rapid development of science and technology, various mechanical equipment and systems are becoming increasingly complex and automated. Once a fault occurs, it may cause serious economic losses or even safety accidents. Therefore, studying fault diagnosis technology is of great significance to ensure the safe operation of equipment, improve production efficiency and promote the sustainable and healthy development of the national economy.
[0003] In recent years, deep learning has been widely applied in the field of fault diagnosis due to its high precision, strong feature extraction ability and strong generalization ability. However, deep learning often requires a large amount of data to train the model. In actual application scenarios, it is difficult to obtain a large amount of fault data due to expensive equipment, low frequency of fault occurrence or difficulty in data collection. This makes it difficult for traditional fault diagnosis methods based on large data to be applicable. Therefore, it is crucial to study fault diagnosis methods under small sample conditions. Under limited sample conditions, how to effectively utilize existing data and improve the accuracy of fault diagnosis has become a problem to be solved in the current fault diagnosis field. SUMMARY
[0004] To solve the above technical problems, the technical solution adopted by the present application is: a small sample bearing fault diagnosis method based on hierarchical contrast learning optimization, comprising the following steps:
[0005] S1: using a sensor to collect bearing vibration signals of different fault categories under different loads, extracting sample data according to a fixed signal length and a sliding window, and then using short-time Fourier transform to obtain time-frequency domain features of the signals;
[0006] S2: dividing the data set according to different loads into a training data set and a test data set, constructing multiple tasks for the training data set and the test data set respectively, each task containing a support set and a query set;
[0007] S3: Construct a small sample bearing fault diagnosis model CMBDC optimized based on hierarchical contrast learning, the CMBDC model comprises: a hierarchical contrast learning module, a parallel attention fusion module, a stacked activation function, and a Brown distance covariance measurement module; wherein the hierarchical contrast learning module is used to calculate the contrast loss by constructing positive and negative sample pairs in the feature extraction process, so that the distance between samples of the same class is closer, and the distance between samples of different classes is farther; the parallel attention fusion module enhances the discrimination between classes by fusing multiple attention mechanisms to learn more comprehensive feature information; the stacked activation function is used to replace the ordinary activation function during feature extraction to enhance the nonlinear expression ability of the network; and the BDC measurement module is used to calculate the similarity between the query sample and each class of the support set, and the sample with high similarity is taken as the classification result of the query sample;
[0008] S4: input the training data set into the CMBDC model, weight and fuse the obtained shallow contrast loss, deep contrast loss, and classification loss calculated by the measurement module to obtain the total loss of the model, calculate the gradient by using the back propagation algorithm, update the network parameters by using the optimization algorithm, continuously optimize the network parameters through multiple training iteration cycles to obtain the final model; and input the test data set into the trained CMBDC model to realize the classification of the test data set.
[0009] The beneficial effects produced by the above technical scheme are as follows:
[0010] The method disclosed by the application utilizes the characteristics of contrast learning that can reduce the distance between samples of the same class and increase the distance between samples of different classes, introduces contrast learning into the feature extraction part, so that the model learns to distinguish the features of positive and negative samples; the designed parallel attention fusion module does not separately input each sample of the support set into the parallel attention fusion module, but splices all samples of the support set into a large tensor and then inputs the large tensor into the parallel attention fusion module, so that the attention mechanism can more comprehensively learn the features of the samples of this class, each attention mechanism can learn different spatial relationships and attention weight distributions, and more rich attention expressions are obtained; the introduction of the stacked activation function enhances the nonlinear expression ability of the model; the BDC measurement method obtains more accurate similarity by measuring the joint distribution between samples; and the classification accuracy of small sample bearing fault diagnosis is effectively improved. BRIEF DESCRIPTION OF DRAWINGS
[0011] The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0012] Figure 1 The flowchart of the method described in the embodiments of the application is shown in the figure;
[0013] Figure 2Model diagram of the method described in the embodiment of the present application. DETAILED DESCRIPTION
[0014] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0015] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced in other ways that are not exactly as described in the present application, and those skilled in the art can make similar extensions without departing from the spirit of the present application, so the present application is not limited to the specific embodiments disclosed below.
[0016] As shown in Figure 1 The present application discloses a small sample bearing fault diagnosis method based on hierarchical contrast learning optimization, specifically comprising the following steps:
[0017] S1: Collect bearing vibration signals of different fault categories under different loads using sensors, extract sample data according to fixed signal length and sliding window, and then obtain time-frequency domain features of the signals using short-time Fourier transform;
[0018] First, the original vibration signal is sampled by sliding window, each sample contains 2500 sampling points, and the step length is 500. Then, the time-frequency domain features of the signal are obtained by short-time Fourier transform.
[0019] S2: Divide the data set according to different loads into a training data set and a test data set, and construct multiple tasks for the training data set and the test data set respectively, each task containing a support set and a query set;
[0020] In this example, the bearing data set (CWRU) of Case Western Reserve University in the United States is selected for experiment, the driving end with a sampling frequency of 12 kHz is selected to construct the data set, a total of 3 loads are selected, 1 hp, 2 hp and 3 hp, each load contains 4 health states, including healthy state, inner ring fault, outer ring fault and rolling element fault, each fault contains 3 different sizes, a total of 10 categories. The experimental data used in the experiment is shown in Table 1:
[0021] Table 1 CWRU data
[0022]
[0023] Each load contains 500 samples of each data, and each sample is 2500 long. The training data set Dtrain 5000 samples in total, test data set D test 5000 samples in total.
[0024] The training data set builds 200 tasks for each training round, each task contains a support set and a query set, which are randomly selected from D train N classes are randomly selected, and K samples are randomly selected from each class to build a support set L samples are randomly selected from the remaining samples in N classes to build a query set
[0025] The test data set builds 20 tasks for each training round, each task contains a support set and a query set, which are randomly selected from D test N classes are randomly selected, and K samples are randomly selected from each class to build a support set L samples are randomly selected from the remaining samples in N classes to build a query set
[0026] In order to verify the performance of the CMBDC model based on hierarchical contrastive learning optimization proposed in the application, six cross-load experiments are designed: 1hp→2hp, 1hp→3hp, 2hp→3hp, 2hp→1hp, 3hp→1hp, 3hp→2hp, wherein 1hp→2hp represents using the data under 1hp to train the model and using the data under 2hp to test the model.
[0027] S3: build a small sample bearing fault diagnosis model CMBDC based on hierarchical contrastive learning optimization;
[0028] The structure diagram of the CMBDC model is shown in Figure 2 The CMBDC model includes: hierarchical contrastive learning module, parallel attention fusion module, stacked activation function, Brown distance covariance measurement module. Among them, the hierarchical contrastive learning module is used to calculate the contrastive loss by constructing positive and negative sample pairs in the feature extraction process, so that the distance between samples of the same class is closer and the distance between samples of different classes is farther; the parallel attention fusion module enhances the discrimination between classes by fusing multiple attention mechanisms to learn more comprehensive feature information; the stacked activation function is used to replace the ordinary activation function when extracting features, which enhances the nonlinear expression ability of the network; the BDC measurement module is used to calculate the similarity between the query sample and each class of the support set, and the larger similarity is taken as the classification result of the query sample.
[0029] The hierarchical contrastive learning module constructed has 6 layers, including 4 convolution blocks and two contrastive learning modules, wherein the first layer and the second layer are convolution blocks, the third layer is a shallow contrastive learning module, the fourth layer and the fifth layer are convolution blocks, and the sixth layer is a deep contrastive learning module.
[0030] The constructed parallel attention fusion module has 6 layers for each attention mechanism. The first and second layers are adaptive average pooling layers, the third layer is a convolutional block, and the fourth and fifth layers are attention generation layers. The compressed feature vector is expanded to the original number of channels through convolution operations, and attention weights are generated through the Sigmoid activation function. The sixth layer calculates the product of the attention weights and the input features to output the final features.
[0031] The introduced stacked activation function is defined as follows:
[0032]
[0033] In the formula, n represents the number of stacked activation functions, a i and b i This represents the scale and bias of each activation function, scaling and shifting the output of the activation functions to avoid simple accumulation.
[0034] The introduced BDC metric model is defined as follows:
[0035] The Euclidean distance matrix between each class of features X in the support set and the features Y in the query sample is calculated as follows:
[0036]
[0037] In the formula, x k x l Let X and L represent two observations of feature X, i.e., two columns of the feature vector. Calculate the Euclidean distance between the k-th column and the l-th column.
[0038] The BDC matrix of X and Y after centering is calculated as follows:
[0039]
[0040] In the formula, the last three terms represent respectively The average row value The column average, and The average of all vectors.
[0041] The value of the BDC metric ρ(X,Y) is calculated using the following expression:
[0042] ρ(X,Y)=tr(X T ,Y)
[0043] In the formula, X and Y represent the BDC matrices of the support set features X and the query set samples Y, respectively.
[0044] S4: input the training data set into the CMBDC model, weight and fuse the obtained shallow contrastive loss, deep contrastive loss and classification loss calculated by the BDC measurement module to obtain the total loss of the model, calculate the gradient by using the back propagation algorithm, update the network parameters by using the optimization algorithm, continuously optimize the network parameters through multiple training iteration cycles to obtain the final model, input the test data set into the trained CMBDC model to realize the classification of the test data set;
[0045] The calculation formulas of the shallow contrastive loss and the deep contrastive loss are as follows:
[0046]
[0047] In the formula, x i,1 ,x i,2 represents a positive sample, x j,k represents a negative sample, and sim represents the similarity between two samples.
[0048] The calculation formula of the classification loss is as follows:
[0049]
[0050] In the formula, x is the real fault category, and y is the fault category predicted by the model.
[0051] The calculation formula of the total loss is as follows:
[0052] L total =αL contrastive +βL c
[0053] In the formula, a is the weight corresponding to the contrastive loss, and β is the weight corresponding to the classification loss.
[0054] In this example, the training rounds are set to 20, each training round contains 200 tasks, the loss function is composed of the contrastive loss and the classification loss, the SGD optimizer is used, the learning rate is 0.001, the loss value gradually decreases during the training process, and the network model finally reaches a stable state.
[0055] The CMBDC model training process is as follows:
[0056] (1) input the training data set into the hierarchical contrastive learning module according to the task to obtain the support set features after shallow contrastive learning and deep contrastive learning The shallow contrastive loss L s , the deep contrastive loss L d ; the query sample does not need to calculate the contrastive loss, and the query sample feature is
[0057] (2) input the support set features query sample feature input into the parallel attention fusion module to obtain the feature after attention selection
[0058] (3) calculate the BDC matrix of the support set feature calculate the BDC matrix of the query sample Finally, calculate obtain the classification result of the query sample Y.
[0059] (4) compare the classification result with the prediction result to calculate the classification loss L c , the classification loss L c , the shallow contrastive loss L s , the deep contrastive loss L d weighted fusion to obtain the total loss L total , calculate the gradient by using the back propagation algorithm, update the network parameters by using the optimization algorithm, train the model, and until the model converges.
[0060] The specific process of using the trained CMBDC model for fault diagnosis is as follows:
[0061] (1) input the test data set into the hierarchical contrastive learning module according to the task, without calculating the shallow contrastive loss and the deep contrastive loss, to obtain the support set feature of each task query set feature
[0062] (2) calculate the BDC matrix of the support set feature query sample feature input into the parallel attention fusion module to obtain the feature after attention selection
[0063] (3) calculate the BDC matrix of the support set feature calculate the BDC matrix of the query sample Finally, calculate obtain the classification result of the query sample Y.
[0064] Table 2 shows the experimental accuracy of different models under cross-load conditions
[0065]
[0066] In order to illustrate the effectiveness of the model of the present application, K nearest neighbor (KNN), ResNet network, matching network (Matching), MLFD, and prototype network (Prototypical) methods are selected for comparison, and 6 cross-load experiments are performed. The comparison experimental results under the condition of 10-way 5-shot are shown in Table 2.
Claims
1. A small sample bearing fault diagnosis method based on hierarchical contrastive learning optimization, the steps of which are as follows: S1: Collect bearing vibration signals of different fault categories under different loads using sensors, extract sample data according to fixed signal length and sliding window, and then obtain time-frequency domain features of the signals using short-time Fourier transform; S2: Divide the data set into training data set and test data set according to different loads, construct multiple tasks for the training data set and the test data set respectively, and each task contains a support set and a query set; S3: Construct a small sample bearing fault diagnosis model CMBDC based on hierarchical contrastive learning optimization, the CMBDC model comprising: a hierarchical contrastive learning module, a parallel attention fusion module, a stacked activation function, and a Brown distance covariance measurement module; wherein the hierarchical contrastive learning module is used to construct positive and negative sample pairs in the feature extraction process, calculate the contrastive loss, make the distance between samples of the same class closer, and make the distance between samples of different classes farther; the parallel attention fusion module enhances the discrimination between classes by fusing multiple attention mechanisms to learn more comprehensive feature information; the stacked activation function is used to replace the ordinary activation function during feature extraction to enhance the nonlinear expression ability of the network; and the BDC measurement module is used to calculate the similarity between the query sample and each class of the support set, and the sample with high similarity is taken as the classification result of the query sample; S4: input the training data set into the CMBDC model, weight and fuse the obtained shallow contrastive loss, deep contrastive loss, and classification loss calculated by the measurement module to obtain the total loss of the model, calculate the gradient using the back propagation algorithm, update the network parameters using the optimization algorithm, continuously optimize the network parameters through multiple training iteration cycles to obtain the final model, and input the test data set into the trained CMBDC model to realize the classification of the test data set.
2. The small sample bearing fault diagnosis method based on hierarchical contrastive learning optimization of claim 1, wherein, Step S2 requires the following: (1) According to the different loads, the data is divided into training data set D train , test data set D test ; (2) Construct multiple tasks for the training data set and the test data set respectively, each task contains a support set and a query set, and the construction method is as follows: An N-way K-shot method is adopted to construct multiple tasks for the training data set, support sets and query sets are constructed for each task, N classes are randomly selected from the total fault categories of the training data set, K samples are randomly selected from each class to form the support set of the task where i represents the i-th task, (x i,j , y i,j ) represents the j-th sample of the i-th task, y i,j is the corresponding label, and k is the number of samples of each task; L samples are randomly selected from the remaining samples of the N classes to form the query set of the task where l represents the number of query samples; multiple tasks are constructed for the test data set in the same way, support sets and query sets are constructed for each task, N classes are randomly selected from the total fault categories of the test data set, K samples are randomly selected from each class to form the support set of the task L samples are randomly selected from the remaining samples of the N classes to form the query set of the task 3. The small sample bearing fault diagnosis method based on hierarchical contrastive learning optimization of claim 1, wherein, Step S3 requires the following: (1) The constructed hierarchical contrast learning module first constructs positive and negative sample pairs. In the N classes of the support set, a class is randomly selected, and two samples are randomly selected to form a positive sample pair where N i represents the randomly selected class, p represents the positive sample pair, x i,1 , x i,2 respectively represent the first positive sample and the second positive sample, and y i is the corresponding label; then, from other classes, samples different from the current sample class are selected to form N-1 negative sample pairs where n represents the negative sample pair, x i,1 represents the positive sample, x j,k represents the negative sample, y i , y j respectively represent the corresponding labels; the InfoNCE contrast loss is used to maximize the similarity score between the positive sample pairs and minimize the similarity score between the negative sample pairs. The calculation formula of the InfoNCE contrast loss is as follows: where x i,1 represents a positive sample, x i,2 represents a negative sample, and sim represents a similarity between two samples. j,k where x i,1 represents a positive sample, x i,2 represents a negative sample, and sim represents a similarity between two samples. j,k where x i,1 represents a positive The hierarchical contrastive learning module has 6 layers, including 4 convolutional blocks and two contrastive learning modules, wherein the first layer and the second layer are convolutional blocks, the third layer is a shallow contrastive learning module, the fourth layer and the fifth layer are convolutional blocks, and the sixth layer is a deep contrastive learning module; (2) The constructed parallel attention fusion module is composed of three coordinate attention mechanisms in parallel, each coordinate attention mechanism decomposes the input feature into two one-dimensional feature vectors along the spatial dimension, representing the information in the vertical and horizontal directions respectively, and aggregates the features along the vertical and horizontal directions; the support set features extracted by the hierarchical contrastive learning module are spliced into a whole vector, and the size of a feature map is c×h×w; k feature maps of each class are combined into k×c×h×w, and then the tensor is reshaped into c×(k×h)×w through axis transformation operation; the transformed features are input into the three attention mechanisms respectively, and the selected results of the three attention mechanisms are weighted and averaged and then output. Each attention mechanism has 6 layers, the first and second layers are adaptive average pooling layers, the third layer is a convolution block, the fourth and fifth layers are attention generation layers, the compressed feature vector is expanded to the original channel number through convolution operation, and the attention weight is generated through the Sigmoid activation function, the sixth layer calculates the product of the attention weight and the input feature, and outputs the final feature; (3) The definition of the stacked activation function is as follows: where n denotes the number of stacked activation functions, a i and b i denote the scale and bias of each activation function, which scale and shift the output of the activation function to avoid simple summation. (4) The definition of the BDC metric module is as follows: 1) Calculate the Euclidean distance matrix of each class of features X of the support set and the feature vector Y of the query sample, and the calculation expression is as follows: where x k , x l represent two observations of feature X, i.e., two columns of the feature vector, the Euclidean distance matrix between the kth column and the lth column is calculated; 2) Calculate the BDC matrix of X and Y after centering, and the calculation expression is as follows: In the formula, the last three terms represent respectively The average row value The column average, and The average of all vectors; 3) Calculate the value of BDC metric ρ(X, Y), and the calculation expression is as follows: p(X, Y) = tr(X T , Y) In the formula, X and Y represent the BDC matrices of the support set features X and the query set sample Y respectively.
4. The small sample bearing fault diagnosis method based on hierarchical contrastive learning optimization of claim 1, wherein, The specific requirements of step S4 are as follows: (1) The training data set is input into the hierarchical contrast learning module according to the task, and the support set features after shallow contrast learning and deep contrast learning are obtained The shallow contrast loss L s , the deep contrast loss L d ; the query sample does not need to calculate the contrast loss, and the query sample feature is obtained (2) support set features query sample features input to the parallel attention fusion module to obtain the features after attention selection (3) Calculate the BDC matrix of each feature of the support set Calculate the BDC matrix of the query sample Finally, calculate Get the classification result of the query sample Y; (4) The classification result is compared with the prediction result to calculate a classification loss L c The classification loss L c , the shallow contrast loss L s , the deep contrast loss L d is weighted and fused to obtain a total loss L total , the gradient is calculated by using a back propagation algorithm, the network parameters are updated by using an optimization algorithm, the model is trained until the model converges (5) The test data set is input into the hierarchical contrast learning module according to the task, without calculating the shallow contrast loss and the deep contrast loss, to obtain the support set features of each task query set features (6) support set features query sample features input to the parallel attention fusion module to obtain the features after attention selection (7) Compute the BDC matrix of the support set features Compute the BDC matrix of the query sample Finally, compute Get the classification result of the query sample Y.
Citation Information
Patent Citations
Rolling bearing fault diagnosis method and device under sample imbalance condition
CN117312941A
Supervised contrast learning network model training multivariate time series data classification method and system
CN117407772A
Small-sample image classification method and system based on metric element learning
CN117975086A