Bearing Fault Diagnosis Method, Device, Electronic Equipment and Storage Medium
By using neural network models to process unbalanced training samples in bearing fault diagnosis, and using feature extraction and classification submodels for fault diagnosis, the problem of poor classification effect in the prior art is solved, and higher diagnostic accuracy and accuracy are achieved.
Patent Information
- Application Number
- CN202210716327.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-06-22
AI Technical Summary
In bearing fault diagnosis, it is difficult for the prior art to effectively deal with unbalanced training samples, resulting in poor classification results, especially in the case of a few types of fault vibration signals, which can easily lead to misclassification and affect the safety and economicality of mechanical equipment.
The pre-trained neural network model is used to convert the vibration signal of the bearing into a feature vector through the feature extraction sub-model, and the classification sub-model is used to troubleshoot based on the feature vector. The target loss function of the neural network model updates the cross entropy loss function based on the weights of each training sample category to adjust the training for unbalanced samples.
It improves the accuracy of bearing fault diagnosis, can effectively classify the vibration signals of bearings in unbalanced training samples, reduce misclassification, and improves the accuracy and reliability of fault diagnosis.
Smart Images

Figure CN115112372B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mechanical equipment fault diagnosis, and particularly to a bearing fault diagnosis method, device, electronic device and storage medium. Background Art
[0002] Modern mechanical equipment has precise structures, complex production regulation and maintenance. It is difficult for humans to effectively predict and prevent mechanical failures. Therefore, the fault diagnosis of mechanical equipment has become an important research topic. For bearings, in real fault diagnosis scenarios, the number of fault samples is scarce. When classifying with unbalanced training samples, the classifier will pay more attention to the majority samples, resulting in poor classification effects. The classifier can refer to a neural network model. For a neural network model, training is usually required. Currently, the training samples have few fault samples and many non-fault samples, resulting in sample imbalance. This situation will cause poor classification effects even if the neural network model is trained. In the case of mechanical equipment being a bearing, the samples are vibration signals. Considering that in industrial monitoring, the fault vibration signals are the minority class, if the fault vibration signals are classified as non-fault categories, it will cause great economic losses and even lead to safety accidents.
[0003] Therefore, a bearing fault diagnosis method is needed to solve the above technical problems. Summary of the Invention
[0004] Embodiments of the present invention provide a bearing fault diagnosis method, device, electronic device and storage medium, which realize the diagnosis of bearing faults and improve the accuracy of diagnosis.
[0005] In a first aspect, embodiments of the present invention provide a bearing fault diagnosis method, the method comprising:
[0006] Input the vibration signal of the bearing into the feature extraction sub-model of a pre-trained neural network model to obtain a feature vector; the neural network model includes a target loss function; the neural network model is obtained by training on multiple training samples, the training samples belong to corresponding sample categories, different sample categories have different weights, and the target loss function is obtained by updating the cross-entropy loss function according to the weights of the sample categories of each training sample;
[0007] Input the feature vector into the classification sub-model of the neural network to obtain multiple probability output values of the vibration signal of the bearing belonging to different sample categories, and obtain the fault situation corresponding to the vibration signal of the bearing based on the highest probability output value. Each probability output value represents the probability of the vibration signal belonging to the sample category, and the sample category represents the fault situation of the bearing vibration signal. The sample category includes fault types and non-fault types.
[0008] In a second aspect, an embodiment of the present invention further provides a bearing fault diagnosis device, which includes:
[0009] A feature vector acquisition module that inputs the vibration signal of the bearing into the feature extraction sub-model of a pre-trained neural network model to obtain a feature vector; the neural network model includes a target loss function; the neural network model is obtained by training multiple training samples, the training samples belong to corresponding sample categories, different sample categories have different weights, and the target loss function is obtained by updating the cross-entropy loss function according to the weights of the sample categories of each training sample;
[0010] A fault diagnosis result acquisition module that inputs the feature vector into the classification sub-model of the neural network to obtain multiple probability output values of the vibration signal of the bearing belonging to different sample categories, and obtains the fault condition corresponding to the vibration signal of the bearing based on the highest probability output value. Each probability output value represents the probability that the vibration signal belongs to the sample category, and the sample category represents the fault condition of the bearing vibration signal. The sample category includes a fault type and a non-fault type.
[0011] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes:
[0012] One or more processors;
[0013] A storage device for storing one or more programs,
[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the bearing fault diagnosis method in any embodiment of the present invention.
[0015] In a fourth aspect, an embodiment of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the bearing fault diagnosis method in any embodiment of the present invention when executed by a computer processor.
[0016] In the technical solution of the embodiment of the present invention, the vibration signal of the bearing is input into the feature extraction sub-model of the pre-trained neural network model to obtain a feature vector; the feature vector is input into the classification sub-model of the neural network to obtain multiple probability output values of the vibration signal of the bearing belonging to different sample categories, and the fault condition corresponding to the vibration signal of the bearing is obtained based on the highest probability output value. Since the target loss function of the neural network model is obtained by updating the cross-entropy loss function according to the weights of each training sample category, and there may be an imbalance problem in the training samples, that is, there are some samples with a large number and some samples with a very small number, and the neural network model in the embodiment of the present invention can adjust the target loss function for the weights of unbalanced samples, that is, obtain an updated target loss function according to the weights of each training sample category. Therefore, the neural network model trains and learns for unbalanced samples, and the classification accuracy of the trained neural network model is higher. Processing the vibration signal according to the neural network model, the obtained fault diagnosis result is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Among them:
[0019] Figure 1 is a schematic flowchart of a bearing fault diagnosis method in Embodiment 1 of the present invention;
[0020] Figure 2 is a schematic flowchart of a method for solving the weight adjustment factor in Embodiment 1 of the present invention;
[0021] Figure 3 is a schematic flowchart of a bearing fault diagnosis test in Embodiment 1 of the present invention;
[0022] Figure 4 is a schematic diagram of the training process of the neural network model in Embodiment 1 of the present invention;
[0023] Figure 5 is a schematic flowchart of a bearing fault diagnosis method in Embodiment 2 of the present invention;
[0024] Figure 6 is a line graph of the F1-score for testing PCS-CGCN in Embodiment 2 of the present invention;
[0025] Figure 7It is a schematic diagram of a confusion matrix for model classification in the second embodiment of the present invention;
[0026] Figure 8 It is a schematic diagram of a time-domain signal with a signal-to-noise ratio of -3dB in the second embodiment of the present invention;
[0027] Figure 9 It is a bar chart comparing the accuracies of various models in the second embodiment of the present invention;
[0028] Figure 10 It is a schematic diagram of an experimental test platform for the IMS dataset in the second embodiment of the present invention;
[0029] Figure 11 It is a bar schematic diagram of the evaluation index of a model in the second embodiment of the present invention;
[0030] Figure 12 (a) It is a loss curve graph of model training before and after introducing transfer learning in the second embodiment of the present invention;
[0031] Figure 12 (b) It is a test accuracy curve graph of model training before and after introducing transfer learning in the second embodiment of the present invention;
[0032] Figure 13 It is a schematic diagram of the structure of a bearing fault diagnosis device in the third embodiment of the present invention;
[0033] Figure 14 It is a schematic diagram of the structure of an electronic device in the fourth embodiment of the present invention. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.
[0035] Before elaborating on the technical solutions of the embodiments of the present invention, an exemplary description of the application scenarios of the embodiments of the present invention will be given first:
[0036] In actual situations, the monitored data often operates in a variable working condition environment throughout the measurement period, and the spatial distribution of the data varies greatly, which will degrade the fault diagnosis performance of the model. Specifically, on the one hand, under variable working conditions, as the speed and load of the rotating machinery change, the generalization performance of the model will decrease. On the other hand, under variable working conditions, different models need to be established to adapt to different working conditions, resulting in high resource consumption. Since the data imbalance problem and the need for variable working condition fault diagnosis coexist in real fault diagnosis scenarios, the embodiments of the present invention propose a method for bearing fault diagnosis, which can still improve the accuracy of bearing fault diagnosis in imbalanced training samples.
[0037] Embodiment 1
[0038] Figure 1 FIG. is a schematic flowchart of a bearing fault diagnosis method provided by an embodiment of the present invention. This embodiment is applicable to the situation of bearing fault diagnosis. This method can be executed by a bearing fault diagnosis device, and the device can be implemented in the form of software and / or hardware.
[0039] As Figure 1 shown, the bearing fault diagnosis method of the embodiment of the present invention specifically includes the following steps:
[0040] S110. Input the vibration signal of the bearing into the feature extraction sub-model of the pre-trained neural network model to obtain a feature vector.
[0041] The neural network model further includes an objective loss function; the neural network model is obtained by training multiple training samples. The training samples belong to corresponding sample categories, and different sample categories have different weights. The objective loss function updates the cross-entropy loss function according to the weights of the sample categories of each training sample. A training sample refers to a sample used during the training of the neural network model. The training samples include multiple sample categories, and each sample category includes multiple samples. The sample category represents the fault condition of the bearing vibration signal. The sample category includes fault types and non-fault types. The sample category can include non-fault sample categories, A-fault sample categories, B-fault sample categories, C-fault sample categories, etc. It should be noted that in the embodiment of the present invention, the fault samples are divided into multiple different fault category samples. In this way, the trained neural network model can accurately obtain the fault type of the bearing according to the vibration signal. To clearly divide different sample categories, sample labels are added to the samples of each sample category, so that each sample category has a corresponding sample label. For example, the sample categories include fault categories and non-fault categories, and the corresponding sample labels are 1 and 0 in sequence. For example, the sample categories include non-fault category, A-fault category, B-fault category, and the corresponding sample labels are 1, 2, 3 in sequence. The vibration signal refers to the vibration signal generated by the bearing during use, and the vibration signal can be collected by a sensor. It is understood that the vibration signal of a faulty bearing is different from that of a non-faulty bearing, and it can be known whether the bearing is faulty from the vibration signal. Further, the fault categories of faulty bearings are different, and the vibration signals are also different. The fault category of the faulty bearing can be known according to the vibration signal. The feature extraction sub-model can process the vibration signal, extract the features in the vibration signal, and obtain a feature vector. The feature extraction sub-model can be a model of a graph convolutional network, or a combined model of a one-dimensional convolutional neural network and a graph convolutional network. If it is a graph convolutional network, the vibration signal is first converted into a graph signal and then input into the graph convolutional network. If it is a combination of a one-dimensional convolutional neural network and a graph convolutional network, the vibration signal is input into the one-dimensional convolutional network, the output one-dimensional feature vector is converted into a graph signal, and then the graph signal is input into the graph convolutional network. The bearing includes sliding bearings, joint bearings, rolling bearings, etc. In the embodiment of the present invention, the bearing can be a rolling bearing.
[0042] Specifically, the vibration signal of the bearing is input into the feature extraction sub-model of the preset trained neural network model to obtain a feature vector, which prepares for the subsequent calculation of the classification layer.
[0043] Further, in the embodiment of the present invention, updating the cross-entropy loss function according to the weights of the sample categories of each training sample includes: in each iterative training process of the neural network model, inputting the weights of the sample categories of each training sample into the following formula to update the cross-loss function; Among them, Loss′ is the loss value of the target loss function, and S i represents the probability output value of the i-th node in the classification nodes of the target loss function. The number of classification nodes is the same as the number of sample categories; the probability output value is expressed as: z i represents the output value of the i-th classification node, and z v represents the output value of the v-th classification node; the weights of the sample categories are obtained through the following formula;
[0044] Among them, ω m represents the weight corresponding to the m-th training sample, cnt() represents the number of training samples of the () category, δ(m) represents the sample category corresponding to the m-th training sample, K represents the total number of all training samples, V represents the number of sample types, h represents the non-fault sample category, f1 represents the first type of fault sample category, and f V-1 represents the fault sample category with the least number of samples, and α is the weight adjustment factor. It should be noted that only the fault sample category with the least number of samples has the weight adjustment factor. For example, cnt() with () being the sample category A represents the number of training samples of sample category A. It should be understood that A can be the sample label. Also, for example, if there are 100 training samples with a sample label of 1, then cnt(1) is equal to 100. It should be noted that in the embodiments of the present invention, the fault sample types are sorted. f1 is the first type of fault sample category, f2 is the second type of fault sample category, and so on, which is convenient for operation and does not cause confusion of the fault sample categories. Among them, the fault sample category with the least number of samples is used as the last sample category in the sorting. That is, f V-1. Of course, this is only an exemplary embodiment. The fault sample categories may not be sorted, but the fault sample category with the least number of samples needs to be distinguished. The fault sample category with the least number of samples has a weight adjustment factor. The way to distinguish the fault sample category with the least number of samples can be in the form of an identifier. The identifier can be a combination of numbers, letters, and numbers and letters. The identifier is not specifically limited here. Different weights are assigned to samples of different categories according to the sample labels. The weights are reduced for non-fault samples. For fault samples with a large number of samples, the proportion of fault samples in all fault samples is used as the weight when calculating the loss value. For the fault samples with the least number, a weight adjustment factor is introduced so that the fault samples with the least number can also obtain the corresponding weight. The weight adjustment factor can be obtained by particle swarm optimization (PSO). The particle swarm algorithm theory is mature and has the advantages of low algorithm complexity and fast convergence speed. At the same time, since the graph convolution network consumes a long time during the training process, if other optimization algorithms are selected, the excessive algorithm complexity will increase the calculation amount of the neural network model, thereby reducing the fault diagnosis effect of the model in the application. Therefore, the particle swarm algorithm is used to solve the optimal value of the weight adjustment factor, and the loss value generated during the training process is set as the fitness function of the particle swarm algorithm iteration. The solution process is as follows: Figure 2 As shown. It should be understood that the particles here refer to the weight adjustment factor, and the population size and the number of iterations are preset. The population size can be set according to expert experience, for example, 3 particles, 5 particles, etc. Then, the particles are randomly initialized, and the particle fitness is evaluated to obtain the global optimum, and it is determined whether the number of iterations is met. If not, the particle speed and position are updated, the fitness function value is calculated again, the particle historical optimal position is updated, and the population global optimal position is updated. After one iteration is completed, the number of iterations is increased by 1, and it is determined again whether the number of iterations is met. If so, the final particle is used as the optimal weight adjustment factor. Of course, the weight adjustment factor can also be obtained by other means, and the embodiment of the present invention does not make specific limitations. The use of a particle swarm algorithm is a preferred solution of the embodiment of the present invention.
[0045] Furthermore, in an embodiment of the present invention, the feature extraction submodel includes a first graph convolutional network, and the initial parameters of the first graph convolutional network are obtained through transfer learning of an offline network model; the initial parameters of the first graph convolutional network of the feature extraction submodel of the neural network model are obtained through transfer learning of the offline network model, including: iteratively training the offline network model, and using the optimal parameters of the offline first graph convolutional network of the trained offline network model as the initial parameters of the first graph convolutional network of the feature extraction submodel of the neural network model; the offline first graph convolutional network has the same network structure as the first graph convolutional network.
[0046] Among them, the first graph convolutional network may refer to a graph convolutional network (GCN) with 32 convolutional kernels. The following second graph convolutional network is the same as the first graph convolutional network, except for the number of convolutional kernels. The initial parameters include weights and biases. Optionally, the network structures of the offline network model and the neural network model are the same. In this way, when training the offline network model, the optimal parameters of the offline first graph convolutional network can be used as the initial parameters of the first graph convolutional network, improving the training efficiency of the neural network model.
[0047] Specifically, by iteratively training the offline network model, the optimal parameters of the offline first graph convolutional network of the trained offline network model are used as the initial parameters of the first graph convolutional network. When training the neural network model, training can be carried out based on the initial parameters, improving the training efficiency and accuracy.
[0048] Exemplarily, when training the offline network model, the ratio of training samples to test samples in the training dataset is 8:2. When training the neural network model, since the number of dataset samples in the target domain is extremely small, in order to approximate the real experimental effect, the ratio of training samples to test samples in the training dataset is set to 2:8. Since there are already initial parameters, the neural network model can be trained based on only a few training samples, and only a small adjustment to the initial parameters is required to complete the training. It should be noted that the other parts in the neural network model are randomly generated initial parameters. Other parts, for example, include a one-dimensional convolutional network before the first graph convolutional network. Optionally, the training samples can be the open-source CWRU bearing dataset.
[0049] Furthermore, the feature extraction sub-model further includes a second graph convolutional network, and the output of the first graph convolutional network is used as the input of the second graph convolutional network; after obtaining the initial parameters of the first graph convolutional network, it further includes: inputting multiple training samples, and iteratively training the neural network model until the number of iterations reaches the preset number of iterations and stops iterating, completing the training of the neural network model.
[0050] Among them, the second graph convolutional network may refer to a graph convolutional network with 64 convolutional kernels.
[0051] Specifically, multiple training samples are input, and the neural network model is iteratively trained. When the number of iterations reaches the preset number of iterations, the iteration stops. The training of the neural network model can reduce the training samples because the initial parameters of the first graph convolutional network are the optimal parameters of the offline first graph convolutional network of the offline network model. Therefore, during training, only a small number of training samples are required. Of course, the specific number of training samples is not limited here. Through the training of the first graph convolutional network and the second graph convolutional network in the neural network model, the training of the neural network model is completed. The loss value of the last iteration is used to determine the weight of each sample category in the loss function, and then the target loss function is determined.
[0052] Further, the feature extraction sub-model further includes a one-dimensional convolutional neural network; the vibration signal of the bearing is input into the feature extraction sub-model of the pre-trained neural network model to obtain a feature vector, including: the vibration signal of the bearing is input into the one-dimensional convolutional neural network to obtain a one-dimensional feature vector; the one-dimensional feature vector is converted into a graph signal, and the graph signal is input into the first graph convolutional network to obtain a first feature vector; the first feature vector is input into the second graph convolutional network to obtain a feature vector.
[0053] Among them, the number of convolution kernels in the one-dimensional convolutional neural network can be 16.
[0054] Specifically, the vibration signal of the bearing is input into the one-dimensional convolutional neural network to obtain a one-dimensional feature vector, the one-dimensional feature vector is converted into a graph signal based on the distance formula, and the graph signal is input into the first graph convolutional network to obtain a first feature vector, and the first feature vector is input into the second graph convolutional network to obtain a feature vector. Through the processing of the one-dimensional convolutional network, the features in the vibration signal can be extracted, and the extracted features can be dimensionally reduced to obtain a one-dimensional feature vector.
[0055] The one-dimensional convolutional network (1D-CNN) includes two parts: a convolutional layer and a pooling layer. The functions of the one-dimensional convolutional network include extracting the feature information of the vibration signal and the dimension reduction function. Optionally, the vibration signal is x, with a length of N, the weight matrix of the convolution kernel of the one-dimensional convolutional network is W, and the activation operation of the convolution kernel is expressed as:
[0056]
[0057] Among them, is the i-th feature of the (l - 1)-th layer, is the convolution kernel weight matrix between the j-th feature of the (l - 1)-th layer and the i-th feature of the l-th layer; is the bias of the i-th feature of the l-th layer; Relu(·) is a non-linear activation function.
[0058] After the convolutional layer is the pooling layer, also known as the max pooling layer. The role of the pooling layer is to compress the information of the convolutional layer and prevent overfitting of the one-dimensional convolutional model. In the embodiments of the present invention, multiple convolutional and pooling operations are required. For example, convolution - pooling - convolution - pooling... convolution - pooling. Learn the feature information of the vibration signal to obtain a feature vector, and perform a flattening operation to convert the multi-dimensional feature into a one-dimensional feature. Each output feature vector can be expressed as:
[0059] Y m =[y1,y2,...,y k ,m∈(0,K],
[0060] where, Y m represents the feature vector of the m-th vibration signal, K represents the number of samples, k is the length of the feature vector, and y represents the feature value.
[0061] Furthermore, in the embodiments of the present invention, converting the one-dimensional feature vector into a graph signal includes: converting the one-dimensional feature vector into a graph signal based on the distance matrix.
[0062] Specifically, the distance matrix can be a matrix of Euclidean distances. By converting the one-dimensional feature vector into a graph signal, it prepares for subsequent graph convolution processing by inputting the graph signal into the first graph convolutional network.
[0063] By calculating the Euclidean distance between feature points, d ij represents the Euclidean distance between i y j and
[0064] d ij =||y i -y j ||,
[0065] obtain the distance matrix of each feature vector,
[0066]
[0067] where, d i,j represents the distance between the i-th feature point and the j-th feature point. D m represents the distance matrix corresponding to the m-th vibration signal.
[0068] A graph signal refers to an undirected graph, and the undirected graph can be defined as where, V, E, respectively represent the vertices, edges, and adjacency matrix of the graph. According to the obtained distance matrix, assign weights to each edge to complete the conversion of the feature graph.
[0069] Define the Laplacian matrix L as follows, where D is the degree matrix, is the adjacency matrix, I k is the k-th order identity matrix, Performing eigenvalue decomposition on L, it can be expressed as: L = UΛU T , Λ = diag([λ0,λ1,...,λ k-1 ) is a diagonal matrix constructed with the eigenvalues of the Laplacian matrix. U is a vector matrix composed of the eigenvectors corresponding to each eigenvalue of the Laplacian matrix. The essence of the Fourier transform is to find the integral of the input signal and the basis vectors. Therefore, the Fourier transform and inverse Fourier transform of the graph signal can be respectively expressed as: where, is the Fourier transform of the graph signal G. The convolution of two graph signals G and g can be expressed as G*g = U(U T G·U T g), by replacing U T g in the formula with gθ, we get G*g = Ugθ·U T G. Since it is very difficult to solve gθ, a general method is to use the Chebyshev polynomial T H iterated H times for replacement fitting, where, is the coefficient vector of the h-th order Chebyshev polynomial, and T h represents the h-th order Chebyshev polynomial. Therefore, the final expression of the graph convolution operation is as follows, where Z is the output after the input graph signal undergoes multiple graph convolution and graph pooling operations.
[0070] Figure 3 shows the entire process of conducting a bearing fault diagnosis experiment without introducing transfer learning. In Figure 3Among them, a rolling bearing is first selected, and vibration signals are collected by sensors and transmitted to a computer. The training samples of the vibration signals are divided. For example, sample 1, sample 2... sample j,... sample N. All the training samples are input into a one-dimensional convolution for feature extraction and dimensionality reduction. The one-dimensional convolution can include two convolution and pooling operations. The training samples are sequentially passed through the C1 convolution layer, the P1 pooling layer, the C2 convolution layer, and the P2 pooling layer. And the convolution results of the one-dimensional convolution are flattened to obtain multiple one-dimensional feature vectors. Optionally, each one-dimensional feature vector refers to the one-dimensional feature vector obtained by flattening a single training sample. The one-dimensional feature vectors of each training sample are converted into graph signals, that is, graph signal 1,... graph signal j,... graph signal N. Each graph signal is input into the first graph convolution network (only the first graph convolution network is shown in the figure, including two convolution layers GC1 and GC2, and the second graph convolution network is not shown. It should be understood that the number of graph convolution networks here can be set according to the actual situation. GP1 and GP2 represent pooling layers). The obtained first feature vector is input into the second graph convolution network. The obtained feature vectors are flattened to obtain the one-dimensional convolution feature vectors of multiple training samples. Then the one-dimensional convolution feature vectors are split into the feature vectors corresponding to each graph signal, and are processed through a fully connected layer and a target loss function to obtain the loss values between each training sample and each sample label. If the loss value does not reach the preset threshold, error backpropagation is performed according to the loss value, and the neural network model is iteratively trained. The weight adjustment factor of the target loss function is iteratively solved and optimized through the PSO algorithm to obtain the optimal weight adjustment factor, complete the training of the neural network model, and output the probability that each sample belongs to each sample label. For example, P(f V-1 |m) represents the probability that the m-th sample belongs to f V-1 . It should be understood that this is a neural network model (PSO optimized Cost-Sensetive Convolution and Graph Convolution Network, PCS-CGCN) without introducing transfer learning of an offline network model. At this time, the hyperparameters of the neural network model can be seen in Table 1. It should be understood that the following network layers are sorted in the order from input to output. For example, first, the input signal enters a convolution layer with a kernel size of 128*1, and then passes through a 2*1 pooling layer. Graph Conv represents a graph convolution layer, and FC represents a fully connected layer.
[0071] Table 1
[0072]
[0073]
[0074] Optionally, the training of the neural network model will be exemplified by Figure 4 for illustrative purposes. Figure 4 is a neural network model that incorporates transfer learning. The network model on the source domain side is an offline network model, and the target domain side is a neural network model, which can also be referred to as an online neural network model. In the offline network model, the input signal is input into a one-dimensional convolution, normalized through a BN (Batch Normalization) layer, and then pooled through a pooling layer (pooling) to obtain a one-dimensional feature vector. In the embodiments of the present invention, the one-dimensional convolutional network includes two convolutional and pooling operations, enabling more comprehensive feature extraction from vibration signals. Subsequently, it passes through a dropout layer and a Flatten layer. The Flatten layer is used to flatten the output result, that is, to convert a multi-dimensional feature vector into a one-dimensional feature vector. Then it passes through an offline first graph convolutional network, including a convolutional layer with 32 convolutional kernels and a pooling layer P2. The output feature vector is input into the dropout layer and then through a fully connected layer FC512, which includes 512 nodes. Finally, through the loss function softmax layer, according to the weight adjustment factor, the offline loss function is adjusted, and the offline network model is iteratively trained. When the number of iterations reaches the preset number of iterations, the training of the offline network model is completed. The optimal parameters of the offline first graph convolutional network of the offline network model are used as the initial parameters of the first graph convolutional network of the neural network model, and the weights of the offline loss function of the offline network model are used as the weights of the loss function of the neural network model. Then, the training of the neural network model is carried out. From Figure 4 it can be known that the difference between the neural network model and the offline network model lies in that the neural network model adds a second graph convolutional network after the first graph convolutional network. When training the neural network model, the initial parameters of other parts except the first graph convolutional network can be randomly initialized, and then the neural network model is trained. Optionally, the number of convolutional operations included in the second graph convolutional network is not specifically limited here and can be the same as the first graph convolutional network, including two convolutional operations.
[0075] Exemplarily, Figure 4 the kernel size, number of kernels, and stride of the network layers from top to bottom of the offline network model in
[0076] Table 2
[0077]
[0078]
[0079] Table 3
[0080]
[0081] It should be understood that in the embodiments of the present invention, the size and quantity of the convolution kernels are not specifically limited and can be adjusted according to actual situations.
[0082] S120. Input the feature vector into the classification sub-model of the neural network to obtain multiple probability output values of the vibration signal of the bearing belonging to different sample categories, and obtain the fault condition corresponding to the vibration signal of the bearing based on the highest probability output value.
[0083] Among them, each probability output value represents the probability that the vibration signal belongs to the sample category. The classification sub-model can be a softmax layer. The classification sub-model is used to calculate the feature vector to obtain the sample category to which the vibration signal corresponding to the feature vector belongs.
[0084] Specifically, the classification layer calculates the feature vector to obtain the probability output value corresponding to each sample category, and obtains the fault condition corresponding to the vibration signal of the bearing from the highest probability output value.
[0085] Further, in the embodiments of the present invention, obtaining the fault condition corresponding to the vibration signal of the bearing based on the highest probability output value includes: using the sample category corresponding to the highest probability output value as the fault diagnosis result of the bearing.
[0086] Specifically, using the sample type corresponding to the highest probability output value as the fault diagnosis result of the bearing can improve the accuracy of the fault diagnosis result.
[0087] In the technical solution of the embodiment of the present invention, by inputting the vibration signal of the bearing into the feature extraction sub-model of the pre-trained neural network model, a feature vector is obtained. The feature vector is input into the classification sub-model of the neural network, and multiple probability output values of the vibration signal of the bearing belonging to different sample categories are obtained. Based on the highest probability output value, the fault condition corresponding to the vibration signal of the bearing is obtained. Each probability output value represents the probability that the vibration signal belongs to the sample category, and the sample category represents the fault condition of the bearing vibration signal. The sample category includes fault types and non-fault types. Since the target loss function of the neural network model is obtained by updating the cross-entropy loss function according to the weights of each training sample category, and there may be an imbalance problem in the training samples, that is, there are some samples with a large number and some samples with a very small number. The neural network model of the embodiment of the present invention can adjust the target loss function for the weights of the imbalanced samples, that is, obtain an updated target loss function according to the weights of each training sample category. Therefore, the neural network model trains and learns the imbalanced samples, and the classification accuracy of the trained neural network model is higher. By processing the vibration signal according to the neural network model, the obtained fault diagnosis result is more accurate.
[0088] Embodiment 2
[0089] Figure 5 The following is a schematic flowchart of a bearing fault diagnosis method provided by an embodiment of the present invention. The embodiment of the present invention is a refinement of step S120 on the basis of the above embodiment. Among them, the same or similar technical terms as the above embodiment will not be described in detail.
[0090] As Figure 5 shown, a bearing fault diagnosis method according to an embodiment of the present invention specifically includes the following steps:
[0091] S210. Input the vibration signal of the bearing into the feature extraction sub-model of the pre-trained neural network model to obtain a feature vector.
[0092] S220. Input the feature vector into the classification sub-model of the neural network to obtain multiple probability output values of the vibration signal of the bearing belonging to different sample categories.
[0093] S230. Use the sample category corresponding to the highest probability output value as the fault diagnosis result of the bearing.
[0094] Among them, after the feature vector is input into the classification layer, probability output values corresponding to the feature vector and each sample category will be obtained. For example, the sample categories include class A1, class B1, class C1, and class D1, and there are four probability output values, which are the probability that the vibration signal belongs to class A1, the probability that the vibration signal belongs to class B1, the probability that the vibration signal belongs to class C1, and the probability that the vibration signal belongs to class D1.
[0095] Specifically, select the highest probability output value from multiple probability output values, and use the sample category corresponding to the highest probability output value as the fault diagnosis result. For example, the sample categories include Class A1, Class B1, Class C1, and Class D1, and there are four probability output values, namely the probability that the vibration signal belongs to Class A1 is 0.2, the probability that the vibration signal belongs to Class B1 is 0.1, the probability that the vibration signal belongs to Class C1 is 0.3, and the probability that the vibration signal belongs to Class D1 is 0.4. The fault diagnosis result of the bearing is obtained as the bearing belongs to Class D1.
[0096] The technical solution of the embodiment of the present invention is to input the vibration signal of the bearing into the feature extraction sub-model of the pre-trained neural network model to obtain a feature vector, input the feature vector into the classification sub-model to obtain a probability output value corresponding to each sample category, obtain a probability output value corresponding to each sample category, and use the sample category corresponding to the highest probability output value as the fault diagnosis result of the bearing. Through the technical solution of the embodiment of the present invention, accurate fault classification of the bearing is realized, whether the bearing has a fault is determined, and moreover, according to the probability output value, it can be accurately known which type of fault the bearing has, improving the accuracy of fault diagnosis. Since the target loss function of the neural network model is obtained by updating the cross-entropy loss function according to the weights of each training sample category, and there may be an imbalance problem in the training samples, that is, there are some samples with a large number and some samples with a very small number, and the neural network model of the embodiment of the present invention can adjust the target loss function for the weights of the imbalanced samples, that is, obtain an updated target loss function according to the weights of each training sample category. Therefore, the neural network model trains and learns for the imbalanced samples, and the classification accuracy of the trained neural network model is higher. Processing the vibration signal according to the neural network model, the obtained fault diagnosis result is more accurate.
[0097] Exemplarily, in order to verify the classification effect and noise robustness of the neural network model PCS-CGCN without introducing transfer learning in the case of uneven sample quantity distribution in the embodiment of the present invention, and to evaluate the effectiveness and accuracy of the neural network model PCS-TCGCN with introduced transfer learning for fault diagnosis, the following experiments are respectively carried out in the embodiment of the present invention. The relevant frameworks of PCS-CGCN and PCS-TCGCN in this example are written in Python tensorflow. All experiments are run on a computer equipped with an i5-6300Q central processing unit and an NVIDIA GeForce GTX 965M GPU.
[0098] It should be noted that both the training samples and test samples in the embodiments of the present invention are data from the CWRU bearing dataset of Case Western Reserve University in the United States. This dataset can be obtained through open-source websites and is widely used in bearing fault diagnosis experiments. Tests were conducted under four rotational speeds of 1797 r / min, 1772 r / min, 1750 r / min, and 1730 r / min and four different loads of 0 hp, 1 hp, 2 hp, and 3 hp. A single-point fault was introduced into the rolling bearing using the electrical discharge machining method. The fault locations include ball fault (BF), outer race fault (OF), and inner race fault (IF), and the fault severities are 0.007 inches, 0.014 inches, 0.021 inches, and 0.028 inches respectively. Then, the acceleration sensor collected vibration signals at a frequency of 12 kHz.
[0099] 1. For the experiment of PCS-CGCN, set the length of each sample to 2048, and divide the vibration signals collected under the 0 hp working condition and 1797 r / min condition to obtain the final unbalanced dataset as shown in Table 4.
[0100] Table 4
[0101]
[0102]
[0103] Test the performance of PCS-CGCN on four datasets. At the same time, to illustrate the feasibility and superiority of PCS-CGCN in the embodiments of the present invention in the classification of unbalanced fault data, comparative experiments were conducted using commonly used network models in this field, including: 1. A fault diagnosis model based on the deep graph convolutional network (DGCN); 2. A fault diagnosis model based on the time-frequency feature bidirectional gated unit (BiGRU); 3. A 1D-CNN fault diagnosis model based on acoustic-vibration data fusion; 4. A support vector machine (SVM) model using the RBF as the kernel function and 24-dimensional features as the input signal. Each experiment was repeated 10 times, and the final average classification accuracy is shown in Table 6, and the F1-score is as Figure 6 shown. The optimal value of the weight factor solved by PSO is shown in Table 5.
[0104] Table 5
[0105]
[0106] Table 6
[0107]
[0108] The experimental results show that in the four datasets, as the imbalance degree of the number of fault samples deepens, the recognition accuracy of different models for faults has changed. Among them, the classification accuracy of the 1D-CNN model has decreased the most severely, reaching 3.93%. Although the classification accuracy of the BiGRU model has only increased by 4.25%, it finally only remains at 96.50%.
[0109] It should be noted that in this example, the classification precision and F1-score are used to illustrate the experimental results. The classification accuracy is widely used to evaluate the performance of the constructed classifier. However, for the imbalanced classification problem, it is unreasonable to use the classification accuracy as a performance indicator to evaluate the quality of the classifier. For example, when 97% of the test set is fault samples and only 3% are normal samples, if the classifier predicts all test samples as fault samples, the classifier will still reach a high classification accuracy of 97%, but this does not mean that the classifier has good performance because it cannot correctly predict the normal samples. It should be understood that the neural network model in the embodiments of the present invention is the classifier. Therefore, in this example, two widely used evaluation indicators are used for classification, including the classification accuracy and F1-score, and the specific calculation formulas are as follows:
[0110]
[0111]
[0112] P represents the classification precision; R represents the recall rate; TP represents the samples with positive labels predicted as positive; TN represents the samples with negative labels predicted as negative; FP represents the samples with negative labels predicted as positive; FN represents the samples with positive labels predicted as negative. The F1-score is a comprehensive manifestation of the classification precision and recall rate, and its numerical value can accurately reflect the classification situation inside the samples.
[0113] In the PCS-CGCN in the embodiments of the present invention, due to the introduction of the cost-sensitive mechanism, it can better cope with the situation of sample quantity change, and its recognition accuracy only drops by 0.49%. At the same time, in the fault classification tasks of the three datasets A, B, and C, the classification accuracy of PCS-CGCN all exceeds 99%. In the fault classification task of dataset D, the average classification accuracy of PCS-CGCN reaches 99.35%, and compared with SVM, the classification accuracy is improved by 4.01%. Compared with deep learning models such as 1D-CNN and DGCN, on the basis of combining their advantages, the recognition accuracy of PCS-CGCN is improved by 3% - 4%. In addition, even if the dataset is seriously imbalanced, PCS-CGCN can achieve a relatively high classification accuracy of the minority class (i.e., F1-score) on the premise of ensuring the overall classification accuracy. In the field of mechanical fault diagnosis, the recognition of the minority class by the model is particularly important, which verifies the effectiveness of the proposed method for classifying imbalanced data again.
[0114] In order to more intuitively observe the classification results of different faults in the imbalanced dataset by various network models, Figure 7 The confusion matrices after the four models complete classification are shown. Since the quantity ratio of healthy samples to fault samples in dataset D reaches 242:9, the performance of the classifier drops, and the model then pays more attention to the majority class of healthy samples and ignores the minority class of fault samples. In the classification confusion matrix of DGCN, 10% of the samples with the true label of BF2 are misclassified as healthy samples, and due to the serious imbalance of the fault dataset, the classification of fault samples is chaotic. And in the classification confusion matrices of 1D-CNN and SVM, the misclassification problem is more serious. In the actual fault diagnosis scenario, such misclassification results will cause mechanical damage and bring huge economic losses to enterprises. However, PSC-CGCN does not have this problem. In the classification confusion matrix of PCS-CGCN, although 5% of BF2 samples and 36% of OF1 samples are misclassified as other fault labels, only post-stage fault type correction by manual is needed, and it will not bring serious consequences. This further shows that the cost-sensitive adjustment layer in PCS-CGCN is feasible for correctly classifying minority class fault samples. And this result is very attractive in the actual bearing fault diagnosis scenario. Therefore, introducing cost-sensitive learning is very effective for improving the overall classification performance of the graph convolutional network.
[0115] 2. Classification experiment of PCS-CGCN on bearing data with background noise: The running signal of the actual bearing contains certain external interference noise, and its sound pressure level will directly affect the performance of the classification model. Therefore, the robustness of the fault diagnosis model to noise is very important. To test the anti-noise interference ability of the PCS-CGCN model, Gaussian white noise is added to the original dataset D, Figure 8It shows the time-domain signal with a signal-to-noise ratio of -3 dB. Similarly, noise test samples with signal-to-noise ratios of -3 dB, -1 dB, 6 dB, and 9 dB are obtained, and the noise robustness of PCS-CGCN is verified based on these samples.
[0116] Complete the fault classification of PCS-CGCN in a noisy environment, and use multiple classification models to compare the performance with the proposed method. The classification accuracy is as Figure 9 shown, and the F1-score is shown in Table 7. Among them, Random dB means that the signal-to-noise ratio of the fault samples is randomly assigned, and its signal-to-noise ratio value ranges from -3 dB to 9 dB. As the noise power increases, the fault recognition accuracy of each model gradually decreases. However, whether it is a single signal-to-noise ratio or a variable signal-to-noise ratio situation, PCS-CGCN has better performance. Table 7 shows the F1-score of each model under noisy samples.
[0117] Table 7
[0118]
[0119] When the signal-to-noise ratio of the fault samples is -3 dB, the recognition accuracy and F1-score of the 1D-CNN, BiGRU, and DGCN models are all lower than 96%, and the recognition accuracy of the SVM model is only about 92%. However, the recognition accuracy of PCS-CGCN reaches 98%, and the F1-score is 97.82%, which means that in a strong noise environment, the PCS-CGCN proposed in the embodiment of the present invention can still complete the classification task well and has good noise robustness. When the signal-to-noise ratio of the fault samples is increased to 9 dB, the recognition accuracy and F1-score of PCS-CGCN both reach about 99%, and the recognition accuracy is improved by 4% - 6% compared with other models. In addition, in the case of variable signal-to-noise ratio, the recognition accuracy of most models is only about 94%, while the recognition accuracy of PCS-CGCN reaches 98%. In a real fault diagnosis environment, the collected fault samples are mostly impure signals, and their signal-to-noise ratios cannot be completely consistent. The variable signal-to-noise ratio situation is closer to the actual application, which also further illustrates that PCS-CGCN is more practical in the field of fault diagnosis.
[0120] 3. PCS-TCGCN experimental verification. To evaluate the effectiveness and accuracy of the PCS-TCGCN model proposed in the embodiment of the present invention for fault diagnosis, the embodiment of the present invention takes rolling bearings as the object and uses bearing data from CWRU and IMS bearing data for experimental verification.
[0121] In an actual fault diagnosis scenario, the established network model needs to have strong generalization ability and be able to adapt to fault diagnosis under different load conditions. The CWRU bearing dataset is divided, and the data under the 0hp load condition is set as the source domain sample, and the data under the 1hp, 2hp, and 3hp load conditions are set as the target domain samples. The differentiation of the sample dataset is shown in Table 8.
[0122] Table 8
[0123]
[0124]
[0125] To illustrate the feasibility and superiority of the PCS-TCGCN model proposed in the embodiments of the present invention in variable working condition fault diagnosis, comparative experiments are carried out using network models commonly used in this field at present, including: (1) In the field of domain adaptation, the transfer component analysis (TCA) based on marginal probability adaptive search for feature subspaces, and KNN is used as the classifier. (2) The joint distribution adaptation method (JDA) that combines marginal probability distribution and conditional probability distribution. (3) The one-dimensional transfer convolutional network (1D-TCNN) based on one-dimensional signals. Each experiment is repeated ten times, and the average value is taken as the final result. The classification accuracies of different network models under variable load conditions are shown in Table 9, and the F1-scores are shown in Table 10.
[0126] Table 9
[0127]
[0128] Table 10
[0129]
[0130] From the classification results in Table 9 and Table 10, it can be seen that the PCS-TCGCN model proposed in the embodiments of the present invention can accurately classify the bearing operating states under different working load conditions, and its classification accuracy is stably maintained above 99%. Compared with the domain transfer-based transfer learning methods of JDA and TCA, the classification accuracy of the transfer learning model proposed in the embodiments of the present invention has been improved by approximately 5%-6%. Compared with 1D-TCNN, the network model proposed in the embodiments of the present invention has better performance in terms of performance, and its classification accuracy has been improved by 3%-4%. Therefore, the above comparison results show that PCS-TCGCN is advanced and feasible in fault diagnosis under variable load conditions. In addition, in the transfer task of 0hp-2hp, the classification accuracy of TCA reaches 96.20%, but in the transfer task of 0hp-1hp, the classification accuracy of TCA is only 93.22%. This means that as the load changes, the time-frequency feature differences under different working conditions increase, resulting in a decrease in the classification accuracy of such feature transfer-based models. The classification accuracy of deep transfer models such as 1D-TCNN finally only maintains at about 96%, which is much lower than the classification accuracy of PCS-TCGCN. This also verifies again that PCS-TCGCN can maintain good stability in variable load fault diagnosis.
[0131] 4. Experiment to verify the generalization performance of the PCS-TCGCN network model: Use the fault signals collected under the working conditions of 1797 hp and a fault diameter of 0.007 inches in the CWRU bearing dataset as source domain samples, and the IMS (Intelligent Maintenance System) bearing dataset as target domain samples. The IMS bearing dataset is also an open-source dataset and can be obtained online. The experimental test platform of the IMS dataset is as Figure 10 shown. It can be seen from the experimental bench that four rolling bearings are installed on the main shaft, and the run-to-failure experiment is carried out under the conditions of a constant 2000 rpm and a radial load of 6000 pounds. Two accelerometers are installed in the vertical and horizontal directions respectively, and vibration signals are collected at a frequency of 20 kHz. Through three run-to-failure experiments, data samples of outer race fault (OF), inner race fault (IF), and ball fault (BF) are collected respectively. Therefore, the IMS bearing dataset contains 3 types of fault signals and normal signals. Considering the differences between the target domain and source domain samples, the fine-tuning samples and test samples of the target domain are divided, as shown in Table 11.
[0132] Table 11
[0133]
[0134] As Figure 11 shown is the bar chart of the evaluation indicators in the two datasets for the four methods. The evaluation indicators include test accuracy and F1-score. InFigure 11 The test accuracy is represented by Accuracy. To ensure the reliability of the experiment, each method was tested 10 times. From Figure 11 it can be seen that when the device changes, the accuracy and F1-score of PCS-TCGCN both reach over 98%. The accuracy and F1-score of other comparative methods for fault diagnosis are both lower than 98%. This means that after introducing transfer learning into PCS-CGCN, even when the mechanical equipment and load conditions change, it can still ensure a high correct rate of fault diagnosis, and can better classify majority-class samples and minority-class samples, making it more practical in daily industrial production.
[0135] Although the training set contains vibration data with the same deceleration trend as the test set, the acquisition environments of the training set and the test set are different. Therefore, there will be certain differences in the sample space, and the change in rotational speed will cause changes in fault characteristics. Traditional joint distribution adaptation (JDA) and transfer component analysis (TCA) models have domain drift problems when the data feature differences are large, which will affect the accuracy of fault diagnosis. Although 1D-TCNN introduces transfer learning, it is affected by the network structure level and network hyperparameters and has poor fault diagnosis effects when the device changes. In summary, PCS-TCGCN with transfer learning introduced can have good fault diagnosis performance when the device changes by using the method of pre-training an online neural network model.
[0136] To further illustrate the influence of adding transfer learning on the training convergence result of PCS-TCGCN, Figure 12 (a) in Figure 12(b) in it plots the loss curves of model training before and after introducing transfer learning. Among them, CGCN uses 1D-CNN for feature dimensionality reduction and feature map reconstruction, and uses the graph convolutional network GCN as a classifier. CGCN (Covolution and Graph Convolution Network) refers to a neural network that combines 1D-CNN and GCN. Obviously, after PCS-TCGCN is pre-trained, the convergence speed has been greatly improved. Compared with PCS-CGCN, when PCS-TCGCN reaches 60 epochs, the loss value and classification accuracy tend to be stable, while PCS-CGCN needs to reach 75 epochs to achieve the same. Moreover, the recognition accuracy of the former is also about 5% higher than that of PCS-CGCN. This means that introducing transfer learning and pre-training the online neural network model under variable device conditions is feasible for improving the convergence speed and classification effect of the model. The GCN and CGCN without introducing cost-sensitive learning and transfer learning have much lower convergence speed and classification accuracy than PCS-CGCN and PCS-TCGCN. This is because the data distributions in the target domain and source domain samples are different and the data distribution is seriously imbalanced. Therefore, the fault diagnosis effect of the model is not good. An epoch means that when training the model, all training data sets have been trained once.
[0137] The embodiment of the present invention proposes a bearing fault diagnosis method to solve the problem that the existing fault diagnosis methods have low classification accuracy for minority class samples. For the PCS-CGCN model, improvements are made on the first graph convolutional network. By introducing cost-sensitive learning, the cross-entropy loss function is adjusted, and the optimal weight adjustment factor is solved based on the PSO algorithm, which can effectively solve the misclassification problem of minority class samples caused by unbalanced fault samples. The CWRU bearing dataset is used for experimental verification. In the unbalanced sample classification experiment, the technical solution of the embodiment of the present invention has an average classification accuracy and F1-score of the model both reaching more than 98%. PCS-CGCN improves the classification accuracy of the minority class on the premise of ensuring the overall classification accuracy, effectively solving the problem of poor classification effect of unbalanced data under the condition of scarce fault samples. Moreover, the embodiment of the present invention uses the CWRU bearing dataset with added Gaussian white noise for experimental verification. In a strong noise environment, the average classification accuracy and F1-score of the technical solution of the embodiment of the present invention are maintained at about 98%. In a weak noise environment, the average classification accuracy and F1-score of PCS-CGCN reach about 99%. In a variable signal-to-noise ratio environment, the average classification accuracy and F1-score of PCS-CGCN still remain above 98%. The experimental results show that PCS-CGCN has strong noise robustness and can be applied to fault diagnosis in real environments.
[0138] Furthermore, in the embodiments of the present invention, transfer learning is introduced into the neural network model, and the weights of the offline first graph convolutional network of the offline network model are used to pre-train the online neural network model, solving the problem of poor fault diagnosis effect of the neural network model under variable working conditions. The CWRU bearing dataset and the IMS bearing dataset are used for verification. In the variable speed experiment, the average classification accuracy and F1-score of PCS-TCGCN are maintained above 99%. In the variable equipment fault diagnosis experiment, the average classification accuracy and F1-score of PCS-TCGCN are around 98%, both superior to other comparison models. The experimental results show that PCS-TCGCN has strong generalization ability and strong practicability in variable working condition fault diagnosis.
[0139] Embodiment III
[0140] Figure 13 FIG. is a schematic structural diagram of a bearing fault diagnosis device provided by an embodiment of the present invention. The bearing fault diagnosis device provided by the embodiment of the present invention can execute the bearing fault diagnosis method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method. The device includes: a feature vector acquisition module 810 and a fault diagnosis result acquisition module 820; where:
[0141] The feature vector acquisition module 810 is configured to input the vibration signal of the bearing into the feature extraction sub-model of the pre-trained neural network model to obtain a feature vector; the neural network model includes a target loss function; the neural network model is obtained by training multiple training samples, and the training samples belong to corresponding sample categories, and different sample categories have different weights. The target loss function is obtained by updating the cross-entropy loss function according to the weights of the sample categories of each training sample; the fault diagnosis result acquisition module 820 is configured to input the feature vector into the classification sub-model of the neural network to obtain multiple probability output values of the vibration signal of the bearing belonging to different sample categories, and obtain the fault condition corresponding to the vibration signal of the bearing based on the highest probability output value. Each probability output value represents the probability that the vibration signal belongs to the sample category, and the sample category represents the fault condition of the bearing vibration signal. The sample category includes the fault type and the non-fault type.
[0142] Furthermore, in the embodiments of the present invention, the feature vector acquisition module 810 is further configured to: in each iterative training process of the neural network model, input the weights of the sample categories of each training sample into the following formula to update the cross-loss function; where Loss′ is the loss value of the target loss function, and S i represents the probability output value of the i-th node in the classification node of the target loss function. The number of classification nodes is the same as the number of sample categories; the probability output value is expressed as: zi represents the output value of the i-th classification node, z v represents the output value of the v-th classification node; the weights of the sample categories are obtained through the following formula;
[0143]
[0144] where, ω m represents the weight corresponding to the m-th training sample, cnt() represents the number of training samples of the () category, δ(m) represents the sample category corresponding to the m-th training sample, K represents the total number of all training samples, V represents the number of sample types, h represents the non-fault sample category, f1 represents the first type of fault sample category, f V-1 represents the fault sample category with the least number of samples, and α is the weight adjustment factor.
[0145] Furthermore, in the embodiment of the present invention, the feature extraction sub-model includes a first graph convolutional network, and the initial parameters of the first graph convolutional network are obtained through transfer learning of an offline network model; the feature vector acquisition module 810 is further configured to: through iterative training of the offline network model, use the optimal parameters of the offline first graph convolutional network of the trained offline network model as the initial parameters of the first graph convolutional network of the feature extraction sub-model of the neural network model; the offline first graph convolutional network has the same network structure as the first graph convolutional network.
[0146] Furthermore, in the embodiment of the present invention, the feature extraction sub-model further includes a second graph convolutional network, and the output of the first graph convolutional network is used as the input of the second graph convolutional network; the device further includes: a model training module, configured to input a plurality of training samples and perform iterative training on the neural network model until the number of iterations reaches a preset number of iterations, stop the iteration, and complete the training of the neural network model.
[0147] Furthermore, in the embodiment of the present invention, the feature extraction sub-model further includes a one-dimensional convolutional neural network; the feature vector acquisition module 810 is further configured to: input the vibration signal of the bearing into the one-dimensional convolutional neural network to obtain a one-dimensional feature vector; convert the one-dimensional feature vector into a graph signal, and input the graph signal into the first graph convolutional network to obtain a first feature vector; input the first feature vector into the second graph convolutional network to obtain a feature vector.
[0148] Furthermore, in the embodiment of the present invention, the feature vector acquisition module 810 is further configured to: convert the one-dimensional feature vector into a graph signal based on the distance matrix.
[0149] Furthermore, in the embodiment of the present invention, the fault diagnosis result acquisition module 820 is further configured to: use the sample category corresponding to the highest probability output value as the fault diagnosis result of the bearing.
[0150] In the technical solution of the embodiment of the present invention, the vibration signal of the bearing is input into the feature extraction sub-model of the pre-trained neural network model to obtain a feature vector, and the feature vector is input into the classification sub-model of the neural network to obtain multiple probability output values of the vibration signal of the bearing belonging to different sample categories. Based on the highest probability output value, the fault condition corresponding to the vibration signal of the bearing is obtained. Each probability output value represents the probability that the vibration signal belongs to the sample category, and the sample category represents the fault condition of the bearing vibration signal. The sample category includes the fault type and the non-fault type. Since the target loss function of the neural network model is obtained by updating the cross-entropy loss function according to the weights of each training sample category, and there may be an imbalance problem in the training samples, that is, there are some samples with a large number and some samples with a very small number. The neural network model of the embodiment of the present invention can adjust the target loss function for the weights of the unbalanced samples, that is, obtain an updated target loss function according to the weights of each training sample category. Therefore, the neural network model trains and learns the unbalanced samples, and the classification accuracy of the trained neural network model is higher. Processing the vibration signal according to the neural network model, the obtained fault diagnosis result is more accurate.
[0151] It should be noted that the various modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional modules are only for the convenience of mutual distinction and do not limit the protection scope of the embodiment of the present invention.
[0152] Embodiment Four
[0153] Figure 14 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Figure 14 It shows a block diagram of an exemplary electronic device 50 suitable for implementing the embodiment mode of the embodiment of the present invention. Figure 14 The displayed electronic device 50 is only an example and should not bring any limitation to the functions and usage scope of the embodiment of the present invention.
[0154] As Figure 14 shown, the electronic device 50 is presented in the form of a general-purpose computing device. The components of the electronic device 50 may include, but are not limited to: one or more processors or processing units 501, a system memory 502, and a bus 503 connecting different system components (including the system memory 502 and the processing unit 501).
[0155] The bus 503 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor bus, or a local bus using any of the several bus structures. By way of example, and not limitation, such architectures include the Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0156] The electronic device 50 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 50, including both volatile and nonvolatile media, removable and non-removable media.
[0157] System memory 502 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 504 and / or cache memory 505. The electronic device 50 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 506 can be used for reading and writing non-removable, nonvolatile magnetic media ( Figure 14 not shown, and typically called a "hard disk drive"). Although Figure 14 not shown in the figures, a disk drive for reading and writing a removable nonvolatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing a removable nonvolatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) can be provided. In these instances, each drive can be connected to the bus 503 by one or more data media interfaces. Memory 502 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of the embodiments of the present invention.
[0158] A program / utility 508 having a set (at least one) of program modules 507 can be stored, for example, in memory 502, such program modules 507 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which examples or some combination thereof may include an implementation of a networking environment. The program modules 507 typically carry out the functions and / or methods of the embodiments described herein.
[0159] The electronic device 50 can also communicate with one or more external devices 509 (such as a keyboard, a pointing device, a display 510, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 50, and / or communicate with any device that enables the electronic device 50 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 511. Moreover, the electronic device 50 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 512. As shown in the figure, the network adapter 512 communicates with other modules of the electronic device 50 through a bus 503. It should be understood that although Figure 14 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 50, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0160] The processing unit 501 executes various functional applications and data processing by running programs stored in the system memory 502, for example, implementing the bearing fault diagnosis method provided by the embodiments of the present invention.
[0161] Embodiment Five
[0162] The embodiments of the present invention also provide a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a bearing fault diagnosis method when executed by a computer processor. The method includes:
[0163] Inputting the vibration signal of the bearing into the feature extraction sub-model of a pre-trained neural network model to obtain a feature vector; the neural network model includes a target loss function; the neural network model is obtained by training a plurality of training samples, the training samples belong to corresponding sample categories, different sample categories have different weights, and the target loss function is obtained by updating the cross-entropy loss function according to the weights of the sample categories of each training sample; inputting the feature vector into the classification sub-model of the neural network to obtain a plurality of probability output values of the vibration signal of the bearing belonging to different sample categories, and obtaining the fault condition corresponding to the vibration signal of the bearing based on the highest probability output value. Each probability output value represents the probability that the vibration signal belongs to the sample category, and the sample category represents the fault condition of the bearing vibration signal. The sample category includes a fault type and a non-fault type.
[0164] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0165] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0166] The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.
[0167] The computer program code for performing the operations of the embodiments of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0168] The above-disclosed is only the preferred embodiment of the present invention. Of course, the scope of rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A bearing fault diagnosis method, characterized in that, Including: Input the vibration signal of the bearing into the feature extraction sub-model of a pre-trained neural network model to obtain a feature vector; the neural network model includes an objective loss function; the neural network model is obtained by training multiple training samples, the training samples belong to corresponding sample categories, different sample categories have different weights, and the objective loss function is obtained by updating the cross-entropy loss function according to the weights of the sample categories of each training sample; Input the feature vector into the classification sub-model of the neural network to obtain multiple probability output values of the vibration signal of the bearing belonging to different sample categories, and obtain the fault condition corresponding to the vibration signal of the bearing based on the highest probability output value. Each probability output value represents the probability that the vibration signal belongs to the sample category, and the sample category represents the fault condition of the bearing vibration signal. The sample category includes fault types and non-fault types; The updating of the cross-entropy loss function according to the weights of the sample categories of each training sample includes: In each iterative training process of the neural network model, input the weights of the sample categories of each training sample into the following formula to update the cross-loss function; Among them, Loss′ is the loss value of the target loss function, and S i represents the probability output value of the i-th node among the classification nodes of the target loss function, and the number of the classification nodes is the same as the number of sample categories; The probability output value is expressed as: z i represents the output value of the i-th classification node, z v is represented as the output value of the v-th classification node; The weights of the sample categories are obtained through the following formula; Among them, ω m represents the weight corresponding to the m-th training sample, cnt() represents the number of training samples of the () category, δ(m) represents the sample category corresponding to the m-th training sample, K represents the total number of all training samples, V represents the number of sample types, h represents the non-fault sample category, f1 represents the first type of fault sample category, f V-1 represents the fault sample category with the least number of samples, and α is the weight adjustment factor.
2. The bearing fault diagnosis method according to claim 1, wherein The feature extraction sub-model includes a first graph convolutional network, and the initial parameters of the first graph convolutional network are obtained through transfer learning of an offline network model; Obtaining the initial parameters of the first graph convolutional network of the feature extraction sub-model of the neural network model through transfer learning of the offline network model includes: Through iterative training of the offline network model, take the optimal parameters of the offline first graph convolutional network of the trained offline network model as the initial parameters of the first graph convolutional network of the feature extraction sub-model of the neural network model; the network structure of the offline first graph convolutional network is the same as that of the first graph convolutional network.
3. The bearing fault diagnosis method according to claim 2, wherein The feature extraction sub-model further includes a second graph convolutional network, and the output of the first graph convolutional network is used as the input of the second graph convolutional network; After obtaining the initial parameters of the first graph convolutional network, it further includes: Input multiple training samples, perform iterative training on the neural network model until the number of iterations reaches a preset number of iterations and stop iterating to complete the training of the neural network model.
4. The bearing fault diagnosis method according to claim 3, characterized in that, The feature extraction sub-model further includes a one-dimensional convolutional neural network; The step of inputting the vibration signal of the bearing into the feature extraction sub-model of a pre-trained neural network model to obtain a feature vector includes: Input the vibration signal of the bearing into the one-dimensional convolutional neural network to obtain a one-dimensional feature vector; Convert the one-dimensional feature vector into a graph signal, and input the graph signal into the first graph convolutional network to obtain a first feature vector; Input the first feature vector into the second graph convolutional network to obtain a feature vector.
5. The bearing fault diagnosis method according to claim 4, characterized in that The step of converting the one-dimensional feature vector into a graph signal includes: Convert the one-dimensional feature vector into a graph signal based on the distance matrix.
6. The bearing fault diagnosis method according to claim 1, wherein The step of obtaining the fault condition corresponding to the vibration signal of the bearing based on the highest probability output value includes: taking the sample category corresponding to the highest probability output value as the fault diagnosis result of the bearing.
7. A bearing fault diagnosis device, characterized in that, Including: A feature vector acquisition module, configured to input the vibration signal of the bearing into the feature extraction sub-model of a pre-trained neural network model to obtain a feature vector; the neural network model includes an objective loss function; the neural network model is obtained by training multiple training samples, the training samples belong to corresponding sample categories, different sample categories have different weights, and the objective loss function is obtained by updating the cross-entropy loss function according to the weights of the sample categories of each training sample; A fault diagnosis result acquisition module, configured to input the feature vector into the classification sub-model of the neural network to obtain multiple probability output values of the vibration signal of the bearing belonging to different sample categories, and obtain the fault condition corresponding to the vibration signal of the bearing based on the highest probability output value. Each probability output value represents the probability that the vibration signal belongs to the sample category, and the sample category represents the fault condition of the bearing vibration signal. The sample category includes a fault type and a non-fault type; The updating of the cross-entropy loss function according to the weights of the sample categories of each training sample includes: In each iterative training process of the neural network model, input the weights of the sample categories of each training sample into the following formula to update the cross-loss function; Among them, Loss′ is the loss value of the target loss function, and S i represents the probability output value of the i-th node among the classification nodes of the target loss function, and the number of the classification nodes is the same as the number of sample categories; The probability output value is expressed as: z i represents the output value of the i-th classification node, z v is represented as the output value of the v-th classification node; The weights of the sample categories are obtained through the following formula; Among them, ω m represents the weight corresponding to the m-th training sample, cnt() represents the number of training samples of the () category, δ(m) represents the sample category corresponding to the m-th training sample, K represents the total number of all training samples, V represents the number of sample types, h represents the non-fault sample category, f1 represents the first type of fault sample category, f V-1 represents the fault sample category with the least number of samples, and α is the weight adjustment factor.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the bearing fault diagnosis method according to any one of claims 1-6.
9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the bearing fault diagnosis method according to any one of claims 1-6 when executed by a computer processor.