Fault self-adaptive diagnosis method of rolling bearing for rotating machinery
By employing stage probability vector embedding and adaptive attention mechanisms in rotating machinery to dynamically adjust feature attention, a multilayer perceptron is constructed for fault diagnosis. This solves the problems of missed and false alarms caused by fixed thresholds in rotating machinery fault diagnosis, and achieves more accurate fault prediction.
Patent Information
- Application Number
- CN202511116179.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-10-28
AI Technical Summary
In the diagnosis of rotating machinery faults, existing technologies often use fixed threshold methods, which can easily lead to early missed detections or late false alarms. Furthermore, neural network-based methods lack explicit modeling of the dynamic process of fault development and cannot adaptively adjust diagnostic criteria.
By mapping stage probability vectors to embedding vectors and combining them with a stage adaptive attention mechanism, feature attention is dynamically adjusted, and a multilayer perceptron is constructed for fault diagnosis. Deep learning is used to optimize feature extraction and diagnosis.
It enables automatic adjustment of diagnostic criteria based on the stage of a fault, improving the initial detection rate, reducing the late false alarm rate, and achieving more accurate predictive maintenance.
Smart Images

Figure CN120846674A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault diagnosis technology for rotating machinery, specifically an adaptive fault diagnosis method for rolling bearings used in rotating machinery. Background Technology
[0002] In the field of rotating machinery fault diagnosis, traditional methods typically employ fixed thresholds or static models to detect fault symptoms. This approach has significant limitations: firstly, in the early stages of a fault, abnormal signals are often weak, and fixed thresholds can easily lead to missed detections, causing delays in optimal maintenance; secondly, in the middle and late stages of a fault, signal characteristics become more pronounced, and the same fixed threshold may become overly sensitive, resulting in false alarms and unnecessary downtime for maintenance. While existing neural network-based diagnostic methods can automatically extract features, they generally lack explicit modeling of the dynamic process of fault development and cannot adaptively adjust diagnostic criteria according to different fault stages (early, middle, and late). Furthermore, although traditional expert systems can distinguish stages using manual rules, rule maintenance is costly and difficult to cover complex operating conditions. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide an adaptive fault diagnosis method for rolling bearings for rotating machinery, which can automatically sense the fault stage, dynamically adjust the evaluation criteria, adapt to the characteristics of industrial data, and balance the initial detection rate and the late false alarm rate, thereby achieving more accurate predictive maintenance.
[0004] The technical solution of this invention is as follows:
[0005] An adaptive fault diagnosis method for rolling bearings used in rotating machinery specifically includes the following steps:
[0006] (1) Collect vibration signals during the operation of the rolling bearing, preprocess the collected vibration signals, and obtain preprocessed vibration signal data.
[0007] (2) Perform time series modeling on the preprocessed vibration signal data, then extract time series features, and then generate stage probability vectors based on the time series features;
[0008] (3) Map the stage probability vector to the stage embedding vector, introduce the stage embedding vector into the stage adaptive attention mechanism, dynamically adjust the attention to different features, and output the optimized feature vector.
[0009] (4) Construct a fault diagnosis network based on deep learning. The fault diagnosis network uses a multilayer perceptron to diagnose the faults of rolling bearings, outputs the probability distribution of various faults of rolling bearings, selects the fault type with the highest probability as the diagnosis result, and confirms the severity of the fault type.
[0010] The preprocessing of the acquired vibration signals includes filtering, noise reduction, and normalization operations.
[0011] After time-series modeling of the preprocessed vibration signal data, time-series features are extracted using an LSTM network or a Transformer network. The extracted time-series features form a feature vector z = [z early ,z mid ,z late ], where z early Representing initial fault characteristics, z mid Representing mid-term fault characteristics, z late This represents a late-stage failure characteristic.
[0012] The generation of stage probability vectors based on time series features is specifically shown in equation (1) below:
[0013]
[0014] In equation (1), p early p represents the probability that the collected vibration signal belongs to an initial fault. mid p represents the probability that the collected vibration signal belongs to a mid-term fault. late P represents the probability that the collected vibration signal belongs to a late-stage fault. early p mid and p late The stage probability vector is composed of these.
[0015] The stage probability vector is calibrated using the Softmax temperature coefficient, and the calculation formula for the calibrated stage probability vector is shown in equation (2) below:
[0016]
[0017] In equation (2), Represents the Softmax temperature coefficient, which is a hyperparameter; This represents the initial failure probability after calibration. This represents the mid-term failure probability after calibration. This represents the probability of late failure after calibration. Representative by and The calibrated stage probability vector is formed.
[0018] The specific steps for mapping the stage probability vector to a stage embedding vector, introducing the stage embedding vector into the stage adaptive attention mechanism to dynamically adjust the attention level to different features, and outputting an optimized feature vector are as follows:
[0019] S31. Map the stage probability vector to the stage embedding vector, as shown in equation (3) below:
[0020]
[0021] In equation (3), W∈R 3×d d represents the learnable embedding matrix, where 3 corresponds to the three fault stages: early, middle, and late; d represents the dimension of the stage embedding vector; T represents the transpose; and W represents the model number. T E represents the transpose of the learnable embedding matrix; stage The stage embedding vector represents the continuity of the fault development stages and has a dimension of R. d×1 ;
[0022] S32. Introduce the stage embedding vector as a dynamic bias term into the stage adaptive attention mechanism, and calculate the attention weights Attention(Q,K,V), as shown in the following equation (4):
[0023]
[0024] In equation (4), Q, K, and V represent the query, key, and value matrices, respectively, which are calculated by multiplying the input data by the corresponding weight matrix. The input data is the stage embedding vector E. stage ;d k Represents the dimension of the key matrix K; U∈R d×n It is a learnable projection matrix, U T Let E be the transpose of U. stage and U T An n-dimensional vector is obtained through multiplication, which is then incorporated into the calculation of attention weights as a dynamic bias term.
[0025] S33. Introduce a stage attention gating mechanism, as shown in the following formula (5):
[0026]
[0027] In equation (5), Att out The optimized feature vector output by the attention gating mechanism in the representative stage, where G represents the gating vector, ⊙ represents element-wise multiplication, and W... g Represents a learnable weight matrix used to optimize the stage embedding vector E. stage Perform a linear transformation; b g It is a bias term used to adjust the result of the linear transformation; σ represents the sigmoid function.
[0028] The aforementioned fault diagnosis network uses a multilayer perceptron for rolling bearing fault diagnosis, outputting the probability distribution of various rolling bearing faults. The specific steps are as follows: The multilayer perceptron consists of an input layer, several hidden layers, and an output layer. First, the optimized feature vector Att...out The input is fed into the input layer of the fault diagnosis network, and then the feature vector Att is optimized. out After passing through several hidden layers, the feature vector Att is continuously refined and optimized through linear transformations of the weight matrix and bias terms combined with nonlinear mappings of the activation function. out The key information related to the fault type is processed, and finally the probability value of each fault type of the rolling bearing is output in the output layer through the Softmax function, thus obtaining the probability distribution of various faults of the rolling bearing. The output dimension of the output layer is consistent with the preset number of fault types.
[0029] The severity of the fault type is determined by combining the fault type in the diagnostic results (i.e., the fault type output) with the calibrated stage probability vector to confirm the severity of the fault type in the initial, middle, and late stages.
[0030] When the fault diagnosis network is trained, the input data is the optimized feature vector obtained after processing the labeled vibration signal data in steps (1)-(3). The label contains the fault type and the corresponding stage information. The objective function is to minimize the cross-entropy loss between the diagnosis result and the actual fault type label. The weight matrix and bias term of each layer of the fault diagnosis network are iteratively updated through the backpropagation algorithm until the loss function converges to the preset threshold, and the trained fault diagnosis network is obtained. The trained fault diagnosis network is used to diagnose the faults of rolling bearings used in rotating machinery.
[0031] Advantages of this invention:
[0032] This invention maps the stage probability vector to a stage embedding vector through a learnable embedding matrix. It automatically adjusts the diagnostic criteria based on the real-time calculated fault probabilities of each stage, resolving the contradiction between the traditional fixed threshold and the early stage faults (requiring low thresholds) and late stage faults (requiring high thresholds). This improves the early detection rate while reducing the late stage false alarm rate. Furthermore, the stage embedding vector is introduced into a stage adaptive attention mechanism to dynamically adjust the degree of attention to different features, thereby explicitly establishing the relationship between the fault stage and the feature weights. This overcomes the uninterpretability of the implicit learning of the fault evolution process in traditional neural networks, making the decision-making process traceable and thus achieving more accurate predictive maintenance. Attached Figure Description
[0033] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] An adaptive fault diagnosis method for rolling bearings used in rotating machinery specifically includes the following steps:
[0036] (1) Data acquisition and preprocessing: The vibration signal during the operation of the rolling bearing is acquired, and the acquired vibration signal is preprocessed (filtering, noise reduction and normalization) to obtain the preprocessed vibration signal data.
[0037] (2) Fault stage identification, which specifically includes the following steps:
[0038] S21. Perform time-series modeling on the preprocessed vibration signal data. Then, use an LSTM network or a Transformer network to extract time-series features. The extracted time-series features form a feature vector z = [z early ,z mid ,z late ], where z early Representing initial fault characteristics, z mid Representing mid-term fault characteristics, z late This represents a late-stage failure characteristic;
[0039] S22. Generate a stage probability vector based on the time series characteristics, and simultaneously use the Softmax temperature coefficient for calibration, as shown in the following formula (2):
[0040]
[0041] In equation (2), Represents the Softmax temperature coefficient, which is a hyperparameter; This represents the initial fault probability after calibration, that is, the probability that the collected vibration signal after calibration belongs to the initial fault. This represents the calibrated mid-term failure probability, that is, the probability that the calibrated vibration signal belongs to a mid-term failure. This represents the probability of a late-failure after calibration, that is, the probability that the acquired vibration signal after calibration belongs to a late-failure. Representative by and The calibrated stage probability vector;
[0042] (3) Dynamic symptom assessment: The stage probability vector is mapped to the stage embedding vector, and the stage embedding vector is introduced into the stage adaptive attention mechanism to dynamically adjust the attention to different features and output an optimized feature vector. The specific steps include:
[0043] S31. Map the stage probability vector to the stage embedding vector, as shown in equation (3) below:
[0044]
[0045] In equation (3), W∈R 3×d d represents the learnable embedding matrix, where 3 corresponds to the three fault stages: early, middle, and late; d represents the dimension of the stage embedding vector; T represents the transpose; and W represents the model number. T E represents the transpose of the learnable embedding matrix; stage The stage embedding vector represents the continuity of the fault development stages and has a dimension of R. d×1 ;
[0046] S32. Introduce the stage embedding vector as a dynamic bias term into the stage adaptive attention mechanism, and calculate the attention weights Attention(Q,K,V), as shown in the following equation (4):
[0047]
[0048] In equation (4), Q, K, and V represent the query, key, and value matrices, respectively, which are calculated by multiplying the input data by the corresponding weight matrix. The input data is the stage embedding vector E. stage ;d k Represents the dimension of the key matrix K; U∈R d×n It is a learnable projection matrix, U T Let E be the transpose of U. stage and U T An n-dimensional vector is obtained through multiplication and incorporated into the calculation of attention weights as a dynamic bias term. The stage embedding vector is introduced as a dynamic bias term into the stage adaptive attention mechanism, so that the attention distribution can be dynamically adjusted according to the current fault stage: in the early stage of the fault, the attention weight of high-frequency feature channels will be automatically enhanced; while in the late stage of the fault, the attention of temporal statistical features will be increased.
[0049] S33. Introduce a stage attention gating mechanism, as shown in the following formula (5):
[0050]
[0051] In equation (5), Att out The optimized feature vector output by the attention gating mechanism in the representative stage, where G represents the gating vector, ⊙ represents element-wise multiplication, and W...g Represents a learnable weight matrix used to optimize the stage embedding vector E. stage Perform a linear transformation; b g It is a bias term used to adjust the result of the linear transformation; σ represents the sigmoid function; Att out It is the attention output result after adjustment by the stage attention gating mechanism. It integrates the attention mechanism's focus on key features and the dynamic adjustment of stage information. It is the final output of dynamic symptom assessment. This output can adaptively adjust the attention to different features according to the fault stage, and can avoid excessive interference from stage information through the gating mechanism when the stage judgment is uncertain, thus ensuring the stability of the basic feature extraction capability.
[0052] (4) Fault Diagnosis: A fault diagnosis network based on deep learning is constructed. The fault diagnosis network uses a multilayer perceptron for rolling bearing fault diagnosis. The multilayer perceptron consists of an input layer, several hidden layers, and an output layer. First, the optimized feature vector Att is... out The input is fed into the input layer of the fault diagnosis network, and then the feature vector Att is optimized. out After passing through several hidden layers, the feature vector Att is continuously refined and optimized through linear transformations of the weight matrix and bias terms combined with nonlinear mappings of the activation function. out The key information related to the fault type is stored in the hidden layer. The number of hidden layers and the number of neurons in each layer can be adjusted according to the actual diagnostic needs (for example, when three hidden layers are set, the number of neurons in the three hidden layers are 256, 128 and 64 respectively). Finally, the probability value of each fault type of the rolling bearing is output through the Softmax function in the output layer, that is, the probability distribution of various faults of the rolling bearing is obtained. The output dimension of the output layer is consistent with the preset number of fault types.
[0053] Then, the fault type with the highest probability is selected as the diagnosis result. By combining the fault type of the diagnosis result, i.e. the fault type output by the diagnosis, with the calibrated stage probability vector, the severity of the fault type in the initial stage, the middle stage, and the late stage is confirmed.
[0054] When training the fault diagnosis network, the input data is the optimized feature vector obtained after processing the labeled vibration signal data in steps (1)-(3). The label contains the fault type (such as inner ring fault, outer ring fault, etc.) and the corresponding stage information. The objective function is to minimize the cross-entropy loss between the diagnosis result and the actual fault type label. The weight matrix and bias term of each layer of the fault diagnosis network are iteratively updated through the backpropagation algorithm until the loss function converges to the preset threshold, and the trained fault diagnosis network is obtained. The trained fault diagnosis network is used to diagnose the faults of rolling bearings used in rotating machinery.
[0055] Performance Analysis:
[0056] Using the 6205-2RS rolling bearing as the experimental object, the entire life cycle fault evolution was simulated on a test bench. Vibration signals covering the normal state and the initial stage (vibration amplitude <0.5g), middle stage (0.5-1.0g), and late stage (≥1.0g) of the fault were collected using an accelerometer (sampling frequency 10kHz). A total of 3000 sets of samples were obtained (2400 sets for training and 600 sets for testing). The fault adaptive diagnosis method (Ours) of this invention was compared with the fault diagnosis method (MLP) based on the existing MLP neural network and the fault diagnosis method (LSTM-MLP) based on the combination of LSTM and MLP. The evaluation indicators included the stage identification accuracy, the fault type diagnosis accuracy, and the F1-Score (an indicator that measures the performance of classification problems and is the harmonic mean of precision and recall) for different stages / types, as shown in Table 1 below. As shown in Table 1, the accuracy of the method of the present invention (Ours) reaches 91.7%, which is 9.2% higher than that of LSTM-MLP. The F1-Score in the initial stage increases from 0.71 to 0.87. The accuracy of fault type diagnosis of the present invention reaches 92.8%, which is significantly higher than that of MLP (76.2%) and LSTM-MLP (83.3%). The F1-Score for inner ring faults increases from 0.64 to 0.88, and the F1-Score for late rolling element faults remains stable at 0.93. The diagnostic performance of the present invention is better in all stages and types of faults throughout the entire cycle.
[0057] Table 1
[0058]
[0059] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An adaptive fault diagnosis method for rolling bearings used in rotating machinery, characterized in that: Specifically, it includes the following steps: (1) Collect vibration signals during the operation of the rolling bearing, preprocess the collected vibration signals, and obtain preprocessed vibration signal data. (2) Perform time series modeling on the preprocessed vibration signal data, then extract time series features, and then generate stage probability vectors based on the time series features; (3) Map the stage probability vector to the stage embedding vector, introduce the stage embedding vector into the stage adaptive attention mechanism, dynamically adjust the attention to different features, and output the optimized feature vector. (4) Construct a fault diagnosis network based on deep learning. The fault diagnosis network uses a multilayer perceptron to diagnose the faults of rolling bearings, outputs the probability distribution of various faults of rolling bearings, selects the fault type with the highest probability as the diagnosis result, and confirms the severity of the fault type.
2. The adaptive fault diagnosis method for rolling bearings for rotating machinery according to claim 1, characterized in that: The preprocessing of the acquired vibration signals includes filtering, noise reduction, and normalization operations.
3. The adaptive fault diagnosis method for rolling bearings for rotating machinery according to claim 1, characterized in that: After time-series modeling of the preprocessed vibration signal data, time-series features are extracted using an LSTM network or a Transformer network. The extracted time-series features form a feature vector z = [z early ,z mid ,z late ], where z early Representing initial fault characteristics, z mid Representing mid-term fault characteristics, z late This represents a late-stage failure characteristic.
4. The adaptive fault diagnosis method for rolling bearings in rotating machinery according to claim 3, characterized in that: The generation of stage probability vectors based on time series features is specifically shown in equation (1) below: In equation (1), p early p represents the probability that the collected vibration signal belongs to an initial fault. mid p represents the probability that the collected vibration signal belongs to a mid-term fault. late P represents the probability that the collected vibration signal belongs to a late-stage fault. early p mid and p late The stage probability vector is composed of these.
5. The adaptive fault diagnosis method for rolling bearings in rotating machinery according to claim 4, characterized in that: The stage probability vector is calibrated using the Softmax temperature coefficient, and the calculation formula for the calibrated stage probability vector is shown in equation (2) below: In equation (2), Represents the Softmax temperature coefficient, which is a hyperparameter; This represents the initial failure probability after calibration. This represents the mid-term failure probability after calibration. This represents the probability of late failure after calibration. Representative by and The calibrated stage probability vector is formed.
6. The adaptive fault diagnosis method for rolling bearings for rotating machinery according to claim 5, characterized in that: The specific steps for mapping the stage probability vector to a stage embedding vector, introducing the stage embedding vector into the stage adaptive attention mechanism to dynamically adjust the attention level to different features, and outputting an optimized feature vector are as follows: S31. Map the stage probability vector to the stage embedding vector, as shown in equation (3) below: In equation (3), W∈R 3×d d represents the learnable embedding matrix, where 3 corresponds to the three fault stages of early, middle and late stages, and d represents the dimension of the stage embedding vector. T represents transpose, W T E represents the transpose of the learnable embedding matrix; stage The stage embedding vector represents the continuity of the fault development stages and has a dimension of R. d×1 ; S32. Introduce the stage embedding vector as a dynamic bias term into the stage adaptive attention mechanism, and calculate the attention weights Attention(Q,K,V), as shown in the following equation (4): In equation (4), Q, K, and V represent the query, key, and value matrices, respectively, which are calculated by multiplying the input data by the corresponding weight matrix. The input data is the stage embedding vector E. stage ;d k Represents the dimension of the key matrix K; U∈r d×n It is a learnable projection matrix, U T Let E be the transpose of U. stage and U T An n-dimensional vector is obtained through multiplication, which is then incorporated into the calculation of attention weights as a dynamic bias term. S33. Introduce a stage attention gating mechanism, as shown in the following formula (5): In equation (5), Att out The optimized feature vector output by the attention gating mechanism in the representative stage, where G represents the gating vector, ⊙ represents element-wise multiplication, and W... g Represents a learnable weight matrix used to optimize the stage embedding vector E. stage Perform a linear transformation; b g It is a bias term used to adjust the result of the linear transformation; σ represents the sigmoid function.
7. The adaptive fault diagnosis method for rolling bearings in rotating machinery according to claim 6, characterized in that: The aforementioned fault diagnosis network uses a multilayer perceptron for rolling bearing fault diagnosis, outputting the probability distribution of various rolling bearing faults. The specific steps are as follows: The multilayer perceptron consists of an input layer, several hidden layers, and an output layer. First, the optimized feature vector Att... out The input is fed into the input layer of the fault diagnosis network, and then the feature vector Att is optimized. out After passing through several hidden layers, the feature vector Att is continuously refined and optimized through linear transformations of the weight matrix and bias terms combined with nonlinear mappings of the activation function. out The key information related to the fault type is processed, and finally the probability value of each fault type of the rolling bearing is output in the output layer through the Softmax function, thus obtaining the probability distribution of various faults of the rolling bearing. The output dimension of the output layer is consistent with the preset number of fault types.
8. The adaptive fault diagnosis method for rolling bearings in rotating machinery according to claim 7, characterized in that: The severity of the fault type is determined by combining the fault type in the diagnostic results (i.e., the fault type output) with the calibrated stage probability vector to confirm the severity of the fault type in the initial, middle, and late stages.
9. The adaptive fault diagnosis method for rolling bearings in rotating machinery according to claim 7, characterized in that: When the fault diagnosis network is trained, the input data is the optimized feature vector obtained after processing the labeled vibration signal data in steps (1)-(3). The label contains the fault type and the corresponding stage information. The objective function is to minimize the cross-entropy loss between the diagnosis result and the actual fault type label. The weight matrix and bias term of each layer of the fault diagnosis network are iteratively updated through the backpropagation algorithm until the loss function converges to the preset threshold, and the trained fault diagnosis network is obtained. The trained fault diagnosis network is used to diagnose the faults of rolling bearings used in rotating machinery.
Citation Information
Patent Citations
Glaucoma image classification method and system based on convolutional neural network
CN118379552A
Rolling bearing fault diagnosis method based on MTF-GN-DARN
CN118535998A
Method and system for generating time series of hybrid diffusion model
CN119829943A
Convolutional Self-encoding Fault Monitoring Method Based on Batch Imaging
US20220044124A1
Cited By
Fault diagnosis and remote monitoring system and method for solar power supply system
CN121055895A