Rolling Bearing Degradation Trend Prediction Method Based on Transfer Learning and Ensemble Learning
By applying transfer learning and integrated learning in rolling bearing degradation trend prediction, degradation trend characteristics under complex operating conditions are extracted, and the problem of low prediction accuracy in the prior art is solved, and higher prediction accuracy and stronger applicability are achieved.
Patent Information
- Application Number
- CN202210321096.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-03-29
AI Technical Summary
The prior art is difficult to effectively extract the degradation trend characteristics in the vibration signals of rolling bearings under complex operating conditions, resulting in low prediction accuracy and inability to meet the needs of industrial applications.
Using transfer learning and ensemble learning methods, sub-model training and model fusion prediction are used to extract degradation trend information under complex operating conditions to improve prediction accuracy. Specific steps include: feature extraction, domain discrimination, prediction training, similarity calculation and weight evaluation.
It improves the accuracy of predicting the degradation trend of rolling bearings under complex working conditions, enhances the generalization ability and practical application value of the model, and can better adapt to diagnostic tasks under different working conditions.
Smart Images

Figure CN114881069B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of degradation trend prediction based on machine learning, and in particular to a rolling bearing degradation trend prediction method based on transfer learning and ensemble learning. Background Art
[0002] Rolling bearings are one of the most widely used components in industrial production, and their health status is directly related to the safe and stable operation of mechanical equipment. However, when operating for a long time under complex working conditions of high speed and high load, rolling bearings are prone to failures such as wear and spalling, which can lead to the rapid degradation of rotating machinery and even cause serious economic losses and casualties. Therefore, it is of great practical significance to predict the degradation trend of rolling bearings.
[0003] At present, domestic and foreign scholars have done a lot of research on the degradation trend prediction problem of rolling bearings. However, in the face of complex and changeable working conditions and strong noise interference, the vibration signal is non-linear, non-stationary and strongly coupled, and the fault characteristics are diverse and not obvious in time scale. Therefore, the time-series variability in the vibration signal of rolling bearings poses higher requirements for the generalization performance of the degradation trend prediction model and poses a severe challenge to the degradation trend prediction based on vibration signals. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a rolling bearing degradation trend prediction method based on transfer learning and ensemble learning, aiming to improve the diversity of feature learning and the generalization performance under variable working conditions, so as to improve the prediction accuracy of the degradation trend of rolling bearings under complex working conditions, thereby better meeting the needs of engineering applications.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A rolling bearing degradation trend prediction method based on transfer learning and ensemble learning, comprising:
[0007] S1. Sub-model training, including:
[0008] S11. Input the original vibration signals of the source domain and the target domain into the feature extractor to extract high-level abstract features;
[0009] S12. Input all the high-level abstract features of the source domain and the target domain into the domain discriminator. The domain discriminator attempts to identify which domain the features come from. By setting a gradient reversal layer between the feature extractor and the domain discriminator, it is urged that the high-level abstract features extracted by the feature extractor have domain invariance;
[0010] S13. Input the high-level abstract features of the source domain into the predictor for future time-step prediction training to make the high-level features have prediction characteristics;
[0011] S14. Repeat S11 to S13 to complete the training of multiple sub-models;
[0012] S2. Model fusion prediction, including:
[0013] S21. Calculate the similarity between the historical actual trend sequence of bearing vibration data and the historical predicted output sequence at the same time step of each sub-model to obtain the similarity coefficient belonging to each sub-model:
[0014] S22. Convert the similarity coefficient into the weight of each sub-model;
[0015] S23. Weight the predicted output sequences of each sub-model at future time steps to obtain the predicted output sequence at future time steps after model fusion.
[0016] A further technical solution is:
[0017] The similarity calculation includes spatial similarity calculation and distribution similarity calculation. The spatial similarity and distribution similarity are added after being normalized respectively to obtain the similarity coefficient;
[0018] Use the dynamic time warping algorithm to calculate the spatial similarity:
[0019]
[0020] s.t. ε k,C ∈ {0, 1}, ε R ∈ {0, 1}
[0021] where the argmin function is used to calculate the value of the variable when the objective function takes the minimum value; q k,C and q R are two data sequences, representing the historical predicted output sequence and the historical actual trend sequence of sub-model k respectively, ε k,C and ε R are the corresponding binary matrices;
[0022] Use Jensen-Shannon divergence to calculate the distribution similarity:
[0023]
[0024]
[0025] where P k,C and P R are the probability distributions corresponding to the two data sequences q k,C and q R respectively, and x is the value in the sequence.
[0026] The calculation formula for converting the similarity coefficient into the weights of each sub - model using a weight evaluator is as follows:
[0027]
[0028] Among them, w k is the weight of sub - model k, μ k is the similarity coefficient of sub - model k obtained after similarity calculation, and K is the total number of sub - models.
[0029] The feature extractor includes three convolutional blocks and a fully - connected layer. Each convolutional block contains a one - dimensional convolutional layer, a BN layer, and a ReLU activation layer.
[0030] The domain discriminator includes two parts each with a fully - connected layer, and the last fully - connected layer is equipped with a softmax activation function to achieve domain classification.
[0031] The beneficial effects of the present invention are as follows:
[0032] The present invention combines transfer learning and ensemble learning and applies them to degradation trend prediction. It can effectively extract the degradation trend information in vibration signals under complex and variable working conditions and strong noise interference, and has a high prediction accuracy. The present invention overcomes the problem that the existing methods have difficulty in extracting bearing degradation features under variable working conditions, resulting in poor prediction performance, and improves the generalization ability and practical application value of the prediction model.
[0033] The present invention calculates the sequence similarity coefficient by jointly calculating the spatial similarity and distribution similarity, thereby weighting the sub - models to form a multi - sub - model fusion network, which has the characteristics of being modular and easy to expand. For different diagnostic tasks, the structure can be adjusted flexibly (by increasing the number of sub - models), and the applicability is stronger.
[0034] Other features and advantages of the present invention will be described in the subsequent specification, and some of them will become obvious from the specification or be understood by implementing the present invention. Brief Description of the Drawings
[0035] Figure 1 is a schematic diagram of the sub - model structure of an embodiment of the present invention.
[0036] Figure 2 is a schematic diagram of model fusion prediction of an embodiment of the present invention.
[0037] Figure 3 is the prediction result of the calculation example of an embodiment of the present invention. Detailed Embodiments
[0038] The following describes the detailed embodiments of the present invention with reference to the accompanying drawings.
[0039] A rolling bearing degradation trend prediction method based on transfer learning and ensemble learning in this embodiment includes:
[0040] S1. Sub-model training, refer to Figure 1 , including:
[0041] S11. Input the original vibration signals of the source domain and the target domain into the feature extractor to extract high-level abstract features;
[0042] S12. Input all the high-level abstract features of the source domain and the target domain into the domain discriminator. The domain discriminator tries to identify which domain the features come from. By setting a gradient reversal layer between the feature extractor and the domain discriminator, it is promoted that the high-level abstract features extracted by the feature extractor have domain invariance;
[0043] S13. Input the high-level abstract features of the source domain into the predictor for future time step prediction training to make the high-level features have prediction characteristics;
[0044] S14. Repeat S11 to S13 to complete the training of multiple sub-models;
[0045] S2. Model fusion prediction, refer to Figure 2 , including:
[0046] S21. Calculate the similarity between the historical actual trend sequence of the bearing vibration data and the historical prediction output sequence at the same time step (t0 - t1) of each sub-model to obtain the similarity coefficient belonging to each sub-model;
[0047] S22. Convert the similarity coefficient into the weight of each sub-model;
[0048] S23. Weight the future time step (t1 - t2) prediction output sequences of each sub-model to obtain the future time step prediction output sequence after model fusion.
[0049] In this embodiment, multiple degradation data sets collected from other bearings are first applied, combined with the vibration data of the target bearing, to respectively learn the common degradation knowledge of the target domain and each source domain, and multiple degradation trend prediction sub-models based on transfer learning are constructed. In the prediction process, the idea of ensemble learning is used to dynamically and adaptively fuse each sub-model, and a model fusion weighting strategy is designed to improve the adaptability of the model to working condition changes.
[0050] The feature extractor is constructed based on a convolutional neural network and includes three convolutional blocks and a fully connected layer. Each convolutional block has a one-dimensional convolutional layer, a BN layer, and a ReLU activation layer.
[0051] Specifically, the convolution kernel parameters (size × number) of the three convolutional layers are 7×32, 5×64, and 3×128 respectively. The final fully connected layer has 512 neurons.
[0052] The number of convolutional layers, the convolution kernel parameters, and the number of neurons in the fully connected layer are not limited to this and can be adjusted specifically according to the actual prediction task.
[0053] The domain discriminator includes two fully connected layers respectively, and the last fully connected layer is equipped with a softmax activation function to achieve domain classification.
[0054] Specifically, the domain discriminator consists of two fully connected layers with 100 neurons and 2 neurons respectively.
[0055] The training method of the sub-model in this embodiment is the "adversarial training" mode:
[0056] The feature extractor extracts high-level abstract features from the original vibration signals from the source domain and the target domain. What the feature extractor attempts to extract are domain-invariant features (i.e., applicable to both the source domain and the target domain). All the high-level abstract features from the source domain and the target domain are input into the domain discriminator. The domain discriminator attempts to classify the features, that is, to determine whether the feature comes from the source domain or the target domain. By setting a gradient reversal layer between the feature extractor and the discriminator, an adversarial game is formed between the feature extractor and the domain discriminator, prompting the features extracted by the feature extractor to be domain-invariant. That is, when the discriminator cannot determine whether the feature comes from the source domain or the target domain, it means that the features extracted by the feature extractor have become domain-invariant.
[0057] Therefore, the purpose of adopting "adversarial training" is to align the data distributions of the source domain and the target domain through adversarial training to extract high-level abstract features with domain-invariance.
[0058] The high-level abstract features of the source domain are input into the predictor for supervised prediction training, so that the high-level features have prediction characteristics.
[0059] After the sub-model is trained, the finally extracted high-level abstract features can not only be used for accurate trend prediction but also have domain-invariance applicable to both the source domain and the target domain.
[0060] The model fusion method in this embodiment is characterized by jointly calculating the spatial similarity and the distribution similarity to form a sequence similarity coefficient, thereby weighting the sub-models.
[0061] Specifically, the similarity calculation includes spatial similarity calculation and distribution similarity calculation. The spatial similarity and the distribution similarity are added after being normalized respectively to obtain the similarity coefficient.
[0062] Specifically, the dynamic time warping (DTW) algorithm is used to calculate the spatial similarity:
[0063]
[0064] s.t. ε k,C ∈ {0, 1}, ε R ∈ {0, 1}
[0065] where the argmin function is used to calculate the variable value when the objective function takes the minimum value; q k,C and q R are two data sequences, representing the historical predicted output sequence and the historical actual trend sequence of sub-model k respectively, ε k,C and ε R are the corresponding binary matrices;
[0066] Specifically, the Jensen-Shannon (JS) divergence is used to calculate the distribution similarity:
[0067]
[0068]
[0069] where P k,C , P R are the probability distributions corresponding to the two data sequences of q k,C , q R respectively, and x is the value in the sequence.
[0070] After the similarity calculation of the sequence, the similarity coefficients {μ1, μ2,..., μ K} belonging to each sub-model are obtained.
[0071] Specifically, the calculation formula for converting the similarity coefficient into the weight of each sub-model by applying the weight evaluator is:
[0072]
[0073] where w k is the weight of sub-model k, μ k is the similarity coefficient of sub-model k obtained after the similarity calculation, and K is the total number of sub-models.
[0074] Specifically, the strategy for weighting the future time step prediction output sequences of each sub-model is: multiply the future time step prediction output of each sub-model by the corresponding weight and sum and average.
[0075] The technical effects of this embodiment are illustrated below through a specific calculation example.
[0076] The experimental analysis data comes from the measured bearing accelerated life experiment data of the Intelligent Maintenance Systems Center jointly established by the University of Wisconsin and the University of Michigan in the United States. The rolling bearing is driven by an AC motor through a friction belt, with the rotational speed maintained at 2000 rpm. The accelerometer sampling frequency is 20 kHz, and samples are taken every 10 minutes, with each acquisition lasting 1 s. Therefore, each sample contains 20,480 data points. The bearing accelerated life experiment was conducted for 7 days, recording the vibration data collected by the accelerometer from the start of operation to the failure of the bearing. Finally, the outer ring of bearing 1 failed, and bearings 2, 3, and 4 all showed degradation. 984 samples were collected for each bearing.
[0077] Taking the data of rolling bearings 1, 3, and 4 as the source domain and the data of rolling bearing 2 as the target domain, with the training set accounting for 75% and the test set accounting for 25%, the test prediction results of the target domain are as Figure 3 shown. The absolute error and mean square error are used as the criteria for measuring the curve accuracy. It can be seen that the proposed model can achieve accurate prediction both in the stable region and in the rapid degradation stage, with the absolute error and mean square error of its curve being 0.2217 and 0.1299 respectively.
[0078] From the prediction results, it can be known that the rolling bearing degradation trend prediction method based on transfer learning and ensemble learning in this embodiment has better prediction performance when facing complex and variable working conditions compared with the traditional degradation trend prediction method, and is more in line with the actual scenario of industrial applications.
[0079] This application is based on transfer learning and ensemble learning, which improves the diversity of the characteristics representing the bearing degradation trend, reduces the risk of overfitting in model prediction, and meets the requirements of the industrial Internet of Things for high prediction accuracy and wide applicability of the degradation prediction model.
[0080] Those of ordinary skill in the art can understand that the above description is only the preferred embodiment of the present invention and is not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A rolling bearing degradation trend prediction method based on transfer learning and ensemble learning, characterized in that, Including: S1. Sub-model training, including: S11. Input the original vibration signals of the source domain and the target domain into the feature extractor to extract high-level abstract features; S12. Input all the high-level abstract features of the source domain and the target domain into the domain discriminator. The domain discriminator tries to identify which domain the features come from. By setting a gradient reversal layer between the feature extractor and the domain discriminator, it prompts the high-level abstract features extracted by the feature extractor to have domain invariance; S13. Input the high-level abstract features of the source domain into the predictor for future time step prediction training to make the high-level features have prediction characteristics; S14. Repeat S11 to S13 to complete the training of multiple sub-models; S2. Model fusion prediction, including: S21. Calculate the similarity between the historical actual trend sequence of the bearing vibration data and the historical prediction output sequence at the same time step of each sub-model respectively to obtain the similarity coefficient belonging to each sub-model; S22. Convert the similarity coefficient into the weights of each sub-model; S23. Weight the future time step prediction output sequences of each sub-model to obtain the future time step prediction output sequence after model fusion; The similarity calculation includes spatial similarity calculation and distribution similarity calculation. The similarity coefficient is obtained by adding the normalized spatial similarity and distribution similarity; The dynamic time warping algorithm is used to calculate the spatial similarity: s.t. ε k,C ∈ {0, 1}, ε R ∈ {0, 1} Among them, the argmin function is used to calculate the variable value when the objective function takes the minimum value; q k,C and q R are two data sequences, representing the historical predicted output sequence and the historical actual trend sequence of the sub-model k respectively, ε k,C and ε R are the corresponding binary matrices; The Jensen-Shannon divergence is used to calculate the distribution similarity: Among them, P k,C , P R are the probability distributions corresponding to the two data sequences of q k,C , q R respectively, and x is the value in the sequence.
2. The method for predicting the degradation trend of rolling bearings based on transfer learning and ensemble learning according to claim 1, wherein, The calculation formula for converting the similarity coefficient into the weights of each sub-model by applying the weight evaluator is as follows: where, w k is the weight of sub-model k, and μ k is the similarity coefficient of sub-model k obtained after similarity calculation, and K is the total number of sub-models.
3. The method for predicting the degradation trend of rolling bearings based on transfer learning and ensemble learning according to claim 1, characterized in that The feature extractor includes three convolutional blocks and a fully connected layer. Each convolutional block has a one-dimensional convolutional layer, a BN layer, and a ReLU activation layer.
4. The method for predicting the degradation trend of a rolling bearing based on transfer learning and ensemble learning according to claim 1, wherein The domain discriminator includes two fully connected layers respectively, and the last fully connected layer has a softmax activation function to achieve domain classification.
Citation Information
Patent Citations
Method for predicting residual life of bearing based on transfer migration
CN112685857A
Rotating machinery small sample fault diagnosis method based on generative adversarial network
CN114091504A