Rolling bearing residual life prediction method, system and device and storage medium
By using the Markov decision-making process model of ResNet network framework and PPO algorithm in the remaining life prediction of rolling bearings, the problems of feature extraction relying on manual experience, instability in the training process, and low gradient propagation efficiency in the prior art are solved, and higher prediction accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510275654.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-10
AI Technical Summary
In the prediction of residual life of rolling bearings, the problem of feature extraction relying on manual experience, unstable training process, and low gradient propagation efficiency in the prior art, resulting in limited prediction accuracy.
Using the Markov decision-making process model based on the ResNet network framework, combined with the PPO algorithm, a dynamic reward function with prediction errors as feedback is designed, and the degradation mode of rolling bearings is adaptively learned, and a dynamic state space and a continuous action space based on Gaussian distribution parameterization is constructed.
It significantly improves the robustness and accuracy of the prediction model, can adaptively capture the nonlinear degradation process under complex operating conditions, and improves the accuracy of residual life prediction.
Smart Images

Figure CN120217849A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mechanical fault prediction, and particularly to a method, system, device and storage medium for predicting the remaining life of a rolling bearing. Background Art
[0002] Rolling bearings are the core components of rotating machinery, and their health status directly affects the reliability and safety of equipment operation. Predicting the remaining life is a key link in predictive maintenance. Traditional methods mainly rely on two technologies: statistical models and traditional deep reinforcement learning algorithms. Among them, statistical models (such as Weibull distribution, Hidden Markov Model) need to manually extract features and are difficult to capture the non-linear degradation process under complex working conditions; traditional deep reinforcement learning (DRL) algorithms (such as DQN, DDPG) are prone to unstable policy updates, gradient disappearance or explosion problems in continuous action spaces, resulting in limited prediction accuracy.
[0003] The deficiencies of the prior art are mainly reflected in:
[0004] Feature extraction depends on manual experience: Traditional methods rely on expert knowledge for signal feature extraction and have poor generalization ability. The training process is unstable: Traditional DRL algorithms are vulnerable to noise interference in complex environments, and it is difficult to control the amplitude of policy updates. Low gradient propagation efficiency: Deep networks are prone to gradient disappearance during training, affecting the model convergence speed and prediction accuracy. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, system, device and storage medium for predicting the remaining life of a rolling bearing, which can adaptively learn the degradation mode of the rolling bearing and significantly improve the robustness and accuracy of the prediction model.
[0006] To achieve the above purpose, the present invention provides the following solutions:
[0007] A method for predicting the remaining life of a rolling bearing, comprising:
[0008] Obtain the vibration signal of the rolling bearing and the corresponding detection time, and construct a dynamic state space and a continuous action space parameterized based on Gaussian distribution;
[0009] Design a dynamic reward function with prediction error as feedback, and based on the ResNet network framework, construct a Markov decision process model according to the state space, the action space and the reward function, and use the PPO algorithm to iteratively train the Markov decision process model to obtain a trained Markov decision process model;
[0010] Use the trained Markov decision process model to predict the remaining life of the rolling bearing.
[0011] Optionally, the state space is represented as In each training cycle (traversal) e, m state-time pairs are randomly sampled as the state set for trajectory prediction; where s i is the vibration signal, y i is the detection time, s t is the current state, y t is the actual remaining life, i is the initial monitoring time point, and T is the final monitoring time point.
[0012] Optionally, the action space is represented as:
[0013] Λ = μ + σ
[0014] where μ is the mean and σ is the standard deviation.
[0015] Optionally, the reward function is represented as:
[0016] r t = -1 × |a t - y t |
[0017] where a t is the predicted value of the agent, and y t is the actual remaining life.
[0018] Optionally, the Markov decision process model includes a policy network and a value network; both the policy network and the value network are constructed based on the ResNet network framework, and the mathematical expression of the residual block structure of the ResNet network framework is: H(x) = F(x) + x, where F(x) represents the residual component and x is the input information.
[0019] Optionally, the ResNet network framework specifically includes:
[0020] Input layer: Accepts a one-dimensional vibration signal and outputs a shape of (bs, 1, s_len(4096));
[0021] Wide convolutional layer: Used to capture short-term dependence features of the signal, and outputs a shape of (bs, 32, 512);
[0022] Batch normalization layer: Used to accelerate model convergence, and outputs a shape of (bs, 32, 512);
[0023] Max pooling layer: Used for dimensionality reduction, and outputs a shape of (bs, 32, 256);
[0024] 8 residual blocks: Each residual block contains a convolutional layer, a batch normalization layer, and an activation function, and the output shapes are (bs, 32, 128), (bs, 64, 64), (bs, 64, 64), (bs, 128, 32), (bs, 128, 32), (bs, 64, 16), (bs, 64, 16), (bs, 32, 8) in sequence;
[0025] Global average pooling layer: used to integrate spatial information, and the output shape is (bs, 32);
[0026] Fully connected layer: outputs the policy value or value estimate, and the output shape is (bs, 2) or (bs, 1).
[0027] Optionally, the objective function of the PPO algorithm is:
[0028]
[0029] where ε is the clipping factor, is the policy ratio, s t is the current state, a t is the prediction value of the agent, θ is the network parameter, π θ is the old policy, π′ θ is the new policy, E t is the expectation for time step t, is the multi-step advantage function, and the calculation formula is:
[0030]
[0031] where γ is the discount factor, L is the multi-step learning length, V φ is the output of the value network.
[0032] The present invention also provides a rolling bearing remaining life prediction system, including:
[0033] Space construction unit, used to obtain the vibration signal of the rolling bearing and the corresponding detection time, and construct a dynamic state space and a continuous action space parameterized based on Gaussian distribution;
[0034] Model construction and training unit, used to design a dynamic reward function with prediction error as feedback, and based on the ResNet network framework, construct a Markov decision process model according to the state space, the action space, and the reward function, and use the PPO algorithm to iteratively train the Markov decision process model to obtain a trained Markov decision process model;
[0035] Model prediction unit, used to predict the remaining life of the rolling bearing by using the trained Markov decision process model.
[0036] The present invention also provides an electronic device, including a memory and a processor. The memory is used for storing a computer program, and the processor runs the computer program to enable the electronic device to execute the rolling bearing remaining life prediction method according to the above.
[0037] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the rolling bearing remaining life prediction method as described above is implemented.
[0038] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0039] The present invention discloses a rolling bearing remaining life prediction method, system, device and storage medium. The method includes obtaining the vibration signal of a rolling bearing and the corresponding detection time, and constructing a dynamic state space and a continuous action space parameterized based on Gaussian distribution; designing a dynamic reward function with prediction error as feedback, and based on the ResNet network framework, constructing a Markov decision process model according to the state space, the action space and the reward function, and using the PPO algorithm to iteratively train the Markov decision process model to obtain a trained Markov decision process model; using the trained Markov decision process model to predict the remaining life of the rolling bearing. The present invention can adaptively learn the degradation mode of the rolling bearing, and significantly improve the robustness and accuracy of the prediction model. Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a basic concept diagram of reinforcement learning in this embodiment;
[0042] Figure 2 It is an architecture diagram of the PPO-ResNet remaining life prediction model in this embodiment;
[0043] Figure 3 It is a structure diagram of the policy value network in this embodiment;
[0044] Figure 4 It is a structure diagram of the residual block in this embodiment;
[0045] Figure 5This is a line graph showing the changes and comparison of the predicted and true values of the remaining useful life (RUL) of a bearing based on PPO-ResNet in this embodiment; among them, part (a) is a schematic diagram of the outer ring; part (b) is a schematic diagram of the rolling element; part (c) is a schematic diagram of the inner ring. Detailed implementation manners
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0047] The purpose of the present invention is to provide a method, system, device and storage medium for predicting the remaining useful life of a rolling bearing, which can adaptively learn the degradation mode of the rolling bearing and significantly improve the robustness and accuracy of the prediction model.
[0048] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0049] As Figure 1 shown, the present invention provides a method for predicting the remaining useful life of a rolling bearing, including:
[0050] Step 100: Obtain the vibration signal of the rolling bearing and the corresponding detection time, and construct a dynamic state space and a continuous action space parameterized based on Gaussian distribution.
[0051] Step 200: Design a dynamic reward function with prediction error as feedback, and based on the ResNet network framework, construct a Markov decision process model according to the state space, the action space and the reward function, and use the PPO algorithm to iteratively train the Markov decision process model to obtain a trained Markov decision process model.
[0052] Step 300: Use the trained Markov decision process model to predict the remaining useful life of the rolling bearing.
[0053] Among them, the PPO algorithm reduces the impact of the update amplitude on the agent's learning process by constraining the ratio between the old and new policies, and introduces an objective function for mini-batch updates and multi-step training. The Markov decision process model includes a policy network and a value network. The output layers of the policy network and the value network are respectively mapped to policy values and value estimates through fully connected layers, and the output dimensions are 2 (policy) and 1 (value) respectively. The final output of the ResNet module compresses the spatial dimension through a global average pooling layer (Global Avgpool) and connects to a fully connected layer to generate a prediction result.
[0054] As a specific implementation, each of the above steps is described in detail.
[0055] State space definition: The state space consists of the vibration signal of the rolling bearing and the monitoring time, and is expressed as where s i is the vibration signal, and y i is the detection time. Each state s t represents the vibration signal and the corresponding detection time at time step t.
[0056] Action space definition: The action space is a continuous space, representing the predicted value of the remaining life by the agent. The action a t is parameterized by a Gaussian distribution and is expressed as Λ = μ + σ, where μ is the mean and σ is the standard deviation. The agent selects an action a t at each time step t according to the current state s t , that is, the predicted value of the remaining life.
[0057] Reward function design: The reward function r t is used to evaluate the prediction accuracy of the agent at time step t, and the calculation formula is r t = -1 × |a t - y t |, where a t is the predicted value of the agent, and y t is the actual remaining life. The goal of the reward function is to guide the agent to optimize the prediction strategy so that its predicted value is as close as possible to the actual value.
[0058] Construction of policy network and value network: The policy network is used to generate the probability distribution of the action a t and is expressed as π θ (a t |s t), where θ is the network parameter. The policy network is constructed based on the ResNet architecture. Through residual connections and Batch Normalization (BN) layers, it effectively alleviates the problem of gradient vanishing, improving the model's convergence speed and generalization performance. The ResNet module contains 8 residual blocks, and the structure of each residual block is H(x) = F(x) + x, where F(x) consists of a convolutional layer, a batch normalization layer, and a ReLU activation function.
[0059] The value network is used to estimate the value V t of the state s φ (s t ), where φ is the network parameter. The value network is also constructed based on the ResNet architecture, with a structure similar to the policy network, and is used to evaluate the long-term return of the current state.
[0060] The policy network and the value network generate trajectory data τ=(s0,a0,r0;s1,a1,r1;…) by interacting with the environment, and optimize the network parameters based on the PPO algorithm.
[0061] Application of the PPO algorithm: The PPO algorithm ensures the stability of the training process by constraining the update amplitude of the old and new policies. Its objective function is:
[0062]
[0063] where: is the policy ratio; is the advantage function, and its calculation method is:
[0064]
[0065] In the formula, ε is the clipping factor. The PPO algorithm gradually updates the policy network parameter θ by optimizing this objective function to maximize the cumulative reward.
[0066] Multi-step reinforcement learning: Through multi-step reinforcement learning, the advantage function is estimated as:
[0067]
[0068] where γ is the discount factor and L is the step size of multi-step learning. Multi-step reinforcement learning can more accurately estimate the long-term return and improve the effect of policy optimization.
[0069] Model Training and Prediction: Initialize the PPO agent and the ResNet network, generate trajectory data by interacting with the environment, and use the PPO algorithm to update the policy network and the value network. During the training process, the agent continuously optimizes the policy to maximize the cumulative reward. In the testing phase, input the vibration signal of the rolling bearing, and the trained model outputs the predicted remaining useful life value. The model can adaptively learn the degradation pattern of the rolling bearing, significantly improving the accuracy and robustness of the prediction.
[0070] The beneficial effects of the present invention are as follows:
[0071] High-precision Prediction: The PPO algorithm significantly improves the RUL prediction accuracy through stable policy updates and by combining the deep feature extraction of ResNet.
[0072] Strong Robustness: Residual connections and BN layers effectively alleviate the problem of gradient disappearance, enhancing the model's anti-interference ability against noise signals.
[0073] Efficient Training: The multi-step reinforcement learning design reduces the sample requirements, and the sample efficiency of PPO is significantly improved compared to traditional DRL algorithms.
[0074] Based on the above technical solutions, the following embodiments are provided.
[0075] Dataset Introduction: The IMS (Intelligent Maintenance System) bearing dataset of the University of Cincinnati in the United States is adopted, covering the full life cycle data of the bearing from healthy to failure, including various fault modes (bearing inner ring, outer ring, rolling elements, etc.) and complete life cycle data, which is suitable for verifying the generalization ability of the model to complex degradation patterns. In the experiment, 4 ZA-2115 type double-row rolling bearings are installed on the shaft, and the shaft is driven by an AC motor at a constant speed of 2000 revolutions per minute. The vibration signal is collected by an acceleration sensor, the sampling frequency is set to 20.48KHz, and one second of data is collected every ten minutes (a total of 20480 data points). The obtained dataset 1 has 2156 files. At the end of the experiment from test to failure, an inner ring fault occurred in bearing 3 and a rolling element fault occurred in bearing 4; dataset 2 has 984 files. At the end of the experiment from test to failure, an outer ring fault occurred in bearing 1.
[0076] Data Preprocessing: In this embodiment, datasets 1 and 2 are used above. Dataset 1 corresponds to the inner ring fault of the bearing (bearing named Bearing_3_3) and the rolling element fault (bearing named Bearing_2_3), and dataset 2 corresponds to the outer ring fault of the bearing (bearing named Bearing_1_2). 70% of the total samples of each dataset are divided into the training set, and 30% are the test set. The fast Fourier transform (FFT) is performed on the time-domain vibration signal to extract the frequency-domain features, highlighting the fault frequency components and making it easier to identify the bearing degradation features.
[0077] Model Training: Using the PPO-ResNet framework, with a one-dimensional vibration signal as the input and the predicted remaining life value as the output. Both the policy network and the value network adopt the ResNet architecture, specifically including:
[0078] Input Layer: Receives a one-dimensional vibration signal and outputs with a shape of (bs, 1, s_len(4096));
[0079] Wide Convolutional Layer: Used to capture short-term dependent features of the signal and outputs with a shape of (bs, 32, 512);
[0080] Batch Normalization Layer: Used to accelerate model convergence and outputs with a shape of (bs, 32, 512);
[0081] Max Pooling Layer: Used for dimensionality reduction and outputs with a shape of (bs, 32, 256);
[0082] 8 Residual Blocks: Each residual block contains a convolutional layer, a batch normalization layer, and an activation function. The output shapes are (bs, 32, 128), (bs, 64, 64), (bs, 64, 64), (bs, 128, 32), (bs, 128, 32), (bs, 64, 16), (bs, 64, 16), (bs, 32, 8) in sequence;
[0083] Global Average Pooling Layer: Used to integrate spatial information and outputs with a shape of (bs, 32);
[0084] Fully Connected Layer: Outputs the policy value or value estimate, with an output shape of (bs, 2) or (bs, 1).
[0085] The clipping factor ε of the PPO algorithm is set to 0.2, the discount factor γ is set to 0.9, and the multi-step learning step size L is set to 5; the training cycle (epoch) is 20 rounds, the number of times the agent interacts with the environment in each cycle (num_episodes) is 5000 times, the learning rate of the policy network (actor_lr) is 0.0001, and the learning rate of the value network (critic_lr) is 0.005. To alleviate the credit assignment problem of the model's delayed rewards, the GAE (Generalized Advantage Estimator) parameter λ is introduced into the advantage function, with a value of 0.9.
[0086] Generate trajectory data by interacting with the environment. After each training cycle, update the network parameters using the PPO objective function to optimize the policy to maximize the cumulative reward.
[0087] Prediction and Verification: Use the trained model for the test set, and compare the prediction results with the actual remaining life values as Figure 5The results show that the predicted curve is highly consistent with the true remaining life value. The evaluation metrics are RMSE (Root Mean Square Error), MAE (Mean Absolute Error), and MAPE (Mean Absolute Percentage Error), and the calculation formulas are as follows:
[0088]
[0089] Among them, a i is the predicted remaining life value, y i is the actual remaining life value, and m is the total number of samples.
[0090] After calculation, the prediction results are as follows:
[0091] Outer ring (Bearing_1_2): RMSE is 0.0391; MAE is 0.0312; MAPE is 13.6290%;
[0092] Rolling element (Bearing_2_3): RMSE is 0.0482; MAE is 0.0341; MAPE is 10.4301%;
[0093] Inner ring (Bearing_3_3): RMSE is 0.0402; MAE is 0.0211; MAPE is 4.3288%;
[0094] Compared with the traditional deep neural network (DNN) and support vector machine (SVM) methods, the prediction error of the present invention is significantly reduced. For example, the MAPE of the inner ring prediction is reduced by about 51% and 64% compared with DNN (8.72%) and SVM (12.15%) respectively.
[0095] Among them, the inner ring has the best prediction performance (MAPE = 4.3288%). Due to its relatively stable degradation mode, ResNet can effectively extract deep frequency domain features; the errors of the outer ring and rolling elements are relatively high, which may be related to their structural complexity and working condition noise interference. As shown in part (a) of Figure 5 , there are multi-band noise peaks in the vibration signal spectrum of the outer ring; the multi-step reinforcement learning design enables the model to capture the long-term degradation trend, further improving the prediction stability of complex components such as rolling elements.
[0096] This embodiment shows that the method for predicting the remaining life of rolling bearings based on PPO-ResNet exhibits excellent performance on different bearing components, especially maintaining high robustness under complex working conditions. The experimental results verify the effectiveness and practicality of the present invention in the remaining life prediction task.
[0097] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the various embodiments, reference can be made to each other.
[0098] In this article, specific examples are used to elaborate on the principles and implementation modes of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation modes and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A method for predicting the remaining life of a rolling bearing, characterized in that: include: Obtain the vibration signal of the rolling bearing and the corresponding detection time, and construct the dynamic state space and the continuous action space parameterized based on Gaussian distribution; Designing a dynamic reward function with prediction error as feedback, and building a Markov decision process model based on the ResNet network framework according to the state space, the action space and the reward function, and iteratively training the Markov decision process model using the PPO algorithm to obtain a trained Markov decision process model; The trained Markov decision process model is used to predict the remaining life of rolling bearings.
2. The method for predicting the remaining life of a rolling bearing according to claim 1, characterized in that: The state space is represented as In each training cycle e, m state-time pairs are randomly sampled As the state set for trajectory prediction; where s i is the vibration signal, y i is the detection time, s t is the current state, y t is the actual remaining life, i is the initial monitoring time point, and T is the final monitoring time point.
3. The method for predicting the remaining life of a rolling bearing according to claim 1, characterized in that: The action space is expressed as: Λ=μ+σ Among them, μ is the mean and σ is the standard deviation.
4. The method for predicting the remaining life of a rolling bearing according to claim 1, characterized in that: The reward function is expressed as: r t =-1×|a t -y t | Among them, a t is the predicted value of the agent, y t The actual remaining life.
5. The method for predicting the remaining life of a rolling bearing according to claim 1, characterized in that: The Markov decision process model includes a policy network and a value network; the policy network and the value network are both constructed based on the ResNet network framework, and the residual block structure mathematical expression of the ResNet network framework is: H(x)=F(x)+x, wherein F(x) represents the residual component and x is the input information.
6. The method for predicting the remaining life of a rolling bearing according to claim 1, characterized in that: The ResNet network framework specifically includes: Input layer: accepts one-dimensional vibration signal, and the output shape is (bs, 1, s_len(4096)); Wide convolutional layer: used to capture the short-term dependency features of the signal, with an output shape of (bs, 32, 512); Batch normalization layer: used to accelerate model convergence, the output shape is (bs, 32, 512); Max pooling layer: used for dimensionality reduction, the output shape is (bs, 32, 256); 8 residual blocks: Each residual block contains a convolutional layer, a batch normalization layer and an activation function, and the output shapes are (bs, 32, 128), (bs, 64, 64), (bs, 64, 64), (bs, 128, 32), (bs, 128, 32), (bs, 64, 16), (bs, 64, 16), (bs, 32, 8); Global average pooling layer: used to integrate spatial information, with an output shape of (bs, 32); Fully connected layer: Outputs policy value or value estimate, with output shape of (bs,2) or (bs,1).
7. The method for predicting the remaining life of a rolling bearing according to claim 1, characterized in that: The objective function of the PPO algorithm is: Where ε is the clipping factor, is the strategy ratio, s t is the current state, a t is the predicted value of the agent, θ is the network parameter, π θ is the old strategy, π′ θ For the new strategy, E t is the expectation at time step t, is a multi-step advantage function, and its calculation formula is: Among them, γ is the discount factor, L is the multi-step learning length, V φ is the output of the value network.
8. A rolling bearing remaining life prediction system, characterized in that: include: A space construction unit, used to obtain the vibration signal of the rolling bearing and the corresponding detection time, and to construct a dynamic state space and a continuous action space parameterized based on Gaussian distribution; A model building and training unit, which is used to design a dynamic reward function with prediction error as feedback, and based on the ResNet network framework, build a Markov decision process model according to the state space, the action space and the reward function, and iteratively train the Markov decision process model using the PPO algorithm to obtain a trained Markov decision process model; The model prediction unit is used to predict the remaining life of the rolling bearing using the trained Markov decision process model.
9. An electronic device, characterized in that: It comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the method for predicting the remaining life of a rolling bearing according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the method for predicting the remaining life of a rolling bearing as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Rolling bearing residual life prediction method based on residual error correction
CN112990524A
Reinforcement learning method for learning internal rewards based on state semantic representation
CN118886476A
A computer-implemented method for training an untrained policy network of an autonomous agent
GB202304559D0