Bearing residual life prediction method based on sparse Gaussian process regression
Through the method based on sparse Gaussian process regression, combined with deep spatiotemporal feature extraction and timestamp information processing, the problem of insufficient accuracy and efficiency of bearing residual life prediction in the prior art is solved, and efficient and accurate bearing residual life prediction and uncertainty quantification are achieved.
Patent Information
- Application Number
- CN202510066930.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-23
AI Technical Summary
The existing bearing residual life prediction methods are insufficient in terms of accuracy and efficiency, and fail to effectively utilize the timestamp information, resulting in uncertainty and complexity of the prediction results.
A method based on sparse Gaussian process regression is adopted, combined with deep spatiotemporal feature extraction and timestamp information processing, a Gaussian process regression model is constructed to achieve accurate prediction and uncertainty quantification of the remaining life of the bearing.
It improves the accuracy and efficiency of bearing residual life prediction, can effectively capture the time evolution laws in the degradation process, provide end-to-end efficient calculations, and significantly improve in uncertainty quantification.
Smart Images

Figure CN120030336A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bearing remaining life prediction, and in particular to a bearing remaining life prediction method based on sparse Gaussian process regression, and in particular to a remaining life prediction method driven by signal data and capable of providing uncertainty quantification. Background Art
[0002] Equipment health management has become a key factor in ensuring system reliability and reducing maintenance costs. Among them, as one of the most critical components of rotating machinery, the prediction of the remaining useful life (RUL) of bearings is of great significance for ensuring the normal operation of equipment, extending service life and optimizing maintenance strategies. Accurate RUL prediction can provide early warning of potential failures and avoid equipment downtime and production losses caused by bearing failure, thus playing a vital role in industrial production. However, uncertainty in the RUL prediction process is a factor that cannot be ignored. The operating environment of the equipment is complex and changeable, and there are many factors affecting RUL and they are random. Therefore, introducing uncertainty estimation in the prediction results can not only provide a reliable time range for decision-making reference, but also improve the robustness and credibility of the prediction model. With the increase in the complexity of the prediction model, the importance of uncertainty estimation in dealing with model uncertainty, data noise and unmeasurable variables has become increasingly prominent. By introducing uncertainty estimation in RUL prediction, the potential risks and errors in the prediction model can be better captured, making the prediction results more reliable and credible. This not only helps to improve the prediction accuracy of the model, but also provides a more reliable decision-making basis for equipment maintenance and management in practical applications. Therefore, studying how to effectively model and evaluate the uncertainty in the prediction results has become one of the important research directions in the current intelligent operation and maintenance field. At present, the mainstream uncertainty estimation methods mainly include Monte Carlo Dropout (MCDropout) and Mean-Variance Estimate (MVE) technology based on deep learning. The MCDropout method simulates the prediction distribution under different network parameters by introducing a Dropout layer in the neural network, thereby realizing uncertainty estimation. The MVE technology estimates uncertainty by simultaneously learning the mean and variance of the predicted values. These methods have achieved certain results in practical applications, but there are still many problems, such as:
[0003] (1) In terms of accuracy, traditional data-driven prediction methods usually only consider the direct mapping relationship between training time series samples and target remaining life prediction values, while ignoring the dynamic correlation between time series samples, resulting in prediction accuracy that is difficult to achieve as expected.
[0004] (2) Existing uncertainty quantification methods for remaining life prediction based on machine learning usually require specialized model design and multiple iterative reasoning, such as the widely used Monte Carlo Dropout method and MVE method. These methods cannot achieve end-to-end one-time calculation, which significantly reduces the efficiency of prediction in engineering applications and is highly complex.
[0005] (3) Existing methods also generally ignore the key role of timestamp information in remaining life prediction and fail to fully utilize the temporal dependency in time series, which further limits the prediction performance and applicability of the model. Summary of the invention
[0006] In view of this, the present invention provides a bearing remaining life prediction method based on sparse Gaussian process regression, which combines time series sample association and timestamp information to improve prediction accuracy and achieve end-to-end efficient computing.
[0007] The present invention provides a method for predicting the remaining life of a bearing based on sparse Gaussian process regression, and the specific steps are as follows:
[0008] Step 1. Process the bearing vibration signal to obtain the time-frequency sequence;
[0009] Step 2. Build a deep spatiotemporal feature extractor to extract deep spatiotemporal features from the time-frequency sequence;
[0010] Step 3. Construct a time information processing module to obtain time features from the time-frequency series;
[0011] Step 4. Construct a Gaussian process regression model based on deep spatiotemporal features and time features; construct a remaining life prediction model by sparse Gaussian process regression;
[0012] Step 5. Use the remaining life prediction model in step 4 to predict the remaining life of the bearing.
[0013] Optionally, the specific steps of step 1 are as follows:
[0014] Step 11. Using continuous wavelet transform to extract the time-frequency characteristics of the bearing vibration signal;
[0015] Step 12. Use the sliding time window to obtain the time-frequency series of time-frequency features.
[0016] Optionally, the specific steps of step 2 are as follows:
[0017] Step 21. Construct a convolutional gated unit network;
[0018] Step 22. Construct a residual convolution gated recurrent unit;
[0019] Step 23. Construct a deep feature extractor based on the convolutional gated unit network and the residual convolutional gated recurrent unit.
[0020] Optionally, the specific steps of step 3 are as follows:
[0021] Step 31. Scale each timestamp corresponding to each element in the time-frequency sequence obtained in step 1 to obtain a scaled timestamp;
[0022] Step 32: Encode the scaled timestamp obtained in step 31 using sine and cosine coding to obtain time coded data;
[0023] Step 33. Construct a time information processing module for feature mapping of time-coded data.
[0024] Optionally, the specific steps of step 3 are as follows:
[0025] Step 41. Use the deep feature extractor of step 2 to obtain deep spatiotemporal features and the time information processing module of step 3 to obtain time features;
[0026] Step 42: splicing the time features along the feature dimension to obtain fusion information;
[0027] Step 43. Based on the fusion information obtained in step 42, a Gaussian process regression model is constructed;
[0028] Step 44. The Gaussian process regression model optimizes the model parameters by minimizing the negative log-likelihood function, thereby maximizing the model's fit to the input depth features and the actual remaining life data, and finally obtaining a trained remaining life prediction model.
[0029] Optionally, the time information processing module is a feedforward neural network encoder module; the feedforward neural network encoder module includes a feedforward neural network layer and layer normalization.
[0030] Optionally, the expression of the time-frequency feature wt(s,τ) of the bearing vibration signal extracted by continuous wavelet transform in step 11 is:
[0031]
[0032] Where wt(s,τ) represents the time and frequency characteristics of the bearing vibration signal obtained by continuous wavelet transform under the scale parameter s and time translation parameter τ; x(t) represents the bearing vibration signal at time t obtained by the acceleration sensor; ψ * represents the complex conjugate of the wavelet basis function ψ.
[0033] Compared with the prior art, the present invention has at least the following beneficial effects:
[0034] (1) The method of the present invention can provide accurate RUL prediction and uncertainty quantification results of bearings. In terms of uncertainty quantification, multiple repeated inferences are not required, and the model structure is more concise and clear.
[0035] (2) The method of the present invention introduces timestamp information as a key input feature, which significantly improves the accuracy of bearing RUL prediction and effectively captures the time evolution law of the degradation process.
[0036] (3) The method of the present invention obtains deep spatiotemporal features and deep timestamp features from the time-frequency feature sequence of the input vibration signal with the help of the designed deep network, and directly uses these two features as the input of Gaussian process regression, thereby realizing the coordinated optimization of network parameters and Gaussian process regression parameters, and achieving accurate end-to-end bearing RUL prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings are only for the purpose of illustrating particular embodiments and are not to be construed as limiting the invention.
[0038] Figure 1 A flowchart of the method for predicting the remaining life of a bearing according to the present invention;
[0039] Figure 2 A residual convolution gated network unit diagram of the bearing remaining life prediction method of the present invention;
[0040] Figure 3 A network structure diagram of spatiotemporal feature extraction based on a residual convolution gated network unit of a bearing remaining life prediction method of the present invention;
[0041] Figure 4 A timestamp coding network structure diagram of the bearing remaining life prediction method of the present invention;
[0042] Figure 5 A diagram showing the prediction results of the bearing remaining life prediction method of the present invention;
[0043] Figure 6 This is a comparison chart of the results of the bearing remaining life prediction method of the present invention, the MCDropout method and the MVE method. DETAILED DESCRIPTION
[0044] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. In addition, the present invention can also be implemented in other ways different from those described herein, and therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0045] A specific embodiment of the present invention, as Figure 1-6 , discloses a bearing remaining life prediction method based on sparse Gaussian process regression, the specific steps are as follows:
[0046] Step 1. Process the bearing vibration signal to obtain the time-frequency sequence;
[0047] Step 11. Use continuous wavelet transform to extract the time-frequency characteristics wt(s,τ) of the bearing vibration signal, expressed as:
[0048]
[0049] Where wt(s,τ) represents the time and frequency characteristics of the bearing vibration signal obtained by continuous wavelet transform under the scale parameter s and time translation parameter τ; x(t) represents the bearing vibration signal at time t obtained by the acceleration sensor; ψ * represents the complex conjugate of the wavelet basis function ψ.
[0050] Furthermore, the scale parameter s controls the stretching and compression of the wavelet basis function and is the control parameter of the frequency characteristic. The time translation parameter τ controls the translation of the wavelet function on the time axis and is the control parameter of the time characteristic. By jointly controlling s and τ, the time and frequency characteristics of the original signal can be obtained, thereby obtaining the time-frequency characteristics of the original signal.
[0051] Step 12. Use the sliding time window to obtain the time-frequency sequence of the time-frequency feature wt(s,τ);
[0052] Furthermore, the overlapping method is combined to enhance the correlation between adjacent elements of the time-frequency series, which enhances the smoothness and consistency of the time-frequency series and ensures that the features captured in adjacent time windows have a certain continuity.
[0053] Specifically, the expression of the time-frequency series within the bearing operation cycle is:
[0054]
[0055] Among them, X (n-T+1) represents the n-T+1th time-frequency series of the bearing; It represents the feature of the jth feature image in the time-frequency sequence, H represents the height of the feature image, W represents the width of the feature image, C represents the number of channels of the feature image, j=1,2,…,n, n represents the total number of time-frequency features obtained, and T represents the total number of feature images obtained.
[0056] The present invention not only retains the time-frequency information of the signal, but also enhances the intrinsic correlation between features by overlapping features in the sequence, which helps the model better capture the potential dynamic changes in the bearing degradation process. Compared with traditional methods, the present invention not only improves the prediction accuracy but also ensures the efficiency and applicability of the model by effectively integrating time windows, sliding windows and wavelet transform technology. In addition, the present invention uses the time window used to obtain the time-frequency feature sequence of the bearing, which overcomes the problem that the bearing is a highly durable mechanical component, the performance degrades slowly, and the performance changes are not obvious within a single sampling cycle.
[0057] Step 2. Build a deep spatiotemporal feature extractor.
[0058] Step 21. Construct a convolutional gated unit network, the expression is:
[0059] r t =σ(W xr *x t +W hr *h t-1 +b r ) (3)
[0060] z t =σ(W xz *x t +W hz *h t-1 +b z ) (4)
[0061]
[0062] Among them, r t represents the reset gate output at time t, which controls the input information x at time t t and the hidden state h at the previous time t-1 t-1 The combination of determines how much information between adjacent moments should be reset; t represents the update gate output at time t, which determines the hidden state h at the previous time t-1 t-1 and the potential new hidden state at time t The degree of mixing; represents the potential new hidden state at time t, which is based on the input information x at time t t and the hidden state h at the previous time t-1 after adjustment by the reset gate t-1 The composition of h represents the intermediate calculation result at time t; t represents the hidden state at time t, which is updated by gate z t The hidden state h at the previous moment k-1 and the potential new hidden state at time t It is obtained by weighted mixing between the past and current inputs, which contains the network's balanced information about the past and current inputs and is the input passed to the next moment; W xr Represents the input information x at time t t With r t The convolution kernel weight between is used to calculate the impact of the current input on the reset gate; W hr Represents the hidden state h at the previous time t-1 t-1 With r t The convolution kernel weight between is used to calculate the influence of the previous hidden state on the reset gate; b r Indicates the reset gate r t The bias term adjusts the output of the reset gate; W xz Represents the input information x at time t t With update gate z t The convolution kernel weight between is used to calculate the impact of the current input on the update gate; W hz represents the hidden state h at time t-1 t-1 With update gate z t The convolution kernel weight between is used to calculate the influence of the previous hidden state on the update gate; b z Indicates the function used to update gate z t The bias term adjusts the output of the update gate; W xh Represents the input information x at time t t and the potential new hidden state at time k The convolution kernel weight between is used to calculate the contribution of the current input to the candidate hidden state; W hh Represents the hidden state h at the previous time t-1 t-1 With the potential new hidden state at time t The convolution kernel weight between is used to calculate the influence of the previous hidden state on the candidate hidden state; b h represents the potential new hidden state for time t The bias term adjusts the calculation result of the potential new hidden state at time t; the operator * represents the convolution operation; ⊙ represents the Hadamard product; tanh(·) and σ(·) represent the hyperbolic tangent function and the sigmoid function, respectively.
[0063] It can be understood that the input information x at time t t It is the time-frequency characteristics of the time-frequency series within the bearing operation cycle.
[0064] The reset gate and update gate of the present invention play a key role in regulating the information flow of the network hidden state. Through the coordinated action of these two gates, the convolutional gating unit can dynamically adjust the flow of information between moments, thereby effectively capturing the time dependencies in the input sequence. This mechanism is particularly critical for tasks such as remaining useful life prediction, because the state degradation process of the equipment is usually accompanied by complex time correlations, and the reset gate and update gate ensure that the model can flexibly handle important patterns in these time series.
[0065] The convolutional gated unit network of the present invention enhances the traditional gated recurrent unit by replacing its fully connected layer with a convolution operation, thereby being able to capture both the spatial structure and the temporal dependency in the data.
[0066] Step 22. Construct a residual convolutional gated recurrent unit.
[0067] The present invention introduces a residual convolution gated recurrent unit into the convolution operation of the convolution gated recurrent unit to achieve deep feature extraction of the time-frequency feature sequence.
[0068] like Figure 2 As shown, the residual convolution gated recurrent unit includes two layers of batch normalization (BN) and two layers of convolutional layers, and an identity connection is added at the input and output of the residual convolution gated recurrent unit.
[0069] The residual convolution gated recurrent unit of the present invention allows the network to capture deep degradation information in bearing vibration data while preventing gradients from vanishing or exploding. The two convolutional layers help to extract different levels of information in time-frequency features and identify degradation patterns in vibration signals over time; batch normalization stabilizes the input data distribution of each layer to ensure a smoother training process. Identity connections allow gradients to be directly transmitted along shortcut paths, preventing gradients from vanishing or exploding in deep networks, while ensuring that key information in time-frequency features is retained in deep networks, enhancing the ability to detect deep degradation.
[0070] Step 23. Construct a deep feature extractor based on the convolutional gated unit network and the residual convolutional gated recurrent unit.
[0071] Specifically, Figure 3 As shown in Figure 1, the deep feature extractor includes stacked residual convolutional gated recurrent units, convolutional gated recurrent unit networks, pooling layers, flattening layers, and fully connected layers to extract deep feature maps from time-frequency feature sequences. The outputs of these layers are then fed into flattening layers and fully connected layers to refine the data features.
[0072] The deep feature extractor of the present invention helps to finely extract and transform temporal and spatial features from raw time-frequency series data so that it is optimized for subsequent analysis or predictive modeling tasks.
[0073] The present invention introduces a residual network structure to form a residual convolution gated recurrent unit to enhance the ability to extract deep features from time-frequency feature sequences. The residual structure includes two layers of batch normalization (BN) and two layers of convolution layers, and adds identity connections at the input and output of the residual. This configuration allows the network to capture deep degradation information in bearing vibration data while preventing the occurrence of gradient vanishing or explosion problems. After implementing the residual convolution gated recurrent unit module, the present invention further stacks multiple residual convolution gated recurrent unit layers with simple convolution gated recurrent unit layers to form a deep feature extractor. The extractor consists of a residual convolution gated recurrent unit layer, a convolution gated recurrent unit layer, and a pooling layer, and its output is then fed into a flattening layer and a fully connected layer to further refine and optimize data features. This architecture is capable of finely extracting and transforming temporal and spatial features from raw time-frequency data, providing high-quality input features for subsequent analysis or predictive modeling tasks.
[0074] Step 3. Construct a time information processing module.
[0075] The timestamp encoder of the proposed time information processing module is as follows Figure 4 shown.
[0076] Step 31. Scale each timestamp corresponding to the time-frequency sequence obtained in step 1 to obtain a scaled timestamp for time t, to ensure that all timestamps fall within the range of (0,1).
[0077] Step 32: Encode the scaled timestamp obtained in step 31 using sine and cosine coding to obtain time coded data;
[0078] Specifically, the expressions for sine and cosine encoding are:
[0079]
[0080] Among them, t ′ represents the timestamp after scaling at time t; k is the dimension index; d model represents the dimension of the encoding, P t′ (.) indicates the encoded scaled timestamp.
[0081] The present invention uses sine and cosine encoding timestamps, which can capture the relative positions of different time points and help understand the time relationship in the sequence. Sine and cosine functions have natural periodicity and can align periodic patterns in vibration signals well through their cyclic properties. This periodicity can be caused by the rotation of mechanical parts such as bearings. Therefore, the encoding method of the present invention can help the model better identify the periodic structure in the signal, especially when the change of the signal presents periodic oscillations, the model can use this property of sine and cosine functions to capture and explain these regular vibration characteristics, thereby improving the recognition ability of periodic patterns. At the same time, the method of the present invention also maintains the distance information between different time points, and can accurately model time dependence.
[0082] Step 33. Construct a time information processing module for feature mapping of time-coded data.
[0083] Furthermore, the time information processing module is a feedforward neural network encoder module, which can promote the adaptive learning of the remaining life prediction model of the present invention to time characteristics.
[0084] Specifically, the feedforward neural network encoder module consists of two feedforward neural network layers and layer normalization. Not only does it normalize the features, it also mitigates the risk of negative variance anomalies, thereby enhancing the stability of model training. In addition, layer normalization promotes smoother gradient flow during optimization, contributing to faster convergence and improved performance.
[0085] Furthermore, the layer normalization h norm The expression is:
[0086]
[0087] Where h represents the deep spatiotemporal features captured by the established deep feature extractor from the time-frequency feature input; E[h] and denote the mean and standard deviation of the deep spatiotemporal feature h in its feature dimension respectively; γ and β denote the trainable parameters for rescaling and offset normalization values respectively.
[0088] The temporal information processing module of the present invention can efficiently capture and encode the temporal features in the sequence, thereby providing more accurate input for the temporal dependency modeling of subsequent models and helping to improve the overall prediction performance.
[0089] Step 4. Construct a remaining life prediction model based on sparse Gaussian process regression.
[0090] Step 41. Use the deep feature extractor of step 2 to obtain deep spatiotemporal features and the time information processing module of step 3 to obtain time features;
[0091] Step 42: Concatenate the time features along the feature dimension to obtain fusion information, expressed as:
[0092] F(X,P t′ )=h(X;W DL )+T(P t′ ; W Time )(10)
[0093] in, Represents the weight parameter W of the deep network based on the residual convolution gated recurrent unit DL The deep spatiotemporal features extracted under the h Represents the dimension of deep spatiotemporal features, represents a real number; represents the weight parameter W from the temporal embedding layer Time The deep temporal features obtained by downscaling the timestamp t′, d t represents the dimension of deep temporal features; X represents the input time-frequency feature sequence; F(X,P t′ ) represents the final fusion information, which means that
[0094] Step 43: construct a Gaussian process regression model based on the fusion information obtained in step 42.
[0095] Furthermore, the main goal of the Gaussian process regression model is to make regression predictions on the input based on the training data and the input to be tested, expressed as:
[0096]
[0097] Among them, F * Indicates the bearing prediction point; f * Indicates the bearing prediction point F * The potential function value of represents the mean of the bearing prediction points; cov(f * ) is the variance of the bearing prediction points; F is the training data; y is the target remaining life prediction value of the training data.
[0098] Furthermore, the predicted mean and variance cov(f * ) is as follows:
[0099]
[0100] cov(f * )=K d (F * ,F * )-K d (F* ,F)[K d (F,F)] -1 K d (F,F * ) (12)
[0101] K d (F * ,F) represents the test point F obtained using the kernel function * Covariance with training data F; K d (F, F) is the covariance matrix of the training data F and the training data F obtained using the kernel function, and y is the target remaining life prediction value of the training data.
[0102] Furthermore, the expression of the kernel function is:
[0103]
[0104] Among them, K d (.) represents the depth kernel function; F(X p ,P t′ ) represents the pth depth fusion information at the time stamp t′ after scaling at time t; F(X q ,P t′ ) represents the qth depth fusion information at the time stamp t′ after scaling at time t; X p represents the pth time-frequency feature sequence; X q represents the qth time-frequency feature sequence; σ f is the output scale, and l is the distance between adjacent points in the scaled input space.
[0105] In the prediction of the Gaussian process regression method, it is assumed that the predicted function value obeys the Gaussian distribution. Based on this distribution, a confidence interval CI with a preset proportion can be obtained to estimate the possible range of the predicted RUL value. The expression is:
[0106]
[0107] Where A represents the constant multiple of the confidence interval corresponding to the preset proportion under the normal distribution.
[0108] Preferably, A=1.96. At this time, the preset ratio is 95%. The 95% confidence interval means that the predicted value has a 95% probability of falling within within the range.
[0109] Step 44. The Gaussian process regression model optimizes the model parameters by minimizing the negative log-likelihood function, thereby maximizing the model's fit to the input depth features and the actual remaining life data, and finally obtaining a trained remaining life prediction model.
[0110] Specifically, the expression of the negative log-likelihood function L is:
[0111]
[0112] Get the derivative of the remaining life prediction model parameters, the expression is:
[0113]
[0114] Among them, θ represents the hyperparameters in the kernel function, including the output scale parameter σ f and scaling parameter l; K d Represents the depth kernel function.
[0115] Furthermore, since the direct calculation of the above formula is inefficient when facing large-scale data, especially the need to invert the covariance matrix, we use the deep kernel function K in the induced point sparsification step 43 d To approximate the calculation, K d Expressed as Further shortened to: Among them, F n represents the deep fusion information of the training data, i.e., the feature representation of the training data extracted by the designed deep network; F m K represents the deep fusion information of the induction point, that is, the feature representation of the induction point extracted by the designed deep network. nm represents the covariance matrix between the training data and the induction points obtained by the kernel function; K nn Represents the covariance matrix obtained by the kernel function between training data;
[0116] Furthermore, in order to select the optimal induction points that can replace the global data, the present invention introduces a variational sparse Gaussian process regression method;
[0117] Specifically, the variational sparse Gaussian process regression method selects m (m<<n) induced point positions X from n obtained time-frequency features. m As a random variable, the Gaussian process regression model obtains an approximate true posterior p(f,f m |y), the expression is:
[0118]
[0119] In the formula, p(f|f m ) is the conditional probability distribution of the Gaussian process, indicating that at the mth induction point position f m In this case, the posterior distribution of the potential function f, q(f m ) is p(f m |y), which represents the distribution of the induced points, and are the mean and variance of the variational posterior, respectively; Represents the deep kernel function after sparsification.
[0120] In order to obtain the true posterior distribution p(f,f m |y), using the variational inference method, whose goal is to minimize the Kullback-Leibler (KL) divergence KL[q(f,f m )‖p(f,f m |y)], the complex p(f,f m |y) is approximated as a variational posterior distribution q(f,f) which is easier to compute m ). KL divergence measures the difference between two distributions. Minimizing this difference can make the approximate distribution as close as possible to the true posterior distribution.
[0121] In specific operations, the problem of minimizing the KL divergence is transformed into maximizing the lower bound of the log-likelihood function, expressed as:
[0122] log p(y)≥E q(f) [log p(y|f)]-KL(q(f m )‖p(f m ))#(20)
[0123] Where log p(y) represents the marginal log-likelihood of the observed data, which is used to measure the degree of fit of the model to the data; E q(f) [.] represents the expectation of the variational distribution q(f); log p(y|f) represents the conditional log-likelihood of the observed data given the potential function f; KL(q(f m )‖p(f m )) represents the variational posterior distribution q(f m ) and the prior distribution p(f m ) is used to measure the difference between the two.
[0124] After using the induced point approximation, the conditional likelihood p(y|f) is expressed as m An approximate distribution of . m Substituting the conditional distribution under , we can get:
[0125]
[0126] Among them, σ represents the standard deviation of the observed data noise, which is used to describe the random noise in the observed data; I represents the unit matrix, whose dimension is the same as that of the observed data y, and is used to represent the covariance structure of the observed noise.
[0127] In order to further accurately approximate log p(y), the present invention introduces two correction terms. The first correction term is the correction term M introduced on the diagonal line. 1 , to correct K nm and The difference between them is expressed as:
[0128]
[0129] Where Tr(·) represents the trace of the matrix.
[0130] The other is the error term M in the compensation approximation process 2 , that is, for K nm With the approximate matrix The difference is expressed as:
[0131]
[0132] Furthermore, the lower bound of the maximized log-likelihood function is further transformed into:
[0133]
[0134] Among them, K nn represents the covariance matrix between training data; q(u) represents the variational posterior distribution of the inducing point u; p(u) represents the prior distribution of the inducing point u;
[0135] Furthermore, according to the above-mentioned sparse operation, the prediction formulas of the original formulas (11) and (12) can be transformed into:
[0136]
[0137] The present invention converts formula (14) into formula (24) to achieve small batch training in deep learning.
[0138] Then, optimization is performed according to formulas (15)-(17) to obtain the optimal parameters of the model.
[0139] Specifically, the training process of the model can be expressed as:
[0140] (1) First, obtain a set of randomly initialized parameters The location of the induction point was defined as the first 50 training points;
[0141] (2) Forward propagation calculation of the remaining life prediction model and calculation of the negative log-likelihood function The gradient of each parameter;
[0142] (3) Update using the Adam optimization algorithm And other parameters;
[0143] (4) Then the above process is iterated continuously until the remaining life prediction model reaches the convergence condition.
[0144] Step 5: Use the remaining life prediction model in step 4 to predict the remaining life of the bearing.
[0145] The first part of the objective function of the present invention represents the log-likelihood of the observed data, and the last two terms correspond to the uncertainty estimate of the induced point and the penalty term of the KL divergence. By maximizing this objective function, the present invention can significantly reduce the computational complexity of the Gaussian process while maintaining high prediction accuracy, thereby making the method applicable to modeling and prediction of large-scale data sets.
[0146] In order to illustrate the effectiveness of the method proposed in this invention, an effectiveness experiment was conducted. Figure 5 , showing the prediction results of the present invention. It can be seen from the figure that the predicted mean is very consistent with the actual remaining life value, and the 95% confidence interval can also cover most of the true values, proving the effectiveness and accuracy of the method.
[0147] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for predicting the remaining life of a bearing based on sparse Gaussian process regression, characterized in that: The specific steps are as follows: Step 1. Process the bearing vibration signal to obtain the time-frequency sequence; Step 2. Build a deep spatiotemporal feature extractor to extract deep spatiotemporal features from the time-frequency sequence; Step 3. Construct a time information processing module to obtain time features from the time-frequency series; Step 4. Construct a Gaussian process regression model based on deep spatiotemporal features and time features; construct a remaining life prediction model by sparse Gaussian process regression; Step 5. Use the remaining life prediction model in step 4 to predict the remaining life of the bearing.
2. The method for predicting the remaining life of a bearing according to claim 1, characterized in that: The specific steps of step 1 are as follows: Step 11. Using continuous wavelet transform to extract the time-frequency characteristics of the bearing vibration signal; Step 12. Use the sliding time window to obtain the time-frequency series of time-frequency features.
3. The method for predicting the remaining life of a bearing according to claim 2, characterized in that: The specific steps of step 2 are as follows: Step 21. Construct a convolutional gated unit network; Step 22. Construct a residual convolution gated recurrent unit; Step 23. Construct a deep feature extractor based on the convolutional gated unit network and the residual convolutional gated recurrent unit.
4. The method for predicting the remaining life of a bearing according to claim 3, characterized in that: The specific steps of step 3 are as follows: Step 31. Scale each timestamp corresponding to each element in the time-frequency sequence obtained in step 1 to obtain a scaled timestamp; Step 32: Encode the scaled timestamp obtained in step 31 using sine and cosine coding to obtain time coded data; Step 33. Construct a time information processing module for feature mapping of time-coded data.
5. The method for predicting the remaining life of a bearing according to claim 4, characterized in that: The specific steps of step 3 are as follows: Step 41. Use the deep feature extractor of step 2 to obtain deep spatiotemporal features and the time information processing module of step 3 to obtain time features; Step 42: splicing the time features along the feature dimension to obtain fusion information; Step 43. Based on the fusion information obtained in step 42, a Gaussian process regression model is constructed; Step 44. The Gaussian process regression model optimizes the model parameters by minimizing the negative log-likelihood function, thereby maximizing the model's fit to the input depth features and the actual remaining life data, and finally obtaining a trained remaining life prediction model.
6. The method for predicting the remaining life of a bearing according to claim 4, characterized in that: The time information processing module is a feedforward neural network encoder module; the feedforward neural network encoder module includes a feedforward neural network layer and layer normalization.
7. The method for predicting the remaining life of a bearing according to claim 2, characterized in that: The expression of the time-frequency characteristics wt(s,τ) of the bearing vibration signal extracted by continuous wavelet transform in step 11 is: Where wt(s,τ) represents the time and frequency characteristics of the bearing vibration signal obtained by continuous wavelet transform under the scale parameter s and time translation parameter τ; x(t) represents the bearing vibration signal at time t obtained by the acceleration sensor; ψ * represents the complex conjugate of the wavelet basis function ψ.