A Rolling Bearing Remaining Life Prediction Method and System Based on Convolutional White Box

Through the method based on convolution white box, combined with time-frequency domain feature extraction and CRATE network architecture, the healthy state is dynamically divided, and the expansion causal convolution attention mechanism and multi-scale convolution are integrated, which solves the problem of insufficient prediction accuracy and interpretability of rolling bearings in the existing technology, and achieves high-precision and stable residual life prediction, which is suitable for electric power, wind power, aerospace and other fields.

CN120145176BActive Publication Date: 2025-07-18SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510615349.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-18
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing rolling bearing residual life prediction method is not very accurate when dealing with vibration signal noise interference, making it difficult to capture long-term dependence information and local degradation characteristics simultaneously, and the model is poor interpretable, which affects prediction accuracy and industrial applications.

Method used

The method based on the convolution white box is adopted, and the time-domain and frequency-domain feature extraction is combined with singular value decomposition and noise reduction, and the healthy state is dynamically divided, and the expansion of causal convolution attention mechanism and multi-scale convolution CRATE network architecture is integrated to synchronously extract the long-term dependence information of bearing signals and the local degradation characteristics.

Benefits of technology

It improves the accuracy and stability of the remaining life prediction of rolling bearings, enhances the local feature extraction capability, improves the interpretability and predictive credibility of the model, is suitable for predictive maintenance of high-value equipment, reduces the risk of unplanned downtime, and extends the service life of key components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145176B_ABST
    Figure CN120145176B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of detecting the remaining service life of bearings, and discloses a rolling bearing remaining life prediction method and system based on convolutional white box. This method extracts time-domain and frequency-domain features from the original vibration signal, performs noise reduction processing, divides the healthy state and the degradation state, and combines the Weibull-MSE loss function to train a deep neural network, so that the prediction of the RUL conforms to the actual degradation process of the bearing; through the convolutional CRATE network architecture that fuses the dilated causal convolutional attention mechanism DCA and the multi-scale convolution MSC, the synchronous extraction of the long-term dependence information and local degradation features of the bearing signal is completed. The present invention improves the accuracy of RUL prediction by combining health state assessment, enhances the local feature extraction ability, improves the local modeling ability of the model through multi-scale information fusion, and improves the interpretability and prediction credibility through the CRATE structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of detecting the remaining service life of bearings, and particularly relates to a method and system for predicting the remaining life of rolling bearings based on convolutional white box. Background Technique

[0002] In modern industrial production, rotating machinery (such as motors, turbines, pumps, and compressors) is widely used in fields such as aerospace, manufacturing, energy, and transportation. During the long-term operation of these devices, the health status of their key component - rolling bearings directly affects the operation efficiency, reliability, and safety of the devices.

[0003] With the development of industrial intelligence, Prognostics and Health Management (PHM) technology has emerged. The core goal of PHM is to discover potential faults of devices in advance through data analysis and modeling, predict their Remaining Useful Life (RUL), so as to reasonably arrange maintenance plans, reduce device downtime, and improve production efficiency. Currently, PHM mainly includes key links such as data acquisition and processing, health status assessment, fault diagnosis, fault prediction, health management, feedback, and iteration. As the core link of fault prediction, RUL prediction can help enterprises carry out maintenance at the right time, avoid unplanned downtime, reduce maintenance costs, and increase the service life of devices. Therefore, researching high-precision RUL prediction methods, especially for key components such as rolling bearings, has important industrial application value.

[0004] Furthermore, rolling bearings are one of the most common rotating support components in mechanical equipment. During long-term operation, factors such as friction, fatigue, and lubrication deterioration will cause it to gradually degrade and eventually fail. Therefore, predicting the remaining life (RUL) of rolling bearings can discover potential faults in advance, extend the device life, and improve production safety.

[0005] Currently, RUL prediction mainly relies on the following types of methods:

[0006] (1) Methods based on physical models.

[0007] This method describes the degradation process of bearings through mathematical or physical models and updates the model parameters by combining with actual monitoring data. For example, using differential equations to describe the material fatigue process or using statistical models to estimate the wear of bearings. Such methods rely on a large amount of prior knowledge, are applicable to specific devices, but it is difficult to establish accurate models under complex working conditions.

[0008] (2) Data-driven methods.

[0009] This method directly extracts features from bearing operation data, learns degradation patterns through machine learning or deep learning algorithms, and predicts the remaining useful life. The specific process includes data preprocessing, feature extraction, model training, and prediction. This type of method avoids complex mathematical modeling and has strong adaptability, but it often lacks the guidance of physical mechanisms, which may lead to prediction errors.

[0010] (3) Based on hybrid methods.

[0011] This method combines physical models with data-driven methods, uses prior knowledge to build an initial model, and optimizes model parameters in a data-driven manner. For example, first establish a physical model to describe the bearing degradation characteristics, and then use machine learning methods to correct the model parameters to make it more in line with the actual working conditions. This method takes into account the interpretability of physical models and the flexibility of data-driven methods, but it is complex to implement and has high requirements for data and domain knowledge.

[0012] Existing improved schemes based on Transformer have solved the problems of global and local information extraction in RUL prediction to a certain extent, but there are still many deficiencies.

[0013] (1) The health state assessment method is imperfect, affecting the prediction accuracy.

[0014] Existing RUL prediction methods usually assume that the bearing degradation process follows a linear or preset fixed curve. When training the model, all bearing data are directly regarded as a unified input, ignoring the individual differences in the operating environment and usage conditions of different bearings. The latest research proposes a method based on correlation coefficient analysis, which dynamically divides the stable period and degradation period of bearings using historical data, thus making up for the limitations of the fixed degradation curve to a certain extent and enhancing the adaptability to individual differences. However, this method does not perform more refined processing on the original historical data. Especially in the case of noise interference in vibration signals, the calculated health assessment state often has errors, affecting the accuracy and stability of RUL prediction.

[0015] (2) It is difficult to capture long-term dependence information and local degradation features simultaneously.

[0016] The closest existing technical solution: Convolutional Transformer (COT).

[0017] Convolutional Transformer (COT) combines CNN and Transformer structures, aiming to extract local and global features simultaneously in RUL prediction to improve the modeling capability of bearing degradation process. Its basic process is as follows: First, before Transformer, CNN is used to extract local features of vibration signals to enhance the capture of bearing degradation patterns; then, long-term dependencies are established through the self-attention mechanism of Transformer to ensure that the model can effectively learn the changing trend of bearing life over time, thereby improving prediction stability.

[0018] Although this method combines the advantages of CNN and Transformer, it still has certain limitations. Due to the limited receptive field of CNN, there may still be information loss when extracting local features of the bearing degradation period, resulting in insufficient characterization of key degradation features. In addition, this structure mainly focuses on time series modeling and lacks full utilization of the spatial characteristics of the signal (such as frequency domain information). RUL prediction not only depends on the degradation trend of the time dimension, but also involves the changes in frequency domain features. The shortcomings of COT in this regard affect the prediction accuracy and generalization ability.

[0019] (3) The model has poor interpretability and is difficult to apply in industrial scenarios.

[0020] The most similar technical solution currently available: Local Enhanced Transformer (MTCT).

[0021] Local enhancement (MTCT) adds a temporal convolution module to the standard Transformer structure to enhance the modeling ability of time series data. The core ideas are as follows: First, an attention mechanism based on temporal convolution is introduced before the Transformer to enhance the ability to capture local features, so that the model pays more attention to the key short-term changes in the bearing degradation process when calculating the attention weight; then, a multi-head self-attention mechanism is used to optimize the feature representation, and the Transformer structure is used to establish long-term dependencies to ensure that the model can effectively learn the changing trend of bearing life over time.

[0022] Although MTCT has improved the modeling ability of local information to a certain extent, its interpretability is still weak. This method still relies on the standard Transformer structure, and the decision-making process of its self-attention mechanism is complex, making it difficult to trace the specific feature effects, making the model's prediction results difficult for engineers or industry experts to intuitively understand. Summary of the invention

[0023] To overcome the problems existing in the related technologies, the disclosed embodiments of the present invention provide a rolling bearing remaining life prediction method and system based on convolutional white box, specifically related to a rolling bearing remaining service life (RUL) prediction method based on convolutional white box Transformer.

[0024] The technical solution is as follows: A rolling bearing remaining life prediction method based on convolutional white box includes the following steps:

[0025] S1, extract time-domain and frequency-domain features from the original vibration signal, and perform noise reduction processing by combining singular value decomposition (SVD) to remove random noise and interference in the signal, and provide input features;

[0026] S2, input the processed data into a deep neural network, and based on the correlation coefficient analysis method, dynamically evaluate the operation process of the bearing, divide the healthy state and the degradation state, and combine the Weibull-MSE loss function to train the deep neural network to make the prediction of RUL conform to the actual degradation process of the bearing;

[0027] S3, through the convolutional CRATE network architecture that fuses the dilated causal convolutional attention mechanism (DCA) and the multi-scale convolution (MSC), complete the synchronous extraction of the long-term dependence information and local degradation features of the bearing signal.

[0028] In step S2, based on the correlation coefficient analysis method, dynamically evaluate the operation process of the bearing, divide the healthy state and the degradation state, and combine the Weibull-MSE loss function to train the deep neural network, including:

[0029] S201, calculate the correlation between the features at each moment after SVD compression and the features at the initial moment using the Pearson correlation coefficient, and take the moment when the correlation starts to be less than 0.90 as the division point of the healthy state , indicating that the bearing enters the degradation stage from the healthy state; the correlation coefficient calculation is shown in formula (1):

[0030] (1)

[0031] In the formula, is the correlation coefficient at moment is the - dimensional singular value, is the - dimensional singular value, is the average value of the - dimensional singular value along the dimension, The singular value obtained by compressing the features at the zero moment, is the singular value obtained by compressing the features at the current moment;

[0032] S202. Construct the RUL life curve. According to the health status of each bearing divided, divide the remaining life labels of each bearing into two parts: the stable state and the rapid decline state; where,

[0033] The stable state is:

[0034] ;

[0035] The rapid decline state is:

[0036] ;

[0037] In the formula, is the remaining life of the bearing at the stable state moment, is the total life of the bearing, is the moment of the bearing segmentation point, is the remaining life of the bearing at the rapid decline state moment, is the current moment;

[0038] S203. Optimize the loss function.

[0039] In step S203, optimizing the loss function includes:

[0040] The Weibull cumulative distribution function CDF is defined as the following formula (2), and different corresponds to the shape factor of different changing trends of the curve, is the characteristic life;

[0041] (2)

[0042] In the formula, is the bearing failure probability at the moment, is the natural constant,

[0043] Combining the Weibull cumulative distribution function with the traditional MSE loss function can obtain the Weibull-MSE loss function; the MSE loss function is shown in formula (3), and the Weibull-MSE loss function formula is as shown in (4);

[0044] (3)

[0045] ;

[0046] (4)

[0047] In the formula, is the MSE loss function, is the time step, is the predicted value, is the hybrid loss function, is the Weibull cumulative distribution function, is the Weibull-MSE loss function, is the Weibull loss function, is the hyperparameter of the weight ratio, is the true RUL label value, is the predicted RUL value, is the actual used time of the bearing, is the predicted used time of the bearing;

[0048] During the training process of the deep neural network, the Weibull-MSE loss function that combines Weibull with the mean square error MSE is used as the loss function in the network. The failure probability of the bearing at different times is fused into the Weibull-MSE loss function through the Weibull cumulative distribution function. The Weibull-MSE loss function is combined with the RUL life curve of the bearing health state evaluation method, so that the prediction of RUL conforms to the actual degradation process of the bearing.

[0049] In step S3, the convolutional CRATE network architecture that fuses the dilated causal convolutional attention mechanism DCA and the multi-scale convolution MSC includes: the dilated causal convolutional attention mechanism DCA is the multi-head subspace convolutional attention that fuses the dilated causal convolution DCC and the CRATE structure, forming the multi-head subspace convolutional sub-attention; through the multi-scale convolution MSC module, the ability to extract spatial features is enhanced.

[0050] Furthermore, the synchronous extraction of the long-term dependence information and local degradation characteristics of the bearing signal is completed, including:

[0051] S301, construct the dilated causal convolution DCC. During the attention calculation process of the Transformer structure, the dilated causal convolution DCC is used to replace the standard self-attention mechanism to capture the local features of the bearing signal;

[0052] S302, embed the multi-scale convolution MSC module. The multi-scale convolution module is embedded in the main loop block of the CRATE structure, placed after the multi-head subspace convolutional attention and before the iterative shrinkage threshold algorithm ISTA to extract the multi-scale features of the bearing signal;

[0053] S303, Regressor optimization and RUL prediction. Use average pooling operation in the regressor along the time dimension to capture long-term dependencies;

[0054] S304, Construct a convolutional CRATE network architecture.

[0055] In step S301, construct a dilated causal convolution DCC, including:

[0056] (a)Exponential dilation factor ;

[0057] (b)Unilateral padding strategy. The extracted feature data is , where is the total number of time steps, and each represents the bearing feature at the -th time step; To complete the convolution operation, pad the data unilaterally, and the padded sequence is , is the padded data, and the number of is determined by , where is the convolution kernel size, is the dilation factor, is the number of padding ;

[0058] In step S302, embed a multi-scale convolution MSC module, including:

[0059] (1)The multi-scale convolution MSC, after the attention mechanism, introduces convolution kernels of different sizes to extract features of different scales in parallel;

[0060] (2)Feature fusion strategy. The input data dimension is , and after feature extraction with different convolution kernel sizes, generates four parallel output results; As the output of the convolution pool, ensure ; is the sequence length, is the dimension of the first convolution module, is the dimension of the second convolution module, is the dimension of the third convolution module;

[0061] (5)

[0062] In the formula, is the output after splicing of multiple convolution modules, is the splicing operation, are the outputs of the first, second, third, and fourth convolution modules, is the selected dimension parameter;

[0063] (3) ReLU activation function and BatchNorm normalization, the formula is:

[0064] (6)

[0065] In the formula, is the activation function, is the batch normalization operation, is the output of the multi-scale convolution module.

[0066] In step S303, the input of the regressor is ; is the input of the regressor at time

[0067] (i) Average pooling is adopted to perform pooling along the time dimension in the regressor to extract long-term dependence information:

[0068] (7)

[0069] In the formula, is the output of the pooling layer, is the layer normalization operation, is the average value operation, is the specified dimension parameter, is the original time series;

[0070] (ii) The fully connected layer outputs the final predicted value, and the high-dimensional features are converted into RUL predicted values through linear mapping:

[0071] (8)

[0072] In the formula, is the weight matrix of the fully connected layer, is the bias vector of the fully connected layer, is the output of the regressor.

[0073] In step S304, a convolutional CRATE network architecture is constructed, including:

[0074] S3041, the encoding reduction principle, defines the encoding rate of a given feature in a specific feature space as follows:

[0075] (9)

[0076] In the formula, is the encoding rate after compression of is the covariance matrix of the feature, representing the correlation between features, is the regularization parameter, To calculate the logarithmic determinant of a specified matrix;

[0077] According to the coding reduction principle, the multi-head subspace convolution attention mechanism and feature sparsification operation of the CRATE structure are derived;

[0078] S3042, perform the transparency improvement of the improved multi-head subspace convolution attention mechanism MSSA, including:

[0079] First, use the subspace attention mechanism to perform feature decomposition, and compress the input bearing features into the local signal model subspace, and the expression is:

[0080] (10)

[0081] In the formula, is the Query matrix in the self-attention, is the Key matrix in the self-attention, is the Value matrix in the sub-attention, is the th non-overlapping subspace;

[0082] Secondly, the transparent attention mechanism is calculated as:

[0083] (11)

[0084] In the formula, is the multi-head self-attention mechanism, is the th subspace, is the th Value matrix, is the activation function, is the th Query matrix, is the th Key matrix, is the Gaussian codebook coding precision, is the feature dimension, is the subspace dimension;

[0085] S3043, feature sparsification based on the ISTA iterative shrinkage threshold algorithm.

[0086] In step S3043, the feature sparsification based on the ISTA iterative shrinkage threshold algorithm includes:

[0087] When the ISTA block receives the output C of the previous module, ISTA compresses and sparsifies the features through the following iterative update process:

[0088] (12)

[0089] Wherein, is the ISTA block output, is the activation function, is the output of the previous module, is the optimization gradient of the target optimization function, is the th iteration, is the target optimization function, is the learning rate.

[0090] Another object of the present invention is to provide a rolling bearing remaining life prediction system based on a convolutional white box, which implements the rolling bearing remaining life prediction method based on the convolutional white box. The system includes:

[0091] A data preprocessing module, which is used to extract time-domain and frequency-domain features from the original vibration signal, perform noise reduction processing by combining singular value decomposition (SVD), remove random noise and interference in the signal, and provide input features;

[0092] A health state evaluation module, which is used to input the processed data into a deep neural network, dynamically evaluate the operation process of the bearing based on the correlation coefficient analysis method, divide the health state and the degradation state, and combine the Weibull-MSE loss function to train the deep neural network to make the prediction of the remaining useful life (RUL) conform to the actual degradation process of the bearing;

[0093] A feature extraction module, which completes the synchronous extraction of long-term dependence information and local degradation features of the bearing signal through a convolutional CRATE network architecture that fuses the dilated causal convolutional attention mechanism (DCA) and the multi-scale convolutional (MSC).

[0094] Combining all the above technical solutions, the beneficial effects of the present invention are as follows:

[0095] Combining health state evaluation to improve the accuracy of RUL prediction: Using the correlation coefficient analysis method to dynamically divide the health state and degradation state of the bearing, avoiding the limitations of the fixed degradation curve, and improving the accuracy of state division.

[0096] Enhancing the local feature extraction ability and accurately modeling the bearing degradation mode: Using the dilated causal convolutional attention mechanism (DCA), introducing dilated convolution in the Transformer structure, improving the ability to capture local degradation features, and making up for the deficiency of the traditional Transformer in local information extraction.

[0097] Multi-scale information fusion to enhance the local modeling ability of the model: By adopting a feature fusion strategy, the convolutional features of different scales are concatenated, enabling the model to capture both global trends and local details simultaneously, and improving the adaptability to complex degradation patterns.

[0098] CRATE structure to improve interpretability and prediction credibility: For the first time, a white-box architecture is applied in the field of RUL prediction, enhancing the transparency and interpretability of the model, which is helpful for the implementation in the re-engineering of deep learning projects.

[0099] Optimize the CRATE regression structure to enhance the stability of RUL prediction: MeanPooling is introduced into the regressor to obtain long-term dependence information, making the RUL prediction more stable and reducing the impact of short-term fluctuations.

[0100] The technical solution proposed by the present invention has good engineering generality, is applicable to the predictive maintenance requirements of high-value equipment such as power, wind power, aerospace, and numerically controlled machine tools, and can significantly enhance the operation stability and safety guarantee ability of industrial systems. It is expected to reduce the risk of unplanned downtime by more than 30% and effectively extend the service life of key components by more than 10%, with significant economic benefits and engineering application value, and has broad industrialization prospects in the fields of equipment health management and intelligent manufacturing.

[0101] The present invention combines the Weibull-MSE loss function with a health assessment method based on dynamic correlation coefficients for the first time to construct an interpretable CRATE network integrating Dilated Causal Convolution (DCA) and Multi-scale Convolution (MSC), effectively overcoming the problems of large prediction deviation, complex debugging, and high implementation threshold existing in traditional methods in actual deployment, and filling the technical gap in the field of a fusion RUL prediction architecture based on "white box + multi-scale + health dynamic recognition".

[0102] In the field of RUL prediction, there have long been bottlenecks such as inaccurate identification of health states, weak local feature extraction capabilities, and poor model deployability. Existing methods mostly rely on static threshold division or fixed degradation curves, making it difficult to adapt to individual differences in bearings. At the same time, the "black box" characteristics of traditional deep networks make it difficult to interpret and optimize prediction results. The present invention introduces a dynamic health assessment method based on the Pearson correlation coefficient and optimizes the degradation fitting accuracy by combining the Weibull-MSE composite loss function, significantly improving the prediction reliability. At the same time, the CRATE integrating convolution achieves the structural transparency of the network structure. This solution effectively alleviates the core difficulties in RUL prediction such as difficult state division, weak degradation modeling, and untraceable prediction, and constructs a technical foundation for evolving towards an industrial-level interpretable prediction model. There are generally two types of biases in existing methods: First, they overly rely on deep learning to automatically extract health state features and ignore the role of expert knowledge in state division and degradation judgment. Second, they lack attention to model interpretability, restricting the popularization and application of "white box" network architectures in industry. The present invention improves the prediction performance and promotes the in-depth application of interpretable models in the field of RUL prediction by introducing a dynamic health state division mechanism, combining expert experience to determine the degradation starting point, and constructing a CRATE framework integrating a convolution structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0103] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure;

[0104] Figure 1 is a flowchart of a method for predicting the remaining useful life of a rolling bearing based on a convolutional white box provided by an embodiment of the present invention;

[0105] Figure 2 is a flowchart of a method for evaluating the health state of a bearing provided by an embodiment of the present invention;

[0106] Figure 3 is a bearing fault change curve, a "bathtub" curve graph provided by an embodiment of the present invention;

[0107] Figure 4 is a schematic diagram of the principle of dilated causal convolution provided by an embodiment of the present invention;

[0108] Figure 5 is a schematic diagram of a multi-scale convolution module provided by an embodiment of the present invention;

[0109] Figure 6 is a schematic diagram of the principle of the convolutional CRATE network architecture provided by an embodiment of the present invention;

[0110] Figure 7 is a graph of the health state division and RUL label of Bearing1_4 provided by an embodiment of the present invention;

[0111] Figure 8 It is the prediction result diagram of Bearing1_1 of the present invention;

[0112] Figure 9 It is the prediction result diagram of Bearing1_2 of the present invention;

[0113] Figure 10 It is the prediction result diagram of Bearing1_3 of the present invention;

[0114] Figure 11 It is the prediction result diagram of Bearing1_4 of the present invention;

[0115] Figure 12 It is the prediction result diagram of Bearing1_5 of the present invention;

[0116] Figure 13 It is the prediction result diagram of Bearing1_6 of the present invention;

[0117] Figure 14 It is the prediction result diagram of Bearing1_7 of the present invention. Detailed implementation manners

[0118] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementations disclosed below.

[0119] The innovation of the present invention lies in: the present invention constructs a CRATE network structure that integrates dilated causal convolution (DCA) and multi-scale convolution (MSC), enhances the ability to synchronously capture long-term dependencies and local degradation features, and ensures the interpretability of the key steps of the prediction method; at the same time, a health state dynamic division mechanism based on the Pearson correlation coefficient is introduced, and the Weibull-MSE loss function is combined to improve the fitting ability of the prediction process to the actual degradation curve of the bearing.

[0120] Example 1, the remaining useful life prediction system for rolling bearings based on convolutional white box provided by the embodiment of the present invention includes:

[0121] A data preprocessing module, which is used to extract time-domain and frequency-domain features from the original vibration signal, and perform noise reduction processing in combination with singular value decomposition (SVD) to remove random noise and interference in the signal, improve the data quality, and provide more stable and accurate input features for subsequent analysis.

[0122] A health status assessment module is used to input the pre - processed data into a deep neural network. Based on the correlation coefficient analysis method, it dynamically assesses the operation process of the bearing, accurately divides the health status and degradation status, and combines the Weibull - MSE loss function for deep neural network training to make the prediction of RUL conform to the actual degradation process of the bearing. This enables the deep neural network to better adapt to the actual degradation law of the bearing and improve the accuracy and robustness of the prediction.

[0123] Through the convolutional CRATE network architecture that fuses the dilated causal convolutional attention mechanism (DCA) and multi - scale convolution (MSC), the synchronous extraction of long - term dependence information and local degradation features of bearing signals is completed;

[0124] The dilated causal convolutional attention mechanism (DCA) is the multi - head subspace convolutional sub - attention formed by fusing the dilated causal convolution (DCC) and the CRATE structure's multi - head subspace convolution attention; through multi - scale convolution (MSC), the ability to extract spatial features is enhanced.

[0125] Example 2, as Figure 1 shown, the method for predicting the remaining useful life of a rolling bearing based on convolutional white - box provided by the embodiment of the present invention includes:

[0126] S1. Extract the time - domain and frequency - domain features of the original vibration signal, and perform noise reduction processing by combining singular value decomposition (SVD) to remove random noise and interference in the signal and provide input features;

[0127] Specifically, it includes:

[0128] S101. Calculate the time - domain statistical features (mean, standard deviation, kurtosis, skewness, etc.) of the original vibration signal to characterize the overall trend of the bearing signal.

[0129] S102. Use the fast Fourier transform (FFT) to extract the frequency - domain features of the vibration signal to enhance the representation ability of the extracted features in the frequency domain of the original vibration signal.

[0130] S103. Use the discrete wavelet transform (DWT) for time - frequency analysis, and enhance the extraction of high - frequency and low - frequency features of the unstable original vibration signal through wavelet transformation to retain the key time - frequency information of bearing faults.

[0131] S104. Normalize the features min - max extracted in steps S101 - S103 to construct the input data of the deep neural network.

[0132] S105. For further dividing the health status subsequently, use singular value decomposition (SVD) for noise reduction, record the singular values obtained from the singular value decomposition and use them to replace the original features to achieve the purpose of removing random noise and compressing the original features.

[0133] S2. Input the processed data into a deep neural network. Based on the correlation coefficient analysis method, dynamically evaluate the operating process of the bearing, divide the healthy state and the degradation state, and combine the Weibull-MSE loss function to train the deep neural network to make the prediction of the remaining useful life (RUL) conform to the actual degradation process of the bearing;

[0134] In the present invention, the deep neural network is a general term for the overall modeling framework for RUL prediction. The innovation of the present invention lies in designing a deep neural network with a specific structure, namely a convolutional neural network based on the CRATE (Coding Rate Autoencoding Transformer Explanation) architecture, which integrates the dilated causal convolution (DCA) and the multi-scale convolution (MSC) modules for simultaneously extracting long-term dependencies and local degradation features. Therefore, the CRATE network itself is the implementation structure of the deep neural network referred to in the present invention, and the specific schematic diagram is as Figure 6 shown.

[0135] Specifically, it includes: S201. Calculate the correlation between the features at each moment after SVD compression and the features at the initial moment using the Pearson correlation coefficient, and take the moment when the correlation starts to be less than 0.90 as the division point p of the healthy state, marking that the bearing enters the degradation stage from the healthy state. The calculation of the correlation coefficient is shown in formula (1);

[0136] (1)

[0137] In the formula, is the correlation coefficient at the moment, is the singular value of the dimension, is the singular value of the dimension, is the average value of the singular values along the dimension, is the average value of the singular values along the dimension, is the singular value obtained by compressing the feature at the zero moment,

[0138] is the singular value obtained by compressing the feature at the current moment;

[0139] In the healthy state, the default life changes very slowly and remains unchanged. Linear degradation is adopted in the degradation stage to enable mapping of the actual degradation process of the bearing. According to the health state of each bearing divided, the remaining life labels of each bearing can be divided into two parts: the stable state and the rapid decline state. The flow of the bearing health state evaluation method is as Figure 2 shown.

[0140] The stable state is:

[0141] ;

[0142] The rapid decline state is:

[0143] ;

[0144] In the formula, is the remaining life of the bearing at the stable state time, is the total life of the bearing, is the bearing segmentation point time, is the remaining life of the bearing at the rapid decline state time, is the current time;

[0145] S203, optimize the loss function.

[0146] In industrial production, it is known that the law of the bearing failure rate changing with time can be summarized as the "bathtub" curve. In the initial stage, the bearing needs to run in new equipment, so the failure rate is relatively high; after entering the normal use period, the bearing will enter a relatively long stable period, and the failure rate is constantly low at this time; after running for a certain period of time, the bearing will inevitably enter the decline period, and the failure rate begins to rise gradually. Among them, the bearing failure change curve, the "bathtub" curve is as Figure 3 shown.

[0147] The Weibull cumulative distribution function (CDF) can exactly represent the three states of the initial stage, normal use period, and decline period of the failure rate changing with time. The CDF is defined as follows (2), different corresponds to the shape factor of different change trends of the curve, is the characteristic life;

[0148] (2)

[0149] In the formula, is the bearing failure probability at time is the natural constant, is the current time;

[0150] Combining the Weibull cumulative distribution function with the traditional MSE loss function can obtain the Weibull-MSE loss function; the MSE loss function is shown in formula (3), and the Weibull-MSE loss function formula is as shown in (4);

[0151] (3)

[0152] ;

[0153] (4)

[0154] In the formula, is the MSE loss function, is the number of time steps, is the predicted value, is the hybrid loss function, is the Weibull cumulative distribution function, is the Weibull-MSE loss function, is the Weibull loss function, is the weight ratio hyperparameter, is the true RUL label value, is the RUL predicted value, is the actual used time of the bearing, is the predicted used time of the bearing;

[0155] Exemplarily, during the training process of the deep neural network, the Weibull-MSE loss function that combines Weibull with mean square error (MSE) is used as the loss function in the network. By means of the Weibull cumulative distribution function, the failure probability of the bearing at different times is fused into the Weibull-MSE loss function, and the Weibull-MSE loss function is combined with the RUL life curve of the bearing health state evaluation method, making the prediction of RUL more in line with the actual degradation process of the bearing.

[0156] S3. Through the convolutional CRATE network architecture that fuses the dilated causal convolutional attention mechanism DCA and the multi-scale convolution MSC, the synchronous extraction of the long-term dependence information and local degradation characteristics of the bearing signal is completed.

[0157] The dilated causal convolutional attention mechanism (DCA) is the multi-head subspace convolutional attention that fuses the dilated causal convolution (DCC) and the CRATE structure, forming the multi-head subspace convolutional sub-attention; through the multi-scale convolution (MSC) module, the ability to extract spatial features is enhanced;

[0158] Among them, DCA (Dilated Causal Convolution Attention Mechanism) enhances the ability to extract local features, compensates for the deficiency of traditional sub-attention mechanisms in modeling locally degraded information, and ensures that the convolutional CRATE network can accurately capture the short-term change trend of bearing signals.

[0159] MSC (Multi-Scale Convolution): Add MSC within the basic CRATE structure to extract signal features at different scales, enhance the adaptability to bearing degradation patterns, enable the deep neural network to simultaneously focus on macroscopic trends and microscopic changes, and improve the comprehensiveness of prediction.

[0160] Exemplarily, the convolutional CRATE network architecture: adopts a mathematically interpretable optimization objective to make the training and inference processes of the deep neural network transparent, and ensures that the prediction results are more credible and valuable for industrial applications.

[0161] Exemplarily, the convolutional CRATE network architecture that fuses the Dilated Causal Convolution Attention Mechanism (DCA) and Multi-Scale Convolution (MSC) completes the synchronous extraction of long-term dependence information and local degradation features of bearing signals, specifically including:

[0162] S301, construct Dilated Causal Convolution (DCC). As Figure 4 The schematic diagram of Dilated Causal Convolution;

[0163] In the attention calculation process of the Transformer structure, use Dilated Causal Convolution (DCC) to replace the standard self-attention mechanism to improve the ability of the basic CRATE network architecture to capture local features of bearing signals. The specific implementation is as follows:

[0164] (a) Exponential dilation factor : Ensure that the model can perceive the local features of bearing signals at different time scales, making it take into account short-term degradation features and long-term trends.

[0165] (b) Unilateral padding strategy: Ensure the causality of time series data, avoid leakage of future time step information, enable the model to truly simulate the degradation process of bearings, and improve the credibility of prediction.

[0166] Due to the causal characteristics of this convolution, it is required that the features at a certain moment in the next layer can only receive the features at the current moment and its previous moments. Due to this causal limitation, when performing the Padding operation on the dilated convolution, in order to keep the sizes of the front and back layers consistent, padding will only be added on one side:

[0167] Assume that the feature data extracted by the present invention is , where is the total number of time steps, and each represents the Bearing features at one time step; to complete the convolution operation, the present invention needs to pad data on one side, and the padded sequence is , To pad the data, The number of is determined by where is the convolution kernel size, is the dilation factor, is the number of

[0168] S302, embed the multi-scale convolution (MSC) module. As shown in Figure 5 Schematic diagram of the multi-scale convolution module;

[0169] The multi-scale convolution module is embedded in the main loop block of the CRATE structure, placed after the multi-head subspace convolution attention and before the iterative shrinkage threshold algorithm (ISTA) to fully extract the multi-scale features of the bearing signal. The specific implementation is as follows:

[0170] (1) Multi-scale convolution (MSC): After the attention mechanism, convolution kernels of different sizes (such as 1×3, 1×5) are introduced to extract features of different scales in parallel to enhance the model's ability to extract multi-scale spatial information and local features. The convolution kernel size distributions of the 4 parallel extraction routes are 1×1; 1×3 and 1×1; 1×5 and 1×1; 1×2 (pooling) and convolution pool (Branch Pool).

[0171] (2) Feature fusion strategy: Concatenate the output features of different convolution kernels to simultaneously utilize local information and global trends and improve the generalization ability of the model.

[0172] Assume that the input data dimension is , and after feature extraction with different convolution kernel sizes, Four parallel output results are generated. To ensure that the feature dimensions before and after entering the multi-scale convolution module do not change, As the output of the convolution pool, it ensures that ; is the sequence length, is the dimension of the 1st convolution module, is the dimension of the 2nd convolution module, is the dimension of the 3rd convolution module;

[0173] (5)

[0174] In the formula, is the output after concatenation of the multi-convolution module, is the concatenation operation, is the output of the 1st, 2nd, 3rd, and 4th convolution modules, is the selected dimension parameter;

[0175] (3) ReLU activation function and BatchNorm normalization: Maintain the non-linear expression ability of the network, accelerate convergence, and reduce the gradient vanishing problem. The formula is expressed as follows:

[0176] (6)

[0177] In the formula, is the output of the pooling layer, is the layer normalization operation, is the average value operation, is the specified dimension parameter, is the original time series;

[0178] S303, Regressor optimization and RUL prediction.

[0179] Compared with the traditional CRATE architecture, the present invention has made improvements in the regressor part:

[0180] In order to better extract global information, the present invention uses average pooling operation along the time dimension in the regressor to help the model better capture long-term dependencies. Assume the input of the regressor is: ; is the input of the regressor at time

[0181] (i) Adopt mean pooling: Pool along the time dimension in the regressor to fully extract long-term dependency information and make the RUL prediction more stable.

[0182] (7)

[0183] In the formula, is the output of the pooling layer, is the layer normalization operation, is the average value operation, is the specified dimension parameter, is the original time series;

[0184] This operation compresses the original time series Z along the time dimension into a feature vector of a fixed length, that is, by calculating the average value along the time dimension direction, the global feature information of all time steps is retained, enhancing the network's capture of long-term information.

[0185] (ii) The fully connected layer outputs the final predicted value: Convert the high-dimensional features into RUL predicted values through linear mapping to improve the model's expression ability and accuracy.

[0186] (8)

[0187] In the formula, is the weight matrix of the fully connected layer, is the bias vector of the fully connected layer, is the output of the regressor.

[0188] Since the remaining useful life (RUL) of a rolling bearing is a definite value, while the features extracted and compressed in a deep learning network are usually high-dimensional vectors. To achieve accurate prediction of RUL, these multi-dimensional features need to be mapped to a scalar-form RUL value through a fully connected layer, so as to complete the regression task of the model and output a clear life prediction result.

[0189] S304. Construct a convolutional CRATE network architecture, as shown in Figure 6 the schematic diagram of the convolutional CRATE network architecture;

[0190] Principles of the convolutional CRATE network architecture: Through the Coding Rate Reduction principle, map the high-dimensional input features to a low-dimensional space, and conduct strict mathematical derivation on the feature distribution to ensure that each step of feature processing of the model has a clear mathematical basis, enhancing the transparent interpretability of the model for the degradation process. Including:

[0191] S3041. Coding reduction principle: The CRATE structure is based on the concept of Coding Rate in information theory, aiming to maximize the compression efficiency of meaningful information in the feature space, ensuring that the model only retains the core information related to RUL prediction and reducing redundant and noise interference.

[0192] Define the coding rate of a given feature X in a specific feature space as follows, where is the coding rate after compression of is the covariance matrix of the feature, representing the correlation between features, is the regularization parameter to prevent numerical instability, is to calculate the log determinant of the specified matrix;

[0193] (9)

[0194] In the formula, is the coding rate after compression of is the covariance matrix of the feature, representing the correlation between features, is the regularization parameter, is to calculate the log determinant of the specified matrix;

[0195] Based on this principle, the multi-head subspace convolution attention mechanism and feature sparsification operation of the CRATE structure are derived.

[0196] S3042, perform the transparency of the multi-head subspace convolution attention mechanism (MSSA).

[0197] The present invention uses an improved multi-head subspace convolution attention mechanism (MSSA), combined with the coding rate reduction target, to achieve the mathematical inferability of the model attention and improve the transparency of the prediction process.

[0198] First, the subspace attention mechanism is used for feature decomposition to compress the input bearing features into a local signal model subspace:

[0199] (10)

[0200] In the formula, is the Query matrix in the self-attention, is the Key matrix in the self-attention, is the Value matrix in the sub-attention, is the th non-overlapping subspace;

[0201] Secondly, the transparent attention mechanism is calculated as:

[0202] (11)

[0203] In the formula, is the multi-head self-attention mechanism, is the kth subspace, is the ith Value matrix, is the activation function, is the ith Query matrix, is the ith Key matrix, is the Gaussian codebook coding accuracy, is the feature dimension, is the subspace dimension;

[0204] Here, the MSSA is basically similar to the multi-head self-attention operator in the standard Transformer, except that the linear operators are all set to be the same as the subspace basis, that is, ;

[0205] S3043, feature sparsification based on the ISTA iterative shrinkage threshold algorithm.

[0206] In the convolutional CRATE network architecture, the Iterative Shrinkage Thresholding Algorithm (ISTA) is used to achieve feature sparsification, retaining the key information most relevant to RUL prediction. When the ISTA block receives the output C from the previous module (here referring to the output of the multi-scale convolutional module), ISTA compresses and sparsifies the features through the following iterative update process, defined as follows:

[0207] (12)

[0208] In the formula, is the output of the ISTA block, is the activation function, is the output of the previous module, is the optimization gradient of the objective optimization function, is the th iteration, is the objective optimization function, is the learning rate.

[0209] As can be seen from the above embodiments, the present invention fully considers the correlation coefficient analysis health state assessment method of time-frequency domain signals, and improves the accuracy of degradation stage division by dynamically evaluating the bearing health state.

[0210] The rolling bearing life prediction method that combines the Weibull loss function with the new health state assessment method enables the model to optimize the RUL prediction by combining physical characteristics, improving the prediction accuracy and stability.

[0211] The white-box Transformer (CRATE) structure combined with convolution enhances the local feature extraction ability through dilated causal convolution (DCA) and multi-scale convolution (MSC), and improves the adaptability to bearing degradation modes.

[0212] For the first time, a fully mathematically inferable architecture and a transparent model are applied in the field of rolling bearing remaining life prediction. Using the CRATE structure makes the prediction process transparent, improving the interpretability and industrial application value.

[0213] Another exemplary one is that in the rolling bearing health state assessment method, time-frequency domain analysis means such as wavelet transform and Hilbert-Huang transform can also be added to common time-frequency domain analysis methods to enhance the feature extraction of vibration signals. In the correlation analysis method, the Pearson correlation coefficient can also be replaced by methods such as cluster analysis, principal component analysis, and dynamic time warping for autocorrelation / cross-correlation / similarity analysis.

[0214] To further illustrate the relevant effects of the embodiments of the present invention, the following experiments are carried out.

[0215] Existing methods usually adopt fixed degradation curves or simple threshold division to determine the bearing health state, ignoring individual differences. The present invention uses the correlation coefficient analysis method of time-frequency domain signals to dynamically divide the health state and degradation state, enabling the model to more accurately capture the true degradation process of the bearing and improving the reliability of RUL prediction.

[0216] Most existing deep learning methods use loss functions such as mean squared error (MSE), which do not fully consider the physical degradation law of bearings, resulting in prediction deviations. The present invention combines Weibull with a new health state evaluation method, fully considering the physical characteristics of rolling bearings, enhancing the model's modeling ability and the prediction accuracy of remaining life.

[0217] Traditional Transformer structures mainly focus on global information and are difficult to effectively model the local degradation characteristics of bearings. The present invention combines the dilated causal convolutional attention mechanism (DCA) with a multi-scale convolutional module to improve the ability to extract short-term degradation information and make the prediction more accurate.

[0218] Most existing Transformer and deep learning models are "black box" structures, making it difficult to understand their internal decision-making processes, which affects their application in industrial scenarios. The present invention first introduces the completely mathematically inferable CRATE structure into RUL prediction, making the model optimization process transparent, improving the credibility of prediction results, and providing a more reliable solution for industrial equipment health management.

[0219] The prediction accuracy of the present invention is improved. After obtaining the preprocessed data features, the present invention evaluates the bearing health state according to the above-mentioned health state division method. See Table 1.

[0220] Table 1 Bearing health state division

[0221]

[0222] According to the health state of each bearing divided, the present invention can divide the remaining life label of each bearing into two parts: stable state and rapid decline state. Among them,

[0223] The stable state is:

[0224] ;

[0225] The rapid decline state is:

[0226] ;

[0227] To more intuitively demonstrate the effectiveness of state division, the present invention compares and shows the original signal, Pearson correlation coefficient, and RUL life curve corresponding to Bearing1_4. Figure 7Health status division and RUL label graph for Bearing1_4, showing the trend consistency between the correlation coefficient curve and the original signal, and transmitting the information to the RUL life curve.

[0228] The prediction results of the present invention for Bearing1_1 to Bearing1_7 are as Figures 8 - 14 shown, and Table 2 shows the specific data. It can be seen from the figure that the method proposed by the present invention has achieved the expected goal on all bearings, distinguished the two working states of the bearings and accurately extracted the local information near the bearing degradation point.

[0229] Table 2 Bearing prediction results

[0230]

[0231] The above is only a relatively preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.

Claims

1. A method for predicting the remaining useful life of a rolling bearing based on convolutional white box, characterized in that, The method includes the following steps: S1. Extract time-domain and frequency-domain features from the original vibration signal, perform noise reduction processing by combining singular value decomposition (SVD) to remove random noise and interference in the signal, and provide input features; S2. Input the processed data into a deep neural network, dynamically evaluate the operation process of the bearing based on the correlation coefficient analysis method, divide the healthy state and the degradation state, and train the deep neural network by combining the Weibull-MSE loss function to make the prediction of the remaining useful life (RUL) conform to the actual degradation process of the bearing; S3. Through the convolutional CRATE network architecture that fuses the dilated causal convolutional attention mechanism (DCA) and the multi-scale convolution (MSC), complete the synchronous extraction of the long-term dependence information and local degradation features of the bearing signal; In step S3, the convolutional CRATE network architecture that fuses the dilated causal convolutional attention mechanism (DCA) and the multi-scale convolution (MSC) includes: The dilated causal convolutional attention mechanism (DCA) is a multi-head subspace convolutional attention that fuses the dilated causal convolution (DCC) and the CRATE structure, forming a multi-head subspace convolutional sub-attention; Through the multi-scale convolution (MSC) module, enhance the ability to extract spatial features; Completing the synchronous extraction of the long-term dependence information and local degradation features of the bearing signal includes: S301. Construct the dilated causal convolution (DCC). During the attention calculation process of the Transformer structure, use the dilated causal convolution (DCC) to replace the standard self-attention mechanism to capture the local features of the bearing signal; S302. Embed the multi-scale convolution (MSC) module. The multi-scale convolution module is embedded in the main loop block of the CRATE structure, placed after the multi-head subspace convolutional attention and before the iterative shrinkage threshold algorithm (ISTA) to extract the multi-scale features of the bearing signal; S303. Optimize the regressor and predict the RUL. Use the average pooling operation along the time dimension in the regressor to capture the long-term dependence relationship; S304. Construct the convolutional CRATE network architecture.

2. The remaining life prediction method of a rolling bearing based on convolutional white box according to claim 1, wherein In step S2, based on the correlation coefficient analysis method, dynamically evaluate the operation process of the bearing, divide the healthy state and the degradation state, and train the deep neural network by combining the Weibull-MSE loss function, including: S201. Calculate the correlation between the features at each moment after SVD compression and the features at the initial moment using the Pearson correlation coefficient. The moment when the correlation starts to be less than 0.90 is used as the division point p of the healthy state, indicating that the bearing enters the degradation stage from the healthy state; The correlation coefficient calculation is shown in formula (1): where r y is the correlation coefficient at time y, x i is the singular value of the i-th dimension of x, y i is the singular value of the i-th dimension of y, is the average value of the singular values of the x extended dimension, is the average value of the singular values of the y along the dimension, x is the singular value obtained by feature compression at the zero time, and y is the singular value obtained by feature compression at the current time; S202. Construct the RUL life curve. According to the healthy state of each bearing divided, divide the remaining life label of each bearing into two parts: the stable state and the rapid decline state; Among them, The stable state is: RUL normal (t)=T - T p The rapid decline state is: RUL rapid (t) = T - t Wherein, RUL normal (t) is the remaining life of the bearing at the stable state at time t, T is the total life of the bearing, T p is the time of the bearing segmentation point, RUL rapid (t) is the remaining life of the bearing at the rapid decline state at time t, and t is the current time; S203. Optimize the loss function.

3. The method for predicting the remaining life of a rolling bearing based on a convolutional white box according to claim 2, wherein In step S203, optimizing the loss function includes: The Weibull cumulative distribution function (CDF) is defined as formula (2) below. Different β corresponds to the shape factor of different curve change trends, and η is the characteristic life; Where \(F(t)\) is the bearing fault probability at time \(t\), \(e\) is the natural constant, and \(t\) is the current time; Combining the Weibull cumulative distribution function with the traditional MSE loss function can obtain the Weibull-MSE loss function; the MSE loss function is shown in formula (3), and the Weibull-MSE loss function formula is shown in (4); L hybrid = L mse + λL weibull where L mse is the MSE loss function, n is the number of time steps, is the predicted value, L hybrid is the hybrid loss function, F() is the Weibull cumulative distribution function, L weibull is the Weibull-MSE loss function, is the Weibull loss function, λ is the weight ratio hyperparameter, t i is the true RUL label value, is the RUL predicted value, T i is the actual operating time of the bearing, is the predicted operating time of the bearing; During the training process of the deep neural network, the Weibull-MSE loss function that combines Weibull with the mean square error MSE is used as the loss function in the network. The fault probability of the bearing at different times is fused into the Weibull-MSE loss function through the Weibull cumulative distribution function. The Weibull-MSE loss function is combined with the RUL life curve of the bearing health state assessment method, making the prediction of RUL conform to the actual degradation process of the bearing.

4. The method for predicting the remaining life of a rolling bearing based on a convolutional white box according to claim 1, characterized in that In step S301, construct the dilated causal convolution DCC, including: (a) The exponential expansion factor \(d = 1, 2, 4, 8\); (b) Unilateral filling strategy, the extracted feature data is X = {x1, x2…x T} where T is the total number of time steps, and each x t represents the bearing feature at the t-th time step; to complete the convolution operation, data is filled unilaterally, and the filled sequence is X pad = {p, p, p…x1, x2…x T}, where p is the filled data, and the number of p is determined by P = (K - 1) / D, where K is the convolution kernel size, D is the dilation factor, and P is the number of filled p; In step S302, embed the multi-scale convolution MSC module, including: (1) After the attention mechanism, the multi-scale convolution MSC introduces convolution kernels of different sizes to extract features of different scales in parallel; (2) Feature fusion strategy. The input data dimension is \(L\times D\). After feature extraction with different convolution kernel sizes, four parallel output results \(L\times D1\), \(L\times D2\), \(L\times D3\), \(L\times D4\) are generated; \(D4\) is the output of the convolution pool, ensuring \(D = D1 + D2 + D3 + D4\); \(L\) is the sequence length, \(D1\) is the dimension of the 1st convolution module, \(D2\) is the dimension of the 2nd convolution module, and \(D3\) is the dimension of the 3rd convolution module; Output = torch.cat((Output1, Output2, Output3, Output4), axis = -1) (5) Where Output is the output after splicing the multi-convolution modules, torch.cat() is the splicing operation, Output1, Output2, Output3, Output4 are the outputs of the 1st, 2nd, 3rd, and 4th convolution modules, and axis is the selected dimension parameter; (3) ReLU activation function and BatchNorm normalization, the formula is: Output = Relu(BatchNorm(Output)) (6) Where Relu() is the activation function, BatchNorm() is the batch normalization operation, and Output is the output of the multi-scale convolution module.

5. The method for predicting the remaining life of a rolling bearing based on convolutional white box according to claim 1, wherein In step S303, the input of the regressor is Z = {z1, z2, z3…z T}; z T is the input of the regressor at time T; (i) Use average pooling to perform pooling along the time dimension in the regressor to extract long-term dependence information: Z′ = LayerNorm(Z.mean(dim = 1)) (7) Where Z′ is the output of the pooling layer, LayerNorm() is the layer normalization operation, mean() is the average value operation, dim is the specified dimension parameter, and Z is the original time series; (ii) The fully connected layer outputs the final prediction value, and the high-dimensional features are converted into the RUL prediction value through linear mapping: Z″ = Z′W + b (8) Wherein, W is the weight matrix of the fully connected layer, b is the bias vector of the fully connected layer, and Z″ is the output of the regressor.

6. The method for predicting the remaining life of a rolling bearing based on convolutional white box according to claim 1, characterized in that, In step S304, a convolutional CRATE network architecture is constructed, including: S3041, the coding reduction principle, defines the coding rate of a given feature X in a specific feature space as follows: Wherein, R(X) is the coding rate after compression of X, ∑x is the covariance matrix of the features, representing the correlation between the features, ω is the regularization parameter, and logdet() is to calculate the logarithmic determinant of the specified matrix; According to the coding reduction principle, the multi-head subspace convolutional attention mechanism and feature sparsification operation of the CRATE structure are derived. S3042, perform the transparency of the improved multi-head subspace convolutional attention mechanism MSSA, including: First, the subspace attention mechanism is used for feature decomposition to compress the input bearing feature X ∈ R T×d into the local signal model subspace, and the expression is: Where Q i is the Query matrix in self-attention, K i is the Key matrix in self-attention, V i is the Value matrix in sub-attention, is the i-th non-overlapping subspace; Secondly, the transparent attention mechanism is calculated as: Wherein, MSSA(X|U k ) is the multi-head self-attention mechanism, U k is the k-th subspace, V i is the i-th Value matrix, softmax() is the activation function, is the i-th Query matrix, K i is the i-th Key matrix, ε is the Gaussian codebook encoding precision, N is the feature dimension, and T is the subspace dimension; S3043, feature sparsification based on the ISTA iterative shrinkage threshold algorithm.

7. The method for predicting the remaining life of a rolling bearing based on convolutional white box according to claim 6, wherein In step S3043, the feature sparsification based on the ISTA iterative shrinkage threshold algorithm includes: When the ISTA block receives the output C from the previous module t+1 / 2 the ISTA compresses and sparsifies the features through the following iterative update process: Where C t+1 is the output of the ISTA block, Relu() is the activation function, C t+1 / 2 is the output of the previous module, is the optimization gradient of the objective optimization function, t is the t-th iteration, f(h) is the objective optimization function, and α is the learning rate.

8. A rolling bearing remaining life prediction system based on convolutional white box, characterized in that, The system implements the rolling bearing remaining life prediction method based on convolutional white box as described in any one of claims 1-7. The system includes: A data preprocessing module, which is used to extract time-domain and frequency-domain features from the original vibration signal, perform noise reduction processing by combining singular value decomposition (SVD), remove random noise and interference in the signal, and provide input features. A health state evaluation module, which is used to input the processed data into a deep neural network, dynamically evaluate the operation process of the bearing based on the correlation coefficient analysis method, divide the health state and the degradation state, and combine the Weibull-MSE loss function to train the deep neural network to make the prediction of RUL conform to the actual degradation process of the bearing. A feature extraction module, which completes the synchronous extraction of the long-term dependence information and local degradation features of the bearing signal through a convolutional CRATE network architecture that fuses the dilated causal convolutional attention mechanism (DCA) and the multi-scale convolution (MSC).

Citation Information

Patent Citations

  • Battery health monitoring and predicting method based on machine learning

    CN118759398A

  • Adaptive stage division and residual life prediction method for multi-stage degraded bearing

    CN119848671A