Bearing residual life prediction method based on multi-teacher element weight knowledge distillation network

By combining a multi-teacher meta-weight knowledge distillation network with multiple teacher models and an adaptive meta-weight strategy gradient learning algorithm, a lightweight student model is constructed, which solves the difficulty of deploying deep neural network models on edge devices and achieves high-precision prediction of the remaining life of bearings.

CN120653962APending Publication Date: 2025-09-16SICHUAN UNIV +1
View PDF 0 Cites 11 Cited by

Patent Information

Application Number
CN202510767552.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing deep neural network models are difficult to deploy on resource-constrained edge computing devices, and the knowledge expression capability of a single teacher model is insufficient, resulting in limited prediction accuracy.

Method used

The Multi-Teacher Meta-Weight Knowledge Distillation Network (MTMWKDN) is adopted. By combining three types of teacher models, LSTM-Attention, TCN-BiLSTM and Transformer, an adaptive meta-weight strategy gradient learning algorithm is designed to construct a lightweight student model to achieve the optimization of knowledge transfer and task fitting.

Benefits of technology

While ensuring prediction accuracy, the model complexity is significantly reduced, enabling it to be actually deployed on edge devices to achieve accurate prediction of the remaining life of bearings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653962A_ABST
    Figure CN120653962A_ABST
Patent Text Reader

Abstract

The invention discloses a bearing residual life prediction method based on a multi-teacher element weight knowledge distillation network, and the method comprises the following steps: (1) carrying out the normalization, noise reduction and segmentation of a multi-dimensional vibration time sequence signal collected by an original sensor, and dividing the signal into a training set and a verification set; (2) inputting the processed training set data into a multi-teacher element weight knowledge distillation network, training teacher models, and fixing parameters of the three teacher models after training is completed; (3) performing knowledge distillation training on the student model by using an adaptive meta-weight strategy gradient learning algorithm; and (4) the student model after distillation training is used for bearing residual life prediction. According to the method, the number of network model parameters is small, the calculation complexity is low, the prediction precision is high, and the method can be actually deployed on edge equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remaining life prediction, and in particular relates to a bearing remaining life prediction method based on a multi-teacher weighted knowledge distillation network. Background Art

[0002] In the field of remaining life prediction, prediction methods can be summarized as physical model-based methods, data-driven methods, and hybrid methods that integrate the first two.

[0003] Physical model-based prediction methods, grounded in knowledge of mechanical principles, material mechanics, and fatigue fracture theory, build mathematical models to analyze key parameters such as device structural characteristics, material properties, and loading conditions, thereby characterizing the degradation process. Commonly used mathematical models include the Paris crack growth model, the Forman-Newman-de Koning model, and the spalling degradation model. While these physical model-based prediction methods have a clear theoretical foundation, actual failure modes are complex, diverse, and overlapping, making it challenging to accurately identify the dominant failure mode. Models constructed based solely on subjective suggestions of a few failure mechanisms often lack reliability and rationality, failing to accurately reflect real-world conditions. Furthermore, model construction relies on manual features derived from expert experience, resulting in poor portability and limited applicability. Thanks to the rapid development of condition monitoring technology, performance degradation data on equipment during operation has become relatively easy to collect. Due to the bottlenecks encountered by physical model methods and the limited availability of data, data-driven approaches have emerged. In comparison, data-driven approaches avoid in-depth physical analysis of the degradation process. Instead, they rely on massive amounts of monitoring data to mine underlying patterns and predict remaining life. This has become an emerging mainstream trend in the field of remaining life prediction. Data-driven prediction methods can be further divided into three technical categories based on modeling approaches: probabilistic statistics, traditional machine learning, and deep learning. Deep learning-based prediction methods use mechanical equipment operating data as a foundation. By constructing deep neural networks, they learn the patterns of change in data throughout the entire life cycle, thereby making accurate predictions of remaining life. Deep learning models can automatically extract complex features that encompass a wealth of information from shallow to deep layers, from the time domain to the frequency domain. With their powerful learning and generalization capabilities, they can accurately capture the degradation patterns of equipment operating conditions. Duan Jiajun et al. designed a multi-scale feature extraction module based on traditional CNNs, mining and integrating degradation features at different scales to achieve Transformer-based prediction of aircraft engine life. Chen et al. introduced an attention mechanism based on LSTM networks to enhance the utilization of ordered information in time series data, thereby improving the accuracy of gear remaining life prediction. Liu Yefeng et al. proposed a bearing life prediction method by integrating a temporal convolutional network with a self-attention mechanism and a bidirectional gated recurrent unit. Experimental results show that the prediction model performs well. Overall, this method deeply analyzes data relationships, performs well in the prediction process, and has great potential for further exploration. Hybrid model-based prediction methods combine physical and data-driven models, leveraging the advantages of both to achieve more accurate equipment remaining life prediction. However, the amount of data required varies between different models. Data-driven models are suitable for large data sets, while physical models are advantageous for small sample sizes. Furthermore, the two models have different principles and some incompatibilities.

[0004] As the demand for prediction accuracy in the industrial sector continues to grow, researchers are turning to deep neural network models to capture complex degradation patterns. However, these models typically have a large number of parameters, making them difficult to deploy on resource-constrained edge computing devices. Therefore, achieving lightweight models while maintaining prediction accuracy has become a key challenge for the implementation of this technology. Summary of the Invention

[0005] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a bearing remaining life prediction method based on a multi-teacher weighted knowledge distillation network with a small number of parameters, low computational complexity, high prediction accuracy, and the ability to be actually deployed on edge devices.

[0006] The object of the present invention is achieved through the following technical solution: a bearing remaining life prediction method based on a multi-teacher weighted knowledge distillation network, comprising the following steps:

[0007] (1) The multi-dimensional vibration time series signal collected by the original sensor is normalized and denoised, and then the signal is segmented to form a standardized data set, which is divided into a training set and a validation set;

[0008] (2) Input the processed training set data into the multi-teacher meta-weight knowledge distillation network to train the teacher model;

[0009] The multi-teacher meta-weighted knowledge distillation network includes three parallel teacher models and one student model. The three teacher models use LSTM-attention, TCN-BiLSTM, and Transformer models respectively. Teacher model 1 includes a sequential LSTM, an attention mechanism, and a fully connected layer. Teacher model 2 includes a sequentially connected temporal convolutional neural network module, a bidirectional long short-term memory network, and a fully connected layer. Teacher model 3 includes an embedding layer, a positional encoder, and three cascaded encoding modules. The positional encoder is used to add the position information of the input data. The feature vector output by the embedding layer is added to the position information output by the positional encoder element by element, and the added features are then input into the encoding module.

[0010] The student model consists of three parallel one-dimensional convolution paths. The three convolution paths have the same structure. Each convolution path includes a convolution layer, an activation function, a convolution layer, an activation function, and a pooling layer connected in sequence. The outputs of the three convolution paths are fused and flattened before being output to the fully connected layer for processing.

[0011] Three teacher models are trained using lifespan label supervision, and the parameters of the three teacher models are fixed after training.

[0012] (3) After fixing the parameters of the three teacher models, the student model is trained with the adaptive meta-weight policy gradient learning algorithm for knowledge distillation; the trained student model is saved;

[0013] (4) The student model trained by distillation is used to predict the remaining life of bearings.

[0014] The beneficial effects of the present invention are as follows: in MTMWKDN, three types of models, LSTM-attention, TCN-BiLSTM, and Transformer, are used as teacher model clusters to complement each other in the feature extraction dimension, enabling the student model to integrate degradation feature information of different scales; at the same time, a student model with strong feature extraction capabilities, a small number of parameters, and low model complexity is constructed; at the same time, the adaptive meta-weight strategy gradient learning algorithm designed dynamically adjusts the weight coefficients of KL loss, soft loss, and hard loss through gradient learning, and can automatically balance the priority of knowledge transfer and task fitting during the training process. Due to the above advantages, the remaining life prediction method based on MTMWKDN of the present invention can obtain a network model with a small number of parameters, low computational complexity, high prediction accuracy, and can be actually deployed on edge devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is the overall framework diagram of the multi-teacher meta-weight knowledge distillation network of the present invention;

[0016] Figure 2 This is the structural diagram of the bidirectional long short-term memory network;

[0017] Figure 3 This is the framework diagram of teacher model 2;

[0018] Figure 4 This is the framework diagram of teacher model 3;

[0019] Figure 5 A structural diagram of the student model;

[0020] Figure 6 This is the training flow chart of the teacher model;

[0021] Figure 7 Schematic diagram of a two-layer optimization problem;

[0022] Figure 8 This is the training flow chart of the student model;

[0023] Figure 9 It is a bearing accelerated life test bench;

[0024] Figure 10 is the prediction result of task AF. DETAILED DESCRIPTION

[0025] In order to solve the problems that prediction models with high model complexity are difficult to actually deploy on edge devices with limited resources, and the single teacher model has insufficient knowledge expression ability and the fixed-weight composite loss function cannot dynamically adjust the contribution rate of each loss item according to the training process, resulting in limited prediction accuracy of the student model, the present invention proposes a rolling bearing remaining life prediction method based on a multi-teacher meta-weight knowledge distillation network (MTMWKDN). The purpose is to obtain a network model with fewer parameters, lower computational complexity, higher prediction accuracy and can be actually deployed on edge devices compared to complex deep neural networks, so as to achieve accurate and stable prediction of the remaining life of rolling bearings.

[0026] Knowledge distillation, as a widely recognized model compression technology, provides an effective solution to the contradiction between model performance and complexity. This technology transfers the knowledge of the complex teacher model to the lightweight student model through the teacher-student model architecture, thereby significantly reducing the complexity of the model while maintaining the prediction performance. Combining the advantages of knowledge distillation, the present invention proposes a multi-teacher meta-weight knowledge distillation network (MTMWKDN). In the multi-teacher meta-weight knowledge distillation network, heterogeneous multi-teacher collaborative distillation is designed, and three types of teacher models are used to enhance the degradation representation ability through differential feature complementarity; KL divergence is introduced as KL loss to assist the student model in learning the feature knowledge of the teacher model, and a composite loss function is designed to integrate KL loss, soft loss and hard loss to ensure the efficiency of knowledge transfer; combined with the idea of ​​meta-learning, an adaptive meta-weight strategy gradient learning algorithm is designed to optimize the weight distribution of KL loss, soft loss and hard loss in real time, and synchronously update the student network parameters, so that the model adapts to the training process of gradually increasing knowledge, thereby improving the prediction accuracy of the student model. The adaptive meta-weight policy gradient learning algorithm updates and optimizes the parameter set and hyperparameter set with the help of a large number of training tasks and a small number of validation tasks. In this way, the student model can better learn the knowledge of different teacher models and adapt to the entire training process, thereby improving the prediction accuracy of the student model.

[0027] The technical solution of the present invention is further described below with reference to the accompanying drawings.

[0028] The present invention provides a bearing remaining life prediction method based on a multi-teacher weighted knowledge distillation network, comprising the following steps:

[0029] (1) The multi-dimensional vibration time series signal collected by the original sensor is normalized and denoised, and then the signal is segmented to form a standardized data set, which is divided into a training set and a validation set;

[0030] (2) Input the processed training set data into the multi-teacher meta-weight knowledge distillation network to train the teacher model;

[0031] The Multi-Teacher Meta-Weight Knowledge Distillation Network (MTMWKDN) consists of three parallel teacher models, student models, a composite loss function, and an adaptive meta-weight policy gradient learning algorithm, such as Figure 1 As shown. The teacher model contains three models with different structures. Different teacher models have their own advantages in capturing different features and patterns of sensor data. The student model aims to be lightweight and is designed with multi-scale convolutional layers to efficiently extract data features of rotating machinery. The composite loss function consists of KL loss, soft loss and hard loss. The composite loss establishes a connection between the teacher model and the student model, and transfers the knowledge of the teacher model to the student model, which helps the student model learn a more comprehensive and rich feature representation, while ensuring the efficiency of knowledge transfer and improving the accuracy of the remaining life prediction of rolling bearings. The adaptive meta-weight strategy gradient learning algorithm updates and optimizes the parameter set and hyperparameter set with the help of a large number of training tasks and a small number of verification tasks. In this way, the student model can better learn the knowledge of different teacher models and adapt to the entire training process, thereby improving the prediction accuracy of the student model.

[0032] In the MTMWKDN model, three types of models, LSTM-Attention, TCN-BiLSTM, and Transformer, are used as teacher models, complementing each other in feature extraction, enabling the student model to integrate degraded feature information at different scales. Teacher Model 1 consists of a sequential LSTM, an attention mechanism, and a fully connected layer. Using the LSTM-Attention mechanism as Teacher Model 1 allows it to capture long-term dependencies and key features. Teacher Model 2 consists of a sequentially connected temporal convolutional neural network module, a bidirectional long short-term memory network, and a fully connected layer, combining the multi-scale feature extraction advantages of the temporal convolutional network with the bidirectional temporal modeling capabilities of the bidirectional long short-term memory network. Teacher Model 3 consists of an embedding layer, a positional encoder, and three cascaded encoding modules. The positional encoder adds positional information to the input data. The feature vector output by the embedding layer is element-wise added to the positional information output by the positional encoder, allowing the element vector at each position to carry both feature and temporal information. This added feature is then input into the encoding module. Each encoding module includes a sequential multi-head attention mechanism, layer normalization, a feedforward neural network, layer normalization, and a fully connected layer. The input and output data of the multi-head attention mechanism are residually connected (element-wise addition) and then layer normalization is performed. The input and output data of the feedforward neural network are residually connected and then layer normalization is performed. Teacher Model 3 is built on the Transformer and uses self-attention to model cross-cycle correlation patterns in data.

[0033] The Long Short-Term Memory Network (LSTM) effectively avoids the problem of gradient vanishing and exploding by virtue of its unique gating mechanism, and performs excellently in processing sequence data. LSTM not only memorizes past state information but also combines current input data to capture long-term data dependencies and learn richer feature representations. However, LSTM has a defect: it cannot directly evaluate the importance weights of different parts of the input sequence. This defect makes it difficult for the model to identify which time series features are more critical to remaining life prediction, which may lead to a decrease in prediction performance. To overcome this problem, the teacher model 1 of the present invention introduces an attention mechanism to enhance the model's ability to capture key information by dynamically evaluating the importance of time series features. Suppose the feature representation learned by the LSTM network for the input data is H = {H1, H2, ..., H m} T , T represents the transposition operation; after the attention mechanism is processed, the feature H at time step i i The importance score is expressed as:

[0034] P i =H i W T

[0035] Where W is the weight matrix; after obtaining the importance score of the i-th eigenvector, normalization is performed:

[0036]

[0037] Where Q i It is H i The attention weight of

[0038] The final feature representation output by the attention mechanism is:

[0039]

[0040] Where, Q={Q1,Q2,…,Q m};

[0041] The LSTM in Teacher Model 1 has a four-layer structure, with 64 hidden units in each layer acting as feature extractors. The features output by the LSTM are then fed into the attention mechanism module. The features output by the attention mechanism are then fed into a fully connected layer for regression processing, resulting in the final prediction result representing the lifespan label. This completes the construction of Teacher Model 1.

[0042] Temporal Convolutional Network (TCN) is a deep learning model specifically designed for processing time series data. It automatically extracts local features from time series by performing convolution operations on the time dimension. Bidirectional Long Short-Term Memory (BiLSTM) is a further development of LSTM. Its network structure is as follows: Figure 2 shown.

[0043] Analyzing the architectural features of BiLSTM, we can find that BiLSTM consists of two LSTM units with the same structure but opposite directions (LSTM forward hidden sequence and LSTM backward hidden sequence); for any time step t, if a small batch of input data is X t , the forward and reverse hidden states of this time step are and And let the hidden layer activation function be φ, then the calculation formula for state update is expressed as:

[0044]

[0045] Where, and is the weight, and is the bias; next, and Perform the splicing operation to obtain the comprehensive hidden state H representing the output layer t For a deep BiLSTM with multiple hidden layers, the concatenated feature information is passed as the input to the next bidirectional layer. After multiple layers of feature extraction and conversion, the network finally calculates the model’s result through the output layer, namely:

[0046] O t =H t W hq +b q

[0047] Where W hq is the weight matrix, b q is the bias vector.

[0048] TCN can capture the local features and global degradation trends of sensor signals through a layered dilated causal convolutional structure, and its receptive field expands exponentially with the depth of the network. BiLSTM can effectively model the bidirectional correlation of degradation trajectories caused by fluctuations in the working conditions of rotating machinery through forward and reverse dual-channel processing. This paper combines the advantages of TCN and BiLSTM to construct a teacher model 2, as shown in the framework diagram. Figure 3 The network structure parameter design is shown in Table 1.

[0049] The process of teacher model 2 is as follows: the input data first uses the TCN module to extract multi-scale time series features, then the BiLSTM module mines the bidirectional time series dependencies, and finally the output of the life prediction task is completed through the fully connected layer. Figure 3 In the figure, circles of different colors in the TCN module represent convolution calculation units in the residual block. Each layer of circles corresponds to the convolution operation in the residual block. By sliding and calculating the convolution kernel, the local features of the input data are extracted. The distribution of green, white, and blue circles reflects the feature transformation process of multi-layer convolution. Multi-scale temporal feature capture is achieved through parameters such as the convolution kernel size (2) and the expansion factor ([1,2,4], which controls the convolution receptive field). The dotted box is an overall TCN module. After being processed by multiple layers of residual blocks and convolution layers, the output is a single feature sequence that integrates multi-scale features. The feature sequence output by the TCN serves as the input of the BiLSTM module, and the BiLSTM module is a multi-layer structure (the number of hidden layers is 3). The output of the previous layer of BiLSTM serves as the input of the next layer, and the features are transmitted layer by layer to deepen the features. The multi-layer BiLSTM is connected sequentially, and each layer further mines the bidirectional temporal information based on the output of the previous layer. Finally, the output of the last layer of BiLSTM will be spliced ​​into a bidirectional vector as the input of the fully connected layer. The fully connected layer receives the comprehensive features output by the BiLSTM module, and through the "output dimension 64→32→1", first maps the features to 64 dimensions, then to 32 dimensions, and finally compresses them to 1 dimension, completing the conversion from feature space to the output of the life prediction task.

[0050] Table 1 Network structure parameters of teacher model 2

[0051]

[0052]

[0053] Transformer’s positional encoding and multi-head attention mechanism break through the sequence length limitation of RNN-type models, process and understand data from a more macro perspective, and provide new perspectives and knowledge to the student model. This paper builds the teacher model 3 based on the encoder part of Transformer, and its framework is as follows: Figure 4 shown.

[0054] Since the input data is temporal, Transformer cannot directly capture the order information of the input data. Therefore, position encoding is used to add the position information of the input data. The specific formula of position encoding is:

[0055]

[0056] Where t represents the time step in the current window, 2i represents an even-numbered bit vector, d represents the dimension of the input data, and 2i+1 represents an odd-numbered bit vector;

[0057] In teacher model 3, the calculation formula of layer normalization is as follows:

[0058]

[0059] Where, The result after layer normalization, x i,j represents the jth element of the i-th time step vector, μ n represents the mean of all elements in the vector at this time step, represents the variance of all elements in the vector at this time step, and ε is a very small constant (usually 10 -5 );

[0060] The feedforward neural network is designed as two linear layers connected by a ReLU activation function. Its function is to first map the output of the multi-head self-attention mechanism and the residual connection to a high-dimensional space, and then map more detailed information in the high-dimensional space to a low-dimensional space through the linear layer to mine deeper features. Finally, the output features are input into the fully connected layer for regression processing.

[0061] Table 2 Network structure parameters of teacher model 3

[0062]

[0063]

[0064] Convolutional Neural Networks (CNNs) automatically and efficiently extract local features from data by performing sliding convolutions on the data using convolutional kernels in convolutional layers. Their core advantage lies in their weight sharing mechanism, whereby the same convolution kernel shares parameters across different spatial locations, significantly reducing the number of network parameters. This lightweight feature is particularly effective when processing high-dimensional time series data, accelerating training convergence while also enabling low-power operation during inference.

[0065] Based on the above advantages, the present invention selects convolutional neural network as the basic architecture of the student model. This choice not only meets the requirements of lightweight model deployment on the device side, but also effectively retains the key information in the degradation signal of rotating machinery through its powerful local feature extraction capability. The student model includes three parallel one-dimensional convolution paths. The three convolution paths have the same structure. Each convolution path includes a convolution layer, an activation function, a convolution layer, an activation function, and a pooling layer connected in sequence. In the feature learning process, the three paths operate in parallel and independently of each other, respectively extracting information from data of short, medium and long time scales, ensuring the integrity of the feature learning process in all directions. The structure of the student model is as follows: Figure 5 The outputs of the three convolutional paths are fused and flattened before being fed into the fully connected layer for processing.

[0066] The parameters and convolutional layer configurations within each convolutional path differ. Taking the first branch as an example: the first convolutional layer is a 1D CNN (2×2×1): 1D convolution is used to initially extract local features from the input. "2×2×1" indicates a kernel size of 2, a stride of 2, and a padding of 1. The first activation function uses LeakyReLU, introducing nonlinear activation to enable the model to learn complex patterns. Compared to ReLU, LeakyReLU preserves small gradients on the negative semi-axis, mitigating vanishing gradients. The second convolutional layer is a 1D CNN (2×2×1): It further extracts more abstract local features from the first layer output, deepening the feature transformation. The second activation function uses LeakyReLU, introducing nonlinearity again to enhance feature representation. The pooling layer uses max pooling: It downsamples the convolution-activated features to retain key features while compressing the data dimensionality and reducing computational effort. Max pooling takes the maximum value within the window to highlight significant features. By varying the kernel size and number of channels, multi-scale feature extraction is achieved. The pooled outputs of the three parallel branches are concatenated to integrate feature information at different scales, allowing the model to capture both local details and long-range dependencies. A flattening operation flattens the fused multidimensional features into a one-dimensional vector, suitable for the input requirements of the subsequent fully connected layer. The fully connected layer receives this flattened one-dimensional feature vector and maps it to the output space through a linear transformation, completing the lifespan prediction task.

[0067] Three teacher models are trained by lifespan label supervision. In this embodiment, the three teacher models all use mean square loss as the loss function for model training. The process is as follows: Figure 6 As shown, the specific training process includes the following steps:

[0068] (2-1) Set the loss accuracy L * and the maximum number of iterations N max , and initialize the number of iteration steps N = 1;

[0069] (2-2) During the current iterative training process, the total loss L is calculated based on the training samples and the teacher model parameters of the current iterative step. 总 , and backpropagate based on the total loss to update the parameters of the teacher model;

[0070] (2-3) Determine L 总 ≤L * Or N≥N max Is it true? If so, the training ends and the parameters of the three teacher models are fixed; otherwise, let N=N+1 and return to step (2-2).

[0071] (3) After fixing the parameters of the three teacher models, the student model is trained with knowledge distillation using an adaptive meta-weight strategy gradient learning algorithm to achieve joint optimization of prediction accuracy and adaptation performance.

[0072] In order to transfer knowledge from the intermediate feature layer to ensure the efficiency of knowledge transfer, the present invention introduces a probability-based knowledge transfer method to match the data probability distribution of the teacher model and the student model in the feature space. Traditional feature extraction methods directly match the actual feature representations of the teacher model and the student model for specific samples. Unlike traditional methods, the probability-based knowledge transfer method transfers the sample probability distribution of the teacher model in the feature space to the student model; by minimizing the difference between the feature probability distributions of the teacher model and the student model, the student model is able to create a feature space geometry similar to that of the teacher model, thereby learning the feature knowledge of the teacher model.

[0073] First, the composite loss function of the multi-teacher meta-weight knowledge distillation network is calculated, including KL loss, soft loss and hard loss;

[0074] KL divergence is often used to measure the difference between two probability distributions and . Therefore, this paper uses KL divergence as KL loss. The specific calculation formula is:

[0075]

[0076] Where n b represents the number of samples, p j|i and q j|i The teacher model and the student model are respectively i and x j The conditional probability distribution on , the calculation formulas are:

[0077]

[0078] Where, and Respectively represent the sample x i and x j The flattened feature vector of the teacher model as input, and They represent the flattened feature vectors of the student model, cos() represents the biased cosine similarity, and the calculation formula is:

[0079]

[0080] Where ||·||2 represents the L2 norm. Compared with the general cosine similarity calculation, biased cosine similarity makes the result range become [0,1] to ensure that the conditional probability distribution is positive.

[0081] Similar to the concept of classification problems, we define a soft label and a hard label, where the soft label is the predicted value of the teacher model and the hard label is the actual lifespan. The training goal of the student model is to minimize the gap with the teacher model and the deviation from the actual value, thus obtaining the soft loss and hard loss. The specific calculation formula is:

[0082]

[0083] Where, represents the predicted value of the teacher model for the sample, represents the predicted value of the sample by the student model, and y represents the true value of the sample;

[0084] Finally, the composite loss function is expressed as:

[0085]

[0086] Where, α i , β i and γ are hyperparameters used to measure the weights of the three loss functions; the composite loss function integrates KL loss, soft loss, and hard loss, thereby ensuring the knowledge transfer efficiency of the distillation network.

[0087] In order to achieve adaptive optimization of the student model during the training process and enable it to more effectively integrate the knowledge features of multiple teacher models, the present invention sets the hyperparameters in the composite loss function as a set of adaptively updated meta-weights. This parameter optimization mechanism based on the meta-learning strategy can not only greatly reduce the workload of manual parameter adjustment, but also significantly improve the prediction accuracy of the model. All hyperparameters of the composite loss function in the multi-teacher meta-weight knowledge distillation network are composed of a hyperparameter set ψ = [α1, α2, α3, β1, β2, β3, γ]; in order to ensure that the meta-weights are updated by gradient descent, they need to be converted into continuous and differentiable functions; specifically, the softmax function is used to perform continuous relaxation operations on the hyperparameter set ψ. After the softmax function takes the value, the value range of the hyperparameter changes from discrete to continuous, and has the property of being differentiable, so that all possible values ​​in the continuous domain can be obtained with the help of gradient descent. The calculation formula is as follows:

[0088]

[0089] All weights and biases of the student model are combined into a parameter set θ = [W, b]. Therefore, the prediction performance of the student model depends on the parameter set θ, and the distillation performance depends on the hyperparameter set ψ. Furthermore, the goal of the adaptive meta-weight policy gradient learning algorithm is to search for the optimal hyperparameter set while ensuring the prediction performance of the student model.

[0090] The objective of the adaptive meta-weight policy gradient learning algorithm is defined as a two-level optimization problem. During the optimization process, the parameter set θ is used as the variables to be optimized in the inner loop, and the hyperparameter set ψ is used as the variables to be optimized in the outer loop. The inner loop first updates the parameters of the student model, and then the outer loop updates the hyperparameters of the composite loss function, with the goal of minimizing the composite loss function.

[0091] The present invention defines the goal of the adaptive meta-weight policy gradient learning algorithm as a two-level optimization problem, that is, the final result is affected not only by the hyperparameter set, but also by the parameter set of the student model, such as Figure 7 As shown. The two-level optimization problem is then defined as:

[0092]

[0093] Where, L train and L val Represent the composite loss functions obtained on the training set and the validation set respectively;

[0094] Since the update of the outer layer parameters usually requires a secondary calculation of the gradient based on the update of the inner layer parameters, the double-layer optimization process significantly increases the computing time and resource consumption. Therefore, a first-order approximation is made for the update of the outer layer parameters, and the calculation formula is:

[0095]

[0096] Where η is the learning rate in the inner optimization; the core idea of ​​the first-order approximation formula is to update the student model parameter θ once to achieve * approximation of

[0097] set up According to the chain rule, we get:

[0098]

[0099] Indicates the gradient of Ψ in the loss function;

[0100] Here, the finite difference formula is used to approximate the quadratic gradient of the second term:

[0101]

[0102] Where, ξ is a scalar whose value is:

[0103]

[0104] In summary, the student model uses the training set to update the parameter set θ, and its update formula is:

[0105]

[0106] Use the validation set to update the hyperparameter set ψ, and its update formula is:

[0107]

[0108] Where λ is the learning rate in the outer layer optimization.

[0109] The training process is as follows Figure 8 As shown, the following steps are included:

[0110] (3-1) Initialize the hyperparameter set ψ and parameter set θ, and set the loss accuracy L * and the maximum number of iterations N max , and initialize the number of iteration steps N = 1;

[0111] (3-2) Calculate θ on the training set * , perform gradient update on θ; then update ψ on the validation set and update θ on the training set to complete one iteration;

[0112] (3-3) Determine L 总 ≤L * Or N≥N max Is it true? If so, the training ends and the optimal θ is obtained; the trained student model is saved for subsequent actual deployment; otherwise, let N=N+1 and return to step (3-2).

[0113] (4) The lightweight student model trained by distillation is directly deployed for bearing remaining life prediction; after inputting data, the quantitative prediction value of the bearing remaining life is output through time series feature extraction and degradation state inference.

[0114] The effects of the present invention are verified by specific experiments below.

[0115] (1) Dataset introduction: In this example, the XJTU-SY rolling bearing dataset of Xi'an Jiaotong University is used for prediction experiments. The detailed structure of the bearing accelerated life test bench is as follows: Figure 9 The relevant parameters of the test bearings are shown in Table 3. During the test, two PCB 352C33 accelerometers mounted vertically and horizontally on the bearing housing were used to collect vibration signals. The sampling frequency was 25.6 kHz, the sampling duration was 1.28 seconds, and the sampling interval was 60 seconds. Detailed information for each test bearing in the XJTU-SY bearing dataset is shown in Table 4. Taking Bearing 1_1 as an example, the total number of samples is 123. Since the sampling interval is 1 minute, the actual lifespan is 123 minutes.

[0116] Table 3 LDK UER204 bearing parameters

[0117] Parameter name Numerical Parameter name Numerical Inner ring raceway diameter 29.30mm Outer ring raceway diameter 39.80mm Bearing center diameter 34.55mm Ball diameter 7.92mm Number of balls 8 contact angle 0° Basic dynamic load rating 12.82KN Basic static load rating 6.65KN

[0118] Table 4 XJTU-SY bearing data set list

[0119]

[0120]

[0121] (2) Division of experimental tasks

[0122] The experimental task arrangement for bearings is shown in Table 5. The pre-training of the teacher model in MTMWKDN uses only the training data, that is, all the bearing data in the training and validation sets. The student model is trained using the training and validation sets divided according to Table 5. The task objective of all experiments is to predict the remaining life of the bearings in the test set.

[0123] Table 5 Case arrangement of bearing life prediction tasks

[0124] Task training set Validation set Test set A Bearing 1-1, bearing 1-2, bearing 1-3 Bearings 1-5 Bearings 1-4 B Bearing 1-1, bearing 1-2, bearing 1-3 Bearings 1-4 Bearings 1-5 C Bearing 2-1, bearing 2-2, bearing 2-3 Bearings 2-5 Bearings 2-4 D Bearing 2-1, bearing 2-2, bearing 2-3 Bearings 2-4 Bearings 2-5 E Bearing 3-1, bearing 3-2, bearing 3-3 Bearings 3-5 Bearing 3-4 F Bearing 3-1, bearing 3-2, bearing 3-3 Bearing 3-4 Bearings 3-5

[0125] (3) Data preprocessing and evaluation indicators

[0126] Before life prediction, time domain features are first extracted from the original horizontal vibration signal, including peak value, mean value, peak-to-peak value, variance, kurtosis, skewness, form factor, margin coefficient, pulse factor and peak factor. Then the time domain signal is converted to the frequency domain, and frequency domain features are extracted, including energy spectrum, spectrum mean, frequency center of gravity and root mean square frequency. The obtained time-frequency domain features are then normalized and used as health indicators that reflect the change in bearing degradation trend. The characteristic analysis based on the vibration signal shows that the vibration data in the horizontal direction contains more abundant effective information than the vertical direction. Therefore, this section only uses the horizontal vibration signal for experimental verification. The health indicators of 10 time steps are selected as input data, and the life ratio P of the 10th time step is used as label data. The calculation formula of the life ratio P is:

[0127]

[0128] Where RUL t is the remaining life of the bearing at time t, and RUL0 is the actual full life of the bearing. For the first nine time steps, a normal distribution is used to fill the data to ensure the integrity of the life.

[0129] In order to quantitatively compare with other methods, this experiment uses five evaluation indicators: root mean square error (RMSE), mean absolute error (MAE), number of parameters, computational complexity (FLOPs), and prediction time. In terms of evaluating prediction accuracy, root mean square error is a commonly used evaluation indicator, and its calculation formula is:

[0130]

[0131] Where n is the total number of samples, y i is the true remaining life value of the i-th sample, Represents the remaining life value predicted by the model.

[0132] The mean absolute error is also an important indicator for evaluating prediction accuracy. Its calculation formula is:

[0133]

[0134] Unlike RMSE, MAE does not excessively amplify the impact of outliers on the overall error due to the square operation, and can more robustly reflect the average deviation of the predicted value from the actual value.

[0135] (4) Remaining life prediction results

[0136] The basic hyperparameters during training are set as follows: batch size is 16, both the teacher model and the student model are trained for 300 epochs, the learning rate of the inner optimization of the student model is 0.001, and the learning rate of the outer optimization is 0.05. The prediction results of the student model (after training) for task AF are as follows: Figure 10 As shown in the figure, the blue line represents the actual remaining life ratio, and the red line represents the predicted remaining life ratio. In the figure, (a) is bearing 1-4; (b) is bearing 1-5; (c) is bearing 2-4; (d) is bearing 2-5; (e) is bearing 3-4; and (f) is bearing 3-5.

[0137] The evaluation indicators of the experimental results are shown in Table 6, where the number of parameters and FLOPs are the values ​​of inputting one test sample. Since the size of the input samples is the same, the number of parameters and the amount of calculation are also the same.

[0138] pass Figure 10 As can be seen from Table 6, the prediction results of the remaining life ratio of most bearings in the six tasks show a fluctuating trend around the true label value, which shows that the prediction results are relatively close to the true value, and preliminarily verifies the effectiveness of the proposed method.

[0139] Table 6 Evaluation indicators of experimental results

[0140]

[0141] (5) Comparative analysis

[0142] To further verify the effectiveness and superiority of the proposed method, this section sets up a comparative experiment. The comparative experiment selects three classic lifespan prediction models: GRU, LSTM, and BiLSTM, as well as the CAKD method and the traditional KD method (using the teacher model cluster and student model of the proposed method).

[0143] The results of the root mean square error (RMSE)-based prediction performance evaluation are shown in Table 7. The experimental results show that across the six tasks, the proposed method achieves minimum RMSE values ​​for most tasks, with a small number of RMSE values ​​approaching the minimum. Furthermore, the proposed method significantly outperforms the comparative method in terms of average RMSE and standard deviation across all tasks, demonstrating that the proposed method, based on MTMWKDN, achieves higher RUL prediction accuracy.

[0144] Table 7 Lifespan prediction results based on RMSE index

[0145] Prediction Methods A B C D E F average value Standard deviation GRU 0.255 0.204 0.193 0.215 0.272 0.178 0.220 0.037 LSTM 0.216 0.148 0.173 0.251 0.186 0.194 0.195 0.036 BiLSTM 0.117 0.098 0.087 0.092 0.124 0.077 0.099 0.018 CAKD 0.092 0.106 0.094 0.124 0.087 0.064 0.095 0.020 Traditional KD 0.076 0.091 0.073 0.095 0.082 0.101 0.086 0.011 Method of the present invention 0.068 0.079 0.072 0.085 0.073 0.068 0.074 0.007

[0146] The prediction performance evaluation results based on mean absolute error are shown in Table 8. Experimental data analysis shows that the proposed method has the smallest MAE values, as well as the mean and standard deviation of the overall task results, across all six prediction tasks. This result further demonstrates that the proposed method based on MTMWKDN has higher remaining useful life prediction accuracy.

[0147] Table 8 Lifespan prediction results based on MAE index

[0148] Prediction Methods A B C D E F average value Standard deviation GRU 0.226 0.134 0.146 0.152 0.175 0.134 0.161 0.035 LSTM 0.173 0.106 0.129 0.211 0.136 0.146 0.150 0.037 BiLSTM 0.089 0.069 0.065 0.081 0.094 0.072 0.078 0.012 CAKD 0.064 0.065 0.071 0.089 0.068 0.059 0.069 0.010 Traditional KD 0.061 0.072 0.063 0.076 0.068 0.081 0.070 0.008 Method of the present invention 0.055 0.061 0.058 0.066 0.064 0.055 0.060 0.005

[0149] Table 9 shows a comparison of prediction results based on the number of parameters, FLOPs, and average prediction time. This comparison demonstrates that the proposed method has a lower model complexity and consumes relatively less prediction time. This combined comparison demonstrates the effectiveness and superiority of the proposed method based on MTMWKDN in achieving lightweight models and high-precision predictions.

[0150] Table 9 Life prediction results based on the number of parameters, FLOPs and average prediction time

[0151]

[0152]

[0153] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A bearing remaining life prediction method based on a multi-teacher weighted knowledge distillation network is characterized by: The following steps are involved: (1) The multi-dimensional vibration time series signal collected by the original sensor is normalized and denoised, and then the signal is segmented to form a standardized data set, which is divided into a training set and a validation set; (2) Input the processed training set data into the multi-teacher meta-weight knowledge distillation network to train the teacher model; The multi-teacher meta-weighted knowledge distillation network includes three parallel teacher models and one student model. The three teacher models use LSTM-attention, TCN-BiLSTM, and Transformer models respectively. Teacher model 1 includes a sequential LSTM, an attention mechanism, and a fully connected layer. Teacher model 2 includes a sequentially connected temporal convolutional neural network module, a bidirectional long short-term memory network, and a fully connected layer. Teacher model 3 includes an embedding layer, a positional encoder, and three cascaded encoding modules. The positional encoder is used to add the position information of the input data. The feature vector output by the embedding layer is added to the position information output by the positional encoder element by element, and the added features are then input into the encoding module. The student model consists of three parallel one-dimensional convolution paths. The three convolution paths have the same structure. Each convolution path includes a convolution layer, an activation function, a convolution layer, an activation function, and a pooling layer connected in sequence. The outputs of the three convolution paths are fused and flattened before being output to the fully connected layer for processing. Three teacher models are trained using lifespan label supervision, and the parameters of the three teacher models are fixed after training. (3) After fixing the parameters of the three teacher models, the student model is trained with the adaptive meta-weight policy gradient learning algorithm for knowledge distillation; the trained student model is saved; (4) The student model trained by distillation is used to predict the remaining life of bearings.

2. The method for predicting remaining life of bearings based on a multi-teacher weighted knowledge distillation network according to claim 1 is characterized in that: The specific implementation method of step (3) is: calculating the composite loss function of the multi-teacher meta-weight knowledge distillation network, including KL loss, soft loss and hard loss; Using KL divergence as KL loss, the specific calculation formula is: Where n b represents the number of samples, p j|i and q j|i The teacher model and the student model are respectively i and x j The conditional probability distribution on ; Define the soft label as the predicted value of the teacher model, and the hard label as the true value of life span; the calculation formulas for soft loss and hard loss are: Where, represents the predicted value of the teacher model for the sample, represents the predicted value of the sample by the student model, and y represents the true value of the sample; Finally, the composite loss function is expressed as: Where, α i , β i and γ are hyperparameters, which are used to measure the weights of the three loss functions; The hyperparameters in the composite loss function are set to a set of adaptively updated meta-weights. All hyperparameters of the composite loss function in the multi-teacher meta-weight knowledge distillation network are combined into a hyperparameter set ψ = [α1, α2, α3, β1, β2, β3, γ]; all weights and biases of the student model are combined into a parameter set θ = [W, b]. The goal of the adaptive meta-weight policy gradient learning algorithm is to search for the optimal hyperparameter set while ensuring the predictive performance of the student model. The goal of the adaptive meta-weight policy gradient learning algorithm is defined as a two-level optimization problem. During the optimization process, the parameter set θ is used as the variable for the inner loop optimization, and the hyperparameter set ψ is used as the variable for the outer loop optimization. The inner loop first updates the parameters of the student model, and then the outer loop updates the hyperparameters of the composite loss function, with the goal of minimizing the composite loss function.

Citation Information

Cited By

  • Method and system for predicting residual life of aero-engine based on dual-channel feature interaction

    CN121117579A

  • Floating type storm power generation system control method based on knowledge distillation and zoning strategy

    CN121165594A

  • GEO spacecraft maneuver detection method considering sample imbalance

    CN121412774A

  • GEO spacecraft maneuver detection method considering sample imbalance

    CN121412774B

  • Voice emotion recognition method based on cross-granularity weight reuse and heterogeneous multitask

    CN121415815A