Bearing life prediction method based on digital twin and LSTM network

By combining digital twins and LSTM networks, simulated data is generated and time series features are extracted, solving the problem of adaptability to complex operating conditions in bearing life prediction in traditional methods. This achieves more accurate bearing life prediction and promotes intelligent maintenance and health management of elevators.

CN120929784APending Publication Date: 2025-11-11CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510833427.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional bearing life prediction methods cannot fully consider the impact of complex working conditions and dynamic environments. Furthermore, machine learning and deep learning methods that rely on large amounts of real data have limited generalization ability in practical applications, making it difficult to accurately predict the remaining service life of bearings.

Method used

A bearing life prediction method based on digital twins and LSTM networks (TwinLSTM-TPA) is adopted. It generates simulated data by constructing a digital twin model of the physical system, extracts time series features by combining 1D-CNN, introduces a time series pattern attention mechanism, and uses a data hybrid training strategy to enhance the model's predictive ability.

Benefits of technology

It provides more accurate bearing life prediction under complex working conditions, improves the model's prediction accuracy and generalization ability, and supports intelligent maintenance and health management of elevators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929784A_ABST
    Figure CN120929784A_ABST
Patent Text Reader

Abstract

A bearing life prediction method based on a digital twinning and LSTM network comprises the steps that firstly, a digital twinning model of a rolling bearing life cycle is established, and a bearing degradation stage is divided into a stable stage, a defect starting stage, a defect expansion stage and a damage stage; actual data and twinborn data are preprocessed, then a self-defined loss function is introduced to balance the proportion of the twinborn data and the real data, and finally the mixed data are input into a TwinLSTM-TPA model; and extracting a time feature matrix from a hidden state output by the LSTM by using 1D-CNN, enhancing the memory and recognition capability of the model on key time features through a TPA mechanism, and finally further mapping the extracted comprehensive features into a predicted value of the bearing life. According to the invention, the stability and adaptability of the prediction result are enhanced. And more accurate bearing life prediction can be provided under complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bearing life prediction and relates to a bearing life prediction method based on digital twin and LSTM network. Background Technology

[0002] The widespread use of elevators relies on ensuring a high level of safety during operation, especially since the elevator traction machine is one of the core components of the elevator system. Bearings, as crucial components of rotating machinery, directly impact the overall performance and safety of the equipment through their operating condition. However, during long-term operation, bearings are subjected to various factors such as load, vibration, temperature, and lubrication conditions, gradually leading to wear and failure. To reduce equipment maintenance costs and minimize losses caused by sudden failures, bearing life prediction has become an important research direction in the field of equipment condition monitoring and health management.

[0003] Traditional bearing life prediction methods primarily rely on empirical formulas and statistical analyses, such as those based on fatigue life theory and probabilistic statistical methods. These methods typically assume a fixed bearing failure model and cannot adequately account for the impact of complex operating conditions and dynamic environments on bearing life. Furthermore, thresholding methods and time-frequency domain analysis based on sensor data such as vibration signals and temperature, while capable of detecting bearing conditions, still lack the ability to model long-term trends and cannot accurately predict the remaining service life of bearings. With the rapid development of sensor technology and artificial intelligence, data-driven machine learning and deep learning methods have gradually become a research hotspot in bearing life prediction. These methods utilize large amounts of bearing operating data to construct data models and uncover bearing degradation patterns, achieving high-precision life prediction. Traditional machine learning and deep learning methods often rely on large amounts of real data for training, but in practical industrial applications, complete bearing life data is difficult to obtain, limiting the model's generalization ability. Digital twin technology provides a new approach to bearing life prediction. By constructing bearing data fusion models, rich simulated data under different operating conditions can be generated, supplementing the deficiencies of real data. Summary of the Invention

[0004] To overcome the shortcomings of existing technologies, this invention proposes a bearing life prediction method based on digital twins and LSTM (long short-term memory) networks (TwinLSTM-TPA). First, a digital twin model of the physical system is constructed, generating rich simulation data under different operating conditions, thus providing more diverse input data for model training. In terms of model architecture, TwinLSTM-TPA employs LSTM to capture long-term dependencies in the bearing degradation process and introduces a temporal pattern attention mechanism to enhance the identification of key temporal features. Furthermore, 1D-CNN (One-Dimensional Convolutional Neural Network) is used to extract key features from time-series data, improving the model's ability to perceive dynamic changes over time. To further improve prediction accuracy and generalization ability, this invention adopts a data hybrid training strategy, uniformly processing the data generated by the digital twin and actual operating data, and concatenating them into a larger and more diverse dataset, thereby enhancing the model's ability to identify different degradation stages and improving the stability and adaptability of the prediction results. The TwinLSTM-TPA model can provide more accurate bearing life prediction under complex operating conditions, providing important technical support for intelligent maintenance and health management of elevators.

[0005] The technical solution adopted by this invention to solve its technical problem is:

[0006] A bearing life prediction method based on digital twin and LSTM network includes the following steps:

[0007] The first step is to establish a digital twin model of the rolling bearing life cycle. This model can flexibly adjust key parameters to adapt to changes in bearing performance under different operating conditions. By inputting parameters, the model can simulate the degradation path of the bearing under actual service conditions. The bearing degradation stages are divided into the stable stage, the defect initiation stage, the defect propagation stage, and the failure stage.

[0008] The second step is to preprocess the real data and twin data to ensure that the twin data and the real data have the same scale, so as to avoid the impact of data distribution differences on model training. Then, a custom loss function is introduced to balance the ratio of twin data and real data. Finally, the mixed data is input into the TwinLSTM-TPA model.

[0009] The third step involves using 1D-CNN to extract the temporal feature matrix from the hidden state output by LSTM, and enhancing the model's ability to remember and recognize key temporal features through the TPA (Temporal Pattern Attention mechanism). Finally, the extracted comprehensive features are further mapped to the predicted value of bearing life.

[0010] 1D-CNN is a deep learning model that is typically used to process one-dimensional sequence data. Its main feature is that it extracts features from the input data through one-dimensional convolutional kernels.

[0011] Furthermore, the process of the first step is as follows:

[0012] Step (1.1) Modeling the stable stage of the rolling bearing: The inner and outer rings of the bearing each have two degrees of freedom, corresponding to the horizontal displacement x, respectively. i x o and vertical displacement y i y o A unit resonator has only one degree of freedom, namely the vertical displacement y. r The equation of motion for the inner ring of the bearing is shown below:

[0013]

[0014] Where, m i For the inner ring quality, c i F is the damping coefficient. i For the external radial load, k i For stiffness, f x and f y These are the nonlinear contact forces applied in the horizontal and vertical directions, respectively. and These represent the velocities of the inner ring in the horizontal and vertical directions, respectively, where g is the acceleration due to gravity.

[0015] The unit resonator is a common physical system widely used in fields such as vibration, acoustics, and electronics. It typically refers to a system that can oscillate naturally at a specific frequency.

[0016] The equation of motion for the outer ring of the bearing is shown below:

[0017]

[0018] Where, m o For the outer ring mass, c o c is the damping coefficient. r k is the stiffness of a unit resonant cavity. o For stiffness, k ry is the damping coefficient between the bearing and the support. r This refers to the displacement of the bearing outer ring support structure. The velocity of the bearing outer ring support structure in the vertical direction;

[0019] The vibration equation of a unit resonator is shown below:

[0020]

[0021] Where, m r For a unit resonator mass, the nonlinear contact force is as follows:

[0022]

[0023] Where, δ j K represents the total deflection at the j-th roller position. j This indicates the contact stiffness between the roller and its trajectory. b Let θ be the number of rollers. i This indicates the angular position of the i-th roller;

[0024] Step (1.2) Modeling the initial stage of rolling bearing defects: During the use of the bearing, the initial stage of outer ring defects is usually caused by tiny surface dents or scratches. These initial defects are caused by improper installation, material defects, or external impacts. In this invention, the defects are simplified as an increase in the roughness of a local area. The formula for the total deflection of the bearing outer ring is as follows:

[0025] δ j =(x i -x o cosθ j +(y i -y o sinθ j +η+ζ+Δζ

[0026] Where η is the local deformation of the rolling element, ζ is the elastic deformation of the bearing outer ring, and Δζ is the increased roughness in the local area, defined as follows:

[0027]

[0028] Where, n j This represents the number of inner ring rotors, j is the rolling element number, and b is the width of the groove pressure. For cross-angle, Let θ be the initial angular position. d ω is the initial angular offset for rolling. c D is the angular velocity of the bearing, t is the time it takes for the rolling elements to rotate, and D is the angular velocity of the bearing. o This refers to the raceway diameter of the bearing's outer ring.

[0029] Step (1.3) Modeling the defect propagation stage of the rolling bearing: During the defect propagation stage, the roughness of the bearing raceway will further increase, and cracks will appear on the surface. The cracks will propagate from the trailing edge of the indentation. When the rolling element passes through the crack, it will release a certain amount of deformation. The total deflection of the overall deformation is shown below:

[0030] δ j =(x i -x o cosθ j +(y i -y o sinθ j +η+ζ1+Δd 1j

[0031] Where, δ j Let ζ1 represent the total deflection of the overall deformation, and Δd represent the surface roughness of the raceway during the defect propagation stage. 1j For the displacement excitation in this stage, when the rolling element passes through the defect, the displacement excitation in this stage is defined as follows:

[0032]

[0033] Where cd is the maximum displacement of the rolling element, as shown below:

[0034]

[0035] Where d is the diameter of the rolling element, and b1 is the chord length of the contact surface of the rolling element on both sides of the guide rail;

[0036] Step (1.4) Modeling the rolling bearing failure stage. In this stage, the width of the defect increases, causing the roll to sink to the bottom of the defect. The deflection in this stage is shown below:

[0037] δ j =(x i -x o cosθ j +(y i -y o sinθ j +η+ζ2+Δd 2j

[0038] Where ζ2 represents the displacement excitation in this stage, and Δd 2j This represents the displacement excitation generated during the failure stage. When the roller passes through the defect in this stage, the displacement excitation is represented as follows:

[0039]

[0040] Where H is the defect height, It is the circumferential cross-sectional angle of the roll at the bottom track of the defect. and These are the angles of roller rotation when entering and leaving the defect, respectively. and The definitions are as follows:

[0041]

[0042] Furthermore, the second step is as follows:

[0043] Step (2.1) Data Preprocessing: Standardization is a commonly used data preprocessing method, and its calculation formula is as follows:

[0044]

[0045] Where z is the standardized distance, x is the original data point, μ is the mean of the data, and σ is the standard deviation of the data. The processed dataset is divided into datasets using random sampling to ensure the balance of the training set and the test set on the twin data and the real data.

[0046] Step (2.2) balances the ratio of twin data to actual data by designing a custom loss function to fully utilize the complementary characteristics of the two types of data. Twin data is generated through a dynamic twin model of the bearing life cycle, possessing rich degradation patterns and high coverage, but it may have certain simulation deviations compared to actual data. While actual data closely reflects real operating conditions, its coverage may be insufficient due to limitations in acquisition conditions and data volume. Therefore, designing a custom loss function for both types of data is crucial for model optimization. The formula for calculating the custom loss function is shown below:

[0047]

[0048] Where N is the total number of samples in the dataset, y sim,i and y true,i These are the true values ​​of the i-th sample in the twin data and the actual data, respectively. and It is an indicator function that indicates whether the data is twin data or actual data: when the data source is twin data... Otherwise, the value is 0; when the data source is actual data. Otherwise, it is 0, w sim and w real These are the weighting coefficients for twin data and actual data, and MSE() represents the mean squared error function;

[0049] Step (2.3) divides the mixed data using a sliding window to ensure that each data sample has the same length. Each data sample has a length of 1000 and a dimension of 1. The divided data samples are used as a new unit signal and input into TwinLSTM-TPA for processing.

[0050] The sliding window method is a commonly used data processing technique, particularly suitable for handling continuous data such as time series and text data. Its core idea is to use a fixed-size window that slides across the data in increments of a certain size, extracting and processing a portion of the data covered by the window each time. The sliding window method effectively captures local features and reduces computational complexity.

[0051] Step (2.4) In the TwinLSTM-TPA network model, the data is first input into the LSTM network for learning and modeling of temporal features. LSTM effectively controls the transmission of information by introducing three gating mechanisms: the input gate, the forget gate, and the output gate. In this model, the number of hidden units, the dimension of the hidden state, and the dimension of the unit state are all 64. The input gate is responsible for determining the input x at the current time step. t and the hidden state h from the previous moment t-1 For the current cell state c t The impact;

[0052] Step (2.5) The input gate performs a linear transformation on the current input data x. t and the hidden state h from the previous moment t-1 With weight matrix W respectively i The product is multiplied and mapped to the dimension of the hidden unit to obtain the feature representation. Subsequently, the bias term b of the input gate is added. i Finally, the linear combination result of the input gates is obtained, and its calculation formula is as follows:

[0053] i t =sig(W i ·[h t-1 ,x t ]+b i )

[0054] Among them, W i It is the weight matrix, b i It is the input gate bias, and sig is the sigmoid activation function;

[0055] Linear transformations primarily refer to operations that map an input vector to an output vector, adhering to the principles of addition and homogeneity. In linear algebra, linear transformations are typically represented by matrix multiplication.

[0056] Step (2.6) Activation function processing: The result of step (2.5) is processed by the sigmoid activation function to obtain the input gate output i. t The activation function is calculated as follows:

[0057]

[0058] The value of each dimension represents the degree of influence of the current time step input on updating the cell state. When the value is close to 1, it means that the current input has a significant impact on the cell state; when the value is close to 0, it means that the current input has a negligible impact on the cell state.

[0059] The sigmoid activation function is one of the common non-linear activation functions in neural networks, widely used in early neural network models, especially in the output layer of binary classification problems. It maps the input signal to a fixed range, typically between 0 and 1.

[0060] Step (2.7) Calculate the forget gate output: The forget gate determines the cell state C of the previous time step. t-1 The formula for determining which information should be discarded and which should be retained is as follows:

[0061] o t =sig(W o ·[h t-1 ,x t ]+b o )

[0062] Input data x t and the hidden state h from the previous moment t-1 It will be compared with the weight matrix W respectively. o Perform a linear transformation to obtain a new vector, and then change the bias term b. o This result is added to the linear combination and then mapped to the interval [0,1] using the Sigmoid activation function, thus obtaining the forget gate value o. t ;

[0063] Among them, W o It is the output gate weight matrix, b o It is the output gate bias, o t It is the value of the forget gate;

[0064] Step (2.8) Calculate the candidate cell state: Candidate cell state The input data is x t The latent information, combined with the hidden state from the previous time step, is calculated using the following formula:

[0065]

[0066] Input data x at the current time step t and the hidden state h from the previous moment t-1 After weight matrix W xc and W hc After a linear transformation, a new latent information vector is generated. Then, the combined information of the current input and the historical state is mapped to the range [-1, 1] through the tanh activation function.

[0067] in, W represents the candidate memory cell state at the current time step. xc and W hc Let b be the weight matrix. c As the bias term, x t h is the input data for the current time step. t-1 This is the hidden state from the previous moment;

[0068] Tanh is an activation function commonly used in neural networks, with an output range of [-1, 1], and is suitable for many neural network architectures.

[0069] Step (2.9) Output hidden state sequence and cell state update: Output through input gate and forget gate, cell state C t The updated value includes historical information and current input information, and the calculation formula is as follows:

[0070]

[0071] h t =o t ⊙tanh(C t )

[0072] The output of the forget gate o t Compared with the previous cell state C t-1 Multiplying by the element determines how much information from the previous time step will be retained; the output i of the input gate... t This controls which information will be added to the current candidate unit state. In LSTM, by combining the outputs of these three gates, the model can flexibly choose to retain, forget, and output different information, thereby effectively capturing long-term dependencies in time series data. When processing time series data, the LSTM model generates a hidden state sequence h. t , representing the feature information of each time step in the sequence;

[0073] Among them, C t-1 ⊙ represents the cell state at the previous time step, and ⊙ represents the element-wise multiplication operation.

[0074] The process of the third step is as follows:

[0075] Step (3.1) Extracting the temporal feature matrix: The goal of 1D-CNN is to convolve the output hidden states of LSTM to extract local features of the time series, thereby enhancing the model's ability to learn from time series data. This process can be achieved by extracting the temporal feature matrix, as shown in the following formula:

[0076]

[0077] in, Let w represent the feature matrix of the i-th sample at the j-th time step, w be the length of the time series, * represent the convolution operation, and C be the feature matrix of the j-th sample. j,T-w+l Here, T represents the kernel size, and l represents the l-th position within the convolution window.

[0078] Convolution is a fundamental operation in signal processing and deep learning. It combines an input signal with a filter to generate an output signal. The purpose of convolution is to extract features or patterns from input data, and it is commonly used in image processing, speech processing, time series analysis, and other fields.

[0079] Step (3.2) Calculate the relevance weights: The process of calculating the relevance weights involves obtaining a set of weight coefficients by evaluating the correlation between the hidden state and the convolutional features. These weight coefficients are used to weight different features, thereby making the model focus more on important time periods or features. The formula for calculating the relevance weights is as follows:

[0080]

[0081] Next, the sigmoid activation function is used to map the correlation values ​​to a weight range. After obtaining the weights, a comprehensive feature vector v is obtained by weighted summation. t The calculation formula is as follows:

[0082]

[0083] in, is a scoring function that measures the similarity between the current hidden state and the features, where m is the total number of features extracted. It is the i-th feature, h t The current hidden state is given, and m is the total number of features extracted. It is the i-th feature, W a This is the attention weight matrix, used for learning matching. and h t The relationship between them. α i Let be the attention weight, representing the feature at time step i. The weight of the contribution. tThis is the comprehensive feature vector, representing the temporal feature representation after attention weighting;

[0084] Step (3.3) uses the TPA mechanism to further map the extracted comprehensive features into predicted bearing life values. The comprehensive feature vector processed by the time pattern attention mechanism is input into the Dense Layer and mapped to the predicted value y through linear transformation. pred ;

[0085] The Dense Layer is a fundamental layer in a neural network, connecting all neurons in the previous layer to every neuron in the current layer. Each neuron receives input from all neurons in the previous layer, performs a weighted summation, and then performs an activation operation to produce an output.

[0086] Preferably, the process of step (3.3) is as follows:

[0087] Step (3.3.1) Generate feature weights: The TPA mechanism calculates the hidden state h′ and the feature map c extracted by convolution. k The correlation between them is calculated using the following formula:

[0088]

[0089] Where, α k For the k-th time step, the extracted feature map c k The weight of the contribution. `exp()` is the exponential function used to ensure that all attention weights are positive, and `score(h′)` is the weight of the contribution. t ,c k The correlation between the hidden state and the k-th feature map is denoted as follows:

[0090] score(h′ t ,c k )=h′ t c k

[0091] This score measures the similarity between the current hidden state and the convolutional feature map; a high score indicates that the current hidden state is highly correlated with the convolutional feature map, while a low score indicates that the correlation between them is weak.

[0092] Among them, the TPA mechanism is an attention mechanism specifically used for extracting and modeling key time-point features in time series data. Its core idea is to dynamically allocate weights based on the current hidden state and the extracted features, highlighting the importance of key time patterns in the sequence, thereby enhancing the model's ability to remember and recognize time series data.

[0093] Step (3.3.2) generates the comprehensive feature vector: using the feature weights calculated in step (3.3.1) and the convolutional feature map c k The weighted summation yields the comprehensive eigenvector Z, calculated using the following formula:

[0094]

[0095] The comprehensive feature vector Z contains the features extracted by convolution, and the TPA mechanism enhances the memory and recognition of important features;

[0096] Step (3.3.3) yields the predicted value y. pred The Dense Layer receives the comprehensive feature vector from the TPA mechanism. Through this layer, the model linearly transforms these features with the learned weights. The Dense Layer then uses a fully connected neural network layer to compute the final predicted value y. pred The calculation formula is as follows:

[0097] y pred =W1·Z+b1

[0098] The output of the linear transformation is the predicted bearing life, which is usually a continuous value.

[0099] Where W1 is the weight matrix of the Dense Layer, and b1 is the bias term of the Dense Layer.

[0100] The DenseLayer, also known as a fully connected layer, is one of the most common layers in neural networks. In deep learning models, the DenseLayer connects all nodes from the previous layer to every neuron in the current layer, with each neuron having weights and biases.

[0101] Compared with existing technologies, the beneficial effects of this invention are as follows: It proposes a bearing life prediction method that integrates digital twins and LSTM networks. The TwinLSTM-TPA model can effectively capture long-term dependencies in the bearing degradation process and enhances the ability to identify key temporal features by introducing a temporal pattern attention mechanism. Furthermore, by utilizing 1D-CNN to extract key features from time-series data, the model's ability to perceive dynamic changes over time is further improved. By adopting a data hybrid training strategy, this invention combines simulated data generated by digital twins with actual operating data to construct a larger and more diverse dataset, improving the model's ability to identify different degradation stages and enhancing the stability and adaptability of the prediction results. TwinLSTM-TPA can provide more accurate bearing life predictions under complex operating conditions, greatly promoting intelligent elevator maintenance and health management, and providing important technical support for improving the stability and safety of equipment operation. Attached Figure Description

[0102] Figure 1 A block diagram illustrating the principle of a bearing life prediction method based on digital twins and LSTM networks is shown. Detailed Implementation

[0103] The present invention will now be further described with reference to the accompanying drawings.

[0104] Reference Figure 1 A bearing life prediction method based on digital twins and LSTM networks includes the following steps:

[0105] The first step is to establish a digital twin model of the rolling bearing life cycle. This model can flexibly adjust key parameters to adapt to changes in bearing performance under different operating conditions and can simulate the degradation path of the bearing under actual service conditions.

[0106] The first step of this invention is to establish a digital twin model of the rolling bearing's life cycle. This model can flexibly adjust key parameters to adapt to changes in bearing performance under different operating conditions. By inputting parameters (such as load and speed), the model can simulate the bearing's degradation path under actual service conditions. The bearing degradation stages are divided into a stable stage, a defect initiation stage, a defect propagation stage, and a failure stage, as follows:

[0107] Step (1.1) Modeling the stable stage of the rolling bearing: The inner and outer rings of the bearing each have two degrees of freedom, corresponding to the horizontal displacement x, respectively. i x o and vertical displacement y i y o A unit resonator has only one degree of freedom, namely the vertical displacement y. r The equation of motion for the inner ring of the bearing is shown below:

[0108]

[0109] Where, m i For the inner ring quality, c i F is the damping coefficient. i For the external radial load, k i For stiffness, f x and f y These are the nonlinear contact forces applied in the horizontal and vertical directions, respectively. and These represent the velocities of the inner ring in the horizontal and vertical directions, respectively, where g is the acceleration due to gravity.

[0110] The unit resonator is a common physical system widely used in fields such as vibration, acoustics, and electronics. It typically refers to a system that can oscillate naturally at a specific frequency.

[0111] The equation of motion for the outer ring of the bearing is shown below:

[0112]

[0113] Where, m o For the outer ring mass, c o c is the damping coefficient. r k is the stiffness of a unit resonant cavity. o For stiffness, k r y is the damping coefficient between the bearing and the support. r This refers to the displacement of the bearing outer ring support structure. The velocity of the bearing outer ring support structure in the vertical direction;

[0114] The vibration equation of a unit resonator is shown below:

[0115]

[0116] Where, m r The unit resonator mass is given. The nonlinear contact force is shown below:

[0117]

[0118] Where, δ j K represents the total deflection at the j-th roller position. j This indicates the contact stiffness between the roller and its trajectory. b θ is the number of rollers. i This indicates the angular position of the i-th roller;

[0119] Step (1.2) Modeling the initial stage of rolling bearing defects: During the use of the bearing, the initial stage of outer ring defects is usually caused by tiny surface dents or scratches. These initial defects are caused by improper installation, material defects, or external impacts. In this invention, the defects are simplified as an increase in the roughness of a local area. The formula for the total deflection of the bearing outer ring is as follows:

[0120] δ j =(x i -x o cosθ j +(y i -y o sinθ j +η+ζ+Δζ

[0121] Where η is the local deformation of the rolling element, ζ is the elastic deformation of the bearing outer ring, and Δζ is the increased roughness in the local area, defined as follows:

[0122]

[0123] Where, n j This represents the number of inner ring rotors, j is the rolling element number, and b is the width of the groove pressure. For cross-angle, Let θ be the initial angular position. d ω is the initial angular offset for rolling. c D is the angular velocity of the bearing, t is the time it takes for the rolling elements to rotate, and D is the angular velocity of the bearing. o This refers to the raceway diameter of the bearing's outer ring.

[0124] Step (1.3) Modeling the defect propagation stage of the rolling bearing: During the defect propagation stage, the roughness of the bearing raceway will further increase, and cracks will appear on the surface. Cracks usually start to propagate from the trailing edge of the indentation. When the rolling element passes through the crack, it will release a certain amount of deformation. The total deflection of the overall deformation is shown below:

[0125] δ j =(x i -x o cosθ j +(y i -y o sinθ j +η+ζ1+Δd 1j

[0126] Where, δ j Let ζ1 represent the total deflection of the overall deformation, and Δd represent the surface roughness of the raceway during the defect propagation stage. 1j For the displacement excitation in this stage, when the rolling element passes through the defect, the displacement excitation in this stage is defined as follows:

[0127]

[0128] Where cd is the maximum displacement of the rolling element, as shown below:

[0129]

[0130] Where d is the diameter of the rolling element, and b1 is the chord length of the contact surface of the rolling element on both sides of the guide rail;

[0131] Step (1.4) Modeling the rolling bearing failure stage: In this stage, the width of the defect increases, causing the roll to sink to the bottom of the defect. The deflection in this stage is shown below:

[0132] δ j =(x i -x o cosθ j +(y i -y o sinθ j+η+ζ2+Δd 2j

[0133] Where ζ2 represents the displacement excitation in this stage, and Δd 2j This represents the displacement excitation generated during the failure stage. When the roller passes through the defect in this stage, the displacement excitation is represented as follows:

[0134]

[0135] Where H is the defect height, It is the circumferential cross-sectional angle of the roll at the bottom track of the defect. and These are the angles of roller rotation when entering and leaving the defect, respectively. and The definitions are as follows:

[0136]

[0137] The second step involves preprocessing the real and twin data to ensure they have the same scale, thus avoiding the impact of data distribution differences on model training. Next, a custom loss function is introduced to balance the ratio of twin to real data. Finally, the mixed data is input into the TwinLSTM-TPA model. The process is as follows:

[0138] Step (2.1) Data Preprocessing: Standardization is a commonly used data preprocessing method, and its calculation formula is as follows:

[0139]

[0140] Where z is the standardized distance, x is the original data point, μ is the mean of the data, and σ is the standard deviation of the data. The processed dataset is divided into two parts using random sampling. This ensures a balance between the training and test sets on the twin dataset and the real data.

[0141] Step (2.2) balances the ratio of twin data to actual data. This invention designs a custom loss function to fully utilize the complementary characteristics of the two types of data. Twin data is generated through a dynamic twin model of the bearing life cycle, possessing rich degradation modes and high coverage, but it may have certain simulation deviations compared to actual data. While actual data closely reflects real operating conditions, its coverage may be insufficient due to limitations in acquisition conditions and data volume. Therefore, designing a custom loss function for both types of data is crucial for model optimization. The formula for calculating the custom loss function is shown below:

[0142]

[0143] Where N is the total number of samples in the dataset, y sim,i and y true,i These are the true values ​​of the i-th sample in the twin data and the actual data, respectively. and It is an indicator function that indicates whether the data is twin data or actual data: when the data source is twin data... Otherwise, the value is 0; when the data source is actual data. Otherwise, it is 0. sim and w real These are the weighting coefficients for twin data and actual data, and MSE() represents the mean squared error function;

[0144] Step (2.3) uses a sliding window to segment the time series data. The sliding window length is set to 1000 and the step size is 100. Each time, a data sequence of length 10000 is processed. After the sliding window, it is ensured that each data sample has the same length. The length of each data sample is 1000 and the dimension is 1. The segmented data sample will be used as a new unit signal and input into TwinLSTM-TPA for processing.

[0145] The sliding window method is a commonly used data processing technique, particularly suitable for handling continuous data such as time series and text data. Its core idea is to use a fixed-size window that slides across the data in increments of a certain size, extracting and processing a portion of the data covered by the window each time. The sliding window method effectively captures local features and reduces computational complexity.

[0146] In step (2.4) of the TwinLSTM-TPA network model, the data is first input into the LSTM network for learning and modeling of temporal features. LSTM effectively controls the transmission of information by introducing three gating mechanisms: the input gate, the forget gate, and the output gate. In this model, the number of hidden units, the dimension of the hidden state, and the dimension of the unit state are all 64. The input gate is responsible for determining the input x at the current time step. t and the hidden state h from the previous moment t-1 For the current cell state c t The impact;

[0147] Step (2.5) The input gate performs a linear transformation on the current input data x. t and the hidden state h from the previous moment t-1 With weight matrix W respectively i The product is multiplied and mapped to the dimension of the hidden unit, resulting in a feature of size 32*64. Then, the bias term b from the input gate is added. i Finally, the linear combination result of the input gates is obtained, and its calculation formula is as follows:

[0148] i t =sig(W i ·[h t-1 ,x t ]+b i )

[0149] Among them, W i It is the weight matrix, b i It is the input gate bias, and sig is the sigmoid activation function;

[0150] Linear transformations primarily refer to operations that map an input vector to an output vector, adhering to the principles of addition and homogeneity. In linear algebra, linear transformations are typically represented by matrix multiplication.

[0151] Step (2.6) Activation function processing: The result of step (2.5) is processed by the sigmoid activation function to obtain the input gate output i. t The activation function is calculated as follows:

[0152]

[0153] The value of each dimension represents the degree of influence of the current time step input on updating the cell state. When the value is close to 1, it means that the current input has a significant impact on the cell state; when the value is close to 0, it means that the current input has a negligible impact on the cell state.

[0154] The sigmoid activation function is one of the common non-linear activation functions in neural networks, widely used in early neural network models, especially in the output layer of binary classification problems. It maps the input signal to a fixed range, typically between 0 and 1.

[0155] Step (2.7) Calculate the forget gate output: The forget gate determines the cell state C of the previous time step. t-1 The formula for determining which information should be discarded and which should be retained is as follows:

[0156] o t =sig(W o ·[h t-1 ,x t ]+b o )

[0157] Input data x t and the hidden state h from the previous moment t-1 It will be compared with the weight matrix W respectively. o Perform a linear transformation to obtain a new vector; then, change the bias term b. oThis result is added to the linear combination and then mapped to the interval [0,1] using the Sigmoid activation function, thus obtaining the forget gate value o. t ;

[0158] Among them, W o It is the output gate weight matrix, b o It is the output gate bias, o t It is the value of the forget gate.

[0159] Step (2.8) Calculate the candidate cell state: Candidate cell state The input data is x t The latent information, combined with the hidden state from the previous time step, is calculated using the following formula:

[0160]

[0161] The current time step has an input x with a shape of 32*1. t and the hidden state h from the previous moment t-1 After weight matrix W xc and W hc After a linear transformation, a new potential information vector of size 32*64 is generated; then, the combined information of the current input and the historical state is mapped to the range [-1,1] through the tanh activation function;

[0162] in, W represents the candidate memory cell state at the current time step. xc and W hc Let b be the weight matrix. c As the bias term, x t h is the input data for the current time step. t-1 This is the hidden state from the previous moment;

[0163] Tanh is an activation function commonly used in neural networks, with an output range of [-1, 1], and is suitable for many neural network architectures.

[0164] Step (2.9) Output hidden state sequence and cell state update: Output through input gate and forget gate, cell state C t The updated value includes historical information and current input information, and the calculation formula is as follows:

[0165]

[0166] h t =o t ⊙tanh(C t )

[0167] The output of the forget gate o tCompared with the previous cell state C t-1 Multiplying by the element determines how much information from the previous time step will be retained; the output i of the input gate... t This controls which information will be added to the current candidate unit state. In LSTM, by combining the outputs of these three gates, the model can flexibly choose to retain, forget, and output different information, thereby effectively capturing long-term dependencies in time series data. When processing time series data, the LSTM model generates a hidden state sequence h. t , representing the feature information of each time step in the sequence;

[0168] Among them, C t-1 ⊙ represents the cell state at the previous time step, and ⊙ represents the element-wise multiplication operation.

[0169] The third step involves using 1D-CNN to extract the temporal feature matrix from the hidden state output by LSTM, and enhancing the model's ability to remember and recognize key temporal features through the TPA (Temporal Pattern Attention) mechanism. Finally, the extracted comprehensive features are further mapped to predicted bearing life values.

[0170] 1D-CNN is a deep learning model that is typically used to process one-dimensional sequence data. Its main feature is that it extracts features from the input data through one-dimensional convolutional kernels.

[0171] The process of the third step is as follows:

[0172] Step (3.1) Extracting the temporal feature matrix: The goal of 1D-CNN is to perform convolution processing on the output hidden state of LSTM and combine it with pooling layer operations to extract local features of the time series, thereby enhancing the model's ability to learn time series data. This process can be achieved by extracting the temporal feature matrix, as shown in the following formula:

[0173]

[0174] The 1D-CNN uses 32 convolutional kernels, a kernel size of 3, a stride of 1, and employs same padding. The extracted temporal feature matrix... The size is 32*100*32;

[0175] in, Let w represent the feature matrix of the i-th sample at the j-th time step, w be the length of the time series, * represent the convolution operation, and C be the feature matrix of the j-th sample. j,T-w+l Here, T represents the size of the convolution kernel, and l represents the l-th position within the convolution window.

[0176] Convolution is a fundamental operation in signal processing and deep learning. It combines an input signal with a filter to generate an output signal. The purpose of convolution is to extract features or patterns from input data, and it is commonly used in image processing, speech processing, time series analysis, and other fields.

[0177] Step (3.2) Calculate the relevance weights: The process of calculating the relevance weights involves obtaining a set of weight coefficients by evaluating the correlation between the hidden state and the convolutional features. These weight coefficients are used to weight different features, thereby making the model focus more on important time periods or features. The formula for calculating the relevance weights is as follows:

[0178]

[0179] Next, the sigmoid activation function is used to map the correlation values ​​to a weight range. After obtaining the weights, a comprehensive feature vector v is obtained by weighted summation. t The calculation formula is as follows:

[0180]

[0181] in, is a scoring function that measures the similarity between the current hidden state and the features, where m is the total number of features extracted. It is the i-th feature, h t The current hidden state is given, and m is the total number of features extracted. It is the i-th feature, W a This is the attention weight matrix, used to learn the matching H. i C and h t The relationship between α i Let be the attention weight, representing the feature at time step i. The weight of the contribution, v t This is the comprehensive feature vector, representing the temporal feature representation after attention weighting;

[0182] Step (3.3) further maps the extracted comprehensive features to the predicted bearing life value through the TPA mechanism. Specifically, the comprehensive feature vector processed by the time pattern attention mechanism is input into the Dense Layer and mapped to the predicted value y through linear transformation. pred .

[0183] The Dense Layer is a fundamental layer in a neural network. It connects all neurons in the previous layer to each neuron in the current layer. Each neuron receives input from all neurons in the previous layer, and the input is weighted, summed, and then activated to produce an output.

[0184] The process of step (3.3) is as follows:

[0185] Step (3.3.1) Generate feature weights: The TPA mechanism calculates the hidden state h′ and the feature map c extracted by convolution. k The correlation between them is calculated using the following formula:

[0186]

[0187] Where, α k For the k-th time step, the extracted feature map c k The weight of the contribution. `exp()` is the exponential function used to ensure that all attention weights are positive, and `score(h′)` is the weight of the contribution. t ,c k The correlation between the hidden state and the k-th feature map is denoted as follows:

[0188] score(h′ t ,c k )=h′ t c k

[0189] This score measures the similarity between the current hidden state and the convolutional feature map. A high score indicates that the current hidden state is highly correlated with the convolutional feature map, while a low score indicates that the correlation between them is weak.

[0190] Among them, the TPA mechanism is an attention mechanism specifically used for extracting and modeling key time-point features in time series data. Its core idea is to dynamically allocate weights based on the current hidden state and the extracted features, highlighting the importance of key time patterns in the sequence, thereby enhancing the model's ability to remember and recognize time series data.

[0191] Step (3.3.2) generates the comprehensive feature vector: using the feature weights calculated in step (3.3.1) and the convolutional feature map c k The weighted summation yields the comprehensive eigenvector Z, calculated using the following formula:

[0192]

[0193] The comprehensive feature vector Z contains the features extracted by convolution, and the TPA mechanism enhances the memory and recognition of important features;

[0194] Step (3.3.3) yields the predicted value y. predThe Dense Layer receives the comprehensive feature vector from the TPA mechanism. Through this layer, the model linearly transforms these features with the learned weights. The Dense Layer then uses a fully connected neural network layer to compute the final predicted value y. pred The calculation formula is as follows:

[0195] y pred =W1·Z+b1

[0196] The output of the linear transformation is the predicted life value of the bearing, which is usually a continuous value;

[0197] Where W1 is the weight matrix of the Dense Layer, and b1 is the bias term of the Dense Layer;

[0198] The DenseLayer, also known as a fully connected layer, is one of the most common layers in neural networks. In deep learning models, the DenseLayer connects all nodes from the previous layer to every neuron in the current layer, with each neuron having weights and biases.

[0199] This embodiment compares and analyzes other bearing life prediction methods with the method of this invention, and verifies the practical applicability of this invention, including the following steps:

[0200] Step 1: Define three bearing life prediction models as follows:

[0201] BiLSTM is a method for predicting the remaining service life of bearings based on degradation detection and LSTM. It extracts time-domain, frequency-domain, and time-frequency-domain features from the full-life bearing diagnostic signals. Monotonicity and trend metrics for each feature are calculated. By ranking the weighted composite indices of the features, the optimal features that fully reflect degradation information are constructed.

[0202] CNN-LSTM is a prediction method based on the combination of CNN and LSTM. First, feature parameters in the time and frequency domains are extracted, and a segmentation function is used to determine the degradation trend. Then, the feature parameters are normalized and used as input to the CNN to extract information. Finally, the deep features are input into the LSTM for prediction.

[0203] CAMTCN is a lifetime prediction model that combines an attention mechanism and a temporal convolutional network (TCN). It also designs an effective adaptive shrinkage (EAS) model to adapt to eliminate morning interference and improve the accuracy of bearing lifetime prediction.

[0204] Step 2, Experimental Dataset

[0205] The bearing life prediction dataset used in this invention is for the SKF-6205 bearing. The dataset includes data from three different operating conditions and collects multiple parameters such as speed, load, temperature, and vibration. The vibration signal is sampled at a frequency of 25.6 kHz, with data collected every 10 seconds, each acquisition lasting 0.1 seconds. By analyzing this data, the bearing's performance at different stages of degradation can be evaluated, and this data can be used for fault detection, diagnosis, and life prediction.

[0206] Step 3: Define evaluation indicators.

[0207] In this invention, to comprehensively evaluate the model's performance, multiple evaluation metrics are employed, including MAE (Mean Absolutes Error), RMSE (Root Mean Square Error), and an error scoring function. The calculation formulas are as follows:

[0208]

[0209] Where, n a y represents the total number of data points. i This is the actual bearing life value. This is a predicted bearing life value. (Error) i A represents the percentage error of RUL in the i-th sample. i According to Error i Dynamically adjusted weights;

[0210] MAE is a commonly used regression loss function used to calculate the average absolute error between predicted and actual values. It is characterized by low sensitivity to outliers, and compared to RMSE, MAE better reflects the average level of actual error.

[0211] Step four: Analyze and compare the results.

[0212] Table 2 shows the comparison results of this invention with other lifetime prediction models;

[0213]

[0214] The results show that TwinLSTM-TPA outperforms the other three models in terms of MAE, RMSE, and Score. Compared with CAMTCN, CNN-LSTM, and BiLSTM, the MAE value decreases by 28.6%, 44.4%, and 58.3%, respectively, and the RMSE value also decreases significantly. CAMTCN exhibits better convergence, but due to the lack of a dynamic attention mechanism, it performs worse than TwinLSTM-TPA in complex time-series data processing. CAMTCN's MAE value is 0.07, slightly lower than TwinLSTM-TPA.

[0215] However, the TwinLSTM-TPA-based method proposed in this invention performs exceptionally well in experiments, achieving an MAE of 0.05. This demonstrates the effectiveness and superiority of the method of constructing a dynamic model of the rolling bearing physical system, generating rich simulation data under different operating conditions, and then balancing the ratio of the two datasets using a custom loss function before inputting the mixed dataset into TwinLSTM-TPA for bearing life prediction. This provides an efficient solution for the field of bearing life prediction.

Claims

1. A bearing life prediction method based on digital twin and LSTM network, characterized in that, The method includes the following steps: The first step is to establish a digital twin model of the rolling bearing life cycle. This model can flexibly adjust key parameters to adapt to changes in bearing performance under different operating conditions. By inputting parameters, the model can simulate the degradation path of the bearing under actual service conditions. The bearing degradation stages are divided into the stable stage, the defect initiation stage, the defect propagation stage, and the failure stage. The second step is to preprocess the real data and twin data to ensure that the twin data and the real data have the same scale, so as to avoid the impact of data distribution differences on model training. Then, a custom loss function is introduced to balance the ratio of twin data and real data. Finally, the mixed data is input into the TwinLSTM-TPA model. The third step involves using 1D-CNN to extract the temporal feature matrix from the hidden state output by LSTM, and enhancing the model's ability to remember and recognize key temporal features through the TPA mechanism. Finally, the extracted comprehensive features are further mapped to the predicted value of bearing life.

2. The bearing life prediction method based on digital twin and LSTM network as described in claim 1, characterized in that, The process of the first step is as follows: Step (1.1) Modeling the stable stage of the rolling bearing: The inner and outer rings of the bearing each have two degrees of freedom, corresponding to the horizontal displacement x, respectively. i x o and vertical displacement y i y o A unit resonator has only one degree of freedom, namely the vertical displacement y. r The equation of motion for the inner ring of the bearing is shown below: Where, m i For the inner ring quality, c i F is the damping coefficient. i For the external radial load, k i For stiffness, f x and f y These are the nonlinear contact forces applied in the horizontal and vertical directions, respectively. and These represent the velocities of the inner ring in the horizontal and vertical directions, respectively, where g is the acceleration due to gravity. The equation of motion for the outer ring of the bearing is shown below: Where, m o For the outer ring mass, c o c is the damping coefficient. r k is the stiffness of a unit resonant cavity. o For stiffness, k r y is the damping coefficient between the bearing and the support. r This refers to the displacement of the bearing outer ring support structure. The velocity of the bearing outer ring support structure in the vertical direction; The vibration equation of a unit resonator is shown below: Where, m r For a unit resonator mass, the nonlinear contact force is as follows: Where, δ j K represents the total deflection at the j-th roller position. j n represents the contact stiffness between the roller and the roller track. b θ is the number of rollers. i This indicates the angular position of the i-th roller; Step (1.2) Modeling the initial stage of rolling bearing defects: The defects are simplified into an increase in roughness in a local area. The formula for the total deflection of the bearing outer ring is as follows: d j =(x i -x o )cosθ j +(y i -y o )sinθ j +η+ζ+Δζ Where η is the local deformation of the rolling element, ζ is the elastic deformation of the bearing outer ring, and Δζ is the increased roughness in the local area, defined as follows: Where, n j This represents the number of inner ring rotors, j is the rolling element number, and b is the width of the groove pressure. For cross-angle, Let θ be the initial angular position. d ω is the initial angular offset for rolling. c D is the angular velocity of the bearing, t is the time it takes for the rolling elements to rotate, and D is the angular velocity of the bearing. o This refers to the raceway diameter of the bearing's outer ring. Step (1.3) Modeling the defect propagation stage of the rolling bearing: During the defect propagation stage, the roughness of the bearing raceway will further increase, and cracks will appear on the surface. The cracks will propagate from the trailing edge of the indentation. When the rolling element passes through the crack, it will release a certain amount of deformation. The total deflection of the overall deformation is shown below: d j =(x i -x o )cosθ j +(y i -y o )sinθ j +η+ζ1+Δd 1j Where, δ j Let ζ1 represent the total deflection of the overall deformation, and Δd represent the surface roughness of the raceway during the defect propagation stage. 1j For the displacement excitation in this stage, when the rolling element passes through the defect, the displacement excitation in this stage is defined as follows: Where cd is the maximum displacement of the rolling element, as shown below: Where d is the diameter of the rolling element, and b1 is the chord length of the contact surface of the rolling element on both sides of the guide rail; Step (1.4) Modeling the rolling bearing failure stage: In this stage, the width of the defect increases, causing the roll to sink to the bottom of the defect. The deflection in this stage is shown below: d j =(x i -x o )cosθ j +(y i -y o )sinθ j +η+ζ2+Δd 2j Where ζ2 represents the displacement excitation in this stage, and Δd 2j This represents the displacement excitation generated during the failure stage. When the roller passes through the defect in this stage, the displacement excitation is represented as follows: Where H is the defect height, It is the circumferential cross-sectional angle of the roll at the bottom track of the defect. and These are the angles of roller rotation when entering and leaving the defect, respectively. and The definitions are as follows:

3. A bearing life prediction method based on digital twin and LSTM network as described in claim 1 or 2, characterized in that, The second step is as follows: Step (2.1) Data Preprocessing: Standardization is a commonly used data preprocessing method, and its calculation formula is as follows: Where z is the standardized distance, x is the original data point, μ is the mean of the data, and σ is the standard deviation of the data. The processed dataset is divided into datasets using random sampling to ensure the balance of the training set and the test set on the twin data and the real data. Step (2.2) Balance the ratio of twin data to actual data. The custom loss function calculation formula is shown below: Where N is the total number of samples in the dataset, y sim,i and y true,i These are the true values ​​of the i-th sample in the twin data and the actual data, respectively. and It is an indicator function that indicates whether the data is twin data or actual data: when the data source is twin data... Otherwise, the value is 0; when the data source is actual data. Otherwise, it is 0, w sim and w real These are the weighting coefficients for twin data and actual data, and MSE() represents the mean squared error function; Step (2.3) divides the mixed data using a sliding window to ensure that each data sample has the same length. Each data sample has a length of 1000 and a dimension of 1. The divided data samples are used as a new unit signal and input into TwinLSTM-TPA for processing. In step (2.4) of the TwinLSTM-TPA network model, the data is first input into the LSTM network for learning and modeling of temporal features. LSTM effectively controls the transmission of information by introducing three gating mechanisms: the input gate, the forget gate, and the output gate. In this model, the number of hidden units, the dimension of the hidden state, and the dimension of the unit state are all 64. The input gate is responsible for determining the input x at the current time step. t and the hidden state h from the previous moment t-1 For the current cell state c t The impact; Step (2.5) The input gate performs a linear transformation on the current input data x. t and the hidden state h from the previous moment t-1 With weight matrix W respectively i The product is multiplied and mapped to the dimension of the hidden unit to obtain the feature representation. Subsequently, the bias term b of the input gate is added. i Finally, the linear combination result of the input gates is obtained, and its calculation formula is as follows: i t =sig(W i ·[h t-1 ,x t ]+b i ) Among them, W i It is the weight matrix, b i It is the input gate bias, and sig is the sigmoid activation function; Step (2.6) Activation function processing: The result of step (2.5) is processed by the sigmoid activation function to obtain the input gate output i. t The activation function is calculated as follows: The value of each dimension represents the degree of influence of the current time step input on updating the cell state. When the value is close to 1, it means that the current input has a significant impact on the cell state; when the value is close to 0, it means that the current input has a negligible impact on the cell state. Step (2.7) Calculate the forget gate output: The forget gate determines the cell state C of the previous time step. t-1 The formula for determining which information should be discarded and which should be retained is as follows: o t =sig(W o ·[h t-1 ,x t ]+b o ) Input data x t and the hidden state h from the previous moment t-1 It will be compared with the weight matrix W respectively. o Perform a linear transformation to obtain a new vector, and then change the bias term b. o This result is added to the linear combination and then mapped to the interval [0,1] using the Sigmoid activation function, thus obtaining the forget gate value o. t ; Among them, W o It is the output gate weight matrix, b o It is the output gate bias, o t It is the value of the forget gate; Step (2.8) Calculate the candidate cell state: Candidate cell state The input data is x t The latent information, combined with the hidden state from the previous time step, is calculated using the following formula: Input data x at the current time step t and the hidden state h from the previous moment t-1 After weight matrix W xc and W hc After a linear transformation, a new latent information vector is generated. Then, the combined information of the current input and the historical state is mapped to the range [-1, 1] through the tanh activation function. in, W represents the candidate memory cell state at the current time step. xc and W hc Let b be the weight matrix. c As the bias term, x t h is the input data for the current time step. t-1 This is the hidden state from the previous moment; Step (2.9) Output hidden state sequence and cell state update: Output through input gate and forget gate, cell state C t The updated value includes historical information and current input information, and the calculation formula is as follows: h t =o t ⊙tanh(C t ) The output of the forget gate o t Compared with the previous cell state C t-1 Multiplying by the element determines how much information from the previous time step will be retained; the output i of the input gate... t This controls which information will be added to the current candidate unit state. In LSTM, by combining the outputs of these three gates, the model can flexibly choose to retain, forget, and output different information, thereby effectively capturing long-term dependencies in time series data. When processing time series data, the LSTM model generates a hidden state sequence h. t , representing the feature information of each time step in the sequence; Among them, C t-1 ⊙ represents the cell state at the previous time step, and ⊙ represents the element-wise multiplication operation.

4. The bearing life prediction method based on digital twin and LSTM network as described in claim 3, characterized in that, The process of the third step is as follows: Step (3.1) Extracting the temporal feature matrix: The goal of 1D-CNN is to convolve the output hidden states of LSTM to extract local features of the time series, thereby enhancing the model's ability to learn from time series data. This process can be achieved by extracting the temporal feature matrix, as shown in the following formula: in, Let w represent the feature matrix of the i-th sample at the j-th time step, w be the length of the time series, * represent the convolution operation, and C be the feature matrix of the j-th sample. j,T-w+l Here, T represents the size of the convolution kernel, and l represents the l-th position within the convolution window. Step (3.2) Calculate the relevance weights: The process of calculating the relevance weights involves obtaining a set of weight coefficients by evaluating the correlation between the hidden state and the convolutional features. These weight coefficients are used to weight different features, thereby making the model focus more on important time periods or features. The formula for calculating the relevance weights is as follows: Next, the sigmoid activation function is used to map the correlation values ​​to a weight range. After obtaining the weights, a comprehensive feature vector v is obtained by weighted summation. t The calculation formula is as follows: in, is a scoring function that measures the similarity between the current hidden state and the features, where m is the total number of features extracted. It is the i-th feature, h t The current hidden state is given, and m is the total number of features extracted. It is the i-th feature, W a This is the attention weight matrix, used for learning matching. and h t The relationship between α i Let be the attention weight, representing the feature at time step i. The weight of contribution, v t This is the comprehensive feature vector, representing the temporal feature representation after attention weighting; Step (3.3) uses the TPA mechanism to further map the extracted comprehensive features into predicted bearing life values. The comprehensive feature vector processed by the time pattern attention mechanism is input into the Dense Layer and mapped to the predicted value y through linear transformation. pred .

5. The bearing life prediction method based on digital twin and LSTM network as described in claim 4, characterized in that, The process of step (3.3) is as follows: Step (3.3.1) Generate feature weights: The TPA mechanism calculates the hidden state h' and the feature map c extracted by convolution. k The correlation between them is calculated using the following formula: Where, α k For the k-th time step, the extracted feature map c k The contribution weights, exp() is the exponential function used to ensure that all attention weights are positive, score(h′) t ,c k The correlation between the hidden state and the k-th feature map is denoted as follows: score(h′ t ,c k )=h′ t c k This score measures the similarity between the current hidden state and the convolutional feature map; a high score indicates that the current hidden state is highly correlated with the convolutional feature map, while a low score indicates that the correlation between them is weak. Step (3.3.2) generates the comprehensive feature vector: using the feature weights calculated in step (3.3.1) and the convolutional feature map c k The weighted summation yields the comprehensive eigenvector Z, calculated using the following formula: The comprehensive feature vector Z contains the features extracted by convolution, and the TPA mechanism enhances the memory and recognition of important features; Step (3.3.3) yields the predicted value y. pred The Dense Layer receives the comprehensive feature vector from the TPA mechanism. Through this layer, the model linearly transforms these features with the learned weights. The Dense Layer then uses a fully connected neural network layer to compute the final predicted value y. pred The calculation formula is as follows: y pred =W1·Z+b1 The output of the linear transformation is the predicted bearing life, which is a continuous value. Where W1 is the weight matrix of the Dense Layer, and b1 is the bias term of the Dense Layer.