Bearing life prediction method fusing time sequence mode mechanism and neural network

By combining LSTM, 1D-CNN and TPA mechanisms, the accuracy and robustness of bearing life prediction are improved, solving the problem of large prediction errors in existing technologies and achieving more accurate life prediction and reasonable maintenance plans.

CN120929783APending Publication Date: 2025-11-11CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510833403.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing bearing life prediction methods suffer from large and difficult-to-handle errors when dealing with complex nonlinear problems, and are unable to adapt to complex operating environments and changing working conditions, resulting in inaccurate predictions.

Method used

The LSTM network is used to capture the long-term dependencies of time series data, combined with 1D-CNN to extract key time features, and the TPA mechanism is introduced to enhance the model's attention to historical moments and current states. Finally, prediction is performed through a Dense Layer.

Benefits of technology

It improves the accuracy and robustness of bearing life prediction, enabling early prediction of bearing life, helping to plan reasonable maintenance times, reduce unnecessary maintenance costs, and ensure efficient and safe equipment operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929783A_ABST
    Figure CN120929783A_ABST
Patent Text Reader

Abstract

A bearing life prediction method fusing a time sequence mode mechanism and a neural network comprises the following steps: step 1, dividing time sequence data of an elevator traction machine bearing through a sliding window, and inputting the divided data into a standard LSTM network in a segmented manner for processing; and 2, extracting a time feature matrix from a hidden state output by the LSTM by using 1D-CNN, and enhancing the memory and recognition capability of the model on key time features through a TPA mechanism. And finally, further mapping the extracted comprehensive features into a predicted value of the service life of the bearing. And step 3, enhancing the memory and identification capability of the model on key time features through a TPA mechanism according to the extracted time features, and finally further mapping the extracted comprehensive features into a predicted value of the bearing life. The method shows relatively high prediction precision and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bearing life prediction and relates to a bearing life prediction method that integrates time-series pattern mechanism and neural network. Background Technology

[0002] By the end of 2023, the total number of special equipment in China reached 21.5891 million units, of which elevators numbered 10.6298 million, accounting for 49.23% of the total. Compared with 2022, the number of elevators achieved a year-on-year growth of 11.88%, maintaining a rapid growth trend for several consecutive years. The widespread use of elevators presupposes that their operation must ensure a high level of safety, especially the elevator traction machine, which, as one of the core components of the elevator system, directly affects the overall safety of the elevator. Among the components of the traction machine, bearings have the highest probability of failure.

[0003] As bearings age, they experience wear, fatigue, and corrosion, leading to performance degradation and even downtime, causing equipment disruptions, economic losses, and safety hazards. Therefore, timely and accurate prediction of bearing remaining life is crucial for ensuring safe equipment operation, improving production efficiency, and reducing maintenance costs. Common bearing life prediction methods include model-based methods, which predict life by establishing physical or mathematical models of bearing wear. These methods typically include statistical distribution models, wear models based on physical laws, and life prediction based on thermodynamics and fatigue models. While these methods provide some theoretical basis, they often have significant errors in practical applications and are difficult to handle complex nonlinear problems. With the development of sensor technology, deep learning methods have become a research hotspot for bearing life prediction in recent years. These methods mainly rely on real-time acquisition of vibration signals, temperature, sound, and other data from the equipment, using signal processing and machine learning algorithms to predict life. Summary of the Invention

[0004] To overcome the shortcomings of existing technologies, this invention proposes a bearing life prediction method integrating a temporal pattern mechanism and a neural network (Dynamic Temporal Pattern Attention Network, abbreviated as DynTPA-NeT). This model captures the long-term dependencies of time-series data using LSTM (long short-term memory), extracts key temporal features using 1D-CNN (One-Dimensional Convolutional Neural Network), and introduces the TPA (Temporal Pattern Attention mechanism) to enhance attention to the correlation between historical moments and the current state, thereby improving the sensitivity and memory capacity for key temporal features. TPA improves the accuracy and robustness of the model's bearing life prediction by weighting features at different time steps. Finally, the model maps the extracted comprehensive features to the predicted bearing life value through a Dense Layer, demonstrating high prediction accuracy and stability.

[0005] The technical solution adopted by this invention to solve its technical problem is:

[0006] A bearing life prediction method integrating temporal pattern mechanisms and neural networks includes the following steps:

[0007] The first step is to divide the time series data of the elevator traction machine bearings into segments using a sliding window and input them into a standard LSTM network for processing.

[0008] The second step involves using 1D-CNN to extract the temporal feature matrix from the hidden states output by the LSTM, and enhancing the model's ability to remember and recognize key temporal features through the TPA (Temporal Pattern Attention) mechanism. Finally, the extracted comprehensive features are further mapped to predicted bearing life values.

[0009] 1D-CNN is a variant of convolutional neural networks specifically designed for processing one-dimensional data. 1D-CNN extracts local features from data through convolution operations and is widely used in tasks that process one-dimensional signals.

[0010] The third step involves enhancing the model's ability to remember and recognize key time features through the TPA mechanism. Finally, the extracted comprehensive features are further mapped to predicted bearing life values; that is, the comprehensive feature vector h after processing by the time pattern attention mechanism. t The input to the Dense Layer is mapped to the predicted value y through a linear transformation. pred.

[0011] Furthermore, the process of the first step is as follows:

[0012] Step (1.1) Data partitioning uses a sliding window method to divide the time series data into equal-length subsequence signals. Each subsequence signal has a length of 1000 and a data dimension of 1. Each subsequence signal is used as a new signal unit and is input into the LSTM network for processing.

[0013] The sliding window method is a commonly used data processing technique, particularly suitable for handling continuous data such as time series and text data. Its core idea is to use a fixed-size window that slides across the data in increments of a certain size, extracting and processing a portion of the data covered by the window each time. The sliding window method effectively captures local features and reduces computational complexity.

[0014] Step (1.2) LSTM effectively manages the flow of information by introducing three gating mechanisms: the input gate, the forget gate, and the output gate. The number of hidden units in LSTM is 64, and the hidden state h t and unit state C t Both have 64 dimensions, and the input gate determines the current input x. t and the hidden state h from the previous moment t-1 Information about the current unit state c t The impact;

[0015] Step (1.3) Calculate the forget gate output: The forget gate determines the cell state C of the previous time step. t-1 The formula for determining which information should be discarded and which should be retained is as follows:

[0016] f t =σ(W f ·[h t-1 ,x t ]+b f )

[0017] Among them, W f It is the forget gate weight matrix, b f It is a forgetting gate bias;

[0018] Forgot gate output f t Control from the previous unit state C t-1 How much information is retained? t A value close to 0 indicates that this part of the information will be forgotten; f t A value close to 1 indicates that this part of the information is retained;

[0019] Step (1.4) Calculate the output gate output: The output gate controls the current unit C. tInformation about the hidden state h from the previous time step t-1 The effect of [the effect] is calculated using the following formula:

[0020] o t =σ(W o ·[h t-1 ,x t ]+b o )

[0021] Input data x t and the hidden state h from the previous moment t-1 It will be through the weight matrix W o A linear transformation is performed to obtain a new vector with a bias of b. o The result of this linear combination is added to the output, and finally, the sigmoid activation function maps the linear combination to the range [0,1] to obtain the output gate value o. t ;

[0022] Among them, W o It is the output gate weight matrix, b o It is the output gate bias;

[0023] Step (1.5) Calculate the candidate cell state: Candidate cell state The input data is x t and the hidden state h from the previous moment t-1 The combined potential information is calculated using the following formula:

[0024]

[0025] Current time step x t and the hidden state h from the previous moment t-1 Through the weight matrix W xc and W hc A linear transformation is used to generate a new latent information vector, and then the tanh activation function is used to restrict the combined information of the current input and the historical state to the range of [-1,1].

[0026] in, For candidate cell states, W xc Is input x t The weight matrix of the candidate memory unit determines the influence of the input data on the candidate state. hc The hidden state h from the previous moment t-1 The weight matrix of the candidate memory unit determines the influence of the previous hidden state on the candidate state, b c This is a bias term used to adjust the calculation results and improve the expressive power of the model;

[0027] Tanh is an activation function commonly used in neural networks, with an output range of [-1, 1], and is suitable for various neural network architectures.

[0028] Step (1.6) Update the cell state: Output the cell state C through the input gate and the forget gate. t The updated value includes historical information and current input information, and the calculation formula is as follows:

[0029]

[0030] Forgot gate output f t Compared with the previous cell state C t-1 Element-wise multiplication indicates how much information from the previous time step is retained; the input gate outputs i. t Control candidate cell state Which information is added to the current cell state? Then, combining the outputs of these three gates, LSTM can dynamically select which information to retain, forget, and output, thereby effectively capturing long-term dependencies in the sequence. When processing time series data, the LSTM model outputs a hidden state sequence h. t , representing the feature information of each time step in the sequence;

[0031] Among them, C t-1 The value represents the cell state at the previous time step, and ⊙ represents the element-wise multiplication operation, used to weight the state update.

[0032] Preferably, the process of step (1.2) is as follows:

[0033] Step (1.2.1) Linear Transformation of the Input Gate: The input gate performs a linear transformation on the input data x. t and the hidden state h from the previous moment t-1 Through the weight matrix W i Mapping yields feature representations at the hidden unit dimension, plus the input gate bias b. i The linear combination result of the input gates is obtained, and its calculation formula is shown below:

[0034] i t =σ(W i ·[h t-1 ,x t ]+b i )

[0035] Among them, W i It is the weight matrix, b i It is the bias, and σ is the sigmoid activation function;

[0036] Linear transformation is a fundamental concept in mathematics, widely applied in algebra, geometry, physics, and many other fields. In linear algebra, a linear transformation refers to a transformation within a vector space that preserves the structural integrity of addition and scalar multiplication operations through a mapping.

[0037] Step (1.2.2) Activation function processing: The result of step (1.2.1) is processed by the sigmoid activation function to obtain the input gate output i. t The activation function is calculated as follows:

[0038]

[0039] The value of each dimension represents the current time step input x. t The degree of contribution when updating the cell state: the closer the value is to 1, the greater the contribution of the current input to the cell state; when the value is close to 0, the contribution of the current input to the cell state is negligible.

[0040] The sigmoid activation function is one of the common non-linear activation functions in neural networks, widely used in early neural network models, especially in the output layer of binary classification problems. It maps the input signal to a fixed range, typically between 0 and 1.

[0041] The second step is as follows:

[0042] Step (2.1) Extracting the temporal feature matrix: The goal of 1D-CNN is to perform convolution operations on the hidden states output by LSTM to extract local features of the time series, thereby enhancing the model's ability to learn time series data and extracting the temporal feature matrix. The formula is shown below:

[0043]

[0044] in, Let w represent the feature matrix of the i-th sample at the j-th time step, w be the length of the time series, * denote the convolution operation, and C be the feature matrix of the j-th sample. j,T-w+l Here, T represents the kernel size, and l represents the l-th position within the convolution window.

[0045] Convolution is a fundamental operation in signal processing, image processing, deep learning, and other fields. In particular, it is a core component of convolutional neural networks. Convolution extracts local features of the input signal by applying a convolution kernel, thereby achieving feature extraction and transformation.

[0046] Step (2.2) Calculate the relevance weights: A set of weight coefficients is obtained by calculating the correlation between the hidden state and the convolutional features, thereby weighting different features and enhancing the model's focus on important time periods or features. The weight calculation formula is as follows:

[0047]

[0048] Then, the sigmoid activation function is used to map the correlation to a weight range, and after obtaining the weights, a weighted sum is used to obtain the comprehensive feature vector v. t The calculation formula is as follows:

[0049]

[0050] in, It is a scoring function that measures the similarity between the current hidden state and the features, h. t The current hidden state is given, and m is the total number of features extracted. It is the i-th feature, W a This is the attention weight matrix, used for learning matching. and h t The relationship between α i Let be the attention weight, representing the feature at time step i. The weight of the contribution. t This is the comprehensive feature vector, representing the temporal feature representation after attention weighting;

[0051] Step (2.3) Calculate the new hidden state: Combine the current hidden state h t and the comprehensive feature vector v t Calculate the new hidden state h t The calculation formula is as follows:

[0052] h t ′=W h′ (W h h t +W v v t )

[0053] Among them, W h′ W h W v The weight matrix is ​​used to linearly combine the current hidden state and the comprehensive feature vector. Through the above steps, 1D-CNN can effectively extract temporal features from the LSTM output. t ′ represents the new hidden state.

[0054] The process of the third step is as follows:

[0055] Step (3.1) Generate feature weights: The TPA mechanism calculates the hidden state h t Feature map c extracted by convolution. k The correlation between them is calculated using the following formula:

[0056]

[0057] Where, α k For the k-th time step, the extracted feature map c k The weight of the contribution, score(h) LSTM ,c k ) is a correlation measure between the hidden state and the k-th feature map, and exp() is an exponential function used to ensure that all attention weights are positive;

[0058] score(h t ′,c k ) = h t 'c k

[0059] This score measures the similarity between the current hidden state and the convolutional feature map. A high score indicates that the current hidden state is highly correlated with the convolutional feature map, while a low score indicates that they are not correlated.

[0060] Among them, the TPA mechanism is an attention mechanism that focuses on the extraction and modeling of key time-point features in time series data. Its core idea is to dynamically allocate weights, highlighting key time patterns in the sequence based on the current hidden state and extracted features, thereby improving the model's ability to remember and recognize time series data.

[0061] Step (3.2) generates the comprehensive feature vector: using the feature weights calculated in step (3.1) and the convolutional feature map c k The weighted summation yields the comprehensive eigenvector z, calculated using the following formula:

[0062]

[0063] The comprehensive feature vector z contains the features extracted by convolution, and the TPA mechanism enhances the memory and recognition of important features;

[0064] Step (3.3) inputs to the Dense Layer: The Dense Layer receives the comprehensive feature vector from the TPA mechanism. Through this layer, the model linearly transforms these features with the learned weights. The Dense Layer then computes the final predicted value y through a fully connected neural network layer. pred The calculation formula is as follows:

[0065] y pred=W1·z+b1

[0066] The output of the linear transformation is the predicted life value of the bearing, which is usually a continuous value;

[0067] Where W1 is the weight matrix of the Dense Layer, and b1 is the bias term of the Dense Layer.

[0068] The DenseLayer, also known as a fully connected layer, is one of the most common layers in neural networks. In deep learning models, the DenseLayer connects all nodes from the previous layer to every neuron in the current layer, with each neuron having weights and biases.

[0069] Compared with existing technologies, the beneficial effects of this invention are: it proposes a bearing life prediction method that integrates a time-series pattern mechanism and a neural network. The DynTPA-NeT model can effectively capture long-term dependencies in time-series data, while using 1D-CNN to extract key temporal features. Finally, the predicted life value is obtained through the TPA mechanism and DenseLayer layer. The DynTPA-NeT model can focus on the correlation between historical moments and the current state during training, thereby enhancing the ability to identify and remember key temporal features. This solves the problem that traditional bearing life prediction methods often rely on limited historical data and empirical knowledge, making it difficult to adapt to complex operating environments and changing working conditions. DynTPA-NeT can not only predict bearing life in advance, but also help engineers plan reasonable maintenance times, avoiding premature or late maintenance. This can greatly reduce unnecessary maintenance costs, ensure efficient equipment operation, and ensure the stability and safety of elevator operation. Attached Figure Description

[0070] Figure 1 This paper presents a framework for a bearing life prediction method that integrates a temporal pattern mechanism and a neural network.

[0071] Figure 2 The network structure of DynTPA-NeT is shown. Detailed Implementation

[0072] The invention will now be further described with reference to the accompanying drawings.

[0073] Reference Figure 1 and Figure 2 A bearing life prediction method integrating temporal pattern mechanism and neural network includes the following steps:

[0074] The first step is to divide the time-series data of the elevator traction machine bearings into segments using a sliding window and input them into a standard LSTM network for processing. The process is as follows:

[0075] Step (1.1) Data partitioning uses a sliding window approach to divide the time series data into equal-length subsequence signals. The sliding window length is set to 1000, and the step size is 100. Each time, a data sequence of length 10000 is processed. After passing through the sliding window, it is ensured that each data sample has the same length. Each data sample has a length of 1000 and a dimension of 1. The partitioned data samples are then used as new unit signals and input into DynTPA-NeT for processing.

[0076] The sliding window method is a commonly used data processing technique, particularly suitable for handling continuous data such as time series and text data. Its core idea is to use a fixed-size window that slides across the data in increments of a certain size, extracting and processing a portion of the data covered by the window each time. The sliding window method effectively captures local features and reduces computational complexity.

[0077] Step (1.2) LSTM effectively manages the flow of information by introducing three gating mechanisms: the input gate, the forget gate, and the output gate. The LSTM has 64 hidden units and 64 hidden states h. t and unit state C t Both have 64 dimensions. The input gate determines the current input x. t and the hidden state h from the previous moment t-1 Information about the current unit state C t The impact.

[0078] Step (1.2.1) Linear Transformation of the Input Gate: The input gate performs a linear transformation on the input data x. t and the hidden state h from the previous moment t-1 Through the weight matrix W i Mapping yields feature representations at the hidden unit dimension, plus the input gate bias b. i The linear combination result of the input gates is obtained. The calculation formula is shown below:

[0079] i t =σ(W i ·[h t-1 ,x t ]+b i )

[0080] Among them, W i It is the weight matrix, b i σ is the bias, and σ is the sigmoid activation function.

[0081] Linear transformation is a fundamental concept in mathematics, widely applied in algebra, geometry, physics, and many other fields. In linear algebra, a linear transformation refers to a transformation within a vector space that preserves the structural integrity of addition and scalar multiplication operations through a mapping.

[0082] Step (1.2.2) Activation function processing: The result of step (1.2.1) is processed by the sigmoid activation function to obtain the input gate output i. t The activation function is calculated as follows:

[0083]

[0084] The value of each dimension represents the current time step input x. t The degree of contribution of the current input to the cell state update. The closer the value is to 1, the greater the contribution of the current input to the cell state. When the value is close to 0, the contribution of the current input to the cell state is negligible.

[0085] The sigmoid activation function is one of the common non-linear activation functions in neural networks, widely used in early neural network models, especially in the output layer of binary classification problems. It maps the input signal to a fixed range, typically between 0 and 1.

[0086] Step (1.3) Calculate the forget gate output: The forget gate determines the cell state C of the previous time step. t-1 Which information should be discarded and which should be retained? The calculation formula is:

[0087] f t =σ(W f ·[h t-1 ,x t ]+b f )

[0088] Among them, W f It is the forget gate weight matrix, b f It is the forget gate bias.

[0089] Forgot gate output f t Control from the previous unit state C t-1 How much information is retained? t A value close to 0 indicates that this part of the information will be forgotten; f t A value close to 1 indicates that this part of the information is retained.

[0090] Step (1.4) Calculate the output gate output: The output gate controls the current unit C. t Information about the hidden state h from the previous time step t-1 The effect of [the effect] is calculated using the following formula:

[0091] ot =σ(W o ·[h t-1 ,x t ]+b o )

[0092] Input data x t and the hidden state h from the previous moment t-1 It will be through the weight matrix W o A linear transformation is performed to obtain a new vector. Bias b o The result of this linear combination is added to the output, and finally, the sigmoid activation function maps the linear combination to the range [0,1] to obtain the output gate value o. t .

[0093] Among them, W o It is the output gate weight matrix, b o It is the output gate bias.

[0094] Step (1.5) Calculate the candidate cell state: Candidate cell state The input data is x t and the hidden state h from the previous moment t-1 The combined potential information. The specific calculation formula is as follows:

[0095]

[0096] The current time step has an input x with a shape of 32*1. t For the hidden state h from the previous time step t-1 Through the weight matrix W xc and W hc After linear transformation, a new latent information vector of size 32*64 is generated. Then, the tanh activation function is used to restrict the combined information of the current input and the historical state to the range [-1, 1].

[0097] in, For candidate cell states, W xc Is input x t The weight matrix of the candidate memory unit determines the influence of the input data on the candidate state. hc The hidden state h from the previous moment t-1 The weight matrix of the candidate memory unit determines the influence of the previous hidden state on the candidate state. c This is a bias term used to adjust the calculation results and improve the expressive power of the model.

[0098] Tanh is an activation function commonly used in neural networks, with an output range of [-1, 1], and is suitable for various neural network architectures.

[0099] Step (1.6) Update the cell state: Output the cell state C through the input gate and the forget gate. t The updated value includes historical information and current input information. The specific calculation formula is as follows:

[0100]

[0101] h t =o t ⊙tanh(C t )

[0102] Forgot Gate Output o t Compared with the previous cell state C t-1 Element-wise multiplication indicates how much information from the previous time step is retained. Input gate outputs i. t Control candidate cell state Which information is added to the current cell state? Then, by combining the outputs of these three gates, LSTM can dynamically select which information to retain, forget, and output, thereby effectively capturing long-term dependencies in the sequence. When processing time series data, the LSTM model outputs a sequence of hidden states.

[0103] Among them, C t-1 The value represents the cell state at the previous time step, and ⊙ represents the element-wise multiplication operation, used to weight the state update.

[0104] The second step involves using 1D-CNN to extract the temporal feature matrix from the hidden states output by the LSTM, and enhancing the model's ability to remember and recognize key temporal features through the TPA (Temporal Pattern Attention) mechanism. Finally, the extracted comprehensive features are further mapped to predicted bearing life values.

[0105] Combination Figure 2 The mixed data first passes through the hidden state h of the LSTM output data sequence. t h t-1 …h t-w+1 h t-w 1D-CNN extracts fixed-length time-series features from the hidden states output by LSTM, and then determines the current time h using a scoring function via a TPA mechanism. t and the previous time h t-w The weights are used to determine the final hidden state h' based on the weights at the current time. t Finally, the remaining life prediction of the bearing is obtained through the Dense Layer.

[0106] 1D-CNN is a variant of convolutional neural networks specifically designed for processing one-dimensional data. It extracts local features from data through convolutional operations and is widely used for tasks involving one-dimensional signals.

[0107] Step (2.1) Extracting the temporal feature matrix: The goal of 1D-CNN is to perform convolution operations on the hidden states output by LSTM to extract local features of the time series, thereby enhancing the model's ability to learn from time series data. Extracting the temporal feature matrix... The formula is shown below:

[0108]

[0109] in, This represents the feature matrix of the i-th sample at the j-th time step. w is the length of the time series, * represents the convolution operation, and C... j,T-w+l Let T be the convolution kernel, and T represent the size of the convolution kernel. l represents the l-th position within the convolution window.

[0110] Convolution is a fundamental operation in signal processing, image processing, deep learning, and other fields. It is a core component of convolutional neural networks. Convolution extracts local features from the input signal by applying a convolution kernel, thereby achieving feature extraction and transformation.

[0111] Step (2.2) Calculate the relevance weights: A set of weight coefficients is obtained by calculating the correlation between the hidden state and the convolutional features, thereby weighting different features and enhancing the model's focus on important time periods or features. The weight calculation formula is as follows:

[0112]

[0113] Then, the sigmoid activation function is used to map the correlations to weight ranges. After obtaining the weights, a weighted summation is performed to obtain the comprehensive feature vector v. t The calculation formula is as follows:

[0114]

[0115] in, It is a scoring function that measures the similarity between the current hidden state and the features. t The current hidden state is given, and m is the total number of features extracted. It is the i-th feature, W a This is the attention weight matrix, used for learning matching. and h t The relationship between them. α i Let be the attention weight, representing the feature at time step i. The weight of the contribution. t The comprehensive feature vector represents the temporal feature representation after attention weighting.

[0116] Step (2.3) Calculate the new hidden state: Combine the current hidden state h t and the comprehensive feature vector v t Calculate the new hidden state h t The calculation formula is as follows:

[0117] h t ′=W h′ (W h h t +W v v t )

[0118] Among them, W h′ W h W v The weight matrix is ​​used to linearly combine the current hidden state and the synthesized feature vector. Through the above steps, 1D-CNN can effectively extract temporal features from the LSTM output. t ′ represents the new hidden state.

[0119] The third step involves enhancing the model's ability to remember and recognize key time features through the TPA mechanism. Finally, the extracted composite features are further mapped to predicted bearing life values. Specifically, the composite feature vector h processed by the time pattern attention mechanism... t The input to the Dense Layer is mapped to the predicted value y through a linear transformation. pred .

[0120] Step (3.1) Generate feature weights: The TPA mechanism calculates h t Feature map c extracted by convolution. k The correlation between them is calculated using the following formula:

[0121]

[0122] Where, α k For the k-th time step, the extracted feature map c k The weight of the contribution. score(h) t ′,c k ) is a measure of the correlation between the hidden state and the k-th feature map. exp() is an exponential function used to ensure that all attention weights are positive.

[0123] score(h t ′,c k ) = h t 'ck

[0124] This score measures the similarity between the current hidden state and the convolutional feature map. A high score indicates that the current hidden state is highly correlated with the convolutional feature map, while a low score indicates that they are not correlated.

[0125] Among them, the TPA mechanism is an attention mechanism that focuses on the extraction and modeling of key time-point features in time series data. Its core idea is to dynamically allocate weights, highlighting key time patterns in the sequence based on the current hidden state and extracted features, thereby improving the model's ability to remember and recognize time series data.

[0126] Step (3.2) generates the comprehensive feature vector: using the feature weights calculated in step (3.1) and the convolutional feature map c k The weighted summation yields the comprehensive eigenvector z. The specific calculation formula is as follows:

[0127]

[0128] The comprehensive feature vector z contains the features extracted by convolution, and the TPA mechanism enhances the memory and recognition of important features.

[0129] Step (3.3) inputs to the Dense Layer: The Dense Layer receives the comprehensive feature vector from the TPA mechanism. Through this layer, the model linearly transforms these features with the learned weights. The Dense Layer then computes the final predicted value y through a fully connected neural network layer. pred The calculation formula is as follows:

[0130] y pred =W1·z+b1

[0131] The output of the linear transformation is the predicted bearing life, which is usually a continuous value.

[0132] Where W1 is the weight matrix of the Dense Layer, and b1 is the bias term of the Dense Layer.

[0133] The DenseLayer, also known as a fully connected layer, is one of the most common layers in neural networks. In deep learning models, the DenseLayer connects all nodes from the previous layer to every neuron in the current layer, with each neuron having weights and biases.

[0134] This embodiment demonstrates an ablation experiment using the method of the present invention and verifies the practical applicability of each module of the present invention, including the following steps:

[0135] Step 1, Experimental Dataset

[0136] The bearing life prediction dataset used in this invention is for the SKF-6205 bearing. The dataset includes data from three different operating conditions and collects multiple parameters such as speed, load, temperature, and vibration. The vibration signal is sampled at a frequency of 25.6 kHz, with data collected every 10 seconds, each acquisition lasting 0.1 seconds. By analyzing this data, the bearing's performance at different stages of degradation can be evaluated, and this data can be used for fault detection, diagnosis, and life prediction.

[0137] Step two, define the evaluation indicators.

[0138] In this invention, to comprehensively evaluate the model's performance, RMSE (Root Mean Square Error) is used as the evaluation metric, calculated using the following formula:

[0139]

[0140] Where N is the number of samples, i represents the i-th sample, and n a y represents the total number of data points. i This is the actual bearing life value. This is the predicted bearing life.

[0141] MAE is a commonly used regression loss function used to calculate the average absolute error between predicted and actual values. It is characterized by low sensitivity to outliers, and compared to RMSE, MAE better reflects the average level of actual error.

[0142] RMSE is a commonly used regression loss function and a frequently used regression evaluation metric used to measure the degree of deviation between model predictions and actual values. RMSE comprehensively reflects the overall level of error across all samples and is more sensitive to larger errors, thus having high reference value in evaluating model performance. The smaller the RMSE value, the closer the model's predictions are to the actual values, and the better its performance.

[0143] Step 3: Analyze the ablation results.

[0144] This invention designs ablation experiments to verify the impact of the TPA mechanism, LSTM module, and 1D-CNN module on the performance of DynTPA-NeT by removing or replacing different modules in the model.

[0145] Table 1 shows the comparative results of the ablation experiments of this invention.

[0146]

[0147]

[0148] Table 1

[0149] Table 1 shows that the error gradually decreased as the modules were improved, validating the necessity of each module. Model 1, as the baseline model, added LSTM and TPA mechanisms but lacked 1D-CNN, achieving MAE and RMSE of 0.18 and 0.14, respectively. This indicates that although the TPA mechanism helps the model focus on important temporal features, the lack of 1D-CNN's feature extraction and noise suppression capabilities for time-series data limits its performance. Model 2 combines LSTM and 1D-CNN. Due to the lack of TPA, Model 2 performs poorly when processing data with complex temporal dependencies, resulting in significantly higher MAE and RMSE than Model 1. This result demonstrates that the prediction model needs the TPA module to extract key time points from the fault data time series. In Model 3, LSTM captures long-term dependency features, 1D-CNN effectively extracts local temporal features from the data, and the TPA mechanism further enhances the model's sensitivity to important time points, ensuring that the model can more accurately predict the remaining service life of the bearing. These factors combined enable Model 3 to achieve the highest accuracy compared to other models when processing complex time-series data. The DynTPA-NeT-based method proposed in this invention demonstrates the effectiveness of dynamic attention mechanisms and provides an efficient solution for bearing life prediction.

Claims

1. A bearing life prediction method integrating temporal pattern mechanism and neural network, characterized in that, The method includes the following steps: The first step is to divide the time series data of the elevator traction machine bearings into segments using a sliding window and input them into a standard LSTM network for processing. The second step is to use 1D-CNN to extract the time feature matrix from the hidden state output by LSTM, and enhance the model's ability to remember and recognize key time features through the TPA mechanism. Finally, the extracted comprehensive features are further mapped to the predicted value of bearing life. The third step is to process the comprehensive feature vector h after the time-pattern attention mechanism. t The input to the DenseLayer layer is mapped to the predicted value y through a linear transformation. pred .

2. The bearing life prediction method integrating temporal pattern mechanism and neural network as described in claim 1, characterized in that, The process of the first step is as follows: Step (1.1) Data partitioning uses a sliding window method to divide the time series data into equal-length subsequence signals. Each subsequence signal has a length of 1000 and a data dimension of 1. Each subsequence signal is used as a new signal unit and is input into the LSTM network for processing. Step (1.2) LSTM effectively manages the flow of information by introducing three gating mechanisms: the input gate, the forget gate, and the output gate. The number of hidden units in LSTM is 64, and the hidden state h t and unit state C t Both have 64 dimensions, and the input gate determines the current input x. t and the hidden state h from the previous moment t-1 Information about the current unit state c t The impact; Step (1.3) Calculate the forget gate output: The forget gate determines the cell state C of the previous time step. t-1 The formula for determining which information should be discarded and which should be retained is as follows: f t =σ(W f ·[h t-1 ,x t ]+b f ) Among them, W f It is the forget gate weight matrix, b f It is a forgetting gate bias; Forgot gate output f t Control from the previous unit state C t-1 How much information is retained in f? t A value close to 0 indicates that this part of the information will be forgotten; f t A value close to 1 indicates that this part of the information is retained; Step (1.4) Calculate the output gate output: The output gate controls the current unit C. t Information about the hidden state h from the previous time step t-1 The effect of [the effect] is calculated using the following formula: the t =σ(W o ·[h t-1 ,x t ]+b o ) Input data x t and the hidden state h from the previous moment t-1 It will be through the weight matrix W o A linear transformation is performed to obtain a new vector with a bias of b. o The result of this linear combination is added to the output, and finally, the sigmoid activation function maps the linear combination to the range [0,1] to obtain the output gate value o. t ; Among them, W o It is the output gate weight matrix, b o It is the output gate bias; Step (1.5) Calculate the candidate cell state: Candidate cell state The input data is x t and the hidden state h from the previous moment t-1 The combined potential information is calculated using the following formula: Current time step x t and the hidden state h from the previous moment t-1 Through the weight matrix W xc and W hc A linear transformation is used to generate a new latent information vector, and then the tanh activation function is used to restrict the combined information of the current input and the historical state to the range of [-1,1]. in, For candidate cell states, W xc Is input x t The weight matrix of the candidate memory unit determines the influence of the input data on the candidate state, W. hc The hidden state h from the previous moment t-1 The weight matrix of the candidate memory unit determines the influence of the previous hidden state on the candidate state, b c This is a bias term used to adjust the calculation results and improve the expressive power of the model; Step (1.6) Update the cell state: Output the cell state C through the input gate and the forget gate. t The updated value includes historical information and current input information, and the calculation formula is as follows: Forgot gate output f t Compared with the previous cell state C t-1 Element-wise multiplication indicates how much information from the previous time step is retained; the input gate outputs i. t Control candidate cell state Which information is added to the current cell state? Then, combining the outputs of these three gates, LSTM can dynamically select which information to retain, forget, and output, thereby effectively capturing long-term dependencies in the sequence. When processing time series data, the LSTM model outputs a hidden state sequence h. t , representing the feature information of each time step in the sequence; Among them, C t-1 The value represents the cell state at the previous time step, and ⊙ represents the element-wise multiplication operation, used to weight the state update.

3. A bearing life prediction method integrating temporal pattern mechanism and neural network as described in claim 1 or 2, characterized in that, The process of step (1.2) is as follows: Step (1.2.1) Linear Transformation of the Input Gate: The input gate performs a linear transformation on the input data x. t and the hidden state h from the previous moment t-1 Through the weight matrix W i Mapping yields feature representations at the hidden unit dimension, plus the input gate bias b. i The linear combination result of the input gates is obtained, and its calculation formula is shown below: i t =σ(W i ·[h t-1 ,x t ]+b i ) Among them, W i It is the weight matrix, b i It is the bias, and σ is the sigmoid activation function; Step (1.2.2) Activation function processing: The result of step (1.2.1) is processed by the sigmoid activation function to obtain the input gate output i. t The activation function is calculated as follows: The value of each dimension represents the current time step input x. t The contribution of the current input to the cell state is indicated by a value closer to 1, which means the current input contributes more to the cell state, and a value closer to 0 means the current input contributes negligibly to the cell state.

4. A bearing life prediction method integrating temporal pattern mechanism and neural network as described in claim 1 or 2, characterized in that, The second step is as follows: Step (2.1) Extracting the temporal feature matrix: The goal of 1D-CNN is to perform convolution operations on the hidden states output by LSTM to extract local features of the time series, thereby enhancing the model's ability to learn time series data and extracting the temporal feature matrix. The formula is shown below: in, Let w represent the feature matrix of the i-th sample at the j-th time step, w be the length of the time series, * denote the convolution operation, and C be the feature matrix of the j-th sample. j,T-w+l Here, T represents the size of the convolution kernel, and l represents the l-th position within the convolution window. Step (2.2) Calculate the relevance weights: A set of weight coefficients is obtained by calculating the correlation between the hidden state and the convolutional features, thereby weighting different features and enhancing the model's focus on important time periods or features. The weight calculation formula is as follows: Then, the sigmoid activation function is used to map the correlation to a weight range, and after obtaining the weights, a weighted sum is used to obtain the comprehensive feature vector v. t The calculation formula is as follows: in, It is a scoring function that measures the similarity between the current hidden state and the features, h. t The current hidden state is given, and m is the total number of features extracted. It is the i-th feature, W a This is the attention weight matrix, used for learning matching. and h t The relationship between α i Let be the attention weight, representing the feature at time step i. The weight of contribution, v t This is the comprehensive feature vector, representing the temporal feature representation after attention weighting; Step (2.3) Calculate the new hidden state: Combine the current hidden state h t and the comprehensive feature vector v t Calculate the new hidden state h t The calculation formula is as follows: h t ′=W h′ (W h h t +W v v t ) Among them, W h′ W h W v The weight matrix is ​​used to linearly combine the current hidden state and the comprehensive feature vector. Through the above steps, 1D-CNN can effectively extract temporal features from the LSTM output. t ′ represents the new hidden state.

5. A bearing life prediction method integrating temporal pattern mechanism and neural network as described in claim 1 or 2, characterized in that, The process of the third step is as follows: Step (3.1) Generate feature weights: The TPA mechanism calculates the hidden state h t Feature map c extracted by convolution. k The correlation between them is calculated using the following formula: Where, α k For the k-th time step, the extracted feature map c k The weight of the contribution, score(h) LSTM ,c k ) is a correlation measure between the hidden state and the k-th feature map, and exp() is an exponential function used to ensure that all attention weights are positive; score(h t ′,c k )=h t ′c k This score measures the similarity between the current hidden state and the convolutional feature map. A high score indicates that the current hidden state is highly correlated with the convolutional feature map, while a low score indicates that they are not correlated. Step (3.2) generates the comprehensive feature vector: using the feature weights calculated in step (3.1) and the convolutional feature map c k The weighted summation yields the comprehensive eigenvector z, calculated using the following formula: The comprehensive feature vector z contains the features extracted by convolution, and the TPA mechanism enhances the memory and recognition of important features; Step (3.3) Input to the Dense Layer: The Dense Layer receives the comprehensive feature vector from the TPA mechanism. Through this layer, the model performs a linear transformation on these features and the learned weights. The Dense Layer calculates the final predicted value y through a fully connected neural network layer. pred The calculation formula is as follows: y pred =W1·z+b1 The output of the linear transformation is the predicted bearing life, which is a continuous value. Where W1 is the weight matrix of the Dense Layer, and b1 is the bias term of the Dense Layer.