Steel belt elevator service life prediction method and system based on multi-head attention and mixed gating

By using the prediction method of multi-head attention and hybrid strobe in steel belt elevators, the steel belt operation data is analyzed, and the problems of low prediction accuracy and inability to consider the diversity of operating conditions in the existing technology are solved, achieving high-precision life prediction and safety improvement.

CN120197348AInactive Publication Date: 2025-06-24HANGZHOU HUAJIAN INTELLIGENT TECHNOLOGY RESEARCH INSTITUTE CO LTD +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510230246.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When predicting the service life of steel belt elevators, the prediction accuracy is low, and the diversity of operating conditions cannot be fully considered, which affects the safety and maintenance costs of the elevator.

Method used

The service life prediction method of steel belt elevators based on multi-head attention and hybrid strobe is adopted. Through data preprocessing, feature weighting, attention calculation, feature extraction and life prediction, the steel belt operation data is comprehensively analyzed to achieve high-precision life prediction.

Benefits of technology

It improves the operating safety of steel belt elevators, reduces maintenance costs, and achieves high-precision life prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197348A_ABST
    Figure CN120197348A_ABST
Patent Text Reader

Abstract

A steel belt elevator service life prediction method based on multi-head attention and mixed gating comprises the following steps: 1) data preprocessing: acquiring vibration signals and load data in operation of a steel belt, performing normalization processing on the vibration signals and the load data, and segmenting the data into time sequence samples through a sliding window; 2) feature weighting: carrying out dynamic weighting on the normalized data through a gating module; 3) attention calculation: capturing correlation between time sequence features through an attention module; 4) feature extraction: performing deep feature extraction on the weighted features through a convolutional neural network module; and 5) life prediction: flattening the extracted features, inputting the flattened features into the full-connection network, and outputting a residual life prediction value of the steel strip. The invention further provides a system for predicting the service life of the steel belt elevator based on multi-head attention and mixed gating. According to the invention, high-precision life prediction can be realized; the operation safety of the elevator is improved; and the maintenance cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of steel belt elevators and relates to a method and system for predicting the service life of steel belt elevators. Background Art

[0002] As the core component of a steel belt elevator, the steel belt undertakes a key transmission function. However, due to the long-term high-load and high-intensity working environment of the steel belt, its wear and degradation will directly affect the safety and service life of the elevator. At present, the prediction of the service life of the steel belt mainly relies on empirical judgment or simple statistical methods, which have the disadvantages of low prediction accuracy and inability to fully consider the diversity of operating conditions. Therefore, developing a method that can accurately predict the service life of the steel belt is of great significance for improving the operating safety of the elevator and reducing maintenance costs. Summary of the Invention

[0003] In order to overcome the deficiencies of the prior art, the present invention provides a method and system for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating, which can effectively combine the gating method and attention mechanism in deep learning, comprehensively analyze the steel belt operation data, and achieve high-precision life prediction; improve the operating safety of the elevator and reduce maintenance costs.

[0004] The technical solution adopted by the present invention to solve its technical problems is:

[0005] A method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating, the method comprising the following steps:

[0006] 1) Data preprocessing: Obtain the vibration signal and load data during the operation of the steel belt, perform normalization processing on them, and divide them into time series samples through a sliding window;

[0007] 2) Feature weighting: Dynamically weight the normalized data through a gating module;

[0008] 3) Attention calculation: Capture the correlation between time series features through an attention module;

[0009] 4) Feature extraction: Perform deep feature extraction on the weighted features through a convolutional neural network module;

[0010] 5) Life prediction: Flatten the extracted features and input them into a fully connected network to output the predicted remaining life value of the steel belt.

[0011] Further, the process of step 1) is:

[0012] 1.1) Data acquisition: Use an acceleration sensor and a strain gauge sensor to collect vibration signals and load signals during the operation of the steel strip. The acceleration sensor is installed at key parts of the steel strip to collect high-frequency vibration signals. The strain gauge sensor is installed in a specific area on the surface of the steel strip to monitor the tension change of the steel strip.

[0013] 1.2) Data preprocessing: Denoise the collected signals. Use a band-pass filter to filter out signals in irrelevant frequency bands and only retain the effective frequency band. Use wavelet transform to eliminate environmental noise. Normalize the signals to a unified numerical range [0,1]. The formula is:

[0014]

[0015] where, X norm is the normalized data, X is the collected effective signal data, min(X) is the minimum value of the channel within the time window, and max(X) is the maximum value of the channel within the time window;

[0016] Extract the time-domain features of the signal, including mean value, standard deviation, root mean square value, peak value, skewness, and kurtosis. Extract the frequency-domain features of the signal, including the main frequency component and the bandwidth energy distribution;

[0017] 1.3) Sliding window segmentation: Divide the signal into time series samples with a fixed length W, and set a certain overlap rate between windows. The formula is:

[0018] X window ={X t , X t+1 ,..., X t+W-1}

[0019] where, W is the window length, X window is the time series sample after sliding window segmentation, X t is the t-th data point in the time series, serving as the starting position of the current window, X t+1 ,..., X t+W-1 are the subsequent data points following the starting point X t , a total of (W - 1) data points, which together with X t form a continuous sequence with a length of W;

[0020] 1.4) Signal splicing: Align the acceleration signal and the strain signal in the time dimension and splice them to form a multi-dimensional input tensor.

[0021] Preferably, in step 1), the acceleration sensor is installed on the upper and lower sides where the steel strip driving wheel contacts the steel strip to capture high-frequency vibration characteristics, and the strain gauge sensor is installed in the high-tension area of the steel strip to monitor the load change.

[0022] In the above 1.2), the noise reduction process includes: using a band - pass filter to retain the effective frequency band signals from 20 Hz to 2000 Hz; performing multi - scale decomposition on the signals through wavelet transform to eliminate environmental noise.

[0023] In the above 1.3), the window length of the sliding window segmentation is 1 second, the corresponding number of sampling points is 10000, and the window overlap rate is 50%.

[0024] Furthermore, the process of step 2) is as follows:

[0025] 2.1) Perform global average pooling on the input tensor to calculate the average value of each channel:

[0026]

[0027] where G c is the global average value of channel c, X t,c is the input value of channel c at time step t, and T is the total number of time steps (Time steps), which is determined by the length of the time series after sliding window segmentation;

[0028] 2.2) Use a two - layer fully - connected network to calculate the feature weights:

[0029] W = Sigmoid(W2·ReLU(W1·G))

[0030] where W is the generated dynamic weight matrix (channel attention weight), W1 and W2 are learnable weight matrices, G is the pooled input vector, and ReLU(·) is an activation function that performs a non - linear transformation on the intermediate features;

[0031] 2.3) Multiply the weight W element - by - element with the input tensor to obtain the weighted features.

[0032] The gating module includes: a global average pooling layer for calculating the average value of each channel of the input tensor and outputting a dimension of a first fully - connected layer for mapping the tensor from dimension to where r is the compression ratio; a second fully - connected layer for mapping the tensor from dimension back to a Sigmoid activation function for normalizing the features output by the fully - connected layer to generate dynamic weighting coefficients.

[0033] Even further, the process of step 3) is as follows:

[0034] 3.1) Use linear transformation to generate query Q, key K, and value V:

[0035] Q = XW Q , K = XW K , V = XW V

[0036] where W Q , W K , W V are learnable parameters, and X is the feature tensor input to the attention module;

[0037] 3.2) Calculate the attention score using the following formula:

[0038]

[0039] where d k is the dimension of the key vector;

[0040] 3.3) Process the attention output through a residual connection and a normalization layer to obtain the enhanced feature.

[0041] The multi-head attention mechanism of the attention module includes: multiple parallel attention heads, and the calculation formula for each attention head is:

[0042]

[0043] where Q i , K i , V i are the query, key, and value of the i-th attention head respectively;

[0044] A concatenation layer for concatenating the outputs of all attention heads:

[0045] Concat([head1,…,head h )

[0046] where h is the number of attention heads;

[0047] A projection layer for mapping the concatenated tensor back to the original feature space, and the formula is:

[0048] MultiHead(Q,K,V) = Concat([head1,…,head h )W O

[0049] where W O is the output weight matrix.

[0050] The process of step 4) is as follows:

[0051] 4.1) Pass through multiple one-dimensional convolutional layers in sequence, and the kernel sizes are 7, 5, and 3 respectively;

[0052] 4.2) After each convolution, a ReLU activation function and a batch normalization layer are connected;

[0053] 4.3) Reduce the time dimension through a max pooling layer to extract high-level features:

[0054] Output = MaxPooling(BatchNorm(ReLU(Conv1D(Input))))

[0055] Among them, Output is the final output tensor, with dimensions (B, C″, T″′), C″ is the number of channels after multiple convolutions, and T″′ is the number of time steps after convolution and pooling;

[0056] MaxPooling is the max pooling operation, and the output dimension is reduced to (B, C′, T″), where T″ = T′ / s, T′ is the number of time steps after convolution, and s is the stride;

[0057] BatchNorm is the batch normalization layer, Conv1D is the one-dimensional convolution operation, Input is the input tensor, with dimensions (B, C, T), B is the batch size (the number of samples processed at one time), C is the number of input channels, and T is the number of time steps (determined by the sliding window length);

[0058] The convolution kernel size of the convolutional neural network module is dynamically adjusted according to the sampling frequency of the time series to extract features in different frequency ranges; the input tensor dimensions are For the three-layer one-dimensional convolution, convolution kernels with sizes of 7, 5, and 3 are used respectively, and the corresponding strides are 2. The dimensions of the output feature tensors of each layer are gradually reduced; the max pooling layer is used to further reduce the time dimension, and the final output tensor dimensions are Among them, C′ is the number of output channels, and T′ is the result of reducing the number of time steps.

[0059] In step 5), a multi-layer fully connected network is used to perform a regression operation on the input features, and the output calculation formula is:

[0060] Output = W n ·ReLU(W n-1 ·…·Input + b n-1 ) + b n

[0061] Among them, Output is the predicted remaining life value of the steel strip, with dimensions (B, 1), representing the predicted life of each sample in the batch; b n is the bias scalar of the nth fully connected layer, W n is the weight matrix of the nth (last) fully connected layer, b n-1 is the bias vector of the (n - 1)th fully connected layer: W n-1is the weight matrix of the (n - 1)-th fully connected layer; Input is the input feature vector with dimension (B, F), where B is the batch size and F is the feature dimension.

[0062] A steel belt elevator service life prediction system based on multi-head attention and hybrid gating, comprising:

[0063] A data preprocessing unit for loading steel belt operation data, performing normalization processing and time series segmentation;

[0064] A feature weighting unit for dynamically weighting the input data through a gating module;

[0065] An attention calculation unit for capturing the cross-time correlation of time series features based on the multi-head attention mechanism;

[0066] A convolution extraction unit for extracting high-level features through a convolutional neural network module;

[0067] A fully connected prediction unit for outputting the predicted remaining life value of the steel belt through a fully connected network.

[0068] The beneficial effects of the present invention are mainly manifested in: being able to effectively combine the gating method and the attention mechanism in deep learning, comprehensively analyzing the steel belt operation data, achieving high-precision life prediction; improving the operation safety of the elevator and reducing the maintenance cost. Description of the Drawings

[0069] Figure 1 is a flowchart of a steel belt elevator service life prediction method based on multi-head attention and hybrid gating. Detailed Embodiments

[0070] The present invention will be further described below with reference to the drawings.

[0071] Referring to Figure 1 , a steel belt elevator service life prediction method based on multi-head attention and hybrid gating, the method comprises the following steps:

[0072] 1) Data preprocessing: Obtain the vibration signal and load data during the operation of the steel belt, perform normalization processing on them, and segment them into time series samples through a sliding window;

[0073] The process of step 1) is as follows:

[0074] 1.1) Data acquisition: Use an acceleration sensor and a strain gauge sensor to collect the vibration signal and load signal during the operation of the steel belt; the acceleration sensor is installed at a key part of the steel belt for collecting high-frequency vibration signals; the strain gauge sensor is installed in a specific area on the surface of the steel belt for monitoring the tension change of the steel belt;

[0075] 1.2) Data preprocessing: Denoise the collected signals. Use a band-pass filter to filter out signals in irrelevant frequency bands and only retain the effective frequency band; use wavelet transform to eliminate environmental noise; normalize the signals to a unified numerical range [0,1]. The formula is:

[0076]

[0077] where X norm is the normalized data, X is the collected effective signal data, min(X) is the minimum value of the channel within the time window, and max(X) is the maximum value of the channel within the time window;

[0078] Extract the time-domain features of the signal, including mean, standard deviation, root mean square value, peak value, skewness, and kurtosis; extract the frequency-domain features of the signal, including the main frequency component and bandwidth energy distribution;

[0079] 1.3) Sliding window segmentation: Segment the signal into time series samples with a fixed length W, and set a certain overlap rate between windows. The formula is:

[0080] X window ={X t ,X t+1 ,...,X t+W-1}

[0081] where W is the window length, X window is the time series sample after sliding window segmentation, X t is the t-th data point in the time series, serving as the starting position of the current window, and X t+1 ,...,X t+w-1 are the subsequent data points following the starting point X t , a total of (W - 1) data points, which together with X t form a continuous sequence of length W.

[0082] 1.4) Signal splicing: Align the acceleration signal and the strain signal in the time dimension and splice them to form a multi-dimensional input tensor.

[0083] Preferably, in step 1), the acceleration sensors are installed on the upper and lower sides where the steel belt driving wheel contacts the steel belt to capture high-frequency vibration characteristics, and the strain gauge sensors are installed in the high-tension area of the steel belt to monitor load changes.

[0084] In 1.2), the denoising process includes: using a band-pass filter to retain the effective frequency band signals from 20Hz to 2000Hz; performing multi-scale decomposition on the signals through wavelet transform to eliminate environmental noise.

[0085] In the above 1.3), the window length of the sliding window segmentation is 1 second, the corresponding number of sampling points is 10,000, and the window overlap rate is 50%.

[0086] 2) Feature weighting: Dynamically weight the normalized data through a gating module;

[0087] The process of step 2) is as follows:

[0088] 2.1) Perform global average pooling on the input tensor to calculate the average value of each channel:

[0089]

[0090] Among them, G c is the global average value of channel c, X t,c is the input value of channel c at time step t, and T is the total number of time steps (Time steps), which is determined by the length of the time series after sliding window segmentation;

[0091] 2.2) Use a two-layer fully connected network to calculate the feature weights:

[0092] W = Sigmoid(W2·ReLU(W1·G))

[0093] Among them, W is the generated dynamic weight matrix (channel attention weight), W1 and W2 are learnable weight matrices, G is the input vector after pooling, ReLU(·) is an activation function, and performs a non-linear transformation on the intermediate features;

[0094] 2.3) Multiply the weight W element-wise with the input tensor to obtain the weighted features.

[0095] The gating module includes: a global average pooling layer for calculating the average value of each channel of the input tensor and outputting a dimension of a first fully connected layer for mapping the tensor from dimension to where r is the compression ratio; a second fully connected layer for mapping the tensor from dimension back to a Sigmoid activation function for normalizing the features output by the fully connected layer to generate dynamic weighting coefficients.

[0096] 3) Attention calculation: Capture the correlation between time series features through an attention module;

[0097] The process of step 3) is as follows:

[0098] 3.1) Use linear transformation to generate query Q, key K, and value V:

[0099] Q = XW Q , K = XW K , V = XW V

[0100] where W Q , W K , W V are learnable parameters, and X is the feature tensor input to the attention module;

[0101] 3.2) Calculate the attention score using the following formula:

[0102]

[0103] where d k is the dimension of the key vector;

[0104] 3.3) Process the attention output through a residual connection and a normalization layer to obtain enhanced features.

[0105] The multi-head attention mechanism of the attention module includes: multiple parallel attention heads, and the calculation formula for each attention head is:

[0106]

[0107] where Q i , K i , V i are the query, key, and value of the i-th attention head respectively;

[0108] A concatenation layer for concatenating the outputs of all attention heads:

[0109] Concat([head1,…,head h )

[0110] where h is the number of attention heads;

[0111] A projection layer for mapping the concatenated tensor back to the original feature space, and the formula is:

[0112] MultiHead(Q,K,V) = Concat([head1,…,head h )W O

[0113] where W O is the output weight matrix.

[0114] 4) Feature extraction: Perform deep feature extraction on the weighted features through a convolutional neural network module;

[0115] The process of step 4) is:

[0116] 4.1) Pass through multiple one-dimensional convolutional layers in sequence, with convolutional kernel sizes of 7, 5, and 3 respectively;

[0117] 4.2) After each convolution, connect a ReLU activation function and a batch normalization layer;

[0118] 4.3) Reduce the time dimension through a max pooling layer to extract high-level features:

[0119] Output = MaxPooling(BatchNorm(ReLU(Conv1D(Input))))

[0120] Where Output is the final output tensor, with dimensions (B, C″, T″′), C″ is the number of channels after multiple convolutions, and T″′ is the number of time steps after convolution and pooling;

[0121] MaxPooling is the max pooling operation, and the output dimension is reduced to (B, C′, T″), where T″ = T′ / s, T′ is the number of time steps after convolution, and s is the stride;

[0122] BatchNorm is the batch normalization layer, Conv1D is the one-dimensional convolution operation, Input is the input tensor, with dimensions (B, C, T), B is the batch size (the number of samples processed at one time), C is the number of input channels, and T is the number of time steps (determined by the sliding window length);

[0123] The convolutional kernel size of the convolutional neural network module is dynamically adjusted according to the sampling frequency of the time series to extract features in different frequency ranges; the input tensor dimensions are The three one-dimensional convolutions use convolutional kernel sizes of 7, 5, and 3 respectively, with corresponding strides of 2, and the output feature tensor dimensions of each layer gradually decrease; the max pooling layer is used to further reduce the time dimension, and the final output tensor dimensions are Where C′ is the number of output channels and T′ is the result of reducing the number of time steps.

[0124] 5) Life prediction: Flatten the extracted features and input them into a fully connected network to output the predicted remaining life value of the steel strip.

[0125] In step 5) above, a multi-layer fully connected network is used to perform a regression operation on the input features, and the output calculation formula is:

[0126] Output = W n ·ReLU(W n-1 ·…·Input + b n-1 ) + b n

[0127] Among them, Output is the predicted remaining life value of the steel strip, with a dimension of (B, 1), representing the predicted life of each sample in the batch; b n is the bias scalar of the nth fully connected layer, and W n is the weight matrix of the nth (last) fully connected layer, and b n-1 is the bias vector of the (n - 1)th fully connected layer: W n-1 is the weight matrix of the (n - 1)th fully connected layer; Input is the input feature vector, with a dimension of (B, F), where B is the batch size and F is the feature dimension.

[0128] In this embodiment, sensor arrangement and data acquisition: At least two acceleration sensors and one strain gauge sensor are arranged on each steel strip, and vibration and load signals during the operation of the steel strip are collected through the acceleration sensors and the strain gauge sensor;

[0129] The acceleration sensors are installed at key positions of the steel strip (such as the contact position between the driving wheel and the steel strip) to collect high-frequency vibration signals during the operation of the steel strip, and are arranged on the upper and lower sides to ensure capturing the complete dynamic response.

[0130] The strain gauge sensor is installed in a specific area on the surface of the steel strip to be used for real-time monitoring of the tension and load changes of the steel strip, and is usually arranged in the high-tension area.

[0131] The signals of the acceleration sensors and the strain gauge sensors are recorded by a high-speed data acquisition system, and the sampling frequency is set to 10 kHz to ensure capturing high-frequency vibration characteristics.

[0132] The collected original signals are as follows:

[0133] Acceleration signal: The collected vibration signal is in the form of a time series, with the unit of m / s 2 , recording the vibration amplitude during the operation of the steel strip;

[0134] Strain signal: Reflecting the load change of the steel strip, with the unit of strain με (microstrain), which is closely related to the life of the steel strip.

[0135] Original signal processing flow: In order to eliminate the noise and redundant information in the original signal and make it suitable for model input, the following preprocessing steps are adopted:

[0136] 1.1. Noise reduction processing, the process is as follows:

[0137] 1.1.1 Filter out signals in irrelevant frequency bands through a band-pass filter, and only retain the effective frequency bands related to the operating state of the steel strip, such as 20 - 2000 Hz;

[0138] 1.1.2 Use wavelet transform to eliminate the interference of environmental noise on the vibration signal;

[0139] 1.2. Normalization is performed as follows:

[0140] 1.2.1 For the dimensional differences of different sensor signals, unified normalization is performed to normalize their amplitude ranges to [0, 1]:

[0141]

[0142] where X norm is the normalized data, X is the collected valid signal data, min(X) is the minimum value of the channel within the time window, and max(X) is the maximum value of the channel within the time window;

[0143] 1.3. Feature extraction is performed as follows:

[0144] 1.3.1 Time-domain feature extraction: Calculate the mean, standard deviation, root mean square value, peak value, skewness, kurtosis, etc. of the vibration signal;

[0145] 1.3.2 Frequency-domain feature extraction: Extract features such as the main frequency component and bandwidth energy distribution of the signal through fast Fourier transform (FFT);

[0146] 1.4. Sliding window segmentation is performed as follows:

[0147] 1.4.1 The signal is segmented according to the set window length (e.g., 1 second, corresponding to 10,000 sampling points) to form time series samples:

[0148] X window ={X t , X t+1 ,..., X t+W-1}

[0149] where W is the window length, and a certain overlap (e.g., 50%) is allowed between windows to increase the data volume; X window is the time series sample after sliding window segmentation, X t is the t-th data point in the time series, serving as the starting position of the current window, X t+1 ,..., X t+w-1 are the subsequent data points following the starting point X t , a total of (W - 1) points, which together with X t form a continuous sequence of length W;

[0150] 1.5. Time series splicing: Align and splice the acceleration signal and the strain signal in the time dimension to form a multi-dimensional input tensor. The tensor dimension is (B, T, C), where B is the batch size (the number of samples input each time), T is the number of time steps (determined by the window length), and C is the number of channels (including the acceleration signal and the strain signal).

[0151] A steel belt elevator service life prediction system based on multi-head attention and hybrid gating, comprising:

[0152] Data preprocessing unit, used to load the steel strip running data, perform normalization processing and time series segmentation;

[0153] A feature weighting unit, used to dynamically weight input data through a gating module;

[0154] Attention calculation unit, used to capture the cross-temporal correlation of time series features based on the multi-head attention mechanism;

[0155] Convolutional extraction unit, used to extract high-level features through convolutional neural network modules;

[0156] The fully connected prediction unit is used to output the remaining life prediction value of the steel strip through a fully connected network.

[0157] The gating module of this embodiment: "Gating" in the present invention refers to a dynamic weighting mechanism. The dynamic channel feature weighting mechanism is a dynamic weighting method for enhancing key features and suppressing redundant information. This mechanism globally models the channel dimension of the input feature, generates a dynamic weight matrix, and realizes adaptive weighting of the importance of features (realizing feature weighted processing of steel strip running data); the attention calculation unit is based on a multi-head attention mechanism to capture the cross-time correlation in the steel strip running data.

[0158] The convolutional neural network module performs multi-layer convolution operations on the weighted feature data to extract high-level features; it includes three layers of one-dimensional convolution, and each layer of convolution uses the ReLU activation function and batch normalization operation; the feature dimension is reduced through the maximum pooling layer to reduce the computational complexity. The fully connected network performs classification and regression operations on the extracted features and outputs the remaining life prediction results of the steel strip; the LeakyReLU activation function is used; the last layer outputs a scalar value as the predicted life of the steel strip.

[0159] The contents described in the embodiments of this specification are merely enumerations of implementation forms of the inventive concept and are for illustrative purposes only. The protection scope of the present invention should not be considered to be limited to the specific forms described in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be thought of by ordinary technicians in this field based on the inventive concept.

Claims

1. A method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating, characterized in that: The method comprises the following steps: 1) Data preprocessing: Obtain the vibration signal and load data of the steel belt during operation, normalize them, and divide them into time series samples through a sliding window; 2) Feature weighting: Dynamically weight the normalized data through the gating module; 3) Attention calculation: Capture the correlation between time series features through the attention module; 4) Feature extraction: deep feature extraction of weighted features through convolutional neural network module; 5) Life prediction: The extracted features are flattened and input into the fully connected network to output the remaining life prediction value of the steel strip.

2. The method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating according to claim 1, characterized in that: The process of step 1) is: 1.1) Data acquisition: Use acceleration sensors and strain gauge sensors to collect vibration signals and load signals during the operation of the steel belt; Acceleration sensors are installed at key locations on the steel belt to collect high-frequency vibration signals; strain gauge sensors are installed at specific areas on the surface of the steel belt to monitor changes in the tension of the steel belt; 1.2) Data preprocessing: The collected signals are subjected to noise reduction processing, and irrelevant frequency band signals are filtered out using a bandpass filter, retaining only the effective frequency band; environmental noise is eliminated using wavelet transform; the signals are normalized to a unified numerical range [0,1], and the formula is: Among them, X norm is the normalized data, X is the collected valid signal data, min(X) is the minimum value of the channel in the time window, and max(X) is the maximum value of the channel in the time window; Extract the time domain characteristics of the signal, including mean, standard deviation, RMS value, peak, skewness and kurtosis; extract the frequency domain characteristics of the signal, including the main frequency component and bandwidth energy distribution; 1.3) Sliding window segmentation: The signal is divided into time series samples according to a fixed length W, and a certain overlap rate is set between windows. The formula is: X window ={X t ,X t+1 ,...,X t+w-1 } Where W is the window length, X window is the time series sample after the sliding window segmentation, X t is the tth data point in the time series, serving as the starting position of the current window, X t+1 ,...,X t+W-1 Following the starting point X t The subsequent data points, a total of (W-1), are t Form a continuous sequence of length W; 1.4) Signal splicing: Align the acceleration signal and strain signal in the time dimension and splice them to form a multi-dimensional input tensor.

3. The method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating according to claim 2, characterized in that: In the step 1), the acceleration sensor is installed on the upper and lower sides of the steel belt driving wheel in contact with the steel belt to capture high-frequency vibration characteristics, and the strain gauge sensor is installed in the high-tension area of ​​the steel belt to monitor load changes; In the above 1.2), the noise reduction process includes: using a bandpass filter to retain the effective frequency band signal of 20 Hz to 2000 Hz; performing multi-scale decomposition of the signal by wavelet transform to eliminate environmental noise; In 1.3), the window length of the sliding window segmentation is 1 second, the corresponding number of sampling points is 10,000, and the window overlap rate is 50%.

4. The method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating according to any one of claims 1 to 3, characterized in that: The process of step 2) is: 2.1) Perform global average pooling on the input tensor and calculate the average value of each channel: Among them, G c is the global average value of channel c, X t,c is the input value of channel c at time step t, and T is the total number of time steps; 2.2) Use a two-layer fully connected network to calculate feature weights: W = Sigmoid(W2·ReLU(W1·G)) Among them, W is the generated dynamic weight matrix, W1 and W2 are learnable weight matrices, G is the input vector after pooling, and ReLU(·) is the activation function, which performs nonlinear transformation on the intermediate features; 2.3) Multiply the weight W by the input tensor element by element to obtain the weighted feature.

5. The method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating according to claim 4, characterized in that: The gating module includes: a global average pooling layer for inputting tensors The average value is calculated for each channel of The first fully connected layer is used to transform the tensor from dimension Map to Where r is the compression ratio; the second fully connected layer is used to transform the tensor from dimension Mapping back Sigmoid activation function is used to normalize the features of the fully connected layer output and generate dynamic weighting coefficients.

6. The method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating according to any one of claims 1 to 4, characterized in that: The process of step 3) is: 3.1) Generate query Q, key K and value V using linear transformation: Q=XW Q ,K=XW K ,V=XW V Among them, W Q , W K , W V is a learnable parameter, X is the feature tensor input to the attention module; 3.2) Calculate the attention score using the following formula: Among them, d k is the dimension of the key vector; 3.3) The attention output is processed through residual connection and normalization layer to obtain enhanced features.

7. The method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating according to claim 6, characterized in that: The multi-head attention mechanism of the attention module includes: multiple parallel attention heads, and the calculation formula of each attention head is: Among them, Q i , K i 、V i are the query, key, and value of the i-th attention head respectively; The concatenation layer is used to concatenate the outputs of all attention heads: Concat([head1,…,head h ]) Where h is the number of attention heads; The projection layer is used to map the concatenated tensor back to the original feature space. The formula is: MultiHead(Q,K,V)=Concat([head1,…,head h ])W O Among them, W O is the output weight matrix.

8. The method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating according to any one of claims 1 to 4, characterized in that: The process of step 4) is: 4.1) Pass through multiple one-dimensional convolutional layers in sequence, with convolution kernel sizes of 7, 5, and 3 respectively; 4.2) Each convolution is followed by a ReLU activation function and a batch normalization layer; 4.3) Reduce the time dimension through the maximum pooling layer and extract high-level features: Output=MaxPooling(BatchNorm(ReLU(Conv1D(Input)))) Among them, Output is the final output tensor, the dimension is (B,C ″ ,T ″′ ),C ″ is the number of channels after multi-layer convolution, T″′ is the number of time steps after convolution and pooling; MaxPooling is the maximum pooling operation, and the output dimension is reduced to (B, C′, T″), where T″=T′ / s, T′ is the time step after convolution, and s is the stride; BatchNorm is the batch normalization layer, Conv1D is the one-dimensional convolution operation, Input is the input tensor, the dimension is (B, C, T), B is the batch size, C is the number of input channels, and T is the number of time steps; The convolution kernel size of the convolutional neural network module is dynamically adjusted according to the sampling frequency of the time series to extract features in different frequency ranges; the input tensor dimension is The three layers of one-dimensional convolution use convolution kernel sizes of 7, 5, and 3 respectively, with a corresponding stride of 2. The dimension of the output feature tensor of each layer is gradually reduced; the maximum pooling layer is used to further reduce the time dimension, and the final output tensor dimension is Where C′ is the number of output channels and T′ is the dimension reduction result of the number of time steps.

9. The method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating according to any one of claims 1 to 4, characterized in that: In step 5), a multi-layer fully connected network is used to perform regression operation on the input features, and the output calculation formula is: Output=W n ·ReLU(W n-1 ·…·Input+b n-1 )+b n Among them, Output is the predicted value of the remaining life of the steel strip, and its dimension is (B,1), which represents the predicted life of each sample in the batch; b n is the bias scalar of the nth fully connected layer, W n is the weight matrix of the nth fully connected layer, b n-1 is the bias vector of the n-1th fully connected layer: W n-1 is the weight matrix of the n-1th fully connected layer; Input is the input feature vector with dimension (B, F), where B is the batch size and F is the feature dimension.

10. A system for implementing the method for predicting the service life of a steel belt elevator based on multi-head attention and hybrid gating as claimed in claim 1, characterized in that: The system comprises: Data preprocessing unit, used to load the steel strip running data, perform normalization processing and time series segmentation; A feature weighting unit, used to dynamically weight input data through a gating module; Attention calculation unit, used to capture the cross-temporal correlation of time series features based on the multi-head attention mechanism; Convolutional extraction unit, used to extract high-level features through convolutional neural network modules; The fully connected prediction unit is used to output the remaining life prediction value of the steel strip through a fully connected network.

Citation Information

Cited By

  • Turnout multi-sensor data feature extraction method and device

    CN121561377A