Old elevator fault diagnosis method based on GCN-BiLSTM
By using the GCN-BiLSTM fusion network, a multi-source heterogeneous dataset was constructed and spatial-temporal feature extraction and fusion were performed. This solved the problem of difficulty in identifying early wear of old elevator bearings, achieved high-precision fault diagnosis and early warning, and improved elevator safety.
Patent Information
- Application Number
- CN202510650406.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies make it difficult to effectively identify early wear and failures in old elevator bearings. Traditional detection methods cannot meet the needs of accurate diagnosis, especially for the problems of lubrication degradation and fitting clearance changes that are unique to old elevators.
A method based on the fusion of graph convolutional neural networks (GCN) and bidirectional long short-term memory networks (BiLSTM) is adopted to construct multi-source heterogeneous datasets, perform spatial feature extraction and time series modeling, and combine graph structure models with multimodal feature fusion to achieve accurate identification and early warning of old elevator failures.
It has achieved high-precision identification and early warning of bearing failures in old elevators, improved the accuracy of fault diagnosis, significantly improved the ability to identify problems such as bearing wear, and reduced safety hazards.
Smart Images

Figure CN120744709A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of elevator fault diagnosis and predictive maintenance, and specifically relates to a fault diagnosis method for old elevators based on a graph convolutional neural network (GCN) and a bidirectional long short-term memory network (BiLSTM). Background Art
[0002] With the acceleration of urbanization in my country, the total number of elevators in use nationwide will exceed 11 million by 2024, of which 900,000 are old elevators with a service life of more than 15 years. These elevators commonly suffer from typical problems such as severe bearing wear, increased clearances between mechanical components, and aging electrical systems, resulting in significantly higher failure rates than newer elevators. Bearings, as core components of elevator traction systems, suffer from wear and deterioration, leading to increased vibration and noise, and decreased operational smoothness. In recent years, numerous safety incidents involving people trapped in elevators, sudden elevator stops, and falls have been reported in older elevators, with bearing failure being the primary cause. Current elevator bearing monitoring suffers from significant technical shortcomings, primarily a lack of real-time vibration signal acquisition capabilities, making it impossible to obtain critical dynamic data on the equipment's operating status, making it difficult to identify early signs of wear. Furthermore, relying on manual inspections and static parameter testing cannot effectively capture progressive deterioration characteristics such as increased bearing clearance and raceway spalling, nor can it assess the impact of these faults on associated components such as the bearing, gearbox, and traction sheave. In particular, for nonlinear degradation problems such as lubrication degradation and changes in fitting clearance that are unique to old elevators, traditional discrete detection methods are completely unable to meet the needs of accurate diagnosis.
[0003] Therefore, a fault diagnosis method for old elevators based on GCN-BiLSTM is proposed. By integrating the spatial feature extraction capability of graph convolutional networks and the temporal modeling advantages of bidirectional long short-term memory networks, accurate fault identification and early warning of old elevators can be achieved. Summary of the Invention
[0004] In view of this, the present invention provides a fault diagnosis method for old elevators based on the fusion of graph convolutional networks and bidirectional long short-term memory networks to solve the problems raised in the above background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] The GCN-BiLSTM-based fault diagnosis method for old elevator bearings includes the following steps:
[0007] Step 1: Construct a multi-source heterogeneous dataset and synchronously collect vibration signals, acoustic emission signals, motor current signals, and bearing temperature signals through the elevator IoT monitoring system. The vibration signal sampling frequency is no less than 12.8 kHz, and the acoustic emission signal acquisition bandwidth is set to 40-100 kHz.
[0008] Pre-process the raw data, use a hardware trigger mechanism to ensure time domain synchronization of multi-source signals, and eliminate abnormal data based on the 3σ criterion to improve data reliability;
[0009] Furthermore, in order to characterize the fault transmission relationship between the various components of the elevator system, a graph structure model based on physical topology is constructed. The key components of the bearing, gearbox, and traction wheel are abstracted as graph nodes. The adjacency matrix is established based on the fault propagation path, and the node features are fused with time-frequency domain indicators to form space-time correlation features.
[0010] Step 2: Design a dedicated graph convolutional network for spatial feature extraction based on the constructed graph structure data;
[0011] First, a graph structure data based on physical connection relationships is constructed, in which the key components of bearings, gearboxes, and traction sheaves are used as graph nodes. An adjacency matrix is established based on the actual mechanical connections and fault propagation paths.
[0012] Considering the differences in coupling between components in old elevators, a learnable weight parameter is introduced into the adjacency matrix, enabling the network to adaptively adjust the connection strength between nodes.
[0013] In terms of network structure design, a three-layer graph convolution layer is used for spatial feature extraction. Each layer of graph convolution operation includes two key processes: node feature transformation and neighborhood information aggregation. In the feature transformation stage, the original node features are linearly mapped through a trainable weight matrix.
[0014] In the information aggregation stage, an improved graph attention mechanism is used to calculate the contribution weights of neighboring nodes, thereby focusing on abnormal vibration conduction paths;
[0015] To address the subtle features of early bearing wear, a residual connection structure is embedded in the last layer of GCN, effectively alleviating the feature degradation problem of deep networks.
[0016] To enhance the model's ability to extract local fault features, an edge feature modeling mechanism is introduced into GCN. Mechanical properties such as the physical distance between components and connection stiffness are encoded as edge features, which participate in graph convolution operations together with node features.
[0017] Considering the widespread noise interference in old elevator monitoring data, the graph-based DropEdge regularization method is used during network training. This method randomly blocks some edge connections, effectively preventing overfitting and enhancing the model's robustness to incomplete topological information.
[0018] Step 3: Construct a multi-level bidirectional LSTM network architecture for temporal feature extraction, and reorganize the spatial feature sequence output by GCN according to the time dimension as the network input;
[0019] In view of the non-stationary characteristics of the fault signal of old elevator bearings, a time-frequency transformation module is added before the input layer. The original vibration signal is converted into a time-frequency image through continuous wavelet transform to enhance the recognition of periodic fault characteristics.
[0020] A three-layer stacked BiLSTM architecture is used for time series feature extraction. Each layer contains two independent LSTM branches, forward and backward, which capture the causality and correlation of fault features through bidirectional information flow.
[0021] In view of the unique impulse response characteristics of bearing faults, a gated expansion mechanism is introduced into the LSTM unit to effectively capture fault pulses of different durations by adjusting the time step interval.
[0022] A dual attention mechanism module is embedded in the network. This includes a temporal attention layer that calculates the feature importance weights at each time step, and a feature attention layer that evaluates the contributions of different sensor channels. This enables the network to adaptively focus on time series segments and signal channels with significant fault characteristics.
[0023] Taking into account the variability of operating conditions of old elevators, a dynamic curriculum learning strategy is introduced during the network training phase. The training data distribution is gradually adjusted from easy to difficult by evaluating the sample complexity.
[0024] The high-dimensional feature vector finally outputted interacts with the spatial topological features extracted by GCN across modalities in the feature fusion layer to form a comprehensive fault representation with both spatial correlation and temporal dynamics.
[0025] Step 4: Construct a multimodal feature fusion and fault diagnosis model based on the attention mechanism. First, deeply couple the spatial topological features extracted by GCN with the temporal dynamic features extracted by BiLSTM.
[0026] A fusion layer based on the cross-attention mechanism is designed to achieve dynamic interaction and complementary enhancement of the two modal features by calculating the correlation weights between spatial features and temporal features.
[0027] For the three typical fault modes of older elevators—bearing wear, gearbox tooth breakage, and traction sheave groove wear—a multi-scale feature extraction module is added after the feature fusion layer. This module uses convolution kernels of 1×3, 1×5, and 1×7 sizes to concurrently extract fault features at different scales. The 1×3 convolution kernel is used to capture the high-frequency impact signal of gear tooth breakage, the 1×5 convolution kernel targets the mid-frequency harmonic characteristics of bearing wear, and the 1×7 convolution kernel captures the long-term trend characteristics of traction sheave groove wear.
[0028] The fused multi-scale features are mapped to the fault category space using the Softmax classification function, and the probability distribution of three types of faults, namely bearing wear, gearbox tooth breakage, and traction sheave groove wear, is output. When the highest probability exceeds the set threshold, it is determined to be the corresponding fault type.
[0029] In order to improve the diagnostic performance of the model under sample imbalance conditions, an improved weighted focal loss function is used in the network training stage, and its expression is:
[0030]
[0031] in is the category weight coefficient, and γ is the adjustment factor. Through dual adjustment, the model pays more attention to minority class fault samples and difficult-to-classify samples;
[0032] At the same time, a transfer learning strategy was introduced, and the parameters of the pre-trained GCN-BiLSTM model were fine-tuned to solve the problem of insufficient specific fault samples of old elevators and improve the diagnostic accuracy under small sample conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 Flowchart of the old elevator fault diagnosis method based on GCN-BiLSTM according to the present invention
[0034] Figure 2 This is the GCN-BiLSTM network structure model diagram described in the present invention DETAILED DESCRIPTION
[0035] The GCN-BiLSTM-based old elevator fault diagnosis method of the present invention is specifically implemented as follows: Figure 1 The technical solution of the present invention is described in detail below with reference to the accompanying drawings so that those skilled in the art can implement the invention with reference to the description.
[0036] Specifically, in step 1, the elevator IoT monitoring system synchronously collects vibration signals (sampling frequency ≥ 12.8kHz), acoustic emission signals (bandwidth 40-100kHz), motor current signals, and bearing temperature signals. To eliminate noise interference, a hardware trigger mechanism is used to ensure time domain synchronization of the multi-source signals, with synchronization error controlled within 1μs. The raw data is preprocessed using the 3σ criterion:
[0037]
[0038] where μ and σ are the signal mean and standard deviation, respectively.
[0039] Furthermore, the elevator mechanical system is abstracted into a graph structure, with key components (bearings, gearboxes, traction wheels) as nodes and mechanical connections and fault conduction paths as edges, to construct an adjacency matrix. ,Node features fuse time domain and frequency domain indicators.
[0040] The final graph representation , where the feature matrix It contains the D-dimensional features of each node, and the adjacency matrix A encodes the interaction relationship between components, preserving the local state and system-level fault propagation characteristics.
[0041] In step 2, a three-layer GCN network architecture is designed based on the multi-source heterogeneous graph data generated in step 1. Each layer of the network is processed in two stages: feature transformation and neighborhood aggregation. The core calculation process is expressed as:
[0042]
[0043] in is the normalized adjacency matrix, is the ReLU activation function, is the trainable parameter matrix.
[0044] In particular, the first layer compresses the original D-dimensional features into a low-dimensional space through a 64-dimensional weight matrix, achieving feature dimensionality reduction while retaining the core pattern of the vibration signal.
[0045] To further enhance the feature expression capability, the second layer introduces an improved graph attention mechanism (GAT), which dynamically calculates the weight coefficients between nodes through a 128-dimensional attention parameter vector:
[0046]
[0047] The design specifically targets the non-uniform coupling characteristics between old elevator components and embeds learnable parameters in the adjacency matrix. Adaptively adjust connection strength.
[0048] In the third layer network design, through the residual connection Maintaining feature integrity and using DropEdge technology to randomly mask edge connections with a probability of 0.3 enhances generalization; the final output 32-dimensional spatial feature vector contains both local fault modes and retains system-level topological relationships, providing optimized input for subsequent BiLSTM timing analysis.
[0049] Step 3 is mainly based on the spatial topological features output by GCN. In this step, a multi-level BiLSTM network is constructed to extract temporal dynamic features.
[0050] First, the 32-dimensional spatial feature sequence is reorganized according to the time window to form a time series input matrix , where T is the time step. To handle the non-stationary nature of the signal, a time-frequency transform module is placed before the input layer, and the periodic features are enhanced by Morlet wavelet transform:
[0051]
[0052] The system employs a three-layer stacked BiLSTM architecture, with each layer containing 128 hidden units. The forward LSTM branch captures the causal temporal relationships of fault features, while the backward LSTM branch models contextual dependencies. In terms of network design, an innovative gated expansion mechanism is introduced, dynamically adjusting the time step interval through learnable parameters to enhance the ability to capture transient impact features.
[0053] To optimize feature selection, the network embeds a dual attention mechanism: a temporal attention layer calculates the importance weight of each time step, while a feature attention layer evaluates the contribution of different sensor channels. These two layers are cascaded to achieve collaborative filtering, significantly improving the expressiveness of key features.
[0054] Finally, the adaptive weighting of spatiotemporal features is achieved through the feature fusion gating mechanism:
[0055]
[0056] in and Represent spatial and temporal features respectively, is a trainable weight matrix. The design ensures the integrity and complementarity of fault features.
[0057] In step 4, fault classification is achieved through multimodal feature fusion. Based on the spatiotemporal features extracted by GCN-BiLSTM, this step first uses the cross-attention mechanism to achieve feature fusion. By calculating the correlation matrix between spatial features and temporal features, a comprehensive fault representation with dynamic weights is generated. The core calculation process is expressed as:
[0058]
[0059] Where Q and K represent the feature matrices of GCN and BiLSTM respectively. is the feature dimension.
[0060] A parallel multi-scale convolution module was designed for five typical fault modes. Three convolution kernels (1×3, 1×5, and 1×7) were used to simultaneously extract local features. After feature concatenation and dimensionality reduction, these features were fed into a Softmax classifier. This design effectively captures the scale characteristics of different faults and significantly improves classification accuracy.
[0061] To deal with the problem of sample imbalance, an improved weighted focal loss function is adopted:
[0062]
[0063] in is the category weight coefficient, Used to adjust the weight distribution of difficult and easy samples. This function effectively alleviates the problem of majority class dominating training.
[0064] The network is initialized by pre-training model parameters, and a layered fine-tuning strategy is adopted under small sample conditions. Specifically, it includes three stages: fine-tuning the classifier with fixed bottom-level parameters, gradually unfreezing middle-level parameters, and global fine-tuning to ensure rapid model convergence and stability.
[0065] When training the GCN-BiLSTM hybrid neural network model, multi-source data from 100 older elevators was collected through an elevator IoT monitoring system to construct a dataset containing three typical fault types. To ensure balanced data distribution, stratified sampling was used to divide the training, validation, and test sets into a ratio of 8:1:1.
[0066] Notably, transfer learning techniques were employed during the pre-training phase. The model was pre-trained using a public bearing fault dataset, with an initial learning rate of 1e-4 set using the AdamW optimizer. This strategy effectively mitigated overfitting issues with small samples. Subsequently, during the fine-tuning phase, after loading the pre-trained weights, a learning rate scheduling strategy with cosine annealing was employed, and dynamic curriculum learning was introduced to improve model convergence stability.
[0067] To further improve model performance, we applied various regularization techniques during training. DropEdge randomly blocked 20% of edge connections, combined with label smoothing to prevent overfitting. An early stopping mechanism terminated training if the validation set loss did not decrease for 15 consecutive rounds. These measures collectively ensured the model's generalization capabilities.
[0068] The parameter configuration in this implementation is shown in Table 1.
[0069] Table 1 Model training parameter settings
[0070]
[0071] In this example, the GCN module uses a three-layer graph convolutional architecture with hidden layer dimensions set to [128, 64, 32] and a four-head attention mechanism. The BiLSTM module, on the other hand, is designed as a three-layer architecture with 128 hidden units. Multi-scale time series modeling is achieved through an expansion gating step size of [1, 3, 5]. Training uses the AdamW optimizer with the Weighted Focal Loss loss function, a batch size of 32, and fine-tuning for 200 epochs.
[0072] Experimental results demonstrate that the method in this embodiment demonstrates superior performance. As shown in Table 3, compared to the baseline model, the GCN-BiLSTM achieves an accuracy of 98.7%, a 23.6% improvement over the traditional SVM. In early bearing wear detection, the F1 score reaches 96.2%, significantly outperforming the LSTM (88.7%).
[0073] Table 2 Model performance comparison results
[0074]
[0075] Furthermore, this embodiment validated the contribution of each module through ablation experiments. Experimental data showed that removing the GCN module resulted in a 7.2% drop in accuracy, while removing the BiLSTM module resulted in a 5.6% drop. This result strongly confirms the necessity of integrating spatial and temporal features. Furthermore, the introduction of the attention mechanism and weighted focal loss improved the accuracy of difficult sample classification by 12.3%.
[0076] Through the detailed description of the above embodiments, those skilled in the art can fully understand the technical solutions of the present invention and implement them. For those skilled in the art, without departing from the core concept of the present invention, various obvious adjustments, modifications or equivalent replacements made to these embodiments should be deemed to be included in the protection scope of the present invention. It should be understood that the scope of protection requested by the present invention should not be limited to the specific implementation examples listed in the specification, but should be based on the maximum scope of protection determined by the technical features defined in the claims and their equivalent alternatives.
Claims
1. A GCN-BiLSTM-based old elevator fault diagnosis method, characterized in that: The following steps are involved: Step 1: Construct a multi-source heterogeneous data set. Vibration signals, acoustic emission signals, motor current signals, and bearing temperature signals are collected synchronously through the elevator IoT monitoring system. The raw data is preprocessed and abnormal data is eliminated based on the 3σ criterion. Step 2: Build a graph structure model based on physical topology, abstract the key components of the elevator into graph nodes, establish an adjacency matrix based on the fault propagation path, and integrate the node features with time-frequency domain indicators to form spatial-temporal correlation features. Step 3: Use a graph convolutional neural network (GCN) to extract spatial features. The GCN includes a multi-layer graph convolution structure. Each layer contains node feature transformation and neighborhood information aggregation, and introduces learnable weight parameters to adaptively adjust the connection strength between nodes. Step 4: Use a bidirectional long short-term memory (BiLSTM) network to extract temporal features. The BiLSTM contains forward and backward LSTM branches and embeds a gated expansion mechanism to enhance the ability to capture transient impact features. Step 5: Based on the attention mechanism, the spatial features extracted by GCN and the temporal features extracted by BiLSTM are fused, and fault classification is implemented through the multi-scale convolution module and Softmax classifier.
2. The method according to claim 1, characterized in that The multi-source signal acquisition in step 1 includes: the sampling frequency of the vibration signal is not less than 12.8kHz; the acquisition bandwidth of the acoustic emission signal is set to 40-100kHz; and a hardware trigger mechanism is used to ensure the time domain synchronization of the multi-source signals, and the synchronization error is controlled within 1μs.
3. The method according to claim 1, characterized in that The construction of the graph structure model in step 2 includes: using bearings, gearboxes, and traction wheels as graph nodes, establishing an adjacency matrix based on mechanical connections and fault propagation paths; and fusing node features with time domain indicators (such as peak value and kurtosis) and frequency domain indicators (such as envelope spectrum features).
4. The method according to claim 1, wherein The GCN network in step 3 includes: a three-layer graph convolution structure, using an improved graph attention mechanism (GAT) to calculate the weights of neighborhood nodes; embedding a residual connection structure in the last layer of GCN to prevent the degradation of deep network features; and using the DropEdge regularization method to randomly mask some edge connections to enhance model robustness.
5. The method according to claim 1, wherein The BiLSTM network in step 4 includes: a three-layer stacked BiLSTM architecture, with each layer containing 128 hidden units; a time-frequency transformation module is added before the input layer to enhance periodic fault characteristics through Morlet wavelet transform; and a dual attention mechanism is embedded, including a time attention layer and a feature attention layer, to adaptively focus on key time sequence segments and sensor channels.
6. The method according to claim 1, characterized in that The multimodal feature fusion in step 5 includes: using a cross-attention mechanism to calculate the correlation weights between spatial features and temporal features; using convolution kernels of three sizes, 1×3, 1×5, and 1×7, to extract multi-scale fault features in parallel; and using a weighted focal loss function to optimize the classification performance of imbalanced samples.
7. The method according to claim 1, characterized in that It also includes a transfer learning strategy: using pre-trained GCN-BiLSTM model parameters for fine-tuning to address the problem of insufficient specific fault samples in old elevators; adopting a layered fine-tuning strategy, including fixing the bottom-level parameters, gradually unfreezing the middle-level parameters, and global fine-tuning in three stages.
Citation Information
Patent Citations
Motor imagery electroencephalogram classification method based on ResCNN-BiGRU
CN117033985A
Power grid fault prediction method based on deep learning
CN118051827A
Electronic load MOS tube burning prediction method based on multi-scale and BILSTM cross fusion
CN118839226A
Battery capacity online estimation method based on model parameters
CN119199563A
Bus travel time prediction method based on space-time diagram attention network
CN119623762A
Cited By
Dry type series reactor turn-to-turn short circuit identification method and system based on bimodal deep learning
CN122221115A