Elevator running state multi-source sensing Internet of Things inspection system
Through multi-source sensing IoT patrol system for elevator operation status, multi-source sensing IoT patrol system with multi-modal sensing acquisition, edge computing and coordinated scheduling of heterogeneous sensors, the problem of limited perception dimensions and untimely response in elevator operation status monitoring is solved, and efficient fault warning and abnormal identification are achieved.
Patent Information
- Application Number
- CN202510501472.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-11
AI Technical Summary
The existing elevator operating status monitoring relies on a single sensor data, lacks multimodal information perception, and the response is not timely. Asynchronous collection of sensors leads to incomplete data fusion and low abnormal recognition accuracy, making it difficult to meet the needs of fine-grained identification in high-frequency use and complex environments.
Build a multi-source perception Internet of Things patrol system for elevator operation status, adopting multi-modal perception acquisition module, edge computing node cluster, heterogeneous sensor coordinated scheduling and attention mechanism-based timing modeling network to realize synchronous acquisition, fusion processing and local discrimination of multi-source data, and has the advantages of comprehensive perception dimensions, low response delay and high abnormal recognition accuracy.
Real-time and accurate identification and fault warning of elevator operating status are achieved, breaking through the problems of response delay and incomplete data fusion of traditional monitoring, improving the system's response timeliness and abnormal identification accuracy in the initial stage of failure, and adapting to the dynamic adjustment ability of different elevator models and environments.
Smart Images

Figure CN120288599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent inspection, and particularly to an Internet of Things inspection system for multi-source perception of elevator operation status. Background Art
[0002] With the continuous increase in urban population density and the rapid growth of high-rise buildings, elevators, as the core equipment for vertical transportation, pose higher requirements for the operation safety and maintenance efficiency in public facility management. Especially in old communities and places with high elevator usage intensity, how to achieve real-time perception, accurate identification, and fault warning of elevator operation status has become a key technical issue in the development of smart cities and Internet of Things applications.
[0003] In the prior art, the inspection of elevator operation status mainly relies on regular manual inspections or traditional single-sensor data collection. Common monitoring methods include acceleration signal analysis, motor current fluctuation detection, or simple abnormal records of logic controllers. These methods have multiple technical limitations: First, the data sources for monitoring are single, lacking comprehensive perception of multi-modal information such as images, audio, and environment, and it is difficult to accurately depict the complex dynamic characteristics during operation. Second, relying on a central server for centralized processing causes problems such as data upload delay and untimely response, and is not suitable for rapid identification in the initial stage of a fault. Third, there is a lack of a coordinated scheduling mechanism among sensors, and different types of sensors are prone to asynchronous phenomena in terms of trigger time, sampling frequency, and transmission path, affecting the integrity and consistency of data fusion.
[0004] In addition, although some studies have attempted to introduce deep learning models for elevator status classification or fault detection, most of the existing solutions still adopt offline training and centralized identification methods, lacking edge computing capabilities and distributed intelligent judgment mechanisms, and cannot effectively support the real-time, hierarchical, and multi-point status perception requirements in the elevator shaft environment. At the same time, most of the existing models ignore the spatial relationship between the sensor deployment location and the time-series data, and fail to achieve accurate modeling based on context and structural distribution, resulting in insufficient accuracy of abnormal identification and a lag in the warning mechanism, affecting the operation safety management effect.
[0005] In terms of operation status determination and abnormal event identification, the current common rules are mostly set based on empirical thresholds or changes in linear statistical indicators, lacking the ability to dynamically adjust to different elevator models, environmental noises, and usage patterns, and cannot meet the fine-grained identification and discrimination requirements in high-frequency usage and complex-structured environments, resulting in problems such as high false alarm rates and diagnostic delays.
[0006] In summary, there is an urgent need for a multi-source perception and inspection system for elevator operation status that integrates multi-modal perception fusion algorithms, edge computing node collaborative deployment, and intelligent time-series recognition mechanisms to solve the key technical problems of limited perception dimensions, untimely response, and low abnormal recognition accuracy in current elevator status monitoring. Summary of the Invention
[0007] An object of the present invention is to propose a multi-source perception Internet of Things inspection system for elevator operation status. The present invention integrates multi-modal perception data collection, edge computing node distributed deployment, and spatio-temporal collaborative perception mechanisms, constructs a time-series modeling network and an abnormal reconstruction discrimination model based on the attention mechanism, realizes synchronous collection and fusion processing of multi-source data through heterogeneous sensor collaborative scheduling, and performs local discrimination and event reporting at the edge node, having the advantages of comprehensive perception dimensions, low response latency, and high abnormal recognition accuracy.
[0008] According to an embodiment of the present invention, a multi-source perception Internet of Things inspection system for elevator operation status, the system includes:
[0009] S1. A multi-modal perception collection module, including multiple types of sensors, for synchronously collecting acceleration data, structural vibration data, noise audio data, environmental temperature and humidity data, and in-car video image data during elevator operation, and generating multi-channel perception data packets through a unified time stamp mechanism;
[0010] S2. A data preprocessing module, for preprocessing the multi-channel perception data packets and outputting a structured multi-modal data sequence;
[0011] S3. An edge computing node cluster, deployed on the top of the elevator car, in the middle of the elevator shaft, and inside the control cabinet, for performing preliminary abnormal discrimination on the structured multi-modal data sequence and generating local event description information;
[0012] S4. A multi-modal feature fusion module, for receiving local event description information from each edge computing node, constructing a fusion feature map based on a time-series alignment and space compensation strategy, and generating a unified feature vector sequence;
[0013] S5. A heterogeneous sensor collaborative scheduling module, aiming at the physical heterogeneity and sampling asynchrony of multiple types of sensors, performing time synchronization control, sampling frequency dynamic adjustment, and acquisition task coordination to make different sensors consistent in trigger conditions, data frequencies, and data transmission paths;
[0014] S6. A status recognition and abnormal annotation module, for inputting the unified feature vector sequence into a time-series modeling network based on the attention mechanism and outputting corresponding elevator operation status labels and abnormal event identification codes.
[0015] Optionally, the S3 specifically includes:
[0016] S31. A data receiving unit, configured to receive a structured multimodal data sequence within a corresponding area, and perform caching and decoding operations on the structured multimodal data sequence;
[0017] S32. A feature extraction unit, configured to perform one-dimensional convolution operations on acceleration data, structural vibration data, and noise audio data, and perform two-dimensional convolution operations on in-car video image data to extract corresponding low-dimensional feature vectors;
[0018] S33. A local anomaly discrimination unit, configured to construct a time series window based on the low-dimensional feature vectors, and calculate a change rate R of the eigenvalue corresponding to the low-dimensional feature vector within each window deviating from the mean value i :
[0019]
[0020] wherein, v i is the i-th eigenvalue within the current window, μ is the mean value of all eigenvalues within this window, σ is the standard deviation, and when R i exceeds a set threshold δ, it is determined as an abnormal point, and after detecting the abnormal point, local event description information is generated;
[0021] S34. A node communication interface, configured to send the local event description information to the multimodal feature fusion module.
[0022] Optionally, the S4 specifically includes:
[0023] S41. An input integration unit, configured to receive local event description information from each edge computing node, and perform preliminary sorting and numbering on the local event description information according to the node deployment location and timestamp order;
[0024] S42. A timing alignment unit, configured to construct a sliding time window based on the global timestamp, perform time synchronization processing on the local event description information of multiple nodes within the same time window, and use an interpolation strategy to complete it if any node event is missing;
[0025] S43. A space compensation unit, configured to calculate a weight factor w j , and apply space compensation to the low-dimensional feature vector x j of each node to generate an adjusted low-dimensional feature vector x' j :
[0026] x′ j = w j ·x j
[0027] wherein, xj represents the low-dimensional feature vector of the j-th node, w j is the weight factor calculated according to the distance from the node to the reference point of the car center;
[0028] S44. A fusion coding unit, configured to input each adjusted low-dimensional feature vector into a feature stacking network structure, and perform feature fusion processing on each adjusted low-dimensional feature vector through a multi-head attention mechanism, and finally generate a unified feature vector sequence.
[0029] Optionally, the S44 specifically includes:
[0030] S441. A feature stacking unit, configured to longitudinally stack the adjusted low-dimensional feature vectors x' j from each edge computing node in the order of the edge computing node numbers to construct an input feature tensor X, where X = [x'1, x'2,..., x' n T , n represents the number of edge nodes, d is the single-node feature dimension;
[0031] S442. A position encoding unit, configured to append a fixed position encoding vector p j to each node position in the input feature tensor X to obtain an encoded vector z with position information j = x' j + p j ;
[0032] S443. A multi-head attention fusion unit, configured to form all the encoded vectors z j into an input feature matrix Z = [z1, z2,..., z n T , and obtain a query matrix Q = ZW Q , a key matrix K = ZW K , and a value matrix V = ZW V through a linear transformation, input the three matrices into the multi-head attention mechanism, and perform parallel attention weight calculation and weighted feature aggregation:
[0033]
[0034] Among them, Attention(Q, K, V) is the attention function, d k is the dimension of the key vector, W Q , W k , W V are linear weight matrices used for mapping during the training process, and softmax is the activation function;
[0035] S444. The fusion output unit is used to splice and linearly map the results output by each attention head to generate a unified feature vector sequence, and output the unified feature vector sequence to the state recognition and anomaly annotation module.
[0036] Optionally, the S6 specifically includes:
[0037] S61. The input encoding unit is used to receive the unified feature vector sequence and construct an input matrix H = [h1, h2,..., h T according to the time step t, where represents the feature vector at the t-th time step, T is the sequence length, and d is the feature dimension;
[0038] S62. The temporal modeling unit is used to input the input matrix H into a temporal network structure with a multi-head attention mechanism to construct a state-aware context relationship and output a context feature vector;
[0039] S63. The state classification unit is used to perform a fully connected transformation and softmax classification on the context feature vector output by the temporal modeling unit, and output a corresponding elevator operation state label set where each s t represents the elevator operation state label corresponding to the t-th time step;
[0040] S64. The anomaly annotation unit is used to construct an anomaly scoring function A t based on the context feature vector, and compare the scoring value at each time step with a set threshold θ to generate an anomaly event identification code:
[0041]
[0042] where, h t is the feature vector at the t-th time step, is the reconstructed feature vector predicted by the model. If A t > θ, it is marked as an abnormal time step and the corresponding anomaly event identification code is output.
[0043] Optionally, the S62 specifically includes:
[0044] S621. The position encoding layer is used to add a position encoding vector p T to each feature vector h t at each time step in the input matrix H = [h1, h2,..., h T t to obtain an encoded vector e t with position information, where e t = h t + p t to form a position encoding sequence E = [e1, e2,..., e T;
[0045] S622, a self-attention encoding layer, is used to input the position encoding sequence E into a multi-head attention mechanism with the same structure as that described in S443. The structure includes a linear mapping process of a query matrix, a key matrix, and a value matrix, as well as an attention weighting calculation mechanism, to form an intermediate representation sequence. The attention calculation method is the same as that of S443 in claim 4;
[0046] S623, a feed-forward network layer, is used to perform feed-forward sub-network operations composed of two linear transformation layers on the vectors at each time step in the intermediate representation sequence. The feed-forward structure form is:
[0047] FFN(x) = max(0, xW1 + b1)W2 + b2
[0048] where x is the input context feature vector at a certain time step, FFN(x) is the output feature vector generated after the vector is processed by two linear transformations and a non-linear activation function, W1 and W2 are linear weight matrices, b1 and b2 are bias terms, and finally the context feature vector is output.
[0049] The beneficial effects of the present invention are:
[0050] (1) By deploying an edge computing node cluster with local computing capabilities at key parts of the elevator, the present invention constructs a node-level anomaly discrimination and event generation mechanism, realizes real-time analysis and preliminary judgment of multi-source data such as structural vibration, noise, and images, effectively breaks through the problem of traditional elevator inspection relying on the central server for processing and high response delay, and significantly enhances the response timeliness and data closed-loop ability of the elevator system in the initial stage of a fault.
[0051] (2) The present invention constructs a spatio-temporal consistency feature fusion mechanism, aligns the local event information uploaded by multiple edge nodes through a sliding time window, and combines the node spatial positions to introduce a weighted compensation strategy to complete multi-modal feature stacking and unified encoding processing based on the attention mechanism, breaks through the technical bottleneck of incomplete feature fusion caused by asynchronous sensor acquisition in the existing system, and improves the data fusion quality and state recognition accuracy.
[0052] (3) The present invention constructs a time series modeling network based on the multi-head attention mechanism, and on this basis, introduces a feature reconstruction error function for anomaly scoring and event annotation to form a unified state recognition and anomaly warning integrated mechanism. The system can automatically generate elevator state labels and anomaly event identification codes according to the changes in the operating state at continuous time steps, breaks through the traditional method relying on static threshold judgment and manual setting of rules, has stronger adaptive ability and refined discrimination ability, and has higher adaptability and recognition accuracy under various working conditions and complex environments. Description of the Drawings
[0053] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. In the accompanying drawings:
[0054] Figure 1 It is a system functional structure framework diagram of an elevator operation status multi-source perception Internet of Things inspection system proposed by the present invention. Detailed implementation manners
[0055] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.
[0056] Reference Figure 1 , an elevator operation status multi-source perception Internet of Things inspection system, the system includes:
[0057] S1. A multimodal perception acquisition module, including multiple types of sensors, is used to synchronously acquire acceleration data, structural vibration data, noise audio data, environmental temperature and humidity data, and video image data inside the car during the operation of the elevator, and generate multi-channel perception data packets through a unified timestamp mechanism;
[0058] In a specific implementation manner of the present invention, the S1 multimodal perception acquisition module includes multiple types of sensor units deployed at key positions such as the top of the elevator car, the motor room, and the shaft wall surface. The sensor units include three-axis acceleration sensors, vibration velocity sensors, digital microphones, temperature and humidity sensors, and high-definition image acquisition modules; among them, the acceleration sensors and vibration sensors are used to obtain the mechanical response characteristics during the operation of the elevator in real time, the microphone is used to collect the audio signals during the operation of the elevator car and the shaft, the temperature and humidity sensors collect the change information of environmental parameters, and the image acquisition module is used to collect the real-time images inside the car; each type of sensor performs data synchronization scheduling through an embedded controller, and uniformly uses a high-precision timestamp mechanism to mark the collected data, constructing a perception data packet with consistent multi-source channels. This data packet contains synchronous data sequences in multiple modalities and serves as the input basis for subsequent edge computing and fusion processing.
[0059] S2. A data preprocessing module, which is used to preprocess the multi-channel perception data packet and output a structured multimodal data sequence;
[0060] In a specific embodiment of the present invention, the S2 data preprocessing module includes a multi-modal data parsing unit, a signal processing unit, and a format standardization unit, which are used to perform classification parsing and preprocessing operations on the multi-channel perception data packets generated by the S1 module; among them, the acceleration data and the structural vibration data are subjected to noise removal and normalization processing through a band-pass filter, and the effective signal segments within the continuous time window are extracted; the noisy audio data is converted into a time-frequency spectrogram form through short-time Fourier transform (STFT) and amplitude normalization is performed; the temperature and humidity data are interpolated and filled in according to the sampling time sequence and outliers are removed; the image data is processed by image enhancement and size standardization, including gray-scale normalization, noise filtering, and spatial resolution resampling; all processed data are realigned according to the time stamp and uniformly output as a structured multi-modal data sequence, which is used as the input data for the subsequent edge computing nodes to perform preliminary feature extraction and discrimination.
[0061] S3. Edge computing node cluster, deployed on the top of the elevator car, in the middle of the elevator shaft, and inside the control cabinet, is used to perform preliminary anomaly discrimination on the structured multi-modal data sequence and generate local event description information;
[0062] In this embodiment, S3 specifically includes:
[0063] S31. Data receiving unit, which is used to receive the structured multi-modal data sequence in the corresponding area and perform caching and decoding operations on the structured multi-modal data sequence;
[0064] S32. Feature extraction unit, which is used to perform one-dimensional convolution operations on the acceleration data, structural vibration data, and noisy audio data, and perform two-dimensional convolution operations on the video image data inside the car to extract the corresponding low-dimensional feature vectors;
[0065] S33. Local anomaly discrimination unit, which is used to construct a time series window based on the low-dimensional feature vectors and calculate the change rate R of the eigenvalue corresponding to the low-dimensional feature vector in each window deviating from the mean value i :
[0066]
[0067] where, v i is the i-th eigenvalue in the current window, μ is the mean value of all eigenvalues in this window, σ is the standard deviation, and when R i exceeds the set threshold δ, it is determined as an abnormal point, and after detecting the abnormal point, local event description information is generated;
[0068] This formula is used to measure the deviation degree of the i-th eigenvalue v i in the current window relative to the mean value μ of all eigenvalues in this window, and is normalized by the standard deviation σ to obtain the feature change rate R iAmong them, μ represents the mean within this time window, reflecting the central tendency of the data, and σ represents the degree of data dispersion. This change rate R i Actually constructs an anomaly metric based on the idea of z-score normalization. When it exceeds the preset threshold δ, it is considered that this feature point shows a significant deviation and is thus determined to be an anomaly point.
[0069] S34, Node communication interface, for sending the local event description information to the multi-modal feature fusion module.
[0070] In this embodiment, by constructing a multi-modal feature fusion module, the local event description information uploaded by different edge nodes is uniformly integrated. The sliding time window mechanism is used to achieve time synchronization of multi-node data. The spatial weighted compensation method guided by node positions is used to enhance the feature collaboration between nodes, and based on the multi-head attention mechanism, deep fusion of cross-modal features is realized, improving the integrity and context consistency of feature representation, being able to more effectively capture potential abnormal patterns during elevator operation, and improving the accuracy of state recognition and the comprehensive processing ability of the system for multi-source perception data.
[0071] S4, Multi-modal feature fusion module, for receiving the local event description information from each edge computing node, constructing a fused feature map based on the time series alignment and spatial compensation strategy, and generating a unified feature vector sequence;
[0072] In this embodiment, S4 specifically includes:
[0073] S41, Input integration unit, for receiving the local event description information from each edge computing node, and performing preliminary sorting and numbering on the local event description information according to the node deployment location and timestamp order;
[0074] S42, Time series alignment unit, for constructing a sliding time window based on the global timestamp, performing time synchronization processing on the local event description information of multiple nodes within the same time window, and using the interpolation strategy to complete the complementation if any node event is missing;
[0075] S43, Spatial compensation unit, for calculating the weight factor w j , and applying spatial compensation to the low-dimensional feature vector x j of each node to generate an adjusted low-dimensional feature vector x' j :
[0076] x′ j =w j ·x j
[0077] Where x jrepresents the low-dimensional feature vector of the j-th node, w j is the weight factor calculated according to the distance from the node to the reference point of the car center;
[0078] is used to perform spatial weighted processing on the low-dimensional feature vectors x of each edge node j where x j represents the original low-dimensional feature vector extracted by the j-th node at the current time step, w j is the compensation coefficient calculated according to the physical distance from the node to the reference point of the car center, which is used to reflect the relative importance of the node in the spatial structure. This weight factor can reflect the layout differences of sensors in the elevator structure, making the features of nodes with more critical spatial positions or more significant physical impacts occupy a higher weight in the overall fusion process, so as to generate the feature vector x j ' .
[0079] S44. The fusion encoding unit is used to input each adjusted low-dimensional feature vector into the feature stacking network structure, and perform feature fusion processing on each adjusted low-dimensional feature vector through the multi-head attention mechanism, and finally generate a unified feature vector sequence.
[0080] The S44 specifically includes:
[0081] S441. The feature stacking unit is used to vertically stack the adjusted low-dimensional feature vectors x' from each edge computing node j in the order of the edge computing node numbers to construct the input feature tensor X, where X = [x'1, x'2,..., x' n T , n represents the number of edge nodes, d is the single-node feature dimension;
[0082] S442. The position encoding unit is used to append a fixed position encoding vector p to each node position in the input feature tensor X j to obtain the encoded vector z with position information j = x' j + p j ;
[0083] S443. The multi-head attention fusion unit is used to form all the encoded vectors z j into the input feature matrix Z = [z1, z2,..., z n T , and obtain the query matrix Q = ZW Q , the key matrix K = ZW K and the value matrix V = ZW V , input the three matrices into a multi-head attention mechanism to perform parallel attention weight calculation and weighted feature aggregation:
[0084]
[0085] Among them, Attention(Q, K, V) is the attention function, and d k is the dimension of the key vector, and W Q , W K , W V are linear weight matrices used for mapping during the training process, and softmax is the activation function;
[0086] is used to implement feature correlation modeling and weighted aggregation operations between different node encoding vectors. Among them, the encoding vectors z j constitute the input feature matrix Z, and the query matrix Q = ZW Q , the key matrix K = ZW K and the value matrix V = ZW V are generated through linear mapping, where W Q , W K , W V are weight parameters learned during the training process; d k represents the dimension of the key vector. This formula first calculates QK T through the inner product to measure the similarity between each pair of features, and scales it through to avoid gradient disappearance. Then, the softmax function is used to normalize the similarity to obtain the attention weight matrix, and finally, it is multiplied by the value matrix V to achieve weighted feature fusion.
[0087] S444. The fusion output unit is used to splice and linearly map the results output by each attention head to generate a unified feature vector sequence, and output the unified feature vector sequence to the state recognition and anomaly annotation module.
[0088] In this embodiment, by constructing a multi-modal feature fusion module composed of input integration, temporal alignment, spatial compensation, and fusion coding, it is possible to uniformly process the local event description information from different edge computing nodes at the time and space levels, ensure the alignment of heterogeneous data within the time window and perform spatial position compensation, thereby constructing a fusion feature representation with higher semantic consistency; during the fusion coding process, a multi-head attention mechanism is introduced to perform weighted modeling on the adjusted feature vectors of each node, which not only improves the sensitivity to key anomaly features but also enhances the system's ability to model complex dependencies between nodes, realizes the deep fusion expression of multi-source information in the elevator operation state, and provides high-quality feature inputs with complete structure and spatio-temporal consistency for subsequent state recognition and anomaly annotation.
[0089] S5. Heterogeneous Sensor Cooperative Scheduling Module. Aiming at the physical heterogeneity and sampling asynchrony of multiple types of sensors, it performs time synchronization control, dynamic adjustment of sampling frequency, and coordination of acquisition tasks to make different sensors coordinated and consistent in triggering conditions, data frequencies, and data transmission paths.
[0090] In a specific embodiment of the present invention, the S5 heterogeneous sensor cooperative scheduling module includes a time synchronization unit, a frequency control unit, and a task coordination unit, which are used to solve the differences in sampling periods, data triggering mechanisms, and transmission strategies among multiple types of sensors. The time synchronization unit uses a method based on a high-precision clock or the Network Time Protocol to uniformly calibrate the timestamps of various sensors to ensure that data can be aligned under the same time reference. The frequency control unit dynamically adjusts the sampling frequencies of different types of sensors according to the load status and current operating stage of the edge computing node. For example, it increases the acceleration and audio sampling densities during the elevator start or brake stage and reduces the amount of redundant data in the stationary state. The task coordination unit reasonably allocates triggering conditions and data reporting paths according to the acquisition priorities and node deployment situations, so that data such as temperature, humidity, vibration, and images maintain load balance and time consistency on the communication link, thereby realizing the efficient cooperation and orderly acquisition of multi-source heterogeneous data.
[0091] S6. State Recognition and Abnormality Annotation Module. It is used to input the unified feature vector sequence into a time series modeling network based on the attention mechanism and output the corresponding elevator operating state label and abnormality event identification code.
[0092] In this embodiment, S6 specifically includes:
[0093] S61. Input Encoding Unit. It is used to receive the unified feature vector sequence and construct an input matrix H = [h1, h2,..., h T according to the time step t, where represents the feature vector at the t-th time step, T is the sequence length, and d is the feature dimension.
[0094] S62. Time Series Modeling Unit. It is used to input the input matrix H into a time series network structure with a multi-head attention mechanism to construct a state-aware context relationship and output a context feature vector.
[0095] S63. State Classification Unit. It is used to perform a fully connected transformation and softmax classification on the context feature vector output by the time series modeling unit and output the corresponding elevator operating state label set where each s t represents the elevator operating state label corresponding to the t-th time step.
[0096] S64. Abnormality Annotation Unit. It is used to construct an abnormality scoring function A based on the context feature vectort , and compare the scoring value at each time step with the set threshold θ to generate an abnormal event identification code:
[0097]
[0098] where h t is the feature vector at the t-th time step, is the reconstructed feature vector predicted by the model. If A t > θ, it is marked as an abnormal time step and the corresponding abnormal event identification code is output.
[0099] is used to evaluate the difference degree between the feature vector h t at the t-th time step and the reconstructed feature vector obtained by model prediction. This formula calculates the Euclidean distance (L2 norm) between the two as the abnormal scoring function A t . The larger the value, the more significant the deviation of the actual state at the current moment from the normal mode learned by the model. If the scoring value A t exceeds the set threshold θ, the current time step is determined to be in an abnormal state, and the corresponding abnormal event identification code is generated accordingly. This method establishes an abnormal detection mechanism based on the principle of reconstruction error, does not rely on fixed rules or manually set templates, and can achieve adaptive abnormal identification.
[0100] The S62 specifically includes:
[0101] S621. A position encoding layer, which is used to add a position encoding vector p T to the feature vector h t at each time step in the input matrix H = [h1, h2,..., h t to obtain an encoded vector e t with position information = h t + p t , and form a position encoding sequence E = [e1, e2,..., e T ;
[0102] S622. A self-attention encoding layer, which is used to input the position encoding sequence E into the multi-head attention mechanism with the same structure as that in S443. The structure includes the linear mapping process of the query matrix, the key matrix and the value matrix, and the attention weighting calculation mechanism to form an intermediate representation sequence. The attention calculation method is the same as that in S443 of claim 4;
[0103] S623. A feed-forward network layer, which is used to perform a feed-forward sub-network operation composed of two linear transformation layers on the vectors at each time step in the intermediate representation sequence respectively. The feed-forward structure form is:
[0104] FFN(x) = max(0, xW1 + b1)W2 + b2
[0105] Among them, x is the input of the context feature vector at a certain time step, FFN(x) is the output feature vector generated after the vector is processed by two layers of linear transformation and non-linear activation function, W1 and W2 are linear weight matrices, b1 and b2 are bias terms, and finally the context feature vector is output.
[0106] In the feed-forward network layer of S623 of the present invention, in the feed-forward calculation formula FFN(x) = max(0, xW1 + b1)W2 + b2 adopted, each parameter has a strict dimension matching relationship: the dimension of the input vector x is 1×d, and the dimension of the first-layer weight matrix W1 is d×d f , the bias term b1 is 1×d f , and the dimension of the intermediate vector after activation remains 1×d f , and then through the second-layer weight matrix W2, its dimension is d f ×d, combined with the bias term is mapped back to the original dimension, and the final output feature vector has the same dimension as the input x, which is 1×d, ensuring that the feed-forward sub-network has good structural closure and residual connection compatibility.
[0107] In this embodiment, by constructing a time series modeling unit including a position encoding layer, a self-attention encoding layer and a feed-forward network layer, the time-dependent information in the unified feature vector sequence is converted into a context-aware feature representation, which can capture the dynamic features and correlation relationships of the elevator operating state changing with time. Among them, the multi-head attention mechanism realizes the differential modeling of the importance of features between different time steps, and the feed-forward network structure enhances the model's ability to express non-linear time series patterns, so as to realize the fine modeling of the evolution process of the operating state and the extraction of context features, improving the system's recognition ability of the operating state changes under complex elevator working conditions and the modeling depth of abnormal evolution trends.
[0108] Example:
[0109] To verify the practicability and effectiveness of the present invention, the "Multi-source Perception Internet of Things Inspection System for Elevator Operation Status" proposed by the present invention was deployed in the elevator management scenario of "Zhonghai Building", a large comprehensive office building in City H. This building is equipped with a total of 12 elevators, with a daily operation frequency exceeding 1,500 times, and the operation time is concentrated in the morning and evening rush hours and the lunchtime period. The maintenance unit of the building's elevator system has long faced problems such as lagging operation status perception, relying on manual labor for fault identification, and high maintenance costs. Especially in the case of frequent use of multiple elevators, sudden operation anomalies are difficult to capture and respond to in a timely manner, posing potential safety hazards. To solve the above problems, between December 2024 and March 2025, the property management party selected 4 elevators mainly serving the high-rise floors to deploy the multi-source perception system of the present invention, continuously monitor and evaluate the operation status, operation environment and local anomalies, and compare and verify with the traditional manual inspection mode.
[0110] After the system deployment was completed, a total of 22 groups of multi-modal acquisition nodes such as acceleration sensors, vibration sensors, microphone modules, high-definition cameras and temperature and humidity sensors were installed on the top of the elevator car, on the wall of the elevator shaft and inside the control cabinet respectively. Each node was connected to the edge computing unit for local data preprocessing and preliminary anomaly judgment, and the system constructed a unified data scheduling mechanism and feature fusion network. All node data was integrated into a structured data packet through local time synchronization, frequency adaptive control and unified timestamp mechanism, and input into the fusion feature generation unit and the time series modeling unit to realize state discrimination, anomaly detection and generate event identification codes locally. The system set the anomaly response threshold based on historical operation data and dynamic context features, and had the ability of continuous learning and self-adaptation.
[0111] During the operation, the system detected that during the morning rush hour on January 15, 2025, the car of a certain elevator had an operation jitter event between the 7th and 8th floors, accompanied by abnormal noises. Traditional manual inspections could not locate this problem on the day of the incident, but the edge nodes of the system identified 3 consecutive time-step anomalies through the acceleration change rate and audio spectrum mutation, and the reconstruction error score AtA_t reached 2.6 times the set threshold, immediately generating an anomaly event identification code and notifying the maintenance personnel through the platform. Subsequent investigations found that the car guide rail was loose. If this problem was not handled in a timely manner, it might lead to abnormal elevator operation postures or even outages. The system helped the maintenance personnel intervene and handle the problem 4 hours in advance, avoiding operation interruptions.
[0112] To quantify the performance of the system in the real operation environment, during the continuous three-month deployment period, the data differences between the system and the traditional inspection method in terms of operation status identification, anomaly event detection, response efficiency and false alarm control were recorded, and the results are shown in Table 1 below.
[0113] Table 1 Comparison table of the deployment effects of the elevator inspection system in Zhonghai Building
[0114]
[0115]
[0116] As can be seen from Table 1, after the deployment of the system of the present invention, the frequency of abnormal detection of elevators has been significantly improved, indicating that the system has stronger sensitivity in identifying minor abnormalities and early fault warning; the response efficiency has been greatly compressed from the traditional manual discovery response time of more than 8 hours to within an average of 0.5 hours, greatly shortening the problem handling cycle. At the same time, through the high-quality feature fusion and context time series analysis mechanism, the system accurately identifies most potential faults, and the false alarm rate of early warning is controlled within 2.3%, far lower than the acceptable engineering level. In addition, the average monthly maintenance work order quantity has decreased by nearly 65%, the abnormal operation of the elevator subjectively felt by users has been greatly reduced, and the number of complaints has decreased by more than 85%.
[0117] In a specific case, on the evening of February 3, 2025, the system detected fluctuations in the structural resonance signal during the operation of an elevator at the bottom of the hoistway. The abnormal frequency band of the audio spectrum continuously appeared between 1.5 kHz and 2 kHz, and the frequency domain change rate exceeded 3.8 times the static mean value. The system immediately calculated the reconstruction error score through the S64 unit, with a value of 1.96, exceeding the set threshold θ = 1.2, and determined it to be abnormal and reported it. After recheck by the maintenance technicians, it was found that the fixing bracket of the position sensor at this location was loose, and the reinforcement treatment was completed in time to avoid further expansion of subsequent faults. If this abnormality was not captured by the system, it might cause false alarms or service interruptions during peak hours.
[0118] To sum up, this embodiment verifies the stability, identification accuracy and intelligence advantages of the system of the present invention in the elevator environment of high-frequency operation and high-rise buildings, especially showing excellent performance in multi-source information fusion, local abnormal detection and time series modeling and discrimination. The system can realize end-side intelligent analysis, edge collaborative decision-making and central abnormal archiving during operation, process the operation state data in a full-process closed-loop manner, significantly improve the intelligent level of elevator operation and maintenance, reduce the risk of fault-induced elevator stoppage, optimize the allocation of operation and maintenance resources, and is an important path for the upgrade of the traditional elevator inspection mode to the intelligent Internet of Things architecture.
[0119] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. An elevator operation status multi-source perception Internet of Things inspection system, characterized in that, The system includes: S1. A multi-modal perception acquisition module, including multiple types of sensors, which is used to synchronously collect acceleration data, structural vibration data, noise audio data, ambient temperature and humidity data, and in-car video image data during the operation of the elevator, and generate multi-channel perception data packets through a unified timestamp mechanism; S2. A data preprocessing module, which is used to preprocess the multi-channel perception data packets and output a structured multi-modal data sequence; S3. An edge computing node cluster, deployed on the top of the elevator car, in the middle of the elevator shaft, and inside the control cabinet, which is used to perform preliminary anomaly discrimination on the structured multi-modal data sequence and generate local event description information; S4. A multi-modal feature fusion module, which is used to receive the local event description information from each edge computing node, construct a fusion feature map based on the time series alignment and spatial compensation strategy, and generate a unified feature vector sequence; S5. A heterogeneous sensor collaborative scheduling module, which, aiming at the physical heterogeneity and sampling asynchrony of multiple types of sensors, performs time synchronization control, sampling frequency dynamic adjustment, and acquisition task coordination to make the collaboration of different sensors consistent in terms of trigger conditions, data frequencies, and data transmission paths; S6. A state recognition and anomaly annotation module, which is used to input the unified feature vector sequence into a time series modeling network based on the attention mechanism and output the corresponding elevator operation state label and anomaly event identification code.
2. The multi-source perception Internet of Things inspection system for elevator operation status according to claim 1, wherein, The S3 specifically includes: S31. A data receiving unit, which is used to receive the structured multi-modal data sequence in the corresponding area and perform caching and decoding operations on the structured multi-modal data sequence; S32. A feature extraction unit, which is used to perform one-dimensional convolution operations on the acceleration data, structural vibration data, and noise audio data, and perform two-dimensional convolution operations on the in-car video image data to extract the corresponding low-dimensional feature vectors; S33. Local anomaly discrimination unit, configured to construct a time series window based on the low-dimensional feature vector and calculate a change rate R of the eigenvalue corresponding to the low-dimensional feature vector within each window deviating from the mean value i :[[-END]] Among them, v i is the i-th eigenvalue in the current window, μ is the mean of all eigenvalues in this window, and σ is the standard deviation. When R i exceeds the set threshold δ, it is determined as an outlier. After detecting an outlier, local event description information is generated; S34. A node communication interface, which is used to send the local event description information to the multi-modal feature fusion module.
3. The multi-source perception Internet of Things inspection system for elevator operation status according to claim 1, characterized in that, The S4 specifically includes: S41. An input integration unit, which is used to receive the local event description information from each edge computing node and perform preliminary sorting and numbering on the local event description information according to the node deployment location and timestamp order; S42. A time series alignment unit, which is used to construct a sliding time window based on the global timestamp, perform time synchronization processing on the local event description information of multiple nodes within the same time window, and use an interpolation strategy to complete it if any node event is missing; S43. A spatial compensation unit, which is used to calculate a weight factor w by combining node spatial position parameters and a physical distribution structure j , and apply spatial compensation to the low-dimensional feature vector x j of each node to generate an adjusted low-dimensional feature vector x' j : x′ j = w j · x j where x j represents the low-dimensional feature vector of the j-th node, and w j is the weight factor calculated according to the distance from the node to the reference point of the car center; S44. A fusion coding unit, which is used to input each adjusted low-dimensional feature vector into a feature stacking network structure, and perform feature fusion processing on each adjusted low-dimensional feature vector through a multi-head attention mechanism, and finally generate a unified feature vector sequence.
4. The multi-source perception Internet of Things inspection system for elevator operation status according to claim 3, characterized in that The S44 specifically includes: S441. Feature stacking unit, which is used to longitudinally stack the adjusted low-dimensional feature vectors x' from each edge computing node j in the order of the edge computing node numbers to construct an input feature tensor X, where X = [x'1, x'2,..., x' n T , n represents the number of edge nodes, and d is the feature dimension of a single node; S442. A position encoding unit for attaching a fixed position encoding vector p to each node position in the input feature tensor X j to obtain an encoded vector z with position information j = x' j + p j ; S443. The multi-head attention fusion unit is used to combine all the encoded vectors z j to form the input feature matrix Z = [z1, z2,..., z n T , and obtain the query matrix Q = ZW Q , the key matrix K = ZW K , and the value matrix V = ZW V . Input the three matrices into the multi-head attention mechanism to perform parallel attention weight calculation and weighted feature aggregation: Among them, Attention(Q, K, V) is the attention function, and d k is the dimension of the key vector, and W Q , W K , W V are the linear weight matrices used for mapping during the training process, and softmax is the activation function; S444. A fusion output unit, which is used to splice and linearly map the results output by each attention head to generate a unified feature vector sequence, and output the unified feature vector sequence to the state recognition and anomaly annotation module.
5. The multi-source perception Internet of Things inspection system for elevator operation status according to claim 4, characterized in that, The S6 specifically includes: S61. An input encoding unit, configured to receive a unified feature vector sequence, and construct an input matrix H = [h1, h2,..., h T according to a time step t from the unified feature vector sequence, where represents a feature vector at the t-th time step, T is the sequence length, and d is the feature dimension; S62, a timing modeling unit, is configured to input an input matrix H into a timing network structure with a multi-head attention mechanism, construct a state-aware context relationship, and output a context feature vector; S63. A state classification unit is configured to perform a fully-connected transformation and softmax classification on the context feature vectors output by the temporal modeling unit, and output a corresponding set of elevator operation state labels where each s t represents the elevator operation state label corresponding to the t-th time step; S64. Anomaly annotation unit, which is used to construct an anomaly scoring function A based on the context feature vector t , and compare the scoring value at each time step with the set threshold θ to generate an anomaly event identification code: Among them, h t is the feature vector at the t-th time step, is the reconstructed feature vector predicted by the model. If A t > θ, it is marked as an abnormal time step and the corresponding abnormal event identification code is output.
6. The multi-source perception Internet of Things inspection system for elevator operating status according to claim 5, characterized in that, The specific composition of the S62 is as follows: S621, position encoding layer, used for inputting matrix H = [h1,h2,...,h T The feature vector h at each time step in ] t Add position encoding vector p t , get the encoding vector e with position information t =h t +p t , forming a position coding sequence E = [e1, e2, ..., e T ]; S622, a self-attention encoding layer, is configured to input a position encoding sequence E into a multi-head attention mechanism with the same structure as that of S443, where the structure includes a linear mapping process of a query matrix, a key matrix, and a value matrix, as well as an attention weighting calculation mechanism, to form an intermediate representation sequence, and the attention calculation method is the same as that of S443 in claim 4; S623, a feed-forward network layer, is configured to perform a feed-forward sub-network operation composed of two linear transformation layers on the vectors at each time step in the intermediate representation sequence, and the feed-forward structure is in the form of: FFN(x) = max(0, xW1 + b1)W2 + b2 where x is the input of the context feature vector at a certain time step, FFN(x) is the output feature vector generated after the vector is processed by two linear transformations and a non-linear activation function, W1 and W2 are linear weight matrices, b1 and b2 are bias terms, and finally the context feature vector is output.
Citation Information
Cited By
Elevator running state detection method and system
CN120553525A
Elevator operation state detection method and system
CN120553525B
Partial discharge multichannel signal real-time synchronous acquisition method based on edge calculation
CN120658775A
Off-grid battery replacement method and system for battery replacement cabinet
CN121224509A
Parking space state monitoring method based on multi-sensor fusion
CN121365355A