Unified supervision system integrating car camera, wearable terminal and cloud platform

By integrating car cameras, wearable terminals, and a cloud platform into a unified monitoring system, the problems of data fragmentation and delayed risk identification during elevator maintenance have been solved, achieving efficient safety monitoring and rapid response, and reducing the risk of accidents.

CN121626792APending Publication Date: 2026-03-10QUICKCO INTELLIGENT TECH (SUZHOU) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Elevator maintenance suffers from fragmented data, lack of unified management, and a lack of collaborative analysis capabilities for multi-source videos. Risk identification is delayed and alarms are not intelligent, leading to a high risk of accidents.

Method used

A unified monitoring system integrating car cameras, wearable terminals, and a cloud platform is adopted. The car camera subsystem acquires multi-view video data, the wearable terminal collects maintenance personnel data, the edge computing node performs time-stamping and spatial registration, and the cloud platform constructs a spatiotemporal panoramic record and performs risk identification, realizing multi-dimensional information closed-loop collection and intelligent identification.

Benefits of technology

It improves the accuracy and response speed of safety monitoring during elevator maintenance, enabling the identification of abnormal events and triggering of protective commands in milliseconds, reducing the probability of accidents and improving the sensitivity and accuracy of risk identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121626792A_ABST
    Figure CN121626792A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent wearing and safety production, in particular to a unified supervision system integrating a car camera, a wearing terminal and a cloud platform and suitable for an elevator maintenance environment, comprising a car camera subsystem for acquiring multi-view video data in an elevator car and in a roof area, generating a video stream with a time scale, and sending the video stream to a cloud platform; uploading to an edge computing node; the wearable terminal subsystem is used for collecting images, voice and action data of a visual angle of a maintainer and uploading the data to the edge computing node; the edge computing node is used for performing time mark synchronization, space registration and event fusion on multi-source data from the car camera subsystem and the wearable terminal subsystem; and the cloud supervision platform is used for recording space-time panorama, identifying risks and generating alarm signals. According to the method, the safety monitoring precision and the response speed in the elevator maintenance process are remarkably improved, it is ensured that abnormal events are recognized in milliseconds, corresponding protection instructions are triggered, and the accident probability caused by delay or information loss in high-risk operation is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent wear and safety production, in particular to a unified monitoring system suitable for elevator maintenance environment, which integrates car camera, wearable terminal and cloud platform. BACKGROUND

[0002] Elevator maintenance is one of the most common closed space high-risk operations in urban building operation and maintenance. The maintenance area usually includes machine room, car, car roof and pit, etc. multi-layer space, closed environment, large vertical span, involving stop ladder sign (LOTO), cleaning, speed limit / maintenance mode switching, door protection, lighting ventilation and emergency rescue, etc. complex operation links.

[0003] At present, elevator maintenance mainly relies on manual experience and on-site operation, although some buildings have installed fixed cameras or video monitoring devices in the car, but there are the following technical problems in general:

[0004] 1. Data is scattered and lacks unified management. Multi-source data generated during maintenance (car fixed camera video, maintenance personnel terminal image, control system operation log, mobile phone communication record) is usually distributed in different devices and systems, lacks unified time scale and correlation index, and it is difficult to form a complete evidence chain, which brings difficulties to post responsibility identification and risk tracing.

[0005] 2. Multi-source video lacks collaborative analysis capability. The car fixed camera is a single view, and the maintenance personnel handheld terminal shooting angle is random and shaking, and they cannot be displayed and labeled synchronously in the same system, and it is difficult to restore the spatial relationship and behavior details of the maintenance site.

[0006] 3. Risk identification lags behind, and the alarm is not intelligent. The existing monitoring devices mainly rely on manual operation or simple threshold judgment, and cannot fuse the elevator controller state quantity (speed, direction, door opening and closing, maintenance mode, etc.) with visual and sensing data, so it is difficult to accurately identify key risk events such as "accidental movement", "door intrusion" and "falling trend".

[0007] At present, the invention patent with publication number CN114119291B discloses a three-dimensional personnel positioning monitoring intelligent construction site monitoring cloud platform, which relates to the technical field of intelligent construction site, including a helmet, a monitoring assembly and a cloud server. The monitoring assembly is placed on the inner side of the front of the helmet, and the monitoring assembly includes a front fixed shell, a lighting lamp, a camera, a thermal imaging head, a receiver, a connection end, a rear fixed shell, a body temperature sensor, a pulse sensor, a positioning module, a labeling module, a receiving module and a voice module.

[0008] Although some existing related patents can realize positioning supervision and monitoring of workers and take into account end-cloud cooperation, the above three common problems still exist. SUMMARY

[0009] Based on the above description, the application provides a unified supervision system integrating a car camera, a wearable terminal and a cloud platform, which significantly improves the safety monitoring accuracy and response speed of the elevator maintenance process, ensures that abnormal events are identified and corresponding protection instructions are triggered within milliseconds, and reduces the probability of accidents caused by delays or information loss in high-risk operations.

[0010] In one aspect, the technical solution for solving the above technical problems is as follows: a unified supervision system integrating a car camera, a wearable terminal and a cloud platform, comprising: a car camera subsystem that acquires multi-view video data of the inside of an elevator car and the top plate area, generates a video stream with a time stamp, and uploads the video stream to an edge computing node through a wireless link; a wearable terminal subsystem that collects images, voice and action data of the maintenance personnel's perspective, and uploads them to the edge computing node through a wireless link; an edge computing node that performs time stamp synchronization, spatial registration and event fusion on multi-source data from the car camera subsystem and the wearable terminal subsystem; a cloud supervision platform that constructs a spatio-temporal panoramic record of the maintenance process, executes a risk identification algorithm and generates an alarm signal; an alarm and instruction module that pushes voice or vibration feedback to the on-site terminal when an abnormal event is detected, and synchronously uploads supervision records.

[0011] Through the above technical solution, multi-dimensional information closed-loop collection and intelligent identification of the elevator maintenance process are realized. The system organically combines car video, maintenance personnel actions and cloud decision-making, breaking through the limitations of traditional reliance on manual monitoring and single video recognition, and can construct a spatio-temporal panoramic record of the work site in real time, forming a complete closed loop from data collection, edge analysis, cloud identification to alarm execution. The system significantly improves the safety monitoring accuracy and response speed of the elevator maintenance process, ensures that abnormal events are identified and corresponding protection instructions are triggered within milliseconds, and reduces the probability of accidents caused by delays or information loss in high-risk operations.

[0012] On the basis of the above technical solution, the application can also be improved as follows.

[0013] Further, the car camera subsystem comprises at least one set of wide-angle cameras and one set of depth cameras, which are spatially registered through a calibration matrix to generate a three-dimensional reconstruction model of the car.

[0014] It should be understood that by introducing the cooperative working mechanism of the wide-angle camera and the depth camera in the car camera subsystem, and using the calibration matrix to realize spatial registration, an accurate three-dimensional reconstruction model of the elevator car interior can be generated; the model can not only realize the spatial positioning and ranging of personnel, tools and car components, but also support subsequent spatial analysis of risk events, such as the relative distance discrimination of maintenance personnel and the door area, the car roof boundary; compared with the traditional two-dimensional video monitoring, the present application can maintain the spatial perception ability of centimeter-level precision in the dynamic environment, provide structured geometric input for cloud risk identification, and improve the explainability and reliability of event detection.

[0015] Further, the wearable terminal subsystem includes a head-mounted or chest-mounted image acquisition device, an inertial measurement unit and a voice recognition module;

[0016] The inertial measurement unit is used to record the posture change of the maintenance personnel, and realizes action recognition in combination with the video frame sequence.

[0017] It should be understood that by setting the head-mounted or chest-mounted image acquisition device and the inertial measurement unit (IMU), the wearable terminal can synchronously acquire the first perspective video and the posture change information of the maintenance personnel; in combination with the voice recognition module, voice interaction and state labeling can be realized without increasing the operation burden. This design significantly improves the integrity and traceability of the maintenance process data, realizes the multi-modal fusion acquisition of the "man-machine-environment" three; using the correspondence between the posture sequence and the video frame, high-risk actions such as bending, reaching and stretching can be accurately identified, providing high-confidence behavior labels for the intelligent recognition model, thereby reducing the false alarm rate.

[0018] Further, the edge computing node is provided with a time tag calibration module and a multi-modal fusion module;

[0019] The time tag calibration module is used to time-align the data uploaded by different devices;

[0020] The multi-modal fusion module jointly analyzes the visual features, action features and control state quantities based on a deep learning model, and outputs an event discrimination result.

[0021] Through the above technical solution, the edge computing node is built-in with the time tag calibration module and the multi-modal fusion module, so that the asynchronous data from different devices can be time-aligned and uniformly expressed before uploading; this structure significantly reduces the multi-terminal communication delay and the time drift problem between different sensing sources, realizes sub-second time tag synchronization; the multi-modal fusion module based on deep learning can simultaneously process visual, action and control state features, and realize intelligent event recognition on the terminal side. This scheme not only reduces the cloud computing load, but also changes the risk identification from passive analysis to active perception, effectively improving the real-time performance and reliability of the system.

[0022] Further, the multi-modal fusion module adopts a space-time attention network structure to weight each modality feature and calculate a comprehensive risk index. ; wherein: is a video feature sequence, is an action and posture feature, is a control system state feature, is a weight coefficient.

[0023] Through the above technical solution, the space-time attention network structure is adopted to weight and fuse the video feature, the posture feature and the control state quantity to calculate the comprehensive risk index R, which can dynamically focus on the key moment and the key part in the complex operation scene. Compared with the traditional weighted average or rule threshold method, the present scheme can adaptively adjust the feature weight, improve the sensitivity of the model to sudden events (such as accidental start, imbalance, slip), and through the continuous change curve of the risk index, the trend prediction and grading determination of the safety state can be realized, which provides a quantitative index basis for the subsequent cloud warning and braking mechanism, and has high robustness and scalability.

[0024] Further, the cloud monitoring platform comprises: a data management module for storing full-process data of the maintenance task; a risk identification module for detecting risk events including unexpected movement of the car, intrusion in the door area, falling trend, and long-time stay; a visualization module for generating a three-dimensional space reconstruction video and providing event playback and evidence chain export functions.

[0025] Through the above technical solution, the cloud monitoring platform realizes the three core functions of data management, risk identification and visualization through modular design. The risk identification module can automatically detect high-risk behaviors such as unexpected movement of the car, personnel falling and intrusion in the door area, and form a discrimination model based on big data learning; the visualization module provides three-dimensional space reconstruction video and event playback function, so that the management personnel can directly restore the accident scene and export the evidence chain; the platform realizes the change from single-point monitoring to full-process and traceable supervision, improves the objectivity and accuracy of accident investigation and responsibility identification, and has significant safety and management benefits.

[0026] Further, the alarm and instruction module feeds back the cloud-generated warning information to the wearable terminal in real time through voice synthesis or vibration feedback to prompt the maintenance personnel to take protective action;

[0027] The system further comprises a supervision terminal for the management personnel to view the real-time state, risk level and historical record of each elevator maintenance operation.

[0028] Through the technical scheme, the alarm and instruction module realizes instant on-site prompt through voice synthesis or vibration feedback, and forms a two-way linkage mechanism of "cloud recognition-terminal intervention". The maintenance personnel can obtain active warning before the danger occurs, and the reaction time is significantly shortened; at the same time, the supervision terminal can view the state and risk level of each elevator maintenance task in real time, and forms a multi-level supervision system; the addition of the module realizes a closed-loop response from risk identification to human-computer interaction, reduces the safety window period caused by information delay, and realizes a safety production mode of "early warning first, protection first".

[0029] In the second aspect, an elevator maintenance risk identification method of a unified supervision system integrating a car camera, a wearable terminal and a cloud platform, uses the unified supervision system integrating the car camera, the wearable terminal and the cloud platform, and is characterized by comprising the following steps: collecting car camera video and wearable terminal image data, and performing time synchronization; extracting video frame features, motion features and control system states; calculating a comprehensive risk index R through a spatio-temporal attention model; when R exceeds a preset threshold, triggering an alarm and a recording mechanism; the risk threshold is updated adaptively according to historical maintenance data, and satisfies: ; wherein .

[0030] Through the technical scheme, the method synchronizes multiple source videos and motion data in time and extracts features, and calculates a comprehensive risk index R through a spatio-temporal attention model; the change of the index reflects the dynamic evolution of the operation safety state; when R exceeds the threshold, the system automatically triggers an alarm and recording, and realizes instant protection response; through the adaptive threshold updating formula, the system can automatically correct the risk criterion according to historical safety samples, and has self-learning ability and environmental adaptability, thereby significantly reducing false positives and false negatives, and improving the stability and intelligent level of long-term operation.

[0031] In the third aspect, a wearable terminal comprises a terminal shell and a plug-in part arranged outside the terminal shell, the plug-in part is in clamping cooperation with a plug-in hole on a standard helmet;

[0032] The terminal shell is provided with: a camera module for collecting image data; a voice module for collecting audio data; a communication module for communicating with an edge computing node and a cloud supervision platform; a storage module for storing image data, audio data and operation logs, and system files; a central processing module for processing data; a power control module and a battery module for providing power to other modules; and a vibration module and a speaker module.

[0033] Through the technical scheme, the wearable terminal adopts a pluggable structure, can be quickly clamped with a standard safety helmet, realizes plug and play and stable wearing, the internal modular layout forms a cooperative working system of image acquisition, voice acquisition, communication, storage and power management, the setting of the vibration module and the loudspeaker module can realize multi-mode alarm feedback, improve the perception reliability of the maintenance personnel in the complex environment of noise or light, the design improves the portability and ergonomic adaptability of the equipment, and ensures the continuous stability of communication and power supply.

[0034] Further, the terminal housing includes a housing one, a housing two and an arc-shaped connecting piece, the housing one and the housing two are connected at two ends of the arc-shaped connecting piece respectively;

[0035] The power supply control module and the battery module are arranged in the housing one;

[0036] The camera module, the voice module, the communication module, the storage module and the central processor module, and the vibration module and the loudspeaker module are all arranged on the PCB and placed in the housing two.

[0037] Through the split structure design of the housing one, the housing two and the arc-shaped connecting piece, the weight can be effectively dispersed, the wearing comfort is optimized, and the independent maintenance of the modules is facilitated; the power supply control and the battery module are arranged in the housing one, so that the electromagnetic interference is reduced; the camera, voice, communication and vibration modules are concentrated in the housing two, so that the signal acquisition and processing unit is formed; the layout makes the overall gravity center of the terminal fit the human curvature, improves the use stability, and facilitates high-intensity work in a small car top environment; compared with the traditional head-mounted camera equipment, the structure of the present application is more compact, the safety is higher, and the maintenance cost is lower.

[0038] Compared with the prior art, the technical scheme of the present application has the following beneficial technical effects:

[0039] 1. The present application fuses the car camera, the wearable terminal and the cloud platform, constructs a unified supervision system covering "man-machine-environment", realizes the space-time panoramic perception of the whole process of elevator maintenance; the car camera subsystem adopts the cooperative calibration technology of wide-angle and depth cameras, can generate a three-dimensional space model in real time, realizes the spatial positioning of the maintenance personnel, equipment and work area; the wearable terminal collects the first visual image, voice and posture information, and realizes action recognition combined with the inertial measurement unit (IMU). The edge computing node is responsible for the time synchronization and spatial registration of multi-modal data, uses a deep learning model for feature fusion and event discrimination; through the above cooperative structure, the present application realizes high-precision alignment and intelligent fusion of multi-source heterogeneous information, so that subtle actions, dangerous postures and equipment abnormalities in the maintenance scene can be quickly captured and recognized, and the sensitivity and accuracy of risk identification are significantly improved;

[0040] 2、The application establishes an intelligent risk identification and feedback closed loop between the edge node and the cloud platform; the spatio-temporal attention network model on the edge side can weight and fuse video, action and control signal, calculate a comprehensive risk index R, and judge the operation safety level according to the dynamic change trend; the cloud supervision platform undertakes data storage, risk identification and visual analysis tasks, has multiple risk detection functions such as unexpected movement, door area intrusion and falling trend, and generates a three-dimensional event playback with an evidence chain; the system adopts an adaptive threshold updating algorithm: ;

[0041] The risk threshold is dynamically adjusted with historical safety samples, and has learning and evolution ability; through voice or vibration feedback mode, the alarm and instruction module can push early warning information to the wearable terminal in real time before the dangerous event occurs, and guide the maintenance personnel to take protective action immediately; the closed loop mechanism realizes intelligent response from risk perception, model calculation to human-computer interaction, greatly shortens the risk response time delay, and improves the active safety level of elevator maintenance operation;

[0042] 3、The hardware and communication architecture of the application adopts modular and distributed design, which not only ensures real-time processing of on-site data, but also ensures long-term stable operation of the system; the wearable terminal adopts a plug-in structure and can be quickly connected with a standard safety helmet, and has protection and comfort; the power control, battery module and signal acquisition, communication module are arranged in the shell one and the shell two, and the ergonomic fit is realized through the arc connecting piece; the edge computing node is deployed near the elevator machine room, has local cache and breakpoint resume functions, and ensures data integrity in the case of network interruption; the cloud platform supports multi-terminal access, task hierarchical management and visual playback, realizes whole process supervision from the operation site to the management terminal; the system forms a "end-edge-cloud" three-layer collaborative system in structure, has advantages of easy maintenance, anti-interference and strong expansion, not only improves the safety control level of elevator maintenance operation, but also provides a replicable intelligent supervision paradigm for other closed space operation scenes. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 It is a general structure schematic diagram of a unified supervision system fusing a car camera, a wearable terminal and a cloud platform of embodiment 1 of the application;

[0044] Figure 2 It is a schematic diagram of a car camera subsystem of embodiment 1 of the application;

[0045] Figure 3 It is a signal acquisition and processing flow schematic diagram of a wearable terminal subsystem of embodiment 1 of the application;

[0046] Figure 4 It is a schematic diagram of an edge computing node of embodiment 1 of the application;

[0047] Figure 5 The cloud supervision platform logical architecture diagram of embodiment 1 of the application;

[0048] Figure 6 The alarm and instruction module running logical diagram of embodiment 1 of the application;

[0049] Figure 7 The system running instance schematic diagram of embodiment 1 of the application;

[0050] Figure 8 The schematic diagram of the wearable terminal cooperating with the helmet of embodiment 3 of the application;

[0051] Figure 9 The structural block diagram of the wearable terminal of embodiment 3 of the application;

[0052] Figure 10 The overall structure schematic diagram of the wearable terminal of embodiment 3 of the application.

[0053] The figure mark: 11, the shell one; 12, the shell two; 13, the arc connecting piece; 14, the plug-in piece; 21, the power control module; 22, the battery module; 31, the camera module; 32, the PCB mainboard; 33, the central processing unit; 34, the storage module; 35, the communication module; 36, the loudspeaker; 37, the sound pickup; 38, the vibration module. DETAILED DESCRIPTION

[0054] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the relevant drawings. The drawings show embodiments of the present application. However, the present application can be implemented in many different forms, and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the specification of the present application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application.

[0056] Embodiment 1:

[0057] A unified supervision system fusing car camera, wearable terminal and cloud platform, the system as a whole adopts an "end-edge-cloud integrated collaborative architecture" to realize intelligent perception, real-time analysis, hierarchical early warning and closed-loop control of the whole process of elevator maintenance; the system is mainly composed of five functional levels: End-side perception layer 、 Edge computing layer 、 Cloud supervision layer 、 Alarm and instruction feedback layer Communication and security management layer 、 End-side perception layer ;

[0058] Each layer is interconnected through a high-speed wireless communication network (5G, Wi-Fi 6 or LoRa enhanced channel), and data interaction is achieved through a unified time synchronization protocol (NTP + PTP hybrid synchronization mechanism) and a data encapsulation standard (MQTT / JSON format); the core goal of the system is to achieve full-factor visualization of maintenance operation environment, personnel behavior and equipment state, and through intelligent collaboration between the edge and the cloud, to identify high-risk states in a timely manner and issue multi-modal warning signals.

[0059] (I) Wearable terminal channel The end-side perception layer includes a car camera subsystem and a wearable terminal subsystem, which is the data source and environmental perception basis of the system.

[0060] The car camera subsystem is deployed in the elevator car, top plate and guide rail area, and includes a wide-angle camera, a depth camera and an auxiliary infrared lighting unit; the subsystem can collect multi-view video data and generate a timestamped video stream V(t); the video signal is pre-processed by a local pre-processing module for noise reduction, frame difference detection and H.265 encoding compression, and the bandwidth occupancy can be controlled within the range of 2-4 Mbps; the system supports dynamic exposure compensation and HDR mode in low light and high light environments to ensure clarity in closed and reflective environments.

[0061] The wearable terminal subsystem is a multi-modal sensing terminal worn by the operator, which integrates a high-definition camera, an inertial measurement unit (IMU), a microphone array and a tactile feedback module; the terminal realizes "man-machine co-view" and "behavior tracking" by collecting real-time operator view video, motion posture and voice signals, and the data is processed by a local embedded processing chip (such as RK3588 or NVIDIA Jetson Nano) for lightweight inference to extract primary feature vectors , and uploaded to the edge computing node through an encrypted channel.

[0062] The edge computing layer, the edge computing node is deployed in the elevator machine room or the floor control box, with GPU / NPU acceleration unit. Its main tasks include: time synchronization and spatial registration of multi-source data to ensure that data from different devices are aligned under the same time reference; performing real-time event recognition and risk pre-screening, marking abnormal segments and uploading to the cloud; providing low-latency local alarm capability, which can directly trigger immediate feedback through the wearable terminal when detecting high-risk postures or falling trends.

[0063] This layer takes edge AI inference as the core, usually achieving a response delay of 100-200 ms, significantly improving the real-time performance and fault tolerance of the system.

[0064] (Three) cloud supervision layer, cloud supervision platform is the brain part of the system, undertake the following tasks: build the time and space panoramic record database of the maintenance process, support cross-elevator, cross-project data traceability and statistical analysis; Run the deep risk identification model, integrate video features, posture features and control system state, calculate the comprehensive risk index R; Realize the hierarchical early warning management and evidence chain generation, provide standardized supervision interface for maintenance supervision department; Through the Web visual interface, real-time picture, risk heat map and historical event playback are presented, cross-platform access (PC / PAD / mobile terminal) is realized.

[0065] The cloud uses a distributed computing framework (Kubernetes+TensorRT) to support multi-elevator parallel monitoring and model dynamic updating.

[0066] (Four) alarm and instruction feedback layer, this layer is responsible for converting the identification results into perceptible feedback to the workers and supervisors; The system supports three types of alarm channels: Supervision terminal channel : voice prompt, vibration feedback or AR prompt in the field of vision; Emergency broadcast channel Communication protocol : cloud dashboard real-time flashing high-risk events and generating records; Data encryption and identity authentication : when detecting major risks (such as falling, clamping), it can automatically link the elevator control system to execute the "emergency stop" instruction.

[0067] This closed-loop mechanism ensures that the alarm response delay is controlled within 1 second, realizing the automation of the whole process of "discovery-feedback-disposal".

[0068] (Five) communication and security management layer, in order to ensure the reliability and data security of the system, the communication layer adopts the following technical measures: Clock synchronization : end-to-edge communication uses MQTT lightweight protocol, edge-to-cloud communication uses HTTPS+WebSocket long connection to ensure real-time and stability; Redundancy and disaster recovery : all transmission contents are encrypted by AES-256, and OAuth 2.0 token mechanism is used to control access; Log audit : multi-device time stamp is realized through the combination of GPS signal and PTP hardware time correction, with an error of less than ±10 ms; Figure 1 : the edge node has local cache function, and can still continue local alarm and data storage in the case of network interruption; More specifically, : the system automatically generates an encrypted log chain for later event tracing and responsibility division.

[0069] (Six) architecture running logic

[0070] System workflow As Figure 2 shown :

[0071] The end-side acquisition module (camera + wearable terminal) starts and sends a heartbeat packet;

[0072] The edge node receives the data stream, performs synchronization, feature extraction and risk preliminary judgment;

[0073] The cloud platform receives the high-risk data packet, and the space-time attention model calculates the risk index R;

[0074] When , the cloud issues an alarm instruction to the wearable terminal and records the event;

[0075] The management terminal receives the alarm in real time and can remotely intervene or trigger the emergency mechanism.

[0076] Health status monitoring: The car camera subsystem, As shown in Adaptive calibration: Fig. 19, The car camera subsystem of the present application is installed in the elevator car and the top plate area, which is used to realize panoramic visualization and three-dimensional space reconstruction of the maintenance operation space, and is the main visual perception unit of the system. The subsystem includes multiple types of camera devices, light compensation units, mounting brackets, synchronous control modules and data transmission interfaces, and maintains high-speed data connection with the edge computing node through wireless or wired link.

[0077] (I) System composition: at least two groups of wide-angle cameras (model can be 120°-160° field of view angle CMOS module) are arranged in the system, which are fixed at the four corners of the car or the center of the top plate, and are used to collect panoramic video data of the inside of the car and the top plate operation area; each video signal is attached with a system timestamp , and is aligned by the time stamp calibration module of the edge node to form a synchronous video stream .

[0078] Depth camera group: the depth camera adopts the principle of structured light or ToF (Time-of-Flight), which is arranged on the central top of the car or the side wall of the counterweight, and is used to obtain the depth information of the operation space ; the depth data resolution is preferably 640×480 or higher, which can generate a point cloud frame rate of 15-30 FPS in real time, providing a spatial basis for three-dimensional reconstruction and risk identification.

[0079] Auxiliary infrared and lighting module: due to the limited light in the car, the system is configured with low-power infrared fill light and adjustable white light LED light source, which supports brightness adaptive adjustment; the light control module is linked with the camera acquisition clock to ensure that the brightness is stable at the exposure moment, thereby reducing the light flicker error.

[0080] Mounting bracket and protection structure: the camera is installed on the top plate and the four corners of the car through a shock-absorbing bracket, and the bracket adopts shock-resistant rubber pads with a vibration frequency of 20-150 Hz; all exposed devices are provided with protective covers (protection level IP65 or above) to prevent dust, oil stains and metal debris from contaminating the lens.

[0081] (II) Spatial calibration and geometric registration: To achieve unified fusion of multi-view data, the system performs spatial calibration and registration during the initial deployment phase, mainly including:

[0082] Internal parameter calibration uses a checkerboard method to determine the focal length of each camera. Principal point coordinates and distortion coefficient Solve the problem;

[0083] External parameter calibration involves calculating the rigid body transformation matrix between the wide-angle camera and the depth camera using laser pointing or manual target marking. ;in, For rotation matrix, It is a translation vector used to achieve spatial coordinate unification;

[0084] Point cloud registration is the process of registering point cloud data generated by a depth camera. Projected onto wide-angle coordinate system: A dense point cloud model under a unified coordinate system is obtained, enabling three-dimensional geometric reconstruction of the car space.

[0085] After calibration, the system can automatically generate a three-dimensional coordinate model of the car, which can be used for worker posture estimation, door intrusion detection, and fall trend detection.

[0086] (III) Video acquisition and data flow control: To ensure the reliability of transmission in a closed metal environment, the camera subsystem adopts a design of distributed acquisition, local buffering, and frame synchronous transmission.

[0087] The data caching mechanism has a built-in cache module in each camera (the circular buffer capacity is about 10 seconds), which can temporarily store and replenish data when the network is briefly interrupted.

[0088] The frame synchronization mechanism uses PTP (Precision Time Protocol) hardware synchronization signals to trigger frames, ensuring that the frame error of multiple video streams does not exceed ±5ms.

[0089] Compression and encapsulation: Video frames are compressed using H.265 / HEVC encoding, combined with TS (Transport Stream) encapsulation to ensure that edge nodes can quickly perform timestamp parsing and stream reassembly after receiving the data.

[0090] Data transmission protocol: The system supports multiple communication methods: wired mode: PoE gigabit Ethernet connection; wireless mode: Wi-Fi 6 (2×2 MIMO) or 5G CPE; data format: uniformly adopts JSON encapsulation header, including frame sequence number, camera number, timestamp and device ID.

[0091] (Four) signal processing and feature extraction, the edge node performs the following preprocessing procedures on the received multi-channel video stream:

[0092] Frame difference and motion detection, using weighted frame difference method and background modeling algorithm to separate static background and dynamic region, and extract motion feature region ;

[0093] Human body detection and skeleton extraction, based on lightweight pose estimation network (such as HRNet-Lite or YOLO-Pose), to identify the key joint coordinates of maintenance personnel , generate skeleton structure diagram ;

[0094] Scene semantic segmentation, using U-Net or SegFormer model to realize semantic segmentation of door area, car wall, floor and tool area, to provide environmental semantic information for risk event discrimination;

[0095] Deep fusion reconstruction, mapping the above results to a unified spatial coordinate system to generate a time series of three-dimensional point clouds , and performing filtering and surface reconstruction on spatial voxels to facilitate visualization and spatio-temporal analysis by the cloud module.

[0096] (Five) Operation and maintenance strategy, in order to improve the long-term stability and field adaptability of the system, the invention introduces a self-diagnosis and dynamic calibration mechanism in the camera subsystem: Abnormal occlusion detection: The system automatically detects the camera temperature, exposure, signal strength and transmission delay every 24 hours, and sends a maintenance reminder if the threshold is exceeded; Maintenance interface: After light changes or device vibration, automatically perform small range pose calibration to correct the calibration matrix ; More specifically, Determine whether the lens is blocked or contaminated by frame difference entropy value and optical flow statistics, and trigger an alarm when the blocking time exceeds 5s; Computing core: The system reserves a USB or Type-C maintenance port for on-site firmware updates or recalibration.

[0097] Technical effects, through the above design, the car camera subsystem can realize: 360° blind area monitoring of the car interior and the top plate work area; visual and depth fusion three-dimensional space reconstruction; low delay, multi-view synchronous transmission and local redundant storage; self-diagnosis capability for device contamination, light abnormalities, etc.

[0098] Compared with traditional single camera or static monitoring, the multi-view structured vision system of the invention can significantly improve the recognition accuracy and timeliness of risk events (such as falling, intrusion, abnormal posture) in the elevator maintenance scene, providing high-quality input for subsequent edge intelligent analysis and cloud risk assessment.

[0099] Communication interface: Edge computing node is deployed near the elevator machine room or control cabinet, which is the data hub and intelligent decision-making front end of the system. It is used for time synchronization, spatial registration, feature extraction, event fusion and risk identification of multi-source data from the car camera subsystem and wearable terminal subsystem. The design goal is to achieve millisecond-level local early warning and efficient uplink transmission without relying on cloud network.

[0100] The node adopts modular hardware and software architecture, which is composed of hardware platform, time calibration module, multi-modal fusion module, risk discrimination module, local cache and communication module, etc.

[0101] (I) Hardware platform composition; Storage and power supply system: Select an edge computing platform with AI inference acceleration unit (such as NVIDIA Jetson Xavier NX, RK3588 or Intel Movidius); CPU frequency ≥ 2.0 GHz, GPU computing power ≥ 1 TFLOPS, memory ≥ 8 GB; support deep learning inference framework such as TensorRT, ONNX Runtime, realize local model loading and real-time calculation; Environmental protection and temperature control: Equipped with dual gigabit Ethernet port (for camera data and uplink data shunting); built-in Wi-Fi 6 and 5G module, realize wireless transmission redundant channel; provide USB-C and RS485 interface, convenient for on-site debugging and control system interconnection; More specifically, Local SSD capacity ≥ 256 GB, used for caching video clips and log files; equipped with UPS voltage stabilizing power module, which can maintain operation for 10-15 minutes in case of power failure, ensuring complete data writing and safe shutdown; Data middle platform layer, Intelligent analysis layer, The shell protection level is ≥ IP54, suitable for high humidity and high dust environment in elevator shaft; built-in temperature control fan and heat dissipation aluminum plate, ensure long time stable operation.

[0102] (II) Time calibration module, multi-device asynchronous acquisition will cause time deviation, this module realizes the time consistency of multi-source data through multi-level calibration algorithm.

[0103] Synchronization mechanism: the system uses NTP (Network Time Protocol) for coarse synchronization (accuracy ±1 s), and combines PTP (Precision Time Protocol) hardware clock to realize fine synchronization (accuracy ±1 ms); record the receiving timestamp of each frame of video, IMU packet and voice frame , correct by the following formula: ;

[0104] Where, is the network delay estimation value, Compensation amount for local clock drift.

[0105] Timing alignment algorithm: sliding window matching method is adopted to calculate the cross-correlation between video key frames and IMU attitude frames, maximizing the inter-frame similarity: ; By finding the maximum correlation delay , automatic alignment is realized; the synchronization error is controlled within ±10 ms.

[0106] (Three) Multimodal fusion module; this module is responsible for joint analysis of visual, attitude, speech and control state features from the car camera and wearable terminal, forming a unified spatio-temporal event representation.

[0107] Input data format: ;

[0108] Among them: : Car multi-view video frame features; : Wearable terminal video frame features; : Attitude and acceleration features; : Elevator control system state quantity (such as door opening and closing, running signal); : Speech semantic label.

[0109] Feature encoding and fusion: video features are reduced to 256 dimensions after being extracted by ResNet50 or MobileViT; attitude and control features are mapped to 128 dimensions by fully connected layer; then input into Spatio-Temporal Attention Network (STAN), calculate the weight coefficient of each modality: ; Among them, learned automatically through attention layer.

[0110] Spatial registration and scene reconstruction: according to the car camera calibration matrix , the wearable terminal video frame is projected into the car coordinate system; the time convolution is performed on the fused feature sequence to generate event embedding vector , which is used for risk discrimination.

[0111] (Four) Risk discrimination module; this module is the core decision unit of the edge node, the main task is to quickly identify potential risk events and trigger local alarm.

[0112] Risk index calculation, comprehensive risk index output by the weighted model, and compared with the dynamic threshold: ; Among them, is the Sigmoid activation function, is the trainable parameter; when , the system immediately triggers an alarm; the threshold Periodically updated by the cloud and fine-tuned according to the following adaptive formula: ; wherein .

[0113] Event classification and local response, risk event classification includes but is not limited to: door area intrusion event: human key points are detected to enter the door area mask; falling trend event: IMU detects a sudden change in vertical acceleration and video depth field change rate > 0.2 m / s; unstable posture event: attitude inclination > 30° and lasts > 2s; long time stay event: human skeleton has no significant displacement for > 60s.

[0114] The system sets the response priority according to the event type, when the recognition result confidence ≥ 0.8, the "local voice alarm" signal can be generated directly on the edge side, without waiting for the cloud to confirm.

[0115] (Five) Local cache and communication mechanism

[0116] Cache strategy: the system uses a ring cache structure to continuously save the last 30 minutes of multi-modal data; when a risk event occurs, the node locks the 30-second segment before and after the event as permanent records and generates an index file; if the cloud is offline, the cache area can save about 48 hours of data.

[0117] Upload mechanism: the edge node and the cloud maintain real-time communication through HTTPS or MQTT over TLS channel;

[0118] Data packet upload adopts hierarchical strategy: high-risk event: upload video + sensor data immediately; low-risk event: only upload summary features and log records.

[0119] Local log and audit: the system automatically generates an encrypted log chain, records the risk discrimination process, alarm time, response status and data check code, and supports later responsibility tracing.

[0120] (Six) Edge intelligence update mechanism, to maintain the long-term adaptability of the algorithm, the edge node and the cloud platform establish a model federated update mechanism: 1. The cloud trains the global model based on historical task data; 2. The edge node uploads local gradient periodically; 3. The cloud performs weighted aggregation update and issues new parameters; 4. The whole process adopts differential privacy protection to ensure the safety of personal data.

[0121] This mechanism enables each node to continuously evolve the risk identification model without revealing the original data.

[0122] Through the above structure and algorithm design, the edge computing node of the application has the following technical advantages: low delay response: risk identification and alarm delay < 500ms; high reliability: even if the cloud is disconnected, it can still run independently and locally alarm; high robustness: can adapt to various network environments and device configurations; evolvability: support cloud edge collaborative model update and parameter self-learning; security compliance: full-link data encryption and log traceability ensure reliable supervision.

[0123] The module realizes the functional transition from passive monitoring to active security intervention in the whole system, and is the core link of realizing intelligent elevator maintenance supervision.

[0124] Visual supervision layer, The cloud supervision platform is the global control and decision-making core of the system, and undertakes functions such as data aggregation, risk identification, threshold adaptive update, visual display and operation and maintenance; The platform is built on a distributed container architecture, and realizes the whole process perception-intelligent identification-classified early warning-evidence preservation-closed loop management of the elevator maintenance process through two-way data interaction with the edge computing node.

[0125] (I) System architecture, the cloud supervision platform adopts a three-layer logical structure of "data platform + intelligent analysis + visual supervision", which corresponds to: Data encryption and access control: Responsible for receiving and storing multi-modal data packets uploaded from edge nodes, including video clips, posture features, control state quantities and event logs; Use distributed database (such as TiDB or MongoDB) to store structured and unstructured data; Cooperate with object storage system (such as MinIO) to manage large capacity video and point cloud files, realize cold and hot layered storage; Establish a maintenance task index table task_id, elevator_id, timestamp, risk_level, support fast query and traceability. Log and audit mechanism: Deploy risk identification models, anomaly detection models and trend prediction modules based on deep learning; Responsible for high-precision identification and risk level evaluation of uploaded data; Support multi-model collaborative reasoning and federated learning parameter aggregation mechanism. Data desensitization and compliance: Provide 3D panoramic playback and BIM superposition interface; Can show risk heat map, operation path trajectory and historical event distribution; Management personnel can view the state of each elevator in real time through the Web or mobile terminal.

[0126] The system runs in a cloud-native environment (Kubernetes + Docker container cluster), which can be horizontally expanded to support parallel monitoring of thousands of elevators.

[0127] (II) Data flow and communication mechanism

[0128] Data uplink, edge nodes upload data packets in HTTPS / MQTT over TLS mode, and the cloud receives and performs identity verification and flow control through the API gateway (API Gateway); each data packet contains a timestamp, device number, event label, and content digest, and the transmission format is as follows:

[0129]

[0130] After data analysis, it is automatically warehoused and distributed to the corresponding task channel.

[0131] Data downlink, the cloud issues alarm instructions, model updates, or threshold adjustment instructions to edge nodes and wearable terminals through the MQTT topic subscription mechanism; instructions use a signature mechanism to prevent forgery or tampering.

[0132] (Three) Risk identification model, the risk identification module in the cloud supervision platform uses a multi-modal spatio-temporal attention network (MM-STAN) structure to accurately identify abnormal behavior and potential dangerous states during elevator maintenance.

[0133] Input feature set: wherein:

[0134] : Multi-view visual features (ResNet feature vectors extracted by edge nodes);

[0135] : Action posture features (IMU quaternion and acceleration sequence);

[0136] : Elevator control signals (door switch, emergency stop, running state, etc.);

[0137] : Speech semantic vector.

[0138] Spatio-temporal attention mechanism, the network consists of three attention layers: time attention layer: calculates the dynamic change weight between consecutive frames; spatial attention layer: focuses on the interaction between human key points and dangerous areas; modal attention layer: automatically adjusts the importance coefficient of vision, action, and control signal.

[0139] Comprehensive risk output: wherein is the Sigmoid function.

[0140] Risk grading standard, the platform divides the risk level according to the risk index

[0141]

[0142] ​(Four) Dynamic threshold adaptive updating algorithm

[0143] The cloud platform has a risk threshold self-learning function based on historical task data to reduce false positive rate and adapt to different workers and environmental conditions.

[0144] Data statistics, statistics of safety sample average risk index in the last N tasks and abnormal sample average index ;

[0145] Threshold updating model, updating risk discrimination threshold according to sliding window decay model: ; wherein, is the smoothing coefficient (recommended 0.8-0.9).

[0146] Self-calibration and individual adaptability, if the same worker triggers a slight false alarm event for three consecutive times, the system automatically adjusts the threshold value by 0.02; if there are two consecutive real risk events that are not recognized, the threshold value is automatically reduced by 0.03. The cloud regularly broadcasts new threshold values to each edge node to ensure consistency across the network.

[0147] (Five) Risk visualization and evidence chain generation, the cloud monitoring platform provides a visual monitoring and event traceability interface, realizing a closed-loop management from risk discovery to evidence preservation.

[0148] Three-dimensional reconstruction display, combining the depth data of the car camera and the perspective of the wearable terminal, a three-dimensional point cloud model of the elevator space is generated in real time; the risk heat map is superimposed in the BIM model, and the dangerous area is highlighted in red; the manager can rotate, zoom in and out to view the worker's trajectory, posture changes and risk point distribution.

[0149] Event playback and evidence chain export, each risk event is automatically associated with video, IMU curve, voice recording and system log; a complete timeline playback and responsibility identification report can be generated; PDF / MP4 format export and encrypted signature are supported to ensure the legal effect of the evidence.

[0150] Data dashboard, providing multi-dimensional statistical views, including task quantity, risk level distribution, device health status, etc.; it can be displayed by time, area, maintenance unit; it supports abnormal trend prediction and risk index time series analysis chart.

[0151] (Six) Safety and privacy protection, in order to ensure the safety and legal compliance of the monitoring system data, the cloud platform designs a multi-layer protection mechanism: More specifically, AES-256 encryption transmission is used in the whole link, and TLS1.3 is used for communication between servers; user identity is authorized through OAuth 2.0, and hierarchical access permission (maintenance personnel, supervisor, system administrator). Voice feedback Vibration feedbackAll operations and alarm event records are recorded into a blockchain log chain, with tamper-proof characteristics; according to the task number or event timestamp, it can quickly trace back to ensure regulatory transparency. Visual prompt Video and voice data are automatically processed for face and voice feature blurring in the cloud to prevent privacy leakage; the platform meets the requirements of GB / T 35273-2020 Personal Information Protection Specification.

[0152] (Seven) System operation and intelligent linkage, the cloud platform can be connected with the city elevator safety supervision center system, realizing cross-regional supervision and abnormal linkage: when detecting a III-level risk event, the system automatically pushes it to the competent department emergency interface; through API connection with the elevator control system, the "emergency stop" or "lock elevator" instruction can be remotely triggered; the platform can automatically generate a safety performance report according to the maintenance frequency for examination and record.

[0153] Through the above design, the cloud supervision platform realizes the following innovative technical effects: multi-modal fusion recognition: comprehensive analysis of video, posture, voice and control signal, improving recognition accuracy and robustness; dynamic self-learning threshold: the risk discrimination model can automatically adapt to different environments according to the operation history; three-dimensional visual supervision: providing an operable interface for spatialized risk presentation and event traceability; full-link security protection: encryption and tamper-proofing from collection to storage; cloud-edge collaborative mechanism: model updates and risk threshold are synchronously distributed, forming a complete closed-loop supervision system.

[0154] In summary, the cloud supervision platform of the present application is significantly superior to the prior art in real-time, intelligence and traceability, and provides reliable data support and intelligent decision basis for the safety control of elevator maintenance work.

[0155] Multi-channel concurrent execution The alarm and instruction module is the safety closed-loop execution unit of the whole system, which undertakes the core functions of risk event feedback, operator prompting, cloud-edge collaborative disposal and emergency linkage control. The module connects the wearable terminal, edge node and cloud platform, realizes the whole process control logic of "risk identification-multimodal feedback-response confirmation-log trace".

[0156] The module is composed of an alarm triggering submodule, an instruction generating submodule, a feedback executing submodule, an emergency linkage submodule and a fault-tolerant communication submodule.

[0157] (One) Overall structure and communication relationship, the alarm and instruction module is located in the logical intermediate layer between the cloud supervision platform and the terminal device, realizing the following through the bidirectional data channel: uplink path: risk identification results and event labels are transmitted from the edge node to the cloud supervision platform; downlink path: the cloud generates alarm instruction packages according to the risk level and sends them to the wearable terminal or the elevator control system.

[0158] Communication adopts MQTT over TLS protocol, and the message topic is designed as follows:

[0159]

[0160] (ii) Alarm triggering logic; alarm triggering is based on comprehensive risk index and event type multi-level judgment. The system adopts a hierarchical alarm strategy, combined with risk level, duration and confidence to make a comprehensive judgment.

[0161] The judgment formula is: ; wherein: : comprehensive risk index; : risk duration; : event recognition confidence; is the weighted coefficient (satisfying ). The alarm level discrimination table is as follows:

[0162]

[0163] The system supports dynamic adjustment of threshold values at all levels to adapt to different maintenance units' safety standards.

[0164] (iii) Instruction generation submodule; when the cloud supervision platform receives the risk events reported by the edge node, the instruction generation process is automatically triggered.

[0165] 1. Instruction package structure, instructions are packaged in unified JSON format, and the fields are as follows:

[0166]

[0167] Among them, "intensity" represents the vibration intensity level (1-3), and "duration" represents the vibration time (milliseconds).

[0168] 2. Instruction classification, voice broadcast instruction: call the text-to-speech (TTS) engine to generate natural speech, support Mandarin, Cantonese or English multi-language broadcast; vibration instruction: encoded as intensity and frequency parameter combination, for example "2-level intensity @ 200Hz"; visual instruction: used for wearable terminal to display AR prompt or flashing warning icon; emergency control instruction: used for linkage elevator control system to trigger mechanical brake or stop lock.

[0169] (iv) Feedback execution submodule, the wearable terminal receives the instruction package and enters the execution process:

[0170] 1. Analysis and verification, the terminal verifies the signature field to ensure that the message source is legal; if the verification fails, it will be discarded and recorded.

[0171] 2. Action execution mechanism; 1. Redundant sending: Voice module automatically plays prompt sound or reads out text according to instruction content; 2. Delay compensation : Linear vibration motor driver generates specific frequency pulse; 3. Offline buffering : LED or head-mounted interface displays red flashing signal; 2. Feedback timeout mechanism : Voice and vibration can be triggered simultaneously, feedback delay <300ms.

[0172] 3. Acknowledgment feedback, after execution is completed, terminal sends acknowledgment package to cloud:

[0173]

[0174] Cloud updates task status accordingly, ensuring feedback chain closed loop.

[0175] (Five) Emergency linkage sub-module, to cope with serious risk events (such as personnel falling, car abnormal movement), the system designs cloud-control system linkage interface.

[0176] 1. Interface design, through industrial communication protocol (Modbus TCP or OPC UA) to interface with elevator control mainboard; Cloud issues "emergency stop instruction package" through secure channel, control system immediately executes brake and locks car operation; At the same time, send SMS or push notification to maintenance supervisor and emergency management terminal.

[0177] 2. Linkage process, (1) Cloud detects L3 risk event; (2) Trigger emergency linkage module to generate control instruction; (3) Edge node and elevator control system confirm instruction status; (4) Control system executes emergency stop→feedback execution result to cloud; (5) Cloud records operation log and generates emergency report.

[0178] 3. Safety mechanism, all linkage instructions need double authentication (digital signature + time token); Only in "maintenance mode" can remote emergency stop be executed, to avoid false triggering; After emergency stop, system needs manual confirmation to unlock.

[0179] (Six) Fault tolerance and delay control mechanism, to ensure stable operation in complex construction environment, the invention sets multiple fault tolerance mechanisms in alarm and instruction module: 3. Delay control formula : Same instruction package can be sent through Wi-Fi and 5G dual channels in parallel, to ensure at least one arrival; Multi-modal feedback integration : Use timestamp comparison and priority queue mechanism to automatically discard lagging instructions, to prevent expired execution; Low latency and high reliability : When network is interrupted, edge node can temporarily store pending alarm packages, and synchronize after network recovery; Emergency linkage intelligentization : If wearable terminal does not return execution status within 3 seconds, system re-sends instruction and counts; Closed-loop traceability mechanism : ; Wherein, is instruction generation delay, for transmission delay, for terminal response delay. The average response time of the actual system is 0.62s.

[0180] (Seven) logging and tracing, each alarm event generates a log chain record in the cloud and the edge node, including: alarm level, timestamp, event type; instruction ID, execution status and delay; confirmation receipt and operator response. The log adopts a hash chain structure to prevent tampering and support later security audit and accident analysis.

[0181] Through the above structure and algorithm, the alarm and instruction module of the application has the following obvious advantages:

[0182] Strong dynamic adaptability : voice, vibration, visual multi-channel alarm triggers simultaneously, intuitive prompt, rapid response; High real-time performance : end-to-end feedback delay is controlled within 1 second, meeting the real-time safety needs of elevator maintenance work; High accuracy Intelligent collaboration : automatic emergency stop and anti-falling response are realized in cooperation with the elevator control system; Panoramic traceability : each alarm has execution and confirmation record, with safety supervision and legal evidence value; Scalability : threshold and alarm intensity can be adjusted in real time according to risk level and individual difference.

[0183] In summary, the application realizes real-time intervention and intelligent protection of the whole process of elevator maintenance by constructing a cloud-edge-end collaborative multi-modal alarm instruction mechanism, significantly improving the safety prevention level and supervision efficiency of the work personnel

[0184] System running instance: to verify the comprehensive performance and safety protection effect of the system, the following describes the system running process, risk identification process and alarm response mechanism in combination with a typical elevator maintenance scene; this embodiment takes an elevator with a rated load of 1000kg and a rated speed of 1.75m / s as the test object, and demonstrates it in the conventional "car roof maintenance" operation mode.

[0185] (One) scene setting, maintenance personnel A wears the wearable terminal (head-mounted camera device + IMU + voice module) of the system and enters the elevator car roof maintenance area; four wide-angle cameras and a set of depth cameras are arranged inside the car and on the top of the shaft, and the edge computing node is deployed in the machine room control cabinet, maintaining 5G wireless connection with the cloud monitoring platform; the maintenance task is assigned a task number and operation plan by the cloud system, and the edge node synchronously initializes the timestamp, and the wearable terminal automatically enters the "task activation mode".

[0186] After the system starts, the following initialization process is automatically completed: 1. Synchronize the clock of each device and establish an encrypted connection; 2. The car camera subsystem starts video acquisition, and the wearable terminal begins to record motion and voice streams; 3. The edge node performs data calibration and frame synchronization, and starts the multi-modal fusion model.

[0187] At this time, the system enters the "active supervision state", and real-time monitoring of the posture, spatial position and working environment changes of the maintenance personnel.

[0188] (II) Running process

[0189] 1. Normal maintenance stage: Maintenance personnel A walks along the top plate of the car to the guide rail area, and the wearable terminal IMU detects that the attitude angle change is less than 10°, the risk index , the system is in a safe state; the cloud supervision platform displays the personnel position and car model superimposed picture in real time, and the risk heat map is a green area; all multi-modal data are uploaded to the cloud after being cached and compressed by the edge node, and the cloud forms a panoramic record of the work.

[0190] 2. Potential risk stage: Door area proximity behavior recognition, when the maintenance personnel is within 30 cm of the car door area, the wide-angle camera detects the intersection of human skeleton key points and door frame area, and the wearable terminal attitude angle (Pitch) reaches 25°;

[0191] The edge node executes the spatio-temporal attention model (STAN) to calculate the risk index: , output , close to the warning threshold ; the system determines that it is "mild risk (L1 level)", the wearable terminal vibrates for 0.5 seconds and prompts by voice: "Close to the door area, please keep a safe distance".

[0192] At the same time, the cloud records the event label door_proximity and updates the risk trend chart.

[0193] 3. High risk stage: Unstable tilt recognition, if the maintenance personnel loses balance at the edge of the door area, the IMU detects that the acceleration peak value |a|=2.9g, the attitude angle Roll=38°, and the duration is 2.4s;

[0194] The edge node calculates the risk comprehensive score: ; the system identifies it as "moderate risk (L2 level)".

[0195] The wearable terminal immediately executes a medium-intensity vibration (1 second @ 200Hz) and prompts by voice: "Detecting attitude imbalance, please pay attention to standing steady"; the risk heat map of the cloud supervision platform interface changes from green to orange, and automatically locks the current video frame ±5 seconds as a key fragment for uploading.

[0196] The management terminal pops up an alarm window to display the maintenance personnel's position and attitude angle change curve.

[0197] 4. Extreme risk stage: fall trend identification and emergency stop linkage. If the maintenance personnel loses stability and the body center of gravity changes suddenly, the IMU detects a vertical acceleration |a_z|>3g, and the depth camera detects a spatial height change >0.25m, the edge node determines a "fall trend event".

[0198] The system calculates , immediately triggers an L3-level alarm: the wearable terminal performs strong vibration (2 seconds x 3 times @ 250Hz), and voice broadcast "detecting fall risk, please stop immediately";

[0199] The edge node sends an emergency stop command to the elevator control system through Modbus TCP, and the mechanical brake executes a delay of less than 150ms;

[0200] The cloud supervision platform generates a red highlight warning, and the event number is EVT20241020_001;

[0201] Synchronously push a message to the maintenance supervisor and the emergency linkage terminal.

[0202] The whole risk identification to braking response process delay is about 0.68 seconds, realizing millisecond-level safety protection.

[0203] (Three) Event data recording and visual playback,

[0204] 1. Data aggregation and storage; after the risk event occurs, the edge node locks the multi-modal data (video, IMU, voice, control signal) before and after the event for 30 seconds, uploads it to the cloud storage bucket; the cloud supervision platform generates an index entry in the database:

[0205]

[0206] The log chain records the whole process of the event, including detection time, alarm triggering time, response time and confirmation status.

[0207] 2. Three-dimensional playback interface: the cloud can reconstruct a three-dimensional model of the car, superimposed with the maintenance personnel's trajectory and attitude vector;

[0208] The risk occurrence point is displayed as a red ball, with a risk index curve;

[0209] The manager can view the video, IMU acceleration waveform and voice record synchronous playback by dragging the timeline.

[0210] 3. Report and analysis: The platform automatically generates event analysis reports, including risk causes, action identification results, alarm execution logs, and emergency stop response times. The report format can be exported as PDF or HTML interactive documents for security assessment and responsibility definition.

[0211] (Four) System stability and performance; in multiple experiments, the system performed as follows:

[0212]

[0213] The system did not have abnormal disconnection or false triggering during the 72-hour continuous operation test, indicating that the architecture of the present application has good stability and robustness.

[0214] (Five) Expansion application example; the system of the present application is not only suitable for single elevator maintenance scene, but also can be extended to multi-elevator cooperation and regional maintenance management through cloud supervision platform.

[0215] When multiple work groups are working in different elevators at the same time, the cloud platform can automatically allocate channels and task numbers;

[0216] The management terminal can view the state of each elevator and the real-time risk heat map under a unified interface;

[0217] Through the cloud data analysis module, the safety behavior model of different maintenance personnel can be evaluated, and the "safety performance index (SPI)" can be constructed for later evaluation.

[0218] In addition, the system can also be interconnected with the city elevator safety supervision database to realize cross-unit risk trend prediction and safety credit evaluation.

[0219] (Six) Summary of technical effects; through the operation verification of this embodiment, it is proved that the performance of the system of the present application in the actual maintenance work has the following significant technical advantages: Figure 8-10 : The whole link delay from risk detection to alarm response is less than 1 second; Camera module Voice module : The comprehensive identification accuracy rate of the spatio-temporal attention model is more than 93%; Communication module : Cloud-edge-end linkage realizes autonomous emergency stop and risk tracing; Storage module : Multi-modal data playback reproduces the whole process to form a usable legal evidence chain; Central processing module : It can manage hundreds of elevators and multiple maintenance personnel in parallel, and support cross-regional supervision deployment.

[0220] In summary, the present application establishes a real-time elevator maintenance risk supervision system based on edge intelligence and cloud cooperation by fusing visual, posture and voice information, which significantly improves the safety, reliability and traceability of maintenance work.

[0221] Example 2:

[0222] An elevator maintenance risk identification method fusing car camera, wearable terminal and cloud platform, using the unified supervision system of the fusing car camera, wearable terminal and cloud platform described in the embodiments, comprising the following steps: S01 collecting car camera video and wearable terminal image data, and performing time synchronization; S02 extracting video frame features, action features and control system states; S03 calculating a comprehensive risk index R through a spatio-temporal attention model; S04 when R exceeds a preset threshold, triggering an alarm and recording mechanism; S05 the risk threshold is updated adaptively according to historical maintenance data, satisfying: ; wherein .

[0223] Specifically: (1) System deployment and data collection; multiple-view high-definition cameras are arranged at the elevator machine room, the top of the car and inside the car to obtain the spatial posture, tool operation and environmental state images of the maintenance personnel; the wearable terminal (including a head-mounted camera, a voice acquisition module and an inertial measurement unit IMU) records the first-view video, voice commands and body action data of the maintenance personnel in real time; the above multi-source data is transmitted synchronously to the edge computing node through a wireless link (such as Wi-Fi 6 or 5G module).

[0224] After the edge computing node receives the multi-source stream data, it first performs time synchronization and spatial registration. Through the time alignment mechanism based on NTP (Network Time Protocol) and the SLAM spatial repositioning algorithm, it ensures that the data from different sensors strictly correspond in time and space, so that the subsequent feature fusion has consistency.

[0225] (2) Feature extraction and state modeling; the synchronized data is input into the feature extraction module, and the system uses a lightweight convolutional neural network (such as MobileNetV3) to extract the appearance features of the video frames, to identify typical operation postures, device components and dangerous actions (such as not wearing a safety belt, reaching out of bounds, approaching rotating components, etc.).

[0226] The IMU data of the wearable terminal generates a three-dimensional action sequence vector after filtering and attitude calculation, and the voice recognition module extracts key semantic instructions (such as "stop", "up", "maintenance complete", etc.) and matches them with the operation time.

[0227] At the same time, the edge node accesses the running state information of the elevator control system, including motor speed, door lock state, car position and safety circuit signal, etc., to construct a unified state vector , wherein represents the image feature, represents the action feature, represents the control state feature.

[0228] (Three) Risk index calculation; the system uses a spatio-temporal attention model to analyze the state vector sequence; the model reflects the influence degree of the risk event in time through the time attention weight

[0229]

[0230] wherein is a multi-modal fusion function used to calculate the risk saliency of the current state; The greater the value of is, the higher the degree of deviation of the operation behavior from the safety category; the system calculates in real time on the edge side and outputs to the cloud platform.

[0231] (Four) Alarm triggering and event recording; when the risk index exceeds the set threshold , the edge node immediately sends a vibration or voice prompt (such as "pay attention to the posture" and "please stay away from the rotating part") to the wearable terminal, and at the same time triggers a video recording and encryption upload mechanism to upload the multi-source image fragments (10 seconds before and after) of the current period to the cloud to form an event record.

[0232] After receiving the abnormal event, the cloud platform automatically generates a time-stamped safety log and synchronizes the data to the supervision center database for subsequent traceability analysis and statistical evaluation.

[0233] (Five) Threshold adaptive update; to reduce the probability of false positives and false negatives, the embodiment sets a threshold dynamic update mechanism; the system periodically calculates the average risk level of the safety operation sample in the cloud , and adaptively adjusts the discrimination threshold according to the following formula:

[0234] ; wherein, is a smoothing coefficient that controls the sensitivity of threshold updating; when the operation environment is stable, the system maintains a higher threshold to avoid frequent alarms; when the environmental noise or dangerous behavior increases, the threshold is automatically lowered to enhance the warning sensitivity.

[0235] (Six) Running process instance; taking elevator car top maintenance as an example: the maintenance personnel turn on the "maintenance mode" and wear the wearable terminal; the system starts collecting video and motion data; if the personnel cross the safety fence or stretch their hands to the edge area of the car top, the "dangerous boundary intrusion" label appears in the video features, the IMU action angle changes abnormally, and the risk index rapidly rises; when​​​​ When the system detects a potential safety hazard, it immediately issues a voice "pay attention to safety" warning and uploads the event to the cloud for record-keeping. If similar events occur multiple times, the system automatically updates the threshold strategy and optimizes the model weights to reduce the false positive rate in subsequent similar scenarios.

[0236] The method and system of the embodiment achieve multi-modal and spatio-temporal integrated risk identification for elevator maintenance processes. The significant technical effects include: 1. Through video and motion feature fusion, potential dangerous operations can be accurately captured, improving the real-time and accuracy of risk identification; 2. Based on the adaptive threshold updating mechanism, the system has the ability of continuous learning and environmental adaptation; 3. Combined with edge node instant feedback and cloud panoramic record, the whole process of "risk identification-alarm-record-self-learning" is realized, greatly improving the safety and intelligence level of elevator maintenance operation.

[0237] Embodiment 3

[0238] Reference Power control module and battery module A wearable terminal includes a terminal housing and a plug-in part provided on the outside of the terminal housing, and a plurality of modules provided in the terminal housing, the plug-in part being engaged with a plug-in hole on a standard helmet; the terminal housing includes a housing one, a housing two and an arc-shaped connecting part, the housing one and the housing two being connected to the two ends of the arc-shaped connecting part respectively; the wearable terminal is detachably fixed on the standard helmet through the plug-in part;

[0239] The terminal housing is provided with: ​ for collecting image data; ​ for collecting audio data, including a sound pickup and an audio processing chip; ​ for communicating with edge computing nodes and cloud monitoring platforms, including 5G, Wi-Fi 6 and Bluetooth modules; ​ for storing image data, audio data, running logs and system files; ​ for processing data; ​ for providing power to other modules; further including a vibration module, a speaker module and an inertial measurement unit.

[0240] The power control module and the battery module are arranged in the housing one, the battery module is a lithium battery, and the battery control module is used for stably inputting power to the battery module and stably outputting power in the battery module; the power control module includes a charging interface, and a through hole corresponding to the charging interface is formed on the housing one.

[0241] The camera module, the voice module, the communication module, the storage module and the central processing unit module, as well as the vibration module and the speaker module are arranged on the PCB board and placed in the housing two; the camera module is arranged at the end of the housing two away from the arc-shaped connecting part.

[0242] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A unified monitoring system fusing a car camera, a wearable terminal and a cloud platform, characterized in that, Comprise: The car camera subsystem acquires multi-view video data of the elevator car interior and the ceiling area, generates a video stream with a time stamp, and uploads the video stream to an edge computing node through a wireless link; The wearable terminal subsystem collects image, voice, and motion data from the maintenance personnel's perspective and uploads them to the edge computing node through a wireless link; The edge computing node synchronizes the time stamps, spatially registers, and event fuses the multi-source data from the car camera subsystem and the wearable terminal subsystem; The cloud monitoring platform constructs a spatio-temporal panoramic record of the maintenance process, executes a risk identification algorithm, and generates an alarm signal; The alarm and instruction module pushes voice or vibration feedback to the on-site terminal when an abnormal event is detected and synchronously uploads the monitoring record. 2.The unified monitoring system of fusion of elevator camera, wearable terminal and cloud platform according to claim 1, characterized in that, The car camera subsystem includes at least one set of wide-angle cameras and one set of depth cameras, which are spatially registered through a calibration matrix to generate a three-dimensional reconstruction model of the car. 3.The unified monitoring system of fusion of elevator camera, wearable terminal and cloud platform according to claim 1, characterized in that, The wearable terminal subsystem includes a head-mounted or chest-mounted image acquisition device, an inertial measurement unit, and a voice recognition module; The inertial measurement unit records the posture changes of the maintenance personnel and realizes motion recognition in combination with the video frame sequence. 4.The unified monitoring system of fusion of elevator camera, wearable terminal and cloud platform according to claim 1, characterized in that, The edge computing node has a time stamp calibration module and a multi-modal fusion module; The time stamp calibration module is used to time-align the data uploaded by different devices; The multi-modal fusion module jointly analyzes visual features, motion features, and control state quantities based on a deep learning model and outputs event discrimination results. 5.The unified monitoring system of fusion of elevator camera, wearable terminal and cloud platform according to claim 1, characterized in that, The multi-modal fusion module uses a spatio-temporal attention network structure to weight each modal feature and calculate a comprehensive risk index: ; wherein: are video feature sequences, are action and pose features, are control system state features, are weight coefficients. 6.The unified monitoring system of fusion of elevator camera, wearable terminal and cloud platform according to claim 1, characterized in that, The cloud monitoring platform includes: A data management module for storing full-process data of maintenance tasks; A risk identification module for detecting risk events including unexpected car movement, door intrusion, falling trend, and long-time stay; A visualization module for generating three-dimensional spatial reconstruction video and providing event playback and evidence chain export functions. 7.The unified monitoring system of fusion elevator camera, wearable terminal and cloud platform according to claim 1, characterized in that, The alarm and instruction module feeds back the cloud-generated early warning information to the wearable terminal in real time through voice synthesis or vibration feedback, prompting the maintenance personnel to take protective action; The system further includes a supervision terminal for management personnel to view the real-time status, risk level, and historical records of each elevator maintenance operation.

8. An elevator maintenance risk identification method of a unified monitoring system of a fusion of a car camera, a wearable terminal and a cloud platform, using the unified monitoring system of the fusion of the car camera, the wearable terminal and the cloud platform according to any one of claims 1-7, characterized in that, Comprise the following steps: Collect car camera video and wearable terminal image data and synchronize the time stamps; Extract video frame features, motion features, and control system states; Calculate the comprehensive risk index R through the spatio-temporal attention model; When R exceeds the preset threshold, trigger the alarm and recording mechanism; The risk threshold is updated adaptively based on historical maintenance data, satisfying: ; wherein .

9. A wearable terminal, comprising: Comprise: A terminal housing and a plug-in part provided on the outside of the terminal housing, which is in clamping cooperation with a plug-in hole on a standard helmet; The terminal housing is provided with: A camera module for collecting image data; A voice module for collecting audio data; A communication module for communicating with the edge computing node and the cloud monitoring platform; A storage module for storing image data, audio data, and operation logs, and system files; A central processing module for processing data; A power supply control module and a battery module are used to provide power to other modules; The terminal further comprises a vibration module and a speaker module. 10.The wearable terminal according to claim 9, wherein, The terminal housing comprises a housing one, a housing two and an arc-shaped connecting piece, the housing one and the housing two are connected to two ends of the arc-shaped connecting piece respectively; The power supply control module and the battery module are arranged in the housing one; The camera module, the voice module, the communication module, the storage module and the central processing unit module, and the vibration module and the speaker module are arranged on the PCB board and placed in the housing two.

Citation Information

Patent Citations

  • A smart construction site supervision cloud platform with three-dimensional personnel positioning supervision

    CN114119291B