Intelligent identification and early warning method and system for dangerous behavior of elevator operator

CN122799487APending Publication Date: 2026-09-22武汉市特种设备检验检测研究院 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610745452.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0008]针对现有技术难以在复杂电梯环境下实时、准确识别多样危险行为且预警滞后的技术问题,本发明提供一种电梯作业人员危险行为智能识别预警方法及系统,通过视频采集单元获取作业区域视频流,利用人体骨骼关键点检测算法提取关键点集合并计算关节角度及运动速度以构建多维特征向量,同时融合可穿戴传感器数据进行多模态分析,将融合特征输入CNN-LSTM混合模型进行危险行为分类,并基于动态阈值判定逻辑结合作业空间位置信息确定风险等级,触发多级预警策略;本发明通过将危险行为转化为可量化的几何与运动特征,并结合多模态数据融合与时空序列分析,显著提升了光照不均、遮挡等复杂场景下的识别准确率与鲁棒性,实现了危险行为发生瞬间的毫秒级分级预警与电梯联动控制,并通过数据闭环管理实现模型迭代与安全决策优化,有效解决了传统监控语义鸿沟大、实时性差及管理断层的问题,大幅降低了电梯作业的安全风险

Benefits of technology

(1)本发明能够解决复杂环境下危险行为特征难以量化与识别精度低的技术难题;通过人体骨骼关键点检测算法将作业人员的姿态、动作幅度及手持物品状态转化为关节角度、运动速度及重心变化等多维量化特征,并结合视觉数据与可穿戴传感器数据的多模态融合分析,消除了传统二维图像在光照不均、遮挡严重场景下的特征退化问题,显著提升了危险行为识别的准确率与鲁棒性,有效克服现有技术无法区分微小动作及正常作业与危险行为的语义鸿沟。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799487A_ABST
    Figure CN122799487A_ABST
Patent Text Reader

Abstract

The application discloses an elevator operator dangerous behavior intelligent identification and early warning method and system, and the method comprises the following steps: acquiring video stream data and sensor auxiliary data of an elevator operation area; performing human body skeleton key point detection on the video stream data, and extracting a human body key point set of a current frame; calculating joint angle features, motion speed features, spatial region features and equipment interaction features based on the human body key point set, and constructing a multi-dimensional video feature vector; performing multi-modal fusion on the video feature vector and the sensor auxiliary data, and inputting the video feature vector and the sensor auxiliary data into a CNN-LSTM hybrid model to output a dangerous behavior category and a probability distribution thereof; when a maximum probability value is greater than or equal to a preset dynamic threshold value, it is determined that a target dangerous behavior occurs, spatial semantic analysis is performed on the target dangerous behavior to determine a risk level; and a corresponding multi-level early warning strategy is triggered according to the risk level; and the application can solve the technical problem of low dangerous behavior recognition precision in a complex light and shielding environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of elevator operation safety protection technology, and more specifically, relates to an intelligent identification and early warning method and system for dangerous behaviors of elevator operators. Background Technology

[0002] With the acceleration of urbanization, the number of elevators in high-rise buildings continues to rise, leading to a surge in the frequency of elevator installation, maintenance, repair, and emergency rescue operations. However, the working environment of elevator shafts, car tops, pits, and machine rooms is characterized by its enclosed, high-altitude, and complex nature, exposing workers to multiple risks such as falls from heights, mechanical crushing, shearing injuries, and being struck by objects. Although the current Special Equipment Safety Law and TSGT5002 Elevator Maintenance and Repair Rules impose strict requirements on operational safety, in practice, existing technologies still struggle to meet the demands of intelligent and refined safety management, specifically manifested in the following deep-seated technical bottlenecks: (1) The semantic gap in behavior recognition and the challenge of capturing subtle movements Dangerous behaviors during elevator operations are highly concealed and diverse. For example, there are very slight visual differences between "moving the escalator cable" and "tidying the cable," and between "bending over for maintenance" and "the initial stage of a slip and fall," and the range of motion is often small. Traditional video surveillance systems can only provide "visual images" and lack the ability to quantitatively analyze human posture. Existing technologies mostly rely on simple motion detection or manual visual recognition, which cannot extract the subtle geometric features of key skeletal points (such as wrists, neck, and torso). This makes it difficult for the system to distinguish between normal working postures and dangerous postures, resulting in a serious semantic gap and high false alarm and false negative rates.

[0003] (2) Feature degradation under complex lighting and shading environments Elevator shaft environments are extremely harsh, with significant uncontrolled variables: uneven lighting, with the deepest parts of the shaft often dark and the car top work area potentially experiencing strong glare or backlighting; severe obstructions, with elevator structural components (such as guide rails, counterweights, and cables) frequently obscuring parts of the workers' bodies. Conventional computer vision algorithms based on 2D RGB images exhibit a sharp decline in feature extraction capabilities under these conditions, failing to consistently produce reliable recognition results. Furthermore, single-viewpoint cameras have numerous blind spots, unable to cover the entire work area.

[0004] (3) Limitations of a single data source and insufficient resistance to interference Most existing solutions rely solely on visual data, lacking the integration of multi-dimensional contextual information. For example, when identifying "not wearing a safety helmet," they cannot combine head posture and position information for cross-validation; when identifying "working at heights," they cannot combine spatial coordinates to determine whether the person is on the designated work surface. This single-dimensional data analysis results in extremely poor robustness of the system in complex backgrounds, making it highly susceptible to interference from shadows and similar objects (such as a dark backpack being misidentified as a safety helmet).

[0005] (4) Lack of real-time performance and the unbreakable nature of the accident chain Elevator accidents often occur instantaneously (e.g., a fall can happen in just a few hundred milliseconds). Current technologies largely rely on a "hindsight" management model, i.e., assigning responsibility through manual inspections or reviewing video footage after the fact. This approach lacks real-time computing and millisecond-level response capabilities, failing to trigger intervention mechanisms at the instant a dangerous act occurs (e.g., a sudden change in center of gravity or crossing a boundary), thus missing the only window of opportunity to prevent an accident and failing to break the causal chain of "dangerous act—accident—injury or fatality."

[0006] (5) Static threshold alarms and lack of hierarchical management mechanism Traditional electronic fences or intrusion detection systems typically use fixed physical thresholds (e.g., alarming upon boundary crossing), failing to dynamically classify risks based on the nature of the behavior. For example, treating "using a mobile phone" (low to medium risk) the same as "falling from a height" (extremely high risk) leads to alarm fatigue or the neglect of critical alerts. Furthermore, existing systems lack closed-loop management capabilities for data across the entire operational process, making it impossible to utilize historical data for model iteration and security trend prediction, thus keeping security management at a purely passive, defensive stage.

[0007] In summary, current technologies have not yet solved the core technical challenges of accurately quantifying minute and dangerous movements in complex environments, real-time fusion of multimodal data, and proactive intervention based on risk levels. Therefore, there is an urgent need to develop an intelligent recognition and early warning solution that integrates deep learning, skeletal keypoint detection, and multi-sensor fusion to address these issues in elevator operation safety. Summary of the Invention

[0008] To address the technical challenges of existing technologies in accurately and in real-time identifying diverse hazardous behaviors in complex elevator environments, and the resulting delays in early warning, this invention provides an intelligent identification and early warning method and system for hazardous behaviors of elevator operators. The system acquires video streams of the work area through a video acquisition unit, extracts a set of key points using a human skeleton key point detection algorithm, and calculates joint angles and movement speeds to construct a multi-dimensional feature vector. Simultaneously, it integrates wearable sensor data for multimodal analysis. The fused features are input into a CNN-LSTM hybrid model for hazardous behavior classification, and the risk level is determined based on dynamic threshold judgment logic combined with work space location information, triggering a multi-level early warning strategy. This invention significantly improves the accuracy and robustness of identification in complex scenarios such as uneven lighting and occlusion by transforming hazardous behaviors into quantifiable geometric and motion features, combined with multimodal data fusion and spatiotemporal sequence analysis. It achieves millisecond-level graded early warning and elevator linkage control at the moment hazardous behavior occurs, and realizes model iteration and safety decision optimization through data closed-loop management. This effectively solves the problems of large semantic gaps, poor real-time performance, and management gaps in traditional monitoring, significantly reducing the safety risks of elevator operations.

[0009] To achieve the above objectives, one aspect of the present invention provides a method for intelligent identification and early warning of dangerous behaviors of elevator operators, comprising the following steps: S1: Acquire video stream data and sensor auxiliary data of the welding operation area; S2: Perform human skeleton key point detection on the video stream data and extract the set of human key points in the current frame; S3: Calculate joint angle features, motion speed features, spatial region features, and equipment interaction features based on the human body key point set, and fuse multi-dimensional features to construct a video feature vector; S4: Perform feature-level fusion of the video feature vector constructed in step S3 and the sensor-assisted data synchronously collected in step S1 to generate a fused feature vector; input the fused feature vector into a pre-trained dangerous behavior recognition model to output dangerous behavior categories and their probability distributions; based on the output probability distribution, extract the maximum probability value and compare it with a preset dynamic threshold to determine whether the behavior is dangerous, and finally output the behavior classification result. S5: When When the threshold is greater than or equal to a preset dynamic threshold, the occurrence of the target dangerous behavior is determined, and spatial semantic analysis of the dangerous behavior category is performed in conjunction with 3DGIS electronic fence data to determine the risk level. S6: Trigger the corresponding multi-level early warning strategy based on the risk level.

[0010] Further, step S1 includes: By deploying special camera modules and sensor nodes in the welding operation area, the system simultaneously acquires anti-strong light narrowband filtered video streams, infrared thermal imaging video streams, depth information video streams, panoramic high-definition video streams, and wide-angle infrared integrated video streams. By integrating sensor nodes into welding equipment and the working environment, personal protective equipment data, environmental perception data, and equipment operating condition data are collected simultaneously. Personal protective equipment data is collected through a smart safety helmet with integrated Hall effect sensors, a six-axis IMU, and RFID positioning tags. Environmental perception data is collected through temperature, humidity, smoke, and harmful gas sensors. Equipment operating condition data is collected through current transformers and voltage transformers. The collected multi-source data is time-stamp aligned and format standardized. The video stream data is decoded into RGB image frame sequences through the edge computing gateway, the sensor auxiliary data is converted into a structured numerical matrix, and the data association relationship is established with millisecond-level timestamps as indexes to form a multimodal dataset containing spatiotemporal synchronization information. Based on real-time collected ambient light intensity, dust concentration, and welding current parameters, the exposure time, gain, and filter switching strategy of the camera module are dynamically adjusted.

[0011] Furthermore, the construction of the visual feature vector in step S3 includes: S31: Calculate the human joint angle features and bone proportion features based on the three-dimensional key point set; S32: Calculate the instantaneous velocity and acceleration characteristics of key points based on the coordinates of 3D key points in multiple consecutive frames; S33: Map the coordinates of the human centroid to a pre-constructed 3D electronic map of the welding operation area, calculate the distance between the centroid projection point and the boundary of the preset safe area; at the same time, extract the relative positional relationship between the key points of the hand and the dangerous area, and generate a spatial occupancygrid feature vector. S34: Based on the key points of the welding equipment separated in step S24, calculate the interaction feature vectors between the hand, welding gun, and mask. S35: Concatenate the geometric features, motion features, position features, and interaction features extracted in steps S31 to S34 to construct a high-dimensional video feature vector; use the Z-score normalization method to normalize the feature vector and eliminate the influence of different dimensions.

[0012] Furthermore, in step S31, the joint angle The expression is: ,in, For the first Each joint angle , , These are the three adjacent key points that constitute a joint; Instantaneous velocity of key points in step S32 The calculation expression is: ,in, In order to be in At that moment, the Spatial coordinates of key points; In order to be in At that moment, the Spatial coordinates of key points; for Time and The time interval between moments.

[0013] Furthermore, the visual feature vector constructed in step S35 A multi-dimensional, layered fusion structure is adopted, and the expression is: in, This represents the vector concatenation operator; Features of joint angles; Kinematic characteristics; Features of a spatial region; For equipment interaction features; Indicates joint angle characteristics; Indicates the instantaneous velocity characteristics of key points; Indicates from the 1st to the 1st The angle values ​​of each key joint at the current moment; Indicates from the 1st to the 1st The speed of movement of each key point at the current moment; Indicates from the 1st to the 1st The acceleration of each key point at the current moment; Indicates from the 1st to the 1st The angular velocity of a key point at the current moment; The distance between the human body and the ground; Distance between human body and danger zone; For area occupancy coding; The distance between the wrist key point and the center point of the welding torch inspection frame; The angle between the welding torch and the normal vector of the working plane; This refers to the proportion of key facial features that are covered by the mask.

[0014] Furthermore, in step S4, the feature vectors are fused. The expression is: ,in, For visual feature vectors; Provide auxiliary data for sensors; For dynamic fusion weights, ; Probability distribution of dangerous behavior categories in step S4 The expression is: in, and These represent the weights and biases of the CNN-LSTM hybrid deep learning model, respectively.

[0015] Furthermore, in step S4, the expression for determining whether an action constitutes a danger is: in, This indicates that a dangerous behavior has been determined to have occurred. 0 indicates that no dangerous behavior was detected; Let be the risk probability distribution vector output by the deep learning model at time t; This indicates taking the maximum probability value from the risk probability distribution vector; if the highest confidence level predicted by the model is greater than or equal to the set threshold. If the highest confidence level is less than the threshold, the result is 1, indicating that a dangerous behavior has occurred and an alarm is triggered; if the highest confidence level is less than the threshold, the result is 1. If the result is 0, the current state is considered safe, and no alarm is triggered; regarding the threshold Special settings: For high-risk behaviors, For low-risk behaviors, ;like If so, then it is determined that a danger has occurred.

[0016] Furthermore, the risk level in step S5 The expression is: 1 indicates Level 1 warning - alert level; 2 indicates Level 2 warning - alert level; 3 indicates Level 3 warning - emergency level; 4 indicates Level 4 warning - rescue level. In step S6, for a Level 1 warning, only a voice prompt is sent to the worker's safety helmet headset: "Please pay attention to safety regulations," and the violation is recorded once, without shutting down the machine; In response to a Level 2 warning, the on-site audible and visual alarm flashes a yellow light, a warning sound is broadcast, and an alarm screenshot is pushed to the on-site safety officer's APP. In response to the Level 3 warning, the red light on site flashes and is accompanied by an alarm siren, automatically triggering the elevator emergency stop circuit to prevent the elevator from moving unexpectedly and causing secondary damage. At the same time, an alarm is displayed on the management office's large screen. In response to a Level 4 alert, in addition to implementing Level 3 alert measures, an emergency call will be automatically made, and the status of personnel will be confirmed using two-way voice communication.

[0017] The second aspect of the present invention provides an intelligent identification and early warning system for dangerous behaviors of elevator operators, used to implement the aforementioned intelligent identification and early warning method for dangerous behaviors of elevator operators, comprising: a multimodal perception layer, an edge computing layer, a data analysis and decision-making layer, and a data closed-loop management layer; The multimodal perception layer is configured to collect video stream data and sensor-aided data from the elevator operating area; The edge computing layer is communicatively connected to the multimodal perception layer and is configured to process the video stream data and sensor-aided data, and output a hazard determination result. The data analysis and decision-making layer is communicatively connected to the edge computing layer and is configured to determine the risk level based on the hazard assessment results and generate multi-level early warning strategies. The data closed-loop management layer is connected to the edge computing layer and the data analysis and decision-making layer, respectively, and is configured to store job data and iteratively optimize the identification model; The multimodal perception layer includes a visual acquisition unit and a wearable sensing unit; the visual acquisition unit includes a special camera module deployed in the welding operation area; the wearable sensing unit includes a smart safety helmet and an environmental monitoring sensor; the smart safety helmet integrates a six-axis IMU, a Hall sensor and an RFID positioning tag; the environmental monitoring sensor includes temperature and humidity, smoke and harmful gas sensors deployed on site; The edge computing layer includes a video preprocessing module, a pose recognition engine, a feature engineering module, a fusion inference module, and a real-time linkage interface. The video preprocessing module is used to denoise, correct distortion, and scale the resolution of the raw video stream. The pose recognition engine is used to detect key points of the human skeleton and construct a set of 3D coordinates. The pose recognition engine is equipped with an optimized HRNet inference model to perform key point detection of the human skeleton and output the time-space coordinates of each key point. 3D coordinate set And execute the occlusion compensation algorithm; The feature engineering module has a built-in feature calculation unit, which is used to calculate joint angle features, motion speed features, spatial region features and equipment interaction features based on the human body key point set, and to fuse multi-dimensional features to construct video feature vectors. Fusion Inference Module: Used to load the trained CNN-LSTM hybrid model, perform multimodal data fusion and dangerous behavior classification, and output probability distribution; Real-time linkage interface: It is hardwired to the elevator control cabinet via GPIO / Relay interface. When a level 3 or higher warning is received, the elevator emergency stop circuit is triggered in milliseconds. The data analysis and decision-making layer includes a risk assessment unit, a multi-level early warning and dispatch unit, and a digital twin visualization module. The risk assessment unit combines 3D GIS electronic fence data to perform spatial semantic analysis of behavior and generate risk levels. 1 indicates a prompt, 2 indicates a warning, 3 indicates an emergency, and 4 indicates a rescue. Multi-level early warning and dispatch units are used to determine risk levels. Different early warning strategies can be invoked, and connections can be made with VoIP voice gateways, SMS / AppPush platforms, and sound and light alarms; The digital twin visualization module is used to build a 3D digital twin model of the elevator shaft, which maps the position and posture of the workers in real time. The data closed-loop management layer includes: a time-series database, which stores massive amounts of key point coordinates, feature vectors, and sensor time-series data; Object storage stores raw video streams, alarm screenshots, and short video clips; Model training and iteration platform: Automatically selects fuzzy samples with confidence levels between 0.5 and 0.7 for manual annotation; regularly uses newly annotated data to fine-tune and update the HRNet and CNN-LSTM models at the edge; Safety Management Portal: Provides a web interface that supports historical alarm queries, safety report generation, and display of safety profiles for operators.

[0018] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: (1) This invention can solve the technical problem of difficulty in quantifying dangerous behavior features and low recognition accuracy in complex environments; by using the human skeleton key point detection algorithm, the posture, movement range and hand-held item status of the worker are transformed into multi-dimensional quantitative features such as joint angle, movement speed and center of gravity change. Combined with the multi-modal fusion analysis of visual data and wearable sensor data, the feature degradation problem of traditional two-dimensional images in scenes with uneven lighting and severe occlusion is eliminated, which significantly improves the accuracy and robustness of dangerous behavior recognition and effectively overcomes the semantic gap of existing technologies that cannot distinguish between small movements and normal operations and dangerous behaviors.

[0019] (2) This invention enables a qualitative leap from passive monitoring to proactive real-time intervention, breaking through the technical bottleneck of delayed early warning. By constructing a dynamic risk assessment mechanism and multi-level early warning strategy based on time-series action sequences, this invention can complete millisecond-level identification and graded response the instant a dangerous behavior occurs (such as slipping or crossing boundaries), and link sound and light alarms, mobile terminal push notifications, and even elevator emergency stop control according to the risk level. This mechanism solves the problem that traditional manual inspections and post-event video playback cannot break the accident chain, putting safety protection forward and greatly reducing the accident rate.

[0020] (3) This invention fills the gap in operational safety management by constructing a closed-loop data security management system; it establishes a closed-loop management mechanism of "monitoring-identification-early warning-optimization" by storing and tracing the entire lifecycle of video streams, key point information and alarm records. By using historical data for model iteration and safety profile generation, it not only solves the problem that existing technologies cannot self-correct false alarms and missed alarms, but also provides scientific data basis for operational process optimization, targeted personnel training and accident root cause analysis, realizing the leap from "experience-driven" to "data-driven" safety management. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating an intelligent identification and early warning method for dangerous behaviors of elevator operators according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of an intelligent identification and early warning system for dangerous behaviors of elevator operators according to an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0023] like Figure 1 As shown, one aspect of the present invention provides an intelligent identification and early warning method for dangerous behaviors of elevator operators, used to monitor, identify, and warn of dangerous behaviors of operators in real time in elevator maintenance, repair, installation, and fault handling scenarios, thereby reducing accident risks and improving the level of work safety management; including the following steps: S1: Acquire video stream data and sensor auxiliary data of the welding operation area; S2: Perform human skeleton key point detection on the video stream data, extract the human key point set of the current frame, and perform occlusion compensation based on skeleton length constraints and motion trends; S3: Calculate joint angle features, motion speed features, spatial region features, and equipment interaction features based on the human body key point set, and fuse multi-dimensional features to construct a video feature vector; S4: Perform multimodal feature-level fusion of the video feature vector constructed in step S3 and the sensor-assisted data synchronously collected in step S1 to generate a fused feature vector; input the fused feature vector into a pre-trained CNN-LSTM hybrid model to output dangerous behavior categories and their probability distributions; based on the output probability distribution, extract the maximum probability value and compare it with a preset dynamic threshold to determine whether the behavior is dangerous, and finally output the behavior classification result; S5: When When the threshold is greater than or equal to a preset dynamic threshold, the occurrence of the target dangerous behavior is determined, and spatial semantic analysis of the dangerous behavior category is performed in conjunction with 3DGIS electronic fence data to determine the risk level. S6: Based on the stated risk level Trigger the corresponding multi-level early warning strategy .

[0024] This invention solves the technical challenge of low accuracy in identifying dangerous behaviors under complex lighting and occlusion environments by quantifying the geometric and kinematic features of human posture and combining multimodal data fusion and spatiotemporal sequence analysis. Through millisecond-level real-time inference and hierarchical early warning mechanisms, it breaks through the early warning lag bottleneck of traditional manual inspections. By constructing a data closed-loop management system of "monitoring-identification-early warning-optimization", it achieves a leap from passive monitoring to proactive prevention, significantly reducing the safety risks of elevator operations.

[0025] The various steps of the present invention will be further described in detail below.

[0026] (1) Multi-source heterogeneous data acquisition and adaptive environment configuration Step S1 aims to solve the perception challenges in the complex environment of elevators through hardware collaboration, specifically including: S11: Parallel acquisition of multi-source heterogeneous data, synchronous acquisition of multi-modal video streams through special camera modules deployed in the welding operation area; the multi-modal video streams include anti-strong light narrowband filter video stream, infrared thermal imaging video stream, depth information video stream, panoramic high-definition video stream and wide-angle infrared integrated video stream; The video stream with strong light resistance and narrow band filtering is filtered by an 808nm or 940nm narrow band filter to remove the 380nm-780nm visible light band of the welding arc, while retaining the near-infrared spectrum to capture personnel limb movements; the infrared thermal imaging video stream is captured by an 8μm-14μm band infrared thermal imager to collect high-temperature heat sources and the outline of human body thermal radiation; the depth information video stream is captured by a ToF or binocular stereo camera to obtain three-dimensional spatial information of the work area; the panoramic high-definition video stream is captured by a PTZ pan-tilt camera to monitor a large-area work area and welding torch operation details; the acquisition of the wide-angle infrared integrated video stream is achieved by deploying wide-angle infrared integrated network cameras on the top of the welding arcade, the top of the elevator car, the pit, and the machine room; the cameras support automatic aperture adjustment to ensure that the target area remains in clear focus as the elevator car moves up and down; the cameras can use RGB, depth, or infrared modes and dynamically adjust according to lighting and environmental complexity. S12: Sensor-assisted data synchronous acquisition. Through sensor nodes integrated into welding equipment and the working environment, personal protective equipment data, environmental perception data, and equipment operating condition data are collected synchronously. Personal protective equipment data is obtained through the smart safety helmet integrating Hall sensors (face mask opening and closing status data), a six-axis IMU (100Hz acquisition of head posture, acceleration, and angular velocity), and RFID positioning tags. Environmental perception data is obtained through temperature, humidity, smoke, and harmful gas sensors. Equipment operating condition data includes current / voltage waveform data of the welding machine / elevator control cabinet collected through current transformers and voltage transformers. S13: Data preprocessing and alignment, performing timestamp alignment and format standardization on the collected multi-source data, decoding the video stream data into RGB image frame sequences through the edge computing gateway, converting sensor auxiliary data into a structured numerical matrix, and establishing data association relationships with millisecond-level timestamps as indexes to form a multimodal dataset containing spatiotemporal synchronization information; S14: Environmental adaptive parameter configuration. Based on real-time collected ambient light intensity, dust concentration and welding current parameters, the exposure time, gain and filter switching strategy of the camera module are dynamically adjusted. For example, when the light intensity is below 15 Lux, the camera automatically switches to infrared imaging mode. When the dust concentration exceeds the threshold, the image denoising algorithm is strengthened to ensure the stability and effectiveness of data acquisition. (2) Human skeleton key point detection and 3D reconstruction based on deep learning Step S2 converts the two-dimensional image information obtained in step S1 into machine-understandable structured data, specifically including: S21: Anti-interference skeletal key point detection. The anti-strong light narrowband filtered video stream obtained in step S1 is input into the pre-trained HRNet (High Resolution Network) model, and the output is the skeletal key point detection of the workers in the current frame. A set of coordinates for key human body points (the coordinate set is...) ;generally or The HRNet model, which includes the nose, eyes, ears, shoulders, elbows, wrists, hips, knees, and ankles, incorporates a synthetic dataset of welding arc light, spatter, and smoke occlusion during the training phase for domain adaptation optimization to improve detection robustness in complex industrial environments. S22: 3D coordinate mapping and scale restoration. To overcome the scale ambiguity caused by perspective projection of 2D images, camera calibration, depth estimation, and coordinate system registration are used to restore the coordinates of 2D pixels. Precisely mapped to the physical three-dimensional world coordinate system of the elevator shaft. ), generating a set of 3D key points (i.e., each key point in time) 3D coordinate set ) S23: Key point completion under occlusion state. When the confidence level of some key points is less than 0.3 due to occlusion by welded parts or equipment, the occlusion compensation algorithm is executed to complete the missing key point coordinates and maintain the temporal continuity of the feature vector. S24: Key point separation of welding equipment: For equipment specific to welding operations (such as welding guns and masks), the instance segmentation model (MaskR-CNN) is used to identify the outline of the equipment and associate it with key points of the human body to distinguish between "hand-held welding gun" and "empty-handed" states, providing feature basis for subsequent identification of dangerous behaviors such as "welding without a mask". Furthermore, in step S22, through camera calibration, depth estimation, and coordinate system registration, the two-dimensional pixels in the image are accurately mapped to the physical three-dimensional world coordinate system of the elevator shaft; the specific steps include: Camera calibration and intrinsic parameter acquisition: Zhang's Method was used, employing a 9×7 checkerboard calibration board. Multiple sets of images were acquired at different angles and positions, and the intrinsic parameter matrix was output. and the coefficient of variation , , Focal length Principal point coordinates; variable coefficients Includes radial distortion and tangential distortion It is used to perform distortion correction on the original image; Depth information acquisition and scale restoration: Since the elevator operation scenario is semi-structured (with a large number of geometric references of known dimensions), the physical depth is obtained using the following method. Includes: Mode A (hardware level), which directly uses an RGB-D depth camera (such as a ToF camera or a structured light camera) to output the physical depth value corresponding to each pixel by emitting infrared light pulses and receiving reflected signals. (Unit: mm); Mode B (algorithm-level / stereo), using a standard RGB camera, scale reconstruction is performed using known reference objects within the elevator shaft; based on the standard height of the elevator car (e.g., ) or guide rail spacing (e.g. (Used as a baseline real scale; measuring the pixel height of a reference object in the image) Combined with real physical height Calculate the scale factor of the current frame. : Combined with camera focal length Calculate the actual distance between the target key points and the camera. : ; Coordinate transformation pipeline: Mapping two-dimensional pixel coordinates to three-dimensional world coordinates requires the following three-level transformation: Step 1: Pixel coordinates → Camera normalized coordinates remove the influence of intrinsic parameters and transform the pixels to the camera coordinate system with the optical center as the origin; Step 2: Camera coordinates → World coordinates (back projection) combined with depth value (From Step 2) Obtain the 3D points in the camera coordinate system. : Step 3: Use the external parameter matrix (rotation matrix) Translation vector The position of the camera relative to the elevator shaft is solved using the PnP algorithm. Transformation to the physical world coordinate system of the elevator shaft : 3D reconstruction based on multi-view geometry: To address the limited field of view of a single camera, the system integrates observation data from multiple cameras: Triangulation: If the same worker appears in the field of view of two cameras at the same time, the system uses epipolar geometric constraints to perform forward intersection of the two two-dimensional observation points and calculate their unique and precise coordinates in three-dimensional space. Point cloud fusion: combining 3D coordinate points from different times and perspectives The data is stitched together to generate the spatiotemporal point cloud trajectory of the workers. Data output and structured representation: The final output shows the time for each key point. 3D coordinate set , ; Among them, a single key point ; Axis: along the direction of the elevator guide rail (horizontal displacement); Axis: Perpendicular to the ground (height); Axis: Perpendicular to the guide rail direction (depth / front / back position); Through the three-dimensional coordinate mapping in this step, the spatial position of the workers can be accurately quantified; for example, it can not only identify "the person is on the car top", but also be accurate to "the person is 0.5 meters away from the edge of the car top", thus providing a precise physical distance basis for subsequent judgment of "risk of falling from height", and solving the fundamental defect that two-dimensional images cannot determine the depth of spatial position; Further, step S23 includes: Spatial interpolation (skeleton constraint): Based on the rigid constraint of human skeleton length, the position of the occluded point is calculated by triangulation using the parent and child nodes of the occluded point as a reference, and a unique solution is selected by combining kinematic constraints.

[0027] Time series forecasting (movement trend): Short-term occlusion: Based on the assumption of uniform motion, the velocity is calculated using the displacement of the first two frames for linear prediction.

[0028] Long-term occlusion: The Kalman filter algorithm is used to predict the current coordinates based on the motion state of the previous N frames, and a smooth transition is achieved when the occlusion ends.

[0029] Contextual reasoning (multiple hypothesis interpolation): using rigid body constraints (if the head is stationary, the torso is relatively stationary) or symmetry mapping (using mirror images of the opposite limbs to complete the reasoning) to assist in reasoning.

[0030] Confidence-weighted fusion: Spatial, temporal, and contextual results are weighted and fused to generate the final completed coordinates, and a "completeness flag" is added to dynamically adjust the weights of subsequent models. The occlusion compensation algorithm includes: Spatial interpolation: Utilizing the rigid constraint of human skeleton length, the position of the occlusion point is estimated through triangulation based on the coordinates of visible parent and child nodes; Temporal prediction: Using the Kalman filter algorithm, the coordinates of the occlusion point in the current frame are predicted based on the motion state of the previous N frames, maintaining the temporal continuity of the feature vector. S231. First, perform occlusion detection and visibility scoring on key points in each frame: set a dual threshold mechanism (such as confidence level). <0.3, marked as complete occlusion, confidence level A score between 0.3 and 0.6 is considered low confidence; confidence level (Mark as visible), input the original keypoint coordinates. and confidence score Generate a visibility mask ; S232, Spatial Interpolation: For isolated occlusion points within a single frame, estimation is performed using rigid constraints on the length of the human skeleton, including: Skeletal length modeling: During the initialization phase, the system calculates the average length of each skeletal segment of the worker based on the first N frames (e.g., the first 30 frames). (such as upper arm length, forearm length, thigh length, etc.), and establish a standard skeletal proportion model; Triangulation completion: using occluded points parent node and child nodes As a benchmark; known and ; by Center of the circle Construct a sphere with radius , and Center of the circle Construct a sphere with radius r, and the intersection of the two spheres is r. The possible candidate set of locations; Taking into account kinematic constraints (such as an elbow flexion range not exceeding 180 degrees), the only reasonable positional coordinates are selected as the completion value. ; S233. For occlusion across multiple consecutive frames, trajectory prediction is performed using the temporal continuity of the video stream; specifically including: For short-term occlusion (<200ms), it is assumed that the human body moves at a constant linear speed during the occlusion period; the two most recent frames before the occlusion are used. and Position, calculate instantaneous velocity Predict the current frame position: ; For long-term occlusion, the Kalman filter algorithm is used, based on the previous... Frame motion state predicts the current coordinates and smoothly transitions when occlusion ends; Kalman filter state equation: Measurements are updated when keypoints are visible; state prediction is performed only when keypoints are occluded; the covariance matrix is ​​used. The weights of the predicted values ​​are dynamically adjusted to smoothly connect with the actual observations when the occlusion ends, avoiding coordinate jumps. S234. For specific dangerous behaviors (such as hands being obstructed by a cable while holding onto it), reasoning is performed using contextual information (multi-hypothesis interpolation based on the association of adjacent joints); including: Rigid body constraints: If it is detected that the worker is wearing a safety helmet (key points of the head are visible) and the IMU sensor data shows that the head is stationary, it is inferred that the obscured torso and limbs should remain relatively stationary, and the relative offset of the previous frame is used for supplementation. Symmetrical mapping: The human limbs are symmetrical; if the key points of the left arm are visible while the right arm is occluded, the system can use the posture mirror mapping of the left arm to complete the posture of the right arm based on the work habit model, and vice versa. S235, Confidence-weighted fusion and feature repair; including: The fusion strategy combines the spatial interpolation results Temporal interpolation results and the results of contextual reasoning Perform a weighted average to generate the final completed coordinates. ; Weight allocation is adjusted according to the duration of occlusion. Short-term occlusion is given higher weight to temporal interpolation, while long-term occlusion is given higher weight to skeletal constraints. Feature vector correction in generating visual feature vectors When completing the task, add a "completeness flag" to the key points. (0 for true, 1 for complete), and dynamically adjust the contribution weight of the feature in this dimension based on this flag during the model inference stage to prevent false features from misleading the classifier; (3) Quantization construction of multi-dimensional behavioral feature vectors Step S3 transforms "posture" into "features" through geometric operations, which is the core of distinguishing between normal operations and dangerous behaviors; by performing geometric and temporal operations on key point data, a high-dimensional visual feature vector is generated. ;include: S31: Static geometric topology feature extraction (this step describes the "shape of the body" used to distinguish different posture categories (e.g., standing vs. bending over)): Calculate the joint angle features and skeletal proportion features of the human body based on the three-dimensional key point set; select three key points constituting the joint to calculate the joint angle using the vector dot product formula, and simultaneously calculate the skeletal length ratio (e.g., upper arm length / forearm length) and the height-to-width ratio of the circumscribed geometry, used to quantitatively describe the static posture of the worker (e.g., bending over, squatting, stretching). S32: Dynamic kinematic feature extraction (This step describes "how the body is moving" and is used to capture the instantaneous state of the action (such as slipping vs. walking slowly)): Based on the coordinates of 3D key points in multiple consecutive frames, calculate the instantaneous velocity and acceleration features of the key points; at the same time, calculate the joint angular velocity to capture the intensity and trend of the action. S33: Spatial location and regional feature extraction, mapping the human body's centroid coordinates (usually the midpoint of the hip) to a pre-constructed 3D electronic map of the welding operation area, calculating the distance between the centroid projection point and the preset safety zone boundary; at the same time, extracting the relative positional relationship between key hand points and dangerous areas (such as the slag spatter area and the high-temperature workpiece area), generating spatial occupancygrid features. S34: Extraction of interactive features of welding equipment. Based on the key points of welding equipment separated in step S24, calculate the interactive feature vectors of the hand, welding gun, and mask. This includes: the minimum Euclidean distance between the hand and the welding gun handle, the angle between the welding gun and the workpiece surface, and the coverage of the face covered by the mask. Combined with the welding machine current / voltage waveform data, verify the logical consistency of "holding the welding gun - powering on - working". S35: Feature Vector Fusion and Normalization: The geometric features, motion features, position features, and interaction features extracted in steps S31 to S34 are concatenated to construct a high-dimensional video feature vector; the Z-score normalization method is used to normalize the feature vector to eliminate the influence of different units (angle, pixel, speed) and ensure the consistency of the numerical scale of each dimension of features. Furthermore, the joint angle calculation in step S31 includes selecting three key points that constitute the joint. (e.g., shoulder, elbow, wrist), calculate joint angles , ,in, For the first Each joint angle , , For three adjacent key points constituting a joint (such as shoulder, elbow, and wrist), the pose information of consecutive frames can be used to calculate features such as range of motion, joint rate of change, and center of gravity change, providing basic data for dangerous behavior identification; the joint angle feature can effectively identify "hand holding a cable" (specific arm bending angle); calculate the "neck-left shoulder-left elbow" angle to determine whether the shoulder is shrugging; and the "left shoulder-left elbow-left wrist" angle to determine whether the hand is reaching for an object; Step S31 also includes calculating the vertical distance between the "top key point" and the "neck key point"; if the distance is less than a threshold (e.g., 20 pixels) and the head key point disappears, it is determined that "no helmet is being worn"; Step S31 also includes calculating the ratio of trunk length (neck to hip) to leg length (hip to ankle) to identify “squatting” behavior (significantly reduced leg length ratio) or “climbing” behavior (limbs extended, abnormal proportion). Step S31 also includes calculating the aspect ratio of the minimum bounding rectangle (AABB) of all key points. The aspect ratio is large when the object is upright and small when it falls (tends to be flat). Furthermore, the instantaneous velocity of the key point in step S32 The calculation includes: for the th A key point: instantaneous speed For this key point, at the current moment relative to the previous frame Bit removal at time intervals It is used to capture the instantaneous high-speed movement when slipping; ,in, In order to be in At that moment, the Spatial coordinates of key points; In order to be in At that moment, the Spatial coordinates of key points; for Time and The time interval between moments; The acceleration calculation formula for the key points in step S32 is as follows: The formula for calculating joint angular velocity is: ; Step S32 also includes calculating the change in the center of gravity. : Calculate the change in the line connecting the center points of the back and knees to determine whether a "fall" or "loss of balance" has occurred; Further, step S33 includes: S331. 3D Electronic Map Construction and Regional Semanticization: Pre-construct a 3D digital twin model of the elevator shaft, dividing the physical space into grid units with safety attributes, including: safety zones (e.g., the center of the car top and the pit maintenance platform, green grid); warning zones (e.g., within 1 meter of the edge of the car top and the landing sill, yellow grid); and restricted areas (e.g., the counterweight area, the traction steel wire rope area, and the area with live electrical control cabinet, red grid). S332, Human Body Center of Mass Projection and Distance Field Calculation: Calculate the coordinates of the human body's center of mass (usually the midpoint between the left and right hip key points) in the world coordinate system based on a set of 3D key points; calculate the shortest Euclidean distance from the center of mass coordinates to the boundaries of various regions; combine the key points of both feet to calculate whether the center of mass projection point falls within the supporting polygon formed by the two feet, and output the stability margin feature; if the center of mass projection point deviates from the supporting polygon, it is judged as "unbalanced". S333. Relative relationship between hands and dangerous areas: For electric welding operations, the distance between key hand points (wrist, fingertips) and high-temperature weld spatter areas (based on welding gun position and physical parabolic model prediction) and high-temperature workpiece areas (based on infrared thermal imaging data) is specifically calculated; it is determined whether key hand points have intruded into the preset "non-contact areas" (such as rotating traction wheels, exposed wiring terminals). S334, Occupancy Grid Feature Encoding: Dividing the job space into... 3D mesh ( This represents the number of grids into which the space is divided along the width (X-axis); This represents the number of grids into which the space is divided along the depth (Y-axis); This represents the number of grids into which the space is divided along the height (Z-axis); based on sensor data, each grid is assigned a state value (0 for idle, 1 for occupied by a human, 2 for occupied by equipment, and 3 for occupied by a hazard); the index of the grid containing the human centroid and the state distribution of its surrounding neighborhood are extracted to generate a spatial occupancy grid feature vector. ; By introducing the Occupancy Grid feature, the system no longer judges the action of "bending over" in isolation, but combines the spatial semantics of "bending over and being located at the edge of the car roof" to accurately determine it as "dangerous leaning over" rather than "normal tool picking up". This can effectively solve the problem of difficult understanding of behavioral intentions in complex backgrounds and significantly reduce the false alarm rate.

[0031] Further, step S34 includes: S341. Equipment key point separation and contour extraction: Identify the pixel-level contours of welding torches, masks, wrenches and other working equipment, and generate binary masks; define specific parts of the equipment (such as the center of the grip of the welding torch, the center of the window of the mask) as virtual key points, and establish dynamic associations with the nearest human hand key points (wrist, fingertips). S342. Hand-equipment distance feature calculation: Calculate the Euclidean distance between the key points of the left / right wrist and the key points of the welding torch handle. If the distance is less than the preset threshold (e.g., 10cm), it is determined to be "holding state". Calculate the intersection-union ratio of the hand key point set and the equipment mask area to determine whether the hand covers the equipment operating surface. S343. Equipment posture and operation feature extraction: Calculate the angle between the line connecting the key points of the welding torch and the direction of gravity (or the normal vector of the workpiece surface). If it is >150° (close to horizontal), it is judged as "illegal flat welding" or "random placement". Calculate the coverage rate of the mask mask on the key points of the face (eyes, nose, mouth). If it is <80% and it is judged as welding action, then trigger the "not wearing a mask" warning. S344. Multimodal logic consistency verification: Combining the sensor-aided data collected in step S1, cross-validate the visual interaction features to eliminate visual illusions; specifically including: Electrical logic verification: If the visual judgment is "welding in progress" (hand holding welding gun and close to the workpiece), then verify the welding machine current / voltage waveform data collected in step S12; if the current is zero (the equipment is not working), then it is judged as "simulated action" or "false detection", and the interactive feature weight is reset; IMU posture verification: If the visual judgment is "holding a heavy object", then the head vibration frequency in the smart helmet IMU data is verified; if the vibration is abnormally violent, it is corrected to "tapping operation" instead of "holding an object at rest"; Furthermore, the visual feature vector constructed in step S35 A multi-dimensional, layered fusion structure is adopted, and the expression is: in, This represents the vector concatenation operator; Features of joint angles; Kinematic characteristics; Features of a spatial region; For equipment interaction features; Indicates joint angle characteristics; Indicates the instantaneous velocity characteristics of key points; Indicates from the 1st to the 1st The angle values ​​of each key joint (15 dimensions) at the current moment are calculated from the coordinates of adjacent key points and are used to describe the posture of the human body (e.g., whether the elbow is bent or the body is tilted). Indicates from the 1st to the 1st The motion speed of each key point (17 dimensions) at the current moment is calculated by removing the coordinates of adjacent frames and calculating the time interval, which is used to describe the speed and intensity of limb movement. Indicates from the 1st to the 1st The acceleration of key points at the current moment is used to capture the intensity of the action and to identify abnormal states such as impact and convulsions. Indicates from the 1st to the 1st The angular velocity of a key point at the current moment; The distance between the human body and the ground, i.e. the vertical height of the center of gravity (hip) from the ground, is used to identify the trend of "working at height" or "fall". The distance between the human body and the danger zone is the distance from the center of mass of the human body to the nearest high-temperature / dangerous grid, calculated based on the OccupancyGrid. This distance is used to determine whether the person has entered the restricted area. To encode the area occupancy, the grid where the personnel are located is One-Hot encoded to determine the specific workstation or area of ​​the personnel in the workshop; The distance between the wrist key point and the center point of the welding gun detection frame is used to determine whether the hand is holding the tool, and to distinguish between "holding the welding gun" and "freehand". The angle between the welding torch and the normal vector of the working plane is used to identify illegal welding angles; The proportion of facial key points covered by the mask is used to identify the violation of "welding without wearing a mask"; Furthermore, in step S35, the Z-score normalization method is used to normalize the concatenated original feature vector. The expression for normalization is: ,in, The feature mean vector calculated on the training dataset. The feature standard deviation vector is calculated on the training dataset; after normalization, the feature vector... The values ​​of each dimension all follow a standard normal distribution with a mean of 0 and a standard deviation of 1, mapping all feature values ​​to the interval [−1,1]. By integrating four-dimensional features—static posture, dynamic trend, spatial location, and equipment interaction—a digital twin comprehensively describing the behavioral state of operators was constructed. Through normalization processing, the model bias problem caused by inconsistent dimensions of multi-source features was solved, significantly improving the accuracy and robustness of hazardous behavior identification.

[0032] (4) Multimodal data fusion and dangerous behavior classification reasoning Step S4 utilizes a deep learning model to process high-dimensional features and achieve accurate classification; Multimodal feature fusion: The video feature vector constructed in step S3 is fused together. Sensor auxiliary data collected synchronously with step S1 Perform feature-level fusion to generate a fused feature vector. The expression is: ,in, For dynamic fusion weights, When the light is good, The value approaches 1 when the shading is severe. Reduce and enhance system robustness; fused feature vector It can significantly improve the accuracy of dangerous behavior recognition in low-light, occluded, or complex action scenarios; wherein, the sensor-assisted data Data including the open / closed status of the welder's mask, the current / voltage waveform of the welding machine, and wearable IMU data are mapped to a high-dimensional feature space through a fully connected layer and then concatenated with the video feature vector. Spatiotemporal Feature Extraction and Hazardous Behavior Classification: Fusing Feature Vectors The input is fed into a pre-trained CNN-LSTM hybrid deep learning model. The CNN module uses a ResNet-18 architecture to extract spatial semantic information (such as pose and equipment interaction state) from single-frame feature vectors. The LSTM module contains 128 hidden units to capture the temporal evolution of feature vectors across multiple frames (such as action trends and speed changes). The model's output layer generates the probability distribution of dangerous behavior categories using a Softmax activation function. : in, and These represent the weights and biases of the CNN-LSTM hybrid deep learning model, respectively. Dynamic threshold determination logic: based on the probability distribution of the output. Extract the maximum probability value and with the preset dynamic threshold The system compares and determines whether a behavior is dangerous, ultimately outputting a behavior classification result (0 or 1); the dynamic threshold... Adaptive adjustment based on the category of dangerous behavior: for high-risk behaviors that may cause immediate harm (such as electric shock, falls). Set a lower value (e.g., 0.75) to improve recall; for progressive or low-risk behaviors (e.g., not wearing a safety helmet, not wearing work clothes, or standing in violation of regulations). Set a higher value (e.g., 0.90) to reduce the false alarm rate; if the maximum probability value is greater than or equal to the dynamic threshold, it is initially determined that a dangerous behavior has occurred, and a preliminary judgment signal is generated. ;otherwise ; The expression for the binary judgment model of dangerous behavior based on dynamic thresholds is as follows: in, This indicates that a dangerous behavior has been determined to have occurred. 0 indicates that no dangerous behavior was detected; Let t be the risk probability distribution vector output by the deep learning model at time t (or the current detection window); This indicates taking the maximum probability value from the risk probability distribution vector; if the highest confidence level predicted by the model is greater than or equal to the set threshold. If the highest confidence level is less than the threshold, the result is 1, indicating that a dangerous behavior has occurred and an alarm is triggered; if the highest confidence level is less than the threshold, the result is 1. If the result is 0, the current state is considered safe, and no alarm is triggered; regarding the threshold Special settings: for high-risk behaviors (such as falling). For low-risk behaviors (such as using mobile phones); To balance the false alarm rate; if (like If the condition is ), then it is determined that a danger has occurred; Behavior sequence verification: To avoid misjudgment of a single frame, continuous... Frame (e.g.) The results of the judgment are statistically analyzed using a sliding window. When the percentage of dangerous behavior judgments within the window exceeds a preset proportion (e.g., 70%), the dangerous behavior is confirmed to have occurred. Otherwise, it is judged as transient interference and the count is reset to further improve the stability of recognition. (5) Risk level assessment step S5 incorporates spatial location information to avoid isolated judgment actions; specifically including S51: Context-based composite verification: To avoid misjudgment by single visual recognition, multi-source context verification is performed on the primary judgment signal. Electrical logic verification: If the judgment is "live work without protection" (risk of electric shock), then verify the welding machine current / voltage waveform data collected in step S1; if the current is zero (equipment is not working), then it is judged as a false alarm, and the primary judgment signal is reset to 0; Spatial location verification: If it is determined to be "violation of high-altitude operation regulations" (fall risk), then verify the spatial occupancygrid features extracted in step S3. If personnel are located within the ground safety grid, the alarm is considered a false alarm, and the primary judgment signal is reset to 0. Equipment status verification: If it is determined that "welding without a mask" (risk of arc burn), then the Hall sensor data of the smart safety helmet is verified; if the mask is actually lowered, then it is determined to be a false alarm and the primary judgment signal is reset to 0. S52: Temporal Continuity Verification: To filter out transient interference (such as light flickering or insect occlusion), the construction length is [missing information]. Frame (e.g.) Sliding time window The cumulative number of frames in the statistics window where the primary determination signal is equal to 1. Only when (in When the confidence level coefficient is 0.7, the dangerous behavior is confirmed to have actually occurred, and the output stability judgment result is equal to 1; otherwise, it is regarded as a transient interference or false detection, and the stability judgment result is equal to 0. S53: Comprehensive evaluation based on behavioral sequences and spatial semantics: Behavioral sequence analysis: Identifying the logic of continuous actions; for example, if the action of "climbing" is followed by "center of gravity being suspended in the air", it is predicted as "risk of falling"; if "bending over" is followed by a long period of stillness, it is judged as "heatstroke / coma" based on IMU data; Spatial semantic analysis: A 3D electronic fence model of the elevator shaft is pre-constructed and divided into safe zones (such as the center of the car top), warning zones (such as within 1 meter of the edge of the car top), and restricted zones (such as the crossbeam and counterweight areas); combined with GPS / UWB positioning data, the semantics of the actions are redefined: if "squatting" is detected and the location is in the "pit", it is judged as "normal maintenance"; if "squatting" is detected and the location is in the "edge of the car top", it is judged as "dangerous leaning", and the risk level is upgraded to high risk; Handheld object interaction verification: For behaviors such as "smoking" or "operating a mobile phone", the YOLOv5 object detection algorithm is used to verify whether there are "mobile phone" or "cigarette" features in the hand area, so as to achieve dual confirmation of action and object; This invention comprehensively assesses different categories of dangerous behaviors by combining behavioral sequences and spatial locations. For example, for illegal work on beams, it simultaneously determines whether the worker is in the designated work area, whether they are wearing a safety belt, and changes in their center of gravity, and calculates the hazard level by combining the amplitude of movements in consecutive frames. For falls, it makes judgments based on sudden changes in the center of gravity, abnormal joint angles, and abnormal movement speeds. For handheld objects (such as smoking or operating a mobile phone), it identifies handheld actions and abnormal operations by combining key hand points with object detection. S54: Risk Level Quantitative Assessment: Based on the verified stability assessment results, and considering the harmfulness and urgency of the behavior, determine the risk level according to predefined mapping rules. Risk level The expression is: 1 indicates Level 1 warning - alert level; 2 indicates Level 2 warning - alert level; 3 indicates Level 3 warning - emergency level; 4 indicates Level 4 warning - rescue level. Level 4 warning: This corresponds to behaviors that may directly lead to serious injury or death, such as electric shock, falls from heights, fires, and explosions. The power supply to the work must be cut off immediately, emergency rescue must be initiated, and the local emergency department must be notified. Level 3 warning: For behaviors that may cause serious injury, such as welding without wearing a mask, operating special equipment in violation of regulations, or falling, an audible and visual alarm must be sounded immediately, work must be stopped, and the on-site safety officer must be notified. Level 2 warning: This applies to behaviors that may cause minor injuries, such as not wearing protective clothing, standing in violation of regulations, smoking, or holding onto cables. On-site voice reminders should be given, violations should be corrected, and records should be kept on file. Level 1 Warning: For habitual violations such as not wearing a safety helmet, using a mobile phone, or not fastening cuffs, only a reminder message is sent to the worker's terminal to provide safety education; By introducing a composite verification of electrical logic, spatial location, and equipment status, false alarms from single visual recognition are effectively eliminated; by combining behavioral sequences with a three-dimensional electronic fence for comprehensive evaluation, the problem of difficult characterization of dangerous behaviors in complex scenarios is solved; and through four-level quantitative grading, precise allocation of early warning resources is achieved.

[0033] (6) Tiered early warning and proactive safety intervention Step S6: Based on risk level Implement differentiated intervention strategies, specifically including: S61. Warning Signal Generation and Strategy Matching: Based on the risk level determined in step S5, the pre-configured multi-level warning strategy library is invoked to generate specific warning execution instructions. The aforementioned early warning strategy library defines the mapping relationship between different risk levels and intervention methods, response time, and notification targets, ensuring that the intensity of early warning measures is proportional to the degree of risk and harm; avoiding alarm fatigue or resource waste caused by "one-size-fits-all" alarms; S62. Tiered Early Warning Information Distribution: Based on the policy library configuration, early warning information is distributed to different terminals through heterogeneous communication networks to achieve precise reach. For Level 1 warning ( (Prompt Level): For those not wearing safety helmets or using mobile phones, only send a voice prompt ("Please pay attention to safety regulations") to the safety helmet headset of the worker, record the violation once, do not trigger equipment shutdown, and avoid interfering with normal operation; For Level II warning ( Warning Level): For behaviors that may cause minor injury, such as smoking or touching cables, the on-site sound and light alarm will flash yellow, a warning sound will be broadcast, and alarm screenshots and short video clips will be pushed to the on-site safety officer's APP to urge on-site correction. For Level 3 warning ( Emergency Level: In response to behaviors that may cause serious injury, such as violations at heights or falls; a red light flashes on site accompanied by an alarm, automatically triggering the elevator emergency stop circuit (through relay linkage control cabinet) to prevent secondary injury caused by accidental elevator movement, and at the same time, an alarm is displayed on the management office's large screen; For Level IV warning ( (Rescue Level): In response to behaviors such as falls and unconsciousness that may directly lead to serious injury or death, in addition to implementing Level 3 warning measures, the system will automatically dial the emergency number and use two-way voice intercom to confirm the person's status, and activate the emergency rescue plan if necessary. S63. On-site equipment linkage control: Based on risk level, execute physical-level safety interventions to achieve "technical prevention" to prevent "human" negligence. Hard-wired linkage: When When the emergency stop signal is activated, the early warning execution unit sends an emergency stop signal to the welding machine / crane control cabinet through the relay output module (DO interface) to forcibly cut off the main power supply or brake and lock the equipment operation. Environmental equipment control: when Furthermore, when a fire hazard is identified, the automatic sprinkler system or smoke exhaust fan will be triggered; when a toxic gas leak is identified, the forced ventilation system will be activated to create a physical safety barrier. S64. Closed-loop feedback and handling tracking: The system records the sending time and receiving status of the early warning command, and initiates a countdown monitoring mechanism. Grace period upgrade: If within the preset grace period (e.g. If no dangerous behavior is detected within 10 seconds (i.e., the stability determination result in step S5 returns to 0), the risk level will be automatically raised by one level (e.g., from...). Rise to (and further notify the superior authority); Event archiving: After the operation is completed, the system automatically generates a "Safety Incident Closed-Loop Report" containing alarm time, location, risk type, and handling results, and archives it to the database as a basis for safety performance evaluation and training; By strictly mapping the risk level of step S5, the precise allocation of early warning resources was achieved; by hard-wired linkage and physical environment control, the accident chain was substantially interrupted; and by the closed-loop feedback mechanism, the continuous monitoring and handling of dangerous conditions were ensured, forming a complete safety management closed loop. (7) Data closed loop and model self-evolution management By storing data throughout the entire lifecycle, driving proactive learning and model iteration through challenging examples, and providing precise training based on security profiles, we construct a closed-loop evolutionary system integrating "data-algorithm-management" to achieve continuous self-optimization of the identification model and a spiral increase in the level of security management. Based on a closed-loop data management layer, a closed-loop optimization mechanism of "data acquisition - model iteration - personnel management" is constructed to solve the problems of model aging and scenario adaptability during long-term system use. Specifically, this includes: The system features full lifecycle storage, which categorizes and archives the multi-source heterogeneous data generated in steps S1 to S6 to ensure data traceability and usability. Time-series data storage stores the original video streams, key point heatmaps, fused feature vectors, and alarm logs in a time-series database (TSDB). Each record is tagged with a millisecond-level timestamp, device ID, and GPS / UWB location tag, supporting rapid retrieval and backtracking by time axis and spatial location. Object storage stores alarm-related short video clips, screenshots, and "Security Incident Closed-Loop Reports" in a distributed object storage system for evidence collection and review. Active Learning and Model Iteration: Establishing a "difficult example sample library" to address the problem of insufficient model generalization ability through active learning mechanisms. Difficult example mining: For "fuzzy samples" with a model confidence level between 0.5 and 0.7, they are automatically uploaded to the cloud-based difficult example sample library; Manual verification and annotation: Security experts conduct manual verification and precise annotation to create a high-quality incremental dataset; Periodic transfer learning: Every quarter, the accumulated incremental dataset is used to perform transfer learning on the HRNet and CNN-LSTM models at the edge. By freezing the bottom feature extraction layer and fine-tuning the top classification layer, the model can continuously adapt to new work scenarios (such as new equipment, process changes), seasonal changes in clothing (such as the impact of heavy winter work clothes on posture recognition) and new violations, so as to achieve continuous evolution of the model. Safety profile generation and management: Based on stored historical data, a multi-dimensional safety evaluation system is built on the safety management portal; the background automatically statistically analyzes the frequency of violations, types of high-risk behaviors, time distribution, and rectification status of each worker, and automatically generates a "Worker Safety Profile"; based on the safety profile analysis results, targeted and customized safety training courseware is pushed (such as pushing case videos of "accidents caused by distraction" for workers who "frequently play with their mobile phones"; and forcing workers who "violate regulations in high-altitude operations" to push theoretical and practical assessments on "correct use of safety belts").

[0034] By storing data throughout the entire lifecycle, it provides tamper-proof data evidence for accident tracing and liability determination; by actively learning and transferring learning, it solves the pain point of traditional AI models becoming outdated as soon as they go online, ensuring a high recognition rate throughout the system's lifecycle; by creating safety profiles and providing customized training, it shifts safety management from "post-incident punishment" to "pre-incident prevention" and "precision education," thereby improving the overall level of safety management.

[0035] like Figure 2 As shown, a second aspect of the present invention provides an intelligent identification and early warning system for dangerous behaviors of elevator operators, used to implement the above method, comprising: a multimodal perception layer, an edge computing layer, a data analysis and decision-making layer, and a data closed-loop management layer; The multimodal perception layer is configured to collect video stream data and sensor-aided data from the elevator operating area; The edge computing layer is communicatively connected to the multimodal perception layer and is configured to process the video stream data and sensor-aided data, and output a hazard determination result. The data analysis and decision-making layer is communicatively connected to the edge computing layer and is configured to determine the risk level based on the hazard assessment results and generate multi-level early warning strategies. The data closed-loop management layer is connected to the edge computing layer and the data analysis and decision-making layer, respectively, and is configured to store job data and iteratively optimize the identification model.

[0036] Furthermore, the multimodal perception layer is responsible for collecting all data and solving the problem of complex environmental perception in elevator shafts, car tops and machine rooms, including a visual acquisition unit and a wearable sensing unit. The visual acquisition unit includes a special camera module deployed in the welding operation area; the special camera module includes a wide-angle infrared integrated camera and a depth camera; the wide-angle infrared integrated camera is deployed at the top and bottom of the shaft, supports automatic aperture and backlight compensation, and automatically switches to infrared mode when the light is insufficient (<15Lux) to acquire 1080P high-definition video stream; the depth camera (ToF / RGB-D) is deployed in the car top operation area to acquire depth information Z to assist in three-dimensional coordinate mapping; The wearable sensing unit includes a smart safety helmet and environmental monitoring sensors; the smart safety helmet integrates a six-axis IMU (accelerometer + gyroscope), a Hall sensor (to detect the opening and closing of the mask) and an RFID positioning tag to collect workers' posture data and identity information in real time; the environmental monitoring sensors include temperature and humidity, smoke and harmful gas sensors deployed on site to help determine the risks of the working environment.

[0037] Furthermore, the edge computing layer is deployed in the industrial edge computing gateway in the elevator machine room or shaft, responsible for low-latency real-time inference and local linkage control; including video preprocessing module, posture recognition engine, feature engineering module, fusion inference module and real-time linkage interface; The video preprocessing module is used to denoise, correct distortion, and scale the resolution of the raw video stream, reducing the bandwidth pressure on data transmission. The pose recognition engine detects key points on the human skeleton and constructs a set of 3D coordinates. It is equipped with an optimized HRNet inference model to perform key point detection and output the temporal coordinates of each key point. 3D coordinate set And execute the occlusion compensation algorithm; The feature engineering module has a built-in feature calculation unit, which is used to calculate joint angle features, motion speed features, spatial region features and equipment interaction features based on the human key point set, and to integrate multi-dimensional features to construct video feature vectors. The fusion inference module is used to load a pre-trained CNN-LSTM hybrid model, perform multimodal data fusion and dangerous behavior classification, and output a probability distribution; specifically, it will... With sensor data Perform weighted fusion to generate fused feature vectors The data is then input into a CNN-LSTM hybrid model to output the probability distribution of dangerous behaviors. ; Real-time linkage interface: It is hardwired to the elevator control cabinet via GPIO / Relay interface. When a level 3 or higher warning is received, the elevator emergency stop circuit is triggered in milliseconds. Furthermore, the data analysis and decision-making layer is deployed on cloud servers or local monitoring centers, including risk assessment units, multi-level early warning and dispatch units, and digital twin visualization modules; The risk assessment unit is used to implement dynamic threshold determination logic; combined with 3D GIS electronic fence data, it performs spatial semantic analysis on behavior; and generates risk levels. (1 indicates a prompt, 2 indicates a warning, 3 indicates an emergency, and 4 indicates a rescue). Multi-level early warning and dispatch units are used to determine risk levels. Different early warning strategies can be invoked, and the system can connect to VoIP voice gateways (for on-site broadcasting), SMS / AppPush platforms (for administrators), and audible and visual alarms. The digital twin visualization module is used to build a 3D digital twin model of the elevator shaft, which maps the position and posture of the workers in real time, enabling intuitive monitoring. Furthermore, the data closed-loop management layer is used for data storage, traceability, and model self-evolution; The data closed-loop management layer includes: Time Series Database (TSDB): Stores massive amounts of key point coordinates, feature vectors, and sensor time series data; Object Storage: Stores raw video streams, alarm screenshots, and short video clips; Model training and iteration platform: Difficult example mining: Automatically filter fuzzy samples with confidence levels between 0.5 and 0.7 for manual annotation; Incremental learning: Regularly fine-tune and update the HRNet and CNN-LSTM models at the edge using newly labeled data; Safety Management Portal: Provides a web interface that supports historical alarm queries, safety report generation (such as "Monthly Violation Statistics"), and display of safety profiles of operators.

[0038] It should be noted that the elevator operator dangerous behavior intelligent identification and early warning system provided in this embodiment can be a computer program (including program code) running on a computer device. For example, the elevator operator dangerous behavior intelligent identification and early warning system is an application software; the elevator operator dangerous behavior intelligent identification and early warning system can be used to execute the corresponding steps in the above-mentioned method provided in the embodiments of this application.

[0039] This invention can automatically identify various dangerous behaviors during elevator operations, including but not limited to: not wearing a safety helmet, touching escalator cables, unauthorized work on high-altitude beams, smoking inside the elevator, using a mobile phone, and falls. By comprehensively analyzing the worker's posture, range of motion, changes in center of gravity, and the state of their held items, the system can accurately distinguish between normal operations and dangerous behaviors in complex work scenarios, which is impossible with traditional manual inspections or single visual monitoring.

[0040] The system of this invention not only utilizes attitude information acquired from video, but also combines sensor data (such as inertial measurement units, accelerometers, or position sensors) for multimodal fusion analysis. By training and recognizing the fused features through deep learning models (CNN, LSTM, or hybrid models), the accuracy and robustness of identifying dangerous behaviors in scenarios with insufficient lighting, occlusion, or minimal movement are improved.

[0041] To address the varying severity of different hazardous behaviors, the system employs a multi-level early warning mechanism. This mechanism triggers alarms the instant a hazardous behavior occurs, including audible and visual alarms, mobile notifications, and management alerts, thereby enabling immediate intervention and reducing safety risks. Compared to traditional methods, this invention provides real-time early warnings before or immediately after a hazardous behavior occurs, effectively improving operational safety.

[0042] This invention quantifies key human features such as position, joint angles, movement speed, and center of gravity changes, and constructs a behavioral sequence model. By analyzing continuous movement sequences, the system can identify high-risk behaviors such as unstable postures and abnormal hand operations during falls or high-altitude operations, achieving dynamic behavior monitoring and trend prediction, which is difficult to achieve in traditional monitoring.

[0043] The system of this invention stores all video, posture information, behavioral characteristics and alarm data, enabling traceability and long-term analysis of dangerous behaviors. By performing statistical analysis and model optimization on historical data, it can continuously improve the accuracy of dangerous behavior identification and provide data support for safety training and work process optimization, forming a complete safety management closed loop.

[0044] The system of this invention is applicable to various work scenarios such as elevator maintenance, repair, installation, and troubleshooting, and can cope with confined spaces, high-altitude operations, uneven lighting, and obstructed environments. Through multimodal fusion and deep learning analysis, this invention can achieve stable and reliable intelligent recognition of dangerous behaviors in various complex environments, which is unmatched by existing technologies.

[0045] This application also provides a computer-readable storage medium storing a computer program that is executed by a processor to implement... Figure 1 The methods provided in each step are detailed in the implementation methods provided in the above steps, and will not be repeated here.

[0046] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for intelligent identification and early warning of dangerous behaviors of elevator operators, characterized in that, Includes the following steps: S1: Acquire video stream data and sensor auxiliary data of the welding operation area; S2: Perform human skeleton key point detection on the video stream data and extract the set of human key points in the current frame; S3: Calculate joint angle features, motion speed features, spatial region features, and equipment interaction features based on the human body key point set, and fuse multi-dimensional features to construct a video feature vector; S4: Perform feature-level fusion of the video feature vector constructed in step S3 and the sensor-assisted data synchronously collected in step S1 to generate a fused feature vector; input the fused feature vector into a pre-trained dangerous behavior recognition model to output dangerous behavior categories and their probability distributions; based on the output probability distribution, extract the maximum probability value and compare it with a preset dynamic threshold to determine whether the behavior is dangerous, and finally output the behavior classification result. S5: When the maximum probability value is greater than or equal to the preset dynamic threshold, it is determined that the target dangerous behavior has occurred, and spatial semantic analysis is performed on the dangerous behavior category in combination with 3D GIS electronic fence data to determine the risk level; S6: Trigger the corresponding multi-level early warning strategy based on the risk level.

2. The intelligent identification and early warning method for dangerous behaviors of elevator operators according to claim 1, characterized in that, Step S1 includes: By deploying special camera modules and sensor nodes in the welding operation area, the system simultaneously acquires anti-strong light narrowband filtered video streams, infrared thermal imaging video streams, depth information video streams, panoramic high-definition video streams, and wide-angle infrared integrated video streams. By integrating sensor nodes into welding equipment and the working environment, personal protective equipment data, environmental perception data, and equipment operating condition data are collected simultaneously. Personal protective equipment data is collected through a smart safety helmet with integrated Hall effect sensors, a six-axis IMU, and RFID positioning tags. Environmental perception data is collected through temperature, humidity, smoke, and harmful gas sensors. Equipment operating condition data is collected through current transformers and voltage transformers. The collected multi-source data is time-stamp aligned and format standardized. The video stream data is decoded into RGB image frame sequences through the edge computing gateway, the sensor auxiliary data is converted into a structured numerical matrix, and the data association relationship is established with millisecond-level timestamps as indexes to form a multimodal dataset containing spatiotemporal synchronization information. Based on real-time collected ambient light intensity, dust concentration, and welding current parameters, the exposure time, gain, and filter switching strategy of the camera module are dynamically adjusted.

3. The intelligent identification and early warning method for dangerous behaviors of elevator operators according to claim 2, characterized in that: The construction of the visual feature vector in step S3 includes: S31: Calculate the human joint angle features and bone proportion features based on the three-dimensional key point set; S32: Calculate the instantaneous velocity and acceleration characteristics of key points based on the coordinates of 3D key points in multiple consecutive frames; S33: Map the coordinates of the human centroid to a pre-constructed 3D electronic map of the welding operation area, calculate the distance between the centroid projection point and the boundary of the preset safe area; at the same time, extract the relative positional relationship between the key points of the hand and the dangerous area, and generate a spatial occupancygrid feature vector. S34: Based on the key points of the welding equipment separated in step S24, calculate the interaction feature vectors between the hand, welding gun, and mask. S35: Concatenate the geometric features, motion features, position features, and interaction features extracted in steps S31 to S34 to construct a high-dimensional video feature vector; use the Z-score normalization method to normalize the feature vector and eliminate the influence of different dimensions.

4. The intelligent identification and early warning method for dangerous behaviors of elevator operators according to claim 3, characterized in that: Joint angle in step S31 The expression is: ,in, For the first Each joint angle , , These are the three adjacent key points that make up the joint.

5. The intelligent identification and early warning method for dangerous behaviors of elevator operators according to claim 3, characterized in that: Instantaneous velocity of key points in step S32 The calculation expression is: ,in, In order to be in At that moment, the Spatial coordinates of key points; In order to be in At that moment, the Spatial coordinates of key points; for Time and The time interval between moments.

6. A method for intelligent identification and early warning of dangerous behaviors of elevator operators according to any one of claims 3-5, characterized in that: The visual feature vector constructed in step S35 A multi-dimensional, layered fusion structure is adopted, and the expression is: in, This represents the vector concatenation operator; Features of joint angles; Kinematic characteristics; Features of a spatial region; For equipment interaction features; Indicates joint angle characteristics; Indicates the instantaneous velocity characteristics of key points; Indicates from the 1st to the 1st The angle values ​​of each key joint at the current moment; Indicates from the 1st to the 1st The speed of movement of each key point at the current moment; Indicates from the 1st to the 1st The acceleration of each key point at the current moment; Indicates from the 1st to the 1st The angular velocity of a key point at the current moment; The distance between the human body and the ground; Distance between human body and danger zone; For area occupancy coding; The distance between the wrist key point and the center point of the welding torch inspection frame; The angle between the welding torch and the normal vector of the working plane; This refers to the proportion of key facial features that are covered by the mask.

7. The intelligent identification and early warning method for dangerous behaviors of elevator operators according to claim 6, characterized in that: In step S4, feature vectors are fused. The expression is: ,in, For visual feature vectors; Provide auxiliary data for sensors; For dynamic fusion weights, ; Probability distribution of dangerous behavior categories in step S4 The expression is: in, and These represent the weights and biases of the CNN-LSTM hybrid deep learning model, respectively.

8. The intelligent identification and early warning method for dangerous behaviors of elevator operators according to claim 7, characterized in that: In step S4, the expression for determining whether an action is dangerous is: in, This indicates that a dangerous behavior has been determined to have occurred. 0 indicates that no dangerous behavior was detected; Let be the risk probability distribution vector output by the deep learning model at time t; This indicates taking the maximum probability value from the risk probability distribution vector; if the highest confidence level predicted by the model is greater than or equal to the set threshold. If the highest confidence level is less than the threshold, the result is 1, indicating that a dangerous behavior has occurred and an alarm is triggered; if the highest confidence level is less than the threshold, the result is 1. If the result is 0, the current state is considered safe, and no alarm is triggered; regarding the threshold Special settings: For high-risk behaviors, For low-risk behaviors, ;like If so, then it is determined that a danger has occurred.

9. A method for intelligent identification and early warning of dangerous behaviors of elevator operators according to any one of claims 1-5, 7, and 8, characterized in that: Risk level in step S5 The expression is: 1 indicates Level 1 warning - alert level; 2 indicates Level 2 warning - alert level; 3 indicates Level 3 warning - emergency level; 4 indicates Level 4 warning - rescue level. In step S6, for a Level 1 warning, only a voice prompt is sent to the worker's safety helmet headset: "Please pay attention to safety regulations," and the violation is recorded once; no shutdown is initiated. In response to a Level 2 warning, the on-site audible and visual alarm flashes a yellow light, a warning sound is broadcast, and an alarm screenshot is pushed to the on-site safety officer's APP. In response to the Level 3 warning, the red light on site flashes and is accompanied by an alarm siren, automatically triggering the elevator emergency stop circuit to prevent the elevator from moving unexpectedly and causing secondary damage. At the same time, an alarm is displayed on the management office's large screen. In response to a Level 4 alert, in addition to implementing Level 3 alert measures, an emergency call will be automatically made, and the status of personnel will be confirmed using two-way voice communication.

10. An intelligent identification and early warning system for dangerous behaviors of elevator operators, characterized in that, The method for intelligent identification and early warning of dangerous behaviors of elevator operators as described in any one of claims 1-9 includes: a multimodal perception layer, an edge computing layer, a data analysis and decision-making layer, and a data closed-loop management layer; The multimodal perception layer is configured to collect video stream data and sensor-aided data from the elevator operating area; The edge computing layer is communicatively connected to the multimodal perception layer and is configured to process the video stream data and sensor-aided data, and output a hazard determination result. The data analysis and decision-making layer is communicatively connected to the edge computing layer and is configured to determine the risk level based on the hazard assessment results and generate multi-level early warning strategies. The data closed-loop management layer is connected to the edge computing layer and the data analysis and decision-making layer, respectively, and is configured to store job data and iteratively optimize the identification model; The multimodal perception layer includes a visual acquisition unit and a wearable sensing unit; the visual acquisition unit includes a special camera module deployed in the welding operation area; the wearable sensing unit includes a smart safety helmet and an environmental monitoring sensor; the smart safety helmet integrates a six-axis IMU, a Hall sensor and an RFID positioning tag; the environmental monitoring sensor includes temperature and humidity, smoke and harmful gas sensors deployed on site; The edge computing layer includes a video preprocessing module, a pose recognition engine, a feature engineering module, a fusion inference module, and a real-time linkage interface. The video preprocessing module is used to denoise, correct distortion, and scale the resolution of the raw video stream. The pose recognition engine is used to detect key points of the human skeleton and construct a set of 3D coordinates. The pose recognition engine is equipped with an optimized HRNet inference model to perform key point detection of the human skeleton and output the time-space coordinates of each key point. 3D coordinate set And execute the occlusion compensation algorithm; The feature engineering module has a built-in feature calculation unit, which is used to calculate joint angle features, motion speed features, spatial region features and equipment interaction features based on the human body key point set, and to fuse multi-dimensional features to construct video feature vectors. Fusion Inference Module: Used to load the trained CNN-LSTM hybrid model, perform multimodal data fusion and dangerous behavior classification, and output probability distribution; Real-time linkage interface: It is hardwired to the elevator control cabinet via GPIO / Relay interface. When a level 3 or higher warning is received, the elevator emergency stop circuit is triggered in milliseconds. The data analysis and decision-making layer includes a risk assessment unit, a multi-level early warning and dispatch unit, and a digital twin visualization module. The risk assessment unit combines 3D GIS electronic fence data to perform spatial semantic analysis of behavior and generate risk levels. 1 indicates a prompt, 2 indicates a warning, 3 indicates an emergency, and 4 indicates a rescue. Multi-level early warning and dispatch units are used to determine risk levels. Different early warning strategies can be invoked, and connections can be made with VoIP voice gateways, SMS / AppPush platforms, and sound and light alarms; The digital twin visualization module is used to build a 3D digital twin model of the elevator shaft, which maps the position and posture of the workers in real time. The data closed-loop management layer includes: a time-series database, which stores massive amounts of key point coordinates, feature vectors, and sensor time-series data; Object storage stores raw video streams, alarm screenshots, and short video clips; Model training and iteration platform: Automatically selects fuzzy samples with confidence levels between 0.5 and 0.7 for manual annotation; regularly uses newly annotated data to fine-tune and update the HRNet and CNN-LSTM models at the edge; Safety Management Portal: Provides a web interface that supports historical alarm queries, safety report generation, and display of safety profiles for operators.