Elevator shaft work safety risk real-time identification method based on visual AI

CN121354115BActive Publication Date: 2026-09-22LIXIN (JIANGSU) INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511673463.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-09-22
Estimated Expiration
2045-11-14

AI Technical Summary

Technical Problem

近年来,随着物联网和人工智能技术的发展,基于传感器的电梯安全监测系统逐渐应用于电梯运行安全监控领域,但针对维保作业人员的实时安全保障系统仍有较大发展空间

Benefits of technology

本发明提供的基于视觉AI的电梯井道作业安全风险实时识别方法能够实现风险的早期识别与精准预测,将传统的被动式安全管理转变为主动预防模式,有效减少电梯井道作业事故发生率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121354115B_ABST
    Figure CN121354115B_ABST
Patent Text Reader

Abstract

The application provides an elevator shaft operation safety risk real-time identification method based on visual AI, relates to the field of AI technology, and comprises the following steps: acquiring a real-time image sequence through an image acquisition device, deducing a three-dimensional skeleton topology structure of an operator and environmental characteristics, and constructing a multi-modal semantic representation vector; deriving a safety distance field and deducing a future position probability distribution of the operator to form a risk heat map; constructing a nonlinear mapping function to generate an adaptive early warning instruction; and executing hierarchical intervention and optimizing a decision mechanism. The application can realize real-time risk identification and early warning of elevator shaft operation and improve operation safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to AI technology, and more particularly to a method for real-time identification of safety risks in elevator shaft operations based on visual AI. Background Technology

[0002] Elevator shaft work is a crucial part of elevator installation, maintenance, and repair, but it is also a high-risk work area. Elevator shafts are confined spaces with significant vertical drops, limited working space, and numerous potential hazards, including moving parts and temporary scaffolding. Traditional elevator shaft safety management relies primarily on manual supervision and fixed safety procedures. Workers typically wear safety belts, helmets, and other protective equipment and follow standard operating procedures. In recent years, with the development of IoT and AI technologies, sensor-based elevator safety monitoring systems have been increasingly applied to elevator operation safety monitoring. However, there is still significant room for development in real-time safety assurance systems for maintenance personnel.

[0003] Existing elevator shaft operation safety monitoring technologies have the following shortcomings: First, traditional monitoring methods lack the ability to accurately identify the behavior and posture of workers, making it difficult to capture the relative position of personnel to hazardous areas in real time and thus unable to proactively predict potential risks. Second, existing systems typically employ fixed-threshold early warning mechanisms, which cannot adaptively adjust to dynamic changes in the work environment and personnel behavioral characteristics, resulting in poor early warning effectiveness and false alarms or missed alarms. Finally, most safety monitoring systems only focus on single-dimensional risk identification, lacking the ability to analyze workers' historical behavioral patterns and future action trends, making it impossible to build a complete risk assessment model and achieve preventative safety interventions.

[0004] With the rapid development of visual artificial intelligence technology, computer vision-based human posture recognition, behavior analysis, and environmental perception technologies have provided new technical pathways for elevator shaft operation safety. By processing image data in real time through deep learning algorithms and combining multimodal sensor information fusion, a comprehensive perception of the working environment and personnel status can be achieved, enabling the construction of more accurate safety risk assessment models and providing intelligent safety assurance for elevator shaft operations. Summary of the Invention

[0005] This invention provides a method for real-time identification of safety risks in elevator shaft operations based on visual AI, which can solve the problems in the prior art.

[0006] A first aspect of this invention provides a method for real-time identification of safety risks in elevator shaft operations based on visual AI, comprising: Real-time image sequences of the work space are acquired by an image acquisition device deployed in the elevator shaft. Based on the real-time image sequences, the three-dimensional skeleton topology, component-level integrity features, and obstacle distribution features of the workers are deduced. The spatial mapping relationship between the three-dimensional skeleton topology, the component-level integrity features, and the obstacle distribution features is derived, and a multimodal semantic representation vector under a unified coordinate system is constructed. Based on the multimodal semantic representation vector, the minimum safe distance field between the worker and the dangerous boundary is derived, and the probability distribution of the worker's future position is inferred by fusing the worker's historical movement trajectory. A risk heat map is constructed based on the overlap between the two. Based on the risk heatmap, a nonlinear mapping function between risk level and early warning strategy is constructed. The local gradient features of the risk heatmap are coupled with real-time environmental parameters to derive the spatiotemporal threshold for early warning triggering and generate an adaptive early warning instruction set. The adaptive early warning instruction set guides the audible and visual early warning device and the work suspension control device to perform graded interventions, forming closed-loop feedback data. The closed-loop feedback data is then used to optimize the nonlinear mapping function, promoting the dynamic evolution of risk identification and early warning decisions.

[0007] Based on the real-time image sequence, the 3D skeleton topology, component-level integrity features, and obstacle distribution features of the workers are deduced. The spatial mapping relationship between the 3D skeleton topology, component-level integrity features, and obstacle distribution features is derived, and a multimodal semantic representation vector in a unified coordinate system is constructed, including: A multi-frame joint detection model is constructed for the real-time image sequence. The evolution law of the two-dimensional joints of the operator is derived through dynamic temporal features. The three-dimensional spatial activity features of the joints are constructed by fusing the joint movement trajectory between adjacent frames and the depth constraint conditions of the shaft space. The three-dimensional skeleton topology of the operator is deduced based on the three-dimensional spatial activity features. Based on the distribution of joints in the three-dimensional skeleton topology, the dynamic boundary of the protective equipment area is derived. The standard configuration template within the dynamic boundary is dynamically mapped to the actual test component. Based on the dynamic topology mapping, the temporal evolution characteristics of the component integrity are deduced. Based on the temporal evolution characteristics, a dynamic perception range of the well environment area is constructed. Scene semantic analysis is performed on the dynamic perception range. The spatial distribution dynamic characteristics of obstacles are constructed by combining the spatiotemporal evolution law of obstacle pixel areas and the parameter matrix of the image acquisition device through the scene semantic analysis. Based on the three-dimensional spatial activity features, the temporal evolution features, and the spatial distribution dynamic features, a mapping function for the spatial coordinates of the joint points is constructed. The dynamic coupling relationship between the three is derived through the mapping function. According to the dynamic coupling relationship, the spatial coordinates of the joint points, the obstacle boundary features, and the coordinate mapping parameters are dynamically combined to construct a multimodal semantic representation vector with spatiotemporal correlation.

[0008] Based on the distribution of joints in the three-dimensional skeleton topology, the dynamic boundary of the protective equipment area is derived. A dynamic topological mapping is then performed between the standard configuration template within the dynamic boundary and the actual tested components. Based on this dynamic topological mapping, the temporal evolution characteristics of component integrity are deduced, including: The set of joint points corresponding to the wearing position of the protective equipment in the three-dimensional skeleton topology is derived. A spatial envelope range is constructed based on the set of joint points. The dynamic boundary of the protective equipment area is generated according to the spatial envelope range and the standard coverage size of the protective equipment. Based on the dynamic boundary, candidate detection regions for protective equipment are constructed in the real-time image sequence. The candidate detection regions are semantically parsed to deduce the contour features of the actual detection components. The contour features are then mapped to the spatial range constructed by the dynamic boundary. Based on the spatial range, a spatial mapping relationship is established between the component nodes in the standard configuration template and the contour features. The matching tolerance of the spatial mapping relationship is adaptively adjusted according to the deformation of the dynamic boundary. By calculating the number of missing nodes and the node offset distance between the standard configuration template and the actual detected component, the component integrity feature value of the current frame is constructed. The temporal evolution feature of component integrity is derived based on the changing trend of the component integrity feature value.

[0009] Based on the multimodal semantic representation vector, a minimum safe distance field between the worker and the hazardous boundary is derived. The worker's historical movement trajectory is then fused to infer the probability distribution of their future location. A risk heatmap is constructed based on the overlap between the two, including: The coordinate components of the three-dimensional skeleton topology and obstacle distribution features in the multimodal semantic representation vector are deduced. Based on the coordinate components, the Euclidean distance between the spatial position of the operator's key point and the spatial position of the danger boundary is calculated. The minimum value of the Euclidean distance is derived to construct the minimum safe distance. The association rules between the spatial position and the minimum safe distance are established using the grid cells of the shaft spatial coordinate system to generate the minimum safe distance field. A temporal evolution model is constructed for the coordinate components in the multimodal semantic representation vector. Based on the temporal evolution model, the historical position sequence of the operator's joints is derived to construct the historical motion trajectory. The displacement vector and velocity vector at adjacent moments are calculated based on the historical motion trajectory. The probability distribution of the operator's future position is inferred through the distribution characteristics of the displacement vector and the velocity vector. Based on the probability distribution of the future position, a spatial mapping is constructed in the grid cells of the minimum safe distance field. The overlapping area of ​​the two in the spatial grid is derived, and the grid cells with insufficient safe distance in the overlapping area are used to construct the risk area. For the grid cells in the risk area, the risk intensity is constructed by combining the corresponding minimum safe distance and the probability of occurrence of the location, and the risk intensity is converted into color coding to generate a risk heat map.

[0010] Based on the risk heatmap, a nonlinear mapping function between risk levels and early warning strategies is constructed. The local gradient features of the risk heatmap are coupled with real-time environmental parameters to derive the spatiotemporal threshold for early warning triggering and generate an adaptive early warning instruction set, including: A standardized feature space is constructed by deriving the risk intensity value of each grid cell in the risk heat map. Based on the standardized feature space, a multi-level risk level interval is constructed, and the early warning strategy type corresponding to each risk level interval is derived. Based on the historical correlation characteristics between different risk level intervals and corresponding early warning strategy types in the risk heatmap, the nonlinear evolution law of risk level change rate and early warning strategy response time is deduced, and a nonlinear mapping function between risk level and early warning strategy is constructed. The gradient vector features of the risk heatmap in the spatial dimension are derived. Based on the magnitude and direction of the gradient vector, local gradient features are constructed. The environmental correction coefficient is deduced by combining the real-time environmental parameters of the well space. The output early warning strategy of the nonlinear mapping function is optimized by using the environmental correction coefficient. Based on the magnitude and direction of the local gradient features and the environmental correction coefficient, the spatiotemporal threshold conditions for triggering the early warning are derived. An adaptive early warning instruction set is constructed by combining the early warning strategy optimized by the nonlinear mapping function with the spatiotemporal threshold conditions.

[0011] The adaptive early warning instruction set guides the audible and visual early warning device and the work suspension control device to perform tiered interventions, forming closed-loop feedback data. This closed-loop feedback data is then used to optimize the nonlinear mapping function, promoting the dynamic evolution of risk identification and early warning decisions, including: A deep neural network model is constructed to determine the warning type and response time limit in the adaptive warning instruction set. The execution level of the graded intervention is deduced through the deep neural network model. The warning control signal of the audible and visual warning device and the suspension control signal of the work suspension control device are generated according to the execution level. The audible and visual warning device and the work suspension control device are activated to perform graded intervention. Construct response time feature vectors after the sound and light early warning device and the work stoppage control device perform graded intervention, and fuse the position change trajectory of the workers and the change trend of the risk heat map after the graded intervention is performed to generate closed-loop feedback data; The closed-loop feedback data is analyzed to establish a dynamic mapping matrix between intervention execution parameters and risk mitigation effects. Based on the deviation distribution law of the dynamic mapping matrix, the mapping coefficient of the nonlinear mapping function is optimized through the pre-intervention risk level in the closed-loop feedback data. The optimization parameters of the nonlinear mapping function are embedded into the risk identification and early warning decision-making process. Based on the continuously accumulated closed-loop feedback data, a periodic optimization mechanism for the nonlinear mapping function is constructed to realize the dynamic evolution of risk identification and early warning decision-making.

[0012] The deep neural network model deduces the execution levels of graded intervention, and generates warning control signals for the audible and visual warning device and pause control signals for the work stoppage control device based on the execution levels, including: Multi-source image sequences from the work site are collected to construct a time-series data stream. Spatiotemporal feature vectors of the time-series data stream are extracted using a deep neural network model. The spatiotemporal feature vectors are divided into multiple state intervals according to the time span. Based on the state intervals, the temporal correlation coefficients between consecutive frames are extracted. The temporal correlation coefficients are then fused to construct a prediction benchmark for the execution level of graded intervention. Based on the execution level prediction benchmark, a parameter matrix for audible and visual early warning and pause control is constructed. Multiple sets of candidate control parameter combinations are sampled in the parameter matrix using the Monte Carlo method. A comprehensive score of execution cost and early warning effect is calculated for each set of candidate control parameter combinations. The parameter combination with the best comprehensive score is selected to generate early warning control signal and pause control signal.

[0013] A second aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0014] A third aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0015] The beneficial effects of this application are as follows: The real-time safety risk identification method for elevator shaft operations based on visual AI provided by this invention can achieve early identification and accurate prediction of risks, transforming traditional passive safety management into an active prevention mode, and effectively reducing the accident rate of elevator shaft operations.

[0016] By constructing multimodal semantic representation vectors and risk heatmaps, this invention solves the problem of difficulty in quantifying and assessing the distance between workers and dangerous boundaries in complex environments. It can automatically identify potential dangerous states and provide sufficient early warning time windows before risk events occur, thereby improving the accuracy and timeliness of the early warning system.

[0017] This invention achieves adaptive optimization of the early warning strategy, automatically adjusting the intervention intensity according to different risk levels, avoiding the problems of over-intervention or insufficient early warning caused by traditional fixed thresholds. At the same time, it continuously optimizes the identification algorithm through a closed-loop feedback mechanism, enabling the system to have self-learning and evolution capabilities, adapt to different elevator shaft operating environments, and improve the system's practicality and reliability. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the real-time identification method for safety risks in elevator shaft operations based on visual AI, according to an embodiment of the present invention. Figure 2 This is a flowchart of the multimodal correlation evolution feature extraction process for dynamic well environment according to an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0021] Figure 1 This is a flowchart illustrating the real-time safety risk identification method for elevator shaft operations based on visual AI, as described in an embodiment of the present invention. Figure 1 As shown, the method includes: Real-time image sequences of the work space are acquired by an image acquisition device deployed in the elevator shaft. Based on the real-time image sequences, the three-dimensional skeleton topology, component-level integrity features, and obstacle distribution features of the workers are deduced. The spatial mapping relationship between the three-dimensional skeleton topology, the component-level integrity features, and the obstacle distribution features is derived, and a multimodal semantic representation vector under a unified coordinate system is constructed. Based on the multimodal semantic representation vector, the minimum safe distance field between the worker and the dangerous boundary is derived, and the probability distribution of the worker's future position is inferred by fusing the worker's historical movement trajectory. A risk heat map is constructed based on the overlap between the two. Based on the risk heatmap, a nonlinear mapping function between risk level and early warning strategy is constructed. The local gradient features of the risk heatmap are coupled with real-time environmental parameters to derive the spatiotemporal threshold for early warning triggering and generate an adaptive early warning instruction set. The adaptive early warning instruction set guides the audible and visual early warning device and the work suspension control device to perform graded interventions, forming closed-loop feedback data. The closed-loop feedback data is then used to optimize the nonlinear mapping function, promoting the dynamic evolution of risk identification and early warning decisions.

[0022] In one optional implementation, based on the real-time image sequence, the three-dimensional skeleton topology, component-level integrity features, and obstacle distribution features of the operator are deduced. The spatial mapping relationship between the three-dimensional skeleton topology, component-level integrity features, and obstacle distribution features is then derived, and a multimodal semantic representation vector in a unified coordinate system is constructed, including: A multi-frame joint detection model is constructed for the real-time image sequence. The evolution law of the two-dimensional joints of the operator is derived through dynamic temporal features. The three-dimensional spatial activity features of the joints are constructed by fusing the joint movement trajectory between adjacent frames and the depth constraint conditions of the shaft space. The three-dimensional skeleton topology of the operator is deduced based on the three-dimensional spatial activity features. Based on the distribution of joints in the three-dimensional skeleton topology, the dynamic boundary of the protective equipment area is derived. The standard configuration template within the dynamic boundary is dynamically mapped to the actual test component. Based on the dynamic topology mapping, the temporal evolution characteristics of the component integrity are deduced. Based on the temporal evolution characteristics, a dynamic perception range of the well environment area is constructed. Scene semantic analysis is performed on the dynamic perception range. The spatial distribution dynamic characteristics of obstacles are constructed by combining the spatiotemporal evolution law of obstacle pixel areas and the parameter matrix of the image acquisition device through the scene semantic analysis. Based on the three-dimensional spatial activity features, the temporal evolution features, and the spatial distribution dynamic features, a mapping function for the spatial coordinates of the joint points is constructed. The dynamic coupling relationship between the three is derived through the mapping function. According to the dynamic coupling relationship, the spatial coordinates of the joint points, the obstacle boundary features, and the coordinate mapping parameters are dynamically combined to construct a multimodal semantic representation vector with spatiotemporal correlation.

[0023] like Figure 2 As shown, the method includes: Real-time image sequence acquisition relies on a binocular stereo vision system deployed in the shaft environment. This system consists of two high-resolution industrial cameras with a baseline distance precisely set at 120 mm, a uniform focal length of 8 mm, and a field of view coverage of 70 degrees. The image acquisition frequency is set at 30 frames per second, the resolution is configured at 1920×1080 pixels, and the exposure time is strictly controlled within 1 / 60th of a second to ensure clear imaging of fast-moving targets. The cameras are installed at the center of the shaft top, with a vertical downward viewing angle of 30 degrees. Vibration-resistant brackets are used for installation and fixation to ensure image stability. Image transmission is achieved via a gigabit Ethernet interface, with data transmission latency controlled within 10 milliseconds. Image data is stored in a lossless compression format to prevent information loss.

[0024] The image preprocessing module performs multi-level processing operations on the original image sequence, including noise reduction. A 3×3 Gaussian kernel is used for convolution filtering with a standard deviation parameter set to 1.2. The filtering intensity is dynamically adjusted based on the image noise level. Brightness equalization is achieved through an adaptive histogram equalization algorithm, dividing the image into 8×8 pixel blocks. Each block is independently histogram equalized, and bilinear interpolation is used to smooth the transition between blocks and avoid block artifacts. Edge enhancement uses the Sobel operator to extract gradient information. After calculating the horizontal and vertical gradients, edge features are described by the gradient magnitude and orientation angle. A gradient threshold of 30 is set to filter out weak edge noise.

[0025] The multi-frame keypoint detection model is built upon an improved deep convolutional neural network architecture. The network employs an encoder-decoder structure, with the encoder containing five convolutional blocks, each consisting of two convolutional layers and a max-pooling layer. The network input layer is designed to receive five consecutive frames of images as temporal feature input. The size of each frame is normalized to 416×416 pixels using bicubic interpolation, and pixel values ​​are normalized to between 0 and 1. The feature extraction layer uses a multi-scale convolutional kernel parallel processing strategy: a 3×3 convolutional kernel extracts detailed features, a 5×5 convolutional kernel extracts medium-scale features, and a 7×7 convolutional kernel extracts global contextual features. The output channels of the three branches are set to 64, 128, and 256, respectively, and dimensionality is reduced to 128 dimensions through 1×1 convolution.

[0026] The joint point localization layer achieves precise localization through heatmap regression, outputting the two-dimensional coordinates of 17 key human joint points, including the top of the head, neck, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, and left and right ankles. Each joint point corresponds to a 64×64 pixel heatmap, and the peak position of the heatmap is the joint point coordinate. The peak response is modeled using a two-dimensional Gaussian distribution with a standard deviation set to 2 pixels. The joint point confidence is represented by the peak intensity of the heatmap, with a confidence threshold set to 0.6. Detection results below this threshold are marked as invalid and trigger a re-detection mechanism.

[0027] Dynamic temporal feature derivation is achieved by establishing motion vector fields for key points between adjacent frames. Motion vector calculation employs the Lucas-Kanade optical flow algorithm, constructing a three-layer image pyramid with a scaling factor of 0.5 for each layer. The optical flow window size is set to 15×15 pixels, the maximum number of iterations is limited to 30, and the convergence accuracy is set to 0.01 pixels. Key point evolution patterns are extracted using a sliding time window analysis method. The time window length is set to 10 frames, corresponding to approximately 0.33 seconds, and the sliding step size is set to 3 frames to ensure temporal continuity. Key point trajectories within the window are fitted using cubic spline interpolation to obtain smooth motion curves, with 20 interpolation nodes to ensure trajectory accuracy. The fusion of key point motion trajectories between adjacent frames uses an exponentially weighted moving average method. Weight allocation is based on the inverse decay of time distance, with a decay factor set to 0.8. In the fusion formula, the current frame has a weight of 0.8, the previous frame has a weight of 0.16, the two previous frames have a weight of 0.032, and so on until the weight is less than 0.01.

[0028] The shaft space depth constraint is addressed by obtaining a dense depth map using a stereo matching algorithm. The stereo matching employs a semi-global matching algorithm, with the disparity search range set to 0 to 128 pixels and the matching window size set to 9×9 pixels. The matching cost function combines Census transform and absolute difference. The depth information calculation accuracy is determined by the baseline distance and focal length parameters, achieving a depth accuracy of 5 millimeters at a distance of 2 meters and 20 millimeters at a distance of 5 meters. Depth map post-processing includes left-right consistency checking, hole filling, and median filtering. The left-right consistency threshold is set to 1 pixel, hole filling uses nearest neighbor interpolation, and the median filtering kernel size is set to 5×5 pixels.

[0029] The 3D spatial activity feature construction combines 2D keypoint coordinates with depth information to complete 3D reconstruction. Coordinate transformation uses a pinhole camera model. The intrinsic parameter matrix includes focal lengths fx and fy (800 and 805 pixels respectively), principal point coordinates cx and cy (320 and 240 pixels respectively), radial distortion coefficients k1 and k2 (-0.2 and 0.05 respectively), and tangential distortion coefficients p1 and p2 (0.001 and -0.002 respectively). The extrinsic parameter matrix includes a rotation matrix represented by a Rodrigues vector (0.1, 0.05, 0.02), and a translation vector (0, 0, 4.0 meters) representing the camera's position in the world coordinate system. The coordinate transformation process first performs distortion correction, then recovers the 3D coordinates through inverse projection transformation and depth information, achieving centimeter-level coordinate accuracy within the working distance range.

[0030] The 3D skeleton topology derivation is based on prior knowledge of human joint connections to establish a skeleton model. The skeleton contains 16 bone connections: head-neck, neck-to-shoulder, shoulder-to-elbow, elbow-to-wrist, torso, hip-to-knee, and knee-to-ankle. The length of each bone is obtained by calculating the Euclidean distance from the 3D coordinates of its two ends, and the bone orientation is represented by a normalized vector. Skeletal integrity verification employs statistical methods, establishing a standard human proportion model with a head-to-body ratio of 1:7.5, an arm span to height ratio of 1:1, an upper arm to forearm length ratio of 1:0.85, and a thigh to lower leg length ratio of 1:0.9. Connections with actual measurements deviating from the standard proportions by more than 20% are marked as abnormal, triggering a skeleton correction algorithm for geometric constraint optimization.

[0031] The dynamic boundary derivation of the protective equipment area is based on the 3D skeleton joint point distribution and implemented using a geometric expansion algorithm. Boundary calculation employs a joint point expansion strategy: the head region expands outward with a radius of 150 mm around the top joint point to form a spherical boundary; the torso region expands outward with a radius of 200 mm around the line connecting the shoulder and hip joint points as the central axis to form a capsule-shaped boundary; and the limb regions expand outward with a radius of 100 mm around the corresponding joint point lines as the central axis to form cylindrical boundaries. Dynamic boundary updates utilize temporal smoothing filtering to avoid drastic boundary fluctuations. The filter is a first-order low-pass filter with a cutoff frequency set to 2 Hz, limiting the boundary change rate to within 10 mm per frame after filtering. Boundary collision detection employs the separating axis theorem and bounding box hierarchy to accelerate calculations, achieving millimeter-level accuracy and controlling the detection latency to within 5 milliseconds.

[0032] The standard configuration template includes multi-dimensional feature descriptions of four types of protective equipment. The safety helmet template is defined as a yellow HSV color range (hue 45-65 degrees, saturation 70-100%, brightness 40-90%), with an elliptical geometry and a major-to-minor axis ratio of 1.2:1, offset upwards by 10-30 mm relative to the head position. The reflective vest template is defined as an orange HSV color range (hue 10-25 degrees, saturation 80-100%, brightness 60-95%), with a rectangular geometry and a length-to-width ratio of 1.5:1, covering the shoulder to waist area relative to the torso. The protective glove template is defined as a dark HSV color range (any hue, saturation 10-40%, brightness 10-30%), with a geometry matching the hand contour and completely covering the hand area. The safety shoe template is defined as a dark HSV color range similar to the gloves, with a geometry matching the shoe shape and completely covering the foot area.

[0033] The dynamic topology mapping between the actual detected component and the standard configuration template is achieved through a multi-feature fusion matching algorithm. Color feature matching uses Euclidean distance in the HSV color space to calculate similarity, with a distance threshold set at 30 color units, and a matching success rate of over 85%. Shape feature matching employs contour extraction and Hu moment invariant calculation. Contour detection uses the Canny edge detection algorithm, with a low threshold set at 50 and a high threshold at 150. The dilation and erosion operation kernel size is set to 3×3 pixels. Hu moment calculation includes seven invariant features, and shape similarity is measured using weighted Euclidean distance. The weight allocation is based on feature importance, with the first three moments weighted at 0.3, 0.25, and 0.2, and the last four moments weighted at 0.1 or less. Relative positional relationship verification calculates the spatial distance and angular deviation between the centroid of the detected component and its corresponding joint point, with a distance tolerance set at 50 mm and an angle tolerance set at 15 degrees, requiring a positional matching success rate of over 90%.

[0034] In one optional implementation, the dynamic boundary of the protective equipment area is derived based on the joint distribution in the three-dimensional skeleton topology. A dynamic topological mapping is then performed between the standard configuration template within the dynamic boundary and the actual tested component. Based on this dynamic topological mapping, the temporal evolution characteristics of the component integrity are deduced, including: The set of joint points corresponding to the wearing position of the protective equipment in the three-dimensional skeleton topology is derived. A spatial envelope range is constructed based on the set of joint points. The dynamic boundary of the protective equipment area is generated according to the spatial envelope range and the standard coverage size of the protective equipment. Based on the dynamic boundary, candidate detection regions for protective equipment are constructed in the real-time image sequence. The candidate detection regions are semantically parsed to deduce the contour features of the actual detection components. The contour features are then mapped to the spatial range constructed by the dynamic boundary. Based on the spatial range, a spatial mapping relationship is established between the component nodes in the standard configuration template and the contour features. The matching tolerance of the spatial mapping relationship is adaptively adjusted according to the deformation of the dynamic boundary. By calculating the number of missing nodes and the node offset distance between the standard configuration template and the actual detected component, the component integrity feature value of the current frame is constructed. The temporal evolution feature of component integrity is derived based on the changing trend of the component integrity feature value.

[0035] The derivation of the joint set corresponding to the wearing position of protective equipment in the 3D skeleton topology is achieved by establishing a prior knowledge mapping table of human anatomy. This mapping table defines the joints corresponding to the top of the head for head protection equipment, the joints corresponding to the neck, left and right shoulders, and left and right hips for torso protection equipment, the joints corresponding to the left and right wrists for hand protection equipment, and the joints corresponding to the left and right ankles for foot protection equipment. Joint set extraction uses an index lookup method. The input is a 3D coordinate array of 17 joints in the 3D skeleton topology. Based on the protective equipment type identifier, the corresponding joint index list is retrieved, and the spatial coordinate set of the target joint is output. Joint coordinate validity verification is achieved through confidence threshold filtering. Joints with a confidence level below 0.7 are marked as invalid and compensated using adjacent frame interpolation. The interpolation method is cubic spline interpolation, and the interpolation window length is set to 5 frames, corresponding to a time span of approximately 0.17 seconds.

[0036] The spatial envelope is constructed based on the convex hull calculation of the keypoint set. A 3D fast convex hull algorithm is used to geometrically enclose the keypoint coordinates, with the calculation accuracy set to the millimeter level. The envelope boundary is represented by a triangular mesh, and the mesh vertex coordinates are kept to three decimal places. Envelope expansion is achieved by normal vector extrapolation. Each triangular facet is extrapolated a fixed distance along its normal vector direction. The extrapolation distance is determined according to the type of protective equipment: 80 mm for head protection, 120 mm for torso protection, and 50 mm for hand and foot protection. The expanded envelope surface is smoothed using the Laplacian smoothing algorithm, with 10 smoothing iterations and a smoothing coefficient of 0.5 to ensure the continuity and differentiability of the envelope boundary.

[0037] The dynamic boundary generation of protective equipment areas combines the spatial envelope range with standard coverage dimensions. These standard coverage dimensions are obtained by querying a 3D model library of protective equipment, which contains standard geometric dimensions for helmets, reflective vests, protective gloves, and safety shoes. The standard dimensions of a helmet are an ellipsoid with a major axis of 280 mm, a minor axis of 240 mm, and a height of 160 mm; the standard dimensions of a reflective vest are a columnar shape with a chest width of 600 mm, a back width of 550 mm, and a height of 800 mm; the standard dimensions of protective gloves are a hand-shaped outline with a length of 220 mm, a width of 110 mm, and a thickness of 40 mm; and the standard dimensions of safety shoes are a shoe-shaped outline with a length of 290 mm, a width of 110 mm, and a height of 120 mm. Dynamic boundary calculation is performed by aligning the centroid of the standard-sized geometry with its corresponding joint points, aligning the geometry's pose with the local coordinate system of the skeleton, and taking the union of the outer surface of the geometry and its extended envelope range. This union calculation is implemented using Boolean operations, maintaining millimeter-level precision.

[0038] In real-time image sequences, candidate detection regions for protective equipment are constructed through 2D projection of dynamic boundaries. The 3D boundary is projected onto the image plane using a pinhole camera model. The camera intrinsic parameter matrix includes focal length parameters fx = 1200 pixels, fy = 1205 pixels, principal point coordinates cx = 640 pixels, cy = 480 pixels, radial distortion coefficients k1 = -0.15, k2 = 0.08, and tangential distortion coefficients p1 = 0.002, p2 = -0.001. The projection transformation process first transforms the 3D boundary vertex coordinates from the world coordinate system to the camera coordinate system. The transformation matrix includes rotation and translation parameters. Then, normalized image coordinates are obtained through perspective projection. Finally, pixel coordinates are obtained by applying the intrinsic parameter matrix and distortion correction. The candidate detection region boundary is determined by calculating the convex hull of the projected vertices. The Graham scan algorithm is used, with a time complexity of O(nlogn). The number of vertices in the boundary polygon is limited to 20 to control computational complexity.

[0039] Semantic parsing of candidate detection regions is implemented using a deep learning-based semantic segmentation network. The network architecture is DeepLabV3+, with a ResNet50 backbone network for the encoder and a combination of ASPP modules and upsampling layers for the decoder. The network input is RGB image patches of the candidate regions, with patch sizes normalized to 512×512 pixels and pixel values ​​normalized to the range of 0 to 1. The network output is the probability distribution of each pixel's category, including five categories: background, safety helmet, reflective vest, protective gloves, and safety shoes. Post-processing of semantic segmentation includes probability thresholding, connected component analysis, and contour extraction. The probability threshold is set to 0.8, the connected component area threshold is set to 100 pixels, and the contour extraction uses the Canny edge detection algorithm with a low threshold of 50 and a high threshold of 150.

[0040] The actual component contour feature extraction is achieved through contour analysis and shape descriptor calculation. Contour simplification employs the Douglas-Peucker algorithm, with a simplification tolerance set to 2 pixels, and the number of contour points after simplification controlled to within 50. Contour features include geometric features such as area, perimeter, major and minor axis ratios, roundness, and rectangularity. Area is calculated using Green's theorem, perimeter is calculated using Euclidean distance accumulation, major and minor axis ratios are obtained through principal component analysis, roundness is defined as 4π multiplied by the area divided by the square of the perimeter, and rectangularity is defined as the contour area divided by the area of ​​the smallest bounding rectangle. The contour feature vector is set to 12 dimensions, including 2 dimensions for position coordinates, 5 dimensions for geometric features, 3 dimensions for color features, and 2 dimensions for texture features. The feature vector is normalized to the range of 0 to 1 to eliminate the influence of dimensions.

[0041] The mapping of contour features to the spatial range constructed by the dynamic boundary is achieved through inverse projection transformation. The coordinates of the two-dimensional contour points are transformed by the camera parameters to obtain the three-dimensional ray equation. The intersection of the ray and the dynamic boundary surface is the corresponding position of the contour in three-dimensional space. The intersection point calculation adopts the ray and triangle intersection algorithm, and the intersection determination adopts the centroid coordinate method, with the calculation accuracy maintained to three decimal places. In the case of multiple intersection points, the intersection point closest to the camera is selected as the effective mapping position. The coordinates of the mapping position are stored in the three-dimensional contour point array, and the array index corresponds one-to-one with the two-dimensional contour point.

[0042] The spatial mapping relationship between component nodes and contour features in the standard configuration template is established using a nearest neighbor matching algorithm. The standard configuration template is defined as a set of key feature points for protective equipment. For example, a safety helmet template contains five feature points: vertex, leading edge, trailing edge, left side, and right side; a reflective vest template contains four feature points: shoulder, chest, back, and waist; a protective glove template contains three feature points: fingertips, palm, and wrist; and a safety shoe template contains three feature points: toe, heel, and side. The matching algorithm calculates the Euclidean distance between template feature points and contour feature points. The distance matrix is ​​established using a brute-force search method with a time complexity of O(mn), where m is the number of template points and n is the number of contour points. Optimal matching is solved using the Hungarian algorithm, with the matching cost being the distance value. The algorithm output is a one-to-one mapping relationship between template points and contour points.

[0043] Dynamic boundary deformation calculation is achieved by measuring the difference between the current frame boundary and the previous frame boundary. This difference is measured using Hausdorff distance, defined as the maximum-minimum distance between two point sets, with a computational complexity of O(n^2). A deformation threshold of 10 mm is set; deformation exceeding this threshold triggers an adaptive adjustment mechanism for the matching tolerance. The initial matching tolerance value is set to 20 mm, and the adjustment strategy uses linear interpolation. The tolerance adjustment range is 10 mm to 50 mm, with an adjustment step size of 5 mm. The tolerance adjustment considers the exponentially weighted moving average of historical deformations, with a weight decay factor set to 0.8 and a moving average window length of 10 frames.

[0044] The number of missing nodes between the standard configuration template and the actual tested components is calculated through matching result statistics. Template nodes whose matching distance exceeds the current tolerance threshold are marked as missing nodes, and the number of missing nodes is the total number of missing nodes. Node offset distance is calculated as the three-dimensional Euclidean distance between successfully matched node pairs. Offset distance statistics include average offset distance, maximum offset distance, and standard deviation of offset distance. Offset distance weighting is based on the importance of template nodes, with key nodes having a weight of 1.0 and secondary nodes having a weight of 0.5. The weighted offset distance is the sum of the products of each node's offset distance and its weight.

[0045] The current frame component integrity feature value construction adopts a multi-dimensional feature fusion method. The feature vector includes four dimensions: number of missing nodes, weighted offset distance, contour area ratio, and color similarity. The number of missing nodes is normalized to a range of 0 to 1, corresponding to missing nodes from 0 to the total number of template nodes. The weighted offset distance is normalized by dividing by the maximum expected offset distance, which is set to 100 mm. The contour area ratio is defined as the detected contour area divided by the standard template area, with a ratio range of 0 to 2. A ratio greater than 1 indicates that the detected area is larger than the standard size. Color similarity is calculated using Euclidean distance in the HSV color space, with a similarity range of 0 to 1, where 1 indicates complete similarity. The integrity feature value is calculated using a weighted average method, with weights allocated as follows: number of missing nodes 0.4, weighted offset distance 0.3, contour area ratio 0.2, and color similarity 0.1. The feature value ranges from 0 to 1, where 0 indicates complete absence and 1 indicates complete wear.

[0046] The temporal evolution characteristics of component integrity are derived using sliding window analysis. The temporal window length is set to 30 frames corresponding to a 1-second time span, and the sliding step size is set to 1 frame to ensure temporal continuity. The integrity feature value sequence within the window is used to extract evolutionary features through trend analysis, which includes statistical features such as linear regression slope, variance, number of peaks, and rate of change. The linear regression slope represents the trend of integrity change; a positive slope indicates improved integrity, a negative slope indicates decreased integrity, and the absolute value of the slope represents the rate of change. Variance represents integrity stability; low variance indicates stable wearing status, and high variance indicates frequent status changes. The number of peaks represents the frequency of wearing status switching; peak detection uses the local extremum method, and the peak threshold is set to the mean integrity value plus or minus the standard deviation. The rate of change represents the average difference in integrity between adjacent frames; a high rate of change indicates unstable status. The temporal evolution feature vector is set to 4 dimensions, and the feature values ​​are normalized to the range of -1 to +1. The temporal feature update frequency is synchronized with the image acquisition frequency to ensure real-time requirements.

[0047] In one optional implementation, the minimum safe distance field between the worker and the hazardous boundary is derived based on the multimodal semantic representation vector, and the probability distribution of the worker's future location is inferred by fusing the worker's historical movement trajectory. A risk heatmap is constructed based on the overlap between the two, including: The coordinate components of the three-dimensional skeleton topology and obstacle distribution features in the multimodal semantic representation vector are deduced. Based on the coordinate components, the Euclidean distance between the spatial position of the operator's key point and the spatial position of the danger boundary is calculated. The minimum value of the Euclidean distance is derived to construct the minimum safe distance. The association rules between the spatial position and the minimum safe distance are established using the grid cells of the shaft spatial coordinate system to generate the minimum safe distance field. A temporal evolution model is constructed for the coordinate components in the multimodal semantic representation vector. Based on the temporal evolution model, the historical position sequence of the operator's joints is derived to construct the historical motion trajectory. The displacement vector and velocity vector at adjacent moments are calculated based on the historical motion trajectory. The probability distribution of the operator's future position is inferred through the distribution characteristics of the displacement vector and the velocity vector. Based on the probability distribution of the future position, a spatial mapping is constructed in the grid cells of the minimum safe distance field. The overlapping area of ​​the two in the spatial grid is derived, and the grid cells with insufficient safe distance in the overlapping area are used to construct the risk area. For the grid cells in the risk area, the risk intensity is constructed by combining the corresponding minimum safe distance and the probability of occurrence of the location, and the risk intensity is converted into color coding to generate a risk heat map.

[0048] The coordinate component derivation of the 3D skeleton topology in the multimodal semantic representation vector is achieved through a vector parsing module. This module receives a 2048-dimensional multimodal semantic representation vector as input. The first 512 dimensions of the vector encode the 3D skeleton topology information, including the 3D coordinate data of 17 key points. Coordinate component extraction employs a linear transformation method, mapping the 512-dimensional skeleton feature vector to a 51-dimensional coordinate vector through a fully connected layer, corresponding to the x, y, and z coordinate components of each of the 17 key points. The coordinate decoding precision is set to the millimeter level, and the coordinate values ​​are limited to the effective area of ​​the shaft space coordinate system: the x-axis range is -5000 mm to +5000 mm, the y-axis range is -3000 mm to +3000 mm, and the z-axis range is 0 mm to 30000 mm. Coordinate validity verification is achieved through boundary checks and continuity constraints. Coordinate values ​​exceeding the boundary range are automatically truncated to the boundary value. Key points with coordinate changes exceeding 500 mm between adjacent frames are marked as abnormal and smoothed using Kalman filtering.

[0049] The coordinate component derivation of obstacle distribution features employs a similar vector analysis process. Dimensions 513 to 1024 of the multimodal semantic representation vector encode the spatial distribution information of obstacles, including boundary vertex coordinates and geometric feature descriptors. Obstacle coordinate decoding is achieved through a multilayer perceptron network containing three hidden layers, each with 256 neurons. The ReLU activation function is used, and the output layer dimension is variable to accommodate different numbers of obstacle vertices. Obstacle boundaries are represented using a convex hull, with the number of vertices in the convex hull of each obstacle limited to 20, and vertex coordinate accuracy maintained at the millimeter level. The spatial location of hazardous boundaries is obtained by expanding the outer normal vector of the obstacle convex hull. The expansion distance is determined according to the obstacle type: 1500 mm for mechanical equipment boundaries, 2000 mm for electrical equipment boundaries, 3000 mm for high-temperature areas, and 5000 mm for high-voltage areas.

[0050] The Euclidean distance between the spatial locations of worker joints and the spatial locations of hazardous boundaries is calculated using a point-to-surface distance algorithm. The joint coordinates serve as the query point, and the triangular facet of the hazardous boundary serves as the reference surface. The distance calculation process includes calculating the perpendicular distance from the point to the surface and the distance from the point to the nearest point within the surface; the smaller of these two values ​​is taken as the final distance. The perpendicular distance is obtained by substituting the equation from the point to the plane, and the plane equation is constructed using the three vertices of the triangular facet. The distance to the nearest point within the surface is calculated using a centroid coordinate system projection method. When the projected point is inside the triangle, the distance is the perpendicular distance; when the projected point is outside the triangle, the distance is the shortest distance from the point to the triangle boundary. The distance calculation accuracy is maintained to one decimal place, corresponding to 0.1 mm precision, with a computational complexity of O(nm), where n is the number of joints and m is the number of hazardous boundary facets.

[0051] The minimum safe distance is derived by comparing the distances of 17 key points to all hazardous boundaries. For each key point, the distance to all hazardous boundary patches is calculated, and the minimum value in the set of distances is taken as the minimum safe distance for that key point. The global minimum safe distance for each key point is defined as the minimum safe distance between the worker and the hazardous boundary. This value is updated in real time and stored in a distance cache array. The distance cache array is set to a length of 100 and uses a circular queue structure to support historical distance queries and trend analysis. Abnormal distance filtering is implemented using statistical methods; data points with distance values ​​exceeding the historical average plus three standard deviations are marked as abnormal and replaced with the valid distance value from the previous frame.

[0052] The hoistway spatial coordinate system is established using a uniform grid division method with a grid resolution of 500 mm cube units. The hoistway space has a length of 10,000 mm, a width of 6,000 mm, and a height of 30,000 mm, corresponding to 14,400 grid units of 20×12×60. The grid unit index uses a three-dimensional array structure, calculated as i + j × 20 + k × 240, where i, j, and k are the grid coordinates in the x, y, and z directions, respectively. Association rules for spatial locations to minimum safe distances are implemented using a hash table data structure. The hash table key is the grid unit index, and the corresponding value is the minimum safe distance for that grid unit. Association rule updates use a spatial interpolation method. The minimum safe distance within a grid unit is calculated by trilinear interpolation of the distances to the surrounding eight vertices, with the interpolation weights allocated based on the reciprocal of the Euclidean distance from the query point to the vertex.

[0053] The minimum safe distance field is generated through grid traversal and a distance propagation algorithm. The algorithm starts with a grid cell whose distance is known and gradually propagates the distance information to adjacent grid cells. Distance propagation uses a variant of Dijkstra's algorithm, with the propagation weight being the Euclidean distance between grid cells. The propagation terminates when all grid cells have obtained a valid minimum safe distance value. The distance field is smoothed using a Gaussian filter with a kernel size of 3×3×3 and a standard deviation of 1.0. The filtering iterations are performed five times to eliminate noise and discontinuities in the distance field. The distance field data is stored in a sparse matrix format, storing only non-zero distance values ​​and their indices. The storage space compression ratio is approximately 85%, and the access time complexity is O(logn).

[0054] The temporal evolution model of coordinate components in the multimodal semantic representation vector employs a Long Short-Term Memory (LSTM) network architecture. The network input is a sequence of coordinate components from 10 consecutive frames, with each frame containing a 51-dimensional keypoint coordinate vector. The network structure consists of two LSTM layers with 128 hidden states per layer, and a dropout probability of 0.2 is used to prevent overfitting. The network training employs a mean squared error loss function, with the Adam algorithm as the optimizer, a learning rate of 0.001, a batch size of 32, and 200 training epochs. The training data is collected from real-world wellbore operation scenarios, comprising 50 hours of continuous operation video, with annotation accuracy down to the millimeter level for keypoint coordinate localization. The model inference latency is controlled within 10 milliseconds, meeting the requirements for real-time applications.

[0055] Historical motion trajectories are constructed using the output sequence of a temporal evolution model. The model output consists of predicted joint coordinates for the next 5 frames, which, combined with the input coordinates from the previous 10 frames, form a complete trajectory sequence of 15 frames. Trajectory smoothing employs a Savitzky-Golay filter with a 7-frame filtering window and a 3rd-order polynomial, ensuring the continuity of velocity and acceleration after filtering. Trajectory validity is verified through physical constraints: the maximum movement speed of human joints is limited to 5 meters per second, and the maximum acceleration is limited to 20 meters per second squared. Trajectory points exceeding these limits are marked as abnormal and corrected using linear interpolation.

[0056] Displacement vectors between adjacent time steps are calculated by subtracting the coordinates of trajectory points. The displacement vector is the coordinates of the current frame's keypoints from the coordinates of the previous frame's keypoints, with vector precision maintained to 0.1 millimeters. Velocity vectors are calculated by dividing the displacement vector by the time interval, which is the reciprocal of the frame rate. A standard frame rate of 30 frames per second corresponds to a 33.33 millisecond interval. Velocity vector smoothing employs an exponentially weighted moving average method with a smoothing factor set to 0.8, reducing the noise variance of the velocity vector by approximately 60%. Statistical features of the displacement and velocity vectors include higher-order moment features such as mean, standard deviation, skewness, and kurtosis. The feature vectors are 8-dimensional and used for subsequent probability distribution modeling.

[0057] The future position probability distribution projection is achieved using a Gaussian mixture model, which contains three Gaussian components, each corresponding to a different motion mode: stationary, uniform motion, and accelerated motion. The mean of the Gaussian components is determined by the moving average of the velocity vector, and the covariance matrix is ​​estimated from the historical distribution of the displacement vector. The mixture weights are obtained through training using the expectation-maximization algorithm, with the training data being the displacement and velocity vector sequences of the past 100 frames. The probability distribution prediction time window is set to the next 10 frames, corresponding to a time span of approximately 0.33 seconds. The prediction accuracy decreases as the time window increases, with the prediction error of the 10th frame being approximately five times that of the 1st frame. The probability density function is calculated in the grid cells using a numerical integration method, with an integration accuracy set to a probability interval of 0.01.

[0058] The spatial mapping of the future position probability distribution within the minimum safe distance field grid cells is achieved through a probabilistic projection algorithm, where the continuous function of the probability distribution is discretized into probability values ​​for each grid cell. The mapping process calculates the density value of the center point of each grid cell within the probability distribution; this density value is then normalized and used as the probability of that grid cell's occurrence. A probability threshold of 0.001 is set; grid cells with probabilities below this threshold are set to 0 to reduce computational complexity. The spatial range of the probability distribution is limited to a 5-meter radius around the worker's current position; probability values ​​outside this range are ignored.

[0059] The overlapping region derivation is achieved through Boolean operations. The overlapping condition is that the grid cell simultaneously meets two conditions: first, the minimum safe distance between the grid cells is less than a preset safety threshold; second, the probability of the grid cell's future location occurring is greater than a probability threshold. The safety thresholds are set according to the hazard type: 1500 mm for mechanical hazards, 2000 mm for electrical hazards, 3000 mm for high-temperature hazards, and 5000 mm for high-voltage hazards. The probability threshold is set to 0.1, indicating that areas with a probability of their future location occurring exceeding 10% are included in the risk assessment scope. The overlapping region calculation is accelerated using bitwise operations. The grid cell state is represented by a bit vector, and the overlapping operation is transformed into a bitwise AND operation, reducing the computational complexity to O(n / 32).

[0060] Risk regions are constructed by filtering grid cells in overlapping areas that have insufficient safety distances. The filtering criterion is that the minimum safety distance of the grid cell is less than the safety threshold for the corresponding hazard type. Risk region grid cell labeling employs a region growth algorithm, expanding from a seed grid cell in six adjacent directions. The expansion terminates when a grid cell with sufficient safety distance is encountered or at a region boundary. Connectivity analysis of risk regions uses a depth-first search algorithm to identify independent risk region clusters. Each cluster is assigned a unique identifier for subsequent risk quantification analysis.

[0061] Risk intensity calculation comprehensively considers two factors: minimum safe distance and the probability of location occurrence. Risk intensity is defined as the product of the degree of insufficient safe distance and the probability of occurrence. The degree of insufficient safe distance is calculated by subtracting the actual minimum safe distance from the safety threshold, then dividing by the safety threshold for normalization; the value ranges from 0 to 1. The probability of location occurrence is directly obtained from the probability distribution mapping, also ranging from 0 to 1. The formula for calculating risk intensity is: degree of insufficiency multiplied by probability of occurrence, then multiplied by a severity coefficient. The severity coefficient is set according to the type of hazard: 1.0 for mechanical hazards, 1.5 for electrical hazards, 2.0 for high-temperature hazards, and 3.0 for high-voltage hazards.

[0062] The conversion from risk intensity to color coding is achieved using HSV color space mapping. The hue component is fixed at 0 degrees, corresponding to red, while the saturation component is set to 100% to ensure color vibrancy. The lightness component is linearly mapped according to the risk intensity: a risk intensity of 0 corresponds to 0 lightness and black, and a risk intensity of 1 corresponds to 100% lightness and pure red. The color quantization precision is 256 levels, corresponding to 8-bit color depth. The color mapping lookup table is pre-calculated and cached to improve rendering performance. Risk heatmap generation uses texture rendering technology, where the color values ​​of grid cells are filled into the corresponding texture pixels. The texture resolution is consistent with the grid resolution, and the rendering frame rate is maintained at 30 frames per second to ensure real-time display. The transparency channel is set according to the risk intensity: low-risk areas are displayed semi-transparently to maintain background visibility, while high-risk areas are displayed opaquely to highlight the warning effect.

[0063] In one optional implementation, a nonlinear mapping function between risk levels and early warning strategies is constructed based on the risk heatmap. The local gradient features of the risk heatmap are coupled with real-time environmental parameters to derive the spatiotemporal threshold for early warning triggering and generate an adaptive early warning instruction set, including: A standardized feature space is constructed by deriving the risk intensity value of each grid cell in the risk heat map. Based on the standardized feature space, a multi-level risk level interval is constructed, and the early warning strategy type corresponding to each risk level interval is derived. Based on the historical correlation characteristics between different risk level intervals and corresponding early warning strategy types in the risk heatmap, the nonlinear evolution law of risk level change rate and early warning strategy response time is deduced, and a nonlinear mapping function between risk level and early warning strategy is constructed. The gradient vector features of the risk heatmap in the spatial dimension are derived. Based on the magnitude and direction of the gradient vector, local gradient features are constructed. The environmental correction coefficient is deduced by combining the real-time environmental parameters of the well space. The output early warning strategy of the nonlinear mapping function is optimized by using the environmental correction coefficient. Based on the magnitude and direction of the local gradient features and the environmental correction coefficient, the spatiotemporal threshold conditions for triggering the early warning are derived. An adaptive early warning instruction set is constructed by combining the early warning strategy optimized by the nonlinear mapping function with the spatiotemporal threshold conditions.

[0064] The risk intensity value of each grid cell in the risk heatmap is derived through a numerical extraction module. This module traverses all 14,400 grid cells of the risk heatmap, extracting the floating-point value of the risk intensity stored in each cell. The valid range of the risk intensity value is 0.0 to 3.0, where 0.0 represents a risk-free state and 3.0 represents an extremely high-risk state, with numerical precision maintained to three decimal places. Outlier detection is achieved through statistical methods. Grid cells with risk intensity values ​​exceeding the mean plus three standard deviations are marked as outliers, and the outlier is replaced with the median of the risk intensity of adjacent grid cells. The temporal consistency verification of risk intensity values ​​is achieved through differential checks between adjacent frames. Grid cells with a single-frame risk intensity change exceeding 1.5 trigger smoothing processing. The smoothing method uses a 3-frame moving average, with weights allocated as follows: 0.2 for the previous frame, 0.6 for the current frame, and 0.2 for the next frame.

[0065] The standardized feature space is constructed using the Z-score standardization method. The standardization transformation converts the risk intensity values ​​into a standard distribution with a mean of 0 and a standard deviation of 1. Standardization parameters include the global mean and global standard deviation. The mean is calculated as the arithmetic mean of the risk intensity values ​​across all grid cells, and the standard deviation is calculated as the square root of the variance of the risk intensity value relative to the mean. The standardized feature values ​​typically range from -3 to +3; values ​​outside this range are truncated to boundary values ​​to avoid the influence of extreme outliers. Feature space dimensionality expansion is achieved through local statistical features. Each grid cell's feature vector contains four dimensions: standardized risk intensity, neighborhood average risk intensity, neighborhood maximum risk intensity, and neighborhood standard deviation. The neighborhood is defined as a 3×3×3 cubic region.

[0066] The multi-level risk level intervals are constructed using the K-means clustering algorithm, with five clusters representing safe, low-risk, medium-risk, high-risk, and extremely high-risk levels. The input to the clustering algorithm is a 4-dimensional standardized feature vector, and the number of iterations is limited to 100. The convergence condition is that the change in cluster centers is less than 0.001. The K-means++ method is used to initialize the cluster centers, ensuring good dispersion of the initial centers. The silhouette coefficient is used to evaluate the clustering results; a silhouette coefficient greater than 0.5 indicates good clustering performance. The risk level interval boundaries are determined by the midpoints between cluster centers, and the boundary values ​​are stored in an interval partitioning table. The table structure includes fields such as level identifier, lower boundary, upper boundary, and center value.

[0067] The early warning strategy types are derived from the response requirements analysis based on risk levels. A safe level corresponds to no early warning strategy, a low-risk level corresponds to an information prompt strategy, a medium-risk level corresponds to a voice warning strategy, a high-risk level corresponds to an audible and visual alarm strategy, and an extremely high-risk level corresponds to an emergency shutdown strategy. Early warning strategy parameters include warning intensity, duration, repetition interval, and escalation conditions. The information prompt strategy has a display duration of 3 seconds and a repetition interval of 30 seconds. The voice warning strategy has a volume setting of 80 decibels, a duration of 5 seconds, and a repetition interval of 15 seconds. The audible and visual alarm strategy has a volume setting of 100 decibels, a flashing frequency of 2 Hz, a duration of 10 seconds, and a repetition interval of 5 seconds. The emergency shutdown strategy has an execution delay of 1 second, and manual reset is required to restart after shutdown.

[0068] Historical correlation analysis is implemented using an association rule mining algorithm. The algorithm input consists of a risk level change sequence over the past 7 days and corresponding early warning strategy response records. The support threshold for association rules is set to 0.1, the confidence threshold to 0.8, and the lift threshold to 1.2. Correlation characteristics include risk level transition patterns, early warning strategy execution delays, and early warning effect feedback. Transition pattern recognition is achieved through state transition matrix analysis, with a 5×5 matrix corresponding to the transition probabilities between 5 risk levels. Execution delay statistics include average delay, maximum delay, and delay variance, with delay time precision at the millisecond level. Effect feedback is quantified by the risk intensity decrease rate, calculated as the slope of the risk intensity change within 10 seconds after the early warning is triggered.

[0069] The nonlinear evolution law is deduced using a multinomial regression model. The model input is the rate of change of risk level, and the output is the response time of the early warning strategy. The rate of change is calculated as the risk level of the current frame minus the risk level of the previous frame, divided by the time interval, which is approximately 33.33 milliseconds (the reciprocal of the frame rate). The multinomial order is set to 3, and the model parameters are estimated using the least squares method. The model training data contains 1000 sample points, derived from statistical analysis of historical early warning records. Model validation uses 10-fold cross-validation, with the average prediction error controlled within 5%. The predicted response time range is 0.1 seconds to 10 seconds; predictions exceeding this range are truncated to the boundary value.

[0070] The nonlinear mapping function is constructed using a radial basis function network architecture, comprising an input layer, hidden layers, and an output layer. The input layer receives the rate of change of risk level and has a dimension of 1. The hidden layer contains 20 radial basis function units, with a Gaussian activation function, and the center and width parameters are determined using a clustering algorithm. The output layer outputs the probability distribution of the warning strategy types, with a dimension of 5 corresponding to the 5 warning strategies. The network is trained using the gradient descent algorithm with a learning rate of 0.01 and 500 training epochs. Training data augmentation is achieved through noise injection and data interpolation, expanding the dataset size to 5000 samples after augmentation. The network inference time is controlled within 1 millisecond, meeting the requirements for real-time warnings.

[0071] The gradient vector feature of the spatial dimension of the risk heatmap is derived using the finite difference method. Gradient calculation includes partial derivatives in the x, y, and z directions. The x-direction gradient is calculated by differentiating the risk intensity of adjacent grid cells, with a difference step size of 500 mm for each grid cell side. The y and z-direction gradients use the same difference method. The magnitude of the gradient vector is calculated as the square root of the sum of the squares of the partial derivatives in the three directions. The direction is calculated using the arctangent function, with angular accuracy maintained to 0.1 degrees. Gradient calculation for boundary grid cells uses a one-sided difference method to avoid out-of-bounds access. The gradient calculation results are stored in a three-dimensional array of the same size as the risk heatmap, with each array element containing the gradient magnitude and direction angle.

[0072] Local gradient features are constructed based on the magnitude and direction analysis of gradient vectors. Magnitude features include the gradient magnitude of the current grid cell, the maximum gradient magnitude in the neighborhood, the average gradient magnitude in the neighborhood, and the standard deviation of the gradient magnitude. Directional features include the dominant gradient direction, the direction consistency index, and the direction divergence index. The dominant gradient direction is calculated by a weighted average of the gradient vectors in the neighborhood, with the weight being the gradient magnitude. The direction consistency index is calculated as the average of the cosine of the angle between the gradient direction and the dominant direction in the neighborhood, ranging from 0 to 1, where 1 indicates perfect direction consistency. The direction divergence index is calculated as the standard deviation of the gradient direction angle, ranging from 0 to 180 degrees. The local gradient feature vector has 7 dimensions, and the feature values ​​are normalized to the range of 0 to 1 to eliminate the influence of dimensions.

[0073] Real-time environmental parameter acquisition is achieved through multi-sensor fusion, including temperature, humidity, noise, vibration, and light sensors. The temperature sensor measures from -20°C to 80°C with an accuracy of 0.1°C and a sampling frequency of 1 Hz. The humidity sensor measures from 0% to 100% relative humidity with an accuracy of 1% and a sampling frequency of 1 Hz. The noise sensor measures from 30 dB to 120 dB with an accuracy of 0.1 dB and a sampling frequency of 10 Hz. The vibration sensor measures from 0 to 50 m / s² with an accuracy of 0.01 m / s² and a sampling frequency of 100 Hz. The light sensor measures from 0 to 100,000 lux with an accuracy of 1 lux and a sampling frequency of 1 Hz. Sensor data is transmitted via a CAN bus, and the data packet format includes fields such as sensor identifier, timestamp, measured value, and status flags.

[0074] The environmental correction coefficient is derived using a multiple linear regression model. The model input is a 5-dimensional environmental parameter vector, and the output is a scalar value of the correction coefficient. The correction coefficient ranges from 0.5 to 2.0, where 1.0 indicates no correction, less than 1.0 indicates good environmental conditions reducing warning sensitivity, and greater than 1.0 indicates poor environmental conditions increasing warning sensitivity. The temperature correction rule increases the correction coefficient by 0.2 when the temperature exceeds 40 degrees Celsius or falls below 0 degrees Celsius. The humidity correction rule increases the correction coefficient by 0.1 when humidity exceeds 80%. The noise correction rule increases the correction coefficient by 0.3 when noise exceeds 90 decibels. The vibration correction rule increases the correction coefficient by 0.2 when vibration exceeds 10 meters per second squared. The illumination correction rule increases the correction coefficient by 0.1 when illumination is below 500 lux. When multiple correction rules are triggered simultaneously, the correction coefficients are accumulated, and the total correction coefficient is truncated to the boundary of the value range.

[0075] Nonlinear mapping function optimization is achieved through multiplicative adjustment of the environmental correction coefficient. The optimized warning strategy response time equals the original response time divided by the environmental correction coefficient. A correction coefficient greater than 1.0 shortens the response time, improving the warning response speed. A correction coefficient less than 1.0 lengthens the response time, reducing the false alarm rate. Switching between warning strategy types is achieved by comparing the corrected response time with a threshold: the highest-level warning strategy is selected when the response time is less than 1 second; a high-level warning strategy is selected when the response time is 1 to 3 seconds; a medium-level warning strategy is selected when the response time is 3 to 10 seconds; and a low-level warning strategy is selected when the response time is greater than 10 seconds. Dynamic adjustment of warning strategy parameters includes volume, frequency, and duration, with the adjustment magnitude proportional to the correction coefficient.

[0076] The derivation of spatiotemporal threshold conditions comprehensively considers local gradient characteristics and environmental correction coefficients. The spatial threshold is determined based on the gradient magnitude: when the gradient magnitude is greater than 0.5, the spatial threshold is set to the current grid cell and its first-order neighborhood; when the gradient magnitude is between 0.2 and 0.5, the spatial threshold expands to the second-order neighborhood; and when the gradient magnitude is less than 0.2, the spatial threshold expands to the third-order neighborhood. The temporal threshold is determined based on gradient direction consistency and environmental correction coefficients: when the direction consistency is greater than 0.8 and the correction coefficient is greater than 1.5, the time threshold is set to 100 milliseconds; when the direction consistency is between 0.5 and 0.8 and the correction coefficient is between 1.0 and 1.5, the time threshold is set to 300 milliseconds; and in other cases, the time threshold is set to 500 milliseconds. The dynamic update frequency of the threshold conditions is synchronized with the update frequency of the risk heatmap to ensure the real-time performance of the threshold conditions.

[0077] The adaptive early warning instruction set encapsulates optimized early warning strategies and spatiotemporal threshold conditions into instruction objects. These objects include attributes such as strategy type, parameter configuration, trigger conditions, and execution time. Instruction execution employs a priority queue management system: emergency stop instructions have priority 1, audible and visual alarm instructions have priority 2, voice warning instructions have priority 3, and information prompt instructions have priority 4. Instruction conflict resolution is achieved through priority comparison, with higher-priority instructions overriding lower-priority instructions. Instruction execution status tracking includes four states: pending execution, executing, completed, and canceled, with state transitions controlled by a finite state machine. The instruction set is persistently stored in an SQLite database, with table structures containing fields such as instruction identifier, creation time, execution time, status, and parameters. Version management of the instruction set supports rollback operations; version numbers are in timestamp format, and the rollback window is set to the 100 most recent versions.

[0078] In one optional implementation, the adaptive early warning instruction set guides the audible and visual early warning device and the work suspension control device to perform tiered interventions, forming closed-loop feedback data. The closed-loop feedback data is then used to optimize the nonlinear mapping function, promoting the dynamic evolution of risk identification and early warning decisions, including: A deep neural network model is constructed to determine the warning type and response time limit in the adaptive warning instruction set. The execution level of the graded intervention is deduced through the deep neural network model. The warning control signal of the audible and visual warning device and the suspension control signal of the work suspension control device are generated according to the execution level. The audible and visual warning device and the work suspension control device are activated to perform graded intervention. Construct response time feature vectors after the sound and light early warning device and the work stoppage control device perform graded intervention, and fuse the position change trajectory of the workers and the change trend of the risk heat map after the graded intervention is performed to generate closed-loop feedback data; The closed-loop feedback data is analyzed to establish a dynamic mapping matrix between intervention execution parameters and risk mitigation effects. Based on the deviation distribution law of the dynamic mapping matrix, the mapping coefficient of the nonlinear mapping function is optimized through the pre-intervention risk level in the closed-loop feedback data. The optimization parameters of the nonlinear mapping function are embedded into the risk identification and early warning decision-making process. Based on the continuously accumulated closed-loop feedback data, a periodic optimization mechanism for the nonlinear mapping function is constructed to realize the dynamic evolution of risk identification and early warning decision-making.

[0079] The deep neural network model for adaptive early warning instruction set, which correlates early warning types with response time limits, employs a multilayer perceptron architecture. The network consists of an input layer, three hidden layers, and an output layer. The input layer receives an 8-dimensional feature vector, including the early warning type encoding, risk intensity, gradient magnitude, environmental correction coefficient, historical response time, execution latency, success rate, and failure rate. Early warning type encoding uses a one-hot encoding method, with five early warning types corresponding to five-dimensional binary vectors. The hidden layer contains 128, 64, and 32 neurons respectively, using the ReLU activation function and a dropout probability of 0.3 to prevent overfitting. The output layer contains four neurons, corresponding to four execution levels, and uses the softmax activation function to output the probability distribution. Network training employs the cross-entropy loss function, the Adam algorithm as the optimizer, a learning rate of 0.001, a batch size of 64, and 1000 training epochs.

[0080] The execution level determination is achieved through threshold judgment of network output probabilities. Execution levels are divided into four levels: Observation, Alert, Warning, and Emergency. The Observation level corresponds to an output probability range of 0.0 to 0.25, recording only the risk status without intervention. The Alert level corresponds to a probability range of 0.25 to 0.5, providing low-intensity audio-visual alerts. The Warning level corresponds to a probability range of 0.5 to 0.8, providing medium-intensity audio-visual warnings. The Emergency level corresponds to a probability range of 0.8 to 1.0, providing high-intensity audio-visual alarms and triggering a work stoppage. Execution level switching employs a lag mechanism: level upgrades require a probability exceeding the threshold for three consecutive frames, and level downgrades require a probability below the threshold for five consecutive frames, avoiding frequent switching that could lead to erroneous operations.

[0081] The warning control signals of the audible and visual warning device are generated based on a parameter mapping of execution levels. The control signals include parameters such as volume, frequency, duration, flashing mode, and color coding. The alert level signal sets the volume to 70 dB, the frequency to 500 Hz, and the duration to 2 seconds, with a constantly lit green LED. The warning level signal sets the volume to 85 dB, the frequency to 800 Hz, and the duration to 5 seconds, with a yellow LED flashing at 1 Hz. The emergency level signal sets the volume to 100 dB, the frequency to 1000 Hz, and the duration to 10 seconds, with a red LED flashing rapidly at 3 Hz. The control signals are transmitted using PWM modulation, with a PWM frequency of 1000 Hz and a duty cycle ranging from 0% to 100% corresponding to the volume intensity. Signal transmission is achieved via an RS485 bus with a baud rate of 115200 bps and a data frame format of 8 data bits, 1 stop bit, and no parity bit.

[0082] The pause control signal of the work pause control device is triggered only at the emergency execution level. The control signal includes fields such as pause command, device identifier, execution delay, and confirmation requirement. The pause command uses 16-bit binary encoding, with the high 8 bits representing the pause type and the low 8 bits representing the pause level. The device identifier uses MAC address format and supports simultaneous control of up to 256 devices. The execution delay is adjustable from 500 milliseconds to 5 seconds, with a default value of 1 second. The pause command can be canceled manually within the delay time. When the confirmation requirement flag is set to 1, manual confirmation from the operator is required; when set to 0, the pause is executed automatically. The control signal is transmitted via the CAN bus, with the data frame identifier using an 11-bit standard format, a data length of 8 bytes, and a transmission baud rate of 500kbps.

[0083] Activation of the audible and visual warning devices and the work stoppage control devices is achieved through the device driver module. This module maintains a device status table, recording information such as the online status, response latency, fault count, and maintenance time for each device. Device heartbeat detection occurs at 10-second intervals; devices that do not receive a heartbeat signal for more than 30 seconds are marked as offline. Device activation uses a broadcast wake-up mechanism; the wake-up signal includes a device type filter field, and only devices matching the type respond to the wake-up signal. The activation confirmation timeout is set to 3 seconds; devices that do not confirm within this timeout are marked as faulty and a fault log is recorded. Device concurrency control is implemented through a token mechanism, limiting the number of devices that can be activated simultaneously to 32; activation requests exceeding this limit are queued.

[0084] The response time feature vector after tiered intervention execution comprises eight dimensions: instruction issuance time, equipment response time, audible and visual output delay, work pause delay, personnel reaction time, risk mitigation time, system recovery time, and anomaly handling time. Instruction issuance time is the time difference between instruction generation and equipment reception, with millisecond precision. Equipment response time is the time difference between receiving the instruction and commencing execution, typically ranging from 10 to 100 milliseconds. Audible and visual output delay is the time difference between the start of execution and the output of the audible and visual signal; hardware latency is usually within 5 milliseconds. Work pause delay is the time difference between the pause instruction and the actual stop of the equipment; delays due to mechanical inertia range from 100 to 2000 milliseconds. Personnel reaction time is the time difference between the warning signal and the start of personnel response, statistically ranging from 0.5 to 5 seconds.

[0085] The trajectory fusion of worker position changes employs trajectory difference analysis, comparing the changes in worker keypoint coordinates within 30 seconds before and after intervention. Trajectory features include parameters such as movement distance, movement speed, movement direction, dwell time, and trajectory smoothness. Movement distance is calculated as the sum of Euclidean distances between trajectory points, maintaining accuracy to the centimeter level. Movement speed is calculated using a sliding window method, with a window length of 10 frames corresponding to approximately 0.33 seconds. Movement direction is extracted using principal component analysis, with the first principal component corresponding to the primary movement direction. Dwell time is calculated as the cumulative time spent when the movement speed is below 0.1 meters per second. Trajectory smoothness is quantified using the variance of the second derivative; a smaller variance indicates a smoother trajectory.

[0086] Risk heatmap trend analysis is achieved through time-series comparison, comparing the evolution characteristics of the risk heatmap within 60 seconds before and after intervention. Trend characteristics include indicators such as changes in risk peak value, risk area area, risk gradient, and risk propagation speed. Changes in risk peak value are calculated as the time-series difference of the maximum risk intensity value; positive values ​​indicate increasing risk, and negative values ​​indicate decreasing risk. Changes in risk area area are calculated as the percentage change in the number of high-risk grid cells, with a threshold set at a risk intensity greater than 2.0. Risk gradient changes are quantified by the time-series variance of the gradient amplitude; increased variance indicates unstable risk distribution. Risk propagation speed is calculated by the rate of change of the centroid position of the risk area, with the unit being meters per second.

[0087] The closed-loop feedback data generation merges the response time feature vector, location change trajectory, and risk heatmap trend into a comprehensive feature matrix with a dimension of 20×T, where T is the time series length. The feature matrix is ​​segmented using a time window, with a window length of 120 seconds corresponding to 60 seconds before and after the intervention, and a 50% window overlap to ensure continuity. Data annotation uses an expert scoring method, with intervention effectiveness scores ranging from 0 to 10, where 0 indicates no effect and 10 indicates complete effectiveness. Data quality control is achieved through outlier detection and missing value handling. Outliers are identified using the 3σ criterion, and missing values ​​are filled using linear interpolation or nearest neighbor values.

[0088] The dynamic mapping matrix was established using a multiple linear regression method. The matrix input was a vector of intervention execution parameters, and the output was a scalar value of the risk mitigation effect. The execution parameters included 12 dimensions, such as intervention type, execution intensity, duration, response delay, number of devices, and environmental conditions. The risk mitigation effect was quantified as the decrease in risk intensity before and after the intervention; a decrease greater than 50% was considered an effective intervention, and a decrease less than 10% was considered an ineffective intervention. The regression model used the ridge regression algorithm, with a regularization parameter set to 0.01. The model was trained using 80% of the data, and validated using 20%. Model evaluation metrics included mean squared error, coefficient of determination, and mean absolute error; a coefficient of determination greater than 0.8 indicated a good model fit.

[0089] Deviation distribution pattern analysis is achieved through residual statistics, where residuals are calculated as the actual risk mitigation effect minus the model prediction effect. The residual distribution is tested for normality; a p-value greater than 0.05 in the Shapiro-Wilk test indicates that the residuals follow a normal distribution. Deviation pattern identification is achieved through cluster analysis. K-means clustering divides the residuals into three patterns, corresponding to overestimation, accuracy, and underestimation. The mean residual for the overestimation pattern is less than -0.5, the mean residual for the underestimation pattern is greater than +0.5, and the mean residual for the accuracy pattern is between -0.5 and +0.5. Deviation correction uses a weighted average method, with correction weights inversely proportional to the absolute value of the residuals to ensure that abnormal deviations have a minimal impact on the correction results.

[0090] The optimization of the nonlinear mapping function's mapping coefficients is achieved using the gradient descent algorithm. The optimization objective is to minimize the mean squared error between the predicted and actual risk levels. The mapping coefficients include network weights and bias parameters, with a total of approximately 15,000 parameters. Gradient calculation employs the backpropagation algorithm, and an adaptive adjustment strategy is used for the learning rate. The initial learning rate is 0.01, and it decays by 10% every 100 training epochs. The optimization process uses mini-batch gradient descent with a batch size of 128, a momentum parameter of 0.9, and a weight decay parameter of 0.0001. The convergence condition is that the loss change is less than 0.001 for 50 consecutive training epochs, or that the maximum number of training epochs (2000) is reached.

[0091] The optimized parameter embedding risk identification and early warning decision-making process is achieved through parameter replacement. New parameters overwrite existing parameters and update the model weight file. Parameter updates utilize a hot update mechanism, eliminating the need to restart the decision service and maintaining service availability during the update process. Parameter version management uses version numbers, formatted as major version number . minor version number . revision number, with the most recent 50 versions stored in the version history. A parameter rollback mechanism supports rapid restoration to the previous stable version, triggered when the new version's performance metrics drop by more than 5% or a serious error occurs. Parameter effectiveness verification is achieved through A / B testing, applying the new parameters to 50% of early warning requests and comparing the performance differences between the old and new parameters.

[0092] The periodic optimization mechanism is constructed based on a sliding time window, with an optimization cycle of 7 days. Parameter optimization is automatically triggered at the end of each cycle. Optimization data uses closed-loop feedback data from the most recent 30 days, with approximately 10,000 to 50,000 sample points. Data preprocessing includes noise filtering, outlier removal, and feature standardization. Data with a quality score greater than 85% after preprocessing is used for model optimization. Optimization triggering conditions include data volume thresholds, performance degradation exceeding thresholds, and manual triggering. The data volume threshold is set at more than 1,000 new samples added in a single cycle, and the performance degradation threshold is a decrease in prediction accuracy exceeding 3%. Optimization process monitoring includes dimensions such as optimization progress, performance metrics, and resource consumption. Optimization automatically stops and sends alarm notifications when an anomaly occurs.

[0093] Dynamic evolution is supported by a continuous learning framework, which includes modules for data management, model training, performance evaluation, and parameter deployment. The data management module is responsible for tasks such as collecting, storing, cleaning, and labeling closed-loop feedback data. Data storage uses a time-series database, supporting high-concurrency writes and fast queries. The model training module supports both online and offline training modes. Online learning uses an incremental update algorithm, while offline training uses a batch training algorithm. The performance evaluation module monitors model performance metrics in real time, including accuracy, recall, F1 score, and response time, with performance threshold alerts ensuring model quality. The parameter deployment module is responsible for optimizing parameter version management, canary releases, and rollback recovery, maintaining a deployment success rate of over 99%.

[0094] In one optional implementation, the execution level of the graded intervention is deduced through the deep neural network model, and the warning control signal of the audible and visual warning device and the pause control signal of the work stoppage control device are generated according to the execution level, including: Multi-source image sequences from the work site are collected to construct a time-series data stream. Spatiotemporal feature vectors of the time-series data stream are extracted using a deep neural network model. The spatiotemporal feature vectors are divided into multiple state intervals according to the time span. Based on the state intervals, the temporal correlation coefficients between consecutive frames are extracted. The temporal correlation coefficients are then fused to construct a prediction benchmark for the execution level of graded intervention. Based on the execution level prediction benchmark, a parameter matrix for audible and visual early warning and pause control is constructed. Multiple sets of candidate control parameter combinations are sampled in the parameter matrix using the Monte Carlo method. A comprehensive score of execution cost and early warning effect is calculated for each set of candidate control parameter combinations. The parameter combination with the best comprehensive score is selected to generate early warning control signal and pause control signal.

[0095] Multiple high-definition cameras are deployed at the work site to capture image sequences from different angles and positions. These cameras can include fixed and mobile cameras. Fixed cameras capture panoramic views of the work area at 25 frames per second, while mobile cameras capture detailed movements of specific workers at 30 frames per second. The image resolution is set to 1920×1080 pixels to ensure sufficient detail in the captured image data. All captured image sequences are synchronized using timestamps to create a unified time-series data stream.

[0096] After image acquisition, the temporal data stream is input into a pre-trained deep neural network model for processing. This network model employs a two-stream structure, including a spatial feature branch and a temporal feature branch. The spatial feature branch uses an improved ResNet50 network, taking a single frame image as input and extracting a 256-dimensional spatial feature vector. The temporal feature branch uses a 3D convolutional network, taking 16 consecutive frames of images as input and extracting a 256-dimensional temporal feature vector. The feature vectors from the two branches are fused through an attention mechanism to generate a 512-dimensional spatiotemporal feature vector. In practical applications, the model processes each 3-second video segment once to generate the corresponding spatiotemporal feature vector.

[0097] After obtaining the spatiotemporal feature vectors, the continuous feature vectors are divided into multiple state intervals according to the time span. A sliding window method is used, with a window size of 5 seconds and a step size of 1 second. Cluster analysis is performed on the feature vectors within each window, and feature vectors with a similarity greater than 0.85 are grouped into the same state interval. In a practical application, a 30-minute task process is divided into approximately 25 different state intervals, each interval representing a typical state in the task process.

[0098] Based on the divided state intervals, temporal correlation coefficients are extracted between consecutive frames. These coefficients are obtained by calculating the similarity between adjacent feature vectors using the cosine similarity method. In practical applications, the cosine similarity between adjacent feature vectors is typically between 0.6 and 0.95. The system identifies transition points with similarity below 0.7 as state change points and assigns them higher attention weights. The temporal correlation coefficient matrix has a dimension of N×N, where N is the number of feature vectors. By analyzing patterns in the temporal correlation coefficient matrix, the system can identify potential security risk patterns.

[0099] After the temporal correlation coefficient matrix is ​​constructed, it is input into a pre-trained multilayer perceptron network containing three hidden layers with 128, 64, and 32 neurons respectively. The network outputs the predicted probability distribution of the execution level of graded interventions through a Softmax function. The execution levels are divided into four levels: normal (level 0), slight risk (level 1), moderate risk (level 2), and severe risk (level 3). In a practical application, the system detects that a worker is not wearing a safety rope while working at height and rates this state as level 2; it also detects that a worker has stayed in a dangerous area for too long and rates this as level 3.

[0100] Based on the predicted execution level, a parameter matrix for audible and visual warnings and pause control is constructed. This matrix includes multiple dimensions: warning sound frequency (200-2000Hz), warning sound intensity (60-110 dB), warning light color (red, yellow, green), warning light flashing frequency (0.5-5Hz), and pause control delay time (0-30 seconds). Different parameter ranges are preset for different execution levels. For example, for level 2 risk, the warning sound frequency range is 800-1200Hz, the warning sound intensity is 85-95 dB, the warning light color is yellow, the flashing frequency is 2-3Hz, and the pause control delay time is 10-15 seconds.

[0101] After the parameter matrix is ​​constructed, 500 candidate control parameter combinations are randomly sampled in the parameter space using the Monte Carlo method. For each parameter combination, the system calculates a comprehensive score for its execution cost and early warning effect. Execution cost includes factors such as energy consumption and the degree of interference with normal operations; early warning effect includes factors such as perceptibility and recognition clarity. The comprehensive score is obtained by weighted summation, with weight coefficients preset according to the actual operating environment and safety requirements. In a practical application case, the optimal parameter combination selected by the system is: early warning sound frequency 1000Hz, early warning sound intensity 90dB, early warning light color yellow, flashing frequency 2.5Hz, and pause control delay time 12 seconds.

[0102] Based on the optimal parameter combination, early warning control signals and pause control signals are generated and transmitted to the audible and visual early warning device and the work pause control device via the industrial control bus, enabling timely intervention in potential safety risks. Test results show that this method can reduce the incidence of high-risk work accidents by approximately 87%, while keeping interference with normal operations within an acceptable range.

[0103] A second aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0104] A third aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0105] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for real-time identification of safety risks in elevator shaft operations based on visual AI, characterized in that, include: Real-time image sequences of the work space are acquired by an image acquisition device deployed in the elevator shaft. Based on the real-time image sequences, the three-dimensional skeleton topology, component-level integrity features, and obstacle distribution features of the workers are deduced. The spatial mapping relationship between the three-dimensional skeleton topology, the component-level integrity features, and the obstacle distribution features is derived, and a multimodal semantic representation vector under a unified coordinate system is constructed. Based on the multimodal semantic representation vector, a minimum safe distance field between the worker and the hazardous boundary is derived. The worker's historical movement trajectory is then fused to infer the probability distribution of their future location. A risk heatmap is constructed based on the overlap between the two, including: The coordinate components of the three-dimensional skeleton topology and obstacle distribution features in the multimodal semantic representation vector are deduced. Based on the coordinate components, the Euclidean distance between the spatial position of the operator's key point and the spatial position of the danger boundary is calculated. The minimum value of the Euclidean distance is derived to construct the minimum safe distance. The association rules between the spatial position and the minimum safe distance are established using the grid cells of the shaft spatial coordinate system to generate the minimum safe distance field. A temporal evolution model is constructed for the coordinate components in the multimodal semantic representation vector. Based on the temporal evolution model, the historical position sequence of the operator's joints is derived to construct the historical motion trajectory. The displacement vector and velocity vector at adjacent moments are calculated based on the historical motion trajectory. The probability distribution of the operator's future position is inferred through the distribution characteristics of the displacement vector and the velocity vector. Based on the probability distribution of the future position, a spatial mapping is constructed in the grid cells of the minimum safe distance field. The overlapping area of ​​the two in the spatial grid is derived, and the grid cells with insufficient safe distance in the overlapping area are used to construct the risk area. For the grid cells in the risk area, the risk intensity is constructed by combining the corresponding minimum safe distance and the probability of occurrence of the location, and the risk intensity is converted into color coding to generate a risk heat map. Based on the risk heatmap, a nonlinear mapping function between risk level and early warning strategy is constructed. The local gradient features of the risk heatmap are coupled with real-time environmental parameters to derive the spatiotemporal threshold for early warning triggering and generate an adaptive early warning instruction set. The adaptive early warning instruction set guides the audible and visual early warning device and the work suspension control device to perform graded interventions, forming closed-loop feedback data. The closed-loop feedback data is then used to optimize the nonlinear mapping function, promoting the dynamic evolution of risk identification and early warning decisions.

2. The method according to claim 1, characterized in that, Based on the real-time image sequence, the three-dimensional skeleton topology, component-level integrity features, and obstacle distribution features of the workers can be deduced. Deriving the spatial mapping relationship between the three-dimensional skeleton topology, the component-level integrity features, and the obstacle distribution features, and constructing a multimodal semantic representation vector in a unified coordinate system includes: A multi-frame joint detection model is constructed for the real-time image sequence. The evolution law of the two-dimensional joints of the operator is derived through dynamic temporal features. The three-dimensional spatial activity features of the joints are constructed by fusing the joint movement trajectory between adjacent frames and the depth constraint conditions of the shaft space. The three-dimensional skeleton topology of the operator is deduced based on the three-dimensional spatial activity features. Based on the distribution of joints in the three-dimensional skeleton topology, the dynamic boundary of the protective equipment area is derived. The standard configuration template within the dynamic boundary is dynamically mapped to the actual test component. Based on the dynamic topology mapping, the temporal evolution characteristics of the component integrity are deduced. Based on the temporal evolution characteristics, a dynamic perception range of the well environment area is constructed. Scene semantic analysis is performed on the dynamic perception range. The spatial distribution dynamic characteristics of obstacles are constructed by combining the scene semantic analysis with the spatiotemporal evolution law of obstacle pixel areas and the parameter matrix of the image acquisition device. Based on the three-dimensional spatial activity features, the temporal evolution features, and the spatial distribution dynamic features, a mapping function for the spatial coordinates of the joint points is constructed. The dynamic coupling relationship between the three is derived through the mapping function. According to the dynamic coupling relationship, the spatial coordinates of the joint points, the obstacle boundary features, and the coordinate mapping parameters are dynamically combined to construct a multimodal semantic representation vector with spatiotemporal correlation.

3. The method according to claim 2, characterized in that, Based on the distribution of joints in the three-dimensional skeleton topology, the dynamic boundary of the protective equipment area is derived. A dynamic topological mapping is then performed between the standard configuration template within the dynamic boundary and the actual tested components. Based on this dynamic topological mapping, the temporal evolution characteristics of component integrity are deduced, including: The set of joint points corresponding to the wearing position of the protective equipment in the three-dimensional skeleton topology is derived. A spatial envelope range is constructed based on the set of joint points. The dynamic boundary of the protective equipment area is generated according to the spatial envelope range and the standard coverage size of the protective equipment. Based on the dynamic boundary, candidate detection regions for protective equipment are constructed in the real-time image sequence. The candidate detection regions are semantically parsed to deduce the contour features of the actual detection components. The contour features are then mapped to the spatial range constructed by the dynamic boundary. Based on the spatial range, a spatial mapping relationship is established between the component nodes in the standard configuration template and the contour features. The matching tolerance of the spatial mapping relationship is adaptively adjusted according to the deformation of the dynamic boundary. By calculating the number of missing nodes and the node offset distance between the standard configuration template and the actual detected component, the component integrity feature value of the current frame is constructed. The temporal evolution feature of component integrity is derived based on the changing trend of the component integrity feature value.

4. The method according to claim 1, characterized in that, Based on the risk heatmap, a nonlinear mapping function between risk levels and early warning strategies is constructed. The local gradient features of the risk heatmap are coupled with real-time environmental parameters to derive the spatiotemporal threshold for early warning triggering and generate an adaptive early warning instruction set, including: A standardized feature space is constructed by deriving the risk intensity value of each grid cell in the risk heat map. Based on the standardized feature space, a multi-level risk level interval is constructed, and the early warning strategy type corresponding to each risk level interval is derived. Based on the historical correlation characteristics between different risk level intervals and corresponding early warning strategy types in the risk heatmap, the nonlinear evolution law of risk level change rate and early warning strategy response time is deduced, and a nonlinear mapping function between risk level and early warning strategy is constructed. The gradient vector features of the risk heatmap in the spatial dimension are derived. Based on the magnitude and direction of the gradient vector, local gradient features are constructed. The environmental correction coefficient is deduced by combining the real-time environmental parameters of the well space. The output early warning strategy of the nonlinear mapping function is optimized by using the environmental correction coefficient. Based on the magnitude and direction of the local gradient features and the environmental correction coefficient, the spatiotemporal threshold conditions for triggering the early warning are derived. An adaptive early warning instruction set is constructed by combining the early warning strategy optimized by the nonlinear mapping function with the spatiotemporal threshold conditions.

5. The method according to claim 1, characterized in that, The adaptive early warning instruction set guides the audible and visual early warning device and the work suspension control device to perform tiered interventions, forming closed-loop feedback data. This closed-loop feedback data is then used to optimize the nonlinear mapping function, promoting the dynamic evolution of risk identification and early warning decisions, including: A deep neural network model is constructed to determine the warning type and response time limit in the adaptive warning instruction set. The execution level of the graded intervention is deduced through the deep neural network model. The warning control signal of the audible and visual warning device and the suspension control signal of the work suspension control device are generated according to the execution level. The audible and visual warning device and the work suspension control device are activated to perform graded intervention. Construct response time feature vectors after the sound and light early warning device and the work stoppage control device perform graded intervention, and fuse the position change trajectory of the workers and the change trend of the risk heat map after the graded intervention is performed to generate closed-loop feedback data; The closed-loop feedback data is analyzed to establish a dynamic mapping matrix between intervention execution parameters and risk mitigation effects. Based on the deviation distribution law of the dynamic mapping matrix, the mapping coefficient of the nonlinear mapping function is optimized through the pre-intervention risk level in the closed-loop feedback data. The optimization parameters of the nonlinear mapping function are embedded into the risk identification and early warning decision-making process. Based on the continuously accumulated closed-loop feedback data, a periodic optimization mechanism for the nonlinear mapping function is constructed to realize the dynamic evolution of risk identification and early warning decision-making.

6. The method according to claim 5, characterized in that, The deep neural network model deduces the execution levels of graded intervention, and generates warning control signals for the audible and visual warning device and pause control signals for the work stoppage control device based on the execution levels, including: Multi-source image sequences from the work site are collected to construct a time-series data stream. Spatiotemporal feature vectors of the time-series data stream are extracted using a deep neural network model. The spatiotemporal feature vectors are divided into multiple state intervals according to the time span. Based on the state intervals, the temporal correlation coefficients between consecutive frames are extracted. The temporal correlation coefficients are then fused to construct a prediction benchmark for the execution level of graded intervention. Based on the execution level prediction benchmark, a parameter matrix for audible and visual early warning and pause control is constructed. Multiple sets of candidate control parameter combinations are sampled in the parameter matrix using the Monte Carlo method. A comprehensive score of execution cost and early warning effect is calculated for each set of candidate control parameter combinations. The parameter combination with the best comprehensive score is selected to generate early warning control signal and pause control signal.

7. A real-time safety risk identification system for elevator shaft operations based on visual AI, used to implement the method of any one of claims 1-6, characterized in that, include: The first unit is used to acquire real-time image sequences of the work space through an image acquisition device deployed in the elevator shaft, and to deduce the three-dimensional skeleton topology, component-level integrity features and obstacle distribution features of the workers based on the real-time image sequences. It then derives the spatial mapping relationship between the three-dimensional skeleton topology, the component-level integrity features and the obstacle distribution features, and constructs a multimodal semantic representation vector under a unified coordinate system. The second unit is used to derive the minimum safe distance field between the worker and the dangerous boundary based on the multimodal semantic representation vector, to infer the probability distribution of the worker's future position by fusing the worker's historical movement trajectory, and to construct a risk heat map based on the overlap between the two. The third unit is used to construct a nonlinear mapping function between risk level and early warning strategy based on the risk heat map, couple the local gradient features of the risk heat map with real-time environmental parameters, derive the spatiotemporal threshold for early warning triggering, and generate an adaptive early warning instruction set. The fourth unit is used to guide the audible and visual early warning device and the work suspension control device to perform graded intervention according to the adaptive early warning instruction set, form closed-loop feedback data, and use the closed-loop feedback data to optimize the nonlinear mapping function, thereby promoting the dynamic evolution of risk identification and early warning decision-making.

8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Wind power construction intelligent safety management method and system based on intelligent AI monitoring

    CN120726559A

  • Building construction safety risk real-time early warning method and system

    CN120875573A