Methods and Systems for Safety Belt Inspection and Spatial Positioning of Construction Workers
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-08-14
AI Technical Summary
[0006]本发明的目的在于提供一种建筑工人安全带巡检与空间定位方法及系统,能够解决现有技术中极端工况下模型鲁棒性差和单目空间投影失真的问题
[0138]1、本发明能降低极端工况下的漏报率:通过ST-HMD模型的时空双触发与多模态补救机制(信息熵动态门控+协方差迹ROI扩增),在目标工地高空且存在脚手架遮挡、姿态扭曲的复杂场景下,目标漏报率相比传统轻量化检测模型有效降低,特别是在重度遮挡场景中,通过sVLM辅网络的上下文巡回推理显著提升了安全带状态的检出率。
Smart Images

Figure CN122345390B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of engineering safety management and computer vision, and in particular to a method and system for inspecting and spatially locating safety belts for construction workers. Background Technology
[0002] Construction work, power infrastructure projects, and other high-altitude operations (such as working near high-rise building edges, climbing tower cranes and scaffolding) have always been a top priority for safety management due to their complex working environments and extremely high risk of falls from heights. In these scenarios, the proper wearing of safety belts by workers is the last line of defense for protecting lives. Traditional safety inspections mainly rely on on-site safety officers' manual patrols or fixed surveillance cameras capturing images at specific points. This method is not only inefficient and lacks comprehensive data coverage, but also prone to blind spots due to the large and complex nature of construction spaces.
[0003] In recent years, multi-rotor drones have been gradually introduced into safety inspections due to their maneuverability. However, in complex high-altitude environments with towering steel pipes, intersecting components, and the need to maintain safe flight distances, traditional drone inspection solutions equipped with conventional vision algorithms still suffer from extremely high false negative rates and the risk of management gaps where "violations are known but not where they are located."
[0004] The core challenges of existing safety belt inspection and spatial positioning technologies for construction workers are twofold: First, existing lightweight deep learning detection models (such as pure convolutional neural networks) heavily rely on complete visible pixel features. Once workers are in a bent-over or twisted posture, or are largely obscured by scaffolding, visual features are broken, making the model prone to "confident false alarms" or complete missed detections. This fails to meet the high reliability requirements of detection under complex unstructured working conditions, and the model exhibits poor robustness under extreme conditions. Second, traditional UAV monocular vision spatial positioning technologies are generally based on the strong assumption that "the construction site is an absolutely ideal horizontal plane," lacking true depth priors. When violators are located on complex three-dimensional terrain with abrupt elevation changes, such as slopes, steps, or deep pits, simple two-dimensional pixel ray projection will produce huge nonlinear distortions, resulting in errors of several meters in the calculated geographic coordinates and distortion of monocular spatial projection.
[0005] Therefore, there is a need to provide a method and system for inspecting and spatially positioning safety belts for construction workers, which can solve the problems of poor model robustness and monocular spatial projection distortion under extreme working conditions in the existing technology. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for inspecting and spatially positioning safety belts for construction workers, which can solve the problems of poor model robustness and monocular spatial projection distortion under extreme working conditions in the prior art.
[0007] This invention is implemented as follows:
[0008] A safety belt inspection and spatial positioning system for construction workers includes a perception layer, an edge computing layer, a platform layer, and an application layer.
[0009] The perception layer includes an unmanned flight platform equipped with a telephoto camera, an RTK positioning module, an IMU attitude sensor, and an obstacle avoidance sensor.
[0010] The edge computing layer is equipped with edge computing devices that support tensor operations and deploys a spatiotemporally triggered multimodal hybrid detection model to process video streams in real time and output personnel positions and seat belt status. The edge computing devices are equipped with GPUs / NPUs. The multimodal hybrid detection model includes two deeply fused AI engines: one is a lightweight CNN main network that incorporates SPD-Conv downsampling and VSS modules, responsible for high frame rate conventional feature extraction; the other is a small visual language model auxiliary network.
[0011] The platform layer includes ground servers that provide path planning services, 3D Gaussian splash 3D modeling services, spatial positioning and calculation services, and a geofencing database. The 3D Gaussian splash 3D modeling service uses pre-flight imagery to render the real scene and extracts a joint constraint matrix composed of absolute ground elevation and local surface normal vectors. The path planning service runs a multi-objective optimization algorithm to generate the optimal flight path that balances focal length-tilt coupling and observation angle penalty. The spatial positioning and calculation service uses pixel coordinates extracted from the edge, combined with the UAV attitude and the joint constraint matrix, to calculate the true ENU coordinates of violators through a micro-plane penetration algorithm. The geofencing database, built on PostGIS, stores compliant high-risk area polygons and uses the GIST index to achieve millisecond-level ST_Contains spatial boundary judgment.
[0012] The application layer includes a 3D visualization platform and a mobile terminal. The 3D visualization platform is responsible for loading 3D reality models, marking the 3D location of the violation targets, and displaying high-definition screenshots of the violations. The mobile terminal is responsible for receiving hierarchical alarm information in real time, guiding safety officers to handle the situation based on precise floor / direction coordinates, and completing the digital closed loop of the entire safety supervision business by uploading rectification feedback.
[0013] A method for safety belt inspection and spatial positioning using a construction worker safety belt inspection and spatial positioning system includes the following steps:
[0014] Step 1: Preliminary environmental geometric feature extraction and viewpoint-aware route planning;
[0015] Step 2: Autonomous aerial inspection and synchronous acquisition and encapsulation of multi-source spatiotemporal data;
[0016] Step 3: Spatiotemporal multimodal detection at the edge and remediation of missed detections due to severe occlusion;
[0017] Step 4: Monocular localization calculation of local surface normal vectors on the platform side and hierarchical control closed loop.
[0018] Step 1 includes the following sub-steps:
[0019] Step 11: Construct a 3D reality model based on the 3D Gaussian splash 3D modeling service at the platform layer, and extract the joint constraint matrix;
[0020] Step 12: Construct a standardized geofencing database;
[0021] Step 13: Focal length-tilt angle coupled route planning based on view perception.
[0022] Step 11 includes the following sub-steps:
[0023] Step 111: The unmanned aerial platform of the perception layer performs a "pre-flight oblique photography" mission over the target construction site to collect a multi-view image sequence covering the entire work surface;
[0024] Step 112: After image acquisition is complete, the images are transmitted back to the platform layer. The platform layer uses 3D Gaussian splashing technology to render a 3D real-world model;
[0025] Step 113: Set up an orthogonal overhead virtual camera;
[0026] A virtual camera is positioned above the 3D reality model, with its optical axis pointing vertically downwards. The camera is located at a safe height above the highest point of the 3D reality model; the absolute elevation of this virtual camera is denoted as... ;
[0027] Step 114: Calculate the depth map for volume rendering;
[0028] For each pixel on the image plane captured by the virtual camera The surface depth value corresponding to this pixel is calculated using the volume rendering formula. The volume rendering formula is:
[0029]
[0030] in, The total number of Gaussian spheres traversed along the ray. For the first The opacity of a Gaussian sphere For the first The distance from the center of the Gaussian sphere to the optical center of the virtual camera;
[0031] The volume rendering formula obtains the depth at which the ray first hits the scene surface by accumulating opacity from front to back;
[0032] Step 115: Generate an absolute elevation map of the Earth's surface;
[0033] Combined with the absolute elevation of the virtual camera and depth map Calculate the absolute elevation of the ground surface corresponding to each pixel:
[0034]
[0035] In the formula, The geographic plane coordinates corresponding to this pixel;
[0036] This yields an absolute elevation map of the entire target construction site.
[0037] Step 116: Extract local surface normal vectors;
[0038] For each point on the absolute elevation map of the Earth's surface, using the three-dimensional coordinates of its neighboring points, a local plane is fitted using the least squares method, and then the unit normal vector of this local plane, i.e., the local surface normal vector, is solved. ;
[0039] Step 117: Construct the joint constraint matrix: Organize the calculated absolute surface elevation and local surface normal vectors into a spatial lookup table according to a spatial grid, denoted as the joint constraint matrix. :
[0040]
[0041] The joint constraint matrix is persistently stored in the platform layer database, and a spatial index is created.
[0042] Step 12 includes the following sub-steps:
[0043] Step 121: Safety management personnel delineate the outer boundary of the high-risk area on the 3D reality model according to relevant safety regulations;
[0044] The outer boundary of a high-risk area includes edges, openings, suspended work platforms, and climbing structures; among them, edges include floor slabs and roof edges, openings include elevator shafts and reserved openings, suspended work platforms include scaffolding and suspended baskets, and climbing structures include ladders and vertical passages.
[0045] Step 122: Convert the outer boundary of the extracted high-risk area into WGS84 geographic coordinates and add the following attributes: area name, risk level, and corresponding safety regulation clauses.
[0046] Step 123: Store the outer boundary data of the above-mentioned high-risk areas into the PostGIS spatial database, and create a GiST index for the geographic geometry field;
[0047] Step 13 includes the following sub-steps:
[0048] Step 131: Establish the following comprehensive optimization objective function:
[0049]
[0050] in, Total path length, which is the sum of Euclidean distances between adjacent waypoints;
[0051] The number of sharp turns is defined as the number of adjacent segments where the rate of change of heading angle exceeds a predetermined threshold.
[0052] : Zoom switching cost; Since the optical zoom camera on the drone requires time and energy to switch between different focal lengths, a cost function is introduced:
[0053]
[0054] in, , The focal length of adjacent viewpoints. This refers to the switching time required for the zoom motor. , These are predetermined weighting coefficients; this cost function is used to encourage gradual changes in focal length between adjacent viewpoints.
[0055] The collision risk penalty is calculated using the following formula:
[0056]
[0057] in, For the first Distance from each waypoint to the nearest obstacle and This is a pre-set normal value; the collision risk penalty causes the flight path to tend to move away from obstacles.
[0058] The perspective penalty is defined as follows:
[0059]
[0060] in, Let be the unit direction vector of the camera's optical axis. This refers to the local surface normal vector of the working surface extracted in step 11;
[0061] All are weighting coefficients, satisfying .
[0062] Step 132: Solve the comprehensive optimization objective function to obtain a three-dimensional waypoint sequence that the UAV can fly. The generated three-dimensional waypoint sequence is sent to the UAV flight platform in the perception layer.
[0063] Step 2 includes the following sub-steps:
[0064] Step 21: After receiving the optimal inspection route calculated in step 13 from the platform layer, the drone in the perception layer automatically unlocks and takes off and flies autonomously along the inspection route. During the flight, the drone's flight control system collects multi-source data in real time at a preset high frequency.
[0065] Step 22: The UAV's onboard controller adds a unified timestamp to each frame of video image, each RTK positioning point, and each set of IMU attitude data;
[0066] Step 23: Encapsulate the acquired multi-source data into a standard data stream according to the principle of time alignment: each frame of image is equipped with its corresponding RTK coordinates and IMU pose;
[0067] Step 24: The encapsulated standard data stream is pushed to the edge computing layer in real time.
[0068] The multi-source data includes: RTK positioning data: used to provide centimeter-level absolute geographic coordinates for the UAV. IMU attitude data: Synchronously records the yaw angle, pitch angle, roll angle, and angular velocity of each axis of the UAV; Video stream: The telephoto camera continuously captures high-definition images at a fixed frame rate and automatically zooms according to the preset focal length value in the inspection route.
[0069] In step 2, the obstacle avoidance sensor data is accessed by the system.
[0070] Step 3 includes the following sub-steps:
[0071] Step 31: First, run a lightweight CNN main network on the edge computing device to perform pre-feature extraction;
[0072] Step 32: Dynamic gating correction of spatial dimension based on information entropy;
[0073] Step 33: Time-based remedies.
[0074] The lightweight CNN main network adopts the following design:
[0075] The input image is appropriately scaled and Mosaic data augmentation strategy is applied.
[0076] An SPD-Conv downsampling module is introduced into the backbone network to preserve fine-grained spatial information of small targets during downsampling;
[0077] The VSS module is introduced to achieve global feature dependency modeling with linear complexity, enhancing the awareness of long-distance context.
[0078] The neck area employs a bidirectional, cross-scale dense connection to efficiently integrate shallow details with deep semantics.
[0079] The detection head uses an anchor-free structure and directly outputs the center pixel coordinates of the 2D bounding box of the person target. Width and height And three types of probability distributions: ,in, These represent "personnel", "wearing seat belts", and "not wearing seat belts", respectively.
[0080] The lightweight CNN main network runs at a high frequency with the video frame rate as the step size, and the inference result of each frame is used as the initial output.
[0081] Step 32 includes the following sub-steps:
[0082] Step 321: For each detected target, calculate the information entropy of its probability distribution:
[0083]
[0084] The value of information entropy The lower the value, the more certain the classification decision of the lightweight CNN main network; the higher the information entropy value, the flatter the probability distribution of the lightweight CNN main network for each category.
[0085] Step 322: Generate dynamic adaptive weights for the lightweight CNN main network based on the information entropy value:
[0086]
[0087] in, This is the preset adjustment coefficient;
[0088] when hour, ;along with Increase It decreased exponentially;
[0089] Step 323: When When the value falls below a certain preset threshold, the system triggers a small visual language model auxiliary network to assist in reasoning.
[0090] The input to this small visual language model auxiliary network includes: image patches of the monitored target area cropped from the original image and structured text prompts. The content of the structured text prompts includes the preliminary category output by the lightweight CNN main network, the pose estimation of the monitored target, the occlusion ratio estimation, and the current target construction site scene description.
[0091] Step 324: After fine-tuning, the small visual language model auxiliary network performs reasoning based on common sense and contextual semantics, and outputs the corrected three-class probability distribution. and the corresponding confidence level;
[0092] Step 325: The seatbelt status is determined by a weighted fusion of the lightweight CNN main network and the small visual language model auxiliary network. The weighted fusion formula is as follows:
[0093]
[0094] Step 33 includes the following sub-steps:
[0095] Step 331: The system maintains the historical state vector of each tracked target. The historical state vector typically contains the position, velocity, size, and rate of change of the detected target on the image plane; state prediction is performed using Kalman filtering.
[0096] Predicted status: ,in, This is the state transition matrix;
[0097] Prediction error covariance matrix: ,in, The process noise covariance matrix; the trace of the covariance matrix Reflecting the uncertainty of the current prediction: the longer the target is lost, the greater the prediction variance. Also bigger;
[0098] Step 332: The edge computing layer calculates based on the trace of the covariance matrix. The cropping range of the region of interest is dynamically adjusted to reflect the uncertainty:
[0099]
[0100] in, To predict the width and height of the target bounding box, This is a preset magnification factor; that is, the higher the prediction uncertainty, the greater the expansion of the cropping frame.
[0101] Step 333: The system forcibly extracts the enlarged region of interest from the image and sends it to a small visual language model auxiliary network for "context retrieval reasoning". The small visual language model auxiliary network determines whether there are personnel matching the historical trajectory based on environmental clues within the region of interest and outputs their seat belt wearing status. If it is determined that there is a violation of not wearing a seat belt, an alarm is generated; otherwise, the judgment continues until the detected target reappears or the maximum number of frames lost is exceeded, and then the tracking is terminated.
[0102] Step 4 includes the following sub-steps:
[0103] Step 41: Solving the spatial microplane penetration under local surface normal vector constraints;
[0104] Step 42: Spatial topology verification and automatic hierarchical alarm;
[0105] Step 43: The platform layer pushes the generated alarm events to the application layer in real time, and the application layer performs 3D visualization and closed-loop with the mobile terminal.
[0106] Step 41 includes the following sub-steps:
[0107] Step 411: The platform layer server receives the pixel coordinates of each detected target. And the corresponding drone time information;
[0108] Step 412: Construct the ray direction vector of the camera mounted on the UAV;
[0109] Using the IMU attitude data of the UAV synchronously transmitted in step 2, combined with the pre-calibrated intrinsic parameter matrix of the camera Based on the camera mounting angle, calculate the camera ray direction vector from the UAV's optical center to the world coordinate system of the detected target pixel. ;
[0110] Step 413: Query the joint constraint matrix;
[0111] Based on the approximate surface projection location corresponding to the pixel coordinates, the joint constraint matrix pre-stored in step 11 of the platform layer is called. Obtain the absolute elevation of the ground at that location. and local surface normal vector Simultaneously, determine the corresponding three-dimensional surface point at this location. ;
[0112] Step 414: Construct the equation of the microtangent plane;
[0113] by Let be a point on the plane, with Using the normal vector, establish the equation of the three-dimensional micro-tangent plane that fits the actual slope of the terrain at that point:
[0114]
[0115] Step 415: Find the intersection of the ray and the plane;
[0116] Specifically, the camera ray parameter equations Substituting into the equation of the three-dimensional micro-tangent plane,
[0117] in, Location of the drone. Let the propagation parameters be to be determined. ,
[0118] The solution yields:
[0119]
[0120] After sorting, we get:
[0121]
[0122] Step 416: Obtain the true 3D coordinates of the target;
[0123] Specifically, the obtained Substituting back into the camera ray parameter equation, we obtain the precise three-dimensional coordinates of the detected target in the ENU coordinate system:
[0124]
[0125] Step 417: Timing smoothing;
[0126] Step 42 includes the following sub-steps:
[0127] Step 421: Calculate the three-dimensional coordinates obtained in Step 41 Convert back to WGS84 latitude and longitude coordinates ;
[0128] Step 422: Call the PostGIS database to query which pre-built electronic fence polygons the coordinate point falls within, and obtain the risk level of the corresponding electronic fence;
[0129] Step 423: Combine the seatbelt status obtained in Step 3 Alarms will be tiered according to the following rules:
[0130] Level 1 Emergency Alarm: Personnel are located within a high-risk electronic fence and their seatbelt status is "not worn"; at this time, the system immediately triggers an audible and visual alarm, sends a high-priority notification to the relevant safety officer, and coordinates with the drone to perform hovering video recording;
[0131] Level 2 warning: Personnel are located within a low-risk single-fence area and their safety belt status is "not worn"; at this time, the system only pushes a warning notification to remind the safety officer to pay attention.
[0132] Compliance Records: If personnel are within any electronic fence but are wearing safety belts, the system will only record compliance information and will not trigger an alarm.
[0133] No event: When personnel are outside all electronic fences, the system only records the personnel's location and does not trigger an alarm.
[0134] All alarm events, along with their location, time, screenshots of violations, and geofence attributes, are stored in the database to form a traceable security log.
[0135] In step 43, the 3D visualization platform loads the 3D reality model generated in step 11. After receiving the alarm, it determines the 3D coordinates of the violation location in the 3D reality model. The offender is marked with a prominent dynamic icon, and an information panel pops up to display the snapshot of the offender, the time, the area they belong to, and the risk level.
[0136] The mobile app / mini-program receives tiered alarm information in real time via cloud messaging services. The tiered alarm information includes the floor, axis number, location description, and geographic coordinates of the person who violated the rules, guiding the on-site safety officer to the designated location quickly. After completing the on-site correction, the safety officer uploads feedback data through the mobile app, forming a closed-loop feedback.
[0137] Compared with the prior art, the present invention has the following advantages:
[0138] 1. This invention can reduce the false negative rate under extreme working conditions: Through the spatiotemporal dual triggering and multimodal remediation mechanism of the ST-HMD model (dynamic gating of information entropy + covariance trace ROI amplification), the false negative rate of the target is effectively reduced compared with the traditional lightweight detection model in complex scenarios such as high altitude of the target construction site with scaffolding obstruction and posture distortion. Especially in the heavily obstructed scenario, the detection rate of the safety belt status is significantly improved by the context-based iterative inference of the sVLM auxiliary network.
[0139] 2. This invention enables centimeter-level 3D positioning without the need for terminal assistance: Relying on the 3D GS joint constraint matrix and micro-cutting plane ray intersection algorithm, the system effectively controls the comprehensive 3D spatial positioning error of violators (the error can be controlled within <20cm) without requiring operators to wear any positioning beacons. This makes the management closed loop of "violation detection - platform alarm - accurate 3D coordinate positioning - rapid on-site correction" possible, completely solving the industry pain point of "knowing the violation but not the location".
[0140] 3. This invention significantly improves inspection efficiency and safety: By coupling a path planning scheme that considers focal length, tilt angle, zoom cost, viewpoint quality, and collision repulsion, the UAV can complete high-definition inspections in an absolutely safe airspace far from high-risk equipment such as tower cranes. Compared to traditional "bow"-shaped close-range inspections, the overall flight range and operational efficiency are effectively improved, while the viewpoint penalty ensures a high-definition capture rate for even small safety belt attachment points. Attached Figure Description
[0141] Figure 1 This is an architecture diagram of the construction worker safety belt inspection and spatial positioning system of the present invention;
[0142] Figure 2 This is a flowchart of the construction worker safety belt inspection and spatial positioning method of the present invention. Detailed Implementation
[0143] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0144] Please see the appendix Figure 1 A safety belt inspection and spatial positioning system for construction workers includes a perception layer, an edge computing layer, a platform layer, and an application layer.
[0145] The perception layer includes an unmanned flight platform equipped with a telephoto camera, an RTK (Real-time kinematic) positioning module, an IMU (Neural Measuring Unit) attitude sensor, and an obstacle avoidance sensor.
[0146] Preferably, the unmanned flight platform can be an industrial-grade collision avoidance drone, which is equipped with a telephoto camera, RTK positioning module, IMU attitude sensor and obstacle avoidance sensor.
[0147] Long-range cameras, such as those supporting high-magnification optical zoom, can capture clear features of small targets like worker safety belts from a safe distance.
[0148] The RTK positioning module is used to provide real-time, interference-resistant centimeter-level WGS84 absolute geographic coordinates.
[0149] The IMU attitude sensor is used to output the yaw, pitch, and roll angles of the UAV at high frequency, providing a spatial reference for subsequent ray rendezvous.
[0150] Obstacle avoidance sensors can employ omnidirectional binocular vision or lidar to ensure the flight safety of unmanned aerial platforms.
[0151] The perception layer, serving as a multi-source data acquisition and maneuvering platform, is the "tentacle" of the entire system in contact with the physical world. It is responsible for executing patrol missions and acquiring real-time multimodal data from the site in all directions. Through an unmanned aerial platform, it autonomously flies over a high-altitude work surface along a planned complex route, simultaneously acquiring high-definition video streams, centimeter-level position coordinates, high-precision attitude information, and distances to surrounding obstacles.
[0152] The edge computing layer is equipped with edge computing devices that support tensor operations and deploys a spatio-temporal hybrid detection model (ST-HMD model) to process video streams in real time and output personnel positions and seat belt status.
[0153] The edge computing device is equipped with a high-performance GPU (Graphics Processing Unit) / NPU (Neural Processing Unit).
[0154] The aforementioned multimodal hybrid detection model (ST-HMD model) comprises two deeply fused AI engines. The first is a lightweight CNN (Convolutional Neural Network) main network that incorporates an SPD-Conv (Space-to-Depth Convolution) downsampling module and a VSS (Visual SourceSafe) module, responsible for high-frame-rate conventional feature extraction. The second is a small visual language model (sVLM) auxiliary network. When the lightweight CNN main network experiences high-entropy confusion (i.e., low confidence) or misses detections due to severe occlusion, the system dynamically extracts contextual region blocks through Kalman filtering covariance, forcibly triggering the small visual language model (sVLM) auxiliary network to perform commonsense reasoning and adaptive fusion, thereby overcoming the perceptual blind spots of traditional visual algorithms.
[0155] As a spatiotemporal multimodal intelligent retrieval engine, the edge computing layer is the "front-end brain" of the entire system. It accelerates the processing of video frames transmitted back by the drone in real time through tensor operations, overcomes complex occlusions by using a spatiotemporal dual-trigger mechanism, and accurately outputs the two-dimensional bounding box (pixel coordinates) of the workers in the picture and their safety belt wearing status. It completes low-latency and high-precision inference of high-concurrency video streams on the airborne or near-end.
[0156] The platform layer includes ground servers that provide path planning services, 3D Gaussian splash (3D GS) 3D modeling services, spatial positioning and calculation services, and geofencing databases.
[0157] Among them, the 3D Gaussian splash 3D modeling service uses pre-flight images to quickly render real scenes and extracts a joint constraint matrix composed of the absolute elevation of the ground surface and the local surface normal vector.
[0158] It provides path planning services and runs multi-objective optimization algorithms to generate optimal routes that take into account focal length-tilt coupling and observation angle penalties.
[0159] The spatial positioning solution service combines the pixel coordinates extracted from the edge with the drone's attitude and joint constraint matrix, and uses a micro-plane penetration algorithm to calculate the true ENU coordinates of the violator. The ENU coordinate system includes the E-axis (East): pointing due east, the N-axis (North): pointing due north, and the U-axis (Up): perpendicular to the local horizontal plane and pointing upwards.
[0160] The geofencing database, built on PostGIS (a PostgreSQL-based object-relational database extension module), is used to store compliant high-risk area polygons. It utilizes GIST (Generalized Search Trees) indexes to achieve millisecond-level ST_Contains (ST_Contains is a commonly used function in spatial databases and Geographic Information Systems (GIS) to determine whether one geometric object completely contains another geometric object) spatial boundary checks.
[0161] The platform layer, as the central hub for 3D calculation and data fusion, is the "computing core and data platform" of the entire system. It is responsible for mapping 2D visual features to physical coordinates in the real world and establishing security control rules. It is used to provide preliminary route planning, 3D real-world elevation extraction, and core monocular ray intersection positioning calculation. It also uses the spatial database to perform topological boundary verification of the electronic fence.
[0162] The application layer includes a 3D visualization platform and mobile terminals.
[0163] Preferably, the 3D visualization platform can be a large-screen 3D visualization platform based on WebGL technology (such as the Cesium.js engine), which is responsible for loading the 3D real-world model, marking the 3D location of the violation target with eye-catching UI (User Interface Design) such as red pulses, and displaying high-definition screenshots of the violation.
[0164] The mobile terminal can be an intelligent mobile terminal equipped with an APP / mini-program, which is responsible for receiving graded alarm information in real time (Level 1 emergency / Level 2 early warning), guiding safety officers to handle the situation based on precise floor / location coordinates, and completing the digital closed loop of the entire safety supervision business by uploading rectification feedback.
[0165] The application layer, serving as the digital twin interaction and closed-loop management terminal, acts as the "command panel" for the entire system. Directly facing safety management personnel, it enables the visualization of data and the issuance of emergency commands. Within the 3D digital twin scenario, it intuitively presents the actual construction site conditions, drone trajectories, and tiered early warning information, and facilitates rapid on-site correction and closed-loop feedback through multi-channel push notifications.
[0166] The system of this invention adopts a four-layer vertical architecture of "perception layer - edge computing layer - platform layer - application layer", with each layer module working collaboratively.
[0167] Please see the appendix Figure 2 A method for inspecting and spatially locating safety belts for construction workers, comprising the following steps:
[0168] Step 1: Preliminary environmental geometric feature extraction and view-aware route planning (led by the platform layer).
[0169] Step 1 is executed at the platform layer (ground server). Its core function is to pre-build a digital geometric model of the target construction site and plan the optimal UAV inspection route accordingly. Step 1 simultaneously extracts the absolute elevation of the ground surface and the local surface normal vector from the 3D reality model as spatial positioning priors, and introduces a "viewpoint penalty term" in the route planning to encourage the UAV to face the observed surface vertically, thereby improving the ability to capture small safety belt attachment points. Step 1 breaks the physical limitation of monocular vision's "lack of depth" by pre-extracting the terrain features of the construction site as geometric priors for 3D positioning and actively planning the cruise trajectory with the best observation viewpoint. This solves the problems of large monocular positioning errors when lacking depth priors and the difficulty in obtaining clear details with fixed viewpoint routes in traditional methods.
[0170] Step 1 includes the following sub-steps:
[0171] Step 11: Construct a 3D reality model based on the 3D Gaussian splash 3D modeling service of the platform layer, and extract the joint constraint matrix.
[0172] Step 11 includes the following sub-steps:
[0173] Step 111: The unmanned aerial vehicle platform of the perception layer performs a "pre-flight oblique photography" mission over the target construction site to collect a multi-view image sequence covering the entire work surface.
[0174] During image acquisition, sufficient forward and lateral overlap should be ensured to guarantee the quality of subsequent 3D reconstruction.
[0175] Step 112: After image acquisition is complete, the images are transmitted back to the platform layer. The platform layer uses 3D Gaussian splashing technology to quickly render a 3D real-world model.
[0176] This 3D Gaussian splashing technique represents a scene as a large number of Gaussian ellipsoids with position, covariance, opacity, and spherical harmonic coefficients. Through differentiable rendering optimization, it can obtain a high-fidelity 3D scene representation in a short time. 3D Gaussian splashing is a conventional technique in this field, and its specific rendering process will not be elaborated here.
[0177] Step 113: Set up the orthogonal overhead virtual camera.
[0178] A virtual camera is positioned above the 3D reality model, with its optical axis pointing vertically downwards. The camera is located at a safe height above the highest point of the 3D reality model (this safe height should ensure complete coverage of the entire work area); the absolute elevation of this virtual camera is denoted as... .
[0179] Step 114: Calculate the depth map for volume rendering.
[0180] For each pixel on the image plane captured by the virtual camera The surface depth value corresponding to this pixel is calculated using the volume rendering formula. The volume rendering formula is:
[0181]
[0182] in, The total number of Gaussian spheres traversed along the ray. For the first The opacity of a Gaussian sphere For the first The distance from the center of the Gaussian sphere to the optical center of the virtual camera.
[0183] The volume rendering formula obtains the depth at which the ray first hits the scene surface by accumulating opacity from front to back.
[0184] Step 115: Generate an absolute elevation map of the Earth's surface.
[0185] Combined with the absolute elevation of the virtual camera and depth map Calculate the absolute elevation of the ground surface corresponding to each pixel:
[0186]
[0187] In the formula, This represents the geographic plane coordinates corresponding to this pixel.
[0188] This yields an absolute elevation map of the entire target construction site.
[0189] Step 116: Extract local surface normal vectors.
[0190] For each point on the absolute elevation map of the Earth's surface, using the three-dimensional coordinates of its neighboring points, a local plane is fitted using the least squares method, and then the unit normal vector of this local plane, i.e., the local surface normal vector, is solved. .
[0191] The local surface normal vector reflects the tilt direction and slope of the local terrain at that point. Preferably, Gaussian weighting can be used in the calculation to enhance the contribution of the neighborhood center.
[0192] Step 117: Construct the joint constraint matrix: Organize the calculated absolute surface elevation and local surface normal vectors into a spatial lookup table according to a spatial grid (the grid size can be set according to the actual accuracy requirements), denoted as the joint constraint matrix. :
[0193]
[0194] The joint constraint matrix is persistently stored in the platform layer database and a spatial index (such as the GiST index) is created to enable quick lookup of the corresponding absolute surface elevation and local surface normal vector based on arbitrary geographic coordinates in subsequent steps. Simultaneously, this joint constraint matrix also serves as the core mathematical support for eliminating "slope projection error" in step 4.
[0195] Step 12: Build a standardized geofencing database.
[0196] Step 12 includes the following sub-steps:
[0197] Step 121: Safety management personnel delineate the outer boundary (geometric polygon) of the high-risk area on the 3D reality model according to relevant safety regulations.
[0198] The outer boundary of the high-risk area includes edges, openings, suspended work platforms, and climbing structures; whereby edges include floor slabs and roof edges, openings include elevator shafts and reserved openings, suspended work platforms include scaffolding and suspended baskets, and climbing structures include ladders and vertical passages.
[0199] Step 122: Convert the outer boundary (geometric polygon) of the extracted high-risk area into WGS84 geographic coordinates (using the existing georegistration information of the 3D reality model), and add attributes to it: area name, risk level (which can be divided into several levels, such as high risk, low risk, etc.), corresponding safety regulations, etc.
[0200] Step 123: Store the outer boundary (geometric polygon) data of the above-mentioned high-risk areas into the PostGIS spatial database, and create a GiST index for the geographic geometry field to support subsequent spatial queries.
[0201] Step 13: Focal length-tilt angle coupled route planning based on view perception.
[0202] Based on the distribution of high-risk areas determined in step 12, the platform layer runs a multi-objective optimization algorithm to generate the optimal three-dimensional waypoint sequence (i.e., inspection route).
[0203] Step 13 includes the following sub-steps:
[0204] Step 131: Establish the following comprehensive optimization objective function:
[0205]
[0206] in, Total path length, which is the sum of Euclidean distances between adjacent waypoints, is used to ensure that the route is as short as possible.
[0207] : Number of sharp turns, defined as the number of adjacent segments where the rate of change of heading angle exceeds a predetermined threshold, used to make the route smoother.
[0208] Zoom switching cost. Since the optical zoom camera on the drone requires time and energy to switch between different focal lengths, a cost function is introduced:
[0209]
[0210] in, , The focal length of adjacent viewpoints. This refers to the switching time required for the zoom motor. , These are predetermined weighting coefficients. This cost function is used to encourage gradual changes in focal length between adjacent viewpoints.
[0211] The collision risk penalty term (based on exponentially decaying collision repulsion) is calculated using the following formula:
[0212]
[0213] in, For the first The distance from each waypoint to the nearest obstacle (extracted based on a digital surface model). and This is a pre-set normal value. This collision risk penalty causes the flight path to tend to move away from obstacles.
[0214] The viewing angle penalty term (which encourages the optical axis to be perpendicular to the measured surface) is defined as follows:
[0215]
[0216] in, It is the unit direction vector of the camera's optical axis (determined by the drone's attitude). This is the local surface normal vector of the working surface extracted in step 11. This perspective penalty term encourages the UAV to adjust its attitude so that the camera optical axis is as perpendicular as possible to the observed surface (i.e., the absolute value of the dot product is close to 1), thereby obtaining the clearest observation perspective, greatly improving the capture rate of small seat belt attachment points, and realizing a comprehensive consideration of active obstacle avoidance and optimal observation perspective.
[0217] These are all weighting coefficients, which can be preset according to actual engineering needs, and usually meet the following requirements. .
[0218] Step 132: Solve the comprehensive optimization objective function to obtain the three-dimensional waypoint sequence (i.e. inspection route) that can be used for UAV flight. The generated three-dimensional waypoint sequence (including the three-dimensional coordinates, expected flight speed, camera focal length and tilt angle of each waypoint) is sent to the UAV flight platform in the perception layer.
[0219] Preferably, a hybrid algorithm integrating diverse ant colony optimization and differential evolution can be used to solve the comprehensive optimization objective function. The solution process is as follows: First, a set of initial candidate paths is generated. Pareto front solutions are obtained through iterative optimization. Then, the path with the minimum comprehensive cost is selected as the final flight path. Finally, the waypoint sequence is smoothed using B-spline curves to obtain a three-dimensional waypoint sequence (i.e., inspection route) suitable for UAV flight.
[0220] Traditional full-coverage route planning is mostly based on the assumption of "orthogonal polygons", which assumes that the camera is vertically downward and the focal length is fixed. This cannot adapt to the actual working conditions of high-altitude inspection, which requires "maintaining a safe distance and high-magnification zoom". Furthermore, it lacks joint quantification of the mechanical cost of camera zoom, viewpoint quality, and dynamic obstacle avoidance.
[0221] This invention specifically addresses the coverage characteristics of telephoto cameras at fixed tilt angles, deriving a trapezoidal coverage model and innovatively incorporating zoom switching costs (focal length abrupt change penalty + switching time penalty) and viewing angle penalty (encouraging the optical axis to be perpendicular to the surface) into a multi-objective optimization framework. Simultaneously, it introduces a collision repulsion term based on exponential decay, enabling the UAV to proactively avoid high-risk obstacles such as tower cranes and scaffolding during the route planning phase. This fills the theoretical gap in existing path planning algorithms for telephoto inspection scenarios, significantly improving the inspection efficiency and image acquisition quality of UAVs while ensuring safety.
[0222] Step 2: Aerial autonomous inspection and synchronous acquisition and encapsulation of multi-source spatiotemporal data (led by the perception layer).
[0223] The core function of step 2 is to synchronously collect and accurately align multi-source spatiotemporal data during the autonomous flight of the drone. Step 2 is strictly executed by the physical hardware of the perception layer, and its goal is to digitize the physical information of the high-altitude work surface, ensuring high synchronization between the timestamp and the spatial coordinate system, and providing a data source for downstream AI.
[0224] Step 2 includes the following sub-steps:
[0225] Step 21: After receiving the optimal inspection route calculated in step 13 from the platform layer, the UAV in the perception layer automatically unlocks and takes off and flies autonomously along the inspection route. During the flight, the UAV's flight control system collects multi-source data in real time at a preset high frequency (e.g., 10Hz).
[0226] The multi-source data includes: RTK positioning data: used to provide centimeter-level absolute geographic coordinates for the UAV. (Usually converted to the ENU coordinate system or directly using WGS84 latitude, longitude, and altitude); IMU attitude data: synchronously records the UAV's yaw, pitch, roll, and angular velocity of each axis. Video stream: The telephoto camera continuously captures high-definition images at a fixed frame rate (e.g., 30fps) and automatically zooms according to the preset focal length value in the inspection route.
[0227] Step 22: The UAV's onboard controller uses a nanosecond-level hardware clock to assign a unified timestamp to each frame of video image, each RTK positioning point, and each set of IMU attitude data.
[0228] Step 23: Encapsulate the acquired multi-source data into a standard data stream according to the principle of time alignment: each frame of image is equipped with its corresponding RTK coordinates and IMU pose at the time closest to the time stamp.
[0229] Step 24: The encapsulated standard data stream is pushed to the edge computing layer in real time via the UAV's onboard high-speed bus.
[0230] In step 2, the data from the obstacle avoidance sensor (LiDAR or binocular vision) is also accessed to the system at an appropriate frequency for dynamic obstacle avoidance during drone flight, ensuring the safe flight of the drone. However, the data from the obstacle avoidance sensor does not directly participate in the multi-source data encapsulation in step 2.
[0231] Step 3: Spatiotemporal multimodal detection and severe occlusion missed detection remediation at the edge (led by the edge computing layer).
[0232] The core function of step 3 is to detect the seat belt status in real time on the edge computing device and perform spatiotemporal remediation for missed detections caused by high occlusion or posture distortion. Step 3 abandons the fixed confidence threshold and adopts dynamic gating driven by information entropy, as well as an adaptive ROI amplification mechanism based on covariance traces, to achieve a closed loop of "self-doubt-semantic help-active retrieval". After receiving the standard data stream returned from step 2 of the perception layer, the multimodal hybrid detection model (ST-HMD model) deployed in the edge computing layer starts working, solving the problems of missed blind spots and inaccurate judgment of complex postures caused by personnel occlusion in complex construction sites.
[0233] Step 3 includes the following sub-steps:
[0234] Step 31: First, run a lightweight CNN main network on the edge computing device to perform pre-feature extraction.
[0235] Preferably, a lightweight CNN main network can adopt the following key designs:
[0236] The input image is appropriately scaled (e.g., to 640×640) and data augmentation strategies such as Mosaic are applied.
[0237] The SPD-Conv downsampling module is introduced into the backbone network to replace the traditional stride convolution, which preserves the fine-grained spatial information of small targets during downsampling.
[0238] The VSS (Visual State Space) module is introduced to achieve global feature dependency modeling with linear complexity, thereby enhancing the perception of long-distance context.
[0239] The neck area employs bidirectional, cross-scale dense connectivity to efficiently integrate shallow details with deep semantics.
[0240] The detection head uses an anchor-free structure and directly outputs the center pixel coordinates of the 2D bounding box of the person target. Width and height And three types of probability distributions: ,in, These represent "personnel", "wearing a seat belt", and "not wearing a seat belt", respectively.
[0241] The lightweight CNN main network operates at a high frequency with the video frame rate as the step size, and the inference results of each frame (including bounding boxes and probability distributions) are used as the initial output. The CNN main network is a common feature extraction method in this field, and will not be elaborated here.
[0242] Step 32: Dynamic gating correction of spatial dimension based on information entropy.
[0243] In traditional methods, when the detection confidence of a CNN network is low, a fixed threshold is often used to discard the data or simply output it. This invention abandons the traditional fixed threshold and introduces information entropy to quantify the degree of disorder in the current prediction of the lightweight CNN main network, and dynamically adjusts the confidence level of the lightweight CNN main network accordingly.
[0244] Step 32 includes the following sub-steps:
[0245] Step 321: For each detected target, calculate the information entropy of its probability distribution:
[0246]
[0247] The value of information entropy The lower the value, the more certain the classification decision of the lightweight CNN main network is; the higher the value of information entropy, the flatter the probability distribution of the lightweight CNN main network for each category (i.e., the higher the degree of "confusion").
[0248] Step 322: Generate dynamic adaptive weights for the lightweight CNN main network based on the information entropy value:
[0249]
[0250] in, This is a preset adjustment coefficient (positive number).
[0251] when hour, ;along with Increase It decreased exponentially.
[0252] Step 323: When When the value falls below a certain preset threshold (indicating that the lightweight CNN main network itself is highly uncertain in its judgment), the system triggers a small visual language model (sVLM) auxiliary network to assist in reasoning.
[0253] The input to this small visual language model (sVLM) auxiliary network includes: an image patch of the monitored target region cropped from the original image (with appropriate expansion of the context); and structured text prompts containing the initial category from the output of the lightweight CNN main network, the pose estimate of the monitored target (which can be provided by an additional keypoint network), the occlusion ratio estimate, and a description of the current target construction site scene.
[0254] Step 324: After fine-tuning, the small visual language model (sVLM) auxiliary network is able to reason based on common sense and contextual semantics, and output the corrected three-class probability distribution. And the corresponding confidence level.
[0255] Step 325: The seatbelt status is determined by a weighted fusion of the lightweight CNN main network and the small visual language model (sVLM) auxiliary network. The weighted fusion formula is as follows:
[0256]
[0257] This weighted fusion formula guarantees that the more confused the lightweight CNN main network becomes (i.e., the more confused it is...), the better. The higher, When the size is smaller, the decision-making power is smoothly transferred to the small visual language model (sVLM) auxiliary network with semantic understanding capabilities, thereby effectively correcting the misjudgment of the lightweight CNN main network under complex occlusion or pose distortion.
[0258] Step 33: Time dimension recovery (context extraction based on covariance trace).
[0259] When the lightweight CNN main network in step 31 fails to output any bounding boxes (i.e., missed detections) because the target is severely occluded (e.g., workers are completely blocked by scaffold beams), the aforementioned spatial dimension correction will not work because there is no corresponding three-class probability distribution. At this point, a time-based remedial mechanism is activated.
[0260] Step 33 includes the following sub-steps:
[0261] Step 331: The system maintains the historical state vector of each tracked target. The historical state vector typically contains the target's position, velocity, size, and rate of change on the image plane. State prediction is performed using Kalman filtering.
[0262] Predicted status: ,in, This is the state transition matrix (which can be set according to the uniform motion model or the uniform acceleration model).
[0263] Prediction error covariance matrix: ,in, The process noise covariance matrix; the trace of the covariance matrix Reflecting the uncertainty of the current prediction: the longer the target is lost, the greater the prediction variance. Also bigger.
[0264] Step 332: The edge computing layer calculates based on the trace of the covariance matrix. The cropping range of the region of interest (ROI) is dynamically adjusted (enlarged) to reflect the uncertainty:
[0265]
[0266] in, To predict the width and height of the target bounding box, This is the preset magnification factor. That is, the higher the prediction uncertainty, the greater the expansion of the clipping frame.
[0267] Step 333: The system forcibly extracts the enlarged region of interest from the image and feeds it into a small visual language model (sVLM) auxiliary network for "context retrieval reasoning". The small visual language model (sVLM) auxiliary network attempts to determine whether there are personnel matching the historical trajectory based on environmental cues within the region of interest (such as the color of clothing on scaffolding, local features of workers, etc.) and outputs their seat belt wearing status. If a violation of not wearing a seat belt is determined, an alarm is generated; otherwise, the judgment continues until the detected target reappears or the maximum allowable number of lost frames is exceeded, at which point the tracking terminates.
[0268] "Context retrieval reasoning" is a routine data processing method of the small visual language model (sVLM) auxiliary network. It is automatically executed by the small visual language model (sVLM) auxiliary network, and its specific operation process will not be described in detail here.
[0269] By employing a spatiotemporal dual-trigger mechanism (information entropy gating in the spatial dimension + covariance trace ROI amplification in the temporal dimension), the detection rate of seat belt status in high-altitude occlusion scenarios is significantly improved.
[0270] Existing cascaded detection relies on rigid confidence thresholds, which are prone to false positives due to overconfidence. This invention uses spatial entropy gating to directly map physical prediction errors (the trace of covariance) to the ROI of the large visual model for field of view capture. The higher the entropy value, the lower the weight; the longer the data is lost, the wider the field of view. When encountering severe occlusion, the system can achieve "self-doubt-expanded search-contextual assistance," deeply binding underlying mathematical uncertainty with AI semantic reasoning. This significantly breaks through the detection limits of existing CNNs under complex occlusion conditions, realizing a multimodal fusion from "static threshold judgment" to "dynamic uncertainty-driven" approaches.
[0271] Step 4: Monocular localization calculation and hierarchical control closed loop of local surface normal vector on the platform side (executed collaboratively by the platform layer and the application layer).
[0272] The core function of step 4 is to convert the two-dimensional pixel coordinates in the image into three-dimensional geographic coordinates in the real world, and to implement hierarchical alarms and terminal closed-loop based on the electronic fence. Step 4 uses the local surface normal vectors provided by the 3D real-scene model to construct a micro-tangent plane, replacing the strong assumption of "absolutely level ground" in traditional algorithms, thereby achieving centimeter-level monocular positioning on terrains such as slopes, steps, and sloping roofs. Step 4 obtains the "two-dimensional pixel coordinates" output from step three of the edge computing layer. and seat belt status Afterwards, the data is sent back to the platform layer, where an innovative micro-cutting plane algorithm is used for spatial mapping, and finally the business loop is completed at the application layer. This solves the engineering problem that a single visual sensor cannot obtain depth information and traditional GPS positioning is difficult to achieve centimeter-level accuracy, and realizes the complete positioning capability from "discovering violations" to "finding violators".
[0273] Step 4 includes the following sub-steps:
[0274] Step 41: Spatial microplane penetration solution under local surface normal vector constraints (executed by the platform layer).
[0275] Step 41 includes the following sub-steps:
[0276] Step 411: The platform layer server receives the pixel coordinates of each detected target. And the corresponding drone time information.
[0277] Step 412: Construct the ray direction vector of the camera mounted on the UAV.
[0278] Specifically, the attitude data (yaw angle, pitch angle, roll angle) of the UAV synchronously transmitted in step 2 is used in conjunction with the pre-calibrated intrinsic parameter matrix of the camera. And camera mounting angle (camera → body rotation matrix) Rotation matrix of machine body → ENU ), calculate the camera ray direction vector from the UAV optical center to the world coordinate system of the detected target pixel. (Normalized).
[0279] After the camera is mounted on the drone, its intrinsic parameter matrix The camera mounting angle can be set according to actual shooting needs, and the camera ray direction vector can then be calculated. This calculation process is a standard method in the field of drones and will not be elaborated upon here.
[0280] Step 413: Query the joint constraint matrix.
[0281] Specifically, based on the approximate projection location of the ground corresponding to the pixel coordinates (estimated through a rough depth assumption), the joint constraint matrix pre-stored in step 11 by the platform layer is called. Obtain the absolute elevation of the ground at that location. and local surface normal vector At the same time, determine the corresponding three-dimensional surface point. (Combination of planar projection coordinates) ).
[0282] Step 414: Construct the equation of the micro-tangent plane.
[0283] Specifically, with Let be a point on the plane, with Using the normal vector, establish the equation of the three-dimensional micro-tangent plane that fits the actual slope of the terrain at that point:
[0284]
[0285] Step 415: Find the intersection of the ray and the plane.
[0286] Specifically, the camera ray parameter equations ( Location of the drone. Let the propagation parameters be to be determined. Substituting into the equation of the three-dimensional micro-tangent plane, we obtain:
[0287]
[0288] After sorting, we get:
[0289]
[0290] Here we need to ensure the denominator (i.e., the ray is not parallel to the ground). If the value is close to 0, a micro-perturbation or a return to the horizontal plane assumption can be used.
[0291] Step 416: Obtain the true three-dimensional coordinates of the target.
[0292] Specifically, the obtained Substituting back into the camera ray parameter equation, we obtain the precise three-dimensional coordinates of the detected target in the ENU coordinate system:
[0293]
[0294] The core advantage of this algorithm is that it no longer assumes that the ground is an ideal horizontal plane, but instead uses the real surface slope and normal vector extracted from the 3D real scene model to construct a 3D micro-cut plane that fits the terrain, thereby completely overcoming the projection error of traditional monocular positioning on terrain such as slopes, steps, and sloping roofs, and achieving centimeter-level spatial positioning.
[0295] Step 417: Timing smoothing.
[0296] For the solution results of multiple consecutive frames of the same detection target, Kalman filtering or moving average can be used to smooth the data in order to suppress noise and jitter.
[0297] Traditional drone monocular vision positioning either assumes that the construction site is an absolutely flat, ideal level surface (resulting in meter-level errors on slopes, steps, and sloping roofs), or relies on workers carrying high-precision RTK terminals (which are costly and cannot monitor workers not wearing positioning beacons), or requires increasing payload and computing power through binoculars or lidar.
[0298] This invention utilizes a 3D real-world model generated using 3D Gaussian splashing technology in the early stages of inspection. From this model, a high-precision absolute elevation matrix of the ground surface is extracted through volume rendering, and local surface normal vectors are simultaneously derived to form a joint constraint matrix. During real-time positioning, instead of assuming a level ground surface, a 3D micro-tangent plane equation conforming to the terrain is constructed using the actual normal vector of the point. The camera ray intersects this 3D micro-tangent plane to solve for the 3D coordinates; this is equivalent to pre-setting a precise "geometric intercept surface" in 3D space for each pixel ray of the UAV camera. By mathematically reducing the dimension of the "prior 3D scene reconstruction" and the "real-time 2D target ray," a monocular camera achieves centimeter-level positioning capabilities comparable to LiDAR.
[0299] Step 42: Spatial topology verification and automatic hierarchical alarm (executed by the platform layer).
[0300] Step 42 includes the following sub-steps:
[0301] Step 421: Calculate the three-dimensional coordinates obtained in Step 41 Convert back to WGS84 latitude and longitude coordinates .
[0302] Step 422: Call the PostGIS database to query which pre-built electronic fence polygons the coordinate point falls within, and obtain the risk level of the corresponding electronic fence (such as high risk, low risk).
[0303] Step 423: Combine the seatbelt status obtained in Step 3 Alarms will be tiered according to the following rules:
[0304] Level 1 Emergency Alarm: Personnel are located within a high-risk electronic fence (such as near an edge, opening, or high-altitude work platform) and their safety belt is not worn. At this time, the system immediately triggers an audible and visual alarm, pushes a high-priority notification to the relevant safety manager, and can coordinate with drones to perform enhanced actions such as hovering video recording.
[0305] Level 2 warning: Personnel are located within a low-risk single-fence area (such as a regular walkway or material storage area) and their safety belt status is "not worn"; at this time, the system only pushes a warning notification to remind the safety officer to pay attention.
[0306] Compliance Records: If personnel are within any electronic fence but are wearing safety belts, the system will only record compliance information and will not trigger an alarm.
[0307] No event: When personnel are outside all electronic fences, the system only records the personnel's location and does not trigger an alarm.
[0308] All alarm events, along with information such as location, time, screenshots of violations, and geofence attributes, are stored in the database to form a traceable security log.
[0309] Step 43: The platform layer pushes the generated alarm events to the application layer in real time, and the application layer performs 3D visualization and closed-loop operation with the mobile terminal (executed by the application layer).
[0310] The aforementioned 3D visualization is executed by a 3D visualization platform (such as a web application based on Cesium.js). The 3D visualization platform loads the 3D reality model generated in step 11 (which needs to be converted to a streaming format such as 3D Tiles). Upon receiving an alarm, it determines the 3D coordinates of the violation location in the 3D reality model. The offender is marked with a prominent dynamic icon (such as a red pulse sphere or flashing light bar) and a pop-up information panel displays the offender's snapshot, time, fence, risk level, etc.
[0311] Preferably, the 3D visualization platform supports multi-view roaming, and managers can click on alarm points to view detailed data and replay historical trajectories.
[0312] The mobile app / mini-program receives tiered alarm information in real time via cloud messaging services (such as MQTT, WebSocket, or mobile push channels). This tiered alarm information includes the floor, axis number, location description, and geographic coordinates of the violating personnel, guiding on-site safety officers to the designated location quickly. After completing on-site corrections (such as requiring workers to wear safety belts or evacuate hazardous areas), the safety officer uploads feedback data, including rectification photos and handling instructions, via the mobile app, forming a closed-loop feedback system.
[0313] All feedback data is automatically synchronized to the application layer database, completing the digital closed loop of security supervision business, and can generate statistical reports (such as violation heat maps, high-incidence period analysis, etc.).
[0314] This invention constructs a complete and highly automated safety belt inspection and spatial positioning system, from early environmental modeling and flight path planning to aerial data acquisition and edge intelligent detection, then to high-precision positioning and hierarchical alarms on the platform, and finally to 3D visualization and mobile closed-loop processing. It can achieve panoramic coverage and high-precision traceability, and can integrate multimodal semantic common sense reasoning (to solve the problem of missed detection due to severe occlusion) and 3D surface geometric constraints (to solve the problem of monocular ranging distortion) for hybrid detection and positioning. It can achieve accurate identification and centimeter-level spatial positioning of personnel wearing safety belts on complex high-altitude work surfaces, and solve the problems of poor model robustness and monocular spatial projection distortion under extreme working conditions, so as to fundamentally open up the high-altitude operation safety management link of "automatic discovery-precise positioning-on-site closed loop".
[0315] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the invention. Therefore, any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A safety belt inspection and spatial positioning system for construction workers, characterized by: It includes the perception layer, edge computing layer, platform layer, and application layer; The perception layer includes an unmanned flight platform equipped with a telephoto camera, an RTK positioning module, an IMU attitude sensor, and an obstacle avoidance sensor. The edge computing layer is equipped with edge computing devices that support tensor operations and deploys a spatiotemporally triggered multimodal hybrid detection model to process video streams in real time and output personnel positions and seat belt status. The edge computing devices are equipped with GPUs / NPUs. The multimodal hybrid detection model includes two deeply fused AI engines: one is a lightweight CNN main network that incorporates SPD-Conv downsampling and VSS modules, responsible for high frame rate conventional feature extraction; the other is a small visual language model auxiliary network. The platform layer includes ground servers that provide path planning services, 3D Gaussian splash 3D modeling services, spatial positioning and calculation services, and a geofencing database. The 3D Gaussian splash 3D modeling service uses pre-flight imagery to render the real scene and extracts a joint constraint matrix composed of absolute ground elevation and local surface normal vectors. The path planning service runs a multi-objective optimization algorithm to generate the optimal flight path that balances focal length-tilt coupling and observation angle penalty. The spatial positioning and calculation service uses pixel coordinates extracted from the edge, combined with the UAV attitude and the joint constraint matrix, to calculate the true ENU coordinates of violators through a micro-plane penetration algorithm. The geofencing database, built on PostGIS, stores compliant high-risk area polygons and uses the GIST index to achieve millisecond-level ST_Contains spatial boundary judgment. The application layer includes a 3D visualization platform and a mobile terminal. The 3D visualization platform is responsible for loading 3D reality models, marking the 3D location of the violation targets, and displaying high-definition screenshots of the violations. The mobile terminal is responsible for receiving hierarchical alarm information in real time, guiding safety officers to handle the situation based on precise floor / direction coordinates, and completing the digital closed loop of the entire safety supervision business by uploading rectification feedback.
2. A method for safety belt inspection and spatial positioning using the construction worker safety belt inspection and spatial positioning system as described in claim 1, characterized in that: Includes the following steps: Step 1: Preliminary environmental geometric feature extraction and viewpoint-aware route planning; Step 2: Autonomous aerial inspection and synchronous acquisition and encapsulation of multi-source spatiotemporal data; Step 3: Spatiotemporal multimodal detection at the edge and remediation of missed detections due to severe occlusion; Step 4: Monocular localization calculation of local surface normal vectors on the platform side and hierarchical control closed loop.
3. The seat belt inspection and spatial positioning method according to claim 2, characterized in that: Step 1 includes the following sub-steps: Step 11: Construct a 3D reality model based on the 3D Gaussian splash 3D modeling service at the platform layer, and extract the joint constraint matrix; Step 12: Construct a standardized geofencing database; Step 13: Focal length-tilt angle coupled route planning based on view perception.
4. The seat belt inspection and spatial positioning method according to claim 3, characterized in that: Step 11 includes the following sub-steps: Step 111: The unmanned aerial platform of the perception layer performs a "pre-flight oblique photography" mission over the target construction site to collect a multi-view image sequence covering the entire work surface; Step 112: After image acquisition is completed, the image is transmitted back to the platform layer; the platform layer uses 3D Gaussian splashing technology to render a three-dimensional real-world model; Step 113: Set up an orthogonal overhead virtual camera; A virtual camera is positioned above the 3D reality model, with its optical axis pointing vertically downwards. The camera is located at a safe height above the highest point of the 3D reality model; the absolute elevation of this virtual camera is denoted as... ; Step 114: Calculate the depth map for volume rendering; For each pixel on the image plane captured by the virtual camera The surface depth value corresponding to this pixel is calculated using the volume rendering formula. The volume rendering formula is: ; in, The total number of Gaussian spheres traversed along the ray. For the first The opacity of a Gaussian sphere For the first The distance from the center of the Gaussian sphere to the optical center of the virtual camera; The volume rendering formula obtains the depth at which the ray first hits the scene surface by accumulating opacity from front to back; Step 115: Generate an absolute surface elevation map; Combined with the absolute elevation of the virtual camera and depth map Calculate the absolute elevation of the ground surface corresponding to each pixel: ; In the formula, The geographic plane coordinates corresponding to this pixel; This yields an absolute elevation map of the entire target construction site. Step 116: Extract local surface normal vectors; For each point on the absolute elevation map of the Earth's surface, using the three-dimensional coordinates of its neighboring points, a local plane is fitted using the least squares method, and then the unit normal vector of this local plane, i.e., the local surface normal vector, is solved. ; Step 117: Construct the joint constraint matrix: Organize the calculated absolute surface elevation and local surface normal vectors into a spatial lookup table according to a spatial grid, denoted as the joint constraint matrix. : ; The joint constraint matrix is persistently stored in the platform layer database, and a spatial index is created. Step 12 includes the following sub-steps: Step 121: Safety management personnel delineate the outer boundary of the high-risk area on the 3D reality model according to relevant safety regulations; The outer boundary of a high-risk area includes edges, openings, suspended work platforms, and climbing structures; among them, edges include floor slabs and roof edges, openings include elevator shafts and reserved openings, suspended work platforms include scaffolding and suspended baskets, and climbing structures include ladders and vertical passages. Step 122: Convert the outer boundary of the extracted high-risk area into WGS84 geographic coordinates and add the following attributes: area name, risk level, and corresponding safety regulation clauses. Step 123: Store the outer boundary data of the above-mentioned high-risk areas into the PostGIS spatial database, and create a GiST index for the geographic geometry field; Step 13 includes the following sub-steps: Step 131: Establish the following comprehensive optimization objective function: ; in, Total path length, which is the sum of Euclidean distances between adjacent waypoints; The number of sharp turns is defined as the number of adjacent segments where the rate of change of heading angle exceeds a predetermined threshold. : Zoom switching cost; Since the optical zoom camera on the drone requires time and energy to switch between different focal lengths, a cost function is introduced: ; in, , The focal length of adjacent viewpoints. This refers to the switching time required for the zoom motor. , These are predetermined weighting coefficients; this cost function is used to encourage gradual changes in focal length between adjacent viewpoints. The collision risk penalty is calculated using the following formula: ; in, For the first Distance from each waypoint to the nearest obstacle and This is a pre-set normal value; the collision risk penalty causes the flight path to tend to move away from obstacles. The perspective penalty is defined as follows: ; in, Let be the unit direction vector of the camera's optical axis. This refers to the local surface normal vector of the working surface extracted in step 11; All are weighting coefficients, satisfying ; Step 132: Solve the comprehensive optimization objective function to obtain a three-dimensional waypoint sequence that the UAV can fly. The generated three-dimensional waypoint sequence is sent to the UAV flight platform in the perception layer.
5. The seat belt inspection and spatial positioning method according to claim 2, characterized in that: Step 2 includes the following sub-steps: Step 21: After receiving the optimal inspection route calculated in step 13 from the platform layer, the drone in the perception layer automatically unlocks and takes off and flies autonomously along the inspection route. During the flight, the drone's flight control system collects multi-source data in real time at a preset high frequency. Step 22: The UAV's onboard controller adds a unified timestamp to each frame of video image, each RTK positioning point, and each set of IMU attitude data; Step 23: Encapsulate the acquired multi-source data into a standard data stream according to the principle of time alignment: each frame of image is equipped with its corresponding RTK coordinates and IMU pose; Step 24: The encapsulated standard data stream is pushed to the edge computing layer in real time.
6. The seat belt inspection and spatial positioning method according to claim 5, characterized in that: The multi-source data includes: RTK positioning data: used to provide centimeter-level absolute geographic coordinates for the UAV. IMU attitude data: Synchronously records the yaw angle, pitch angle, roll angle, and angular velocity of each axis of the UAV; Video stream: The telephoto camera continuously captures high-definition images at a fixed frame rate and automatically zooms according to the preset focal length value in the inspection route. In step 2, the obstacle avoidance sensor data is accessed by the system.
7. The seat belt inspection and spatial positioning method according to claim 2, characterized in that: Step 3 includes the following sub-steps: Step 31: First, run a lightweight CNN main network on the edge computing device to perform pre-feature extraction; Step 32: Dynamic gating correction of spatial dimension based on information entropy; Step 33: Time-based remedies.
8. The seat belt inspection and spatial positioning method according to claim 7, characterized in that: The lightweight CNN main network described above adopts the following design: The input image is appropriately scaled and Mosaic data augmentation strategy is applied. An SPD-Conv downsampling module is introduced into the backbone network to preserve fine-grained spatial information of small targets during downsampling; The VSS module is introduced to achieve global feature dependency modeling with linear complexity, enhancing the awareness of long-distance context. The neck area employs a bidirectional, cross-scale dense connection to efficiently integrate shallow details with deep semantics. The detection head is an anchor-free structure, directly outputting the center pixel coordinates of the 2D bounding box of the person target. Width and height And three types of probability distributions: ,in, These represent "personnel", "wearing seat belts", and "not wearing seat belts", respectively. The lightweight CNN main network runs at a high frequency with the video frame rate as the step size, and the inference result of each frame is used as the initial output. Step 32 includes the following sub-steps: Step 321: For each detected target, calculate the information entropy of its probability distribution: ; The value of information entropy The lower the value, the more certain the classification decision of the lightweight CNN main network; the higher the information entropy value, the flatter the probability distribution of the lightweight CNN main network for each category. Step 322: Generate dynamic adaptive weights for the lightweight CNN main network based on the information entropy value: ; in, This is the preset adjustment coefficient; when hour, ;along with Increase It decreased exponentially; Step 323: When When the value falls below a certain preset threshold, the system triggers a small visual language model auxiliary network to assist in reasoning. The input to this small visual language model auxiliary network includes: image patches of the monitored target area cropped from the original image and structured text prompts. The content of the structured text prompts includes the preliminary category output by the lightweight CNN main network, the pose estimation of the monitored target, the occlusion ratio estimation, and the current target construction site scene description. Step 324: After fine-tuning, the small visual language model auxiliary network performs reasoning based on common sense and contextual semantics, and outputs the corrected three-class probability distribution. and the corresponding confidence level; Step 325: The seatbelt status is determined by a weighted fusion of the lightweight CNN main network and the small visual language model auxiliary network. The weighted fusion formula is as follows: ; Step 33 includes the following sub-steps: Step 331: The system maintains the historical state vector of each tracked target. The historical state vector typically contains the position, velocity, size, and rate of change of the detected target on the image plane; state prediction is performed using Kalman filtering. Predicted status: ,in, This is the state transition matrix; Prediction error covariance matrix: ,in, The process noise covariance matrix; the trace of the covariance matrix Reflecting the uncertainty of the current prediction: the longer the target is lost, the greater the prediction variance. Also bigger; Step 332: The edge computing layer uses the trace of the covariance matrix. The cropping range of the region of interest is dynamically adjusted to reflect the uncertainty: ; in, To predict the width and height of the target bounding box, This is a preset magnification factor; that is, the higher the prediction uncertainty, the greater the expansion of the cropping frame. Step 333: The system forcibly extracts the enlarged region of interest from the image and sends it to a small visual language model auxiliary network for "context retrieval reasoning". The small visual language model auxiliary network determines whether there are personnel matching the historical trajectory based on environmental clues within the region of interest and outputs their seat belt wearing status. If it is determined that there is a violation of not wearing a seat belt, an alarm is generated; otherwise, the judgment continues until the detected target reappears or the maximum number of frames lost is exceeded, and then the tracking is terminated.
9. The seat belt inspection and spatial positioning method according to claim 2, characterized in that: Step 4 includes the following sub-steps: Step 41: Solving the spatial microplane penetration under local surface normal vector constraints; Step 42: Spatial topology verification and automatic hierarchical alarm; Step 43: The platform layer pushes the generated alarm events to the application layer in real time, and the application layer performs 3D visualization and closed-loop with the mobile terminal.
10. The seat belt inspection and spatial positioning method according to claim 9, characterized in that: Step 41 includes the following sub-steps: Step 411: The platform layer server receives the pixel coordinates of each detected target. And the corresponding drone time information; Step 412: Construct the ray direction vector of the camera mounted on the UAV; Using the IMU attitude data of the UAV synchronously transmitted in step 2, combined with the pre-calibrated intrinsic parameter matrix of the camera Based on the camera mounting angle, calculate the camera ray direction vector from the UAV's optical center to the world coordinate system of the detected target pixel. ; Step 413: Query the joint constraint matrix; Based on the approximate surface projection location corresponding to the pixel coordinates, the joint constraint matrix pre-stored in step 11 of the platform layer is called. Obtain the absolute elevation of the ground at that location. and local surface normal vector Simultaneously, determine the corresponding three-dimensional surface point at this location. ; Step 414: Construct the equation of the microtangent plane; by Let be a point on the plane, with Using the normal vector, establish the equation of the three-dimensional micro-tangent plane that fits the actual slope of the terrain at that point: ; Step 415: Find the intersection of the ray and the plane; Specifically, the camera ray parameter equations Substituting into the three-dimensional micro-tangent plane equation, in, Location of the drone. Let the propagation parameters be to be determined. , The solution yields: ; After sorting, we get: ; Step 416: Obtain the true 3D coordinates of the target; Specifically, the obtained Substituting back into the camera ray parameter equation, we obtain the precise three-dimensional coordinates of the detected target in the ENU coordinate system: ; Step 417: Timing smoothing; Step 42 includes the following sub-steps: Step 421: Calculate the three-dimensional coordinates obtained in Step 41 Convert back to WGS84 latitude and longitude coordinates ; Step 422: Call the PostGIS database to query which pre-built electronic fence polygons the coordinate point falls within, and obtain the risk level of the corresponding electronic fence; Step 423: Combine the seatbelt status obtained in Step 3 Alarms will be tiered according to the following rules: Level 1 Emergency Alarm: Personnel are located within a high-risk electronic fence and their seatbelt status is "not worn"; at this time, the system immediately triggers an audible and visual alarm, sends a high-priority notification to the relevant safety officer, and coordinates with the drone to perform hovering video recording; Level 2 warning: Personnel are located within a low-risk single-fence area and their safety belt status is "not worn"; at this time, the system only sends a warning notification to remind the safety officer to pay attention. Compliance Records: If personnel are within any electronic fence but are wearing safety belts, the system will only record compliance information and will not trigger an alarm. No event: When personnel are outside all electronic fences, the system only records the personnel's location and does not trigger an alarm. All alarm events, along with their location, time, screenshots of violations, and geofence attributes, are stored in the database to form a traceable security log. In step 43, the 3D visualization platform loads the 3D reality model generated in step 11. After receiving the alarm, it determines the 3D coordinates of the violation location in the 3D reality model. The offender is marked with a prominent dynamic icon, and an information panel pops up to display the snapshot of the offender, the time, the area they belong to, and the risk level. The mobile app / mini-program receives tiered alarm information in real time via cloud messaging services. The tiered alarm information includes the floor, axis number, location description, and geographic coordinates of the person who violated the rules, guiding the on-site safety officer to the designated location quickly. After completing the on-site correction, the safety officer uploads feedback data through the mobile app, forming a closed-loop feedback.
Citation Information
Patent Citations
Improved safety belt detection method
CN106295601A
Construction site unsafe behavior detection method and system based on image recognition
CN120599519A