Unmanned aerial vehicle intelligent inspection method for construction site operation safety
By deploying drones and base stations on the construction site, establishing a time synchronization and spatial coordinate conversion model, and realizing multi-perspective re-shooting and spatiotemporal data alignment, the difficult problems of dynamic occlusion and violation identification in the construction site safety supervision system were solved, and the intelligent level of construction safety and the reliability of responsibility definition were improved.
Patent Information
- Application Number
- CN202510799538.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-16
AI Technical Summary
The existing construction site safety supervision system has difficulty in dynamically sensing changes in construction structures and obstructed areas, and cannot effectively identify violations, making it difficult to define accident responsibility.
By deploying drones, personnel positioning equipment and base stations, a time synchronization network and spatial coordinate conversion model are established to achieve multi-perspective re-shooting and spatiotemporal data alignment, a spatiotemporal graph model is constructed to identify violations and generate safety analysis reports.
It has achieved comprehensive coverage inspections of high-altitude and hidden work areas, enhanced the visual restoration and causal analysis of violation incidents, provided a complete chain of evidence, and provided comprehensive data support for accident responsibility definition and safety warnings.
Smart Images

Figure CN120653015A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drones, and in particular to an intelligent drone inspection method for construction site operation safety. Background Art
[0002] With the acceleration of urbanization and the increasing complexity of building structures, high-risk construction scenarios such as high-altitude construction, deep foundation pits, and areas densely populated with large equipment are proliferating, posing a significant challenge to traditional safety supervision models. The construction industry is a high-incidence area for safety accidents. According to industry statistics, approximately 67% of accidents are caused by the failure to promptly identify hidden dangers or effectively trace violations. The shortcomings of existing technology systems in dynamic risk perception, multi-source evidence chain construction, and three-dimensional traceability have become a core bottleneck hindering the intelligent upgrade of construction safety.
[0003] For example, the existing Chinese patent with publication number CN117369532A discloses a construction site safety inspection drone and intelligent identification method, including a wireless transmission module, a scheduling module, a ground control station, a read-write module, a storage unit and a drone swarm, and also includes a control unit; the drone swarm is electrically connected to the control unit through the wireless transmission module, and the control unit is electrically connected to the scheduling module, involving the field of drone task dispatch and command technology, and the drone swarm is controlled by the wireless transmission module, the scheduling module, the ground control station, the read-write module, and the storage unit, and the corresponding task instructions can be sent to the drones in a timely and accurate manner; the drones equipped with resources that can complete the new task and are in a task-free state are arranged from near to far according to their distance from the task location; the arranged drones are grouped into groups of n and added to the task queue, so as to ensure the rapid completion of the task and achieve efficient scheduling of the drones, thereby improving the efficiency of task completion.
[0004] For example, the existing Chinese patent with publication number CN118833423A discloses an intelligent dam inspection drone and an inspection method, including a body, a detection mechanism is provided on the lower side of the body, a control unit is installed in the body, propellers are installed on the upper sides of four brackets, a floating plate is provided on the upper side of the body, an air bag connected to the mounting hole is provided on the upper side of the floating plate, a rotating frame is rotatably installed between the two mounting plates, an exhaust pipe is connected to the outer side of the rotating frame, an air outlet pipe is connected between the rotating frame and the mounting hole, an air inlet pipe is also connected to the mounting hole, a storage box is filled with sodium peroxide powder, and a water inlet pipe is installed with a water inlet one-way valve. The present invention uses the body to follow the dam body and inspect the dam body laterally and alternately downward. During the inspection process, the height between the body and the lower liquid level is gradually reduced. When a crash occurs, the depth of the body falling into the water can be reduced. At the same time, the gas stored in the air bag can be sprayed into the water to carry the body to swim in the water, so that the body swims towards the shore in the water.
[0005] The shortcomings of the above patents are:
[0006] On the one hand, although monitoring systems are installed in high-altitude dangerous areas on construction sites, such as scaffolding, deep foundation pits, and near tower cranes, these systems rely solely on fixed routes and image recognition, making it difficult to dynamically detect changes in obstructions. Furthermore, drones only fly fixed routes and camera angles. Once the construction structure changes, the system cannot detect new risk points.
[0007] On the other hand, current positioning devices can determine a worker's location and speed, but they lack a unified timestamp and spatial coordinate integration mechanism with drone inspection videos. For example, if a worker climbs without a safety belt, the system only knows they are in a dangerous area, but fails to detect their failure to take safety precautions. Alternatively, if a worker illegally crosses a restricted area, the video captures their movement, but the positioning data isn't immediately linked. Consequently, if a worker encounters a problem, there's no complete chain of evidence for subsequent safety analysis and accountability.
[0008] To this end, the present invention proposes a drone intelligent inspection method for construction site operation safety to solve the above-mentioned problems. Summary of the Invention
[0009] In view of the shortcomings of the existing technology, the present invention provides a drone intelligent inspection method for construction site operation safety to solve the problems raised in the above background technology.
[0010] To achieve the above objectives, the present invention is implemented through the following technical solutions: A drone intelligent inspection method for construction site safety, comprising:
[0011] Step 1: deploy drones, personnel positioning equipment and reference stations, and establish a time synchronization network and spatial coordinate conversion model through the reference stations;
[0012] Step 2: When the drone performs an inspection task, it detects the blocked area in the image in real time based on the preset route, generates a dynamic obstacle avoidance path based on the three-dimensional model, and controls the drone to perform multi-view supplementary shooting;
[0013] Step 3: aligning the multi-view supplementary images with the positioning data collected by the personnel positioning device in time and space, and establishing a mapping relationship between video image coordinates and geographic coordinates;
[0014] Step 4: constructing a spatiotemporal graph model including personnel positions and action features based on the spatiotemporal aligned data, and identifying the violation behavior sequence through the spatiotemporal graph model;
[0015] Step 5: Reconstructing a three-dimensional scene by fusing the multi-view supplementary images, and superimposing the illegal behavior sequence onto the three-dimensional scene to generate a spatiotemporal trajectory of the illegal event;
[0016] Step 6: Generate a security analysis report based on the spatiotemporal trajectory of the violation event, and feed the security analysis report back to the security management platform.
[0017] Preferably, in step 1, deploying drones, personnel positioning equipment and reference stations, and establishing a time synchronization network and a spatial coordinate conversion model through the reference stations further includes:
[0018] Sub-step 1.1: deploy reference stations throughout the construction site, install GNSS receivers and rubidium atomic clock modules, and activate the network timing service of the reference stations;
[0019] Sub-step 1.2: Synchronize the clocks of the drone and the personnel positioning device through the precision clock protocol and calculate the clock offset:
[0020]
[0021] Where Δt is the clock offset, t master is the atomic clock time of the reference station, t slave is the local time of the drone or personnel equipment, d down is the downlink signal transmission distance, d up is the uplink signal transmission distance, c is the speed of light;
[0022] Sub-step 1.3, calibrate the coordinate transformation relationship between the drone camera and the reference station based on the checkerboard calibration plate, and solve the external parameter matrix:
[0023]
[0024] in, is the coordinate transformation matrix, K is the camera intrinsic parameter matrix, R is the rotation matrix, t is the translation vector, is the three-dimensional coordinate of the i-th corner point of the calibration plate, is the image pixel coordinate of the corresponding corner point;
[0025] Sub-step 1.4, verify the spatial coordinate transformation residuals:
[0026]
[0027] Among them, ∈ is the spatial calibration residual, is the coordinate of the jth feature point collected by the lidar, is the true coordinate of the jth feature point measured by the GNSS receiver of the base station, and N is the number of verification points;
[0028] When ∈<∈ max If not, repeat sub-step 1.3.
[0029] Preferably, in step 2, when the drone performs an inspection task, the drone detects the blocked area in the image in real time based on a preset route, generates a dynamic obstacle avoidance path according to the three-dimensional model, and controls the drone to perform multi-view supplementary shooting, further comprising:
[0030] Sub-step 2.1: collect image streams in real time through the camera carried by the drone, input the occlusion detection model to identify the occlusion area, and output the coordinate set of the center of the occlusion area. and confidence scores occ ≥0.7,
[0031] in, is the image horizontal coordinate of the center point of the i-th occluded area, is the image ordinate of the center point of the i-th occluded area;
[0032] Sub-step 2.2: Based on the three-dimensional model and the center coordinates of the occluded area, call the improved RRT* algorithm to generate a dynamic obstacle avoidance path and calculate the path cost function:
[0033]
[0034] Among them, C path is the total cost function, α, β, γ are weight coefficients, d goal is the Euclidean distance from the current path point to the target point, M is the total number of obstacles within the current path planning range, is the shortest distance between the path point and the kth obstacle, v(t) is the linear velocity vector of the UAV at time t, and dt is the discrete time step;
[0035] Sub-step 2.3: Generate drone control instructions based on the dynamic obstacle avoidance path to drive the drone to perform multi-view retake. The control equation is:
[0036]
[0037] Among them, u(t) is the attitude adjustment of the UAV, e p (t) is the position error vector, K p , K d , K i are the proportional, differential and integral gain parameters of the PID controller, τ is the integral variable, and dτ is the differential time element.
[0038] Preferably, in step 3, the multi-view supplementary images are temporally and spatially aligned with the positioning data collected by the personnel positioning device to establish a mapping relationship between video image coordinates and geographic coordinates, further comprising:
[0039] Sub-step 3.1, compensating and interpolating the video frame timestamps of the multi-view retaken images and the timestamps of the personnel positioning device to generate a synchronized time series:
[0040] t align =t raw +Δt cam -Δt sensor ,
[0041] Among them, t align To align the timestamp, t raw is the original data timestamp, Δt cam is the camera acquisition delay, Δt sensor Transmission delay for positioning equipment;
[0042] Sub-step 3.2: Based on the spatial coordinate transformation model of step 1, the pixel coordinates of the multi-view retaken images are mapped to the global coordinate system, and the projection relationship is calculated:
[0043]
[0044] Among them, X global 、Y global 、Z global is the global coordinate, is the coordinate transformation matrix, K is the camera intrinsic parameter matrix, (u, v) is the image pixel coordinate;
[0045] Sub-step 3.3, jointly optimize the personnel positioning data and the mapped image coordinates to minimize the spatiotemporal alignment residual:
[0046]
[0047] in, RTK coordinates of personnel positioning equipment, is the global coordinate, λ is the spatiotemporal weight adjustment factor, is the timestamp of the i-th data point, The original timestamp recorded by the positioning device for the i-th person.
[0048] Preferably, in step 4, a spatiotemporal graph model including personnel positions and action features is constructed based on the spatiotemporal aligned data, and the sequence of illegal behaviors is identified through the spatiotemporal graph model, further comprising:
[0049] Sub-step 4.1, define the nodes and edges of the spatiotemporal graph:
[0050] Node feature vector:
[0051] in, is the feature vector of node v in the space-time graph at time t,
[0052] is the geographic coordinate, is the plane velocity component, is the vertical acceleration;
[0053] Edge feature matrix:
[0054] in, is the edge feature between node u and node v in the space-time graph at time t, is the distance between people, is the velocity vector angle, I interact It is the interactive flag;
[0055] Sub-step 4.2, extract spatiotemporal features through spatiotemporal graph convolutional network:
[0056]
[0057] in, is the feature vector of node v after the l+1th layer of graph convolution, F(v) is the spatiotemporal neighborhood of node v, W spa 、W tem is the spatial / temporal convolution weight matrix, σ is the LeakyReLU activation function, is the feature vector of the neighborhood node u in layer l, is the feature vector of the neighborhood node u in the l-τ layer, d u is the degree of node u;
[0058] Sub-step 4.3, output the violation probability and generate the sequence:
[0059]
[0060] when When it is determined to be a violation,
[0061] in, is the probability of violation at time t, W cls is the classification layer weight matrix, is the feature vector of node v in layer L.
[0062] Preferably, in step 5, fusing the multi-view supplementary images to reconstruct a three-dimensional scene, and superimposing the illegal behavior sequence on the three-dimensional scene to generate a spatiotemporal trajectory of illegal events further includes:
[0063] Sub-step 5.1, performing multi-view stereo reconstruction on the multi-view retaken images and calculating a pixel depth map:
[0064]
[0065] Where D(u,v) is the depth value of the image pixel (u,v), I k is the image data of the k-th retake angle, [u, v, d] T is a three-dimensional point represented by homogeneous coordinates, I ref is the reference view image, T k is the camera pose matrix of the k-th retake perspective, and π(·) is the camera projection function;
[0066] Sub-step 5.2, fuse the multi-view depth maps to generate a dense point cloud and construct a surface mesh:
[0067]
[0068] Among them, M is the surface mesh model, PoissonRecon is the Poisson surface reconstruction algorithm, and D i is the number of depth maps of the i-th perspective, is the divergence operator, is the geographic coordinate, V is the indicator function gradient field, and ρ is the point cloud density constraint;
[0069] Sub-step 5.3, interpolate the illegal behavior sequence into the three-dimensional scene in time and space to generate a four-dimensional trajectory:
[0070]
[0071] Among them, Traj(t) is the four-dimensional space-time trajectory function, Slerp(·) is the spherical linear interpolation operator, is the initial posture state expressed by quaternion, is the end time pose state expressed by quaternion, t start is the starting timestamp of the violation sequence, t end The end timestamp of the violation sequence.
[0072] Preferably, in step 6, generating a security analysis report based on the spatiotemporal trajectory of the violation event and feeding the security analysis report back to the security management platform further includes:
[0073] Sub-step 6.1: Perform causal reasoning analysis on the spatiotemporal trajectory of the violation event and construct a dynamic Bayesian network to calculate the root cause probability:
[0074]
[0075] Among them, P(C root |Traj) is the root cause C under the condition of the given spatiotemporal trajectory Traj of the violation event root The posterior probability of |, C root As the potential root cause, E iis the evidence node, C is the root cause set, Pa(E i ) is the evidence node E in the Bayesian network i The parent node set of P(C root ) is the root cause C root The prior probability of P(c) is the probability of root cause c in the root cause set C, and P(E i |Pa(E i )) is the parent node Pa(E i ) occurs under the condition that the evidence node E i The conditional probability of
[0076] Sub-step 6.2: Generate a structured safety report based on the root cause probability and define the report content template:
[0077] R={VID,t start ,t end ,Type,P risk ,SceneSnapshot},
[0078] Among them, VID is the unique identifier of the violation event, P risk For the highest risk probability,
[0079] SceneSnapshot is the keyframe index of the 3D scene, and Type is the violation type identifier.
[0080] Sub-step 6.3, update the violation detection model parameters through the federated learning framework:
[0081]
[0082] Among them, L m is the local loss function of the mth construction site, λ is the model aggregation regularization coefficient, is the global model parameter after k+1 rounds of federated learning, η is the learning rate, is the global model parameter after k rounds of federated learning, is the local model parameter of the mth construction site, is the gradient of the local loss function at the mth construction site.
[0083] Preferably, the personnel positioning device in step 1 is a smart helmet integrating dual-frequency RTK, UWB and IMU, the positioning data update frequency is not less than 10Hz, and the horizontal positioning accuracy is better than 1 cm;
[0084] In step 2, the occlusion detection adopts an improved YOLOv8 network, adds an occlusion feature branch to the detection head, and outputs a binary mask of the occlusion area and a confidence score.
[0085] A terminal device includes a processor and a memory, wherein the memory stores a computer program. When the program is executed by the processor, an intelligent drone inspection method for construction site operation safety is implemented.
[0086] A storage medium stores a computer program, which, when executed by a processor, implements a drone intelligent inspection method for construction site operation safety.
[0087] The present invention provides a drone intelligent inspection method for construction site safety. It has the following beneficial effects:
[0088] 1. The present invention adopts dynamic occlusion perception and multi-view adaptive re-shooting technology solutions. Through the high-precision spatial coordinate system and real-time occlusion detection mechanism constructed by the base station, combined with dynamic obstacle avoidance path planning, it can achieve comprehensive coverage inspection of high-altitude and hidden work areas. Compared with the static detection solution that relies on fixed routes in the existing technology, it solves the problem of missed detection caused by dynamic occlusion of construction structures, effectively eliminates the misjudgment of safety hazards caused by blind spots, and significantly improves the monitoring capability of high-risk areas deep in scaffolding and under cantilevered platforms, avoiding the false safety state of "detection completed but risks not identified".
[0089] 2. The present invention adopts a multi-source data spatiotemporal fusion and behavior trajectory chain modeling technology solution. Through a unified time reference and spatial coordinate conversion model, it deeply integrates the drone video stream and personnel positioning data, and constructs a spatiotemporal graph model to analyze personnel interaction behavior, thereby realizing the visualization restoration and causal analysis of the entire process of violation events. Compared with the existing technology of isolated detection of static violation actions, it solves the problem of incomplete evidence chain caused by the separation of operator behavior and positioning data, significantly enhances the ability to trace violation paths, time periods and collaborative behaviors, and provides comprehensive and coherent data support for accident responsibility definition and safety warning. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] Figure 1 Flowchart of the present invention. DETAILED DESCRIPTION
[0091] To help those skilled in the art understand the present invention, the following will provide a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only partial embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0092] The present invention is described in detail below with reference to the accompanying drawings:
[0093] Example:
[0094] Please see the attached Figure 1 The embodiment of the present invention provides a drone intelligent inspection method for construction site safety, comprising:
[0095] Step 1: Deploy drones, personnel positioning equipment, and base stations, and establish a time synchronization network and spatial coordinate conversion model through the base stations;
[0096] Sub-step 1.1: Deploy base stations throughout the construction site, install GNSS receivers and rubidium atomic clock modules, and activate the network timing service of the base stations.
[0097] Sub-step 1.2: Synchronize the clocks of the drone and the personnel positioning device through the precision clock protocol and calculate the clock offset:
[0098]
[0099] Where Δt is the clock offset, t master is the atomic clock time of the reference station, t slave is the local time of the drone or personnel equipment, d down is the downlink signal transmission distance, d up is the uplink signal transmission distance, c is the speed of light;
[0100] Sub-step 1.3, calibrate the coordinate transformation relationship between the drone camera and the reference station based on the checkerboard calibration plate, and solve the external parameter matrix:
[0101]
[0102] in, is the coordinate transformation matrix, K is the camera intrinsic parameter matrix, R is the rotation matrix, t is the translation vector, is the three-dimensional coordinate of the i-th corner point of the calibration plate, is the image pixel coordinate of the corresponding corner point;
[0103] Sub-step 1.4, verify the spatial coordinate transformation residuals:
[0104]
[0105] Among them, ∈ is the spatial calibration residual, is the coordinate of the jth feature point collected by the lidar, is the true coordinate of the jth feature point measured by the GNSS receiver of the base station, and N is the number of verification points;
[0106] When ∈<∈ max If the calibration is completed, re-execute sub-step 1.3;
[0107] Step 2: When the drone performs an inspection mission, it detects the blocked areas in the image in real time based on the preset route, generates a dynamic obstacle avoidance path based on the 3D model, and controls the drone to take supplementary photos from multiple perspectives.
[0108] Sub-step 2.1: Use the camera on the drone to collect image streams in real time, input the occlusion detection model to identify the occlusion area, and output the coordinate set of the center of the occlusion area. and confidence scores occ ≥0.7,
[0109] in, is the image horizontal coordinate of the center point of the i-th occluded area, is the image ordinate of the center point of the i-th occluded area;
[0110] In sub-step 2.2, based on the 3D model and the center coordinates of the occluded area, the improved RRT* algorithm is used to generate a dynamic obstacle avoidance path and calculate the path cost function:
[0111]
[0112] Among them, C path is the total cost function, α, β, γ are weight coefficients, d goal is the Euclidean distance from the current path point to the target point, M is the total number of obstacles within the current path planning range, is the shortest distance between the path point and the kth obstacle, v(t) is the linear velocity vector of the UAV at time t, and dt is the discrete time step;
[0113] In sub-step 2.3, the drone control instructions are generated based on the dynamic obstacle avoidance path to drive the drone to perform multi-view retakes. The control equation is:
[0114]
[0115] Among them, u(t) is the attitude adjustment of the UAV, e p (t) is the position error vector, K p , K d , K i are the proportional, differential, and integral gain parameters of the PID controller, τ is the integral variable, and dτ is the differential time element;
[0116] Step 3: Temporally and spatially align the multi-view supplementary images with the positioning data collected by the personnel positioning device to establish a mapping relationship between the video image coordinates and the geographic coordinates;
[0117] Sub-step 3.1: Compensate and interpolate the video frame timestamps of the multi-view supplementary images and the timestamps of the personnel positioning device to generate a synchronized time series:
[0118] t align=t raw +Δt cam -Δt sensor ,
[0119] Among them, t align To align the timestamp, t raw is the original data timestamp, Δt cam is the camera acquisition delay, Δt sensor Transmission delay for positioning equipment;
[0120] Sub-step 3.2: Based on the spatial coordinate transformation model in step 1, map the pixel coordinates of the multi-view supplementary images to the global coordinate system and calculate the projection relationship:
[0121]
[0122] Among them, X global 、Y global 、Z global is the global coordinate, is the coordinate transformation matrix, K is the camera intrinsic parameter matrix, (u, v) is the image pixel coordinate;
[0123] Sub-step 3.3, jointly optimize the person positioning data and the mapped image coordinates to minimize the spatiotemporal alignment residual:
[0124]
[0125] in, RTK coordinates of personnel positioning equipment, is the global coordinate, λ is the spatiotemporal weight adjustment factor, is the timestamp of the i-th data point, The original timestamp recorded by the positioning device for the i-th person;
[0126] Step 4: Based on the spatiotemporal aligned data, a spatiotemporal graph model containing the location and action features of the personnel is constructed, and the sequence of illegal behaviors is identified through the spatiotemporal graph model;
[0127] Sub-step 4.1, define the nodes and edges of the spatiotemporal graph:
[0128] Node feature vector:
[0129] in, is the feature vector of node v in the space-time graph at time t,
[0130] is the geographic coordinate, is the plane velocity component, is the vertical acceleration;
[0131] Edge feature matrix:
[0132] in, is the edge feature between node u and node v in the space-time graph at time t, is the distance between people, is the velocity vector angle, I interact It is the interactive flag;
[0133] Sub-step 4.2, extract spatiotemporal features through spatiotemporal graph convolutional network:
[0134]
[0135] in, is the feature vector of node v after the l+1th layer of graph convolution, F(v) is the spatiotemporal neighborhood of node v, W spa 、W tem is the spatial / temporal convolution weight matrix, σ is the LeakyReLU activation function, is the feature vector of the neighborhood node u in layer l, is the feature vector of the neighborhood node u in the l-τ layer, d u is the degree of node u;
[0136] Sub-step 4.3, output the violation probability and generate the sequence:
[0137]
[0138] when When it is determined to be a violation,
[0139] in, is the probability of violation at time t, W cls is the classification layer weight matrix, is the feature vector of node v in layer L;
[0140] Step 5: Reconstruct the 3D scene by fusing the multi-view supplementary images, and superimpose the violation sequence onto the 3D scene to generate the spatiotemporal trajectory of the violation event;
[0141] Sub-step 5.1, perform multi-view stereo reconstruction on the multi-view retaken images and calculate the pixel depth map:
[0142]
[0143] Where D(u,v) is the depth value of the image pixel (u,v), I k is the image data of the k-th retake angle, [u, v, d] T is a three-dimensional point represented by homogeneous coordinates, I ref is the reference view image, T kis the camera pose matrix of the k-th retake perspective, and π(·) is the camera projection function;
[0144] Sub-step 5.2, fuse the multi-view depth maps to generate a dense point cloud and construct a surface mesh:
[0145]
[0146] Among them, M is the surface mesh model, PoissonRecon is the Poisson surface reconstruction algorithm, and D i is the number of depth maps of the i-th perspective, is the divergence operator, is the geographic coordinate, V is the indicator function gradient field, and ρ is the point cloud density constraint;
[0147] Sub-step 5.3, interpolate the violation sequence into the three-dimensional scene to generate a four-dimensional trajectory:
[0148]
[0149] Among them, Traj(t) is the four-dimensional space-time trajectory function, Slerp(·) is the spherical linear interpolation operator, is the initial posture state expressed by quaternion, is the end time pose state expressed by quaternion, t start is the starting timestamp of the violation sequence, t end is the end timestamp of the violation sequence;
[0150] Step 6: Generate a security analysis report based on the spatiotemporal trajectory of the violation event and feed the security analysis report back to the security management platform;
[0151] Sub-step 6.1: Perform causal reasoning analysis on the spatiotemporal trajectory of the violation event and construct a dynamic Bayesian network to calculate the root cause probability:
[0152]
[0153] Among them, P(C root |Traj) is the root cause C under the condition of the given spatiotemporal trajectory Traj of the violation event root The posterior probability of |, C root As the potential root cause, E i is the evidence node, C is the root cause set, Pa(E i ) is the evidence node E in the Bayesian network i The parent node set of P(C root ) is the root cause C root The prior probability of P(c) is the probability of root cause c in the root cause set C, and P(E i |Pa(E i)) is the parent node Pa(E i ) occurs under the condition that the evidence node E i The conditional probability of
[0154] Sub-step 6.2: Generate a structured safety report based on the root cause probability and define the report content template:
[0155] R={VID,t start ,t end ,Type,P risk ,SceneSnapshot},
[0156] Among them, VID is the unique identifier of the violation event, P risk For the highest risk probability,
[0157] SceneSnapshot is the keyframe index of the 3D scene, and Type is the violation type identifier.
[0158] Sub-step 6.3, update the violation detection model parameters through the federated learning framework:
[0159]
[0160] Among them, L m is the local loss function of the mth construction site, λ is the model aggregation regularization coefficient, is the global model parameter after k+1 rounds of federated learning, η is the learning rate, is the global model parameter after k rounds of federated learning, is the local model parameter of the mth construction site, is the gradient of the local loss function at the mth construction site.
[0161] Step 1: By building a high-precision spatiotemporal reference network, a unified spatiotemporal mapping between drones, personnel, and equipment, and the geographic coordinate system, is achieved. The distributed deployment of reference stations and the atomic clock timing mechanism eliminate clock drift and spatial reference deviation between multiple devices, providing a millimeter-level precision spatiotemporal foundation for subsequent dynamic inspections and data fusion. A closed-loop calibration process combining checkerboard calibration and lidar verification ensures a highly reliable conversion between the camera coordinate system and the global geographic coordinate system, effectively resolving image positioning distortion caused by calibration errors in traditional solutions and establishing a precise spatial framework for 3D reconstruction of occluded areas and restoration of behavioral trajectories.
[0162] Step 2 utilizes dynamic perception and adaptive path planning technology to significantly enhance proactive inspection capabilities in complex construction environments. A real-time occlusion detection model, combined with an improved path search algorithm, enables the drone to intelligently identify dynamic blind spots obstructed by scaffolding meshes and equipment, generating an optimized trajectory that balances obstacle avoidance and visual coverage. A multi-perspective recapture mechanism proactively adjusts flight attitude and shooting angle to ensure comprehensive coverage of high-risk areas such as overhanging structures and equipment gaps. This overcomes the bottleneck of missed inspections caused by a single viewing angle in traditional fixed-route models, comprehensively improving the reliability and timeliness of hidden danger detection.
[0163] Step 3 uses multi-source data spatiotemporal alignment technology to construct a pixel-level cross-modal mapping system. A bidirectional delay compensation mechanism eliminates temporal misalignment between video streams and positioning data. A spatial projection model accurately maps image pixel coordinates to a geographic coordinate system. A joint optimization algorithm further corrects data residuals, forming a temporally and spatially consistent multimodal data pool. This step overcomes the limitations of traditional methods of isolated analysis of video and positioning data, providing a seamless data foundation for collaborative modeling of human behavior and spatial scenarios, supporting the full-factor correlation and tracing of violations.
[0164] The spatiotemporal graph model constructed in step 4 deeply integrates multidimensional dynamic features, enabling precise capture and serialized analysis of complex violations. Node feature vectors integrate individual position, velocity, and acceleration information, while edge feature matrices encode group interactions. The spatiotemporal graph convolutional network, by extracting features from both the spatiotemporal and temporal domains, overcomes the fragmented limitations of traditional single-frame image analysis. This model effectively identifies continuous, high-risk actions such as climbing, lingering, and illegal collaboration, and constructs a complete "cause-process-result" behavioral evidence chain, significantly enhancing the intelligence of safety supervision and the interpretability of violation determinations.
[0165] Step 5 uses multi-view 3D reconstruction and spatiotemporal trajectory fusion technology to achieve a holographic visualization of the violation. The 3D reconstruction algorithm, combined with surface mesh optimization, accurately restores the three-dimensional form of the complex structures of scaffolding and equipment gaps. Four-dimensional trajectory interpolation technology dynamically correlates the violation sequence with the 3D scene, generating a trajectory map that integrates time, space, and posture. This step transcends the limitations of two-dimensional planar trajectory representation and supports tracing the spatial evolution of violations from any perspective and time point, providing an intuitive, multi-dimensional decision-making basis for accident responsibility determination and safety strategy optimization.
[0166] Step 6 builds a closed-loop intelligent analysis system to empower the entire process, from data collection to management decision-making. A dynamic causal inference model, combined with a domain knowledge base, accurately locates the potential root causes of violations. A federated learning framework collaboratively optimizes model parameters through data from multiple construction sites, continuously improving the system's ability to generalize to new violation patterns. Structured report templates integrate 3D scene snapshots with behavioral trajectory chains to generate safety analysis files with strong evidentiary power. This step completely revolutionizes the inefficient traditional manual tracing model, forming a self-evolving, verifiable intelligent safety management ecosystem and driving construction safety management and control to a higher level of proactive prevention and precise intervention.
[0167] A terminal device includes a processor and a memory. The memory stores a computer program. When the program is executed by the processor, an unmanned aerial vehicle (UAV) intelligent inspection method for construction site operation safety is implemented.
[0168] A storage medium stores a computer program, which, when executed by a processor, implements an intelligent drone inspection method for construction site safety.
[0169] By integrating high-computing processors and customized storage architectures, the terminal device enables real-time edge computing and multi-source data fusion processing for construction site safety inspection tasks. The device's built-in spatiotemporal synchronization module seamlessly connects to the base station network, ensuring that data streams from heterogeneous devices such as drones and personnel positioning terminals are collaboratively analyzed with nanosecond time errors and centimeter-level spatial accuracy. Dynamic obstacle avoidance path planning and multi-view image reshooting algorithms run directly on the device side, breaking through the latency bottleneck of traditional solutions that rely on cloud computing, compressing the response time for blocked area detection to seconds, and meeting the real-time safety warning needs of complex scenarios such as high-altitude operations and large-scale machinery-intensive areas. The edge computing architecture further reduces network dependence, ensuring that violation trajectories and three-dimensional reconstruction data can continue to be generated even in weak network environments on construction sites, building all-weather, all-terrain autonomous inspection capabilities.
[0170] The storage medium solidifies the program code and pre-trained model parameters of the drone's intelligent inspection algorithm to form a security management knowledge carrier that can be deployed across platforms. The medium's embedded multi-source data spatiotemporal alignment engine and federated learning interface support the secure migration of core data assets such as three-dimensional reconstruction models and violation feature libraries to different construction site management systems, enabling rapid replication and collaborative optimization of inspection experience. The program code uses containerized packaging technology, which can be flexibly adapted to various edge computing devices and cloud servers to ensure synchronous processing of full-link data from the drone flight control terminal to the headquarters management platform. The storage medium has version traceability and incremental update functions, and continuously improves the accuracy of occlusion recognition and the completeness of trajectory restoration by remotely pushing algorithm iteration packages, forming a self-evolving intelligent inspection ecosystem.
[0171] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A drone intelligent inspection method for construction site safety, characterized in that: include: Step 1: deploy drones, personnel positioning equipment and reference stations, and establish a time synchronization network and spatial coordinate conversion model through the reference stations; Step 2: When the drone performs an inspection task, it detects the blocked area in the image in real time based on the preset route, generates a dynamic obstacle avoidance path based on the three-dimensional model, and controls the drone to perform multi-view supplementary shooting; Step 3: aligning the multi-view supplementary images with the positioning data collected by the personnel positioning device in time and space, and establishing a mapping relationship between video image coordinates and geographic coordinates; Step 4: constructing a spatiotemporal graph model including personnel positions and action features based on the spatiotemporal aligned data, and identifying the violation behavior sequence through the spatiotemporal graph model; Step 5: Reconstructing a three-dimensional scene by fusing the multi-view supplementary images, and superimposing the illegal behavior sequence onto the three-dimensional scene to generate a spatiotemporal trajectory of the illegal event; Step 6: Generate a security analysis report based on the spatiotemporal trajectory of the violation event, and feed the security analysis report back to the security management platform.
2. The method for intelligent inspection of construction site operations by drones according to claim 1, characterized in that: In step 1, deploying drones, personnel positioning equipment, and reference stations, and establishing a time synchronization network and a spatial coordinate conversion model through the reference stations further includes: Sub-step 1.1: deploy reference stations throughout the construction site, install GNSS receivers and rubidium atomic clock modules, and activate the network timing service of the reference stations; Sub-step 1.2: Synchronize the clocks of the drone and the personnel positioning device through the precision clock protocol and calculate the clock offset: Where Δt is the clock offset, t master is the atomic clock time of the reference station, t slave is the local time of the drone or personnel equipment, d down is the downlink signal transmission distance, d up is the uplink signal transmission distance, c is the speed of light; Sub-step 1.3, calibrate the coordinate transformation relationship between the drone camera and the reference station based on the checkerboard calibration plate, and solve the external parameter matrix: in, is the coordinate transformation matrix, K is the camera intrinsic parameter matrix, R is the rotation matrix, t is the translation vector, is the three-dimensional coordinate of the i-th corner point of the calibration plate, is the image pixel coordinate of the corresponding corner point; Sub-step 1.4, verify the spatial coordinate transformation residuals: Among them, ∈ is the spatial calibration residual, is the coordinate of the jth feature point collected by the lidar, is the true coordinate of the jth feature point measured by the GNSS receiver of the base station, and N is the number of verification points; When ∈<∈ max If not, repeat sub-step 1.
3.
3. The method for intelligent inspection of construction site operations by drones according to claim 1, characterized in that: In step 2, when the drone performs an inspection task, the drone detects the blocked area in the image in real time based on the preset route, generates a dynamic obstacle avoidance path according to the three-dimensional model, and controls the drone to perform multi-view supplementary shooting, further comprising: Sub-step 2.1: collect image streams in real time through the camera carried by the drone, input the occlusion detection model to identify the occlusion area, and output the coordinate set of the center of the occlusion area. and confidence scores occ ≥0.7, in, is the image horizontal coordinate of the center point of the i-th occluded area, is the image ordinate of the center point of the i-th occluded area; Sub-step 2.2: Based on the three-dimensional model and the center coordinates of the occluded area, call the improved RRT* algorithm to generate a dynamic obstacle avoidance path and calculate the path cost function: Among them, C path is the total cost function, α, β, γ are weight coefficients, d goal is the Euclidean distance from the current path point to the target point, M is the total number of obstacles within the current path planning range, is the shortest distance between the path point and the kth obstacle, v(t) is the linear velocity vector of the UAV at time t, and dt is the discrete time step; Sub-step 2.3: Generate drone control instructions based on the dynamic obstacle avoidance path to drive the drone to perform multi-view retake. The control equation is: Among them, u(t) is the attitude adjustment of the UAV, e p (t) is the position error vector, K p , K d , K i are the proportional, differential and integral gain parameters of the PID controller, τ is the integral variable, and dτ is the differential time element.
4. The method for intelligent inspection of construction site operations by drones according to claim 1, characterized in that: In step 3, the multi-view supplementary images are temporally and spatially aligned with the positioning data collected by the personnel positioning device to establish a mapping relationship between video image coordinates and geographic coordinates, further comprising: Sub-step 3.1, compensating and interpolating the video frame timestamps of the multi-view retaken images and the timestamps of the personnel positioning device to generate a synchronized time series: t align =t raw +Δt cam -Δt sensor , Among them, t align To align the timestamp, t raw is the original data timestamp, Δt cam is the camera acquisition delay, Δt sensor Transmission delay for positioning equipment; Sub-step 3.2: Based on the spatial coordinate transformation model of step 1, the pixel coordinates of the multi-view retaken images are mapped to the global coordinate system, and the projection relationship is calculated: Among them, X global 、Y global 、Z global is the global coordinate, is the coordinate transformation matrix, K is the camera intrinsic parameter matrix, (u, v) is the image pixel coordinate; Sub-step 3.3, jointly optimize the personnel positioning data and the mapped image coordinates to minimize the spatiotemporal alignment residual: in, RTK coordinates of personnel positioning equipment, is the global coordinate, λ is the spatiotemporal weight adjustment factor, is the timestamp of the i-th data point, The original timestamp recorded by the positioning device for the i-th person.
5. The method for intelligent inspection of construction site operations by unmanned aerial vehicles according to claim 1, characterized in that: In step 4, a spatiotemporal graph model including personnel positions and action features is constructed based on the spatiotemporal aligned data, and a sequence of illegal behaviors is identified through the spatiotemporal graph model, further comprising: Sub-step 4.1, define the nodes and edges of the spatiotemporal graph: Node feature vector: in, is the feature vector of node v in the space-time graph at time t, is the geographic coordinate, is the plane velocity component, is the vertical acceleration; Edge feature matrix: in, is the edge feature between node u and node v in the space-time graph at time t, is the distance between people, is the velocity vector angle, I interact It is the interactive flag; Sub-step 4.2, extract spatiotemporal features through spatiotemporal graph convolutional network: in, is the feature vector of node v after the l+1th layer of graph convolution, F(v) is the spatiotemporal neighborhood of node v, W spa 、W tem is the spatial / temporal convolution weight matrix, σ is the LeakyReLU activation function, is the feature vector of the neighborhood node u in layer l, is the feature vector of the neighborhood node u in the l-τ layer, d u is the degree of node u; Sub-step 4.3, output the violation probability and generate the sequence: when When it is determined to be a violation, in, is the probability of violation at time t, W cls is the classification layer weight matrix, is the feature vector of node v in layer L.
6. The method for intelligent inspection of construction site operations by drones according to claim 1, characterized in that: In step 5, fusing the multi-view supplementary images to reconstruct a three-dimensional scene, and superimposing the illegal behavior sequence onto the three-dimensional scene to generate a spatiotemporal trajectory of illegal events, further comprising: Sub-step 5.1, performing multi-view stereo reconstruction on the multi-view retaken images and calculating a pixel depth map: Where D(u,v) is the depth value of the image pixel (u,v), I k is the image data of the k-th retake angle, [u, v, d] T is a three-dimensional point represented by homogeneous coordinates, I ref is the reference view image, T k is the camera pose matrix of the k-th retake perspective, and π(·) is the camera projection function; Sub-step 5.2, fuse the multi-view depth maps to generate a dense point cloud and construct a surface mesh: Among them, M is the surface mesh model, PoissonRecon is the Poisson surface reconstruction algorithm, and D i is the number of depth maps of the i-th perspective, is the divergence operator, is the geographic coordinate, V is the indicator function gradient field, and ρ is the point cloud density constraint; Sub-step 5.3, interpolate the illegal behavior sequence into the three-dimensional scene in time and space to generate a four-dimensional trajectory: Among them, Traj(t) is the four-dimensional space-time trajectory function, Slerp(·) is the spherical linear interpolation operator, is the initial posture state expressed by quaternion, is the end time pose state expressed by quaternion, t start is the starting timestamp of the violation sequence, t end The end timestamp of the violation sequence.
7. The method for intelligent inspection of construction site operations by drones according to claim 1, characterized in that: In step 6, generating a security analysis report based on the spatiotemporal trajectory of the violation event and feeding the security analysis report back to the security management platform further includes: Sub-step 6.1: Perform causal reasoning analysis on the spatiotemporal trajectory of the violation event and construct a dynamic Bayesian network to calculate the root cause probability: Among them, P(C root |Traj) is the root cause C under the condition of the given spatiotemporal trajectory Traj of the violation event root The posterior probability of |, C root As the potential root cause, E i is the evidence node, C is the root cause set, Pa(E i ) is the evidence node E in the Bayesian network i The parent node set of P(C root ) is the root cause C root The prior probability of P(c) is the probability of root cause c in the root cause set C, and P(E i |Pa(E i )) is the parent node Pa(E i ) occurs under the condition that the evidence node E i The conditional probability of Sub-step 6.2: Generate a structured safety report based on the root cause probability and define the report content template: R={VID,t start ,t end ,Type,P risk ,SceneSnapshot}, Among them, VID is the unique identifier of the violation event, P risk For the highest risk probability, SceneSnapshot is the keyframe index of the 3D scene, and Type is the violation type identifier. Sub-step 6.3, update the violation detection model parameters through the federated learning framework: Among them, L m is the local loss function of the mth construction site, λ is the model aggregation regularization coefficient, is the global model parameter after k+1 rounds of federated learning, η is the learning rate, is the global model parameter after k rounds of federated learning, is the local model parameter of the mth construction site, is the gradient of the local loss function at the mth construction site.
8. The method for intelligent inspection of construction site operations by using a drone according to claim 1, characterized in that: The personnel positioning device in step 1 is a smart helmet that integrates dual-frequency RTK, UWB and IMU, with a positioning data update frequency of no less than 10Hz and a horizontal positioning accuracy better than 1 cm; In step 2, the occlusion detection adopts an improved YOLOv8 network, adds an occlusion feature branch to the detection head, and outputs a binary mask of the occlusion area and a confidence score.
9. A terminal device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores a computer program, and when the program is executed by the processor, an unmanned aerial vehicle intelligent inspection method for construction site operation safety as described in any one of claims 1 to 8 is implemented.
10. A storage medium, characterized in that: A computer program is stored, and when the program is executed by a processor, an unmanned aerial vehicle intelligent inspection method for construction site operation safety as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Construction site operation safety inspection unmanned aerial vehicle and intelligent identification method
CN117369532A
Dam intelligent inspection unmanned aerial vehicle and inspection method
CN118833423A
Intelligent construction site management system and management method suitable for epidemic situation period
CN111383132A
Human body abnormal behavior recognition method under transformer substation video monitoring based on posture estimation
CN116912930A
Image rendering method and device and electronic equipment
CN117576283A
Cited By
Multi-mode full-autonomous inspection method and system for electric unmanned aerial vehicle
CN120909339A
Multi-device cooperative control system and method for intelligent capsule bin of construction site
CN121115647A
Unmanned aerial vehicle inspection and AI potential safety hazard identification method for construction site
CN121482658A
Unmanned aerial vehicle inspection and ai security risk identification method for construction site
CN121482658B
Intelligent agent hidden maneuvering method, system and device and readable storage medium
CN121680453A