Unmanned aerial vehicle inspection and ai security risk identification method for construction site

By fine-tuning the pre-trained model and generating a construction site-specific safety hazard identification model, and combining image data to calculate the three-dimensional spatial location, the problem of poor model adaptability of the UAV inspection system on construction sites was solved. This enabled accurate positioning and trend analysis of safety hazards, improving identification accuracy and response efficiency.

CN121482658BActive Publication Date: 2026-04-10重庆爵木建设集团有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing drone inspection systems on construction sites suffer from poor model adaptability, high false alarm and false alarm rates, inability to respond intelligently, and lack of spatial and temporal correlation in identifying safety hazards, making it difficult to perform accurate positioning and trend analysis.

Method used

By fine-tuning the pre-trained model using a specific labeled dataset from the construction site, a construction site-specific safety hazard identification model is generated. Based on periodic full-domain coverage and high-risk area information, drone flight paths are generated to identify and dynamically execute inspection tasks in real time. Combined with image data, three-dimensional spatial location is calculated to generate a four-dimensional hazard event data package.

Benefits of technology

It enables precise location and trend analysis of safety hazards at construction sites, improving identification accuracy and response efficiency, and facilitating efficient handling by management personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482658B_ABST
    Figure CN121482658B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of safety identification, and in particular to a kind of unmanned aerial vehicle inspection and AI safety hazard identification method for construction site, first, through the self-learning mechanism of site, using active learning screening and manual annotation, quickly generate the exclusive AI identification model adapted to specific construction site.Subsequently, in the intelligent inspection stage, unmanned aerial vehicle is based on the dynamic route of fusion historical data and construction plan flight, and runs exclusive model to carry out real-time identification;When finding hidden danger, automatically adjust the path to carry out fine investigation and evidence collection, and generate four-dimensional hidden danger warning work order containing accurate three-dimensional position and time through coordinate mapping.Significantly improve the accuracy, positioning accuracy and system adaptive capacity of safety hazard identification, facilitate efficient disposal for management personnel.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of safety identification, in particular to a UAV inspection and AI safety hazard identification method for construction sites. BACKGROUND

[0002] With the advancement of smart construction site construction, using UAVs for automatic inspection has become an important means to improve the efficiency of construction safety management. The existing technology usually uses a UAV that pre-plans a flight path to collect site images or videos, and uses a visual model based on deep learning to analyze the images offline to identify typical safety hazards such as not wearing a safety helmet, not wearing a safety rope, fire smoke, and lack of edge protection.

[0003] However, the existing technical solutions have obvious deficiencies: first, the AI identification model is usually trained on a general data set or limited historical data, and the environment, layout, and construction phase of different construction sites differ greatly, resulting in a significant decrease in the identification accuracy and adaptability of the model in new scenarios, with high false positive and false negative rates. Second, the traditional inspection mode is an open-loop process that separates "collection" and "analysis", and the UAV is only used as a data collection tool and cannot respond intelligently based on real-time analysis results. Third, the identified safety hazards are usually isolated static image events that lack association with the three-dimensional spatial location and time dimension of the construction site, making it difficult to accurately locate and analyze trends, and making it difficult for management personnel to efficiently handle them. SUMMARY

[0004] The purpose of the present application is to provide a UAV inspection and AI safety hazard identification method for construction sites, which accurately locates and analyzes trends to facilitate efficient handling by management personnel.

[0005] To achieve the above-mentioned purpose, the present application provides a UAV inspection and AI safety hazard identification method for construction sites, comprising the following steps:

[0006] Obtain an initial image set of the target construction site, label the key sample set selected from the initial image set, and fine-tune the pre-trained general safety hazard identification model using the generated construction site-specific labeled data set to generate a construction site-exclusive safety hazard identification model;

[0007] Based on the periodic global coverage requirement, historical hazard probability heat map, and current high-risk operation area information, generate a main flight path and at least one inspection sub-task for the UAV;

[0008] The video stream collected by the unmanned aerial vehicle flying along the main route in real time is identified by using the construction site specific safety hazard identification model, and when an abnormal target meeting the preset condition is identified, the unmanned aerial vehicle is controlled to dynamically perform the corresponding inspection subtask to obtain image data of the target.

[0009] Based on the image data and the unmanned aerial vehicle pose data, the actual position of the target in the three-dimensional space of the construction site is calculated, a four-dimensional hazard event data packet is generated, and a warning work order is generated.

[0010] The method further comprises:

[0011] Receiving on-site verification feedback data for the four-dimensional hazard event data packet, and based on the on-site verification feedback data, incrementally training the construction site specific safety hazard identification model.

[0012] The method further comprises:

[0013] The initial image set is identified by using the general safety hazard identification model to obtain an initial identification result containing target categories, bounding boxes, and initial confidence.

[0014] Images of identified targets with initial confidence in a preset fuzzy interval are filtered as first-class candidate samples.

[0015] Images containing multiple dense targets, or targets with a background contrast lower than a preset threshold, are filtered as second-class candidate samples.

[0016] The first-class candidate samples and the second-class candidate samples are merged to form the set of key samples to be verified after deduplication processing.

[0017] The method further comprises:

[0018] The feature extraction network parameters of the general safety hazard identification model are frozen, and only the classification and regression network heads are unfrozen and trained, and the construction site specific annotation dataset is used for training.

[0019] The feature extraction network parameters within a set output range are gradually unfrozen, and different learning rates are applied to different layer groups, and false positives and false negatives generated during the annotation process are used for collaborative training to generate the construction site specific safety hazard identification model.

[0020] The method further comprises:

[0021] generate a main route based on a preset grid route;

[0022] generate the historical hidden danger probability heat map according to historical hidden danger data, and determine a region center point with a heat value exceeding a first threshold value as a first type of dynamic interest point;

[0023] obtain current high-risk operation region information, and determine a region center point as a second type of dynamic interest point;

[0024] generate an inspection sub-task package for the first type of dynamic interest point with a distance greater than a preset offset threshold value from the main route, or the second type of dynamic interest point reaching a set priority.

[0025] The method further comprises:

[0026] insert the first type of dynamic interest point with a distance less than a preset offset threshold value from the main route into the main route through a path smoothing algorithm.

[0027] When an abnormal target meeting preset conditions is identified, the UAV is controlled to dynamically perform a corresponding inspection sub-task to obtain target image data, including:

[0028] When a target recognition confidence output by the construction site-specific safety hidden danger recognition model exceeds a preset second threshold value, and a target risk category reaches a set category, the abnormal target is determined to meet the preset conditions;

[0029] According to a current pose of the UAV and a position of the abnormal target in an image, a geographical region of the abnormal target is calculated.

[0030] An instant exploration sub-task is generated with the geographical region as a center.

[0031] The UAV is controlled to pause execution of the main route, switch to execution of the instant exploration sub-task, and return to a break point of the main route to continue execution after the instant exploration sub-task is completed.

[0032] According to the image data and UAV pose data, an actual position of the target in a three-dimensional space of the construction site is calculated, including:

[0033] Multiple key evidence images in the image data and corresponding UAV pose data are obtained.

[0034] Target feature points in each key evidence image are matched with surface texture features of a pre-constructed three-dimensional model of a real scene of the construction site.

[0035] Based on the principle of multi-view geometric intersection, the three-dimensional geographic coordinates of the target in the construction site real scene three-dimensional model coordinate system are solved by using the matched feature points and the corresponding unmanned aerial vehicle pose data.

[0036] The four-dimensional hidden danger event data packet and the early warning work order are generated, including:

[0037] The three-dimensional geographic coordinates, the identification timestamp, the hidden danger type and the identification confidence of the target are encapsulated as a core data layer.

[0038] The evidence image or video clip in the image data is associatedly stored, and a three-dimensional scene snapshot centered on the three-dimensional geographic coordinates is automatically generated and encapsulated as an evidence layer.

[0039] The information of the current inspection task identification, the hidden danger discovery stage and the construction area where the target is located are associatedly encapsulated as a context layer.

[0040] The core data layer, the evidence layer and the context layer are integrated to form the four-dimensional hidden danger event data packet, and the early warning work order is generated based on the four-dimensional hidden danger event data packet.

[0041] Based on the on-site verification feedback data, the incremental training of the construction site exclusive safety hidden danger identification model includes:

[0042] The identification results and images labeled as false positives, and the identification results labeled as false negatives are stored in the incremental training sample library combined with the obtained supplementary labeled images.

[0043] Based on the set period, samples are extracted from the incremental training sample library and mixed with the set samples of the construction site specific labeled data set, and the construction site exclusive safety hidden danger identification model is periodically retrained to update the model parameters.

[0044] The application provides a construction site-oriented unmanned aerial vehicle inspection and AI safety hazard identification method. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required to be used in the embodiments or the prior art description will be briefly introduced.

[0046] Fig. 1 is a step schematic diagram of a construction site-oriented unmanned aerial vehicle inspection and AI safety hazard identification method provided by the application.

[0047] Fig. 2 is a flowchart of a construction site-oriented unmanned aerial vehicle inspection and AI safety hazard identification method provided by the application. DETAILED DESCRIPTION

[0048] Hereinafter, exemplary embodiments will be described in detail with reference to the accompanying drawings. In the following description, unless otherwise specified, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present application.

[0049] The terms used in the present application are merely for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "an" and "the" used in the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein means and includes any or all possible combinations of one or more associated listed items.

[0050] It should be understood that, although the terms first, second, third, etc. can be employed in this application to describe various information, the information is not to be limited to these terms. These terms are only used to distinguish one category of information from another. For example, without departing from the scope of the application, first information could also be referred to as second information, and, similarly, second information can also be referred to as first information. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining."

[0051] Referring to Figs. 1-2 The application provides a UAV inspection and AI safety hazard identification method for construction sites, comprising the following steps:

[0052] S101, obtain the initial image set of the target site, label the key sample set filtered from the initial image set, and fine-tune the pre-trained general safety hazard identification model using the generated site-specific labeled data set to generate a site-exclusive safety hazard identification model.

[0053] Specifically, before the UAV takes off, the system needs to import the digital base map of the site, which is usually a building information model or a real scene three-dimensional model generated through pre-flight survey. Based on this model, the task planning module automatically performs the following operations:

[0054] Draw the collection boundary: determine the three-dimensional geographic range of the site to be inspected according to the model.

[0055] Generate a standardized grid flight path: the system generates a flight path covering the entire collection boundary at a set flight height (e.g. 80 meters) in a parallel "bow" shape. The core parameters of this path are the forward overlap rate and the lateral overlap rate, both of which are set to a high standard (e.g. 80%). High overlap rate ensures that the same object is photographed by multiple consecutive photos, which is the basis for subsequent detailed sample analysis and potential three-dimensional reconstruction.

[0056] Set the collection parameters: automatically set the fixed parameters of the UAV camera, such as aperture, shutter speed, and ISO, and uniformly set the trigger mode to equal time interval or equal distance interval to ensure the consistency of image brightness and scale.

[0057] The UAV is equipped with a dual-frequency GNSS module and a high-precision inertial navigation system, and performs fully automatic flight according to the planned route.

[0058] The UAV flies along the route, the gimbal keeps the lens vertical downward (or locks according to the preset angle), and automatically and smoothly triggers the camera shutter according to the planned parameters, performing grid shooting. At the precise moment of each shutter trigger, the flight control system synchronously records a set of spatio-temporal metadata, which at least includes:

[0059] High-precision position information: latitude, longitude, and altitude.

[0060] High-precision attitude information: pitch angle, roll angle, and heading angle.

[0061] Timestamp: UTC time accurate to milliseconds.

[0062] Camera identification: associated with the camera and lens parameters used.

[0063] During this process, each collected JPEG or RAW format image file is immediately bound to the corresponding set of spatiotemporal metadata, forming an image-pose data pair.

[0064] After the flight mission is completed, all image-pose data pairs are transmitted to the edge computing server on the worksite through a high-speed link. The system performs automatic verification to check data integrity (such as whether the image is damaged or the pose data is missing), and stores the verified data in the initial standardized image library. At this point, a complete set of georeferenced, comprehensive coverage, and standardized format initial images is constructed.

[0065] The goal of sample screening is to intelligently select a subset of images that are most valuable for improving the performance of the current model from the vast number of initial images, to minimize the workload of subsequent manual annotation, and to achieve efficient and targeted optimization of the model. This process relies on the inference of a pre-trained base model and an active learning strategy. A pre-set general safety hazard identification base model (such as a neural network model trained on a large public safety dataset) is called to perform batch and automated forward inference on all images in the initial image set. The model analyzes each image and outputs a series of initial recognition results. Each result includes: the recognized target object class (such as "safety helmet", "person", "smoke"), the target's location bounding box in the image, and an initial confidence score representing the model's certainty (a value between 0 and 1). All these initial recognition results of the images are structured and stored in the initial inference result library.

[0066] According to the pre-set screening strategy, key samples are automatically selected from the "initial inference result library". There are two main strategies in parallel channels:

[0067] Channel A: Select the first type of candidate samples. Traverse all initial recognition results and check the initial confidence score for each recognized target. Set a confidence ambiguity interval (for example, 0.3 to 0.7). Mark all recognition targets with confidence scores within this interval as "ambiguous targets". The system traces back to the original images where these "ambiguous targets" are located. Any image containing at least one "ambiguous target" is marked as a candidate ambiguous sample, which is the first type of candidate sample. This type of image reflects the model's uncertainty in the current worksite scenario and is a boundary case that the model needs to focus on learning.

[0068] Channel B: Screen the second type of candidate samples while analyzing those models that show high confidence (e.g., greater than 0.7) in their recognition results. The screening logic focuses on the complexity of the target and the information content of the scene. For example, the system will preferentially screen out images that contain multiple small and dense targets (such as a group of workers), targets that are partially occluded, or scenes with low contrast between the background and the target. More specifically, the system will look for images in which a target is identified as a "person" but is not simultaneously identified with high confidence as wearing a "hard hat." Such images are labeled as candidate complex scene samples, i.e., the second type of candidate samples. They are rich in information that helps the model learn to distinguish details and complex contextual relationships. This is achieved in the following ways:

[0069] "Contains multiple small and dense targets": Calculate the ratio of the total area of all the bounding boxes of the identified targets in a single image to the area of the image. If the ratio is below a set threshold (e.g., 0.1) but the number of targets exceeds a threshold (e.g., 5), it is determined to be a "small and dense target" scene. Alternatively, directly calculate the average IoU (Intersection over Union) between targets. If it is higher than a threshold, it is determined to be dense.

[0070] "Target is partially occluded": In the initial recognition results output by the general model, for each target's bounding box, calculate the semantic segmentation confidence map of the pixels inside it. If there is a continuous region (e.g., area ratio exceeds 30%) with a confidence below another threshold, it can be determined that the target is partially occluded. More simply, it can be determined by checking whether the bounding box intersects with the image edge or overlaps with other target boxes in a large area.

[0071] "Low contrast between background and target": Extract the pixels inside the target bounding box (foreground) and the pixels in a certain area around it (background), and calculate their color histograms (such as RGB space histograms) or texture features (such as LBP features). Calculate the cosine distance or Bhattacharyya distance between the foreground and background feature vectors. If the distance value is below a pre-set threshold, it is determined to be low contrast.

[0072] These criteria can be implemented through image processing libraries (such as OpenCV) and simple numerical comparison code.

[0073] Merge the first type of candidate samples and the second type of candidate samples, and through image hashing or feature matching technology, quickly remove duplicate images from the merged sample set (e.g., adjacent frames with similar images due to high overlap), ensuring that the selected sample set has the greatest information diversity. Finally, a key sample set is generated that is much smaller in size than the original image library but highly concentrated in value. This set contains both cases that the model is confused about and complex scenes that contain rich learning information, providing accurate targets for the next stage of efficient manual annotation.

[0074] The set of key samples to be verified is loaded into the collaborative labeling workbench and dispatched to authorized safety inspectors in the form of a task queue. The interactive interface of the workbench is specially designed, mainly containing three core areas:

[0075] Main visual area: Highlight the current key sample image to be labeled.

[0076] Algorithm suggestion area: Parallel display all initial recognition results of the image by the general base model, including bounding box, class label and initial confidence recorded when screened.

[0077] Labeling operation and feedback area: Provide simplified tool buttons (confirm, modify box, draw new box, select category, delete) and feedback options (true / false positive / miss).

[0078] The safety inspector does not need to start from scratch, but quickly reviews and corrects the content of the "algorithm suggestion area". This is a decision-making closed loop:

[0079] For the recognition box with high initial confidence (such as >0.7), the safety inspector's primary judgment is whether the target exists and the category is correct. For example, the algorithm identifies a red bucket as "flame", and the safety inspector can correct it to "junk". Operationally, only need to click the wrong recognition box, and select the correct category from the drop-down list.

[0080] For the recognition box with confidence in the ambiguous interval, the safety inspector needs to judge whether it is a real hidden danger and the specific type. For example, a fuzzy "safety helmet" box with a confidence of 0.5, the safety inspector needs to confirm whether it is a person without a safety helmet or other objects. By clicking the operation to "confirm" or "delete".

[0081] After any recognition box is confirmed or modified, the safety inspector can supplement the key attributes for it. For example, for a "person" target, supplement its "whether to wear a safety helmet" and "whether to wear a safety rope" state attributes. These attributes will be stored as a composite label (such as "person_ without safety helmet") together, providing a more fine-grained learning goal for the model.

[0082] After reviewing the algorithm suggestions, the safety inspector needs to use his professional experience to check whether there are any safety hazards in the image that the algorithm has completely missed (i.e. missed detection). If there is a missed detection, the safety inspector uses the "draw new box" tool to draw a bounding box around the target in the image. Then, select the corresponding label (such as "missing protective net" or "illegal open flame") from the standard hidden danger category list. This process directly enriches the model's understanding of the unique hidden dangers of the construction site.

[0083] For each processed image, the system encapsulates all the bounding boxes, final class labels, supplementary attributes that have been manually confirmed or corrected, along with the unique identifier of the image, into a structured standard annotation file (e.g. JSON format). All such files are stored in the site-specific annotation dataset.

[0084] The images and annotation files in the site-specific annotation dataset are randomly divided into training set, validation set and local test set according to a preset ratio (e.g. 8:1:1). Only the training set is subjected to light online data augmentation specific to the site scene, such as random brightness adjustment, simulated slight haze, small angle rotation, etc., to improve the model's robustness to changes in lighting and viewing angle, but avoid using strong augmentation that may distort the site geometry (e.g. large flip).

[0085] Load the pre-trained general safety hazard recognition base model. First, freeze all deep layers in the model's backbone network that are used for general feature extraction, only replace or randomly initialize the model's terminal classification head and regression head (network layers responsible for final target classification and positioning), and unfreeze the parameters of these layers. Use the training set of the site-specific annotation dataset to train the model with a lower learning rate. This stage has fewer training rounds, and the goal is to let the model's "decision layer" preliminarily adapt to the class distribution and positioning characteristics of the new site without damaging the learned general features. During training, the validation set is used to monitor performance to prevent overfitting.

[0086] Gradually unfreeze several shallower layers in the model's backbone network. These layers are responsible for capturing more basic and more scene-specific features (e.g. edges, textures), and different learning rates are applied to different layer groups. A relatively high learning rate is used for the newly unfrozen shallow layers to quickly adapt to site-specific features, and a lower learning rate is used for the fine-tuned classification head for fine-tuning. The preliminary misreporting sample library is added as a special validation set or negative sample to the training loop, consciously strengthening the model's learning on cases where it has made mistakes, directly and specifically improving the model's recall and precision in the site. The reserved local test set and preliminary misreporting sample library are used to comprehensively evaluate the fine-tuned model. Core indicators include not only average precision but also performance improvement on key samples. When the model's performance on the validation set and local test set reaches a stable and meets the preset threshold, training stops. The model is exported in an optimized format suitable for edge inference, officially named the site-specific safety hazard recognition model (initial version), and automatically deployed to the UAV on-board computing unit or edge inference server connected to it, replacing the original general base model, completing a complete "on-site self-learning" closed loop. The entire process from start to deployment is controlled within a few hours, achieving rapid transformation from "general" to "dedicated". Among them, the "preliminary misreporting sample library" has two main sources:

[0087] Model fine-tuning stage produces: When training (fine-tuning) the model with site-specific annotation data set, the data set is divided into training set, validation set and test set. When the model is evaluated on the validation set and test set, all the samples that are misidentified (i.e. "false positives" that identify negative samples as positive samples, and "false negatives" that fail to identify positive samples) and their annotation information are automatically recorded and stored in a temporary database.

[0088] Artificially annotated link produces: When the safety officer uses the collaborative annotation workbench to review the "to-be-verified key sample set", his operation behavior (such as marking the model's suggested frame as "false positive", or manually drawing a frame to annotate the target that the model has not identified, i.e. "false negative") will be captured by the system. These false positives and false negatives confirmed by artificial confirmation, together with their correct annotations, will be immediately stored in the sample library.

[0089] This library is automatically built by the system in the initialization stage, and is the natural product of the model training and manual review process.

[0090] S102, based on the periodic global coverage requirement, the historical hidden danger probability heat map and the current high-risk operation area information, generate the main route of the unmanned aerial vehicle and at least one inspection sub-task.

[0091] Specifically, the acquisition method of the periodic global coverage baseline: this information is a built-in rule of the system. Based on the established initial standardized grid route, a standard global coverage route covering the entire site geographical boundary is automatically generated according to the inspection frequency set by the safety management regulations (such as once a day). This route serves as the basic skeleton of this task.

[0092] The acquisition method of the hidden danger probability heat map: from the hidden danger space-time database, extract the three-dimensional position coordinates of all identified hidden danger events in a certain period (such as the past 7 days). Generation process: using kernel density estimation algorithm, conduct spatial statistical analysis on these historical hidden danger points on the two-dimensional plane or three-dimensional real scene model of the site. The more concentrated the hidden danger points in an area, the "hotter" (such as red) the color displayed on the heat map, representing a higher risk probability. This map is not a static picture, but a dynamic geographic information layer with probability value. Among them, "hidden danger space-time database": it is a structured relational database or time series database table. Whenever the system completes a complete hidden danger identification and positioning process (i.e. S104 step), a "four-dimensional hidden danger event data package" is generated, and the core data (time stamp, three-dimensional coordinates, hidden danger type, confidence, task ID, etc.) of the data package will be extracted and written into the corresponding table of "hidden danger space-time database" as an entry. Therefore, this database is obtained by the system continuously accumulating and storing in each inspection task. As the system runs, historical data naturally forms.

[0093] The way to obtain high-risk operation area information: through data interface docking with project management system or building information model, the construction progress plan of the day is automatically imported. From it, the high-risk operation activities being carried out or about to be carried out (such as tower crane hoisting, scaffold erection, foundation pit excavation) and their designated operation area range (usually represented by three-dimensional polygon) are parsed. These operation areas are associated with the standard hidden danger category library, and their preset risk types (such as "personnel intrusion forbidden area" and "unfastened safety rope" corresponding to hoisting area) are marked.

[0094] According to the historical hidden danger data, the historical hidden danger probability heat map is generated, and the regions (hot spot clusters) with heat values exceeding the first threshold value are calculated. The geometric center or density peak point of each hot spot cluster is calculated, and the center point is determined as the first type of dynamic interest point, and the priority is automatically sorted according to the heat value; the current high-risk operation area information is obtained, and the center point or key corner point of the circumscribed rectangle thereof is converted, and the region center point is determined as the second type of dynamic interest point, and the priority is set according to the operation risk level (such as special danger and high danger). Wherein, the heat map generation is to call the existing geographic information analysis library (such as ArcGIS Engine, PostGIS or Python's scipy.stats.gaussian_kde). The input is the (X, Y) coordinate set of all hidden danger points in the past N days in the "hidden danger space-time database". A smooth circular influence area is generated around each hidden danger point using the kernel density estimation algorithm. After superimposing all the influence areas, a continuous probability density distribution map of the entire construction site plane is obtained, that is, the heat map. The heat value is represented by a color gradient.

[0095] First type of dynamic interest point determination:

[0096] A heat value threshold T_hot is set.

[0097] The connected component analysis is performed on all pixel regions with values greater than T_hot on the heat map, and a plurality of "hot spot clusters" are obtained.

[0098] For each "hot spot cluster", the geometric center (centroid) of the pixel set is calculated, or the pixel point with the maximum heat value is directly found as the density peak point. This point is determined as the "first type of dynamic interest point".

[0099] The priority P1 of the point can be directly mapped to the average heat value of the hot spot cluster.

[0100] Second type of dynamic interest point determination:

[0101] The "current high-risk operation area information" obtained from the project management system interface is usually given in the form of a polygon boundary coordinate string.

[0102] Calculate the minimum bounding rectangle of the polygon, or directly calculate the barycenter of the polygon. Determine the center point of the rectangle or the barycenter of the polygon as the "second type of dynamic interest point".

[0103] The priority P2 of the point is determined by the preset operation risk level, for example, "extra-high risk" operation corresponds to priority 5, "high risk" corresponds to 4, and so on.

[0104] Based on the standard global coverage flight route as the basic skeleton, the distance and energy consumption required to deviate from the main flight route to fly to each first type of dynamic interest point are evaluated. For interest points with acceptable deviation cost, the local path of the main flight route is dynamically adjusted to smoothly pass near the interest point, ensuring that the UAV can fly over these hotspot areas at the best viewing angle (usually top view or side view). In the final generated main flight route, each waypoint contains not only position and height information, but also a preset standard collection action (such as "hover at waypoint A for 5 seconds and take a photo"). For those inserted interest point near waypoints, more detailed preset collection actions are bound, such as performing a circle flight above hotspot H1 and continuously recording video. For each first type of dynamic interest point IP1, calculate its Euclidean distance D to the nearest waypoint WP_near on the main flight route. At the same time, the time increment Δt and energy consumption increment ΔE required for the UAV to deviate from WP_near to IP1 and return to the original flight route can be estimated. Set a comprehensive cost threshold C_max (which can be calculated based on distance, time, and energy consumption weighting). If Cost(D, Δt, ΔE) < C_max, then insert:

[0105] Find WP_near and its previous and next waypoints WP_prev, WP_next on the main flight route.

[0106] Use a cubic spline interpolation algorithm or a Bezier curve algorithm to generate a new path segment that smoothly passes near IP1 (with a set offset tolerance) from WP_prev to WP_next.

[0107] Replace the straight line segment from WP_prev to WP_next with the new path segment, and re-discretize to generate a new waypoint sequence to ensure flight smoothness and safety.

[0108] Generate an independent predefined fine inspection sub-task package for the first type of dynamic interest point that is more than a preset offset threshold away from the main flight route, or for the second type of dynamic interest point with extremely high priority (such as an area where super-high layer hoisting is being carried out). The "sub-task package" is a data structure or JSON object that issues a set of explicit and executable fine exploration instructions to the UAV flight control system. Trigger the generation process:

[0109] Condition A (Distance-based): The 2D plane Euclidean distance between the interest point and the nearest waypoint of the main route is greater than a preset offset threshold D_max (e.g., 30 meters).

[0110] Condition B (Priority-based): The preset priority of the interest point reaches or exceeds a system-set "very high" level threshold P_high (e.g., corresponding to "special danger" operation area).

[0111] The data packet contains the following key fields, which can be obtained and filled from the system's existing information or preset templates:

[0112] Target area geometric description: Based on the interest point coordinates, a preset exploration range (such as a circular radius R or a rectangular area) is set according to the hazard type or operation nature.

[0113] Refined flight path: Based on the above area, a standard exploration path is planned. For example, for a planar area, a "bow-shaped" path can be used, and for a point-like target, a "surrounding" or "close hovering" path can be used. The path is defined by a series of ordered three-dimensional waypoint coordinate sequences.

[0114] Exclusive shooting action parameters: Define the actions of the UAV after it arrives at the area, including: the specific height to descend to, the target pitch / yaw angle of the gimbal, the camera zoom level, the shooting / video recording mode and duration, etc.

[0115] Associated attributes: Associate back to the original interest point ID, priority, expected trigger condition (such as triggering when the main model recognition confidence is greater than 0.8) of the generated sub-task packet.

[0116] At the software level, the process is implemented through the following steps:

[0117] Step 1 (Condition Judgment): In the task planning stage, the system iterates through all the determined dynamic interest point list, and applies the above condition A and condition B judgment to each point.

[0118] Step 2 (Template Calling and Parameter Filling): For each interest point marked as needing to generate a sub-task packet, the system calls the corresponding predefined task template according to its type (such as whether it is a "hoisting area" or a "historical fire hotspot"). The template is a task framework containing variable placeholders. Then, the system fills the specific coordinates of the interest point, the preset area size, the priority, etc. into the corresponding variables of the template.

[0119] Step 3 (Package Formatting and Loading): The filled-in task data is serialized into a standard data format (such as JSON, XML, or a specific binary instruction format) that the flight control system can recognize, forming an independent "sub-task package" file or data block. Finally, this package is added to the dynamic task queue of the current inspection task and marked as "waiting for trigger."

[0120] This sub-task package contains:

[0121] Exploration path: A refined flight path from the nearest access point on the main route to the area (such as a zigzag or loop path).

[0122] Exploration action: Detailed shooting instructions (such as "approach to 20 meters height, gimbal pitch angle -30 degrees, take high-definition photos").

[0123] Trigger condition: The sub-task package is pre-loaded into the flight control system and is not immediately executed, but waits for the trigger signal from the real-time recognition module in step S103.

[0124] At this point, the output is: an intelligent fusion of the main route and a set of pre-defined refined inspection sub-task packages, which together form the complete task plan for this inspection.

[0125] S103, using the construction site-specific safety hazard recognition model to identify the video stream collected by the unmanned aerial vehicle flying along the main route in real time, and when an abnormal target that meets the preset conditions is identified, controlling the unmanned aerial vehicle to dynamically execute the corresponding inspection sub-task to obtain the image data of the target.

[0126] Specifically, the unmanned aerial vehicle flight controller accurately executes the intelligent fusion of the main route. At the same time, the on-board visual acquisition unit continuously acquires high-definition video stream, and the edge computing unit synchronously runs the construction site-specific safety hazard recognition model for real-time analysis. The model infers each frame of video and outputs real-time recognition results (target category, bounding box, confidence) with time and space stamps. When the following two conditions are met at the same time, it is determined as an effective abnormal event, triggering a dynamic response:

[0127] High confidence: The confidence of identifying a specific hazard category exceeds a second threshold (such as >0.85).

[0128] Spatial aggregation or severity: A single target has extremely high risk (such as "open fire"), or multiple hazard targets of the same category appear in a frame (such as dense "no safety helmet" personnel), forming a risk cluster.

[0129] Once triggered, the system immediately performs: combining the current precise pose data of the UAV, the pixel coordinates of the abnormal target in the image, quickly solving the approximate geographical coordinates; around the coordinates, dynamically generating an immediate fine exploration sub-task. The core of this task is to obtain clearer and multi-angle evidence, and the task parameters include: a temporary exploration area centered on the target point, a fine flight mode executed in this area (for example: "centered on the target point, radius 10 meters, height dropped to 30 meters, execute half-ring flight"), and targeted shooting instructions (such as "keep gimbal on target point, record 10 seconds of high-definition video"). Among them, the input of quickly solving the approximate geographical position: the pixel coordinates (u, v) of the target in the image, the current pose of the UAV (position [X_u, Y_u, Z_u], attitude angle [roll, pitch, yaw]), camera intrinsic matrix K and installation offset.

[0130] Calculation: According to the pinhole camera model, convert the pixel coordinates (u, v) into a normalized ray direction vector relative to the camera optical center. Use the attitude matrix R of the UAV to rotate this vector to the world coordinate system (NEU coordinate system). Assuming that the target is located on the ground (Z=0 plane), the intersection of this ray and the ground plane can be calculated to obtain the approximate ground coordinates of the target (X_est, Y_est, 0). This is a standard monocular vision ranging and positioning problem.

[0131] Take the above (X_est, Y_est) as the center, generate a circular area with a radius of R (such as 10 meters) as the "temporary exploration area".

[0132] Plan a path from the current position of the UAV to the upper space of the area, and specify the fine flight mode of "hovering gaze" or "small radius circling" in the upper space of the area.

[0133] Generate shooting instructions such as "align gimbal to (X_est, Y_est) point, zoom to 2 times, record 10 seconds of high-definition video".

[0134] Send a "task suspension and intervention" request to the flight controller. The flight controller temporarily suspends the current main flight line task, records the breakpoint position, and accepts new dynamic sub-task instructions. The flight controller controls the UAV to leave the original main flight line, fly to the temporary exploration area along the optimal path, and accurately execute the pre-designed fine flight mode and shooting instructions. At the same time, the AI model performs secondary confirmation analysis on the higher quality images returned.

[0135] When the subtask is completed (e.g., a surround shot is completed), or the preset maximum execution time is reached, the dynamic subtask manager notifies the flight controller. The flight controller calculates the optimal path to return to the breakpoint of the original main route (usually directly flying back or flying to the next adjacent waypoint) according to the global task optimization algorithm, controls the UAV to smoothly return, and seamlessly continues the main route task from the breakpoint or the nearest valid waypoint. During the entire dynamic adjustment process, if a preloaded predefined subtask package is triggered, the more complex preset path in the package is directly executed. When the dynamic subtask is completed, the flight controller (running the path planning algorithm) performs the following steps:

[0136] Input: current position of UAV P_current, main route breakpoint P_break, subsequent waypoint list of main route [WP_next1, WP_next2,...], known obstacle map.

[0137] Planning:

[0138] Strategy A (direct return): Calculate the straight-line path from P_current to P_break. Use A* algorithm or rapid random tree algorithm to plan a collision-free path considering static obstacles.

[0139] Strategy B (go to next waypoint): Calculate the path from P_current to WP_next1. Compare the total path cost (distance, time) of returning to P_break and going to WP_next1, and select the point with lower cost as the target point P_target.

[0140] Output and execution: Convert the planned path from P_current to P_target into a series of control instructions to control the UAV to smoothly fly to P_target, and then automatically continue the original task.

[0141] S104, based on the image data and UAV pose data, calculate the actual position of the target in the three-dimensional space of the construction site, generate a four-dimensional hidden danger event data package and a warning work order.

[0142] Specifically, after completing an investigation, the required data needs to be synchronized and collected, including:

[0143] Key evidence images / video clips: multiple frames of high-definition images or a short video taken by the UAV during dynamic adjustment (e.g., hovering, surrounding), with the clearest image quality and the most direct angle to the hidden danger target.

[0144] High-precision spatio-temporal metadata: The spatio-temporal metadata from the UAV flight control system, which strictly corresponds to each frame of the evidence image, including the precise GNSS position, IMU attitude angle, and camera focal length, principal point, and other intrinsic parameters.

[0145] Site reality 3D model: The pre-constructed site reality 3D model with geographic reference (generated by photogrammetry) or building information model is called from the system.

[0146] A key evidence image with the clearest view is selected, and the corresponding spatio-temporal metadata and camera parameters are used to calculate a rough photogrammetric ray direction. This ray direction line is preliminarily projected onto the surface of the site reality 3D model to obtain a rough 3D initial position. Specifically, the method for quickly calculating the approximate geographic position is used, but a clearer image and more accurate pose data are used to calculate a more accurate photogrammetric ray (a 3D ray). The ray is intersected with the pre-stored high-precision site reality 3D model (Mesh grid model) to obtain the first 3D intersection point as the initial value for optimization iteration.

[0147] Visual feature points are automatically extracted from the hazard target area of the key evidence image. At the same time, a corresponding model surface view with real texture is generated around the 3D initial position on the reality 3D model. Through computer vision algorithms, the image feature points are quickly matched with the feature points of the model surface view. Specifically, the SIFT, SURF, or ORB feature extraction algorithm is used to extract dozens to hundreds of stable local feature points and their descriptors from the region centered on the target in the key evidence image. In the 3D modeling software or game engine, a virtual "model snapshot" image with texture is generated by real-time rendering of the reality 3D model according to the approximate view angle when the UAV took the evidence image (estimated from the initial position and pose). The evidence image and the rendered "model snapshot" image are matched using the FLANN matcher or brute-force matcher, and the RANSAC algorithm is used to remove false matches to obtain a set of reliable 2D (evidence image)-2D (model snapshot) feature point correspondence relationships. Since each pixel of the model snapshot corresponds to a 3D point on the model, this set of correspondence relationships actually establishes a 2D (evidence image)-3D (reality model) correspondence relationship.

[0148] Since the UAV obtains multiple frames of pictures taken from different angles during dynamic exploration, the system will repeat the above feature matching process. Using the matching results obtained from multiple different angles, joint calculation is performed through the principles of spatial resection and forward intersection. This is similar to determining the position of a point using multiple observation stations, which can significantly improve the positioning accuracy, especially the accurate calculation of the elevation coordinates of the hidden point in the vertical direction. Finally, the accurate three-dimensional geographic coordinates of the hidden target in the global coordinate system of the construction site are output. Specifically, for the same hidden target, the three-dimensional initial position is repeatedly calculated, and the process of quickly matching the image feature points with the feature points of the model surface view is repeated, obtaining multiple (≥2) evidence images taken from different angles, and multiple sets of 2D-3D corresponding relationships corresponding to each image. For a single image, using its multiple sets of 2D-3D corresponding relationships, the accurate pose of the UAV when taking the image can be calculated (i.e. solving a PnP problem). The commonly used function is solvePnP or solvePnPRansac in OpenCV; when the refined pose of at least two images is obtained, the observation rays (determined by 2D pixel points and camera optical center) of the same target 3D point are used, and the best intersection point of these rays in space is calculated through the least squares method. The intersection point is the final accurate three-dimensional coordinates of the hidden target.

[0149] After obtaining the accurate coordinates, a four-dimensional hidden event data package containing complete evidence chain and context is automatically packaged. The data package adopts a unified structured format (such as JSON) and contains the following hierarchical information:

[0150] Core data layer:

[0151] Event unique identifier: automatically generated global unique ID.

[0152] Accurate three-dimensional coordinates: fused calculation of (X, Y, Z).

[0153] Timestamp: the accurate time when the hidden danger was first identified.

[0154] Hidden danger type: specific category identified by AI (such as "not wearing safety helmet").

[0155] Recognition confidence: the highest confidence value output by the model.

[0156] Target number: the number of hidden targets of the same type identified.

[0157] Evidence layer:

[0158] Evidence image segment index: associated storage of the file path of the multiple frames of high-definition images or short video segments containing the hidden danger.

[0159] Three-dimensional scene screenshot: The system automatically generates a three-dimensional scene snapshot with attention markers centered on precise three-dimensional coordinates from the real three-dimensional model, intuitively showing the location of the hidden danger in the overall environment.

[0160] Context layer:

[0161] Associated inspection task ID: The number of this flight task.

[0162] Discovery stage: Marking whether it is discovered in "main route flight" or "dynamic sub-task".

[0163] Associated construction area: Automatically associated BIM components or construction partition information according to coordinates.

[0164] Automatically generate intuitive visual warning work orders from four-dimensional hidden danger event data packages, and push them to project management platforms and mobile terminals of site safety officers in real time. The work order is accurately positioned on the three-dimensional map, and is accompanied by on-site pictures, video clips and detailed descriptions to guide personnel to quickly handle. According to the warning information, go to the site to check and confirm, and feed back the verification results through the mobile terminal application. The feedback is divided into three categories: "confirmed to be true" (true positive), "system false alarm" (false positive), or "system missed" (labeled hidden danger not recognized). This step introduces key human judgment. All human feedback results and their corresponding original image data are automatically collected into the "incremental training sample library". Periodically (such as every week), use this sample library to perform incremental learning training on the "construction site-specific safety hidden danger recognition model", so as to realize the continuous iteration and performance optimization of the AI model driven by real business feedback, forming a complete self-evolution closed loop.

[0165] The following beneficial effects can be achieved:

[0166] Improve recognition accuracy and adaptability: By building a "construction site self-learning" mechanism, the AI model can quickly adapt to the characteristics of a specific construction site, significantly improving the accuracy and recall rate of safety hazard identification.

[0167] Realize intelligent guidance inspection: The UAV can dynamically adjust the inspection path and shooting focus according to the real-time analysis results of AI or the historical hidden danger heat map, realizing active and detailed inspection of "suspect and explore".

[0168] Generate spatiotemporally related warning information: Integrate the hidden dangers identified in the two-dimensional image with the three-dimensional real scene model and timestamp of the construction site to generate a digital warning work order containing accurate location, type, time, and image evidence, greatly facilitating the positioning and review of management personnel.

[0169] Form a system autonomous closed loop: Through the "human-machine cooperative verification" feedback link, continuously optimize the AI model and UAV inspection strategy, so that the entire system has the ability to continuously improve itself, reducing long-term operation and maintenance costs.

[0170] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the application embrace any and all variations of the present application that fall within the scope of the general inventive concept as defined in the claims and that the application may be practiced otherwise than is specifically described herein. Any reference to claim limitations in the specification shall not be construed as applying only to the specific claim section. The specification and drawings should be considered in a descriptive sense only and not limiting on the scope of the present application.

[0171] It is to be understood that the application is not limited to the precise construction described in the specification above and shown in the drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the application.

Claims

1. A method for drone inspection and AI-based safety hazard identification at construction sites, characterized in that, Includes the following steps: An initial image set of the target construction site is obtained, and the key sample set to be verified selected from the initial image set is labeled to obtain a construction site labeled dataset. The pre-trained general safety hazard identification model is fine-tuned using the construction site labeled dataset to generate a construction site-specific safety hazard identification model. Based on the requirement of periodic full coverage, the heat map of historical hidden danger probability, and the information of current high-risk operation areas, the main flight path of the UAV and at least one inspection sub-task are generated. The construction site-specific safety hazard identification model is used to identify the video stream collected in real time by the UAV flying along the main flight path. When an abnormal target that meets the preset conditions is identified, the UAV is controlled to dynamically execute the corresponding inspection sub-task to obtain the image data of the target. Based on the image data and UAV pose data, the actual position of the target in the three-dimensional space of the construction site is calculated, and a four-dimensional hidden danger event data package and early warning work order are generated. A set of key samples to be verified is selected from the initial image set, including: The general safety hazard identification model is used to identify the initial image set to obtain an initial identification result that includes the target category, bounding box, and initial confidence level. Images containing the identified target whose initial confidence level is within a preset ambiguity range are selected as the first type of candidate samples; Images containing multiple dense targets, or where the contrast between the target and the background is below a preset threshold, are selected as the second type of candidate samples. The first type of candidate samples and the second type of candidate samples are merged and deduplicated to form the key sample set to be verified.

2. The method for drone inspection and AI-based safety hazard identification at construction sites as described in claim 1, characterized in that, The method further includes: Receive on-site verification feedback data for the four-dimensional hidden danger event data packet, and incrementally train the site-specific safety hazard identification model based on the on-site verification feedback data.

3. The method for drone inspection and AI-based safety hazard identification at construction sites as described in claim 1, characterized in that, The pre-trained general safety hazard identification model is fine-tuned using the generated construction site annotation dataset to generate a construction site-specific safety hazard identification model, including: Freeze the feature extraction network parameters of the general safety hazard identification model, unfreeze and train only its classification and regression network heads, and use the construction site annotation dataset for training; The feature extraction network parameters within the set output range are gradually unfrozen, and differentiated learning rates are applied to different layers. At the same time, the false positives and false negatives generated during the annotation process are used for collaborative training to generate the construction site-specific safety hazard identification model.

4. The method for drone inspection and AI-based safety hazard identification at construction sites as described in claim 1, characterized in that, Based on the requirement of periodic full-area coverage, historical hazard probability heat maps, and current high-risk operation area information, a main flight path for the UAV and at least one inspection sub-task are generated, including: The main route is generated based on the preset grid route. The historical hazard probability heat map is generated based on historical hazard data, and the center point of the area where the heat value exceeds the first threshold is determined as the first type of dynamic interest point; Obtain information on the current high-risk work area and determine the center point of the area as a second type of dynamic point of interest; The first type of dynamic interest points that are more than a preset offset threshold away from the main route, or the second type of dynamic interest points that have reached a set priority, will be used to generate inspection sub-task packages.

5. The method for drone inspection and AI-based safety hazard identification at construction sites as described in claim 4, characterized in that, The method further includes: The first type of dynamic interest points that are less than a preset offset threshold from the main route are inserted into the main route using a path smoothing algorithm.

6. The method for drone inspection and AI-based safety hazard identification at construction sites as described in claim 1, characterized in that, When an abnormal target that meets preset conditions is identified, the drone is controlled to dynamically execute the corresponding inspection sub-task to obtain image data of the target, including: When the target identification confidence level output by the construction site-specific safety hazard identification model exceeds the preset second threshold and the target risk category reaches the set category, it is determined to be an abnormal target that meets the preset conditions; Based on the current pose of the UAV and the location of the abnormal target in the image, the geographical region of the abnormal target is calculated; Generate real-time reconnaissance subtasks centered on the said geographic region; The drone is controlled to pause the execution of the main flight path, switch to the execution of the real-time reconnaissance sub-task, and return to the interruption point of the main flight path to continue execution after completion.

7. The method for drone inspection and AI-based safety hazard identification at construction sites as described in claim 1, characterized in that, Based on the image data and UAV pose data, the actual position of the target in the three-dimensional space of the construction site is calculated, including: Acquire multiple key evidence images from the image data and their corresponding UAV pose data; The target feature points in each frame of the key evidence image are matched with the surface texture features of a pre-constructed 3D model of the construction site. Based on the principle of multi-view geometric intersection, the three-dimensional geographic coordinates of the target in the coordinate system of the construction site real scene three-dimensional model are calculated by using multiple successfully matched feature points and the corresponding UAV pose data.

8. The method for drone inspection and AI-based safety hazard identification at construction sites as described in claim 7, characterized in that, Generate a four-dimensional hidden danger event data package and an early warning work order, including: The target's three-dimensional geographic coordinates, identification timestamp, hazard type, and identification confidence level are encapsulated into a core data layer; The evidence images or video clips in the image data are associated and stored, and a three-dimensional scene snapshot centered on the three-dimensional geographic coordinates is automatically generated and encapsulated as an evidence layer; The information related to the current inspection task identifier, the stage of hazard discovery, and the construction area where the target is located is encapsulated as a context layer. The core data layer, evidence layer, and context layer are integrated to form the four-dimensional hidden danger event data package, and an early warning work order is generated based on the four-dimensional hidden danger event data package.

9. The method for drone inspection and AI-based safety hazard identification at construction sites as described in claim 2, characterized in that, Incremental training is performed on the construction site-specific safety hazard identification model based on the on-site inspection feedback data, including: The identification results and images labeled as false positives, as well as the identification results labeled as false negatives, are combined with the acquired supplementary labeled images and stored in the incremental training sample library. Samples are extracted from the incremental training sample library at a set periodic interval and mixed with the set samples from the construction site annotation dataset to periodically retrain the construction site-specific safety hazard identification model in order to update the model parameters.

Citation Information

Patent Citations

  • Monitoring method and system for full life cycle of transformer substation project

    CN120634140A

  • Unmanned aerial vehicle intelligent inspection method for construction site operation safety

    CN120653015A