Real-time three-dimensional modeling method for engineering construction site based on unmanned aerial vehicle aerial image

By using real-time 3D modeling methods based on UAV aerial imagery, combined with visual-inertial synchronous positioning and deep semantic features, the problem of unstable model accuracy at construction sites was solved, enabling precise quantification of construction progress and safety optimization decisions, and improving the dynamic monitoring capabilities of construction sites.

CN121120938BActive Publication Date: 2026-05-15BEIJING HSINCHU LANYUE ELECTRIC POWER ENGINEERING SERVICES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING HSINCHU LANYUE ELECTRIC POWER ENGINEERING SERVICES CO LTD
Filing Date
2025-09-04
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies for 3D modeling of engineering construction sites, the accuracy of the model is affected by dynamic targets on site, making it impossible to accurately identify construction activities. The lack of multi-source information fusion and forward-looking planning leads to inaccurate assessment of construction progress, and safety is not considered as an optimization objective.

Method used

A real-time 3D modeling method based on UAV aerial imagery is adopted. A 3D spatial point cloud model is generated through visual-inertial synchronous positioning and mapping algorithms. Combined with depth semantic features and cross-view semantic consistency verification, low-confidence areas are identified. In addition, a quantitative analysis report on construction progress is generated by combining construction logic, and UAV flight parameters are optimized to achieve intelligent decision-making.

Benefits of technology

It enables real-time tracking and precise quantification of dynamic changes at the construction site, provides high-confidence data support, and achieves precise closed-loop management from data acquisition to decision optimization, thereby improving the accuracy of construction progress assessment and the coverage of safety monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120938B_ABST
    Figure CN121120938B_ABST
Patent Text Reader

Abstract

The application discloses an engineering construction site real-time three-dimensional modeling method based on unmanned aerial vehicle aerial image, relates to the technical field of computer vision, and comprises the following steps: collecting oblique image data and high-frequency pose data, acquiring sparse key point features, deep semantic features and high-frequency time sequence features, generating a three-dimensional space point cloud model, identifying a low-confidence area, comparing the three-dimensional space point cloud model with a dynamic semantic three-dimensional model of the last period as a reference model to identify a semantic change area and a semantic state conversion conforming to a preset construction logic, constructing a dynamic semantic three-dimensional model of the current period, and generating an unmanned aerial vehicle flight parameter adjustment instruction. The application improves the three-dimensional modeling quality through cross-view semantic consistency verification, interprets progress changes based on the construction logic, combines a building information model and time sequence prediction, solves a multi-objective optimization function, decides unmanned aerial vehicle collection work, and provides high-confidence data support for construction site management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, specifically to a method for real-time 3D modeling of engineering construction sites based on UAV aerial imagery. Background Technology

[0002] In the field of engineering construction, the construction site is the specific physical space where various engineering projects are implemented. It covers the entire process from project commencement to completion and acceptance, involving work areas such as land leveling, foundation construction, main structure construction, and decoration and finishing. It usually includes different construction elements, such as building material storage points, parking and work areas for construction machinery and equipment, and the activity range of construction personnel, and is in a dynamic state as the project progresses.

[0003] To address the lack of timely and accurate status awareness of dynamically changing construction sites, existing technologies employ drone aerial photography for 3D modeling. However, issues arise during visual synchronous positioning and mapping. Model accuracy is affected by dynamic targets on-site. Periodic model comparisons only identify surface geometric and semantic changes, failing to interpret data changes as specific construction activities. Furthermore, subsequent data acquisition and optimization lack intelligent decision-making mechanisms that integrate multi-source information for proactive prediction and forward-looking planning. Construction safety is also not considered as an optimization objective in aerial planning. Consequently, inaccurate quantitative assessments of construction progress and a lack of high-confidence data support for project management decisions are encountered. Summary of the Invention

[0004] The purpose of this invention is to provide a real-time 3D modeling method for engineering construction sites based on UAV aerial imagery, in order to solve the problems mentioned in the background art.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a real-time 3D modeling method for engineering construction sites based on UAV aerial imagery, comprising the following S1-S7:

[0006] S1. Collect and preprocess the tilted image data of the current period and the high-frequency pose data of the UAV itself to obtain the preprocessed image dataset and pose dataset.

[0007] S2. Obtain sparse keypoint features and deep semantic features from the image dataset, and obtain high-frequency temporal features from the pose dataset;

[0008] S3. Based on sparse key point features and high-frequency temporal features, a visual-inertial synchronous positioning and mapping algorithm is used to generate a three-dimensional spatial point cloud model to describe the construction site in the current cycle. The three-dimensional spatial points in the three-dimensional spatial point cloud model are subjected to cross-view semantic consistency verification to identify low confidence areas.

[0009] S4. Project the 3D spatial points onto the 2D image and combine them with deep semantic features to assign semantic labels to the 3D spatial points in order to generate a 3D spatial point cloud model with semantic labels.

[0010] S5. Compare the semantically labeled 3D spatial point cloud model with the dynamic semantic 3D model of the previous cycle, which serves as the baseline model, to identify semantic change areas and semantic state transition areas that conform to the preset construction logic, and construct the dynamic semantic 3D model of the current cycle. In the first cycle, the semantically labeled 3D spatial point cloud model is used as the initial baseline model.

[0011] S6. By comparing the current period's dynamic semantic 3D model with the benchmark model, a quantitative analysis report of construction progress is generated.

[0012] S7. Overlay the dynamic semantic 3D model and the baseline model, and determine the key areas of concern based on the quantitative analysis report of construction progress, low confidence areas and semantic state transition areas. For the key areas of concern, generate UAV flight parameter adjustment instructions by solving the flight parameter optimization cost function.

[0013] A further improvement of the technical solution of the present invention is that the UAV includes a flight platform, a sensing payload and a positioning and navigation system, wherein the flight platform is a multi-rotor UAV;

[0014] The sensing payload includes a multi-axis stabilized gimbal and a high-resolution digital camera mounted on the multi-axis stabilized gimbal. The multi-axis stabilized gimbal is used to control the tilt angle of the digital camera.

[0015] The positioning and navigation system includes a differential global navigation satellite system receiver, a high-frequency inertial measurement unit, and a hardware time synchronization unit;

[0016] The process of deploying drones to conduct periodic flights at the construction site, collecting and preprocessing oblique image data and the drone's own high-frequency pose data for each period, to obtain preprocessed image and pose datasets includes:

[0017] Global Navigation Satellite System (GNSS) reference stations are deployed at control points with known three-dimensional coordinates at the construction site. These reference stations are used to send differential correction data to the differential GNSS receiver.

[0018] The periodic flight routes of the UAV are planned in advance through the ground control station. The flight routes include waypoints and the tilt angles corresponding to the waypoints.

[0019] During the flight of the UAV, the attitude of the high-resolution digital camera is adjusted by a multi-axis stabilization gimbal so that the high-resolution digital camera can acquire tilted image data at preset time intervals. The UAV flies autonomously according to a periodic flight path.

[0020] High-frequency pose data is generated by acquiring satellite observation values ​​and inertial measurement values ​​through a differential global navigation satellite system receiver and a high-frequency inertial measurement unit, respectively.

[0021] Each time oblique image data is acquired, a timestamp is recorded in the high-frequency pose data through a hardware time synchronization unit;

[0022] Lens distortion correction is performed on tilted image data based on the internal calibration parameters of the high-resolution digital camera;

[0023] By fusing satellite observations and inertial measurements in real time using a fusion algorithm, a six-degree-of-freedom motion trajectory for the UAV is generated. The six-degree-of-freedom motion trajectory is used as high-frequency pose data, where... For position vectors, For velocity vector, For attitude quaternions, and These are the zero bias vectors of the accelerometer and gyroscope, respectively. This is a transpose operation;

[0024] The corrected tilted image data and high-frequency pose data are synchronized and aligned according to the timestamp, so that each frame of tilted image data is associated with the high-frequency pose data at the moment of shooting, forming a preprocessed image dataset and pose dataset.

[0025] A further improvement to the technical solution of this invention lies in the following: the process of obtaining sparse keypoint features and deep semantic features from image datasets, and obtaining high-frequency temporal features from pose datasets, includes:

[0026] For each tilted image in the image dataset, the key points that maintain stability under image rotation and scale changes are identified and located by comparing the intensity difference between the center pixel and the ring-shaped pixels in the neighborhood of the center pixel.

[0027] Within a local image patch surrounding each keypoint, pixel pairs are selected, and a binary intensity test is performed on each pair to generate a binary vector as the keypoint descriptor. The binary intensity test calculation process includes comparing the intensity values ​​of the selected pixel pairs. If the intensity value of the first pixel is less than the intensity value of the second pixel, the test result is 1; if the intensity value of the first pixel is not less than the intensity value of the second pixel, the test result is 0. The specific calculation formula is as follows:

[0028] ;

[0029] in, and The first The intensity value of a group of pixel pairs. For the first binary vector generated The value of each bit;

[0030] By combining the two-dimensional coordinates of keypoints in oblique image data and the descriptors of keypoints, sparse keypoint features are obtained.

[0031] The oblique image data is input into a pre-trained convolutional neural network with an encoder-decoder structure;

[0032] The oblique image data is encoded using an encoder in a convolutional neural network. The encoding process includes abstracting the oblique image data layer by layer and converting it into a high-dimensional feature map.

[0033] High-dimensional feature maps are used as deep semantic features, and are used to distinguish the numerical representation of abstract image patterns of different construction elements.

[0034] The six-degree-of-freedom motion trajectories in the pose dataset are used as high-frequency temporal features.

[0035] A further improvement to the technical solution of this invention lies in the following: the process of generating a three-dimensional spatial point cloud model describing the construction site in the current period, based on sparse key point features and high-frequency temporal features, using a visual-inertial synchronous positioning and mapping algorithm, includes:

[0036] The initialization process for performing visual-inertial simultaneous localization and mapping (VIXMR) technology includes:

[0037] Select consecutive initial image frames from the image dataset and obtain sparse keypoint features and high-frequency temporal features corresponding to the initial image frames;

[0038] By comparing the descriptors contained in the sparse keypoint features, a keypoint correspondence is established between the initial image frames;

[0039] Based on the correspondence between key points and high-frequency temporal characteristics, a joint optimization algorithm is used to simultaneously solve for the position vector and attitude quaternion corresponding to the initial image frame in the six-degree-of-freedom motion trajectory, the three-dimensional spatial coordinates of each corresponding key point in the initial three-dimensional spatial point cloud model, and the zero bias vectors of the accelerometer and gyroscope in the six-degree-of-freedom motion trajectory, so as to construct the three-dimensional spatial point cloud model.

[0040] For each subsequent image frame after the initialization process is completed, incremental pose tracking and mapping are performed. Incremental pose tracking and mapping includes:

[0041] Match the sparse keypoint features of the current image frame with the 3D spatial points that already exist in the 3D spatial point cloud model;

[0042] Using the inertial measurement values ​​between the previous image frame and the current image frame, motion prediction is performed on the camera pose of the current image frame. The camera pose includes the position vector and the attitude quaternion.

[0043] Using the camera pose as the initial value, the precise camera pose of the current image frame is solved by minimizing the reprojection error of the three-dimensional space points, and the precise camera pose is added to the motion trajectory of the high-resolution digital camera.

[0044] The global optimization and map update process is performed periodically. The global optimization and map update process includes:

[0045] A factor graph is constructed by using the camera pose in the motion trajectory of the updated high-resolution digital camera, the 3D spatial points in the 3D spatial point cloud model, and the visual constraints and inertial constraints transformed from sparse keypoint features and high-frequency temporal features as factor nodes.

[0046] The factor map is solved using a nonlinear optimization algorithm to minimize the joint cost function, which includes reprojection error and inertial measurement error between continuous camera poses, to output a globally consistent high-resolution digital camera motion trajectory and 3D point cloud model. The calculation formula is as follows:

[0047] ;

[0048] in, Let be the joint cost function to be minimized. It is a set of state variables that include the camera pose. A set of landmarks containing the coordinates of points in three-dimensional space. For camera pose Road sign Reprojection error, For continuous camera pose and Inertial measurement error between and These are the covariance matrices for visual measurements and inertial measurements, respectively. and These are the sets of visual constraints and the sets of inertial constraints, respectively.

[0049] For sparse keypoints that are stably tracked under camera pose but are not included in the 3D spatial point cloud model, triangulation is performed to generate the 3D spatial coordinates of the sparse keypoints, and the generated 3D spatial points are added to the 3D spatial point cloud model to obtain the 3D spatial point cloud model of the construction site in the current cycle.

[0050] A further improvement to the technical solution of this invention lies in the process of performing cross-view semantic consistency verification on three-dimensional spatial points in a three-dimensional spatial point cloud model to identify low-confidence regions, which includes:

[0051] For each newly generated 3D spatial point in the 3D spatial point cloud model, the set of camera poses of the observed 3D spatial points is determined based on the motion trajectory of the high-resolution digital camera.

[0052] For each camera pose in the camera pose set, the 3D spatial point is projected onto the corresponding oblique image data to obtain the view projection coordinates;

[0053] Based on the view projection coordinates, deep semantic features are extracted from their respective high-dimensional feature maps, and temporary semantic labels are generated for the projection points under each view, forming a set of semantic labels.

[0054] The semantic label set is statistically analyzed and compared. When the inconsistency of semantic labels in the semantic label set exceeds a preset threshold, the region where the three-dimensional spatial point is located is marked as a low-confidence region.

[0055] A further improvement to the technical solution of this invention lies in the process of projecting three-dimensional spatial points onto a two-dimensional image and combining them with depth semantic features to assign semantic labels to the three-dimensional spatial points, thereby generating a three-dimensional spatial point cloud model with semantic labels, including:

[0056] The camera projection transformation process involves calculating the two-dimensional pixel coordinates of a three-dimensional point in the image dataset using camera projection transformation.

[0057] For each 3D point in the 3D point cloud model, the camera pose of the observed 3D point is determined based on the motion trajectory of the high-resolution digital camera.

[0058] For each determined camera pose, a camera projection model is used to transform the three-dimensional spatial coordinates of a point into two-dimensional pixel coordinates in the corresponding image frame under the camera pose. The calculation formula for the camera projection model is as follows:

[0059] ;

[0060] in, As a scale factor, Two-dimensional pixel coordinates, For the internal calibration parameter matrix of a high-resolution digital camera, Let the rotation matrix be the camera pose. Let be the translation vector of the camera pose. This represents the camera extrinsic parameter matrix, which is composed of the rotation matrix and the translation vector. Three-dimensional spatial coordinates;

[0061] Assigning semantic labels to each point in three-dimensional space involves the following process:

[0062] Based on the two-dimensional pixel coordinates, the deep semantic feature vector of the two-dimensional pixel coordinates is queried and extracted from the high-dimensional feature map;

[0063] The deep semantic feature vector is transformed into a fused feature vector through a feature fusion algorithm.

[0064] The fused feature vector is input into the classifier, which outputs the probability that a point in the three-dimensional space belongs to each of the preset construction element categories.

[0065] The construction element category with the highest probability among all construction element categories is used as the semantic label of the three-dimensional spatial point.

[0066] The process continues until all three-dimensional spatial points in the three-dimensional spatial point cloud model are assigned semantic labels, resulting in a three-dimensional spatial point cloud model with semantic labels. In this model, each three-dimensional spatial point contains its corresponding three-dimensional spatial coordinates and semantic labels.

[0067] A further improvement to the technical solution of this invention lies in: comparing the semantically labeled 3D spatial point cloud model with the dynamic semantic 3D model of the previous period, which serves as a baseline model, to identify semantically changing regions and regions with semantic state transitions conforming to preset construction logic, and constructing the dynamic semantic 3D model of the current period includes:

[0068] Determine whether a dynamic semantic 3D model from the previous cycle exists;

[0069] If there is no dynamic semantic 3D model for the previous period, the current period is determined to be the first period, and the 3D spatial point cloud model with semantic labels is designated as the dynamic semantic 3D model for the current period.

[0070] If a dynamic semantic 3D model from the previous cycle exists, then the current cycle is determined not to be the first cycle. The dynamic semantic 3D model from the previous cycle is used as the baseline model. The 3D spatial point cloud model with semantic labels is compared with the baseline model to obtain the change state identifier.

[0071] Based on the change state identifier, construct a dynamic semantic 3D model for the current cycle;

[0072] The constructed dynamic semantic 3D model for the current cycle is archived to serve as the baseline model for the next cycle.

[0073] The process of comparing a semantically labeled 3D point cloud model with a benchmark model includes:

[0074] Using the nearest neighbor search algorithm, the corresponding relationships between 3D spatial points in the semantically labeled 3D spatial point cloud model of the current period are found and established in the benchmark model.

[0075] For each pair of three-dimensional spatial points consisting of corresponding relationships, a change recognition calculation is performed to obtain change state identifiers. These change state identifiers include geometric change, semantic change, and semantic state transition identifiers. The execution process of the change recognition calculation includes:

[0076] Calculate the Euclidean distance between two 3D spatial points in a 3D spatial point pair. When the Euclidean distance is greater than a preset geometric change threshold, mark the 3D spatial point pair as having undergone geometric change.

[0077] Compare the semantic labels of two 3D spatial points in a 3D spatial point pair. When the semantic labels are inconsistent, mark the 3D spatial point pair as having a semantic change.

[0078] When a three-dimensional spatial point pair is marked as a semantic change, it is further determined whether the change of the corresponding semantic label conforms to the valid process flow rules in the preset construction process logic knowledge base. If the change of the semantic label conforms to the valid process flow rules, then a semantic state transition identifier is added to the three-dimensional spatial point pair.

[0079] The construction process of the dynamic semantic 3D model for the current cycle includes:

[0080] The data structure is based on the current period's semantically labeled 3D spatial point cloud model;

[0081] For each three-dimensional point in the basic data structure, add a change status identifier.

[0082] A further improvement to the technical solution of this invention lies in the process of generating a quantitative analysis report of construction progress by comparing the current period's dynamic semantic 3D model with the baseline model, which includes:

[0083] Receive the semantic region category specified by the user as the computation target;

[0084] If the current period is not the first period, three-dimensional spatial points that match the semantic labels and semantic region categories are selected from the dynamic semantic three-dimensional model and the baseline model of the current period to form the target point cloud of the current period and the baseline target point cloud.

[0085] On a horizontal two-dimensional plane covering the current period target point cloud and the reference target point cloud, define a plane with a preset area. A regular grid composed of grid cells;

[0086] Project the current target point cloud and the reference target point cloud onto a regular grid, respectively. For each grid cell in the regular grid, take the average elevation value of the 3D spatial points falling within the grid cell, and calculate the current elevation value. and benchmark elevation value Generate the current digital surface model and the baseline digital surface model;

[0087] Calculate the volume change of each grid cell in the regular grid one by one. and the volume change of each grid cell. Perform algebraic summation to obtain the total volume change of semantic region categories;

[0088] The total volume change, semantic region category, and cycle time information are integrated into a quantitative analysis report of construction progress.

[0089] A further improvement to the technical solution of this invention lies in the fact that the process of overlaying and rendering a dynamic semantic 3D model and a baseline model includes:

[0090] In the user terminal's graphical rendering interface, load the current cycle's dynamic semantic 3D model and baseline model.

[0091] Establish rendering rules, which are used to map the change state identifiers of 3D spatial points in the dynamic semantic 3D model to preset visual styles;

[0092] According to the rendering rules, the dynamic semantic 3D model of the current cycle is colored. Among them, the regions with geometric and semantic changes are displayed with preset highlight colors.

[0093] The numerical data from the quantitative analysis report of construction progress are displayed in the form of charts and graphs overlaid on the corresponding positions in the graphical rendering interface.

[0094] A further improvement to the technical solution of this invention lies in: determining key areas of concern based on the quantitative analysis report of construction progress, low-confidence areas, and areas of semantic state transition; and generating UAV flight parameter adjustment instructions for these key areas of concern by solving the flight parameter optimization cost function.

[0095] Obtain the construction operation plan and safety risk area data for the next cycle from the building information model of the engineering construction.

[0096] Analyze the quantitative analysis report of construction progress, identify the semantic region categories where the total volume change exceeds the preset analysis threshold, and determine the three-dimensional spatial points in the dynamic semantic three-dimensional model of the current period that match the semantic region categories as the progress key point cloud;

[0097] From the dynamic semantic 3D model of the current cycle, extract 3D spatial points marked as low-confidence regions to form a low-confidence point cloud;

[0098] Extract the three-dimensional spatial points with attached semantic state transition labels from the dynamic semantic three-dimensional model of the current cycle to form a state transition point cloud;

[0099] Based on a dynamic semantic 3D model of historical cycles, a temporal prediction model based on long short-term memory network is used to generate a heatmap of expected semantic state transitions for the next cycle, and point clouds of expected key regions are extracted from the semantic state transition heatmap.

[0100] The point cloud of progress key points, low confidence points, state transition points, expected key area points, and points corresponding to the data of safety risk areas are merged, and the spatial range boundary of the merged point cloud is calculated. The calculated spatial range boundary is defined as the key focus area for the next collection cycle.

[0101] Assign high confidence weights to low confidence point clouds in key areas of concern, assign high progress association weights to state transition point clouds, and assign high security monitoring weights to point clouds corresponding to data in security risk areas.

[0102] A flight parameter optimization cost function is constructed with the optimization objectives of maximizing model confidence, schedule correlation, and safety monitoring coverage, and with acquisition efficiency as a constraint. Model confidence is an indicator used to quantitatively evaluate the reliability of regional information in the 3D spatial point cloud model, and is calculated through cross-view semantic consistency. Schedule correlation is an indicator used to measure the degree of correlation between changes in the construction site area and the project construction progress, and is jointly determined by the actual semantic state transitions, the work content obtained from the construction work plan, and the semantic state transition heatmap generated by the time-series prediction model. Safety monitoring coverage is an indicator used to evaluate the effectiveness of the acquisition operation in monitoring the preset safety control area obtained from the building information model.

[0103] The flight parameter optimization cost function is solved by an optimization algorithm to obtain a set of optimized flight parameters that minimize the cost value. The optimized flight parameters include flight altitude, camera tilt angle and flight path density for key areas of interest.

[0104] The optimized flight parameters will be integrated into a single command for adjusting the UAV flight parameters.

[0105] Due to the adoption of the above technical solution, the technical progress achieved by this invention compared to the prior art is as follows:

[0106] This invention provides a real-time 3D modeling method for engineering construction sites based on UAV aerial imagery. By introducing cross-view semantic consistency verification during the 3D modeling process, it can identify and mark low-confidence areas in the 3D spatial point cloud model, solving the problem of unstable model accuracy caused by interference from dynamic targets on site, and providing a high-confidence data foundation for subsequent quantitative analysis.

[0107] This invention provides a real-time 3D modeling method for engineering construction sites based on UAV aerial imagery. By combining periodic comparisons with preset construction procedure logic, it can intelligently interpret the original semantic changes into effective construction state transitions, realizing dynamic tracking and precise quantification of site changes, and transforming ambiguous construction progress into objective and traceable construction activity reports.

[0108] This invention provides a real-time 3D modeling method for engineering construction sites based on UAV aerial imagery. By combining building information modeling, construction operation planning and time-series prediction models, and incorporating safety monitoring coverage into the optimization objective, the method transforms the adjustment of UAV flight parameters into a forward-looking intelligent decision-making process that serves both safety and schedule. This achieves precise closed-loop management from data acquisition and analysis to problem-driven operation optimization. Attached Figure Description

[0109] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0110] Figure 1 The flowchart illustrates the real-time 3D modeling method for engineering construction sites based on UAV aerial imagery provided by this invention. Detailed Implementation

[0111] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0112] Example 1, such as Figure 1 As shown, this invention provides a real-time 3D modeling method for engineering construction sites based on UAV aerial imagery, including the following S1-S7:

[0113] S1. By deploying drones to conduct periodic flights at the construction site, for each period, the tilt image data of the current period and the high-frequency pose data of the drone itself are collected and preprocessed to obtain the preprocessed image dataset and pose dataset.

[0114] In some embodiments, the drone includes a flight platform, a sensing payload, and a positioning and navigation system, wherein the flight platform is a multi-rotor drone.

[0115] In some embodiments, the sensing payload includes a multi-axis stabilization gimbal and a high-resolution digital camera mounted on the multi-axis stabilization gimbal, the multi-axis stabilization gimbal being used to control the tilt angle of the digital camera.

[0116] In some embodiments, the positioning and navigation system includes a differential global navigation satellite system receiver, a high-frequency inertial measurement unit, and a hardware time synchronization unit.

[0117] In some embodiments, a Global Navigation Satellite System (GNSS) reference station is deployed at a control point with known three-dimensional coordinates at the construction site. The GNSS reference station is used to transmit differential correction data to a differential GNSS receiver.

[0118] In some embodiments, a ground control station pre-plans periodic flight routes for the UAV, which include waypoints and the tilt angles corresponding to the waypoints.

[0119] In some embodiments, during the flight of the UAV, the attitude of the high-resolution digital camera is adjusted by a multi-axis stabilization gimbal so that the high-resolution digital camera acquires tilted image data at preset time intervals, wherein the UAV flies autonomously according to a periodic flight path.

[0120] In some embodiments, satellite observations and inertial measurements are acquired by a differential global navigation satellite system receiver and a high-frequency inertial measurement unit, respectively, to generate high-frequency pose data.

[0121] In some embodiments, a timestamp is recorded in the high-frequency pose data each time tilted image data is acquired, using a hardware time synchronization unit.

[0122] In some embodiments, lens distortion correction is performed on tilted image data based on the internal calibration parameters of the high-resolution digital camera.

[0123] In some embodiments, a fusion algorithm is used to process satellite observations and inertial measurements in real time to generate a six-degree-of-freedom motion trajectory for the UAV. The six-degree-of-freedom motion trajectory is used as high-frequency pose data, where... For position vectors, For velocity vector, For attitude quaternions, and These are the zero bias vectors of the accelerometer and gyroscope, respectively. This is for the transpose operation.

[0124] In some embodiments, the corrected tilt image data and high-frequency pose data are time-stamped and aligned, so that each frame of tilt image data is associated with the high-frequency pose data at the moment of shooting, forming a preprocessed image dataset and pose dataset.

[0125] S2. Obtain sparse keypoint features and deep semantic features from the image dataset, and obtain high-frequency temporal features from the pose dataset.

[0126] In some embodiments, for each tilted image data in the image dataset, key points that maintain stability under image rotation and scale changes are identified and located by comparing the intensity difference between the center pixel and the ring-shaped pixels in the neighborhood of the center pixel.

[0127] In some embodiments, within a local image patch surrounding each keypoint, pixel pairs are selected, and a binary intensity test is performed on each pixel pair to generate a binary vector as a descriptor for the keypoint. The calculation process of the binary intensity test includes comparing the intensity values ​​of the selected pixel pairs. If the intensity value of the first pixel is less than the intensity value of the second pixel, the test result is 1; if the intensity value of the first pixel is not less than the intensity value of the second pixel, the test result is 0. The specific calculation formula is as follows:

[0128]

[0129] in, and The first The intensity value of a group of pixel pairs. For the first binary vector generated The value of each bit.

[0130] In some embodiments, sparse keypoint features are obtained by combining the two-dimensional coordinates of keypoints in oblique image data and the descriptors of keypoints.

[0131] In some embodiments, tilted image data is input into a pre-trained convolutional neural network with an encoder-decoder structure.

[0132] In some embodiments, the oblique image data is encoded by an encoder in a convolutional neural network, the encoding process including abstracting the oblique image data layer by layer and converting it into a high-dimensional feature map.

[0133] In some embodiments, high-dimensional feature maps are used as deep semantic features, and these high-dimensional feature maps are used to distinguish the numerical representation of abstract image patterns of different construction elements.

[0134] In some embodiments, the six-degree-of-freedom motion trajectories in the pose dataset are used as high-frequency temporal features.

[0135] S3. Based on sparse key point features and high-frequency temporal features, a visual-inertial synchronous positioning and mapping algorithm is used to generate a three-dimensional spatial point cloud model to describe the construction site in the current cycle. The three-dimensional spatial points in the three-dimensional spatial point cloud model are then subjected to cross-view semantic consistency verification to identify low-confidence areas.

[0136] In some embodiments, an initialization process for visual-inertial simultaneous localization and mapping (VILM) technology is performed, the initialization process including:

[0137] In some embodiments, consecutive initial image frames are selected from the image dataset, and sparse keypoint features and high-frequency temporal features corresponding to the initial image frames are obtained.

[0138] In some embodiments, keypoint correspondences are established between initial image frames by comparing the descriptors contained in sparse keypoint features.

[0139] In some embodiments, based on the correspondence between key points and high-frequency temporal characteristics, a joint optimization algorithm is used to simultaneously solve for the position vector and attitude quaternion corresponding to the initial image frame in the six-degree-of-freedom motion trajectory, the three-dimensional spatial coordinates of each corresponding key point in the initial three-dimensional spatial point cloud model, and the zero bias vectors of the accelerometer and gyroscope in the six-degree-of-freedom motion trajectory, so as to construct a three-dimensional spatial point cloud model.

[0140] In some embodiments, for each subsequent image frame after the initialization process is completed, incremental pose tracking and mapping processing is performed, which includes:

[0141] In some embodiments, sparse keypoint features of the current image frame are matched with 3D spatial points that already exist in the 3D spatial point cloud model.

[0142] In some embodiments, the camera pose of the current image frame is predicted using inertial measurement values ​​between the previous image frame and the current image frame. The camera pose includes a position vector and an attitude quaternion.

[0143] In some embodiments, the precise camera pose of the current image frame is solved by minimizing the reprojection error of the three-dimensional spatial points, using the camera pose as the initial value, and the precise camera pose is added to the motion trajectory of the high-resolution digital camera.

[0144] In some embodiments, a global optimization and map update process is performed periodically, which includes:

[0145] In some embodiments, a factor graph is constructed by using the camera pose in the motion trajectory of the updated high-resolution digital camera, the three-dimensional spatial points in the three-dimensional spatial point cloud model, and the visual constraints and inertial constraints transformed from sparse keypoint features and high-frequency temporal features as factor nodes.

[0146] In some embodiments, a nonlinear optimization algorithm is used to solve the factor map, minimizing the joint cost function that includes reprojection error and inertial measurement error between continuous camera poses, to output a globally consistent high-resolution digital camera motion trajectory and 3D spatial point cloud model. The specific calculation formula is as follows:

[0147]

[0148] in, Let be the joint cost function to be minimized. It is a set of state variables that include the camera pose. A set of landmarks containing the coordinates of points in three-dimensional space. For camera pose Road sign Reprojection error, For continuous camera pose and Inertial measurement error between and These are the covariance matrices for visual measurements and inertial measurements, respectively. and These are the sets of visual constraints and the sets of inertial constraints, respectively.

[0149] In some embodiments, for sparse keypoints that are stably tracked under camera pose but are not included in the three-dimensional spatial point cloud model, triangulation is performed to generate the three-dimensional spatial coordinates of the sparse keypoints, and the generated three-dimensional spatial points are added to the three-dimensional spatial point cloud model to obtain the three-dimensional spatial point cloud model of the construction site in the current cycle.

[0150] In some embodiments, for each newly generated 3D spatial point in the 3D spatial point cloud model, the set of camera poses of the observed 3D spatial points is determined based on the motion trajectory of a high-resolution digital camera.

[0151] In some embodiments, for each camera pose in the camera pose set, a three-dimensional spatial point is projected onto the corresponding tilted image data to obtain the view projection coordinates.

[0152] In some embodiments, deep semantic features are extracted from the corresponding high-dimensional feature maps based on the view projection coordinates, and temporary semantic labels are generated for the projection points under each view to form a semantic label set.

[0153] In some embodiments, the semantic label set is statistically analyzed and compared. When the inconsistency of semantic labels in the semantic label set exceeds a preset threshold, the region where the three-dimensional spatial point is located is marked as a low-confidence region.

[0154] S4. Project the 3D spatial points onto the 2D image and combine them with deep semantic features to assign semantic labels to the 3D spatial points in order to generate a 3D spatial point cloud model with semantic labels.

[0155] In some embodiments, the two-dimensional pixel coordinates of a three-dimensional spatial point in an image dataset are calculated through camera projection transformation. The camera projection transformation process includes:

[0156] In some embodiments, for each three-dimensional spatial point in the three-dimensional spatial point cloud model, the camera pose of the observed three-dimensional spatial point is determined based on the motion trajectory of a high-resolution digital camera.

[0157] In some embodiments, for each determined camera pose, a camera projection model is used to transform the three-dimensional spatial coordinates of a point in three-dimensional space into two-dimensional pixel coordinates in the corresponding image frame under the camera pose. The specific calculation formula of the camera projection model is as follows:

[0158]

[0159] in, As a scale factor, Two-dimensional pixel coordinates, For the internal calibration parameter matrix of a high-resolution digital camera, Let the rotation matrix be the camera pose. Let be the translation vector of the camera pose. This represents the camera extrinsic parameter matrix, which is composed of the rotation matrix and the translation vector. These are three-dimensional spatial coordinates.

[0160] In some embodiments, each three-dimensional spatial point is assigned a semantic label, and the assignment process includes:

[0161] In some embodiments, based on the two-dimensional pixel coordinates, the deep semantic feature vector of the two-dimensional pixel coordinates is queried and extracted from the high-dimensional feature map.

[0162] In some embodiments, a feature fusion algorithm is used to transform deep semantic feature vectors into fused feature vectors.

[0163] In some embodiments, the fused feature vector is input into a classifier, which outputs the probability that a three-dimensional spatial point belongs to a preset category of each construction element.

[0164] In some embodiments, the construction element category with the highest probability among all construction element categories is used as the semantic label of a three-dimensional spatial point.

[0165] In some embodiments, semantic labels are assigned to all three-dimensional spatial points in the three-dimensional spatial point cloud model to obtain a semantically labeled three-dimensional spatial point cloud model, wherein the three-dimensional spatial points in the semantically labeled three-dimensional spatial point cloud model contain corresponding three-dimensional spatial coordinates and semantic labels.

[0166] S5. Compare the semantically labeled 3D spatial point cloud model with the dynamic semantic 3D model of the previous cycle, which serves as the baseline model, to identify semantically changing regions and regions that conform to the preset construction logic of semantic state transitions, and construct the dynamic semantic 3D model of the current cycle. In the first cycle, the semantically labeled 3D spatial point cloud model is used as the initial baseline model.

[0167] In some embodiments, it is determined whether a dynamic semantic 3D model from the previous cycle exists.

[0168] In some embodiments, if there is no dynamic semantic 3D model for the previous period, the current period is determined to be the first period, and the 3D spatial point cloud model with semantic labels is designated as the dynamic semantic 3D model for the current period.

[0169] In some embodiments, if a dynamic semantic 3D model of the previous period exists, it is determined that the current period is not the first period. The dynamic semantic 3D model of the previous period is used as the benchmark model, and the 3D spatial point cloud model with semantic labels is compared with the benchmark model to obtain the change state identifier.

[0170] In some embodiments, a dynamic semantic 3D model for the current period is constructed based on the change state identifier.

[0171] In some embodiments, the constructed dynamic semantic 3D model for the current period is archived as a baseline model for the next period.

[0172] In some embodiments, the process of comparing a semantically labeled 3D spatial point cloud model with a benchmark model includes:

[0173] In some embodiments, a nearest neighbor search algorithm is used to find and establish corresponding relationships between 3D spatial points in a semantically labeled 3D spatial point cloud model in the current period and in a baseline model.

[0174] In some embodiments, for each pair of three-dimensional spatial points consisting of correspondences, a change recognition calculation is performed to obtain a change state identifier, wherein the change state identifier includes geometric change, semantic change, and semantic state transition identifiers. The execution process of the change recognition calculation includes:

[0175] In some embodiments, the Euclidean distance between two three-dimensional spatial points in a three-dimensional spatial point pair is calculated, and when the Euclidean distance is greater than a preset geometric change threshold, the three-dimensional spatial point pair is marked as having undergone geometric change.

[0176] In some embodiments, the semantic labels of two three-dimensional spatial points in a three-dimensional spatial point pair are compared, and when the semantic labels are inconsistent, the three-dimensional spatial point pair is marked as semantically changed.

[0177] In some embodiments, when a three-dimensional spatial point pair is marked as a semantic change, it is further determined whether the change of the corresponding semantic label conforms to the valid process flow rules in the preset construction process logic knowledge base. If the change of the semantic label conforms to the valid process flow rules, then an additional semantic state transition identifier is added to the three-dimensional spatial point pair.

[0178] In some embodiments, the process of constructing the dynamic semantic 3D model for the current period includes:

[0179] In some embodiments, the basic data structure is a three-dimensional spatial point cloud model with semantic labels for the current period.

[0180] In some embodiments, a change status identifier is attached to each three-dimensional spatial point in the basic data structure.

[0181] S6. By comparing the current period's dynamic semantic 3D model with the benchmark model, a quantitative analysis report on construction progress is generated.

[0182] In some embodiments, a semantic region category specified by the user as the computation target is received.

[0183] In some embodiments, when the current period is not the first period, three-dimensional spatial points that match the semantic labels and semantic region categories are selected from the dynamic semantic three-dimensional model and the baseline model of the current period, respectively, to form the target point cloud of the current period and the baseline target point cloud.

[0184] In some embodiments, on a horizontal two-dimensional plane covering the current period target point cloud and the reference target point cloud, a plane with a preset area is defined. A regular grid composed of grid cells.

[0185] In some embodiments, the current periodic target point cloud and the reference target point cloud are projected onto a regular grid, and for each grid cell in the regular grid, the average elevation value of the three-dimensional spatial points falling within the grid cell is taken, and the current elevation value is calculated accordingly. and benchmark elevation value Generate the current digital surface model and the baseline digital surface model.

[0186] In some embodiments, the volume change of each grid cell in the regular grid is calculated one by one. and the volume change of each grid cell. By performing algebraic summation, the total volume change of the semantic region categories is obtained.

[0187] In some embodiments, the total volume change, semantic region category, and cycle time information are integrated into a quantitative analysis report of construction progress.

[0188] S7. Overlay the dynamic semantic 3D model and the baseline model, and determine the key areas of concern based on the quantitative analysis report of construction progress, low confidence areas and semantic state transition areas. For the key areas of concern, generate UAV flight parameter adjustment instructions by solving the flight parameter optimization cost function.

[0189] In some embodiments, the dynamic semantic 3D model and the baseline model for the current period are loaded in the graphical rendering interface of the user terminal.

[0190] In some embodiments, rendering rules are established to map the change state identifiers of three-dimensional spatial points in a dynamic semantic three-dimensional model to a preset visual style.

[0191] In some embodiments, the dynamic semantic 3D model of the current cycle is colored according to the rendering rules, wherein regions with geometric and semantic changes are displayed using preset highlight colors.

[0192] In some embodiments, the numerical data in the construction progress quantitative analysis report are overlaid on the corresponding position in the graphical rendering interface in the form of charts.

[0193] In some embodiments, the construction work plan and safety risk area data for the next cycle are obtained from the building information model of the engineering construction.

[0194] In some embodiments, the construction progress quantitative analysis report is analyzed to identify semantic region categories where the total volume change exceeds a preset analysis threshold, and the three-dimensional spatial points in the dynamic semantic three-dimensional model of the current period that match the semantic region categories are identified as the progress key point cloud.

[0195] In some embodiments, three-dimensional spatial points marked as low-confidence regions are extracted from the dynamic semantic three-dimensional model of the current period to form a low-confidence point cloud.

[0196] In some embodiments, three-dimensional spatial points with attached semantic state transition identifiers are extracted from the dynamic semantic three-dimensional model of the current cycle to form a state transition point cloud.

[0197] In some embodiments, a dynamic semantic 3D model based on historical cycles is used to generate a semantic state transition heatmap for the next cycle through a time-series prediction model based on a long short-term memory network, and to extract point clouds of expected key regions from the semantic state transition heatmap.

[0198] In some embodiments, the progress key point cloud, low confidence point cloud, state transition point cloud, expected key area point cloud and point cloud corresponding to the security risk area data are merged, and the spatial range boundary of the merged point cloud is calculated. The calculated spatial range boundary is defined as the key focus area for the next collection cycle.

[0199] In some embodiments, high confidence weights are assigned to low-confidence point clouds within a key area of ​​concern, high progress association weights are assigned to state transition point clouds, and high security monitoring weights are assigned to point clouds corresponding to data in security risk areas.

[0200] In some embodiments, a flight parameter optimization cost function is constructed with the optimization objectives of maximizing model confidence, schedule correlation, and safety monitoring coverage, and constrained by acquisition efficiency. Here, model confidence is an indicator used to quantitatively evaluate the reliability of regional information in a 3D spatial point cloud model, calculated through cross-view semantic consistency. Schedule correlation is an indicator used to measure the degree of correlation between changes in the construction site area and the project construction progress, determined by the actual semantic state transitions, the work content obtained from the construction work plan, and the semantic state transition heatmap generated by the time-series prediction model. Safety monitoring coverage is an indicator used to evaluate the effectiveness of the acquisition operation in monitoring the preset safety control area obtained from the building information model.

[0201] In some embodiments, the flight parameter optimization cost function is solved by an optimization algorithm to obtain a set of optimized flight parameters that minimize the cost value. The optimized flight parameters include flight altitude, camera tilt angle and flight path density for key areas of interest.

[0202] In some embodiments, flight parameters are optimized and integrated into UAV flight parameter adjustment commands.

[0203] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for real-time 3D modeling of engineering construction sites based on UAV aerial imagery, characterized in that, Includes the following steps: S1. Collect and preprocess the tilted image data of the current period and the high-frequency pose data of the UAV itself to obtain the preprocessed image dataset and pose dataset. S2. Obtain sparse keypoint features and deep semantic features from the image dataset, and obtain high-frequency temporal features from the pose dataset; S3. Based on the sparse key point features and the high-frequency temporal features, a visual-inertial synchronous positioning and mapping algorithm is used to generate a three-dimensional spatial point cloud model to describe the construction site in the current cycle, and cross-view semantic consistency verification is performed on the three-dimensional spatial points in the three-dimensional spatial point cloud model to identify low confidence areas. S4. Project the three-dimensional spatial points onto the two-dimensional image, and combine the depth semantic features to assign semantic labels to the three-dimensional spatial points, so as to generate the three-dimensional spatial point cloud model with semantic labels. S5. The three-dimensional spatial point cloud model with semantic labels is compared with the dynamic semantic three-dimensional model of the previous cycle, which serves as the reference model, to identify semantic change areas and semantic state transition areas that conform to the preset construction logic, and to construct the dynamic semantic three-dimensional model of the current cycle. In the first cycle, the three-dimensional spatial point cloud model with semantic labels is used as the initial reference model. S6. By comparing the current dynamic semantic 3D model with the benchmark model, a construction progress quantitative analysis report is generated; S7. Overlay the dynamic semantic 3D model and the baseline model, and determine the key areas of concern based on the construction progress quantitative analysis report, the low confidence area and the semantic state transition area. For the key areas of concern, generate UAV flight parameter adjustment instructions by solving the flight parameter optimization cost function. The process of performing cross-view semantic consistency verification on the 3D spatial points in the 3D spatial point cloud model to identify low-confidence regions includes: For each newly generated 3D spatial point in the 3D spatial point cloud model, the set of camera poses of the observed 3D spatial point is determined based on the motion trajectory of the high-resolution digital camera. For each camera pose in the set of camera poses, the three-dimensional space point is projected onto the corresponding tilted image data to obtain the view projection coordinates; Based on the view projection coordinates, the deep semantic features are extracted from the corresponding high-dimensional feature maps, and temporary semantic labels are generated for the projection points under each view, forming a semantic label set. The semantic tag set is statistically analyzed and compared. When the inconsistency of the semantic tags in the semantic tag set exceeds a preset threshold, the region where the three-dimensional spatial point is located is marked as a low-confidence region.

2. The method for real-time 3D modeling of engineering construction sites based on UAV aerial imagery according to claim 1, characterized in that, The unmanned aerial vehicle (UAV) includes a flight platform, a sensing payload, and a positioning and navigation system, wherein the flight platform is a multi-rotor UAV. The sensing payload includes a multi-axis stabilization gimbal and a high-resolution digital camera mounted on the multi-axis stabilization gimbal, wherein the multi-axis stabilization gimbal is used to control the tilt angle of the digital camera. The positioning and navigation system includes a differential global navigation satellite system receiver, a high-frequency inertial measurement unit, and a hardware time synchronization unit; The process of acquiring and preprocessing the tilted image data of the current period and the high-frequency pose data of the UAV itself to obtain the preprocessed image dataset and pose dataset includes: A Global Navigation Satellite System (GNSS) reference station is deployed at a control point with known three-dimensional coordinates at the construction site of the project. The GNSS reference station is used to send differential correction data to the differential GNSS receiver. The periodic flight path of the UAV is pre-planned through the ground control station. The flight path includes waypoints and the tilt angles corresponding to the waypoints. During the flight of the UAV, the attitude of the high-resolution digital camera is adjusted by the multi-axis stabilization gimbal so that the high-resolution digital camera acquires tilt image data at preset time intervals, wherein the UAV flies autonomously according to the periodic flight path; The differential global navigation satellite system receiver and the high-frequency inertial measurement unit respectively acquire satellite observation values ​​and inertial measurement values ​​to generate high-frequency pose data; Each time the tilted image data is acquired, a timestamp is recorded in the high-frequency pose data by the hardware time synchronization unit. Lens distortion correction is performed on the tilted image data based on the internal calibration parameters of the high-resolution digital camera; The satellite observations and inertial measurements are processed in real time by a fusion calculation algorithm to generate the six-degree-of-freedom motion trajectory of the UAV, and the six-degree-of-freedom motion trajectory is used as the high-frequency pose data. The corrected tilted image data and the high-frequency pose data are synchronized and aligned according to the timestamp, so that each frame of the tilted image data is associated with the high-frequency pose data at the corresponding moment of shooting, forming a preprocessed image dataset and pose dataset.

3. The method for real-time 3D modeling of engineering construction sites based on UAV aerial imagery according to claim 2, characterized in that, The process of obtaining sparse keypoint features and deep semantic features from the image dataset, and obtaining high-frequency temporal features from the pose dataset, includes: For each of the tilted image data in the image dataset, by comparing the intensity difference between the center pixel and the ring-shaped pixels in the neighborhood of the center pixel, the key points that maintain stability under image rotation and scale changes are identified and located. Within a local image patch surrounding each keypoint, pixel pairs are selected, and a binary intensity test is performed on each pair of pixel pairs to generate a binary vector as a descriptor for the keypoint. By combining the two-dimensional coordinates of the key points in the oblique image data and the descriptors of the key points, sparse key point features are obtained; The tilted image data is input into a pre-trained convolutional neural network with an encoder-decoder structure; The oblique image data is encoded by the encoder in the convolutional neural network. The encoding process includes abstracting the oblique image data layer by layer and converting it into a high-dimensional feature map. The high-dimensional feature map is used as a deep semantic feature, and the high-dimensional feature map is used to distinguish the numerical representation of abstract image patterns of different construction elements; The six-degree-of-freedom motion trajectories in the pose dataset are used as high-frequency temporal features.

4. The method for real-time 3D modeling of engineering construction sites based on UAV aerial imagery according to claim 3, characterized in that, The process of generating a three-dimensional spatial point cloud model describing the construction site in the current cycle, based on the sparse key point features and the high-frequency temporal features, using a visual-inertial synchronous positioning and mapping algorithm, includes: Select consecutive initial image frames from the image dataset, and obtain the sparse keypoint features and the high-frequency temporal features corresponding to the initial image frames; By comparing the descriptors contained in the sparse keypoint features, a keypoint correspondence is established between the initial image frames; Based on the key point correspondence and the high-frequency time sequence characteristics, the position vector and attitude quaternion corresponding to the initial image frame in the six-degree-of-freedom motion trajectory, the three-dimensional spatial coordinates of each corresponding key point in the initial three-dimensional spatial point cloud model, and the zero bias vectors of the accelerometer and gyroscope in the six-degree-of-freedom motion trajectory are solved simultaneously through a joint optimization algorithm to construct a three-dimensional spatial point cloud model. Using the inertial measurement values ​​between the previous image frame and the current image frame, motion prediction is performed on the camera pose of the current image frame, wherein the camera pose includes the position vector and the pose quaternion. Using the camera pose as the initial value, the precise camera pose of the current image frame is solved by minimizing the reprojection error of the three-dimensional spatial points, and the precise camera pose is added to the motion trajectory of the high-resolution digital camera. The joint cost function, which includes the reprojection error and the inertial measurement error between the continuous camera poses, is minimized by a nonlinear optimization algorithm to output a globally consistent motion trajectory of the high-resolution digital camera and the three-dimensional spatial point cloud model. For sparse key points that are stably tracked under the camera pose but are not included in the three-dimensional spatial point cloud model, triangulation is performed to generate the three-dimensional spatial coordinates of the sparse key points, and the generated three-dimensional spatial points are added to the three-dimensional spatial point cloud model to obtain the three-dimensional spatial point cloud model of the construction site in the current cycle.

5. The method for real-time 3D modeling of engineering construction sites based on UAV aerial imagery according to claim 1, characterized in that, The process of projecting the three-dimensional spatial points onto a two-dimensional image and combining the depth semantic features to assign semantic labels to the three-dimensional spatial points in order to generate a semantically labeled three-dimensional spatial point cloud model includes: For each three-dimensional point in the three-dimensional point cloud model, the camera pose of the observed three-dimensional point is determined based on the motion trajectory of the high-resolution digital camera. For each determined camera pose, a camera projection model is used to transform the three-dimensional spatial coordinates of the three-dimensional spatial points into two-dimensional pixel coordinates in the corresponding image frame under the camera pose. Based on the two-dimensional pixel coordinates, the deep semantic feature vector of the two-dimensional pixel coordinates is queried and extracted from the high-dimensional feature map; The deep semantic feature vector is transformed into a fused feature vector through a feature fusion algorithm; The fused feature vector is input into a classifier, which outputs the probability that the three-dimensional spatial point belongs to each preset construction element category. The construction element category with the highest probability among the probabilities of each construction element category is used as the semantic label of the three-dimensional spatial point. The process continues until semantic labels are assigned to all three-dimensional spatial points in the three-dimensional spatial point cloud model, resulting in a semantically labeled three-dimensional spatial point cloud model. The three-dimensional spatial points in the semantically labeled three-dimensional spatial point cloud model contain the corresponding three-dimensional spatial coordinates and the semantic labels.

6. The method for real-time 3D modeling of engineering construction sites based on UAV aerial imagery according to claim 5, characterized in that, The process of comparing the semantically labeled 3D spatial point cloud model with the dynamic semantic 3D model of the previous period, which serves as a baseline model, to identify semantically changing regions and regions with semantic state transitions conforming to preset construction logic, and constructing the dynamic semantic 3D model of the current period includes: Determine whether a dynamic semantic 3D model from the previous cycle exists; If there is no dynamic semantic 3D model in the previous period, the current period is determined to be the first period, and the 3D spatial point cloud model with semantic labels is designated as the dynamic semantic 3D model in the current period. If a dynamic semantic 3D model from the previous period exists, then the current period is determined not to be the first period. The dynamic semantic 3D model from the previous period is used as the baseline model. The 3D spatial point cloud model with semantic labels is compared with the baseline model to obtain the change state identifier. Based on the changed state identifier, construct the dynamic semantic three-dimensional model for the current period; The constructed dynamic semantic 3D model for the current cycle is archived as a baseline model for the next cycle. The process of comparing the semantically labeled 3D spatial point cloud model with the benchmark model includes: Using the nearest neighbor search algorithm, the corresponding relationships between the three-dimensional spatial points in the three-dimensional spatial point cloud model with semantic labels in the current period are found and established in the benchmark model. For each pair of three-dimensional spatial points formed by the aforementioned correspondence, a change recognition calculation is performed to obtain a change state identifier, wherein the change state identifier includes geometric change, semantic change, and semantic state transition identifiers. The execution process of the change recognition calculation includes: Calculate the Euclidean distance between two three-dimensional spatial points in the three-dimensional spatial point pair. When the Euclidean distance is greater than a preset geometric change threshold, mark the three-dimensional spatial point pair as the geometric change. By comparing the semantic labels of the two three-dimensional spatial points in the three-dimensional spatial point pair, when the semantic labels are inconsistent, the three-dimensional spatial point pair is marked as the semantic change; When the three-dimensional spatial point pair is marked as the semantic change, it is further determined whether the change of the corresponding semantic tag conforms to the valid process flow rules in the preset construction process logic knowledge base. If the change of the semantic tag conforms to the valid process flow rules, then the semantic state transition identifier is added to the three-dimensional spatial point pair. The construction process of the dynamic semantic 3D model in the current period includes: The three-dimensional spatial point cloud model with semantic labels in the current period is used as the basic data structure; For each of the three-dimensional spatial points in the basic data structure, attach the change state identifier.

7. The method for real-time 3D modeling of engineering construction sites based on UAV aerial imagery according to claim 6, characterized in that, The process of generating a quantitative analysis report of construction progress by comparing the current period's dynamic semantic 3D model with the baseline model includes: Receive the semantic region category specified by the user as the computation target; If the current period is not the first period, the three-dimensional spatial points that match the semantic label and the semantic region category are selected from the dynamic semantic three-dimensional model and the baseline model of the current period respectively, to form the target point cloud of the current period and the baseline target point cloud; On a horizontal two-dimensional plane covering the current period target point cloud and the reference target point cloud, a regular grid composed of grid cells with a preset area is defined; The current periodic target point cloud and the reference target point cloud are projected onto the regular grid, and for each grid cell in the regular grid, the average elevation value of the three-dimensional space points falling into the grid cell is taken, and the current elevation value and the reference elevation value are calculated respectively to generate the current digital surface model and the reference digital surface model. The volume change of each grid cell in the regular grid is calculated one by one, and the volume changes of each grid cell are summed algebraically to obtain the total volume change of the semantic region category. The total volume change, the semantic region category, and the periodic time information are integrated into a quantitative analysis report of construction progress.

8. The method for real-time 3D modeling of engineering construction sites based on UAV aerial imagery according to claim 7, characterized in that, The process of overlaying and rendering the dynamic semantic 3D model and the baseline model includes: In the graphical rendering interface of the user terminal, the dynamic semantic 3D model and the baseline model for the current cycle are loaded. Establish rendering rules, which are used to map the change state identifiers of three-dimensional spatial points in the dynamic semantic three-dimensional model to a preset visual style; According to the rendering rules, the dynamic semantic 3D model of the current period is colored, wherein the regions with geometric changes and semantic changes are displayed using a preset highlight color; The numerical data from the construction progress quantitative analysis report are displayed in the form of charts and graphs overlaid on the corresponding positions in the graphical rendering interface.

9. The method for real-time 3D modeling of engineering construction sites based on UAV aerial imagery according to claim 7, characterized in that, The process of determining key areas of concern based on the construction progress quantitative analysis report, the low-confidence areas, and the semantic state transition areas, and then generating UAV flight parameter adjustment instructions for these key areas of concern by solving the flight parameter optimization cost function, includes: Obtain the construction operation plan and safety risk area data for the next cycle from the building information model of the engineering construction. Analyze the construction progress quantitative analysis report, identify the semantic region category where the total volume change exceeds the preset analysis threshold, and determine the three-dimensional spatial points in the dynamic semantic three-dimensional model of the current period that match the semantic region category as the progress key point cloud; From the dynamic semantic 3D model of the current period, extract the 3D spatial points marked as the low confidence region to form a low confidence point cloud; From the dynamic semantic 3D model of the current cycle, extract the 3D spatial points that have been attached with the semantic state transition identifier to form a state transition point cloud; Based on the historical cycle, the dynamic semantic 3D model generates a semantic state transition heatmap expected for the next cycle through a time-series prediction model based on a long short-term memory network, and extracts the expected key region point cloud from the semantic state transition heatmap. The progress key point cloud, the low confidence point cloud, the state transition point cloud, the expected key area point cloud, and the point cloud corresponding to the security risk area data are merged, and the spatial range boundary of the merged point cloud is calculated. The calculated spatial range boundary is defined as the key focus area for the next collection cycle. Assign high confidence weights to the low confidence point clouds within the key attention area, assign high progress association weights to the state transition point clouds, and assign high security monitoring weights to the point clouds corresponding to the security risk area data. A flight parameter optimization cost function is constructed with the optimization objectives of maximizing model confidence, progress correlation and safety monitoring coverage, and with data acquisition efficiency as a constraint. The flight parameter optimization cost function is solved by an optimization algorithm to obtain a set of optimized flight parameters that minimize the cost value. The optimized flight parameters include flight altitude, camera tilt angle, and flight path density for the key area of ​​interest. The optimized flight parameters are then integrated into UAV flight parameter adjustment commands.