Historic building surveying method and system based on multi-unmanned aerial vehicle cooperation
By employing a multi-UAV collaborative flight and adaptive acquisition strategy, a unified image dataset is generated and 3D sparse reconstruction is performed. This solves the problems of data alignment difficulties and insufficient model quality in historical building surveying, and achieves efficient and accurate 3D real-scene model reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LIAONING PROVINCIAL BUILDING DESIGN & RES INST
- Filing Date
- 2026-05-19
- Publication Date
- 2026-06-16
AI Technical Summary
Existing technologies for 3D mapping of historical buildings suffer from difficulties in aligning multi-view image data, complex and inefficient data processing, inability to handle complex geometric shapes, and lack of ability to evaluate model quality in real time and dynamically adjust acquisition strategies, resulting in hollow models, blurred textures, or geometric distortions.
By scheduling multiple drones equipped with different sensors to perform autonomous orbiting flights, a collaborative image dataset with a unified timestamp and spatial coordinates is generated. This dataset is then used for extracting corresponding feature points and performing 3D sparse reconstruction. The geometric complexity distribution of building surfaces is calculated, and adaptive planning is used to supplement flight tracks for secondary image acquisition. Finally, a high-precision 3D real-world model is generated.
It achieves spatiotemporal synchronization of multi-source information, simplifies the data processing flow, improves the accuracy and detail of the 3D model, avoids the time and energy consumption of blindly repeating scans, and enhances the level of surveying automation.
Smart Images

Figure CN122223243A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multi-UAV collaborative 3D mapping technology, specifically a method and system for historical building mapping based on multi-UAV collaboration. Background Technology
[0002] 3D mapping of historical buildings is a crucial step in cultural relic preservation and digital archiving. Current technologies typically employ a single drone to take multiple flights along a pre-set path, or use multiple identical drones to sequentially acquire images of the target. These methods result in temporal and spatial discrepancies in the acquired image data, leading to difficulties in aligning multi-view images and complex and inefficient data processing procedures.
[0003] Fixed flight paths or simple repetitive flights are insufficient to handle the complex geometry of historical building surfaces. Areas rich in detail, such as decorative reliefs, damaged structures, and eaves and brackets, often exhibit hollow models, blurred textures, or geometric distortions after initial reconstruction. Existing solutions lack the ability to assess model quality in real time and dynamically adjust acquisition strategies during flight, often requiring manual intervention and experience-based follow-up flights, which is not only inefficient but also lacks systematic assurance of coverage for high-precision areas.
[0004] Systematically scheduling heterogeneous UAV resources to achieve efficient and synchronous data acquisition, and then intelligently identifying and accurately supplementing data collection in weak areas of reconstruction, has become a key challenge in improving the automation level and model accuracy of historical building surveying. This requires the technical solution to be able to coordinate multi-source data and possess the ability to autonomously analyze and make decisions based on preliminary reconstruction results. Summary of the Invention
[0005] This invention aims to solve at least one of the technical problems existing in the prior art; Therefore, this invention proposes a historical building surveying method based on multi-UAV collaborative surveying, including: Multiple drones were deployed to autonomously circle the target historical building, and each drone was controlled to simultaneously collect multi-view image data of the target historical building. Each drone was equipped with a different combination of sensors. Spatiotemporal benchmarks are unified for multi-view image data collected by multiple drones to generate a collaborative image dataset with unified timestamps and spatial coordinates. The collaborative image dataset is subjected to corresponding feature point extraction and 3D sparse reconstruction processing to generate an initial 3D sparse point cloud model of the target historical building. Based on the initial three-dimensional sparse point cloud model, the geometric complexity distribution of the target historical building surface is calculated; Based on the geometric complexity distribution, an adaptive supplementary acquisition trajectory is generated for each UAV, and the UAV is controlled to perform secondary image acquisition in areas with high geometric complexity along the supplementary acquisition trajectory to generate a supplementary image dataset. By fusing the collaborative image dataset and the supplementary image dataset, dense matching and surface reconstruction calculations are performed to generate a 3D real-world model of the target historical building.
[0006] Furthermore, the method of scheduling multiple drones to autonomously orbit the target historical building and controlling each drone to simultaneously collect multi-view image data of the target historical building includes: Obtain the digital surface model of the target historical building and the preset orbital flight radius to generate a basic orbital trajectory; Based on the number of drones participating in the collaboration, the basic orbital trajectory is divided into equal phases, and each drone is assigned a unique starting phase point and orbital flight direction. For each UAV, plan an approach trajectory from the takeoff point to its corresponding starting phase point, and a departure trajectory from the end phase point back to the takeoff point. Then, stitch together the corresponding segments of the approach trajectory, the basic orbital trajectory, and the departure trajectory to form the complete flight mission trajectory for each UAV. During flight, unified collaborative control commands are used to synchronously trigger all UAVs to collect images at preset spatial points on their respective flight paths, ensuring the temporal and spatial correlation of multi-view image data.
[0007] Furthermore, the process of unifying the spatiotemporal reference of multi-view image data collected by multiple drones to generate a collaborative image dataset with unified timestamps and spatial coordinates includes: Receive raw image data packets uploaded by each drone. Each raw image data packet contains image data, drone positioning and attitude determination system data at the time of acquisition, and device local timestamp. Based on a preset time synchronization protocol, the local time of each UAV is calibrated to a unified system time reference, and the local timestamp of the image is converted into a unified system timestamp according to the calibration relationship. Using the UAV positioning and attitude determination system data corresponding to each image after time stamp conversion, and combined with the sensor's internal orientation parameters, the external orientation elements of all images are uniformly calculated to the same spatial coordinate system. The image data that has undergone timestamp conversion and spatial coordinate system unification is indexed and organized according to the unified system timestamp order to generate the collaborative image dataset.
[0008] Further, the step of extracting corresponding feature points and performing 3D sparse reconstruction processing on the collaborative image dataset to generate an initial 3D sparse point cloud model of the target historical building includes: The scale-invariant feature transform operator is used to extract the images in the collaborative image dataset to obtain a set of local feature descriptors for each image; Based on the similarity of feature descriptors, feature matching is performed among all images to establish the correspondence between image pairs of the same feature points. Using multi-view geometric constraints, relative orientation calculations are performed on image pairs containing the corresponding relationship of the same-name feature points to restore the relative position and pose between the images; An incremental motion recovery structure algorithm is used to gradually add images, and bundle adjustment is used to optimize the exterior orientation elements and three-dimensional spatial coordinates of corresponding feature points of all images to form the initial three-dimensional sparse point cloud model.
[0009] Furthermore, the calculation of the geometric complexity distribution of the target historical building surface based on the initial three-dimensional sparse point cloud model includes: In the initial three-dimensional sparse point cloud model, a three-dimensional regular voxel mesh is established to divide the space into multiple voxel units; The number of sparse point cloud points falling into each voxel unit is counted and used as the point density of the corresponding voxel. Calculate the degree of directional change between the normal vectors of each point within each voxel, and use it as the rate of change of the normal vector of the corresponding voxel; The geometric complexity value of the corresponding voxel is obtained by weighted calculation based on the point density and normal change rate of each voxel. The geometric complexity values of all voxels are mapped back to three-dimensional space to form a geometric complexity distribution field covering the target historical building.
[0010] Furthermore, generating adaptive re-collection tracks for each UAV based on the geometric complexity distribution includes: In the geometric complexity distribution field, a complexity threshold is set, and regions with geometric complexity values higher than the complexity threshold are selected and marked as regions to be supplemented. For each area to be supplemented, its three-dimensional spatial bounding box is calculated, and based on a preset safety distance, an outer observation envelope that completely surrounds the area to be supplemented is generated. Assign one or more areas to be sampled to each UAV, and generate a series of discrete observation viewpoints on the corresponding outer observation envelope plane of each assigned area, in accordance with the requirements of image overlap and observation angle. Multiple observation points assigned to the same UAV are used to plan paths according to the principle of shortest flight distance, generating the UAV's supplementary data collection trajectory.
[0011] Furthermore, the controlled UAV performs secondary image acquisition along the supplementary acquisition trajectory for areas with high geometric complexity, generating a supplementary image dataset, including: The replenishment flight path is sent to the corresponding UAV, and the UAV is controlled to fly autonomously along the replenishment flight path; During the flight of the UAV, when it reaches the preset observation viewpoint on the supplementary data acquisition track, the onboard sensors are automatically triggered to acquire images of the designated area to be supplemented. Record the UAV positioning and attitude determination system data and system timestamp during the secondary image acquisition, and bind them to the acquired image data; After completing all flight missions for supplementary data acquisition, the image data and corresponding positioning and attitude information obtained by all UAVs during the secondary acquisition phase are summarized to generate the supplementary image dataset.
[0012] Furthermore, the process of fusing the collaborative image dataset and the supplementary image dataset, performing dense matching and surface reconstruction calculations, and generating a 3D reality model of the target historical building includes: The supplementary image dataset is merged with the collaborative image dataset, and global bundle adjustment is performed on all merged images to obtain the exterior orientation elements of all optimized images. Based on the optimized exterior orientation elements, multi-view stereo dense matching is performed on the surface of the target historical building. The corresponding three-dimensional spatial coordinates of each pixel are calculated to generate a high-density three-dimensional dense point cloud. Statistical filtering and radius filtering are performed on the three-dimensional dense point cloud to remove outlier noise points; The denoised 3D dense point cloud is processed into a triangular mesh to construct a triangular mesh surface model of the target historical building. The color information of the image is mapped onto each triangular face of the triangular mesh surface model to generate a 3D real-world model with realistic textures; The process of performing multi-view stereo dense matching on the surface of the target historical building, and calculating the corresponding three-dimensional spatial coordinates for each pixel, includes: On each image, a regularly distributed image-square matching grid is established; For each grid point on the reference image, under its epipolar geometric constraints, search for corresponding points on multiple adjacent images; A semi-global matching algorithm is used to calculate the optimal disparity value for each grid point on the reference image; Based on the optimal parallax value, the exterior orientation elements of the reference image, and the exterior orientation elements of the adjacent images, the forward intersection is used to calculate the object-space three-dimensional coordinates corresponding to the grid points. By traversing all grid points on all images, the three-dimensional spatial coordinates are calculated, and all coordinate points are aggregated to form the high-density three-dimensional dense point cloud.
[0013] Furthermore, the process of triangulating the denoised 3D dense point cloud to construct a triangular mesh surface model of the target historical building includes: The denoised 3D dense point cloud is spatially partitioned, and a 3D spatial index structure is constructed. Based on the theory of Poisson surface reconstruction, an indicator function is constructed starting from the three-dimensional dense point cloud and its normal vector; Extract the isosurface of the indicator function; the isosurface is the potential surface of the target historical building. The extracted isosurfaces are discretized into a mesh structure consisting of a series of vertices and triangular patches; The mesh structure is adaptively simplified and smoothed, and the mesh topology is optimized while preserving the features to generate the final triangular mesh surface model.
[0014] Furthermore, the present invention also includes a historical building surveying system based on multi-UAV collaboration, the system including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, it implements the steps of the historical building surveying method based on multi-UAV collaboration described above.
[0015] Compared with the prior art, the beneficial effects of the present invention are: By scheduling a swarm of drones equipped with different sensors to perform synchronized autonomous orbiting flights, multi-view, multi-spectral imagery data of the target building can be acquired in a single operation. This method achieves spatiotemporal synchronization of multi-source information from the data acquisition source, eliminating the data alignment and fusion challenges associated with traditional time-sharing and device-based acquisition. The resulting collaborative imagery dataset possesses an inherently unified spatiotemporal reference, simplifying subsequent data processing and providing a high-quality heterogeneous data foundation for constructing information-rich, geometrically and texturally consistent 3D models.
[0016] After the initial collaborative flight and generation of an initial 3D sparse point cloud model, the algorithm automatically analyzes the model surface, calculates the geometric complexity of each region, and accurately identifies areas rich in detail or structurally complex, such as reliefs and eaves. Based on this complexity distribution, the system then plans and executes a dedicated data acquisition path for each UAV, directed towards high-complexity areas. This adaptive acquisition strategy, based on model feedback, precisely targets limited secondary flight resources to weak points in model reconstruction, achieving targeted high-resolution data augmentation for complex structures. This not only improves the completeness and accuracy of the final model in terms of detail but also avoids the time and energy consumption caused by blindly and repeatedly scanning the entire building. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the steps of the historical building surveying method based on multi-UAV collaboration described in this invention. Figure 2 A flowchart illustrating multi-UAV orbital flight and synchronized image acquisition; Figure 3 A flowchart for 3D sparse reconstruction of a collaborative image dataset; Figure 4 A schematic diagram of the two-dimensional thermal distribution of the geometric complexity of the surface of a historical building; Figure 5 The curves show the impact of smooth iteration of triangular meshes on average side length and mesh quality. Detailed Implementation
[0018] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] See Figure 1 Multiple drones were deployed to autonomously orbit a target historical building, simultaneously acquiring multi-view image data from each drone, with each drone equipped with a different sensor combination. The multi-view image data acquired by the drones underwent spatiotemporal benchmarking to generate a collaborative image dataset with a unified timestamp and spatial coordinates. Corresponding feature point extraction and 3D sparse reconstruction were performed on the collaborative image dataset to generate an initial 3D sparse point cloud model of the target historical building. Based on the initial 3D sparse point cloud model, the geometric complexity distribution of the target historical building's surface was calculated. According to the geometric complexity distribution, adaptive supplementary acquisition tracks were generated for each drone, which was then controlled to perform secondary image acquisition along these tracks in areas with high geometric complexity, generating a supplementary image dataset. The collaborative and supplementary image datasets were fused, and dense matching and surface reconstruction calculations were performed to generate a 3D real-world model of the target historical building.
[0020] See Figure 2In one embodiment of the present invention, multiple drones are scheduled to autonomously orbit a target historical building, and each drone is controlled to synchronously acquire multi-view image data of the target historical building, wherein each drone carries a different combination of sensors. The multi-view image data acquired by the multiple drones undergoes spatiotemporal benchmark unification processing to generate a collaborative image dataset with a unified timestamp and spatial coordinates. The collaborative image dataset is then processed with corresponding feature point extraction and 3D sparse reconstruction to generate an initial 3D sparse point cloud model of the target historical building. Based on the initial 3D sparse point cloud model, the geometric complexity distribution of the target historical building's surface is calculated. According to the geometric complexity distribution, an adaptive supplementary acquisition trajectory is generated for each drone, and the drone is controlled to perform secondary image acquisition along the supplementary acquisition trajectory for areas with high geometric complexity, generating a supplementary image dataset. The collaborative image dataset and the supplementary image dataset are fused, and dense matching and surface reconstruction calculations are performed to generate a 3D real-world model of the target historical building.
[0021] The system acquires a digital surface model of the target historical building and a preset orbital flight radius. Based on the outer contour of the digital surface model, a circular baseline path is generated in the horizontal plane. This baseline path is then copied and connected vertically according to preset flight altitude layers to form a spiraling basic orbital track. Depending on the number of participating drones, the basic orbital track is divided into equal phase segments. For example, when using four drones, the initial phase circle of the basic orbital track is divided into four 90-degree phase intervals. Each drone is assigned a unique starting phase point and orbital flight direction, ensuring that the starting phase points of drones A, B, C, and D are evenly spaced along the circumference, and their flight directions are set to the same direction. An approach track is planned for each drone from its takeoff point to its corresponding starting phase point. The takeoff point is set in the same airspace at a safe distance from the target historical building. The approach track uses a straight path. A departure track is also included, returning from the end phase point to the takeoff point. The corresponding segments of the approach track, the basic orbital track, and the departure track are then pieced together in chronological order to form the complete flight mission track for each drone. During flight, unified collaborative control commands, such as trigger pulses based on high-precision time protocol network synchronization, are used to synchronously trigger all UAVs to acquire images at preset spatial points on their respective flight paths. The preset spatial points are calculated and determined according to the image overlap requirements to ensure the temporal and spatial correlation of multi-view image data.
[0022] In some embodiments, raw image data packets uploaded by each UAV are received. Each raw image data packet contains image data, UAV positioning and attitude system data at the time of acquisition, and the device's local timestamp. Based on a preset time synchronization protocol, before the flight mission, the ground control station provides time synchronization to all UAVs, establishes a deviation model between the local time of each UAV and the time of the ground control system, calibrates the local time of each UAV to a unified system time reference, and converts the device's local timestamp of the images into a unified system timestamp according to the calibration relationship. Using the UAV positioning and attitude system data corresponding to each image after timestamp conversion, combined with the sensor's internal orientation parameters, including camera focal length, principal point coordinates, and lens distortion coefficients, the external orientation elements of all images, i.e., the position and attitude of the images at the time of capture, are uniformly calculated to the same spatial coordinate system. This spatial coordinate system adopts the WGS84 geodetic coordinate system and the corresponding projected coordinate system. Optionally, the calculation process is implemented through collinearity equations, as shown in the following formula:
[0023] in: These are the coordinates of the image point in the image coordinate system. These are the coordinates of the principal point of the image. It's the camera's focal length. These are the coordinates of the object point in the object space coordinate system. These are the coordinates of the photography center in the object space coordinate system. It is the image's pose rotation matrix. This is a scaling factor. It can be understood that the formula links the image's image-side coordinates with its object-side spatial coordinates through the location of the camera center and the image's pose. Image data that has undergone timestamp transformation and spatial coordinate system unification is indexed and organized according to the unified system timestamp order to generate a collaborative image dataset. The collaborative image dataset is a database or structured file list containing image file paths, unified system timestamps, and exterior orientation element parameters.
[0024] See Figure 3In one embodiment of the present invention, multiple drones are scheduled to autonomously orbit a target historical building, and each drone is controlled to synchronously acquire multi-view image data of the target historical building, wherein each drone carries a different combination of sensors. The multi-view image data acquired by the multiple drones undergoes spatiotemporal benchmark unification processing to generate a collaborative image dataset with a unified timestamp and spatial coordinates. The collaborative image dataset is then processed with corresponding feature point extraction and 3D sparse reconstruction to generate an initial 3D sparse point cloud model of the target historical building. Based on the initial 3D sparse point cloud model, the geometric complexity distribution of the target historical building's surface is calculated. According to the geometric complexity distribution, an adaptive supplementary acquisition trajectory is generated for each drone, and the drone is controlled to perform secondary image acquisition along the supplementary acquisition trajectory for areas with high geometric complexity, generating a supplementary image dataset. The collaborative image dataset and the supplementary image dataset are fused, and dense matching and surface reconstruction calculations are performed to generate a 3D real-world model of the target historical building.
[0025] Scale-invariant feature transformation (SMT) operators are used to extract features from images in the collaborative image dataset. The SMT extraction process includes constructing a Gaussian difference scale space, detecting extrema, accurately locating keypoints, assigning keypoint orientations, and generating feature descriptors, resulting in a set of local feature descriptors for each image. This set is a 128-dimensional list of feature vectors. Based on the similarity of the feature descriptors, feature matching is performed among all images using the nearest neighbor distance ratio method. The Euclidean distance between the feature descriptor to be matched and the candidate feature descriptors is calculated. A distance ratio threshold is set; when the ratio of the nearest neighbor distance to the second nearest neighbor distance is less than this threshold, the matching pair is accepted, establishing a correspondence between image pairs of corresponding feature points. Using multi-view geometric constraints, relative orientation calculations are performed on image pairs containing corresponding feature points. By solving the essential matrix or fundamental matrix, the relative position and pose between images are recovered, completing the relative orientation of the image pairs. An incremental motion recovery structure algorithm is adopted. Starting from the initial image pair, images are added step by step. The new images are matched with the reconstructed point cloud by feature matching and spatial forward intersection. Bundle adjustment is used to optimize the exterior orientation elements and the three-dimensional spatial coordinates of corresponding feature points of all images. Bundle adjustment achieves optimization by minimizing the reprojection error, forming an initial three-dimensional sparse point cloud model. The initial three-dimensional sparse point cloud model contains a series of three-dimensional spatial point coordinates and corresponding feature description information.
[0026] In some embodiments, a three-dimensional regular voxel mesh is established in the initial three-dimensional sparse point cloud model. The three-dimensional regular voxel mesh divides the three-dimensional space where the target historical building is located into cubic grid cells with fixed side lengths. Each cubic grid cell is called a voxel cell. The number of sparse point cloud points falling into each voxel cell is counted, and the number of sparse point cloud points in each voxel cell is used as the point density of the corresponding voxel. The degree of directional change between the normal vectors of each point in each voxel is calculated. The normal vectors are obtained by fitting the local plane of the points in the voxel. The degree of directional change between the normal vectors is quantified by calculating the variance or standard deviation of the angle between the normal vectors, which is used as the normal change rate of the corresponding voxel. Based on the point density and normal change rate of each voxel, the geometric complexity value of the corresponding voxel is obtained by weighted calculation. The formula used for weighted calculation is:
[0027] in: Representative voxels The geometric complexity value, Representative voxels The normalized value of the point density. Representative voxels The normalized value of the normal rate of change, and It is a pre-set weighting coefficient, and satisfies It can be understood that the geometric complexity value comprehensively reflects the discreteness and local curvature changes of the surface point cloud. Mapping the geometric complexity values of all voxels back to three-dimensional space, and associating the geometric complexity value with the center coordinates of each voxel unit, forms a three-dimensional scalar field covering the target historical building, i.e., the geometric complexity distribution field. Optionally, for voxel units with no sparse points falling into their field, their geometric complexity value is assigned a background value.
[0028] In one embodiment of the present invention, multiple drones are scheduled to autonomously orbit a target historical building, and each drone is controlled to synchronously acquire multi-view image data of the target historical building, wherein each drone carries a different combination of sensors. The multi-view image data acquired by the multiple drones undergoes spatiotemporal benchmark unification processing to generate a collaborative image dataset with a unified timestamp and spatial coordinates. The collaborative image dataset is then processed with corresponding feature point extraction and 3D sparse reconstruction to generate an initial 3D sparse point cloud model of the target historical building. Based on the initial 3D sparse point cloud model, the geometric complexity distribution of the target historical building's surface is calculated. According to the geometric complexity distribution, an adaptive supplementary acquisition trajectory is generated for each drone, and the drone is controlled to perform secondary image acquisition along the supplementary acquisition trajectory in areas with high geometric complexity, generating a supplementary image dataset. The collaborative image dataset and the supplementary image dataset are fused, and dense matching and surface reconstruction calculations are performed to generate a 3D real-world model of the target historical building.
[0029] In the geometric complexity distribution field, a complexity threshold is set. This threshold is a predefined scalar value used to distinguish between high-complexity and low-complexity regions. Voxel elements with geometric complexity values higher than the complexity threshold are selected from the geometric complexity distribution field. The continuous regions formed by these selected voxel elements in three-dimensional space are marked as regions to be re-sampled. For each region to be re-sampled, its three-dimensional bounding box is calculated. The three-dimensional bounding box is defined by the minimum and maximum coordinate values of the region in the three axes of the Cartesian coordinate system. Based on a preset safety distance, a fixed distance is extended outside the six outer surfaces of the three-dimensional bounding box to generate a hexahedral outer observation envelope that completely surrounds the region to be re-sampled. One or more areas to be acquired are assigned to each UAV. The allocation strategy is based on the principle of spatial proximity, assigning adjacent areas to the same UAV. For each assigned area, a series of discrete observation viewpoints are generated on its corresponding outer observation envelope plane, in accordance with the requirements of image overlap and observation angle. The generation of observation viewpoints ensures that the images taken along any flight direction have a forward overlap and lateral overlap of no less than a set percentage with the images of adjacent viewpoints.
[0030] In some embodiments, multiple observation viewpoints assigned to the same UAV are path-planned according to the shortest flight distance principle. A traveling salesman problem (TSP) algorithm is used to calculate the shortest flight path to all designated observation viewpoints, and the path points are connected by a smooth curve to generate a supplementary acquisition track for the UAV. The supplementary acquisition track is then distributed to the corresponding UAV, controlling it to fly autonomously along the track. During flight, when the UAV reaches a preset observation viewpoint on the supplementary acquisition track, its onboard sensors automatically trigger image acquisition of the designated area to be acquired. The UAV's positioning and attitude determination system data and system timestamp during the secondary image acquisition are recorded and bound to the acquired image data. After completing the flight tasks of all supplementary acquisition tracks, all image data obtained by the UAVs during the secondary acquisition phase and their corresponding positioning and attitude determination information are summarized to generate a supplementary image dataset. The organization of the supplementary image dataset is the same as that of the collaborative image dataset. Optionally, flight speed and turning radius constraints of the UAV are considered during path planning. The three-dimensional coordinates of the observation viewpoints are also included. It is generated by the following formula:
[0031]
[0032]
[0033] in: and These represent the minimum and maximum values of the three-dimensional bounding box of the area to be sampled, along the X, Y, and Z axes, respectively. It is a preset safe distance. These are parameters between 0 and 1, used to parameterize and determine specific locations on the outer observation envelope by systematically changing... By setting a value and applying constraints to a specified surface (such as the front surface, back surface, etc.), a series of discrete observation point positions covering the entire outer observation envelope can be generated. In essence, the formula defines a method for generating point coordinates on the outer observation envelope (an expanded bounding box).
[0034] In one embodiment of the present invention, multiple drones are scheduled to autonomously orbit a target historical building, and each drone is controlled to synchronously acquire multi-view image data of the target historical building, wherein each drone carries a different combination of sensors. The multi-view image data acquired by the multiple drones undergoes spatiotemporal benchmark unification processing to generate a collaborative image dataset with a unified timestamp and spatial coordinates. The collaborative image dataset is then processed with corresponding feature point extraction and 3D sparse reconstruction to generate an initial 3D sparse point cloud model of the target historical building. Based on the initial 3D sparse point cloud model, the geometric complexity distribution of the target historical building's surface is calculated. According to the geometric complexity distribution, an adaptive supplementary acquisition trajectory is generated for each drone, and the drone is controlled to perform secondary image acquisition along the supplementary acquisition trajectory in areas with high geometric complexity, generating a supplementary image dataset. The collaborative image dataset and the supplementary image dataset are fused, and dense matching and surface reconstruction calculations are performed to generate a 3D real-world model of the target historical building.
[0035] The supplementary image dataset and the collaborative image dataset are merged. The merged image set includes all images acquired during the initial orbital flight and all images acquired during the secondary supplementary flight. Global bundle adjustment is performed on all merged images. The exterior orientation elements and 3D coordinates of tie points of all images acquired during the initial and secondary acquisitions are used as overall parameters. Iterative optimization is performed with the objective function of minimizing the sum of squared reprojection errors of all feature points to obtain the optimized exterior orientation elements of all images. The optimized exterior orientation elements of all images have higher consistency and accuracy. Based on the optimized exterior orientation elements, multi-view stereo dense matching is performed on the surface of the target historical building. The corresponding 3D spatial coordinates of each pixel are calculated to generate a high-density 3D dense point cloud. The number of points in the 3D dense point cloud is increased by an order of magnitude compared to the initial 3D sparse point cloud model. Statistical filtering and radius filtering are performed on the 3D dense point cloud. Statistical filtering calculates the distribution of the average distance from each point to its nearest neighbor and removes points whose distance exceeds a set standard deviation. Radial filtering counts the number of points in the neighborhood of a specified radius and removes points whose number of points in the neighborhood is less than a set threshold, thus removing outlier noise points. The denoised 3D dense point cloud is triangulated to construct a triangular mesh surface model of the target historical building. The color information of the image is mapped onto each triangular face of the triangular mesh surface model to generate a 3D real-world model with realistic texture. The color information is obtained by calculating the visibility and projection relationship of each triangular facet on various images and fusing the color values of multiple images.
[0036] In some embodiments, multi-view stereo dense matching is performed on the surface of the target historical building. The process involves calculating the corresponding 3D spatial coordinates for each pixel. On each image, a regularly distributed image-side matching grid is established, with the grid size set to a fixed pixel interval. For each grid point on the reference image, corresponding points are found on multiple adjacent images under their epipolar geometric constraints. Adjacent images are selected based on their baseline length and overlap area with the reference image. A semi-global matching algorithm is used to calculate the optimal disparity value for each grid point on the reference image. The semi-global matching algorithm aggregates matching costs along multiple paths and uses dynamic programming to solve for the optimal disparity of each pixel. Based on the optimal disparity value, the exterior orientation elements of the reference image, and the exterior orientation elements of adjacent images, forward intersection is used to calculate the object-side 3D spatial coordinates corresponding to the grid points. All grid points on all images are traversed to complete the calculation of 3D spatial coordinates, and all coordinate points are aggregated to form a high-density 3D dense point cloud. It can be understood that the cost aggregation function of the semi-global matching algorithm uses a matching cost that incorporates the absolute difference of pixel gradients, and its smoothing term penalty function is:
[0037] in: This represents the change in disparity between adjacent pixels. Smoothing penalty applied at the time and It is the penalty coefficient, and , This is an indicator function; it takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. (The formula contains...) This represents the absolute difference in disparity between two adjacent pixels. Optionally, the penalty coefficient can be adaptively adjusted during the matching process based on the texture richness of the image. and The values are shown in Table 1, which illustrates typical parameter settings for the two filtering methods when processing dense 3D point clouds.
[0038] Table 1: Examples of Typical Parameters for Filtering 3D Dense Point Clouds Filtering methods Parameter name Parameter Description Example values Statistical Filtering Number of nearest neighbors Number of nearest neighbors used to calculate average distance 50 Statistical Filtering Standard deviation multiple The standard deviation multiplier used to set the distance threshold 1.5 radius filtering Search radius Spherical search radius (unit: meters) used to count neighborhood points 0.05 radius filtering Minimum number of points threshold Minimum number of points that should be included in the neighborhood 10 It is understandable that the parameter values listed in Table 1 can be adjusted according to the actual point density and noise level of the 3D dense point cloud.
[0039] See Figure 4 The study reveals the distribution characteristics of the geometric complexity values of the target historical building's surface in planar space: Spatial distribution characteristics: The complexity values exhibit a concentric circle gradient decay pattern centered on the building's center (near the origin). The geometric complexity value is highest (close to 105) in the central region (approximately -10m to 10m in the X / Y range), gradually decreasing towards the periphery, with the complexity value approaching 0 in the edge region (where the absolute value of the X / Y coordinates is greater than 50m). This distribution pattern highly matches the typical spatial form of historical buildings, characterized by "complex central structure and relatively simple peripheral outline." Calculation basis: This distribution field is generated based on an initial 3D sparse point cloud model. By dividing the space into a regular voxel grid, the point density and normal variation rate within each voxel are statistically analyzed and weighted to obtain the geometric complexity value. The point density reflects the sampling richness of the 3D structure in the region, while the normal variation rate directly characterizes the richness of surface curvature and detailed features. Engineering significance: The distribution results provide a quantitative basis for the adaptive generation of UAV re-collection tracks. By setting a complexity threshold, high-value areas to be re-collected are selected (as shown in the red and orange core areas in the figure). Based on the 3D bounding boxes of these areas, an outer observation envelope is generated, thereby guiding UAVs to conduct targeted secondary image acquisition of building core areas with high geometric complexity and rich details. Ultimately, this improves the reconstruction accuracy and detail integrity of the 3D reality model.
[0040] In one embodiment of the present invention, multiple drones are scheduled to autonomously orbit a target historical building, and each drone is controlled to synchronously acquire multi-view image data of the target historical building, wherein each drone carries a different combination of sensors. The multi-view image data acquired by the multiple drones undergoes spatiotemporal benchmark unification processing to generate a collaborative image dataset with a unified timestamp and spatial coordinates. The collaborative image dataset is then processed with corresponding feature point extraction and 3D sparse reconstruction to generate an initial 3D sparse point cloud model of the target historical building. Based on the initial 3D sparse point cloud model, the geometric complexity distribution of the target historical building's surface is calculated. According to the geometric complexity distribution, an adaptive supplementary acquisition trajectory is generated for each drone, and the drone is controlled to perform secondary image acquisition along the supplementary acquisition trajectory in areas with high geometric complexity, generating a supplementary image dataset. The collaborative image dataset and the supplementary image dataset are fused, and dense matching and surface reconstruction calculations are performed to generate a 3D real-world model of the target historical building.
[0041] The denoised 3D dense point cloud is spatially partitioned, and a 3D spatial index structure is constructed. This index structure uses an octree, recursively dividing the point cloud space into eight sub-cubes until the number of points in each sub-cube falls below a set threshold. Based on Poisson surface reconstruction theory, starting from the 3D dense point cloud and its normal vectors, the normal vector of each point in the 3D dense point cloud is calculated. The normal vector can be obtained by estimating the distribution of its neighboring points through principal component analysis. An indicator function is constructed, which is positive inside the object, negative outside the object, and zero on the object's surface. The isosurfaces of the indicator function are extracted; these isosurfaces represent the potential surfaces of the target historical building. A moving cube algorithm is used to traverse the 3D spatial mesh to extract isosurfaces with an indicator function value of zero. These isosurfaces are approximated by multiple triangular facets. The extracted isosurfaces are discretized into a mesh structure composed of a series of vertices and triangular facets. The vertices in the mesh structure lie on the isosurfaces, and the triangular facets connect these vertices to approximate the surface. The mesh structure is adaptively simplified and smoothed. Adaptive simplification reduces the number of triangular faces by iteratively shrinking the edges that contribute the least to the model shape. Smoothing uses the Laplacian smoothing algorithm to adjust the position of vertices to improve mesh quality. The mesh topology is optimized while preserving features to generate the final triangular mesh surface model.
[0042] In some embodiments, constructing the indicator function is central to Poisson surface reconstruction. The indicator function is obtained by solving a Poisson equation. Solving the Poisson equation involves discretizing the three-dimensional space into a voxel mesh and solving for each voxel. Indicator function gradient field Set as the normal vector field of the three-dimensional dense point cloud. The indicator function is solved by minimizing the difference between the gradient field and the vector field. In a concrete implementation, the indicator function can be obtained by solving a linear system of the following form:
[0043] in: It is the gradient operator. It is the interpolated vector field of the normal vectors of a 3D dense point cloud on discrete voxels. This minimization problem can be understood as ultimately transforming into solving a system of linear equations concerning the values of indicator functions on the voxels. Optionally, to accelerate the solution process, a multigrid method or a conjugate gradient method can be used to solve the above linear system. In the smoothing process, Laplace smoothing achieves a smoothing effect by shifting the position of each vertex towards the average of its neighboring vertex positions; the number of smoothing iterations is set according to the initial quality of the mesh. The final triangular mesh surface model is a watertight set of triangular facets with a reasonable topological structure, capable of accurately representing the geometry of the target historical building.
[0044] See Figure 5 During the triangular mesh smoothing iteration process, the average side length and mesh quality score exhibit a significant negative correlation evolution trend. Specifically, as the number of smoothing iterations gradually increases from 0 to 5, the average side length of the triangular facets continuously decreases from the initial 2.50 cm to 1.70 cm. This change directly reflects the process by which the Laplace smoothing algorithm iteratively shrinks the vertex neighborhood, making the mesh cell size tend to be uniform. At the same time, the mesh quality score steadily increases from 0.65 to 0.90, indicating that the smoothing process effectively optimizes the shape and topology of the triangular facets and reduces the proportion of deformed facets. From the evolutionary analysis, the first 3 iterations are the key stage for performance improvement: the rate of decrease in average side length and the rate of increase in quality score are both at a high level, indicating that the smoothing algorithm is most efficient at correcting mesh defects during this stage. However, in the 3rd to 5th iterations, the changes in both indicators slow down significantly, reflecting the diminishing marginal effect of smoothing. When the mesh quality reaches a certain threshold, the contribution of further iterations to side length uniformity and shape optimization decreases significantly, while also avoiding the loss of architectural geometric features caused by over-smoothing. This curve relationship provides a quantitative basis for the post-processing of triangular mesh surface models: depending on the target application scenario (such as high-precision texture mapping or lightweight visualization), the optimal balance point between the average side length and mesh quality can be selected, thereby improving the rendering and analysis efficiency of the model while ensuring geometric accuracy.
[0045] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A method for surveying and mapping historical buildings based on multi-UAV collaboration, characterized in that, The method includes: Multiple drones were deployed to autonomously circle the target historical building, and each drone was controlled to simultaneously collect multi-view image data of the target historical building. Each drone was equipped with a different combination of sensors. Spatiotemporal benchmarks are unified for multi-view image data collected by multiple drones to generate a collaborative image dataset with unified timestamps and spatial coordinates. The collaborative image dataset is subjected to corresponding feature point extraction and 3D sparse reconstruction processing to generate an initial 3D sparse point cloud model of the target historical building. Based on the initial three-dimensional sparse point cloud model, the geometric complexity distribution of the target historical building surface is calculated; Based on the geometric complexity distribution, an adaptive supplementary acquisition trajectory is generated for each UAV, and the UAV is controlled to perform secondary image acquisition in areas with high geometric complexity along the supplementary acquisition trajectory to generate a supplementary image dataset. By fusing the collaborative image dataset and the supplementary image dataset, dense matching and surface reconstruction calculations are performed to generate a 3D real-world model of the target historical building.
2. The historical building surveying method based on multi-UAV collaborative mapping according to claim 1, characterized in that, The method of scheduling multiple drones to autonomously circle the target historical building and controlling each drone to simultaneously collect multi-view image data of the target historical building includes: Obtain the digital surface model of the target historical building and the preset orbital flight radius to generate a basic orbital trajectory; Based on the number of drones participating in the collaboration, the basic orbital trajectory is divided into equal phases, and each drone is assigned a unique starting phase point and orbital flight direction. For each UAV, plan an approach trajectory from the takeoff point to its corresponding starting phase point, and a departure trajectory from the end phase point back to the takeoff point. Then, stitch together the corresponding segments of the approach trajectory, the basic orbital trajectory, and the departure trajectory to form the complete flight mission trajectory for each UAV. During flight, unified collaborative control commands are used to synchronously trigger all UAVs to collect images at preset spatial points on their respective flight paths, ensuring the temporal and spatial correlation of multi-view image data.
3. The historical building surveying method based on multi-UAV collaborative mapping according to claim 1, characterized in that, The process of unifying the spatiotemporal reference of multi-view image data collected by multiple drones to generate a collaborative image dataset with unified timestamps and spatial coordinates includes: Receive raw image data packets uploaded by each drone. Each raw image data packet contains image data, drone positioning and attitude determination system data at the time of acquisition, and device local timestamp. Based on a preset time synchronization protocol, the local time of each UAV is calibrated to a unified system time reference, and the local timestamp of the image is converted into a unified system timestamp according to the calibration relationship. Using the UAV positioning and attitude determination system data corresponding to each image after time stamp conversion, and combined with the sensor's internal orientation parameters, the external orientation elements of all images are uniformly calculated to the same spatial coordinate system. The image data that has undergone timestamp conversion and spatial coordinate system unification is indexed and organized according to the unified system timestamp order to generate the collaborative image dataset.
4. The historical building surveying method based on multi-UAV collaborative mapping according to claim 1, characterized in that, The step of extracting corresponding feature points and performing 3D sparse reconstruction on the collaborative image dataset to generate an initial 3D sparse point cloud model of the target historical building includes: The scale-invariant feature transform operator is used to extract the images in the collaborative image dataset to obtain a set of local feature descriptors for each image; Based on the similarity of feature descriptors, feature matching is performed among all images to establish the correspondence between image pairs of corresponding feature points. Using multi-view geometric constraints, relative orientation calculations are performed on image pairs containing the corresponding relationship of the same-name feature points to restore the relative position and pose between the images; An incremental motion recovery structure algorithm is used to gradually add images, and bundle adjustment is used to optimize the exterior orientation elements and three-dimensional spatial coordinates of corresponding feature points of all images to form the initial three-dimensional sparse point cloud model.
5. The historical building surveying method based on multi-UAV collaborative mapping according to claim 1, characterized in that, The calculation of the geometric complexity distribution of the target historical building surface based on the initial three-dimensional sparse point cloud model includes: In the initial three-dimensional sparse point cloud model, a three-dimensional regular voxel mesh is established to divide the space into multiple voxel units; The number of sparse point cloud points falling into each voxel unit is counted and used as the point density of the corresponding voxel. Calculate the degree of directional change between the normal vectors of each point within each voxel, and use it as the rate of change of the normal vector of the corresponding voxel; The geometric complexity value of the corresponding voxel is obtained by weighted calculation based on the point density and normal change rate of each voxel. The geometric complexity values of all voxels are mapped back to three-dimensional space to form a geometric complexity distribution field covering the target historical building.
6. The historical building surveying method based on multi-UAV collaborative mapping according to claim 1, characterized in that, The step of generating adaptive re-collection tracks for each UAV based on the geometric complexity distribution includes: In the geometric complexity distribution field, a complexity threshold is set, and regions with geometric complexity values higher than the complexity threshold are selected and marked as regions to be supplemented. For each area to be supplemented, its three-dimensional spatial bounding box is calculated, and based on a preset safety distance, an outer observation envelope that completely surrounds the area to be supplemented is generated. Assign one or more areas to be sampled to each UAV, and generate a series of discrete observation viewpoints on the corresponding outer observation envelope plane of each assigned area, in accordance with the requirements of image overlap and observation angle. Multiple observation points assigned to the same UAV are used to plan paths according to the principle of shortest flight distance, generating the UAV's supplementary data collection trajectory.
7. The historical building surveying method based on multi-UAV collaborative mapping according to claim 6, characterized in that, The controlled UAV performs secondary image acquisition along the supplementary acquisition trajectory for areas with high geometric complexity, generating a supplementary image dataset, including: The replenishment flight path is sent to the corresponding UAV, and the UAV is controlled to fly autonomously along the replenishment flight path; During the flight of the UAV, when it reaches the preset observation viewpoint on the supplementary data acquisition track, the onboard sensors are automatically triggered to acquire images of the designated area to be supplemented. Record the UAV positioning and attitude determination system data and system timestamp during the secondary image acquisition, and bind them to the acquired image data; After completing all flight missions for supplementary data acquisition, the image data and corresponding positioning and attitude information obtained by all UAVs during the secondary acquisition phase are summarized to generate the supplementary image dataset.
8. The historical building surveying method based on multi-UAV collaborative mapping according to claim 1, characterized in that, The process of fusing the collaborative image dataset and the supplementary image dataset, performing dense matching and surface reconstruction calculations, and generating a 3D reality model of the target historical building includes: The supplementary image dataset is merged with the collaborative image dataset, and global bundle adjustment is performed on all merged images to obtain the exterior orientation elements of all optimized images. Based on the optimized exterior orientation elements, multi-view stereo dense matching is performed on the surface of the target historical building. The corresponding three-dimensional spatial coordinates of each pixel are calculated to generate a high-density three-dimensional dense point cloud. Statistical filtering and radius filtering are performed on the three-dimensional dense point cloud to remove outlier noise points; The denoised 3D dense point cloud is processed into a triangular mesh to construct a triangular mesh surface model of the target historical building. The color information of the image is mapped onto each triangular face of the triangular mesh surface model to generate a 3D real-world model with realistic textures; The process of performing multi-view stereo dense matching on the surface of the target historical building, and calculating the corresponding three-dimensional spatial coordinates for each pixel, includes: On each image, a regularly distributed image-square matching grid is established; For each grid point on the reference image, under its epipolar geometric constraints, search for corresponding points on multiple adjacent images; A semi-global matching algorithm is used to calculate the optimal disparity value for each grid point on the reference image; Based on the optimal parallax value, the exterior orientation elements of the reference image, and the exterior orientation elements of the adjacent images, the forward intersection is used to calculate the object-space three-dimensional coordinates corresponding to the grid points. By traversing all grid points on all images, the three-dimensional spatial coordinates are calculated, and all coordinate points are aggregated to form the high-density three-dimensional dense point cloud.
9. The historical building surveying method based on multi-UAV collaborative mapping according to claim 8, characterized in that, The process of triangulating the denoised 3D dense point cloud into a triangular mesh to construct a triangular mesh surface model of the target historical building includes: The denoised 3D dense point cloud is spatially partitioned, and a 3D spatial index structure is constructed. Based on the theory of Poisson surface reconstruction, an indicator function is constructed starting from the three-dimensional dense point cloud and its normal vector; Extract the isosurface of the indicator function; the isosurface is the potential surface of the target historical building. The extracted isosurfaces are discretized into a mesh structure consisting of a series of vertices and triangular patches; The mesh structure is adaptively simplified and smoothed, and the mesh topology is optimized while preserving the features to generate the final triangular mesh surface model.
10. A historical building surveying system based on multi-UAV collaboration, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the historical building surveying method based on multi-UAV collaboration as described in any one of claims 1 to 9.