An urban building high-fidelity three-dimensional reconstruction method based on air-ground multi-source data fusion
By employing a multi-source air-ground data fusion method, utilizing FAST-LIVO2 and factor graph optimization models, PPFNet, PointNetLK, and 3D Gaussian sputtering technology, the problems of low point cloud accuracy and poor model integrity in multi-source air-ground data fusion were solved, achieving high-fidelity 3D reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA UNIV OF GEOSCIENCES (WUHAN)
- Filing Date
- 2026-04-09
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies suffer from low point cloud accuracy, poor model integrity, and insufficient real-time performance in multi-source air-ground data fusion. Especially in large and complex building environments, SLAM mapping algorithms are prone to point cloud drift and geometric distortion. Traditional registration relies on manual feature extraction, which is inefficient and has large errors, and has limited applicability to system integration.
The model is optimized by combining FAST-LIVO2 lidar odometry with factor map. Global nonlinear optimization is performed by IMU pre-integration residual, lidar point cloud ICP matching residual, visual reprojection residual and GNSS RTK positioning residual. Point cloud registration is performed by combining PPFNet and PointNetLK. High-fidelity reconstruction of the building model is achieved by using 3D Gaussian sputtering technology.
It improves point cloud accuracy and model integrity, achieves seamless coverage of all building elements, enhances visual realism and real-time interactivity, and meets the real-time requirements of high-precision 3D modeling.
Smart Images

Figure CN122454080A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional geographic information and point cloud processing technology, and more specifically, to a method for high-fidelity three-dimensional reconstruction of urban buildings by fusing multi-source data from air and ground. Background Technology
[0002] Ground point cloud reconstruction is a 3D data generation technology based on a ground-based mobile platform using SLAM (Simultaneous Localization and Mapping) technology. This technology generates discrete environmental point clouds through LiDAR scanning or multi-view image matching, and simultaneously uses sensor data such as inertial measurement units for pose estimation and trajectory optimization to stitch together a continuous point cloud map. It complements aerial point clouds generated by UAV aerial surveying, which mainly reflect building roofs and upper facades, in terms of perspective and data characteristics. However, existing technologies have the following drawbacks: 1. Limitations of ground-based data acquisition equipment: Strict control of movement speed and distance is required, otherwise it is easy to cause motion distortion of the lidar or insufficient point cloud density, resulting in low operational flexibility. Complex structures require manual marking of key frames, which relies on experience and is inefficient. 2. In large and complex building environments, SLAM mapping algorithms are prone to point cloud drift and geometric distortion due to accumulated sensor errors. At the same time, the generated point cloud density is uneven and noise is significant, affecting the registration accuracy and integrity with aerial point cloud data. 3. The back-end optimization of SLAM mapping algorithms is highly complex, with a large computational load, limited real-time performance, and imperfect dynamic interference handling. Dynamic point removal relies on motion consistency score calculation, which may lead to misjudgment of objects with sudden movement, resulting in incomplete point clouds.
[0003] Air-to-ground fusion refers to the technical process of high-precision matching, stitching, and joint analysis of 3D point clouds generated by ground mobile devices (such as robots and vehicles) using SLAM technology and point clouds generated by UAV aerial photography. Its core objective is to construct a seamless, unified 3D digital scene covering both air and ground perspectives, serving precise mapping, environmental perception, and intelligent decision-making. However, existing technologies have the following drawbacks: 1. Registration relies on manual feature extraction: Traditional methods require manual feature extraction during point cloud registration, which is not only inefficient but also prone to introducing subjective errors, affecting registration accuracy. 2. System integration and applicability limitations: Poor cross-platform compatibility; the conversion between the local coordinate system of the ground point cloud and the UAV's WGS84 coordinate system relies on a fixed process, making it difficult to adapt to other sensor combinations and limiting the technology's widespread application. 3. Low accuracy and efficiency of air-to-ground fusion: The overall fusion process has significant shortcomings in both accuracy and efficiency, failing to meet the real-time requirements of high-precision 3D modeling.
[0004] Overcoming the comprehensive limitations of existing methods in terms of geometric integrity, visual realism, and real-time interactivity is an urgent problem to be solved. Summary of the Invention
[0005] The purpose of this invention is to provide a high-fidelity 3D reconstruction method for urban buildings by fusing multi-source data from air and ground, which can improve point cloud accuracy, model integrity and reconstruction fidelity.
[0006] This invention provides a method for high-fidelity 3D reconstruction of urban buildings by fusing multi-source data from air and ground, characterized by comprising the following steps:
[0007] S1: Collect ground-based data and drone-based data respectively through ground-based equipment and drone-based equipment to obtain building-related point cloud and image data, and complete the spatiotemporal alignment of ground-based sensors and preprocessing of raw point clouds from drones. S2: Based on the point cloud and image data related to the building, the motion state and pose information of the sensor are estimated using the FAST-LIVO2 lidar odometry. A factor graph optimization model is added to the back end of the FAST-LIVO2 algorithm. The IMU pre-integration residual, the lidar point cloud ICP matching residual, the visual reprojection residual, the loop closure constraint residual, and the GNSS RTK positioning residual are embedded into the factor graph optimization model for global nonlinear optimization to obtain a high-precision ground point cloud map. S3: Based on the ground point cloud map, first use PPFNet to extract the global geometric features of the UAV point cloud and the ground point cloud and complete the coarse registration, then use PointNetLK to iteratively optimize the coarsely registered point cloud for fine registration, and obtain the fused unified point cloud. S4: Based on the fused unified point cloud, a 3D Gaussian sputtering model is constructed and the parameters are initialized. LiDAR depth information is introduced as a supervision signal. The depth features and local geometric features are fused to form a unified geometric representation. High-fidelity 3D reconstruction of urban buildings is achieved through joint supervision training with multiple loss functions, and a complete building model is output.
[0008] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for high-fidelity 3D reconstruction of urban buildings by fusing air and ground multi-source data.
[0009] The method for high-fidelity 3D reconstruction of urban buildings by fusing multi-source air and ground data provided in this invention has the following beneficial effects: This invention addresses the comprehensive limitations of existing methods in terms of geometric integrity, visual realism, and real-time interactivity. It generates a fine point cloud of the building facade by collecting data from a ground-based mobile device and accurately registers it with an aerial point cloud obtained by a drone. Finally, it achieves high-fidelity reconstruction of the building model from geometry to appearance using Gaussian sputtering technology based on multi-source data from both air and ground. Specifically, this invention constructs a factor graph optimization model, embedding various residual factors, lap-loop factors, and GNSS RTK global positioning factors into the model. It then uses nonlinear optimization methods for joint optimization, improving the accuracy and robustness of the SLAM algorithm, thereby enhancing point cloud accuracy. Furthermore, by fusing ground point clouds and UAV point clouds, this invention complements the perspective blind spots of single data sources at the data level, effectively solving the problems of missing building facades or incomplete roof models. This achieves seamless coverage of all building geometric elements, thus improving model integrity. Finally, this invention skips the traditional process of generating coarse meshes from point clouds and then performing texture mapping. Using 3D Gaussian sputtering technology, it tightly integrates the fused precise point clouds with multi-view images, directly optimizing a physically realistic appearance model from the data. This significantly improves detail reproduction and visual fidelity, thereby enhancing the model's reconstruction fidelity.
[0010] In summary, this invention utilizes a point cloud reconstruction method for building facades and its fusion technology with UAV point clouds to achieve blind-spot-free coverage of building geometry by coordinating the facade details of ground SLAM point clouds with the top surface information of UAV point clouds. Furthermore, it innovatively uses the fused unified point cloud as the initialization structure and constraint condition for Gaussian sputtering, enabling efficient and high-fidelity rendering from any perspective. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart of the urban building high-fidelity 3D reconstruction method based on multi-source air-ground data fusion provided by the present invention; Figure 2 This is a schematic diagram illustrating the air-to-ground multi-source point cloud registration and fusion effect provided by the present invention; Figure 3 This is a comparison image of the Gaussian sputtering initial point cloud provided by the present invention; Figure 4 This is a rendering of a 3D reconstruction of a building based on Gaussian sputtering using multi-source data from both air and ground, provided by the present invention. Detailed Implementation
[0012] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0013] Figure 1A schematic diagram of the high-fidelity 3D reconstruction method for urban buildings using multi-source air-ground data fusion according to this embodiment is shown. In this embodiment, the high-fidelity 3D reconstruction method for urban buildings using multi-source air-ground data fusion includes the following steps: S1: Collect ground-based data and drone-based data respectively through ground-based equipment and drone-based equipment to obtain building-related point cloud and image data, and complete the spatiotemporal alignment of ground-based sensors and preprocessing of raw point clouds from drones. In one exemplary embodiment, the specific process of ground-based data acquisition is as follows: a Livox Mid-360 lidar, a combined 180° wide field-of-view imaging system, and a GNSS RTK positioning system are used to scan around the building. In the time dimension, timestamps are unified based on a hardware synchronization triggering mechanism. In the spatial dimension, a joint calibration method between sensors is established to achieve external parameter conversion constraints. The moving speed during scanning is ≤0.5m / s, the distance between the device and the building facade is maintained at 1.5-3m, the overlap rate of adjacent scan zones is ≥30%, and a zigzag path is used to supplement the scan for complex structures and the building feature points are automatically marked.
[0014] In one exemplary embodiment, the specific process of data acquisition by the UAV is as follows: a circular route is planned with the building center as the center and a radius of 15-20m. The flight altitude covers the scanning boundary from the top of the building to the ground. The super high-rise building is scanned in layers according to height intervals and the layer overlap is ensured. The UAV, which integrates lidar and visible light sensors, synchronously collects spatial point clouds and optical images and records GNSS position information. The original point cloud is then preprocessed by noise removal, data filtering and sampling simplification.
[0015] S2: Based on the point cloud and image data related to the building, the motion state and pose information of the sensor are estimated using the FAST-LIVO2 lidar odometry. A factor graph optimization model is added to the back end of the FAST-LIVO2 algorithm. The IMU pre-integration residual, the lidar point cloud ICP matching residual, the visual reprojection residual, the loop closure constraint residual, and the GNSS RTK positioning residual are embedded into the factor graph optimization model for global nonlinear optimization to obtain a high-precision ground point cloud map.
[0016] In one exemplary embodiment, the state vector of the factor graph optimization model is defined as: , in, For location, Wield quaternion pose, For speed, and For IMU zero bias, For visual feature points; The back-end optimization objective function of the factor graph optimization model is: , Where X represents the state variable to be optimized; For IMU pre-integrated residuals; ICP matching residuals for laser point clouds; For visual reprojection residuals; For lap-loop constraint residuals; For GNSS RTK positioning residuals.
[0017] In one exemplary embodiment, in step S2, a scene feature vector is generated using a bag-of-words model, and the feature similarity between two frames is... Loopback verification is initiated at the specified time, and the loopback cost function is: , in, The cost function is the loop-closure cost function. This represents a robust kernel function; This represents the relative pose between the current frame and the loopback frame. The pose is transmitted through the keyframe graph; The laser point cloud ICP matching residual is: , in, This represents the laser point cloud ICP matching residual; This represents the laser point cloud pose transformation matrix from frame (i-1) to frame (i). and Let i represent the laser point cloud of frame i-1 and the laser point cloud of frame i. Point-to-surface ICP optimization is completed by combining normal vector constraints.
[0018] S3: Based on the ground point cloud map, first use PPFNet to extract the global geometric features of the UAV point cloud and the ground point cloud and complete the coarse registration. Then use PointNetLK to iteratively optimize the coarsely registered point cloud for fine registration to obtain the fused unified point cloud.
[0019] In one exemplary embodiment, the coarse registration process is as follows: S311. Perform voxel downsampling and normalization preprocessing on the UAV point cloud Pu and the ground point cloud Ps and estimate the normal vector. S312. Calculate the relative position and normal vector relationship between each point and its neighboring points using PPFNet, aggregate local geometric information and extract global features; S313. Calculate the feature vector matching degree based on cosine similarity, and filter reliable matching point pairs through RANSAC; S314. Calculate and center the centroids based on reliable matching point pairs, and solve for the rotation matrix R and translation vector t through singular value decomposition of the covariance matrix to construct the coarse registration transformation matrix: , Coarse registration completed.
[0020] In one exemplary embodiment, the fine registration process is as follows: S321. PointNet++ is used as the backbone network for feature extraction to extract a 1024-dimensional global feature descriptor for the two point clouds after coarse registration. S322. The inverse combination algorithm based on the Lucas-Kanade algorithm iteratively solves the optimal transformation parameters. The maximum number of iterations is set to 50, the initial learning rate is 0.01, and the learning rate is adaptively adjusted. The iteration stops when the feature residual change rate between the last two iterations is less than 1e-6 or the maximum number of iterations is reached. S323. Introduce a dual verification mechanism for feature consistency, which simultaneously verifies the geometric spatial distance and PointNet feature spatial distance of the transformed point pairs, filters valid correspondences, and completes fine registration.
[0021] S4: Based on the fused unified point cloud, a 3D Gaussian sputtering model is constructed and the parameters are initialized. LiDAR depth information is introduced as a supervision signal. The depth features and local geometric features are fused to form a unified geometric representation. High-fidelity 3D reconstruction of urban buildings is achieved through joint supervision training with multiple loss functions, and a complete building model is output.
[0022] In one exemplary embodiment, the accumulated LiDAR point cloud from multiple frames is projected onto an image to construct a dense depth map. The constructed depth information is mapped into high-dimensional feature information using a U-Net encoder. A cross-attention mechanism is then used to fuse the high-dimensional depth feature information with local geometric features of the LiDAR point cloud, including curvature and normal vectors. The fused feature information is used to guide 3D Gaussian sputtering reconstruction. The total loss function jointly supervised by multiple loss functions during training is: , , , in, This is the depth mean square error loss; For L1 loss; For difference structure similarity loss; The loss is the normal vector. The curve loss is represented by α, β, and γ; these are the loss weights. Indicates the number of LiDAR point clouds; This represents the predicted depth of the i-th point; This represents the LiDAR point cloud depth at the i-th point; Represents structural similarity loss; Represents the true value of an image; This indicates the rendering of the image.
[0023] In one exemplary embodiment, each Gaussian primitive in the 3D Gaussian sputtering model includes position coordinates, covariance matrix, RGB color features, and opacity; the distribution of Gaussian primitives is adaptively adjusted based on point cloud density, increasing primitive density in feature-rich regions and sparsening Gaussian primitives in planar regions.
[0024] In some embodiments, the above-described method for high-fidelity 3D reconstruction of urban buildings through multi-source air-ground data fusion can also be implemented in the following ways.
[0025] Figure 2 This is a schematic diagram illustrating the effect of air-ground multi-source point cloud registration and fusion provided by the present invention; in this embodiment, the method for high-fidelity 3D reconstruction of urban buildings by air-ground multi-source data fusion includes the following steps: I. Multi-source data acquisition from air and ground Ground-based data collection: (1) The ground data acquisition equipment uses Livox Mid-360 lidar to achieve 360° horizontal omnidirectional coverage, and constructs a combined 180° large field of view imaging system and a high-precision real-time dynamic differential positioning system (GNSS RTK) through the spatial pose collaborative layout of industrial cameras to scan and collect data along urban streets and building perimeters.
[0026] (2) Sensor spatiotemporal alignment: In the time dimension, a high-precision timestamp unification strategy based on hardware synchronization triggering mechanism is used to eliminate the asynchronicity and time drift of sensor data acquisition; in the spatial dimension, a joint calibration method between sensors is established to clarify the external parameter conversion constraints between multi-source data and realize the accurate alignment of multimodal acquisition data under a unified reference system.
[0027] (3) Building facade scanning operation: move horizontally along the building facade, and scan the height to cover the ground to the 2nd-3rd floor. The overlap rate of adjacent scan zones is ≥30%. For complex structures, use a zigzag path to supplement the scan to ensure the integrity of features. (4) Data acquisition control: The moving speed is ≤0.5m / s to avoid laser radar motion distortion and to keep the distance between the equipment and the facade 1.5-3m; automatically mark the building corners, doors and windows and other feature points as a reference for subsequent loop detection.
[0028] Data collection from drones: (1) Flight route planning: Circular flight: With the building center as the center, the radius is 15-20m, and the flight altitude covers the scanning boundary from the top of the building to the ground. For super high-rise buildings, layered scanning is carried out at certain height intervals, and necessary overlapping areas are ensured between layers.
[0029] (2) Data acquisition: The UAV platform integrates lidar and visible light sensor to simultaneously acquire spatial point cloud and optical image of the target area. All aerial data must record accurate GNSS location information.
[0030] (3) Data preprocessing: The quality of the acquired raw point cloud is optimized, including noise removal, data filtering and sampling simplification.
[0031] II. Ground Point Cloud Map Construction Front-end odometry: The FAST-LIVO2 lidar odometry is used to estimate the motion state and pose information of the sensor.
[0032] Backend optimization: A factor graph optimization model was added to the backend of the FAST-LIVO2 algorithm to optimize and correct the sensor pose and point cloud coordinates through nonlinear optimization.
[0033] (1) Definition of state vector: Tightly coupled laser-inertial-vision fusion:
[0034] in: Location), Quaternion posture) speed), IMU zero bias) (Visual feature points).
[0035] (2) Backend optimization objective function, minimizing multi-source residuals:
[0036] in, IMU pre-integral residuals are used to compensate for motion distortion based on median integrals. Laser point cloud ICP matching residual; Visual reprojection residual.
[0037] (3) Generate scene feature vectors using the bag-of-words (BoW) model. When the feature similarity between two frames is... Initiate loopback verification at the specified time; cyclic cost function
[0038] in, This represents the relative pose between the current frame and the loopback frame. The pose is transmitted through the keyframe graph.
[0039] (4) LIO inter-frame constraint optimization Laser point cloud ICP matching residual:
[0040] Point-to-surface ICP optimization with normal vector constraints (5) GNSS RTK factor fusion The system uses a GNSS RTK real-time dynamic differential positioning system to obtain the absolute world coordinates of the ground acquisition equipment and transforms them into the local coordinate system of the ground point cloud to provide global positioning constraints, correct point cloud coordinates and sensor trajectory drift.
[0041] The residual factor, inter-frame constraint factor, loop closure factor, and GNSS RTK constraint factor mentioned above are embedded into the factor graph optimization model for global optimization.
[0042] Backend optimization objective function:
[0043] in: - IMU pre-integral residuals - Laser point cloud ICP matching residual - Visual reprojection residuals - Cyclic constraint residuals - GNSS RTK positioning residual.
[0044] This method effectively solves the cumulative drift problem of the FAST-LIVO2 front-end odometer in large-scale building facade scenarios through multi-factor collaborative optimization, and provides a high-precision unified coordinate reference for air-to-ground-cloud fusion through GNSS RTK.
[0045] III. Point cloud registration based on multi-source air-ground data: (1) Coarse registration: Global feature matching based on PPFNet To overcome the challenges of large differences in point cloud density, noise, and initial position deviations, a deep learning-based global feature extraction and matching strategy is adopted. The Point Pair Feature Network (PPFNet) is used to extract global geometric features of the point cloud, and a robust correspondence between point clouds is learned through a deep learning network. The specific steps are as follows: The input UAV point cloud Pu and SLAM ground point cloud Ps are preprocessed for standardization: voxel downsampling is used to unify the point cloud density, reduce the amount of data while retaining the main geometric structure and improving the efficiency of subsequent calculations; then, normal vector estimation is performed to assign geometric information representing the local surface orientation to each point and construct high-level geometric features.
[0046] PPFNet feature extraction: For each point, calculate its relative position and normal vector relationship with neighboring points to aggregate local geometric information; Feature matching: Calculate the matching degree between feature vectors based on cosine similarity; Initial transformation estimation: Filter reliable matching point pairs through RANSAC; Output coarse registration transformation matrix Tcoarse.
[0047] Tcoarse is calculated based on reliable matching pairs filtered by RANSAC, resulting in a set of reliable matching pairs filtered by RANSAC. The final step is to calculate the optimal rigid body transformation matrix Tcoarse, which consists of a rotation matrix R and a translation vector t, minimizing the alignment error between the transformed UAV point cloud and the SLAM ground point cloud. This problem can be transformed into a least-squares optimization problem, namely, finding R and t that minimize the following equation, with the specific steps as follows:
[0048] Let the set of reliable matching pairs after RANSAC filtering be . ,in , There are N pairs in total.
[0049] Calculate the centroid of the point cloud: ,
[0050] Centralized point pair calculation: Centroid-free point pairs ,
[0051] Covariance Matrix and Singular Value Decomposition (SVD): Calculate and decompose the covariance matrix H.
[0052] in Let U and V denote the outer product, where U and V are orthogonal matrices and Σ is a diagonal matrix. If det(U)det(V^T)>0, then the rotation matrix R is: R=V·U^T If det(U)det(V^T)<0, the sign of the last column of V needs to be adjusted: R=V·diag(1,1,sign(det(U)))·U^T Translation vector: t=qR·p The final rigid body transformation of the coarse registration transformation matrix is:
[0053] (2) Fine registration: Iterative optimization based on PointNetLK By combining PointNet's global feature representation with the optimized framework of the Lucas-Kanade algorithm, high-precision registration can still be maintained in scenarios with repetitive textures on building facades. The computational efficiency is improved by more than 50% compared to traditional ICP. Specific steps are as follows: An enhanced PointNet++ is used as the backbone network for feature extraction to extract global feature descriptors for two point clouds. Its multi-scale grouping and hierarchical sampling mechanism can better capture geometric information from fine local structure to macroscopic overall shape, and finally output a 1024-dimensional global feature vector.
[0054] This paper employs an efficient optimization strategy within the LK algorithm framework, using an inverse combination algorithm to iteratively solve for the optimal transformation parameters. This algorithm linearizes the nonlinear optimization problem, providing an incremental update of the transformation parameters in each iteration. In this embodiment, the maximum number of iterations is set to 50, and an adaptively adjusted learning rate (initially 0.01) is used, automatically reducing the step size to improve accuracy as convergence approaches. The convergence condition is set to a feature residual change rate less than 1e-6 between the current and next iterations, or reaching the maximum number of iterations.
[0055] A dual verification mechanism for feature consistency is introduced, which examines both the geometric distance between transformed point pairs and their distance in the PointNet feature space. Only when the correspondence of both constraints is satisfied is the result considered valid, thereby further enhancing the robustness of registration in challenging scenarios.
[0056] (3) Registration effect illustration and parameter example Table 1 shows an example configuration in one embodiment of the present invention. Those skilled in the art can adjust the parameters according to specific scenarios, but this will not affect the implementation of the technical solution of the present invention.
[0057] Table 1: Examples of Main Parameter Settings in the Point Cloud Registration Stage
[0058] (4) Performance index comparison is shown in Table 2; Table 2: Comparison of Performance Indicators
[0059] IV. Gaussian Sputtering for Air-Ground Fusion: Gaussian sputtering is applied to the image for efficient and high-fidelity rendering, ultimately outputting a complete (city) architectural model. (1) Construction and parameter initialization of Gaussian sputtering model 3D Gaussian primitives are generated from the merged point cloud as the basic elements for rendering. Each Gaussian unit contains the following parameters: position coordinates Covariance matrix Color characteristics Opacity ; The distribution of Gaussian primitives is adaptively adjusted based on point cloud density. Primitive density is increased in feature-rich regions and appropriately thinned out in planar regions to improve computational efficiency. The initial point cloud effect is as follows. Figure 3 As shown.
[0060] (2) Deep supervision and multimodal feature fusion LiDAR depth information is introduced as an additional supervisory signal. This module not only directly supervises the depth but also fuses depth features with local geometric features to form a unified geometric representation.
[0061] The depth data used in this invention is the depth data of the center of the laser 3D point cloud relative to the camera. The depth from the center of the point cloud to the camera is essentially the coordinate value of the point cloud in the camera's Z-axis direction after transforming it from the LiDAR coordinate system to the camera coordinate system. To improve the density of the depth data, depth accumulation is performed using multiple frames of data, and erroneous depths are automatically filtered when projecting LiDAR points onto the image plane.
[0062] To ensure consistent representation of depth information and local geometric features during fusion, this invention designs a depth-supervised network that maps the original depth information into a high-dimensional feature representation. The encoder employs U-Net to perform feature extraction, noise reduction, and abstract representation of the depth data.
[0063] The encoded depth features are fused with the extracted local geometric features (curvature, normal vectors) to form a unified geometric representation. Specifically, a fusion module employing a cross-attention mechanism ensures that the information from each modality complements each other, preserving local shape details while providing global depth constraints. The fused features will then guide subsequent 3D reconstruction networks to improve the overall accuracy and robustness of the reconstruction.
[0064] During training, mean squared error (MSE) is used as the primary loss for deep supervision:
[0065] in Indicates the depth of network prediction. This represents the true depth obtained from LiDAR after preprocessing. The final total loss is:
[0066]
[0067]
[0068] in and These are L1 loss and Differentiable Structural Similarity Loss, respectively. , and The weights of each loss term. This design ensures that the network is adequately supervised at both the depth and local geometry levels, thereby achieving high-quality, detailed reconstruction.
[0069] (3) Rendering effect comparison and implementation parameter examples like Figure 4 The image shown is a 3D reconstruction rendering of a building based on Gaussian sputtering from multiple sources of air and ground data.
[0070] Table 3 provides an example of parameters for the Gaussian sputtering model in one embodiment of the present invention. Those skilled in the art can reproduce the three-dimensional Gaussian rendering process described in the present invention based on the above parameters.
[0071] Table 3: Parameter Examples of Gaussian Sputtering Model
[0072] This embodiment provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for high-fidelity 3D reconstruction of urban buildings using multi-source air-ground data fusion.
[0073] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for high-fidelity 3D reconstruction of urban buildings using multi-source data fusion from both air and ground sources, characterized in that... Includes the following steps: S1: Collect ground-based data and drone-based data respectively through ground-based equipment and drone-based equipment to obtain building-related point cloud and image data, and complete the spatiotemporal alignment of ground-based sensors and preprocessing of raw point clouds from drones. S2: Based on the point cloud and image data related to the building, the motion state and pose information of the sensor are estimated using the FAST-LIVO2 lidar odometry. A factor graph optimization model is added to the back end of the FAST-LIVO2 algorithm. The IMU pre-integration residual, the lidar point cloud ICP matching residual, the visual reprojection residual, the loop closure constraint residual, and the GNSS RTK positioning residual are embedded into the factor graph optimization model for global nonlinear optimization to obtain a high-precision ground point cloud map. S3: Based on the ground point cloud map, first use PPFNet to extract the global geometric features of the UAV point cloud and the ground point cloud and complete the coarse registration, then use PointNetLK to iteratively optimize the coarsely registered point cloud for fine registration, and obtain the fused unified point cloud. S4: Based on the fused unified point cloud, a 3D Gaussian sputtering model is constructed and the parameters are initialized. LiDAR depth information is introduced as a supervision signal. The depth features and local geometric features are fused to form a unified geometric representation. High-fidelity 3D reconstruction of urban buildings is achieved through joint supervision training with multiple loss functions, and a complete building model is output.
2. The method for high-fidelity 3D reconstruction of urban buildings by fusion of multi-source air and ground data according to claim 1, characterized in that, The specific process of ground-based data acquisition is as follows: a Livox Mid-360 lidar, a combined 180° wide field-of-view imaging system, and a GNSS RTK positioning system are used to scan around the building. In the time dimension, timestamps are unified based on a hardware synchronization triggering mechanism. In the spatial dimension, a joint calibration method between sensors is established to achieve external parameter conversion constraints. The moving speed during scanning is ≤0.5m / s, the distance between the equipment and the building facade is maintained at 1.5-3m, and the overlap rate of adjacent scan zones is ≥30%. For complex structures, a zigzag path is used to supplement the scan and automatically mark building feature points.
3. The method for high-fidelity 3D reconstruction of urban buildings by fusion of air and ground multi-source data according to claim 1, characterized in that, The specific process of data collection by the UAV is as follows: a circular route is planned with the building center as the center and a radius of 15-20m. The flight altitude covers the scanning boundary from the top of the building to the ground. The super high-rise building is scanned in layers according to the height interval and the layer overlap is ensured. The UAV, which integrates lidar and visible light sensors, synchronously collects spatial point clouds and optical images and records GNSS position information. The original point cloud is then preprocessed by noise removal, data filtering and sampling simplification.
4. The method for high-fidelity 3D reconstruction of urban buildings by fusing multi-source air and ground data according to claim 1, characterized in that, The state vector of the factor graph optimization model is defined as follows: , in, For location, Wield quaternion pose, For speed, and For IMU zero bias, For visual feature points; The back-end optimization objective function of the factor graph optimization model is: , Where X represents the state variable to be optimized; For IMU pre-integrated residuals; ICP matching residuals for laser point clouds; For visual reprojection residuals; For lap-loop constraint residuals; For GNSS RTK positioning residuals.
5. The method for high-fidelity 3D reconstruction of urban buildings by fusion of air and ground multi-source data according to claim 1, characterized in that, In step S2, scene feature vectors are generated using the bag-of-words model. When the feature similarity between two frames is... Loop closure verification is initiated at the specified time, and the loop closure cost function is: , in, The cost function is the loop-closure cost function. This represents a robust kernel function; The relative pose between the current frame and the loopback frame. The pose is transmitted through the keyframe graph; The laser point cloud ICP matching residual is: , in, This represents the residual of laser point cloud ICP matching; This represents the laser point cloud pose transformation matrix from frame (i-1) to frame (i). and Let i represent the laser point cloud of frame i-1 and the laser point cloud of frame i. Point-to-surface ICP optimization is completed by combining normal vector constraints.
6. The method for high-fidelity 3D reconstruction of urban buildings by fusion of air and ground multi-source data according to claim 1, characterized in that, The specific process of coarse registration is as follows: S311. Perform voxel downsampling and normalization preprocessing on the UAV point cloud Pu and the ground point cloud Ps and estimate the normal vector. S312. Calculate the relative position and normal vector relationship between each point and its neighboring points using PPFNet, aggregate local geometric information and extract global features; S313. Calculate the feature vector matching degree based on cosine similarity, and filter reliable matching point pairs through RANSAC; S314. Calculate and center the centroids based on reliable matching point pairs, and solve for the rotation matrix R and translation vector t through singular value decomposition of the covariance matrix to construct the coarse registration transformation matrix: , Coarse registration completed.
7. The method for high-fidelity 3D reconstruction of urban buildings by fusion of air and ground multi-source data according to claim 1, characterized in that, The specific process of fine registration is as follows: S321. PointNet++ is used as the backbone network for feature extraction to extract a 1024-dimensional global feature descriptor for the two point clouds after coarse registration. S322. The inverse combination algorithm based on the Lucas-Kanade algorithm iteratively solves the optimal transformation parameters. The maximum number of iterations is set to 50, the initial learning rate is 0.01, and the learning rate is adaptively adjusted. The iteration stops when the feature residual change rate between the last two iterations is less than 1e-6 or the maximum number of iterations is reached. S323. Introduce a dual verification mechanism for feature consistency, which simultaneously verifies the geometric spatial distance and PointNet feature spatial distance of the transformed point pairs, filters valid correspondences, and completes fine registration.
8. The method for high-fidelity 3D reconstruction of urban buildings by fusion of multi-source air and ground data according to claim 1, characterized in that, In step S4, the accumulated LiDAR point cloud from multiple frames is projected onto the image to construct a dense depth map. The constructed depth information is mapped into high-dimensional feature information using a U-Net encoder. Then, a cross-attention mechanism is used to fuse the high-dimensional depth feature information with the local geometric features of the LiDAR point cloud, including curvature and normal vectors. The fused feature information is used to guide 3D Gaussian sputtering reconstruction. The total loss function jointly supervised by multiple loss functions during training is: , , , in, This is the depth mean square error loss; For L1 loss; For difference structure similarity loss; The loss is the normal vector. For curvature loss; α, β, and γ are the loss weights; Indicates the number of LiDAR point clouds; This represents the predicted depth of the i-th point; This represents the LiDAR point cloud depth at the i-th point; Represents structural similarity loss; Represents the true value of an image; This indicates the rendering of the image.
9. The method for high-fidelity 3D reconstruction of urban buildings by fusion of air and ground multi-source data according to claim 1, characterized in that, Each Gaussian primitive in the 3D Gaussian sputtering model contains position coordinates, covariance matrix, RGB color features, and opacity. The distribution of Gaussian primitives is adaptively adjusted based on point cloud density, increasing primitive density in feature-rich regions and sparsening Gaussian primitives in planar regions.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the high-fidelity 3D reconstruction method for urban buildings using multi-source data fusion from air and ground as described in any one of claims 1-9.