Urban-level real scene three-dimensional modeling method based on air-ground multi-source data
Through the urban-level real-life three-dimensional modeling method of multi-source data in open-ground data, the Transformer point cloud registration model and ground SLAM technology are used to solve the problems of insufficient ground accuracy and difficulty in data update in urban-level three-dimensional modeling, and high-precision and automated three-dimensional reconstruction and update are achieved.
Patent Information
- Application Number
- CN202511018404.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-07-23
AI Technical Summary
The existing technology has insufficient ground accuracy, problems with the bottom layer of facade occlusion and data update difficulties in urban-level three-dimensional modeling, which is difficult to meet the needs of refined applications.
The urban-level real-life three-dimensional modeling method based on air-ground multi-source data is adopted. By constructing ground and aerial point clouds, semantic labels, local geometric features and color features are extracted, and cross-point cloud feature interaction is achieved using the Transformer point cloud registration model, and precise registration and data updates are performed in combination with ground SLAM technology.
The geometric accuracy and detail richness of the urban real-life three-dimensional model on the ground and building facade are improved, and automated and efficient three-dimensional reconstruction and model updates are realized.
Smart Images

Figure SMS_83 
Figure SMS_94 
Figure SMS_95
Abstract
Description
Technical Field
[0001] The present invention relates to the field of photogrammetry and remote sensing technology, and in particular to a city-level real-scene three-dimensional modeling method based on multi-source data of space and ground. Background Art
[0002] Oblique photogrammetry uses drones equipped with multi-angle cameras to rapidly capture image data of urban features from both vertical and oblique perspectives. Combining photogrammetry principles with computer vision algorithms (such as structured light motion recovery and feature point matching) generates high-resolution 3D models of real-world scenes. This technology efficiently covers large areas, generating models with rich texture information and realistic visuals. It is widely used in urban 3D modeling, urban planning, digital twins, disaster assessment, and other fields. However, it still has the following limitations: 1. Insufficient ground accuracy: Due to the limited aerial shooting angle and obstruction by ground objects (such as trees and buildings), the model's geometric accuracy and detail performance in the ground area are poor, making it difficult to meet the needs of refined applications; 2. Deformation and occlusion of the bottom facade: The bottom facade of a building is often blocked by dynamic or static obstructions such as vehicles, pedestrians, and vegetation, resulting in data loss. This can cause the generated model to have blurred textures or distorted geometric structures, affecting the integrity and authenticity of the model. 3. Difficulty in data updating: Urban features (such as buildings and roads) change frequently. Updating traditional oblique photography models requires re-conducting large-scale aerial photography and data processing, which is a complex process with a long cycle and high cost.
[0003] Terrestrial laser scanning (TLS) uses a terrestrial laser scanner to emit laser beams to acquire high-precision 3D point cloud data of surface features. This technology generates high-resolution local 3D models through multi-station scanning and point cloud registration. Terrestrial laser scanning offers sub-centimeter accuracy and can capture geometric details of complex features. It is particularly well-suited for high-precision modeling of static scenes and is widely used in areas such as building facade measurement, cultural heritage preservation, and detailed urban modeling. However, it still has the following limitations: 1. Low operating efficiency: Terrestrial laser scanning requires point-by-point data collection at multiple fixed sites, which takes a long time to cover a large area. Especially in complex urban environments, equipment transportation and site setup consume a lot of manpower and time. 2. Complex data processing: The amount of point cloud data generated by laser scanning is huge, requiring complex pre-processing (such as denoising, registration, and segmentation) and post-processing (such as model reconstruction and texture mapping). This requires high computing resources and professional skills, and requires a lot of manual operation; 3. Limited scope of application: Terrestrial laser scanning is limited by the field of view of the equipment and terrain conditions. It is difficult to efficiently obtain complete data on the tops of high-rise buildings or large-scale objects. In addition, it is easily disturbed by moving objects in dynamic environments, resulting in incomplete data.
[0004] Chinese patent CN113643434A discloses a 3D modeling method, intelligent terminal, and storage device based on air-ground collaboration. The 3D modeling method based on air-ground collaboration includes: S101: acquiring aerial 3D laser point cloud data, ground 3D laser point cloud data, and oblique image data of a modeling target, fusing the aerial 3D laser point cloud data with the ground 3D laser point cloud data to form an air-ground laser fusion point cloud, and forming a dense matching point cloud based on the oblique image data; S102: adjusting the coordinates of objects in the air-ground laser fusion point cloud so that the coordinates of the same objects in the air-ground laser fusion point cloud and the dense matching point cloud are consistent; S103: constructing a 3D model of the modeling target based on the air-ground laser fusion point cloud and the oblique image data after the adjusted coordinates. Although the air-ground laser fusion point cloud has fewer occlusion areas than the oblique image, reducing data acquisition blind spots and making the details of the 3D model clearer, it still relies too much on fixed ground stations, resulting in limitations in usage scenarios or scope. Summary of the Invention
[0005] Therefore, the technical problem to be solved by the present invention is to overcome the defects in the existing technology and propose a city-level real-scene 3D modeling method based on multi-source data of air and ground, which performs local 3D reconstruction and fine-tuning of the building facades and ground parts of existing real-scene 3D city models (especially aerial oblique photography models).
[0006] In order to solve the above technical problems, the present invention provides a city-level real-scene 3D modeling method based on multi-source data of vacant land, comprising: A city-level real-scene 3D modeling method based on multi-source data of vacant land includes the following steps: 1) Construct ground point cloud and aerial point cloud, and extract semantic labels of ground image and aerial image respectively , local geometric features and color features And back-project to the corresponding ground point cloud and aerial point cloud, perform preliminary coordinate conversion and alignment, and unify the ground point cloud data to the aerial point cloud coordinate reference; 2) Calculate the fused multimodal features of ground point cloud and aerial point cloud separately , ,in, Semantic tags , local geometric features and color features The weight matrix, is the Sigmoid activation function, To cascade high-dimensional comprehensive features on the feature channel dimension, is the Hadamard product operation; 3) Aerial point cloud and ground point cloud Build separately and The local neighborhood of and , respectively extract the central features corresponding to the local neighborhood and Then, the feature interaction between point clouds is realized through the Transformer point cloud registration model. The first point in the aerial point cloud Points, The first point in the ground point cloud points, where the attention weight for, , in, and is a learnable linear transformation matrix that transforms the fused multimodal features of the aerial point cloud into and fusion multimodal features of ground point clouds Mapping to query and keyspace, is the feature dimension, scaling factor To stabilize the gradient, is the relative position encoding function, which converts the coordinate difference Mapped to a scalar bias term, the Softmax function is used to achieve probability distribution conversion; 4) For the selected 3D reconstruction area, the ground point cloud is transformed into the coordinate system of the aerial oblique photography model, and data replacement and texture updates are performed.
[0007] As a preferred embodiment, a hierarchical geometric encoder is constructed based on the PointNet++ model to extract local geometric features of the point cloud. ; Use lightweight convolutional network (MobileNetV3) to directly extract color features .
[0008] As a preferred embodiment, color features are extracted The illumination invariance constraint is introduced to enhance the robustness of color features.
[0009] As a preferred embodiment, based on the dense correspondence between the point clouds directly output by the above-mentioned Transformer point cloud registration model, the transformation matrix between the two sets of point clouds is solved. After obtaining the accurate transformation matrix, the ground point cloud is transformed into the coordinate system of the aerial oblique photography model.
[0010] As a preferred embodiment, for the 3D reconstruction area selected for air-ground fusion 3D reconstruction, ground point cloud data is preferentially used. If the aerial oblique photography model is in point cloud format, the registered ground point cloud is directly merged into the aerial point cloud. For overlapping areas, selective retention or fusion is performed based on point cloud density and confidence. If the aerial oblique photography model is a mesh, delete the low-precision mesh of the corresponding area in the original model, reconstruct the mesh of the area based on the ground point cloud, and smoothly stitch it with the retained model part.
[0011] As a preferred embodiment, its training loss function includes a ternary loss function that considers geometric alignment loss, semantic consistency loss and feature separability loss, wherein the geometric alignment loss uses an improved Chamfer distance to measure the registration error; the semantic consistency loss is used to constrain the semantic label consistency of matching point pairs; and the feature separability loss enhances the feature differentiation between overlapping and non-overlapping regions through contrastive learning.
[0012] As a preferred embodiment, during model training, the overlap of point cloud samples in all data sets is pre-calculated and divided into different training stages according to the overlap range. The model is first trained on data sets with low overlap, and then gradually transferred to data sets with high overlap for training.
[0013] The present invention has the following advantages: By accurately aligning the ground point cloud map constructed by ground mobile SLAM (Simultaneous Localization and Mapping) technology and the point cloud data of the drone oblique photography model, a Transformer point cloud registration model is designed. This model introduces semantic labels, color features, and local geometric features into the Transformer model based on the domain cross-attention mechanism, which can achieve accurate registration of heterogeneous point clouds with large differences in air-ground perspectives and low overlap; it can solve the problems of existing oblique photography urban 3D models in terms of missing details and insufficient accuracy in the ground and building facade areas, as well as the complex, long-cycle, and low-automation process of updating real-life 3D models. DETAILED DESCRIPTION
[0014] The technical invention of the present invention will be described clearly and completely below. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.
[0015] The present invention provides a city-level real-scene 3D modeling method based on multi-source data of vacant land, comprising the following steps: 1) Construct ground point clouds and aerial point clouds respectively, and then extract semantic labels of ground images and aerial images respectively , local geometric features and color features And back-project to the corresponding ground point cloud and aerial point cloud, perform preliminary coordinate conversion and alignment, and unify the ground point cloud data to the aerial point cloud coordinate reference; Specifically, the present invention designs a multimodal feature encoding module based on image semantics, color, and point cloud geometry to achieve effective representation of local features of point clouds. A pre-trained semantic segmentation network (DeepLabv3+) is used to extract pixel-level semantic labels from images. , and map the semantic labels to the point cloud space. If the point cloud here is a lidar point cloud collected synchronously with the camera, the three-dimensional point cloud can be projected onto the image plane through the following projection formula to complete the association between the point cloud and the semantic information. , Among them, the original coordinates of the laser point are , It is a homogeneous coordinate representation constructed for projection calculation, which facilitates the projection transformation from three-dimensional to two-dimensional through matrix multiplication. Laser point Homogeneous coordinates projected onto the 2D image plane, represents the horizontal projection coordinate, Indicates vertical projection coordinates, and the "1" at the end is used for homogeneous coordinate transformation; is the depth scaling factor of the laser point, which is used to normalize the projection result. The camera intrinsic parameter is K, and the external parameters of the laser radar and camera are , which needs to be determined in advance through calibration (e.g., checkerboard method). This invention extracts semantic object information from images through deep learning methods and then projects it onto a laser point cloud based on sensor extrinsic parameters. This allows for the use of alternative image semantic extraction models and algorithms, as well as for direct extraction of semantic information from laser point clouds.
[0016] At the same time, based on the existing PointNet++ model, a hierarchical geometric encoder is constructed to extract the local geometric features of the point cloud. The PointNet++ model implements a hierarchical geometric encoder. This paper retrains the original model design to meet the requirements for local geometric feature extraction in space-ground fusion 3D reconstruction applications. Specifically, the point cloud data is first downsampled to extract multiple center points, which are then divided into multiple local regions based on the center points. For each local region, the local geometric features within it are extracted. Then, the local features are aggregated over a larger area into a higher-dimensional feature vector, thereby increasing the model's receptive field and ultimately achieving a global geometric understanding of the entire point cloud scene.
[0017] The color features of the point cloud are directly extracted using a lightweight convolutional network (MobileNetV3). , and projected into the corresponding point cloud space using the same method as the semantic labeling. Considering that the colors of heterogeneous point clouds may be inconsistent, an illumination invariance constraint is introduced to enhance the robustness of color features. For example, by converting the RGB color values of the point cloud to the Lab color space, which effectively separates brightness information from chromaticity information, the illumination invariance constraint is constructed primarily based on the Euclidean distance of the chromaticity channel. In this way, the illumination inconsistency problem of heterogeneous point clouds can be suppressed to a certain extent, and more robust color features can be extracted.
[0018] 2) Calculate and fuse multimodal features In order to effectively integrate these heterogeneous features and adaptively adjust the weight contribution of different features, this paper adopts a feature fusion mechanism based on dynamic attention weights. Acting on semantic tags , local geometric features and color features , after linear combination, and then through the Sigmoid activation function , and get a dynamic attention weight vector between 0 and 1: ; Semantic tags , local geometric features and color features The weight matrix indicates the importance of each modal feature in the fusion, and reflects the contribution of each modal feature in specific tasks and scenarios. The Sigmoid activation function The function is to generate normalized attention weights, giving the network the ability to dynamically adjust the feature fusion ratio. Subsequently, the original feature vectors (semantic, geometric, color) are cascaded in the feature channel dimension to obtain a higher-dimensional comprehensive feature. ; Finally, the attention weight vector obtained by the above calculation is Hadamard-producted with the concatenated feature vector element by element ( ) operation to complete the adaptive adjustment of the contribution of each modal feature and obtain the fused multimodal feature , , F represents the fusion multimodal features, which is an overview. All fusion multimodal features in the present invention are calculated by this formula, such as is the i-th multimodal encoding feature of the aerial point cloud.
[0019] 3) Aerial point cloud and ground point cloud Build separately and The local neighborhood of and , respectively extract the neighborhood corresponding center features and ,Then, The first point in the aerial point cloud Points, The first point in the ground point cloud points, and realizes feature interaction between point clouds through the Transformer architecture, that is, mapping the multimodal features of aerial point clouds and ground point clouds to query and key spaces respectively, and calculating the attention weights, where the attention weights for,
[0020] in, and is a learnable linear transformation matrix that fuses the multimodal features of the aerial point cloud Fusion of multimodal features with ground point cloud Mapped to query and key spaces, is the feature dimension, scaling factor To stabilize the gradient, is the relative position encoding function, which converts the coordinate difference Mapped into a scalar bias term, the coordinate difference is converted into The mapping is a scalar bias term to enhance geometric consistency; the Softmax function is used to transform probability distributions. The standard Transformer architecture typically performs global attention calculations on all input features, resulting in computational complexity that grows quadratically with data size. In contrast, the neighborhood cross-attention mechanism constructs local neighborhoods of point clouds, limiting the scope of attention calculations and effectively reducing computational complexity. This makes it more suitable for processing large-scale point cloud data while preserving the spatial information of local structures.
[0021] 4) Transform the ground point cloud into the coordinate system of the aerial oblique photography model, and perform data replacement and texture updates.
[0022] Specifically, the 3D model fusion and update includes the following steps: 41) Data Alignment: The Transformer point cloud registration model described above can directly output dense correspondences between point clouds. The robust RANSAC algorithm can be used to solve the transformation matrix between the two point clouds. Once the precise transformation matrix is obtained, the ground point cloud is transformed into the coordinate system of the aerial oblique photography model, achieving precise alignment of air-ground data.
[0023] 42) Data Replacement: For the selected target update area, ground point cloud data with higher accuracy and richer detail is prioritized. If the aerial oblique photography model is in point cloud format, the registered ground point cloud is directly merged into the aerial point cloud. Overlapping areas can be selectively retained or fused based on point cloud density and confidence (derived from SLAM covariance). When constructing a ground point cloud map, the SLAM algorithm estimates the sensor's pose and uses a covariance matrix to represent the uncertainty of the pose estimate. The trace of the covariance matrix (the sum of all elements on the main diagonal) is used to represent the confidence score of the ground point cloud. Lower values indicate lower uncertainty and higher confidence. By setting a confidence threshold, only point clouds with high-confidence poses are retained, resulting in higher accuracy and reliability. After the ground point cloud confidence is filtered, the density of the airborne point clouds in the overlapping area is calculated separately. If the density is comparable (the difference in point cloud density is less than a certain threshold), the two point clouds are fused; otherwise, the denser point cloud is retained. The point cloud density is calculated using the k-nearest neighbor method.
[0024] If the aerial oblique photography model is a mesh, delete the low-precision mesh of the corresponding area in the original model, reconstruct the mesh of the area based on the ground point cloud, and smoothly stitch it with the retained aerial model part.
[0025] 43) Texture Update: Using high-resolution camera images and accurate poses collected from the ground, texture mapping is performed on the updated model area to generate high-definition facade and ground textures.
[0026] 44) Consistency processing: Check and process the topological errors, cracks or overlaps that may exist in the fused model to ensure the manifold and watertightness of the model.
[0027] In order to achieve accurate registration of ground SLAM point clouds and aerial oblique photography point clouds, the present invention first uses the GNSS geographic coordinate information of each aerial point cloud to perform preliminary coordinate conversion and alignment, and unify the aerial point clouds to the same geographic coordinate reference; then uses the designed Transformer point cloud registration model to achieve accurate matching of point clouds, and can realize automatic and high-precision 3D reconstruction of the ground and building facades.
[0028] This invention can realize air-ground fusion mapping based on UAV oblique photography data and ground-based LiDAR-camera-INS-GNSS RTK positioning data, making full use of the advantage of UAV oblique photography's wide coverage area and having the following characteristics: 1. High precision: Combining the high precision of close-range ground measurement with the global coverage of aerial measurement, the geometric accuracy and detail richness of the urban real-scene 3D model on the ground and building facades are significantly improved.
[0029] 2. High degree of automation: Utilizing simultaneous localization and mapping (SLAM) technology and deep learning-based point cloud registration technology, it reduces manual operations and improves the efficiency of 3D reconstruction and model updating.
[0030] 3. Strong robustness: The point cloud registration method based on the Transformer model combines image semantic information and point cloud geometric information, and can handle space-ground matching problems in complex urban scenes such as large perspective differences, low overlap, and lighting changes.
[0031] In the model training of the present invention, in order to suppress the noise interference of a large number of non-overlapping areas in the cross-view heterogeneous point cloud, the present invention uses a ternary loss function that considers geometric alignment loss, semantic consistency loss, and feature separability loss. Among them, the geometric alignment loss uses an improved Chamfer distance to measure the registration error; the semantic consistency loss is used to constrain the semantic label consistency of the matching point pairs; the feature separability loss enhances the feature differentiation of overlapping and non-overlapping areas through contrastive learning. Therefore, the total loss function for:
[0032]
[0033]
[0034]
[0035] in, They are geometric alignment loss, semantic consistency loss and feature separability loss, is the weight of each loss of geometric alignment loss, semantic consistency loss and feature separability loss, is the total loss function, For aerial point cloud, is the ground point cloud, The first point in the aerial point cloud Points, The first point in the ground point cloud points, T is the transformation matrix between point clouds, representing the aerial point cloud To ground point cloud The geometric transformation of is the inverse transform of T, representing the ground point cloud To the aerial point cloud The geometric transformation of Indicates that Aerial point cloud To ground point cloud The geometric transformation of Will Ground point cloud arrive The geometric transformation of is the candidate matching pair of semantic labels, s is the semantic label of the point cloud, is the semantic label of the i-th point in the aerial point cloud, is the semantic label of the jth point in the aerial point cloud, is the i-th multimodal encoding feature of the aerial point cloud, is the jth multimodal encoding feature of the ground point cloud, is the multimodal encoding feature of the k-th point in the ground point cloud, where i, j, and k are used to represent elements in a set, such as ∈ ,express is a point in the aerial point cloud, representing an index value. (i,j)∈ , indicating that i and j are candidate matching pairs of semantic labels γ is a hyperparameter in deep learning models, representing a threshold. For feature pairs that should match, the distance between their feature values should be within the threshold; for feature pairs that should not match, the distance between their feature values should be outside the threshold. This parameter is set based on empirical values and adjusted based on model performance.
[0036] At the same time, the present invention adopts a progressive learning strategy during training, that is, by controlling the degree of overlap of point clouds, it can achieve optimized transition learning from coarse to fine. This strategy divides the difficulty of data according to the overlap of heterogeneous point clouds in the air and ground. The higher the overlap, the lower the difficulty of the data set, and vice versa. The overlap is determined by the intersection over union (IoU). By precalculating the overlap of point cloud samples in all data sets, the overlap is that the air and ground point clouds in the training data set are all well-aligned, and the overlap can be calculated directly based on their three-dimensional coordinate information. And the training is divided into different stages according to the range of overlap. Specifically, at the beginning of training, data with a higher degree of overlap (such as IoU ≥ 0.7) is selected to ensure that the model can more easily learn stable and accurate feature matching relationships; then the overlap threshold is gradually lowered to increase the complexity and challenge of the training data.
[0037] The training is divided into K stages, and the overlap threshold of the kth stage is:
[0038] in, is the minimum point cloud overlap, is the maximum point cloud overlap. In each stage of training, point cloud samples that meet the following formula are dynamically selected from the data set as the training set , to achieve progressive training from coarse alignment to fine optimization, this strategy can accelerate the training process and improve the overall performance of the model.
[0039]
[0040] in, For aerial point cloud, is the ground point cloud, is the overlap of point cloud samples, The model first dynamically selects data samples within the overlap range of the current stage for training. By adjusting the loss function weights and learning rate, the model's ability to handle complex, low-overlap scenarios is gradually strengthened. As training progresses, the data complexity and difficulty are gradually increased, ultimately achieving a steady transition from simple to complex scenarios.
[0041] As one of the specific implementation methods, the construction of ground point cloud includes the following steps: 111) Data Collection: Use a handheld or vehicle-mounted mobile measurement system that integrates high-precision laser radar (LiDAR), inertial measurement unit (IMU), visible light camera, and real-time dynamic differential positioning system (GNSS RTK) to perform mobile scanning along urban streets, building perimeters, and other areas, and simultaneously collect radar point clouds, inertial navigation data, camera image sequences, and GNSS RTK positioning data.
[0042] 112) Sensor spatiotemporal calibration: Before or after data collection, the relative poses (i.e., sensor extrinsics) and timestamps between the lidar, IMU, and camera must be accurately calibrated and synchronized to ensure that the sensor extrinsics are accurate and the timestamps are synchronized.
[0043] 113) SLAM 3D Mapping: This framework uses a factor graph-optimized laser SLAM framework to construct accurate ground point cloud maps. The framework is divided into two parts: Front-end odometry: Use the FAST-LIO2 lidar odometry to estimate the sensor's motion state and pose information.
[0044] Back-end optimization: A factor graph model is constructed, including IMU pre-integration factors, GNSSR TK global positioning factors, lidar odometry factors, and loop closure detection factors. Nonlinear optimization is used to optimize the sensor pose and point cloud coordinates to eliminate the accumulated error introduced by the odometry over long distances. Back-end optimization using a factor graph model effectively addresses the inherent accumulated error of the front-end odometry, thereby building a globally consistent, high-precision map.
[0045] Point cloud coloring: After obtaining a high-precision, high-density ground point cloud map with global geographic coordinates, the LiDAR point cloud is colored using image and sensor external parameters to generate a ground color point cloud map. That is, the ground point clouds processed by this invention are all color point cloud maps.
[0046] The air-ground fusion 3D reconstruction method designed in the present invention can realize air-ground fusion mapping based on UAV oblique photography data and ground-based lidar-camera-inertial navigation-GNSS RTK positioning data, fully utilizing the advantage of the large coverage range of UAV oblique photography and realizing automated and high-precision 3D reconstruction of the ground and building facades.
[0047] The present invention adopts mobile lidar SLAM technology and fuses image data to generate a color point cloud, which can be replaced by other inventions such as ground-based lidar (TLS) scanning and motion structure recovery (SfM) technology.
[0048] The construction of aerial point cloud includes the following steps: 121) Data acquisition: Use drones equipped with lidar and visible light cameras to obtain aerial point cloud and image data of the target area. Ensure that the aerial data contains GNSS geographic coordinates.
[0049] 122) Data Preprocessing: Perform data preprocessing operations such as denoising, filtering, and downsampling on the original aerial point cloud. If a LiDAR point cloud is not available, a structure-from-motion (SfM) method is used to generate color point cloud data from drone aerial imagery.
[0050] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will readily appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. A city-level real-scene 3D modeling method based on multi-source data of vacant land, characterized by: The following steps are included: 1) Construct ground point cloud and aerial point cloud, and extract semantic labels of ground image and aerial image respectively , local geometric features and color features And back-project to the corresponding ground point cloud and aerial point cloud, perform preliminary coordinate conversion and alignment, and unify the ground point cloud data to the aerial point cloud coordinate reference; 2) Calculate the fused multimodal features of ground point cloud and aerial point cloud separately , ,in, Semantic tags , local geometric features and color features The weight matrix, is the Sigmoid activation function, To cascade high-dimensional comprehensive features on the feature channel dimension, is the Hadamard product operation; 3) Aerial point cloud and ground point cloud Build separately and The local neighborhood of and , The first point in the aerial point cloud Points, The first point in the ground point cloud points, respectively extract the central features corresponding to the local neighborhood and Then, the feature interaction between point clouds is realized through the Transformer point cloud registration model, where the attention weight for, , in, and is a learnable linear transformation matrix that transforms the fused multimodal features of the aerial point cloud into and fusion multimodal features of ground point clouds Mapping to query and keyspace, is the feature dimension, scaling factor To stabilize the gradient, is the relative position encoding function, which converts the coordinate difference Mapped to a scalar bias term, the Softmax function is used to achieve probability distribution conversion; 4) For the selected 3D reconstruction area, the ground point cloud is transformed into the coordinate system of the aerial oblique photography model, and data replacement and texture updates are performed.
2. The city-level real-scene 3D modeling method based on multi-source data of vacant land according to claim 1 is characterized in that: Based on the PointNet++ model, a hierarchical geometric encoder is constructed to extract local geometric features of point clouds. ; Use lightweight convolutional network to directly extract color features .
3. The city-level real-scene 3D modeling method based on multi-source data of vacant land according to claim 2 is characterized in that: Extracting color features The illumination invariance constraint is introduced to enhance the robustness of color features.
4. The city-level real-scene 3D modeling method based on multi-source data of vacant land according to claim 1 is characterized in that: Based on the dense correspondence between the point clouds directly output by the Transformer point cloud registration model mentioned above, the transformation matrix between the two sets of point clouds is solved. After obtaining the accurate transformation matrix, the ground point cloud is transformed into the coordinate system of the aerial oblique photography model.
5. The city-level real-scene 3D modeling method based on multi-source data of vacant land according to claim 1 is characterized in that: For the 3D reconstruction area selected for air-ground fusion 3D reconstruction, ground point cloud data is preferred. If the aerial oblique photography model is in point cloud format, the registered ground point cloud is directly merged into the aerial point cloud. For overlapping areas, selective retention or fusion is performed based on point cloud density and confidence. If the aerial oblique photography model is a mesh, delete the low-precision mesh of the corresponding area in the original model, reconstruct the mesh of the area based on the ground point cloud, and smoothly stitch it with the retained model part.
6. The city-level real-scene 3D modeling method based on multi-source data of vacant land according to claim 1 is characterized in that: Its training loss function includes a ternary loss function that considers geometric alignment loss, semantic consistency loss, and feature separability loss. Among them, the geometric alignment loss uses an improved Chamfer distance to measure the registration error; the semantic consistency loss is used to constrain the semantic label consistency of matching point pairs; and the feature separability loss enhances the feature differentiation between overlapping and non-overlapping regions through contrastive learning.
7. The city-level real-scene 3D modeling method based on multi-source data of vacant land according to claim 6 is characterized in that: During model training, the overlap of point cloud samples in all data sets is pre-calculated and divided into different training stages according to the overlap range. The model is first trained on data sets with low overlap, and then gradually transferred to data sets with high overlap for training.
Citation Information
Patent Citations
Manufacturing method of true digital ortho map (TDOM) based on light detection and ranging (LiDAR) point cloud and aerial image
CN103017739A
Fine three-dimensional modeling method and system based on air-ground multi-source data fusion
CN119273853A
Adaptive depth source channel joint coding method for three-dimensional point cloud wireless transmission
CN119729018A
Cited By
Urban building three-dimensional visual lightweight system based on data fusion
CN120765877A
Radar signal processing method and device based on artificial intelligence
CN120802204A
Bridge model establishment method based on multi-source point cloud data
CN120805272A
Air-ground integrated three-dimensional scene automatic modeling method, system and device and storage medium
CN121010718A