Urban building three-dimensional visual lightweight system based on data fusion
Through point cloud data comparison, vector analysis and neural network matching, the modeling authenticity and stability issues of urban building three-dimensional models are solved, and dynamic updating and efficient three-dimensional reconstruction are achieved, which is suitable for urban planning and disaster monitoring.
Patent Information
- Application Number
- CN202511277536.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-09
AI Technical Summary
In existing technologies, the authenticity and uniformity of urban building three-dimensional models rely on random filling, which affects the accuracy of texture mapping. The building models are difficult to update dynamically. In addition, there is a large difference between high-altitude and ground perspectives, and data cannot be directly fused. As a result, the stability and integrity of the three-dimensional reconstruction are insufficient.
Through point cloud data comparison and vector analysis, missing building point clouds are removed, new buildings are added, octree blocking is used to optimize nearest neighbor search, the drone altitude is lowered layer by layer, and neural networks and semantic matching are used to improve rendering realism, ensuring that the model is consistent at the geometric and semantic levels and achieving dynamic updates.
It reduces the computational complexity of point cloud processing and improves the quality of cross-view 3D reconstruction. It is suitable for dynamic scenarios such as urban planning and disaster monitoring, ensuring the authenticity and stability of the model.
Smart Images

Figure CN120765877A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of building three-dimensional visualization, and in particular to a three-dimensional visual lightweight system for urban buildings based on data fusion. Background Art
[0002] The urban building three-dimensional visual lightweight system based on data fusion is a system that uses an airborne lidar system to reconstruct the model of urban buildings.
[0003] Among the existing approximate solutions, for example, CN115937439B discloses a method, device, and electronic device for constructing a three-dimensional model of an urban building. This solution addresses the technical problem of high labor costs required to construct a three-dimensional model of an urban building. First, building attributes are generated based on satellite remote sensing images and AI interpretation results. Multiple vector outlines of the same building are grouped through vector intersection judgment and Hausdorff distance calculation. Based on the geographic location and height data of the vector outlines, a facade white mold is generated through stretching modeling. The texture scaling ratio and repetition pattern are automatically determined in combination with the building height and window width. The facade texture data is mapped to the white mold by a technical means. This solves the problem of inconsistent attributes of different parts of a building in the existing technology and achieves the technical effect of avoiding the high cost of manual adjustment of texture mapping. However, there are still technical problems in that the authenticity and uniformity of the modeling rely on random filling, which affects the accuracy of the texture mapping and makes it difficult to dynamically update the building model.
[0004] For example, CN118365819A discloses a method and apparatus for three-dimensional reconstruction of urban buildings. This solution addresses the technical problem of being unable to efficiently collect data in scenarios such as rainy days and at night due to limitations of the lidar imaging principle. First, millimeter-wave radar is used to replace the lidar. The center of the array radar system is used as the rotation center to obtain echo signals at various angles. The echo signals are imaged into main and secondary images through the BP algorithm (BackProjection). The three-dimensional coordinates are calculated based on the RD equation (RangeDopplerEquation). The three-dimensional data is output through neighborhood matching. This technical approach solves the problem of inconsistent attributes of different parts of a building in the existing technology and achieves the technical effect of avoiding the high cost of manual adjustment of texture mapping. However, there are still technical problems that the existing solution mainly uses a single main image, a single secondary image, or a small number of secondary images for matching, with limited information fusion. The stability and integrity of the three-dimensional reconstruction still have room for improvement. There is also a large difference in perspective between high altitude and ground, and data cannot be directly fused. Summary of the Invention
[0005] In view of the above situation, in order to overcome the defects of the existing technology, the present invention provides a three-dimensional visual lightweight system for urban buildings based on data fusion. In view of the problem of inconsistent attributes of different parts of buildings in the existing technology, the technical effect of avoiding the high cost of manual adjustment of texture mapping is achieved. However, there are still technical problems that the authenticity and uniformity of modeling rely on random filling, which affects the accuracy of texture mapping and the building model is difficult to dynamically update. This solution uses point cloud data comparison and vector analysis to eliminate disappeared building point clouds and add new buildings, ensuring that the model of non-changing parts remains unchanged while reducing the computational complexity of large-scale point cloud processing. Octree block optimization is used to optimize the nearest neighbor search to further reduce the computational complexity of point cloud matching. By utilizing the verification and model correction mechanism, it is more suitable for dynamic scenarios such as urban planning and disaster monitoring compared to traditional static city models. In view of the fact that existing technologies mainly use single main image, single secondary image or a small number of secondary image matching, information fusion is limited, the stability and integrity of 3D reconstruction still have room for improvement, and there is a large difference between high-altitude and ground perspectives, and data cannot be directly fused, this solution starts from the drone's perspective, gradually reduces the altitude, and renders the intermediate perspective virtual image through the initial rough model, gradually narrowing the feature gap with the ground perspective, and improving the rendering realism through pixel difference, structural similarity and semantic matching, ensuring that the model is consistent with the real scene at the geometric, structural and semantic levels, and improving the quality of cross-perspective 3D reconstruction.
[0006] The technical solution adopted by the present invention is as follows: The present invention provides a three-dimensional visual lightweight system for urban buildings based on data fusion, which includes a real image acquisition module, a bridging module and a building object update module;
[0007] The real image acquisition module acquires drone images and ground-view images from real photos taken by drones and on the ground, and the drone images and ground-view images form an image dataset;
[0008] The bridging module adopts a bridging method to bridge the feature gap between the drone and ground perspectives to construct a three-dimensional model of urban buildings;
[0009] The building object update module adopts a building object update method to identify changes in building form, update the three-dimensional model of urban buildings, and track the geometric and semantic changes of buildings.
[0010] Furthermore, the bridging method specifically includes the following steps:
[0011] Step A1: Initialization. Specifically, since the perspective differences between drone images are small, a neural network is used to first construct an initial rough building layout model based on the drone photos from a high-altitude perspective.
[0012] Step A2: Iteratively generate intermediate perspectives. Specifically, starting from the altitude of the drone, gradually decreasing the altitude, using the rough building layout model to render a virtual image at the current iteration altitude, adding the virtual image to the image dataset, aligning the image dataset into a globally consistent 3D spatial coordinate system, and obtaining the current camera pose.
[0013] Step A3: Virtual training, specifically, using a 3D Gaussian splattering technique to render a lower-level virtual image based on the image dataset and the camera pose, using a loss function to enhance the realism of the virtual image, generating a ground-level virtual image, using a semantic segmentation network to determine the semantics of buildings in the virtual image, and training the neural network by comparing the drone image, the ground-view image, and the virtual image, optimizing the rough building layout model, and ultimately constructing a 3D urban building model. The loss function formula used is as follows: ;
[0014] Where, represents the total loss function, The pixel difference score that represents the color difference of each pixel in the virtual image and the real photo, represents the structural similarity score between the virtual image and the real photo, Indicates the matching degree of semantic judgment results between virtual images and real photos.
[0015] Furthermore, the building object updating method specifically includes the following steps:
[0016] Step B1: predefinition, specifically, converting the image data set used in establishing the urban building three-dimensional model into point cloud data, recording the data as a first point cloud set, and recording the urban building three-dimensional model corresponding to the first point cloud set as a first building model;
[0017] Step B2: New data collection, specifically, using an airborne lidar system to emit laser pulses over the city to obtain the current 3D spatial point dataset, recorded as the second point cloud;
[0018] Step B3: Preprocessing, specifically, performing denoising and coordinate registration on the second point cloud set, and using a semantic segmentation algorithm to distinguish buildings in the second point cloud set;
[0019] Step B4: Vector intersection is used to determine the existence of buildings based on the spatial overlap of the first point cloud set and the second point cloud set. Specifically, the area where there is no building corresponding to the first point cloud set in the second point cloud set is marked as a missing building, and the area where there is no building corresponding to the second point cloud set in the first point cloud set is marked as a new building. The first point cloud set and the second point cloud set are divided into blocks using an octree to reduce the amount of nearest neighbor search calculations. The nearest neighbor search algorithm is used to obtain the vector distance of the building between the first point cloud set and the second point cloud set. The addition of a building and the modification of the roof structure will directly reflect the change in roof height and shape. The change of vertical distance in the Z-axis direction can reflect the change more intuitively. In comparison, the facade changes such as opening and closing windows, adding or removing decorative components have less impact on the overall geometric update of the building. Therefore, the projection value of the vector distance in the Z-axis direction is calculated. Since the roof point of the building is theoretically perpendicular to the ground, its projection value is higher, while the normal vector of the facade point is mostly horizontal, and the projection value is lower. Therefore, a first threshold is set, and only the point cloud corresponding to the projection value above the first threshold is retained. The facade points of the building are filtered, and connected component clustering is performed on the point cloud to obtain and retain point cloud clusters. The buildings corresponding to the point cloud clusters are marked as changed buildings.
[0020] Step B5: Verification phase, used to improve the accuracy of the model and construct the changing buildings in the urban building 3D model.
[0021] In step B5, the verification phase specifically includes the following steps:
[0022] Step B51: constructing a changed building in the three-dimensional urban building model, and verifying whether the new building and the disappeared building are constructed in the three-dimensional urban building model based on the point cloud corresponding to the new building in the second point cloud set;
[0023] Step B52: removing point clouds corresponding to the disappeared buildings and the newly added buildings from the first point cloud set, and verifying whether the non-newly added buildings and the disappeared buildings in the three-dimensional urban building model remain unchanged based on the first point cloud set;
[0024] Step B53: Modify the urban building 3D model according to the verification results.
[0025] The beneficial effects achieved by the present invention using the above scheme are as follows:
[0026] (1) In view of the problem of inconsistent attributes of different parts of buildings in existing technologies, the technical effect of avoiding the high cost of manual adjustment of texture mapping is achieved. However, there are still technical problems that the authenticity and uniformity of modeling rely on random filling, which affects the accuracy of texture mapping and the building model is difficult to dynamically update. This solution uses point cloud data comparison and vector analysis to eliminate missing building point clouds and add new buildings to ensure that the non-changing part of the model remains unchanged. At the same time, it reduces the computational cost of large-scale point cloud processing. It uses octree block optimization and nearest neighbor search optimization to further reduce the computational cost of point cloud matching. It uses verification and model correction mechanisms. Compared with traditional static city models, it is more suitable for dynamic scenarios such as urban planning and disaster monitoring.
[0027] (2) In view of the technical problems that the existing technology mainly adopts the matching of single main image, single auxiliary image or a small number of auxiliary images, the information fusion is limited, the stability and integrity of 3D reconstruction still have room for improvement, the high altitude and ground perspectives are very different, and the data cannot be directly fused, this solution starts from the drone perspective, gradually reduces the height, and renders the intermediate perspective virtual image through the initial rough model, gradually narrowing the feature gap with the ground perspective, and improving the rendering realism through pixel difference, structural similarity and semantic matching, ensuring that the model is consistent with the real scene at the geometric, structural and semantic levels, and improving the quality of cross-perspective 3D reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Fig. 1 A module connection diagram of a data fusion-based urban building three-dimensional visual lightweight system provided by the present invention;
[0029] Fig. 2 is a flowchart of the bridging method;
[0030] Fig. 3 A flowchart of the method for updating building objects.
[0031] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0033] Example 1: See Figs. 1-3This embodiment provides a three-dimensional visual lightweight system for urban buildings based on data fusion, which includes a real image acquisition module, a bridging module, and a building object update module;
[0034] The real image acquisition module acquires drone images and ground-view images from real photos taken by drones and on the ground, and the drone images and ground-view images form an image dataset;
[0035] The bridging module adopts a bridging method to bridge the feature gap between the drone and ground perspectives to construct a three-dimensional model of urban buildings;
[0036] The building object update module adopts a building object update method to identify changes in building form, update the three-dimensional model of urban buildings, and track the geometric and semantic changes of buildings.
[0037] Example 2: See Figs. 1-2 This embodiment is based on the above embodiment, and the bridging method specifically includes the following steps:
[0038] Step A1: Initialization. Specifically, since the perspective differences between drone images are small, a neural network is used to first construct an initial rough building layout model based on the drone photos from a high-altitude perspective.
[0039] Step A2: Iteratively generate intermediate perspectives. Specifically, starting from the altitude of the drone, gradually decreasing the altitude, using the rough building layout model to render a virtual image at the current iteration altitude, adding the virtual image to the image dataset, aligning the image dataset into a globally consistent 3D spatial coordinate system, and obtaining the current camera pose.
[0040] Step A3: Virtual training, specifically, using a 3D Gaussian splattering technique to render a lower-level virtual image based on the image dataset and the camera pose, using a loss function to enhance the realism of the virtual image, generating a ground-level virtual image, using a semantic segmentation network to determine the semantics of buildings in the virtual image, and training the neural network by comparing the drone image, the ground-view image, and the virtual image, optimizing the rough building layout model, and ultimately constructing a 3D urban building model. The loss function formula used is as follows: ;
[0041] Where, represents the total loss function, The pixel difference score that represents the color difference of each pixel in the virtual image and the real photo, represents the structural similarity score between the virtual image and the real photo, Indicates the matching degree of semantic judgment results between virtual images and real photos.
[0042] Example 3: See Figs. 1-3 This embodiment is based on the above embodiment, and the building object updating method specifically includes the following steps:
[0043] Step B1: predefinition, specifically, converting the image data set used in establishing the urban building three-dimensional model into point cloud data, recording the data as a first point cloud set, and recording the urban building three-dimensional model corresponding to the first point cloud set as a first building model;
[0044] Step B2: New data collection, specifically, using an airborne lidar system to emit laser pulses over the city to obtain the current 3D spatial point dataset, recorded as the second point cloud;
[0045] Step B3: Preprocessing, specifically, performing denoising and coordinate registration on the second point cloud set, and using a semantic segmentation algorithm to distinguish buildings in the second point cloud set;
[0046] Step B4: Vector intersection is used to determine the existence of buildings based on the spatial overlap of the first point cloud set and the second point cloud set. Specifically, the area where there is no building corresponding to the first point cloud set in the second point cloud set is marked as a missing building, and the area where there is no building corresponding to the second point cloud set in the first point cloud set is marked as a new building. The first point cloud set and the second point cloud set are divided into blocks using an octree to reduce the amount of nearest neighbor search calculations. The nearest neighbor search algorithm is used to obtain the vector distance of the building between the first point cloud set and the second point cloud set. The addition of a building and the modification of the roof structure will directly reflect the change in roof height and shape. The change of vertical distance in the Z-axis direction can reflect the change more intuitively. In comparison, the facade changes such as opening and closing windows, adding or removing decorative components have less impact on the overall geometric update of the building. Therefore, the projection value of the vector distance in the Z-axis direction is calculated. Since the roof point of the building is theoretically perpendicular to the ground, its projection value is higher, while the normal vector of the facade point is mostly horizontal, and the projection value is lower. Therefore, a first threshold is set, and only the point cloud corresponding to the projection value above the first threshold is retained. The facade points of the building are filtered, and connected component clustering is performed on the point cloud to obtain and retain point cloud clusters. The buildings corresponding to the point cloud clusters are marked as changed buildings.
[0047] Step B5: Verification phase, used to improve the accuracy of the model and construct the changing buildings in the urban building 3D model.
[0048] Example 4: See Figs. 1-3 This embodiment is based on the above embodiment. In step B5, the verification phase specifically includes the following steps:
[0049] Step B51: constructing a changed building in the three-dimensional urban building model, and verifying whether the new building and the disappeared building are constructed in the three-dimensional urban building model based on the point cloud corresponding to the new building in the second point cloud set;
[0050] Step B52: removing point clouds corresponding to the disappeared buildings and the newly added buildings from the first point cloud set, and verifying whether the non-newly added buildings and the disappeared buildings in the three-dimensional urban building model remain unchanged based on the first point cloud set;
[0051] Step B53: Modify the urban building 3D model according to the verification results.
[0052] Example 5: See Figs. 1-3 This embodiment is based on the above embodiment. The image dataset includes a training set and a test set. The training set contains building images collected from the ground perspective and the high-altitude perspective of a drone. The test set contains real images with altitudes decreasing by two hundred meters.
[0053] Example 6: See Figs. 1-3 This embodiment is based on the above embodiment. In step A2, the height is lowered layer by layer, specifically by two hundred meters per layer.
[0054] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0055] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
[0056] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.
Claims
1. A three-dimensional visual lightweight system for urban buildings based on data fusion, characterized by: It includes real image acquisition module, bridging module and building object update module; The real image acquisition module acquires drone images and ground-view images from real photos taken by drones and on the ground, and the drone images and ground-view images form an image dataset; The bridging module adopts a bridging method to bridge the feature gap between the drone and ground perspectives to construct a three-dimensional model of urban buildings; The building object update module adopts a building object update method to identify changes in building form, update the three-dimensional model of urban buildings, and track the geometric and semantic changes of buildings.
2. The urban building three-dimensional visual lightweight system based on data fusion according to claim 1 is characterized in that: The bridging method specifically includes the following steps: Step A1: Initialization. Specifically, since the perspective differences between drone images are small, a neural network is used to first construct an initial rough building layout model based on the drone photos from a high-altitude perspective. Step A2: Iteratively generate intermediate perspectives. Specifically, starting from the altitude of the drone, gradually decreasing the altitude, using the rough building layout model to render a virtual image at the current iteration altitude, adding the virtual image to the image dataset, aligning the image dataset into a globally consistent 3D spatial coordinate system, and obtaining the current camera pose. Step A3: Virtual training, specifically, using 3D Gaussian splattering technology to render a lower-altitude virtual image based on the image dataset and the camera pose, using a loss function to improve the realism of the virtual image, generating a virtual image at ground level, using a semantic segmentation network to determine the semantics of buildings in the virtual image, and training the neural network by comparing the drone image, the ground-view image, and the virtual image, optimizing the rough building layout model, and ultimately constructing a 3D model of urban buildings.
3. The data fusion-based 3D visualization lightweight system for urban buildings according to claim 2 is characterized in that: The building object updating method specifically includes the following steps: Step B1: predefinition, specifically, converting the image data set used in establishing the urban building three-dimensional model into point cloud data, recording the data as a first point cloud set, and recording the urban building three-dimensional model corresponding to the first point cloud set as a first building model; Step B2: New data collection, specifically, using an airborne lidar system to emit laser pulses over the city to obtain the current 3D spatial point dataset, recorded as the second point cloud; Step B3: Preprocessing, specifically, performing denoising and coordinate registration on the second point cloud set, and using a semantic segmentation algorithm to distinguish buildings in the second point cloud set; Step B4: Vector intersection, used to determine the existence of buildings based on the spatial overlap of the first point cloud set and the second point cloud set. Specifically, areas in the second point cloud set where there are no buildings corresponding to the first point cloud set are marked as disappeared buildings, and areas in the first point cloud set where there are no buildings corresponding to the second point cloud set are marked as newly added buildings. An octree is used to partition the first and second point cloud sets into blocks. A nearest neighbor search algorithm is used to obtain the vector distance of the buildings between the first and second point cloud sets. The projection value of the vector distance in the Z-axis direction is calculated. A first threshold is set, and only point clouds corresponding to projection values above the first threshold are retained. Connected component clustering is performed on the point clouds to obtain and retain point cloud clusters, and the buildings corresponding to the point cloud clusters are marked as changed buildings. Step B5: Verification phase, used to improve the accuracy of the model and construct the changing buildings in the urban building 3D model.
4. The data fusion-based 3D visualization lightweight system for urban buildings according to claim 3 is characterized in that: In step B5, the verification phase specifically includes the following steps: Step B51: constructing a changed building in the three-dimensional urban building model, and verifying whether the new building and the disappeared building are constructed in the three-dimensional urban building model based on the point cloud corresponding to the new building in the second point cloud set; Step B52: removing point clouds corresponding to the disappeared buildings and the newly added buildings from the first point cloud set, and verifying whether the non-newly added buildings and the disappeared buildings in the three-dimensional urban building model remain unchanged based on the first point cloud set; Step B53: Modify the urban building 3D model according to the verification results.
Citation Information
Patent Citations
Urban real scene three-dimensional modeling method based on multi-source geographic information coupling
CN116129067A
City building three-dimensional model monomer reconstruction method based on point cloud
CN116310192A
3D modeling method based on digital twin cities
CN118799488A
Air-ground image matching method and device, equipment and storage medium
CN120125742A
Urban building change detection method based on unmanned aerial vehicle video and three-dimensional model
CN120259916A
Cited By
Urban three-dimensional terrain rapid modeling method and system based on AI point cloud semantic segmentation
CN121788748A