A method based on 3D Gaussian for large-scale scene reconstruction

Through a 3D Gaussian-based method, the drone image data and perspective filtering strategy are used to solve the problem of model capacity, uneven division and memory usage in large-scale scene reconstruction, and efficient and fast urban-level large-scene modeling and real-time rendering are achieved.

CN119991939BActive Publication Date: 2025-07-18ZHONGKE XINGTU DIGITAL EARTH HEFEI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510012317.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-07-18
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing technology has problems such as limited model capacity, uneven scene division, too long training and rendering time, and too large memory usage in large scale scenario reconstruction, resulting in low reconstruction quality and efficiency.

Method used

Using a 3D Gaussian method, the three-dimensional reconstruction of urban-level large scenes is carried out through drone multi-view image data processing, point cloud data blocking, viewing angle filtering and parallel training, combined with 3D Gaussian Splatting technology.

Benefits of technology

It realizes efficient city-level large-scenario modeling, shortens training time, avoids memory overflow, improves reconstruction quality, and supports real-time rendering on front-end browsers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991939B_ABST
    Figure CN119991939B_ABST
Patent Text Reader

Abstract

The present invention discloses a method based on 3D Gaussian for large-scale scene reconstruction, which includes the following steps: S1, collecting multi-view image data of drones; S2, inputting the collected image data into colmap, and using colmap to generate the pose graph of the camera and the scene point cloud data; S3, removing noise points or outliers in the point cloud data obtained by colmap; S4, dividing the processed point cloud data into blocks; S5, using a view filtering strategy to filter redundant views during the scene reconstruction process; S6, transmitting the processed data information into 3D Gaussian Splatting for parallel training of scene blocks; S7, rendering the scene blocks on the browser and fusing them into the entire large scene. The present invention realizes an efficient three-dimensional reconstruction method for large scenes, and the modeling results can be rendered and fused in real time on the front-end page.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross - application of computer vision and computer graphics, and particularly relates to a method based on 3D Gaussian for large - scale scene reconstruction. Background Art

[0002] Large - scale scene modeling at the urban scale is an important research direction in the fields of computer vision and computer graphics. 3D modeling at the urban scale allows for the digital representation of the geographical information of a city and its surrounding environment. By creating accurate 3D models to simulate the appearance, structure, and terrain of a city, it supports fields such as urban planning, land management, traffic planning, architectural design, and visualization simulation. 3D modeling technology provides basic technical support for the development of digital twins and the metaverse, and promotes the construction of smart cities and the digital economy.

[0003] Neural Radiance Fields (NeRF) technology has become the mainstream method for creating high - quality 3D scene visual effects. NeRF is a new type of 3D reconstruction technology based on deep learning. It learns and represents the 3D geometry and radiance characteristics of a scene through a neural network model, thereby generating high - quality continuous views from sparse input viewpoints.

[0004] In recent years, 3D Gaussian distribution (3DGS) has become a promising alternative solution. Compared with NeRF, it shows remarkable performance in terms of visual quality and rendering speed, and can achieve realistic real - time rendering at a speed of more than 30 frames per second at 1080p resolution. Most academic explorations on 3DGS focus on improving the rendering quality of small scenes and memory optimization.

[0005] Both NeRF and 3D Gaussian Splattiing are trying many methods to extend them to large scenes. For example, Block-NeRF divides the urban scene into multiple blocks and assigns corresponding training images to each block. Mega-NeRF adopts a divide-and-conquer strategy, divides the scene with a grid, and trains each scene block with an MLP. Although the quality rendered by these methods has been improved compared to traditional methods, the problems of missing details and high training time still exist. Subsequently, 3DGS also proposed solutions for the training and modeling of large scenes. VastGaussian divides the scene into m×n grid cells according to the projection position of the camera on the ground plane, and each scene contains a similar number of training views. CityGuassian customizes a contraction function to make the Gaussian more concentrated and then performs grid division. Most of the above methods can be abstracted as grid division schemes. Although such a division strategy is simple and effective, since the overall scene is an irregular area, if only relying on the grid to divide, there will be a large difference in the size of the scene division area, and the data volume of some boundary blocks is too small, which is prone to overfitting due to too little data volume. That is, the irregular area and irregular point cloud distribution lead to unbalanced workloads of different blocks. This causes the rendering quality to decline, the training resources not to be fully utilized, and the training efficiency to be greatly reduced, thus affecting the overall modeling level.

[0006] At the same time, when 3DGS is applied to large-scale scene reconstruction, the original 3DGS has high memory requirements for storage and means higher memory overhead during training. The rich details of large-scale scenes require a large number of 3D Gaussian functions. Directly applying 3DGS to large-scale scenes will result in out-of-memory errors or too low reconstruction quality. For example, a 32GB GPU can optimize approximately 11 million 3D Gaussians, while a small garden scene with an area of less than 100 square meters requires approximately 5.8 million 3D Gaussians in Mip-NeRF 360 for high-fidelity reconstruction.

[0007] In summary, there are currently some problems in three-dimensional reconstruction of large scenes.

[0008] (1) The model capacity is limited and cannot train large-scale urban-level scenes.

[0009] (2) Using sparse grid division for the scene, the scene division is uneven, affecting the reconstruction quality.

[0010] (3) The training and rendering time is too long, and the memory occupancy of the scene is too large. Summary of the Invention

[0011] The present invention proposes a new scene division and perspective filtering strategy for an efficient method of three-dimensional reconstruction of large urban scenes, solving the above problems.

[0012] The present invention provides a method for large-scale scene reconstruction based on 3D Gaussian, and the specific solution is as follows:

[0013] A method for large-scale scene reconstruction based on 3D Gaussian includes the following steps:

[0014] S1. Collect multi-view image data of drones;

[0015] S2. Input the collected image data into colmap, and use colmap to generate the pose graph of the camera and the scene point cloud data;

[0016] S3. Remove the noise points or outliers in the point cloud data obtained by colmap;

[0017] S4. Divide the processed point cloud data into blocks;

[0018] S5. Use the perspective filtering strategy to filter the redundant perspectives in the scene reconstruction process;

[0019] S6. Input the processed data information into 3D Gaussian Splatting for parallel training of scene blocks;

[0020] S7. Render the scene blocks on the browser and fuse them into the entire large scene.

[0021] Preferably, the drones in step S1 fly at the same height.

[0022] Preferably, step S3 uses the Statistical Outlier Removal method to remove the noise points or outliers in the point cloud data obtained by colmap.

[0023] Preferably, step S4 uses the K-Means algorithm to divide the processed point cloud data into blocks.

[0024] Preferably, the specific steps of step S4 are as follows:

[0025] S41. Randomly select the initial block centers;

[0026] S42. For each point, calculate its distance d to each block center, and assign the point to the block with the closest distance;

[0027] S43. Update the block center points;

[0028] S44. Iteratively update the block centers.

[0029] Preferably, the specific method for filtering redundant perspectives in step S5 is: determine a threshold θ, where P is the total number of sparse points in the scene block, V is the total number of perspectives of all relevant cameras, and α is a coefficient between 0 and 1 used to adjust the tightness of the threshold to adapt to different scenarios; if the number of sparse points in the block corresponding to the perspective of a certain associated camera is greater than the threshold, then retain this perspective to assist in the initialization of the scene block, and if it is less than the threshold, then discard it.

[0030] The present invention also discloses a computer-readable storage medium with a computer program stored thereon. After the computer program runs, it executes the method described in any one of the above.

[0031] The present invention also discloses a computer system, including a processor and a storage medium. The storage medium has a computer program stored thereon. The processor reads and runs the computer program from the storage medium to execute the method described in any one of the above.

[0032] The beneficial effects of the present invention are as follows:

[0033] The present invention provides a method for real-time 3D modeling of urban-level large scenes using UAV images. By proposing a new scene partitioning strategy and perspective filtering strategy, a more efficient large scene modeling method is achieved. At the same time, the present invention uses 3DGaussianSplatting for parallel training of partitions, improving the training efficiency of the scene, shortening the training time, and avoiding memory overflow. In addition, the scene of the present invention has the ability to be rendered in real time on a front-end browser, realizing flexible and fast urban-level large scene modeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0037] For example Figure 1 , a method based on 3D Gaussian for large-scale scene reconstruction includes the following steps:

[0038] S1. Collect multi-view image data of the unmanned aerial vehicle (UAV). Specifically, select an area and use the UAV to perform a constant-speed cruise on the scene. The UAV takes pictures and collects image data of the scene at the same height according to the specified route.

[0039] S2. Input the collected image data into the colmap software, and use the colmap software to generate the pose graph of the camera and the initial point cloud data of the scene.

[0040] S3. Use the StatisticalOutlierRemoval method to remove the noise points or outliers in the point cloud data obtained by colmap.

[0041] Specifically, SOR is a point cloud denoising method based on statistics, which is used to remove noise or abnormal points from the point cloud data.

[0042] Threshold = μ + α·σ

[0043] Where Threshold: the threshold, which is the standard for judging whether a point is an outlier. If the average distance between a certain point and its neighbors exceeds this threshold, it is considered an outlier. μ represents the overall average density of the point cloud data. It measures the overall density of the point cloud. α is a user-defined parameter, representing a multiple of the standard deviation, which is used to control the sensitivity of identifying outliers. A larger α value will retain more points, while a smaller α value will more strictly remove outliers. σ is used to describe the degree of dispersion of the distances between points. The larger the standard deviation, the greater the local density difference of the point cloud.

[0044] S4. Use the K-Means algorithm to partition the processed point cloud data. Through the camera-to-world transformation, the coordinates of each camera are (x i , y i ).

[0045] Specifically, the specific steps of step S4 are as follows:

[0046] S41. Randomly select initial block centers: It is necessary to randomly select n initial block centers. Each center is represented as where j represents different scenario blocks. To avoid local optimal solutions caused by uneven initial selection of block centers, K-means++ is used. First, a point is randomly selected as the first block center, and then, based on the existing center points, points farther away from them are selected to gradually increase the block centers, trying to disperse the initial center points as much as possible.

[0047] S42. Assign points to the nearest block: For each point, calculate its distance d to each block center and assign the point to the nearest block.

[0048]

[0049] Update block centers: Once all points are assigned to a certain block, the center points of each block need to be updated next. The center point of each block is updated to the average value of all points in the block. The new center is:

[0050]

[0051] where C j represents the set of points included in the j-th block, and |C j | is the number of points in the block.

[0052] S43. Update block center points: Once all points are assigned to a certain block, the center points of each block need to be updated next. The center point of each block is updated to the average value of all points in the block. The new center is:

[0053]

[0054] where C j represents the set of points included in the j-th block, and |C j | is the number of points in the block.

[0055] S44. Iteratively update block centers: Until the positions of the block centers no longer change significantly. Usually, it is judged by whether the moving distance of the block centers is less than a certain threshold.

[0056]

[0057] where represents the old point center, and ∈ is the set convergence condition, usually a very small value.

[0058] Finally, through this process, the points in the point cloud will be gradually divided into n blocks, and the number of points in each block is as uniform as possible.

[0059] S5. Use a view filtering strategy to filter redundant views during the scene reconstruction process.

[0060] Specifically, before sending the segmented scene into the 3DGS for training, the point cloud and camera views within the scene block are insufficient to achieve high-quality reconstruction of this scene block. The point cloud within the block is associated with camera views outside the block. These associated camera views must be added to provide more initialization information for the training optimization of this block. However, since the visible spaces of the associated camera views for this block are different, not every camera view provides a positive effect. If the visible space of the view is too small, it will instead affect the reconstruction quality of the model. To address this issue, we propose a view filtering strategy to determine whether to add the view of this camera. Specifically, a threshold θ is determined.

[0061]

[0062] Where P is the total number of sparse points within the scene block, V is the total number of views of all relevant cameras, and α is a coefficient between 0 and 1 that can be used to adjust the tightness of the threshold to adapt to different scenes.

[0063] If the number of sparse points corresponding to the view of a certain associated camera in the block is greater than the threshold, then retain this view to assist in the initialization of this scene block. If it is less than the threshold, then discard it. Finally, after optimizing the scene block, the 3DGS data generated outside the scene block will be removed.

[0064] S6. Input the processed data information into 3DGaussianSplatting for parallel training of scene blocks. Through parallel training, improve the training efficiency of the scene, shorten the training time, and avoid the situation of memory overflow.

[0065] S7. Render the scene blocks on the browser and fuse them into the entire large scene. Use the ThreeJS framework to build a front-end display page to visualize the scene blocks after being trained by 3DGS, and multiple blocks can be loaded simultaneously, enabling the scene blocks to be stitched together into a complete large scene.

[0066] The present invention provides a method for real-time three-dimensional modeling of urban-level large scenes using drone images. By proposing a new scene segmentation strategy and view filtering strategy, a more efficient large scene modeling method is achieved. At the same time, the present invention uses 3D Gaussian Splatting for segmented parallel training, improving the training efficiency of the scene, shortening the training time, and avoiding the situation of memory overflow. In addition, the scene of the present invention has the ability to be rendered in real time on the front-end browser, realizing flexible and fast urban-level large scene modeling.

[0067] The present invention also discloses a computer-readable storage medium and a computer system. Specifically, a computer program is stored on a computer-readable storage medium. After the computer program runs, it executes the method described in any one of the above. A computer system includes a processor and a storage medium. A computer program is stored on the storage medium. The processor reads and runs the computer program from the storage medium to execute the method described in any one of the above.

[0068] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0069] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein can be implemented using a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.

[0070] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read from, and write to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.

[0071] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code. The computer-readable medium includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. The storage media may be any available media that can be accessed by a computer. By way of example and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a web site, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc where disks typically reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0072] The foregoing description of the disclosure has been provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0073] Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method based on 3D Gaussian for large-scale scene reconstruction, characterized in that, It includes the following steps: S1. Collect multi-view image data of the drone; S2. Input the collected image data into colmap, and use colmap to generate the pose graph of the camera and the scene point cloud data; S3. Remove the noise points or outliers in the point cloud data obtained by colmap; S4. Use the K-Means algorithm to block the processed point cloud data; through the camera-to-world transformation, the coordinates of each camera are (x i , y i ); The specific steps are as follows: S41, Randomly select the initial block centers: It is necessary to randomly select n initial block centers, and each center is represented as where j represents different scenario blocks; to avoid the local optimal solution caused by uneven initial selection of block centers, K-means++ is used to first randomly select a point as the first block center, and then select points farther away from the existing center points according to the existing center points, gradually increasing the block centers to disperse the initial center points; S42. Assign points to the nearest block: For each point, calculate the distance d from it to the center of each block, and assign the point to the nearest block; Update block centers: Once all points are assigned to a certain block, the center point of each block needs to be updated next; the center point of each block is updated to the average value of all points in the block, and the new center is as follows: where C j denotes that the j-th block contains a set of points, and |C j | is the number of points in that block; S43, Update the center point of the block: Once all points are assigned to a certain block, the center point of each block needs to be updated next; the center point of each block is updated to the average value of all points in this block, and the new center is as follows: where C j represents the set of points contained in the j-th block, and |C j | is the number of points in that block; S44. Iteratively update the block center: Until the position of the block center no longer changes significantly, usually judged by whether the moving distance of the block center is less than a certain threshold: Among them represents the old point center, and ∈ is the set convergence condition; Finally, through this process, the points in the point cloud will be gradually divided into n blocks, and the number of points in each block is as uniform as possible; S5. Use a perspective filtering strategy to filter redundant perspectives during the scene reconstruction process. The specific method is as follows: Determine a threshold θ, where P is the total number of sparse points in the scene block, V is the total number of perspectives of all relevant cameras, and α is a coefficient between 0 and 1 used to adjust the tightness of the threshold to adapt to different scenes. If the number of sparse points in the block corresponding to the perspective of a certain associated camera is greater than the threshold, then retain this perspective to assist in the initialization of the scene block; if it is less than the threshold, then discard it. S6. Input the processed data information into 3D Gaussian Splatting for parallel training of scene blocks; S7. Render the scene blocks on the browser and fuse them into the entire large scene.

2. The method according to claim 1, wherein: The drone in step S1 flies at the same altitude.

3. The method according to claim 1, wherein: In step S3, the StatisticalOutlier Removal method is used to remove the noise points or outliers in the point cloud data obtained by colmap.

4. The method according to claim 1, characterized in that: In step S4, the K-Means algorithm is used to partition the processed point cloud data.

5. A computer-readable storage medium, characterized in that: A computer program is stored on a medium. After the computer program runs, it executes the method described in any one of claims 1 to 4.

6. A computer system, characterized in that: It includes a processor and a storage medium. A computer program is stored on the storage medium. The processor reads and runs the computer program from the storage medium to execute the method described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-angle consistent plane detection and analysis method for monocular video scene three dimensional structure

    CN106570507A

  • Grid curved surface reconstruction system for scene understanding

    CN110009671A