Method for large-scale scene reconstruction based on 3D Gaussian

By adopting new scene division and perspective filtering strategies in large-scale scene reconstruction and combining 3D Gaussian technology, the problems of limited model capacity, uneven scene division, too long training and rendering time and too large memory in the existing technology are solved, and efficient and fast large-scene modeling and rendering are achieved.

CN119991939AActive Publication Date: 2025-05-13ZHONGKE XINGTU DIGITAL EARTH HEFEI CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510012317.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-13
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

In the large-scale scene reconstruction, the existing technology has problems such as limited model capacity, uneven scene division, too long training and rendering time, and too large memory usage.

Method used

A 3D Gaussian method is proposed to achieve efficient reconstruction of large-scale scenes through the collection and processing of multi-view image data of drones, combined with colmap and 3D Gaussian Splatting technology.

Benefits of technology

It improves the efficiency and quality of large-scene modeling, shortens training time, avoids memory overflow, and realizes the ability to render in real-time on front-end browsers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991939A_ABST
    Figure CN119991939A_ABST
Patent Text Reader

Abstract

The invention discloses a method for reconstructing a large-scale scene based on 3D Gaussian. The method comprises the following steps: S1, collecting multi-view image data of an unmanned aerial vehicle; s2, inputting the collected image data into a colmap, and generating a pose map of a camera and scene point cloud data by using the colmap; s3, noise points or outliers in the point cloud data obtained by the colmap are removed; s4, partitioning the processed point cloud data into blocks; s5, filtering redundant visual angles in the scene reconstruction process by using a visual angle filtering strategy; s6, the processed data information is transmitted into 3D Gaussian Splitting, and parallel training of scene blocks is carried out; and S7, rendering the scene blocks on the section of browser, and fusing the scene blocks into a whole large scene. The efficient large-scene three-dimensional reconstruction method is realized, and the modeling result can be rendered and fused in real time on a front-end page.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-application of computer vision and computer graphics, and in particular to a method for large-scale scene reconstruction based on 3D Gaussian. Background Art

[0002] Large-scale scene modeling at the city scale is an important research direction in the field of computer vision and computer graphics. 3D modeling at the city scale allows for digital representation of geographic information of the city and its surroundings. By creating accurate 3D models to simulate the appearance, structure, and terrain of the city, it supports fields such as urban planning, land management, transportation planning, architectural design, and visualization simulation. 3D modeling technology provides basic technical support for the development of digital twins and metaverses, and promotes the construction of smart cities and digital economy.

[0003] Neural Radiance Fields (NeRF) technology has become a mainstream method for creating high-quality 3D scene visual effects. NeRF is a new 3D reconstruction technology based on deep learning, which learns and represents the 3D geometry and radiation properties of the scene through a neural network model, thereby generating high-quality continuous perspectives from sparse input perspectives.

[0004] In recent years, 3D Gaussian distribution (3DGS) has emerged as a promising alternative solution. Compared with NeRF, it has shown significant performance in visual quality and rendering speed, and can achieve realistic real-time rendering at more than 30 frames per second at 1080p resolution. Most academic explorations on 3DGS focus on improving rendering quality and memory optimization for small scenes.

[0005] NeRF and 3D Gaussian Splattiing are trying many methods to extend them to large scenes. For example, Block-NeRF divides the urban scene into multiple blocks and assigns corresponding training images to each block. Mega-NeRF adopts a divide-and-conquer strategy, divides the scene with grids, and uses an MLP to train each scene block. Although the quality of rendering by these methods has improved compared with traditional methods, the problems of missing details and high training time still exist. Later, 3DGS also proposed solutions for the training and modeling of large scenes. VastGaussian divides the scene into m×n grid units according to the projection position of the camera on the ground plane, and makes each scene contain a similar number of training views. CityGuassian customizes a shrinkage function to make the Gaussian more clustered, and then divides it into grids. Most of the above methods can be abstracted as grid division schemes. Although such a division strategy is simple and effective, since the overall scene is an irregular area, if it is only divided by grids, the size of the scene division area will be too different, and the amount of data for some boundary blocks will be too small, which is prone to overfitting due to too little data. That is, irregular regions and irregular point cloud distribution lead to unbalanced workloads in different blocks, which results in reduced rendering quality and insufficient utilization of training resources, greatly reducing training efficiency and thus affecting the overall modeling level.

[0006] At the same time, when 3DGS is applied to large-scale scene reconstruction, the original 3DGS has high requirements for storage memory and means higher memory overhead during training. The rich details of large-scale scenes require a large number of 3D Gaussian functions. Directly applying 3DGS to large-scale scenes will lead to out-of-memory errors or low reconstruction quality. For example, a 32GB GPU can optimize about 11 million 3D Gaussians, while a small garden scene with an area of ​​less than 100 square meters requires about 5.8 million 3D Gaussians in Mip-NeRF 360 to achieve high-fidelity reconstruction.

[0007] In short, there are some problems in the current three-dimensional reconstruction of large scenes.

[0008] (1) The model capacity is limited and cannot train large-scale city-level scenarios.

[0009] (2) The scene is divided into uneven parts using sparse grids, which affects the reconstruction quality.

[0010] (3) The training rendering time is too long and the scene memory usage is too large. Summary of the invention

[0011] The present invention proposes a new scene division and perspective filtering strategy, which is used for a method of efficiently performing three-dimensional reconstruction on large city-level scenes, and solves the above-mentioned problems.

[0012] The present invention provides a method for large-scale scene reconstruction based on 3D Gaussian, and the specific scheme is as follows:

[0013] A method for large-scale scene reconstruction based on 3D Gaussian includes the following steps:

[0014] S1, collects multi-view image data from drones;

[0015] S2, input the collected image data into colmap, and use colmap to generate the camera pose graph and scene point cloud data;

[0016] S3, remove noise points or outliers in the point cloud data obtained by colmap;

[0017] S4, dividing the processed point cloud data into blocks;

[0018] S5, using the perspective filtering strategy to filter redundant perspectives in the scene reconstruction process;

[0019] S6, passing the processed data information into 3D Gaussian Splatting for parallel training of scene blocks;

[0020] S7 renders the scene blocks on that section of the browser and merges them into the entire large scene.

[0021] Preferably, the UAV in step S1 keeps flying at the same altitude.

[0022] Preferably, step S3 uses a Statistical Outlier Removal method to remove noise points or outlier points in the point cloud data obtained by colmap.

[0023] Preferably, step S4 uses a K-Means algorithm to divide the processed point cloud data into blocks.

[0024] Preferably, step S4 specifically comprises the following steps:

[0025] S41, randomly select the initial block center;

[0026] S42, for each point, calculate the distance d from it to the center of each block, and assign the point to the block with the closest distance;

[0027] S43, updating the block center point;

[0028] S44, iteratively update the block center.

[0029] Preferably, the specific method of filtering redundant perspectives in step S5 is: determining a threshold θ, Where P is the total number of sparse points in the scene block, V is the total number of viewing angles of all related cameras, and α is a coefficient between 0 and 1, which is used to adjust the tightness of the threshold to adapt to different scenes; if the number of sparse points in the block corresponding to the viewing angle of an associated camera is greater than the threshold, then this viewing angle is retained to assist in the initialization of the scene block, and if it is less than the threshold, it is discarded.

[0030] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. After the computer program is run, any of the above methods is executed.

[0031] The present invention also discloses a computer system, including a processor and a storage medium, wherein a computer program is stored on the storage medium, and the processor reads and runs the computer program from the storage medium to execute any of the methods described above.

[0032] The beneficial effects of the present invention are:

[0033] The present invention provides a method for real-scene three-dimensional modeling of large city-level scenes using drone images. By proposing a new scene segmentation strategy and a perspective filtering strategy, a more efficient large-scene modeling method is achieved. At the same time, the present invention uses 3D Gaussian Splatting for block parallel training, which improves the training efficiency of the scene, shortens the training time, and avoids memory overflow. In addition, the scene of the present invention has the ability to be rendered in real time on the front-end browser, realizing flexible and fast large-scene modeling of the city. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0035] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] like Figure 1 , a method for large-scale scene reconstruction based on 3D Gaussian, comprising the following steps:

[0038] S1, collecting multi-view image data from drones. Specifically, an area is selected, and a drone is used to cruise the scene at a constant speed, and the drone takes pictures and collects image data of the scene at the same height along a specified route.

[0039] S2, input the collected image data into the colmap software, and use the colmap software to generate the camera pose graph and the initial point cloud data of the scene.

[0040] S3, use the Statistical Outlier Removal method to remove noise points or outliers in the point cloud data obtained by colmap.

[0041] Specifically, SOR is a statistically based point cloud denoising method used to remove noise or outliers from point cloud data.

[0042] Threshold=μ+α·σ

[0043] Where Threshold: threshold, the standard for judging whether a point is an outlier. If the average distance between a point and its neighbors exceeds the threshold, it is considered an outlier. μ represents the overall average density of the point cloud data. It measures the overall density of the point cloud. α is a user-defined parameter that represents a multiple of the standard deviation and is used to control the sensitivity of identifying outliers. A larger α value will retain more points, while a smaller α value will remove outliers more strictly. σ is used to describe the degree of discreteness of the distance between points. The larger the standard deviation, the greater the difference in local density of the point cloud.

[0044] S4, use K-Means algorithm to divide the processed point cloud data into blocks. Through camera-to-world transformation, the coordinates of each camera are (x i ,y i ).

[0045] Specifically, step S4 includes the following steps:

[0046] S41, randomly select initial block center: n initial block centers need to be randomly selected. Each center is represented by Where j represents different scene blocks. In order to avoid the local optimal solution caused by uneven initial selection of block centers, K-means++ is used to randomly select a point as the first block center, and then select points farther away from the existing center points, gradually increase the block centers, and try to disperse the initial center points.

[0047] S42, point assignment to the nearest block: For each point, calculate its distance d to the center of each block, and assign the point to the block with the nearest distance.

[0048]

[0049] Update block center: Once all points are assigned to a block, the next step is to update the center point of each block. The center point of each block is updated to the average value of all points in the block. for:

[0050]

[0051] Among them C j represents the set of points contained in the jth block, |C j | is the number of points in the block.

[0052] S43, Update block center point: Once all points are assigned to a block, the center point of each block needs to be updated. The center point of each block is updated to the average value of all points in the block. for:

[0053]

[0054] Among them C j represents the set of points contained in the jth block, |C j | is the number of points in the block.

[0055] S44, iteratively update the block center: until the position of the block center no longer changes significantly. This is usually determined by whether the moving distance of the block center is less than a certain threshold.

[0056]

[0057] in Represents the old point center, ∈ is the set convergence condition, which is usually a very small value.

[0058] Finally, through this process, the points in the point cloud are gradually divided into n blocks, and the number of points in each block is as uniform as possible.

[0059] S5, uses the view filtering strategy to filter out redundant viewpoints in the scene reconstruction process.

[0060] Specifically, before sending the divided scene into 3DGS for training, the point cloud and camera perspective within the scene block are not sufficient to achieve high-quality reconstruction of this scene block. The point cloud within the block will be associated with the camera perspective outside the block. These associated camera perspectives must be added to provide more initialization information for the training optimization of this block. However, since the associated camera perspectives are different for the visible space of this block, not every camera provides a positive effect. If the visible space of the perspective is too small, it will affect the reconstruction quality of the model. To address this problem, we proposed a perspective filtering strategy to determine whether to add the perspective of this camera. Specifically, a threshold θ is determined.

[0061]

[0062] Where P is the total number of sparse points in the scene block, V is the total number of viewing angles of all relevant cameras, and α is a coefficient between 0 and 1 that can be used to adjust the tightness of the threshold to adapt to different scenes.

[0063] If the number of sparse points in the block corresponding to the view of a certain associated camera is greater than the threshold, this view is retained to assist in the initialization of the scene block. If it is less than the threshold, it is discarded. Finally, after the scene block is optimized, the 3DGS data generated outside the scene block will be removed.

[0064] S6, the processed data information is passed to 3D Gaussian Splatting for parallel training of scene blocks. Through parallel training, the training efficiency of the scene is improved, the training time is shortened, and memory overflow is avoided.

[0065] S7, renders the scene blocks on the browser and merges them into a large scene. The front-end display page is built using the ThreeJS framework to visualize the scene blocks trained by 3DGS, and multiple blocks can be loaded at the same time, so that the scene blocks can be spliced ​​into a complete large scene.

[0066] The present invention provides a method for real-scene three-dimensional modeling of large city-level scenes using drone images. By proposing a new scene segmentation strategy and a perspective filtering strategy, a more efficient large-scene modeling method is achieved. At the same time, the present invention uses 3D Gaussian Splatting for block parallel training, which improves the training efficiency of the scene, shortens the training time, and avoids memory overflow. In addition, the scene of the present invention has the ability to be rendered in real time on the front-end browser, realizing flexible and fast large-scene modeling of the city.

[0067] The present invention also discloses a computer-readable storage medium and a computer system, wherein a computer-readable storage medium stores a computer program, and after the computer program is run, any of the above methods is executed. A computer system includes a processor and a storage medium, wherein the storage medium stores a computer program, and the processor reads and runs the computer program from the storage medium to execute any of the above methods.

[0068] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. The technician may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as resulting in a departure from the scope of the present invention.

[0069] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in cooperation with a DSP core, or any other such configuration.

[0070] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read and write information from / to the storage medium. In an alternative, a storage medium may be integrated into a processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and the storage medium may reside in a user terminal as discrete components.

[0071] In one or more exemplary embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented as a computer program product in software, each function may be stored on or transmitted by a computer-readable medium as one or more instructions or codes. Computer-readable media include both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one place to another. Storage media may be any available medium that can be accessed by a computer. As an example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer. Any connection is also properly referred to as a computer-readable medium. For example, if the software is transmitted from a website, a server, or other remote source using a coaxial cable, a fiber optic cable, a twisted pair, a digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of the medium. Disk and disc as used herein include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, wherein disk often reproduces data magnetically, while disc reproduces data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0072] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein, but should be granted the widest scope consistent with the principles and novel features disclosed herein.

[0073] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for large-scale scene reconstruction based on 3D Gaussian, characterized in that: The following steps are involved: S1, collects multi-view image data from drones; S2, input the collected image data into colmap, and use colmap to generate the camera pose graph and scene point cloud data; S3, remove noise points or outliers in the point cloud data obtained by colmap; S4, dividing the processed point cloud data into blocks; S5, using the perspective filtering strategy to filter redundant perspectives in the scene reconstruction process; S6, passing the processed data information into 3D Gaussian Splatting for parallel training of scene blocks; S7 renders the scene blocks on that section of the browser and merges them into the entire large scene.

2. The method according to claim 1, characterized in that: The UAV in step S1 keeps flying at the same altitude.

3. The method according to claim 1, characterized in that: Step S3 uses the Statistical OutlierRemoval method to remove noise points or outliers in the point cloud data obtained by colmap.

4. The method according to claim 1, characterized in that: Step S4 uses the K-Means algorithm to divide the processed point cloud data into blocks.

5. The method according to claim 1, characterized in that The specific steps of step S4 are: S41, randomly select the initial block center; S42, for each point, calculate the distance d from it to the center of each block, and assign the point to the block with the closest distance; S43, updating the block center point; S44, iteratively update the block center.

6. The method according to claim 1, characterized in that The specific method of filtering redundant perspectives in step S5 is: determining a threshold θ, Where P is the total number of sparse points in the scene block, V is the total number of viewing angles of all related cameras, and α is a coefficient between 0 and 1, which is used to adjust the tightness of the threshold to adapt to different scenes; if the number of sparse points in the block corresponding to the viewing angle of an associated camera is greater than the threshold, then this viewing angle is retained to assist in the initialization of the scene block, and if it is less than the threshold, it is discarded.

7. A computer-readable storage medium, characterized in that: A computer program is stored on the medium, and after the computer program is run, the method according to any one of claims 1 to 6 is executed.

8. A computer system, characterized in that: The method comprises a processor and a storage medium, wherein a computer program is stored in the storage medium, and the processor reads and runs the computer program from the storage medium to execute the method as claimed in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-angle consistent plane detection and analysis method for monocular video scene three dimensional structure

    CN106570507A

  • Grid curved surface reconstruction system for scene understanding

    CN110009671A

  • Rapid radiation field reconstruction method under sparse view angle input

    CN115170741A

  • Large-scene high-fidelity live-action three-dimensional modeling method capable of realizing unbounded expansion

    CN118608706A