High-quality large-scale scene reconstruction method based on 3D Gaussian sputtering

By dividing large scenes into multiple grid cells and using point cloud adaptive interpolation and global-local Gaussian decoder to dynamically adjust and merge 3D Gaussian distributions, the problem of block boundary inconsistency in large-scale three-dimensional scene reconstruction in the prior art is solved, and a higher quality and consistent three-dimensional model reconstruction is achieved.

CN119942014AActive Publication Date: 2025-05-06SHENYANG AEROSPACE UNIVERSITY

Patent Information

Application Number
CN202510105229.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing large-scale three-dimensional scene reconstruction method based on 3D Gaussian sputtering ignores the relationship between blocks when processing each block, resulting in inconsistencies at the block boundaries, affecting the reconstruction quality.

Method used

A high-quality large-scene reconstruction method based on 3D Gaussian sputtering is adopted. By dividing the large scene into multiple grid cells, the point cloud points in each grid cell are refined using the point cloud adaptive interpolation module, the 3D Gaussian parameters are calculated in combination with the global-local Gaussian decoder, and the Gaussian distribution is dynamically adjusted through the Gaussian refinement module. Finally, the 3D Gaussian distribution of the grid cells is merged by weighted average to generate a complete three-dimensional model.

Benefits of technology

It improves the consistency and quality of large-scene reconstruction, ensures consistent Gaussian parameter prediction between different grid cells, reduces redundant information, and improves the accuracy and integrity of the three-dimensional model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942014A_ABST
    Figure CN119942014A_ABST
Patent Text Reader

Abstract

The invention provides a high-quality large scene reconstruction method based on 3D Gaussian sputtering. The method comprises the following steps: dividing a large scene into a plurality of grid units; refining the point cloud points in the boundary of each grid unit by using a point cloud adaptive interpolation module; calculating 3D Gaussian parameters inside and outside the boundary of each grid unit by using a global-local Gaussian decoder; performing refining processing on the 3D Gaussian distribution through a Gaussian refining module; and combining the 3D Gaussian distributions of all the grid units into a complete large-scene 3D Gaussian distribution in a weighted average manner, and finally generating a complete large-scene three-dimensional model. According to the high-quality large-scale scene reconstruction method, the consistency and quality of large-scale scene reconstruction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 3D Gaussian reconstruction, and in particular to a high-quality large scene reconstruction method based on 3D Gaussian sputtering. Background Art

[0002] Large-scale 3D scene reconstruction is crucial for autonomous driving, virtual reality, environmental monitoring, and aerial surveying. 3D Gaussian sputtering (3DGS) has attracted attention for its high reconstruction quality and fast rendering speed. With the rapid development of 3D Gaussian, some methods have been developed to achieve large-scale 3D scene reconstruction based on 3D Gaussian sputtering. These methods usually adopt a divide-and-conquer strategy, dividing large scenes into multiple independent blocks, then processing each individual block, and finally merging the processed blocks. Although these methods effectively solve the problem of large-scale 3D scene reconstruction, the independent processing of each block will ignore the relationship between each block, resulting in inconsistency at the block boundary.

[0003] Therefore, it becomes an urgent problem to propose a high-quality large scene reconstruction method based on 3D Gaussian sputtering to improve the consistency and quality of large scene reconstruction. Summary of the invention

[0004] In view of this, the present invention provides a high-quality large scene reconstruction method based on 3D Gaussian sputtering to solve the problems existing in the prior art.

[0005] The present invention provides a high-quality large scene reconstruction method based on 3D Gaussian sputtering, comprising:

[0006] Collecting a large scene image from multiple perspectives and performing image preprocessing to obtain a preprocessed large scene image from multiple perspectives, wherein images from adjacent perspectives overlap;

[0007] According to the preprocessed multi-view large scene image, COLMAP is used to obtain the camera pose data and point cloud data corresponding to the multi-view large scene image;

[0008] Divide the scene ground plane into multiple grid cells, where each grid cell has a clear boundary and coordinate range;

[0009] Determine the contributing camera of each grid unit, wherein, for any grid unit, a group of cameras that contribute the most to it is selected as the contributing camera of the grid unit according to the visibility criterion;

[0010] Expand the point cloud points within the boundary of each grid unit outward to obtain the extended point cloud data outside the boundary of each grid unit;

[0011] Use the point cloud adaptive interpolation module to refine the point cloud points within the boundary of each grid unit to obtain a more refined point cloud representation;

[0012] Use a pre-trained point cloud segmentation model to segment the refined point cloud data within the boundary of each grid unit and the extended point cloud data outside the boundary of each grid unit to obtain segmented point cloud data;

[0013] According to the segmented point cloud data, the 3D Gaussian parameters inside and outside the boundary of each grid cell are calculated using the global-local Gaussian decoder to obtain the 3D Gaussian distribution inside and outside the boundary of each grid cell;

[0014] Dynamically adjust the 3D Gaussian distribution within each grid cell boundary using a Gaussian refinement module, wherein the Gaussian refinement module is used to dynamically allocate the number of Gaussians according to local geometric complexity, allocate more Gaussians to areas with complex geometric details, and delete redundant Gaussians;

[0015] Delete the 3D Gaussian distribution outside the boundary of each grid cell to obtain the optimized 3D Gaussian distribution of each grid cell;

[0016] The weights of all adjacent grid cells are determined according to the contribution camera of each grid cell, and the weights of all adjacent grid cells and the optimized 3D Gaussian distribution parameters are weighted averaged to obtain the 3D Gaussian distribution of the merged complete scene;

[0017] A 3D Gaussian rendering technique is used to generate a three-dimensional model according to the 3D Gaussian distribution of the complete scene.

[0018] Preferably, the image preprocessing includes denoising, normalization, and cropping and scaling.

[0019] Further preferably, the steps of determining the contributing camera of any grid unit A are as follows:

[0020] Projecting the boundary of the grid unit A onto the image plane of the camera, and calculating the projected area of ​​the grid unit A in each camera viewing angle;

[0021] Obtain the total number of pixels based on the camera resolution and calculate the total area of ​​the camera image;

[0022] Calculate the ratio of the projected area of ​​grid cell A to the total area of ​​the camera image in each camera view;

[0023] A plurality of cameras with the highest ratios are selected as contributing cameras of the grid unit A.

[0024] Further preferably, the specific steps of using the point cloud adaptive interpolation module to refine the point cloud points within the boundary of each grid unit are as follows:

[0025] Create two empty collections, one for storing the indexes of the points that have been processed, and the other for storing the points generated after interpolation;

[0026] Traverse each point in the initial point cloud and interpolate. The interpolation method for any point is as follows:

[0027] Find the K nearest neighbor points of the current point;

[0028] Perform 3D Voronoi partitioning using an incremental algorithm based on the K nearest neighbor points to obtain Voronoi polygons. The vertices of the Voronoi polygons are potential interpolation points.

[0029] Use Wasserstein distance to evaluate the topological difference between the set of K nearest neighbor points and the set containing the Voronoi polygon vertices. If the evaluation result shows that these vertices will not destroy the topological structure, the vertices of the Voronoi polygon are added to the set as interpolation points. Otherwise, perform 2D Voronoi interpolation.

[0030] The points generated after interpolation are merged with the original point cloud to obtain the refined point cloud.

[0031] Further preferably, the number of K nearest neighbor points of the current point is determined according to the sparsity of the point cloud, wherein the higher the curvature of the region, the more nearest neighbor points are selected.

[0032] Further preferably, the specific steps of 2D Voronoi interpolation are as follows:

[0033] Use principal component analysis to project the original K nearest neighbor points onto a two-dimensional plane;

[0034] Use the incremental algorithm to perform Voronoi division on the two-dimensional plane to obtain Voronoi polygons;

[0035] Map the vertices of the two-dimensional Voronoi polygon back to three-dimensional space and add the mapped vertices to the collection as interpolation points.

[0036] Further preferably, the specific steps of obtaining the 3D Gaussian distribution inside and outside the boundary of each grid unit are as follows:

[0037] Use the local Gaussian decoder to predict the 3D Gaussian parameters of each grid unit, and use the predicted 3D Gaussian parameter rendered image to compare with the real image, calculate the reconstruction loss, and then update the parameters of the local Gaussian decoder according to the reconstruction loss;

[0038] Update the parameters of the global Gaussian decoder according to the local Gaussian decoders of all grid cells after the updated parameters;

[0039] Use the global Gaussian decoder to predict the 3D Gaussian parameters of each grid unit, then calculate the difference between the Gaussian parameters predicted by the local Gaussian decoder and the global Gaussian decoder for each grid unit, and then update the parameters of the local Gaussian decoder according to the difference to make it closer to the prediction of the global Gaussian decoder;

[0040] The local Gaussian decoder with updated parameters is used to re-predict the 3D Gaussian parameters of each grid cell to obtain the 3D Gaussian distribution inside and outside the boundary of each grid cell.

[0041] Further preferably, the specific steps of using the Gaussian refinement module to dynamically adjust the 3D Gaussian distribution within each grid cell boundary are as follows:

[0042] Using a key point scorer, a key point score map of each input view is calculated from the image features, wherein the key point score map is used to reflect the geometric complexity of different regions in the image and can indicate the importance of each region in the image;

[0043] In the key point score map, for areas with higher scores, the Gaussian center is further subdivided into multiple smaller centers, and for areas with lower scores, the transparency and scaling of the corresponding Gaussian center are gradually reduced until it is completely deleted.

[0044] Further preferably, the weights of adjacent grid units are determined according to the number of cameras shared between the adjacent grid units. The more cameras shared between adjacent grid units, the higher the weights. If the contributing cameras of two adjacent grid units both include camera A, camera A is the shared camera of the two adjacent grid units.

[0045] The high-quality large scene reconstruction method based on 3D Gaussian sputtering provided by the present invention can improve the consistency and quality of large scene reconstruction. The method first divides the large scene into multiple grid units, and then uses a point cloud adaptive interpolation module to refine the point cloud points within the boundary of each grid unit, wherein the point cloud adaptive interpolation module can solve the problem of sparse point clouds in low curvature areas, while keeping the topological structure of the scene unchanged, and uses a global-local Gaussian decoder to calculate the 3D Gaussian parameters within and outside the boundary of each grid unit. The global-local Gaussian decoder can ensure that the Gaussian parameter predictions between different grid units remain consistent, thereby effectively solving the consistency problem between each grid unit, and then the 3D Gaussian distribution is refined by the Gaussian refinement module. By dynamically increasing or decreasing the Gaussian according to the complexity of the area, the quality of the 3D Gaussian can be improved and redundant information can be reduced. Finally, the 3D Gaussian distributions of all grid units are merged into a complete large scene 3D Gaussian distribution by weighted averaging, and finally a complete large scene three-dimensional model is generated. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0047] Figure 1 It is a flow chart of a high-quality large scene reconstruction method based on 3D Gaussian sputtering provided by the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in order to avoid blurring the present invention due to unnecessary details, only the processing steps closely related to the scheme of the present invention are shown in the accompanying drawings, and other details that are not closely related to the present invention are omitted.

[0049] like Figure 1 As shown, the present invention provides a high-quality large scene reconstruction method based on 3D Gaussian sputtering, comprising the following steps:

[0050] S1: collecting a large scene image from multiple perspectives and performing image preprocessing to obtain a preprocessed large scene image from multiple perspectives, wherein images from adjacent perspectives overlap;

[0051] Among them, multi-perspective large scene images can be collected by drone cameras. Specifically, the drone is controlled to fly around the center of the large scene at a constant speed, and the drone camera is controlled to capture scene images from multiple angles, and ensure that there is enough overlap between adjacent perspective images;

[0052] The image preprocessing includes denoising, normalization, and cropping and scaling;

[0053] The denoising process is as follows: first, select a suitable denoising algorithm (such as Gaussian filtering, median filtering, bilateral filtering) according to the noise type of the image; then set the parameters of the denoising algorithm (such as filter size, number of iterations) according to the noise level and resolution of the image; finally, use an image processing library (such as OpenCV or Pillow) to denoise the image;

[0054] The normalization process is as follows: the pixel values ​​of the denoised image are normalized to the range of [0,1] using the image processing library;

[0055] The steps of cropping and scaling are as follows: Use the image processing library to crop and scale the normalized image to a fixed size, such as 256×256 pixels;

[0056] S2: Using COLMAP to obtain the camera pose data and point cloud data corresponding to the multi-view large scene image based on the preprocessed multi-view large scene image;

[0057] S3: Divide the scene ground plane into a plurality of grid cells, for example, using m×n grids, where each grid cell has a clear boundary and coordinate range;

[0058] The specific steps of S3 are as follows:

[0059] S31: Select a suitable grid division scheme according to the size and complexity of the scene to ensure that the size of each grid unit is moderate, neither too large to affect the calculation nor too small to affect the efficiency;

[0060] S32: Divide the scene ground plane into a plurality of grid units according to the selected grid division scheme, each grid unit having a clear boundary and coordinate range;

[0061] S4: determining the contributing camera of each grid unit, wherein for any grid unit, a group of cameras that contribute the most to it is selected as the contributing cameras of the grid unit according to a visibility criterion;

[0062] The steps for determining the contributing camera of any grid cell A are as follows:

[0063] S41: Projecting the boundary of the grid unit A onto the image plane of the camera, and calculating the projection area of ​​the grid unit A in each camera viewing angle;

[0064] S42: Obtain the total number of pixels according to the camera resolution and calculate the total area of ​​the camera image;

[0065] S43: Calculate the ratio of the projection area of ​​the grid unit A to the total area of ​​the camera image in each camera viewing angle;

[0066] S44: selecting a plurality of cameras (the number of which can be adjusted according to the size of the scene) with the highest ratio as contributing cameras of the grid unit A;

[0067] S5: Expand the point cloud points within the boundary of each grid unit outward according to a certain ratio, for example, 20%, to obtain extended point cloud data outside the boundary of each grid unit;

[0068] S6: Use the point cloud adaptive interpolation module to refine the point cloud points within the boundary of each grid unit to obtain a more refined point cloud representation. The point cloud adaptive interpolation module can solve the problem of sparse point clouds in low curvature areas while keeping the topological structure of the scene unchanged.

[0069] The specific steps of S6 are as follows:

[0070] S61: Create two empty sets, one for storing the indexes of the points that have been processed, and the other for storing the points generated after interpolation;

[0071] S62: Traverse each point in the initial point cloud and perform interpolation, wherein the interpolation method of any point is as follows:

[0072] S621: Find the K nearest neighbor points of the current point. The selection of the K nearest neighbor points needs to consider the sparsity of the point cloud. Fewer nearest neighbor points can be selected in low curvature areas, while more nearest neighbor points are required in high curvature areas.

[0073] S622: Perform 3D Voronoi partitioning using an incremental algorithm based on K nearest neighbor points to obtain Voronoi polygons, where the vertices of the Voronoi polygons are potential interpolation points;

[0074] S623: Use Wasserstein distance to evaluate the topological structure difference between the set of K nearest neighbor points and the set containing the Voronoi polygon vertices. If the evaluation result shows that these vertices will not destroy the topological structure, the vertices of the Voronoi polygon are added to the set as interpolation points. Otherwise, perform 2D Voronoi interpolation.

[0075] Among them, the specific steps of 2D Voronoi interpolation are as follows:

[0076] S6231: Project the original K nearest neighbor points onto a two-dimensional plane using principal component analysis (PCA);

[0077] S6232: Perform Voronoi partitioning on a two-dimensional plane using an incremental algorithm to obtain Voronoi polygons;

[0078] S6233: Map the vertices of the two-dimensional Voronoi polygon back to three-dimensional space and add the mapped vertices as interpolation points to the set;

[0079] S63: merging the points generated after interpolation with the original point cloud to obtain a refined point cloud;

[0080] S7: Use the pre-trained point cloud segmentation model to segment the refined point cloud data within the boundary of each grid unit and the extended point cloud data outside the boundary of each grid unit to obtain segmented point cloud data;

[0081] S8: According to the segmented point cloud data, a global-local Gaussian decoder is used to calculate the 3D Gaussian parameters inside and outside the boundary of each grid cell, and the 3D Gaussian distribution inside and outside the boundary of each grid cell is obtained. The global-local Gaussian decoder can ensure that the Gaussian parameter predictions between different grid cells are consistent, thereby effectively solving the consistency problem between each grid cell.

[0082] The specific steps of S8 are as follows:

[0083] S81: Use the local Gaussian decoder to predict the 3D Gaussian parameters of each grid unit, and use the predicted 3D Gaussian parameter rendered image to compare with the real image, calculate the reconstruction loss, and then update the parameters of the local Gaussian decoder according to the reconstruction loss;

[0084] S82: updating the parameters of the global Gaussian decoder according to the local Gaussian decoders of all grid units after the parameters are updated;

[0085] S83: using the global Gaussian decoder to predict the 3D Gaussian parameters of each grid unit, then calculating the difference between the Gaussian parameters predicted by the local Gaussian decoder and the global Gaussian decoder for each grid unit, and then updating the parameters of the local Gaussian decoder according to the difference to make them closer to the prediction of the global Gaussian decoder;

[0086] S84: re-predicting the 3D Gaussian parameters of each grid unit using the local Gaussian decoder with updated parameters to obtain the 3D Gaussian distribution inside and outside the boundary of each grid unit;

[0087] S9: dynamically adjusting the 3D Gaussian distribution within the boundary of each grid unit using a Gaussian refinement module, wherein the Gaussian refinement module is used to dynamically allocate the number of Gaussians according to local geometric complexity, allocate more Gaussians to areas with complex geometric details, and delete redundant Gaussians;

[0088] The specific steps for S9 are as follows:

[0089] S91: using a key point scorer, calculating a key point score map of each input view from the image features, wherein the key point score map is used to reflect the geometric complexity of different regions in the image and can indicate the importance of each region in the image;

[0090] S92: In the key point score map, for a region with a higher score, further subdivide the Gaussian center into multiple smaller centers to more accurately capture local geometric details and subtle changes in the image, and for a region with a lower score, gradually reduce the transparency and scaling of the corresponding Gaussian center until it is completely deleted, wherein deleting redundant information from the Gaussian set can avoid Gaussian center overlap and improve the efficiency and accuracy of the model;

[0091] S10: deleting the 3D Gaussian distribution outside the boundary of each grid unit to obtain the optimized 3D Gaussian distribution of each grid unit to prevent boundary artifacts and ensure the quality of the merged scene;

[0092] S11: Determine the weights of all adjacent grid cells according to the contributing camera of each grid cell, and perform weighted averaging on the weights of all adjacent grid cells and the optimized 3D Gaussian distribution parameters to obtain the 3D Gaussian distribution of the merged complete scene;

[0093] The weights of adjacent grid units are determined according to the number of cameras shared between the adjacent grid units. The more cameras shared between the adjacent grid units, the higher the weights. If the contributing cameras of two adjacent grid units both include camera A, camera A is the shared camera of the two adjacent grid units.

[0094] S12: Generate a three-dimensional model according to the 3D Gaussian distribution of the complete scene using a 3D Gaussian rendering technique.

[0095] The high-quality large scene reconstruction method based on 3D Gaussian sputtering provided by the present invention can improve the consistency and quality of large scene reconstruction. The method first divides the large scene into multiple grid units, and then uses a point cloud adaptive interpolation module to refine the point cloud points within the boundary of each grid unit, wherein the point cloud adaptive interpolation module can solve the problem of sparse point clouds in low curvature areas, while keeping the topological structure of the scene unchanged, and uses a global-local Gaussian decoder to calculate the 3D Gaussian parameters within and outside the boundary of each grid unit. The global-local Gaussian decoder can ensure that the Gaussian parameter predictions between different grid units remain consistent, thereby effectively solving the consistency problem between each grid unit, and then the 3D Gaussian distribution is refined by the Gaussian refinement module. By dynamically increasing or decreasing the Gaussian according to the complexity of the area, the quality of the 3D Gaussian can be improved and redundant information can be reduced. Finally, the 3D Gaussian distributions of all grid units are merged into a complete large scene 3D Gaussian distribution by weighted averaging, and finally a complete large scene three-dimensional model is generated.

[0096] It should be noted that the purpose of publishing the embodiments is to help further understand the present invention, but those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the contents disclosed in the embodiments, and the scope of protection claimed by the present invention shall be subject to the scope defined in the claims.

Claims

1. A high-quality large scene reconstruction method based on 3D Gaussian sputtering, characterized in that: include: Collecting a large scene image from multiple perspectives and performing image preprocessing to obtain a preprocessed large scene image from multiple perspectives, wherein images from adjacent perspectives overlap; According to the preprocessed multi-view large scene image, COLMAP is used to obtain the camera pose data and point cloud data corresponding to the multi-view large scene image; Divide the scene ground plane into multiple grid cells, where each grid cell has a clear boundary and coordinate range; Determine the contributing camera of each grid unit, wherein, for any grid unit, a group of cameras that contribute the most to it is selected as the contributing camera of the grid unit according to the visibility criterion; Expand the point cloud points within the boundary of each grid unit outward to obtain the extended point cloud data outside the boundary of each grid unit; Use the point cloud adaptive interpolation module to refine the point cloud points within the boundary of each grid unit to obtain a more refined point cloud representation; Use a pre-trained point cloud segmentation model to segment the refined point cloud data within the boundary of each grid unit and the extended point cloud data outside the boundary of each grid unit to obtain segmented point cloud data; According to the segmented point cloud data, the 3D Gaussian parameters inside and outside the boundary of each grid cell are calculated using the global-local Gaussian decoder to obtain the 3D Gaussian distribution inside and outside the boundary of each grid cell; Dynamically adjust the 3D Gaussian distribution within each grid cell boundary using a Gaussian refinement module, wherein the Gaussian refinement module is used to dynamically allocate the number of Gaussians according to local geometric complexity, allocate more Gaussians to areas with complex geometric details, and delete redundant Gaussians; Delete the 3D Gaussian distribution outside the boundary of each grid cell to obtain the optimized 3D Gaussian distribution of each grid cell; The weights of all adjacent grid cells are determined according to the contribution camera of each grid cell, and the weights of all adjacent grid cells and the optimized 3D Gaussian distribution parameters are weighted averaged to obtain the 3D Gaussian distribution of the merged complete scene; A 3D Gaussian rendering technique is used to generate a three-dimensional model according to the 3D Gaussian distribution of the complete scene.

2. The high-quality large scene reconstruction method based on 3D Gaussian sputtering according to claim 1 is characterized in that: The image preprocessing includes denoising, normalization, and cropping and scaling.

3. The high-quality large scene reconstruction method based on 3D Gaussian sputtering according to claim 1 is characterized in that: The steps to determine the contributing camera for any grid cell A are as follows: Projecting the boundary of the grid unit A onto the image plane of the camera, and calculating the projected area of ​​the grid unit A in each camera viewing angle; Obtain the total number of pixels based on the camera resolution and calculate the total area of ​​the camera image; Calculate the ratio of the projected area of ​​grid cell A to the total area of ​​the camera image in each camera view; A plurality of cameras with the highest ratios are selected as contributing cameras of the grid unit A.

4. The high-quality large scene reconstruction method based on 3D Gaussian sputtering according to claim 1 is characterized in that: The specific steps for using the point cloud adaptive interpolation module to refine the point cloud points within the boundary of each grid cell are as follows: Create two empty collections, one for storing the indexes of the points that have been processed, and the other for storing the points generated after interpolation; Traverse each point in the initial point cloud and interpolate. The interpolation method for any point is as follows: Find the K nearest neighbor points of the current point; Perform 3D Voronoi partitioning using an incremental algorithm based on the K nearest neighbor points to obtain Voronoi polygons. The vertices of the Voronoi polygons are potential interpolation points. Use Wasserstein distance to evaluate the topological difference between the set of K nearest neighbor points and the set containing the Voronoi polygon vertices. If the evaluation result shows that these vertices will not destroy the topological structure, the vertices of the Voronoi polygon are added to the set as interpolation points. Otherwise, perform 2D Voronoi interpolation. The points generated after interpolation are merged with the original point cloud to obtain the refined point cloud.

5. The high-quality large scene reconstruction method based on 3D Gaussian sputtering according to claim 4 is characterized in that: The number of K nearest neighbor points of the current point is determined according to the sparsity of the point cloud. The higher the curvature of the region, the more nearest neighbor points are selected.

6. The high-quality large scene reconstruction method based on 3D Gaussian sputtering according to claim 4, characterized in that: The specific steps of 2DVoronoi interpolation are as follows: Use principal component analysis to project the original K nearest neighbor points onto a two-dimensional plane; Use the incremental algorithm to perform Voronoi division on the two-dimensional plane to obtain Voronoi polygons; Map the vertices of the two-dimensional Voronoi polygon back to three-dimensional space and add the mapped vertices to the collection as interpolation points.

7. The high-quality large scene reconstruction method based on 3D Gaussian sputtering according to claim 1 is characterized in that: The specific steps to obtain the 3D Gaussian distribution inside and outside the boundaries of each grid cell are as follows: Use the local Gaussian decoder to predict the 3D Gaussian parameters of each grid unit, and use the predicted 3D Gaussian parameter rendered image to compare with the real image, calculate the reconstruction loss, and then update the parameters of the local Gaussian decoder according to the reconstruction loss; Update the parameters of the global Gaussian decoder according to the local Gaussian decoders of all grid cells after the updated parameters; Use the global Gaussian decoder to predict the 3D Gaussian parameters of each grid unit, then calculate the difference between the Gaussian parameters predicted by the local Gaussian decoder and the global Gaussian decoder for each grid unit, and then update the parameters of the local Gaussian decoder according to the difference to make it closer to the prediction of the global Gaussian decoder; The local Gaussian decoder with updated parameters is used to re-predict the 3D Gaussian parameters of each grid cell to obtain the 3D Gaussian distribution inside and outside the boundary of each grid cell.

8. The high-quality large scene reconstruction method based on 3D Gaussian sputtering according to claim 1, characterized in that: The specific steps to use the Gaussian refinement module to dynamically adjust the 3D Gaussian distribution within each grid cell boundary are as follows: Using a key point scorer, a key point score map of each input view is calculated from the image features, wherein the key point score map is used to reflect the geometric complexity of different regions in the image and can indicate the importance of each region in the image; In the key point score map, for areas with higher scores, the Gaussian center is further subdivided into multiple smaller centers, and for areas with lower scores, the transparency and scaling of the corresponding Gaussian center are gradually reduced until it is completely deleted.

9. The high-quality large scene reconstruction method based on 3D Gaussian sputtering according to claim 1, characterized in that: The weights of adjacent grid cells are determined according to the number of cameras shared between the adjacent grid cells. The more cameras shared between adjacent grid cells, the higher the weights. If the contributing cameras of two adjacent grid cells both include camera A, camera A is the shared camera of the two adjacent grid cells.

Citation Information

Patent Citations

  • Three-dimensional scene representation method and device, storage medium and program product

    CN118799524A

  • Three-dimensional Gaussian sputtering optimization method for pose-free input

    CN119006678A

  • Hierarchical gaussian mixture model-based fast and robust robot three-dimensional reconstruction method

    WO2022095302A1

Cited By

  • Large-scene high-efficiency high-quality three-dimensional reconstruction method and device based on Gaussian sputtering

    CN121213820A