Sparse visual angle three-dimensional reconstruction method based on verification image guided Gaussian number control

Through the generative new perspective generation model and the Gaussian number control training strategy guided by the verification set, the overfitting problem in sparse perspective three-dimensional reconstruction is solved, and a sparse perspective three-dimensional reconstruction method with high-quality new perspective rendering and resource-saving resource saving is realized.

CN120259568AInactive Publication Date: 2025-07-04HANGZHOU DIANZI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510741084.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing sparse view 3D reconstruction method has overfitting problems when generating new view images, resulting in a decline in rendering quality and high computing resources and storage requirements, making it difficult to achieve high-quality three-dimensional reconstruction under sparse view conditions.

Method used

A new perspective generation model is used to generate new perspective images, and a verification set is formed through a distorted image filtering strategy. Combined with the Gaussian number control training strategy guided by the verification set, including model reduction and early stop strategy, the number of Gaussian balls is dynamically adjusted to avoid overfitting.

Benefits of technology

High-quality new perspective rendering with sparse perspective three-dimensional reconstruction is realized, reducing model storage requirements and training rendering time, and improving the applicability and wide use of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259568A_ABST
    Figure CN120259568A_ABST
Patent Text Reader

Abstract

The invention discloses a sparse view angle three-dimensional reconstruction method based on verification image guided Gaussian number control. A training image of a sparse view angle is used as input, a new view angle image is obtained by adopting a generative new view angle generation model, a distorted image filtering strategy is proposed, and a verification set of a subsequent training stage is obtained. Meanwhile, a training strategy for guiding Gaussian number control through a verification set is designed, the number of Gaussian balls is supervised by introducing the verification set, and a model reduction strategy and an early stop strategy are adopted, so that the number of the Gaussian balls is effectively controlled to be in an optimal interval. According to the method, the excellent new view angle rendering quality is realized, the model storage requirement is remarkably reduced, the training and rendering speed is improved, the method is easy to deploy on other sparse view angle three-dimensional reconstruction methods, and the applicability and wide usability of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, relates to three-dimensional reconstruction with sparse viewpoints, and specifically relates to a method for three-dimensional reconstruction with sparse viewpoints based on generative verification image-guided Gaussian number control. Background Art

[0002] In the fields of computer vision and three-dimensional reconstruction, three-dimensional reconstruction with sparse viewpoints is an extremely challenging technical problem. Compared with the dense-view two-dimensional images relied on by general three-dimensional reconstruction, the core goal of three-dimensional reconstruction with sparse viewpoints is to use two-dimensional images with limited viewpoints to quickly generate the geometric structure and appearance information of a three-dimensional scene, which plays an important role in fields such as virtual reality, augmented reality, intelligent driving, cultural heritage protection, and medical imaging. The advantage of this technology is that it can overcome the dependence on a large number of dense-view images in traditional methods, reduce the complexity and cost of data acquisition, and at the same time maintain high-quality reconstruction results.

[0003] Currently, three-dimensional reconstruction usually uses methods such as Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) for modeling and rendering. NeRF learns the structure of a 3D scene through a neural network and predicts the radiance of the scene, thereby achieving high-quality rendering of novel views. 3DGS is an emerging technology in the field of three-dimensional reconstruction. It replaces the implicit neural network of NeRF with an explicit representation based on 3D Gaussians, represents the scene with millions of explicit Gaussians, and performs fast and effective rasterization rendering through splatting, greatly improving the real-time rendering efficiency.

[0004] Although the existing technologies can achieve high-quality real-time rendering in dense-view three-dimensional scenes, they require a large number of densely captured images for effective scene reconstruction, which limits their practicality and popularity. To address this limitation, three-dimensional reconstruction with sparse viewpoints has been the focus of attention. Since the number of training images provided by sparse viewpoints is limited, some information in the scene cannot be fully captured, which can lead to the generation of misrepresentations. It can correctly represent the scene under the training viewpoints, but it will perform poorly in novel view rendering, showing problems such as geometric distortion, inaccurate hole filling, and issues like "floaters" (high-density regions at irregular positions in the scene) and "background collapse" (the background is represented by artifacts with similar appearances close to the camera).

[0005] Recent 3DGS methods have explored using sparse-view images as input for 3D reconstruction. Although these methods have achieved some improvements under sparse-view conditions, overfitting still exists. As the number of Gaussians increases, the maximum number of Gaussians in the reconstruction process improves the rendering quality of training images; however, for images of test views, the rendering quality initially increases but then decreases, showing obvious overfitting. Therefore, how to optimize 3D reconstruction for sparse views, reduce redundant and incorrect 3D representations, while meeting the requirements of reducing storage and computing resources and improving the effect of few-view 3D reconstruction has become an urgent technical problem to be solved.

[0006] In recent years, diffusion models have demonstrated excellent capabilities in generating high-quality images. They can generate images of new views based on text prompts or reference images, highlighting their potential to be extended to three-dimensional space. For example, ViewCrafter combines a video diffusion model with point cloud reconstruction technology to generate new views of a scene. 3DGS-Enhancer uses a video diffusion model and trajectory interpolation to enhance the unbounded 3DGS representation. Although these methods have enriched 3D content, they have not solved a fundamental problem: the inconsistency between the real world and the generated images, which often contain hallucinations and distorted details. Since 3D reconstruction aims to create an accurate digital replica of reality, it is crucial to avoid including such hallucinatory content. Therefore, how to make full use of such generative novel view synthesis (NVS) models to assist the 3D reconstruction process has great technical value. Summary of the Invention

[0007] Aiming at the deficiencies of the prior art, the present invention proposes a sparse-view 3D reconstruction method based on verified image-guided Gaussian number control. Specifically, the present invention takes sparse-view training images as input, uses a generative novel view synthesis (NVS) model, namely a generative NVS model, to obtain new-view images, and proposes a distorted image filtering strategy to obtain a verification set for the subsequent training stage. At the same time, we design a training strategy for verified set-guided Gaussian number control. By introducing the verification set for Gaussian sphere number supervision, adopting a model reduction strategy and an early stopping strategy, the Gaussian sphere number is effectively controlled within the optimal range.

[0008] In a first aspect, the present invention provides a sparse-view 3D reconstruction method based on verified image-guided Gaussian number control, specifically including the following steps:

[0009] Step 1: Divide the dense-view benchmark dataset for 3D reconstruction into a training set and a test set for sparse-view 3D reconstruction tasks.

[0010] Step 2: Use the generative NVS model to generate new perspective images. Through the designed distorted image filtering strategy, the final validation set views are obtained.

[0011] Taking the training set under sparse perspectives as input, use ViewCrafter in the generative NVS model to generate synthetic images of new perspectives as candidate images for the validation set.

[0012] A distorted image filtering strategy is designed. By calculating the per-pixel reprojection error between each training set image and the candidate images of the validation set, distorted images with low geometric consistency are excluded. The finally retained candidate images of the validation set constitute the validation set.

[0013] Step 3: Through Structure from Motion (SfM), obtain the camera poses of the validation set under unified camera intrinsics and a sparse point cloud with a better global geometric structure.

[0014] Resize and crop the images of the training set and the test set to be the same size as the images of the validation set. Design a joint initialization method. By combining the sparse input training set images and the generated validation set images, an initial point cloud is generated. The joint initialization requires 2 times of Structure from Motion (SfM).

[0015] Specifically, the input of the first SfM is the training set and test set images, generating the camera intrinsics and the camera poses corresponding to each image; the input of the second SfM is the training set, the validation set, and the camera poses obtained from the first SfM. When extracting features, set the camera intrinsics as the input camera intrinsics, thereby obtaining the camera poses of the validation set under unified camera intrinsics and the sparse point cloud generated by the training set plus the validation set.

[0016] Step 4: Create a Gaussian model for training and initialize the Gaussian sphere for the sparse point cloud obtained in Step 3.

[0017] Step 5: Design a training strategy for controlling the number of Gaussians guided by the validation set.

[0018] The training strategy includes introducing the validation set for Gaussian sphere number supervision, a model reduction strategy, and an early stopping strategy. By the feedback of the validation set during the iterative training process, the number of Gaussian spheres is dynamically adjusted to avoid overfitting.

[0019] Specifically, the training strategy divides the iterative training process of 3DGS into 2 stages. By setting variables: the number of rounds for validation set supervision to stop, the first stage and the second stage are divided; at the same time, set the validation set supervision interval. In the first stage, at the beginning of the iterative training, set the optimal index for validation set supervision When the Gaussian densification starts, if the current iteration number is exactly an integer multiple of the validation set supervision interval, perform validation set supervision and calculate the validation set supervision metrics. If then update and record the optimal Gaussian number as the Gaussian number at the current iteration. In the second stage, based on Gaussian densification, a model reduction strategy and an early stopping strategy are designed. The model reduction strategy reduces the Gaussian number to a non-overfitting state by randomly discarding Gaussian spheres.

[0020] The early stopping strategy is as follows: After randomly discarding Gaussian spheres, still perform the original Gaussian densification and keep supervising the Gaussian number after each densification operation. Once the Gaussian number is greater than the optimal Gaussian number after the current iteration, immediately perform the random discarding of Gaussian spheres in the model reduction strategy.

[0021] Loop and execute the model reduction strategy and the early stopping strategy until the end of the Gaussian densification stage, thereby controlling the number of Gaussian spheres within the controllable range of the optimal Gaussian number and achieving a non-overfitting state.

[0022] Step Six: Integrate the above steps into the existing sparse-view 3D reconstruction method to optimize the sparse-view 3D reconstruction effect and obtain better new-view rendering quality.

[0023] Integrate the training strategies for generating validation set images using the generative NVS model and controlling the Gaussian number guided by the validation set designed in the above steps into the architecture of the sparse-view 3DGS method (such as SparseGS, FSGS, DNGaussian, CoR-GS, etc.) based on the original 3DGS and capable of controlling the Gaussian number by changing the densification gradient threshold. During the operation, follow the principle of "lossless integration", that is, ensure that each inserted improvement measure does not interfere with or damage the main structure of the original 3DGS model framework.

[0024] On the second aspect, the present invention provides a sparse-view 3D reconstruction system based on validation image-guided Gaussian number control, including a dataset acquisition module, a validation set view generation module, a motion construction module, a Gaussian model creation module, a training strategy module, and a sparse-view 3D reconstruction module.

[0025] The described dataset acquisition module divides the dense-view benchmark dataset for 3D reconstruction into a training set and a test set for the sparse-view 3D reconstruction task.

[0026] The described validation set view generation module uses the generative NVS model to generate new-view images based on the training set views; through the designed distortion image filtering strategy, eliminate the distortion images with low geometric consistency to obtain the final validation set views.

[0027] The described motion construction module scales and crops the image sizes of the training set and the test set to be the same as those of the validation set. Then, two motion constructions (Structure from Motion, SfM) are performed: the first time, the input is the training set and the test set, and new camera intrinsics and camera poses are obtained; the second time, the input is the training set, the validation set, and the camera poses obtained from the first SfM, and the camera poses of the validation set under unified camera intrinsics and a sparse point cloud with a better global geometric structure are obtained.

[0028] The described Gaussian model creation module uses the jointly initialized sparse point cloud obtained by the motion construction module as the input in the Gaussian sphere initialization stage to create a Gaussian model for training.

[0029] The described training strategy module adopts a training strategy of validating set-guided Gaussian number control, which includes introducing the validation set for Gaussian sphere number supervision, a model reduction strategy, and an early stopping strategy, and dynamically adjusts the number of Gaussian spheres through the feedback of the validation set in the iterative training process, so as to avoid overfitting.

[0030] The described sparse-view three-dimensional reconstruction module adopts a sparse-view three-dimensional reconstruction method optimized by the dataset acquisition module, the validation set view generation module, the motion construction module, the Gaussian model creation module, and the training strategy module to achieve sparse-view three-dimensional reconstruction.

[0031] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the implementation manner of the first aspect above.

[0032] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the implementation manner of the first aspect above.

[0033] The present invention has the following beneficial effects:

[0034] In order to alleviate the overfitting phenomenon existing in sparse-view three-dimensional reconstruction, the present invention adopts a generative novel view synthesis model and a distortion image filtering strategy to obtain realistic novel view images, which are used as the validation set to obtain a more accurate sparse point cloud in the initialization stage. At the same time, a training strategy of validating set-guided Gaussian number control is designed, which effectively controls the number of Gaussian spheres and alleviates the overfitting problem of sparse-view three-dimensional reconstruction. The present invention not only achieves excellent novel view rendering quality, but also significantly reduces the model storage requirements, improves the training and rendering speed, is easy to be deployed on other sparse-view three-dimensional reconstruction methods, and improves the applicability and wide usability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Flow chart for generating validation set images by the generative NVS model and the validation set-guided Gaussian quantity control strategy.

[0036] Figure 2 Comparison chart of synthetic images output by the generative NVS model during the generation process of the validation set view for Scenario 1.

[0037] Figure 3 Comparison chart of sparse point clouds using only the training set with sparse views and using the training set plus the validation set during the motion construction stage for Scenario 1.

[0038] Figure 4 Qualitative comparison chart of rendered images and real images of the test set in the comparative experiment of three-dimensional reconstruction methods with different scenarios and different sparse views. Specific implementation manners

[0039] The technical solution of the present invention will be further explained and illustrated in conjunction with the accompanying drawings;

[0040] Embodiment 1

[0041] The process of generating validation set images by the generative NVS model and the validation set-guided Gaussian quantity control strategy is as Figure 1 shown, and specifically includes the following steps:

[0042] Step 1: First, select a certain number of pictures from the dense view benchmark dataset for three-dimensional reconstruction as the training set for the sparse view three-dimensional reconstruction task, and the remaining pictures are used as the test set. In the embodiment, the Mip-NeRF 360 dataset, the LLFF dataset, and the tandt dataset are used; for the Mip-NeRF 360 dataset, the number of training sets is set to 12; for the LLFF dataset, the number of training sets is set to 3; for the tandt dataset, the number of training sets is set to 24. Thus, the training set and the test set under the sparse view are obtained for further validation set view generation and initialization.

[0043] Step 2: Generation of the validation set view.

[0044] Taking the training set under the sparse view as the input, use ViewCrafter in the generative NVS model to generate synthetic images of new views. The number of input images is N. Assuming the number of generated images is A, then for each pair of adjacent input images, X new views (X is default selected as 25 in this embodiment) will be generated between them. Due to the pre-training model limitation of ViewCrafter, the generated image size is 1024×576 (pixels).

[0045] It is not reasonable to directly use these synthetic images as the validation set because some of the synthetic views have significant hallucinations and distorted details, such as Figure 2 as shown, and they need to be excluded from the candidate images for the validation set. For this purpose, a distorted image filtering strategy is designed to exclude distorted images with low geometric consistency by calculating the per-pixel reprojection error between each training set image and the candidate images for the validation set.

[0046] Specifically, first, feature extraction is performed on all training set images and candidate images for the validation set, and Scale-Invariant Feature Transform (SIFT) points and corresponding descriptors are extracted from each image. For each pair of training set images and candidate images for the validation set, feature matching is performed using the FLANN algorithm to obtain corresponding feature point pairs. Then, the essential matrix is estimated using the RANSAC algorithm, and this matrix represents the relative rotation and translation relationship between the two views.

[0047] Based on the estimated essential matrix , a confidence map between the reprojected image and the original input image is calculated. First, using the camera intrinsic matrix and the relative motion parameters (rotation matrix and translation vector ), the projection position of each pixel in the candidate image for the validation set in the training set view is calculated. The pixel content of the candidate image for the validation set is back-projected onto the training set view to generate a reprojected image. For the empty pixel regions that appear during the back-projection process, bilinear interpolation is used for filling. The pixel difference between the reprojected image and the training set image is calculated to generate a per-pixel confidence map. After obtaining the reprojection pixel confidence maps between each candidate image for the validation set and all training set images, the number of low-confidence pixels in the reprojected views is counted , and the training set image corresponding to the minimum number of low-confidence pixels is found as the closest training set view to this candidate image for the validation set. A hallucination determination threshold is set, which is 0.1 in the embodiment. If the number of low-confidence pixels between this closest training set view and the candidate image for the validation set is excessive, it is determined that this candidate image for the validation set contains too much hallucination content and should be discarded. The finally retained candidate images for the validation set constitute the validation set.

[0048] The obtained validation set is only used to optimize the initial Gaussian point cloud during the initialization phase and for supervision during the subsequent training phase, and is not directly used for training, so as to avoid interfering with the training set obtained from real scenes.

[0049] Step 3. Resize and crop the images of the training set and the test set to be the same size as the images of the validation set, which is 1024×576 (pixels), for subsequent motion construction. In traditional sparse Figure 3 In the DGS method, the generation of the initial point cloud often only depends on a limited number of input training set images, which limits the richness of the initial geometric information. To obtain accurate image poses and make full use of the synthetic images of the validation set, a joint initialization method is designed to generate the initial point cloud by combining the sparse input training set images and the generated validation set images. Although the validation set images contain some hallucination details, they provide a large amount of correct geometric structures and richer 3D geometric information. The joint initialization requires 2 times of Structure from Motion (SfM), both of which are implemented using the COLMAP tool.

[0050] Specifically, the input for the first motion construction is the training set and the test set images, which generates the camera poses corresponding to each image in the camera intrinsics; the input for the second motion construction is the training set, the validation set, and the camera poses obtained from the first motion construction. When extracting features, set the camera intrinsics as the input camera intrinsics, so as to obtain the camera poses of the validation set under the unified camera intrinsics and the sparse point cloud generated by the training set plus the validation set. Compared with the sparse point cloud generated only using the training set, it has a better global geometric structure. Figure 3 It is a comparison diagram of the initial point cloud using only the sparse view training set and the sparse point cloud using the training set plus the validation set in the motion construction stage.

[0051] Step 4. Create a Gaussian model and initialize the sparse point cloud obtained in Step 3 as Gaussian spheres and store them in the Gaussian model.

[0052] Each Gaussian sphere in the Gaussian model contains four key attributes: coordinates, covariance matrix, opacity, and spherical harmonic functions. Among them, the coordinates represent the position of the Gaussian sphere in the three-dimensional scene, the covariance matrix determines the shape of the Gaussian sphere, and the opacity and the spherical harmonic functions work together to represent the color information of the Gaussian sphere.

[0053] The construction of the above Gaussian spheres is consistent with the basic 3DGS. Although there may be certain differences between the Gaussian models constructed by different sparse view 3D reconstruction methods and the basic 3DGS during the Gaussian sphere initialization stage, each method needs to convert the sparse point cloud into the Gaussian spheres constructed by their respective methods. This method does not involve changes to the Gaussian sphere initialization process in different sparse view 3D reconstruction methods, and only uses the sparse point cloud after joint initialization in Step 3 as the input for the Gaussian sphere initialization stage.

[0054] Step 5. Design a training strategy for guiding the control of the number of Gaussians in the validation set.

[0055] This includes introducing a validation set for Gaussian sphere quantity supervision, a model reduction strategy, and an early stopping strategy. The number of Gaussian spheres is dynamically adjusted through the feedback of the validation set during the iterative training process to avoid overfitting. Although the validation set images contain a small amount of hallucinations, their impact is negligible when used for quantitative evaluation. The generated validation set images are used as a substitute for the real images, and the difference between the rendered images of the 3DGS under the validation set view and the validation set images is calculated to construct the validation set supervision metric. , and its calculation formula is as follows:

[0056]

[0057] Among them, represents the number of images in the validation set, is the validation set image, represents the rendered image under the validation set view. Since the Peak Signal-to-Noise Ratio (PSNR) and the validation set supervision metric are inversely correlated, during the training process, whether overfitting occurs is judged by monitoring the change of , and the optimal number of Gaussians is determined.

[0058] Specifically, the training strategy of validation set-guided Gaussian quantity control divides the iterative training process of 3DGS into two stages, and divides the first stage and the second stage by setting the variable: the validation set supervision stop round, which is 5000 iterations; setting the variable: the validation set supervision interval, which is 100 iterations. In the first stage, at the beginning of the iterative training, the optimal validation set supervision metric is set. When Gaussian densification starts, if the current iteration number is exactly an integer multiple of the validation set supervision interval, validation set supervision is performed, and the validation set supervision metric is calculated. If , then is updated, and the optimal number of Gaussians is recorded as the number of Gaussians at the current iteration. In the second stage, based on the original Gaussian densification, a model reduction strategy and an early stopping strategy are designed.

[0059] The model reduction strategy is specifically as follows: Since the Gaussian densification process is not modified in the first stage, this still leads to overfitting, but the optimal number of Gaussians obtained through introducing the validation set for supervision in the first stage is less than the number of Gaussians at the current iteration. Therefore, the model reduction strategy is carried out. The model reduction strategy randomly discards Gaussian spheres to reduce the number of Gaussians to the previous non-overfitting state. The model reduction strategy also sets the variable: the over-quantity reduction ratio , set to 0.2. After randomly discarding Gaussian spheres each time, this variable decreases by 0.05 until it reaches 0. Therefore, when randomly discarding Gaussian spheres each time, the specific number of discarded spheres The calculation formula is as follows:

[0060]

[0061] The early stopping strategy is specifically as follows: After randomly discarding Gaussian spheres, still perform the original Gaussian densification and maintain the supervision of the number of Gaussians after each densification operation. Once after the current iteration, the number of Gaussians is greater than the optimal number of Gaussians , immediately perform the random discarding of Gaussian spheres in the model reduction strategy.

[0062] Loop and execute the model reduction strategy and the early stopping strategy until the end of the Gaussian densification stage, thereby achieving the control of the number of Gaussian spheres within the controllable range of the optimal number of Gaussians , reaching a non-overfitting state, and realizing the improvement of the sparse-view 3D reconstruction quality and the acceleration of the training speed.

[0063] Step 6: Integrate the training strategies for generating validation set images and controlling the number of validation set-guided Gaussians designed in the foregoing steps into the sparse-view 3DGS method (such as SparseGS, FSGS, DNGaussian, CoR-GS, etc.) architecture based on the original 3DGS and capable of controlling the number of Gaussians by changing the densification gradient threshold. During the operation, follow the principle of "lossless integration", that is, ensure that each inserted improvement measure does not interfere with or damage the main structure of the original 3DGS model framework. Thanks to this scientific integration method and the consumption of computing resources and storage space saved in the foregoing steps, the entire 3D reconstruction process can be quickly and conveniently optimized and upgraded, obtaining better new-view rendering quality, faster training and rendering speeds, and improving the applicability and wide usability of the original sparse-view 3DGS method.

[0064] Example 2

[0065] To further introduce the sparse-view 3D reconstruction method with generative validation image-guided Gaussian number control, taking the sparse-view 3DGS method with SparseGS as the benchmark, after integrating Steps 1 to 5 of Example 1, the operation process is as follows:

[0066] Process 1: Data preparation.

[0067] Obtain a dense view benchmark dataset for 3D reconstruction and divide it into a training set and a test set for the sparse view 3D reconstruction task; use a generative NVS model to generate new view images based on the training set views. Since some synthesized views have significant distortion, they need to be excluded from the validation set images. Through the designed distorted image filtering strategy, distorted images with low geometric consistency are eliminated to generate low-distortion validation set views. The image sizes of the training set and the test set are scaled and cropped to be consistent with the validation set image size. Perform two motion constructions (Structure from Motion, SfM): the first input is the training set and the test set to obtain a new camera pose participating in the camera; the second input is the training set, the validation set, and the camera pose obtained by the first SfM, and the validation set camera pose under the unified camera intrinsic parameters and a sparse point cloud with better global geometric structure are obtained.

[0068] Process 2: Initialization phase.

[0069] According to the SparseGS method, a 3D Gaussian model is created, and the sparse point cloud obtained in process 1 is initialized as a Gaussian sphere and stored in the 3D Gaussian model. The rest of the initialization process and the setting of initialization parameters are consistent with the SparseGS method.

[0070] Process 3: Iterative training.

[0071] In the iterative training process of the SparseGS method, a designed validation set is added to guide the training strategy of Gaussian number control. Specifically, during the training process of the original SparseGS method, the 3D Gaussian sphere is first projected to different training set perspectives according to the given training set camera pose. And it is rendered through the rasterizer to obtain the rendered training set image as well as the depth map and visibility mask. By calculating the sum of the loss function, the gradient backpropagation step and the optimization step are performed to update the parameters of the 3D Gaussian model. And in the Gaussian densification stage, the Gaussian sphere is cloned, split and eliminated according to the properties of the 3D Gaussian sphere. In the Gaussian densification stage and optimization step of the original SparseGS method, the designed validation set is added to guide the training strategy of Gaussian number control without changing the rest of the training process.

[0072] For the Gaussian densification stage, the SparseGS method sets the Gaussian densification stage from 500 iterations to 18,000 iterations. The present invention designs a training strategy for the control of the number of Gaussians guided by the validation set, and divides the original Gaussian densification stage into two stages. The first stage is from 500 iterations to 5,000 iterations. Based on the optimization steps of the SparseGS method, the validation set is added for supervision to obtain the optimal number of Gaussians. ; In the second stage, from the 5001st iteration to the 18000th iteration, the Gaussian densification step of the SparseGS method is replaced by repeatedly executing the model reduction strategy and the early stopping strategy until the end of the Gaussian densification stage. This realizes controlling the number of Gaussian spheres within the controllable range of the optimal Gaussian number to alleviate the overfitting phenomenon of the original SparseGS method.

[0073] After the Gaussian densification stage ends, continue the iterative training according to the SparseGS method, continuously optimizing the parameters of the 3D Gaussian model until the set number of training iterations (30000 iterations in the SparseGS method) is reached to end the training. After the training ends, the obtained 3D Gaussian model can accurately reflect the geometric structure and appearance details of the scene corresponding to the sparse view, presenting an excellent 3D reconstruction effect.

[0074] Process 4: Novel view rendering. The 3D Gaussian model obtained after iterative training needs to be rendered with a novel view test set to evaluate the quality of scene reconstruction. Use the rasterizer in the SparseGS method to render the test set, project the 3D Gaussian model onto a 2D plane, and obtain the rendered image of the novel view.

[0075] Finally, taking the SparseGS method as a benchmark, the proposed sparse view 3D reconstruction method with generative verification image-guided Gaussian number control achieves better novel view rendering quality, faster training and rendering speeds, and improves the applicability and wide usability of the original sparse view 3DGS method.

[0076] To demonstrate the improvement of the proposed sparse view 3D reconstruction method with generative verification image-guided Gaussian number control on the reconstruction effect of the original sparse view 3D reconstruction method, large-scale comparative experiments are carried out. Three benchmark datasets (all re-scaled and cropped) are selected, namely the Mip-NeRF360 dataset, the tandt dataset, and the LLFF dataset; and 4 sparse view 3D reconstruction methods, namely FSGS, CoR-GS, DNGaussian, and SparseGS. Four metrics, namely peak signal-to-noise ratio PSNR, structural similarity SSIM, perceptual difference LPIPS, and the number of Gaussian spheres Number (in units of 1000, written as K), are selected to evaluate the results, as shown in Table 1:

[0077] Table 1

[0078]

[0079] In contrast, when combined with our proposed method, all sparse-view 3D reconstruction methods show significant performance improvements on three different datasets. Besides better rendering quality, the memory consumption is also greatly reduced. For example, when combining our method with CoR-GS on the LLFF dataset, the PSNR of the test views is increased by 1.7 dB, while the storage cost is reduced by 40%. Figure 4 The qualitative comparison graphs of the rendered images and the ground truth images of the test sets for different scenarios and different sparse-view 3D reconstruction methods in the comparative experiment further show that our method significantly reduces the number of Gaussians and enhances details in some regions.

[0080] To illustrate the positive impacts of the two major technical innovations in the proposed generative verification image-guided Gaussian number control strategy, namely the joint initialization of the validation set and the training set (hereinafter referred to as joint initialization for short in the following table) and the validation set supervised Gaussian number control (hereinafter referred to as number control for short in the following table), an ablation experiment was conducted to compare the sparse-view 3D reconstruction results of this method and the original SparseGS in the re-scaled and cropped Mip-NeRF 360 dataset, and five metrics, namely the peak signal-to-noise ratio PSNR, structural similarity SSIM, perceptual difference LPIPS, the number of Gaussian spheres Number (in units of 1000, written as K), and training time Train (min), were selected to evaluate the results, as shown in Table 2:

[0081] Table 2

[0082] Joint Initialization Quantity Control PSNR↑ SSIM↑ LPIPS↓ Num.↓ Train(min) 16.247 0.424 0.468 601K 32.0 √ 17.077 0.464 0.439 609K 31.8 √ 16.646 0.448 0.455 499K 25.9 √ √ 17.189 0.479 0.446 377K 28.5

[0083] It can be seen from the results that both the joint initialization of the validation set and the training set and the validation set supervised Gaussian number control have improvements in evaluation metrics compared to the original SparseGS, effectively alleviating the overfitting problem in sparse-view 3D reconstruction. In addition, this method shortens the training time and reduces the consumption of storage space determined by the number of Gaussians.

[0084] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, they can still make several substitutions or modifications to these described embodiments, and these substitution or modification methods should all be regarded as belonging to the protection scope of the present invention.

[0085] The parts not detailed in the present invention belong to the well-known technologies in the art.

Claims

1. A sparse-view three-dimensional reconstruction method based on verified image-guided Gaussian number control, characterized in that It includes the following steps: Step 1: Divide the dense view benchmark dataset for 3D reconstruction into a training set and a test set for sparse view 3D reconstruction tasks; Step 2: Use the generative NVS model to generate new view images; through the designed distorted image filtering strategy, obtain the final validation set views; Step 3: Through Structure from Motion (SfM), obtain the camera poses of the validation set under unified camera intrinsics and a sparse point cloud with a better global geometric structure; Step 4: Create a Gaussian model for training and initialize the Gaussian sphere for the sparse point cloud obtained in Step 3; Step 5: Design a training strategy for controlling the number of Gaussians guided by the validation set; Step 6: Integrate the above steps into the existing sparse view 3D reconstruction method to optimize the sparse view 3D reconstruction effect and obtain better new view rendering quality.

2. The sparse-view three-dimensional reconstruction method based on verified image-guided Gaussian number control according to claim 1, wherein The specific operations of Step 2 are as follows: Taking the training set under sparse views as the input, use ViewCrafter in the generative NVS model to generate synthetic images of new views as candidate images for the validation set; Design a distorted image filtering strategy. By calculating the per-pixel reprojection error between each training set image and the candidate images of the validation set, exclude the distorted images with low geometric consistency; the finally retained candidate images of the validation set constitute the validation set.

3. The sparse-view three-dimensional reconstruction method based on verified image-guided Gaussian number control according to claim 1, wherein The specific operations of Step 3 are as follows: Scale and crop the image sizes of the training set and the test set to be the same as that of the validation set images. Design a joint initialization method to generate an initial point cloud by combining the sparse input training set images and the generated validation set images; The joint initialization requires 2 times of Structure from Motion; Specifically, the input of the first Structure from Motion is the training set and test set images, and the camera intrinsics corresponding to each image and the camera poses are generated; the input of the second Structure from Motion is the training set, the validation set, and the camera poses obtained from the first Structure from Motion. When extracting features, set the camera intrinsics as the input camera intrinsics, thereby obtaining the camera poses of the validation set under unified camera intrinsics and the sparse point cloud generated by the training set plus the validation set.

4. The sparse-view three-dimensional reconstruction method based on verified image-guided Gaussian number control according to claim 1, wherein The training strategy for controlling the number of Gaussians guided by the validation set is specifically as follows: The training strategy includes introducing the validation set for Gaussian sphere number supervision, a model reduction strategy, and an early stopping strategy. Dynamically adjust the number of Gaussian spheres through the feedback of the validation set during the iterative training process to avoid overfitting; The training strategy divides the iterative training process of 3DGS into two stages. By setting variables: the number of rounds for validation set supervision to stop, the first stage and the second stage are divided; at the same time, the interval for validation set supervision is set; in the first stage, at the beginning of iterative training, the optimal index for validation set supervision is set , after Gaussian densification starts, if the current iteration round is exactly an integer multiple of the interval for validation set supervision, validation set supervision is performed, and the validation set supervision index is calculated , if , then update , and record the optimal number of Gaussians as the number of Gaussians under the current iteration; In the second stage, based on Gaussian densification, a model reduction strategy and an early stopping strategy are designed; the model reduction strategy reduces the number of Gaussians to a non-overfitting state by randomly discarding Gaussian spheres; The early stopping strategy is specifically: after randomly discarding Gaussian spheres, still perform Gaussian densification and maintain the supervision of the number of Gaussians after each densification operation. Once after the current iteration, the number of Gaussians is greater than the optimal number of Gaussians, immediately perform the random discarding of Gaussian spheres in the model reduction strategy; Loop and execute the model reduction strategy and the early stopping strategy until the end of the Gaussian densification stage, thereby realizing the control of the number of Gaussian spheres within the controllable range of the optimal number of Gaussians and achieving a non-overfitting state.

5. The sparse-view three-dimensional reconstruction method based on verification image-guided Gaussian number control according to claim 1, wherein, Integrate the training strategy of generating validation set images and controlling the number of Gaussian distributions in the validation set using the generative NVS model designed in the above steps into the sparse-view 3DGS method architecture based on the original 3DGS, which can control the number of Gaussian distributions by changing the densification gradient threshold; During the operation, follow the principle of "lossless integration", that is, ensure that each inserted improvement measure does not interfere with or damage the main structure of the original 3DGS model framework.

6. A sparse-view three-dimensional reconstruction system based on verified image-guided Gaussian number control, characterized in that, It includes a dataset acquisition module, a validation set view generation module, a motion construction module, a Gaussian model creation module, a training strategy module, and a sparse-view three-dimensional reconstruction module; The dataset acquisition module divides the dense-view benchmark dataset for three-dimensional reconstruction into a training set and a test set for sparse-view three-dimensional reconstruction tasks; The validation set view generation module uses the generative NVS model to generate new view images based on the training set views; by designing a distortion image filtering strategy, eliminate the distorted images with low geometric consistency to obtain the final validation set views; The motion construction module scales and crops the image sizes of the training set and the test set to be the same as the image size of the validation set; Then perform motion construction SfM twice: the first input is the training set and the test set to obtain new camera intrinsics and camera poses; the second input is the training set, the validation set, and the camera poses obtained from the first SfM to obtain the camera poses of the validation set under unified camera intrinsics and a sparse point cloud with a better global geometric structure; The Gaussian model creation module uses the jointly initialized sparse point cloud obtained by the motion construction module as the input in the Gaussian sphere initialization stage to create a Gaussian model for training; The training strategy module adopts a training strategy of controlling the number of Gaussian spheres in the validation set, which includes introducing the validation set for Gaussian sphere number supervision, a model reduction strategy, and an early stopping strategy, and dynamically adjusts the number of Gaussian spheres through the feedback of the validation set during the iterative training process to avoid overfitting; The sparse-view three-dimensional reconstruction module adopts a sparse-view three-dimensional reconstruction method optimized by the dataset acquisition module, the validation set view generation module, the motion construction module, the Gaussian model creation module, and the training strategy module to achieve sparse-view three-dimensional reconstruction.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in any one of claims 1-5.

8. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements the method described in any one of claims 1-5.

Citation Information

Patent Citations

  • Novel view angle synthesis method based on Gaussian splash and fusing learnable basis function

    CN118505541A

  • Sparse view angle complex outdoor scene three-dimensional reconstruction method, electronic equipment and storage medium

    CN119832161A

  • Underwater three-dimensional scene reconstruction method and device based on three-dimensional Gaussian splashing and underwater imaging model

    CN120014164A

  • Method for reconstruction of magnetic resonance images

    US20120008843A1