Small sample new view synthesis method based on reinitialized three-dimensional Gaussian splashing

By employing a small-sample new view synthesis method based on reinitialized 3D Gaussian splashing, and utilizing sparse 3D point cloud reconstruction and a multi-stage sampling strategy, the initialization problem of new view synthesis under sparse perspectives is solved, and high-quality new view generation is achieved.

CN121190646APending Publication Date: 2025-12-23ZHEJIANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511302739.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing technologies suffer from artifacts, structural mismatches, and geometric distortions in the synthesis of new views from sparse perspectives due to the initialization dependence on dense point clouds. Furthermore, existing methods struggle to achieve high-quality new view synthesis without the aid of external data.

Method used

A small-sample new view synthesis method based on reinitialized 3D Gaussian splashing is adopted. Through sparse 3D point cloud reconstruction, spatial expansion hybrid sampling, intra-view detail sampling and cross-view consistency sampling, an optimized path is gradually constructed to form a high-quality new view synthesis.

Benefits of technology

Without relying on extended views and dense point clouds, it significantly improves structural fidelity and detail accuracy, achieving efficient and high-quality new view synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190646A_ABST
    Figure CN121190646A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample new view synthesis method based on reinitialization three-dimensional Gaussian splashing. The method comprises the following steps: firstly, acquiring sparse three-dimensional point cloud and camera internal and external parameters from a training view through a motion recovery structure algorithm; then, sampling points are generated in the point cloud bounding box by adopting a spatial expansion hybrid sampling strategy, and a coarse-grained Gaussian set is constructed and optimized; thirdly, obtaining a rendered image through Gaussian initialization and splash rendering, calculating pixel importance through depth errors and transmissivity, generating a fine-grained Gaussian set through back projection, and optimizing the fine-grained Gaussian set; and finally, calculating a sampling probability based on the cross-view contribution degree, screening key Gaussian distribution, and optimizing to form a final Gaussian set, thereby realizing high-quality new view synthesis. According to the method, the problem of sparsity difference is solved by eliminating extended view dependence, the multi-view consistency and local geometric details of a new view scene are improved, and high-quality new view synthesis of sparse data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of novel view synthesis methods, and particularly relates to a novel view synthesis method for small samples based on reinitialized three-dimensional Gaussian splashing. Background Technology

[0002] The novel view synthesis technique aims to generate scene images observed from any unseen perspective, based on a finite number of input viewpoint images. This technology has significant application value in fields such as virtual reality, augmented reality, immersive communication, and robot vision.

[0003] In recent years, Neural Radiation Field (NeRF) has utilized implicit neural networks to learn continuous representations of scenes. However, it exhibits limitations such as high overfitting risk, long training time, and difficulties in real-time rendering under sparse viewpoint conditions, making it unsuitable for efficient deployment and resource-constrained environments. In contrast, 3D Gaussian scatter plots, by representing scenes with explicit Gaussian volumes and combining them with differentiable rendering algorithms, can achieve efficient real-time rendering and output quality similar to NeRF. However, its high dependence on SfM point cloud estimation for initialization makes it still prone to problems such as floating artifacts, background mismatch, and geometric distortion under sparse viewpoint conditions.

[0004] Despite this, current mainstream 3DGS methods still heavily rely on dense point clouds generated by motion-recovery algorithms under extended views as initialization data. This presents a "sparseness inconsistency" problem compared to the setup that uses only sparsely trained views for compositing. This data leakage in initialization not only violates the fairness requirements of few-view compositing but may also introduce artifacts, structural mismatches, and other issues, weakening the method's versatility and practical applicability. Some methods attempting to circumvent this problem, such as using random point clouds or view-independent points for initialization, while avoiding extended view leakage, significantly reduce geometric quality and rendering effects. Therefore, how to achieve stable and high-quality Gaussian initialization using only sparsely trained views without introducing external data has become a key challenge in current few-view compositing. Summary of the Invention

[0005] The technical problem to be solved by this invention is how to complete high-quality new view synthesis using only small samples without relying on extended views and dense point clouds, and provides a small sample new view synthesis method based on reinitialized 3D Gaussian splashing.

[0006] To achieve the above-mentioned objectives, the present invention specifically adopts the following technical solution:

[0007] A method for synthesizing small-sample new views based on re-initialized 3D Gaussian splashing includes the following steps:

[0008] S1. Use the structure-of-motion algorithm to obtain sparse 3D point cloud, camera extrinsic matrix, and camera intrinsic matrix based on the pre-acquired training view; wherein, the camera extrinsic matrix contains rotation matrix and translation vector, and the camera intrinsic matrix contains focal length and principal point coordinates;

[0009] S2. By using a spatial expansion hybrid sampling strategy, a set of uniform and random sampling points is generated within the expanded bounding box of the sparse 3D point cloud, which is used to construct and optimize a coarse-grained Gaussian set.

[0010] S3. A 3D Gaussian set is formed by initializing the sparse 3D point cloud using a Gaussian initialization strategy. The 3D Gaussian set is then splashed with the camera extrinsic and intrinsic parameters of the trained view to obtain the rendered image. Based on the optimized coarse-grained Gaussian set, an in-view detail sampling re-initialization strategy is used to calculate the composite importance score of each pixel in the rendered image based on depth error and transmittance, thereby selecting key pixels. The selected key pixels are then back-projected into 3D scene points to form a fine-grained Gaussian set and optimized.

[0011] S4. Adopting cross-view Figure 1 The consistent sampling reinitialization strategy calculates the contribution score of each Gaussian in the optimized fine-grained Gaussian set across all training views. After converting the calculated contribution scores into sampling probabilities, key Gaussian distributions are selected from the optimized fine-grained Gaussian set according to the sampling probabilities to form the final Gaussian set and optimize it, thus completing the synthesis of new views of the scene.

[0012] Based on the above scheme, each step can be implemented in the following preferred manner.

[0013] As a preferred embodiment, in S1, the specific steps of obtaining the sparse 3D point cloud, camera extrinsic matrix, and camera intrinsic matrix using the structure-of-motion-reconstruction algorithm include: firstly, extracting feature points of the training view using the SIFT feature detection algorithm and establishing cross-view matching relationships to form matching point pairs; then, filtering matching point pairs based on the random sampling consensus algorithm and the five-point method to eliminate incorrect matches, and calculating the fundamental matrix based on the filtered matching point pairs to obtain the initial camera pose of the training view; using the initial camera pose and the filtered matching point pairs, obtaining the initial 3D point coordinates through triangulation, optimizing all camera poses and 3D point coordinates through bundle adjustment to minimize reprojection errors, and finally outputting the camera extrinsic matrix, camera intrinsic matrix, and sparse 3D point cloud for providing geometric constraints.

[0014] As a preferred option, the specific process of S2 is as follows:

[0015] S21: The sparse 3D point cloud reconstructed by the structure-of-motion algorithm is used as the initial sparse point cloud. The bounding box of the initial sparse point cloud is spatially expanded by scaling along the X-axis, Y-axis and Z-axis using preset expansion coefficients to form an expanded space; where the expansion coefficient is a real number greater than 1.

[0016] S22: Within the extended space, a hybrid sampling method combining uniform grid sampling and random sampling is used to generate sampling points with color attributes, forming a sampling point set. The coordinates of the generated sampling points are used as the initial position parameters of the coarse-grained Gaussian set, and the color attributes of the sampling points are used as the initial color parameters of the coarse-grained Gaussian set. Based on the generated sampling points, other parameters of the coarse-grained Gaussian set are obtained through a Gaussian initialization strategy, forming a coarse-grained Gaussian set. This set is then optimized using a training view to form an optimized coarse-grained Gaussian set. The number of sampling points generated by uniform grid sampling is obtained by rounding down the product of the preset uniform sampling coefficient and the three-axis side length of the extended space. The number of sampling points generated by random sampling is obtained by rounding down the product of the preset random sampling coefficient and the three-axis side length of the extended space.

[0017] As a preferred approach, S3's in-view detail sampling reinitialization strategy includes:

[0018] S31: For each pixel in the rendered image of each training viewpoint, obtain the importance score of the pixel calculated based on the depth error and the importance score calculated based on the transmittance, and sum the two importance scores by weight to obtain the composite importance score of the pixel.

[0019] S32: Based on the calculated composite importance score of each pixel, a preset number of key pixels are selected from the rendered image of each training viewpoint using the importance sampling method, and the key pixels are mapped to three-dimensional scene points.

[0020] S33: The obtained 3D scene point coordinates are used as the initial position parameters of the fine-grained Gaussian set, and the color value of the key pixel is used as the initial color parameter of the fine-grained Gaussian set. The other parameters of the fine-grained Gaussian set are obtained based on the generated key pixels through the Gaussian initialization strategy, forming a fine-grained Gaussian set. The set is then optimized through the training view to form an optimized fine-grained Gaussian set.

[0021] Preferably, in S31, for a pixel in the rendered image, the importance score of the pixel based on the depth error is calculated by comparing the rendered depth map with the pseudo-ground depth map generated by the pre-trained monocular depth estimation network; wherein, the rendered depth map is obtained by splashing a three-dimensional Gaussian set with the camera extrinsic matrix and the camera intrinsic matrix of the trained view.

[0022] As a preferred option, for a single pixel Its importance score is calculated based on depth error. The calculation formula is:

[0023]

[0024] in, This indicates a normalization operation; Represents pixels in the rendering depth map The depth value; Represents pixels in the pseudo-true depth map The depth value; This represents the L2 norm.

[0025] Preferably, in S31, for a pixel in the rendered image, the importance score of that pixel, calculated based on transmittance, is determined by analyzing the cumulative transmittance of the opacity parameter during the blending process.

[0026] As a preferred option, for a single pixel Its importance score is calculated based on transmittance. The specific calculation formula is as follows:

[0027]

[0028]

[0029] in, Indicating the first coarse-grained Gaussian set An opacity parameter of Gaussian along the ray direction; Indicates reaching the th Cumulative transmittance before Gauss; This represents the total number of Gaussians in the coarse-grained Gaussian set; Indicates the first An opacity parameter of Gaussian along the ray direction.

[0030] Preferably, in S32, for a key pixel Map it to 3D scene points using the following steps:

[0031] S321: First, calculate the depth value of the key pixel using the following formula. :

[0032]

[0033] in, and Represented as the first and A fixed opacity value preset by Gaussian; Indicates the first Gaussian coordinates; and These represent the rotation matrix and translation vector in the camera extrinsic matrix, respectively. Indicates the center position of the camera; This represents the total number of Gaussians in the coarse-grained Gaussian set; Represents the L2 norm;

[0034] S322: Combine the key pixel coordinates with its depth value, back-project the key pixel into a 3D scene point through the camera extrinsic matrix, and obtain the coordinates of the 3D scene point.

[0035] As a preferred option, cross-view in S4 Figure 1 Consistent sampling re-initialization strategies include:

[0036] S41: For the optimized fine-grained Gaussian set, the first... Gaussian Calculate its contribution score across all training views using the following formula. Then, the probability is normalized to convert the contribution score into a sampling probability:

[0037]

[0038] in, Indicates the coverage discriminant function, when The projection covers the pixels in the training view. hour, Otherwise, it is 0; This represents the set consisting of all training views; This represents a training view; Indicates the first An opacity parameter of Gaussian along the ray direction; Indicates reaching the th Cumulative transmittance before Gauss;

[0039] S42: Based on the calculated sampling probability, a preset number of Gaussian distributions are randomly selected from the optimized fine-grained Gaussian set using the importance sampling method. The sampled Gaussian distributions are used as key Gaussian distributions, and the position parameters and color features of the key Gaussian distributions are used as the initialization parameters of the final Gaussian set. Other parameters of the final Gaussian set are obtained based on the generated Gaussian distributions through the Gaussian initialization strategy. The final Gaussian set is then formed and optimized by training the view to form the optimized final Gaussian set, thus completing the synthesis of the new view of the scene.

[0040] Compared with the prior art, the present invention has the following advantages:

[0041] This invention proposes a novel small-sample view synthesis method based on re-initialized 3D Gaussian splashing, with a core design of a two-stage coarse-to-fine re-initialization strategy. The first stage employs a spatial expansion hybrid sampling algorithm, sampling points at high density in an expanded 3D space using a hybrid approach to ensure global structure coverage and initialization path convergence. The second stage adopts a view-driven detail-aware sampling strategy, guiding pixel sampling for perceiving local details based on the pre-optimization results of the previous stage, and combining intra-view and cross-view Gaussian point selection mechanisms to achieve high-precision synthesis. This invention effectively alleviates the sparsity inconsistency problem and improves structure restoration and detail fidelity without relying on expanded views and dense point clouds. Validation results show that this method significantly outperforms existing methods on multiple standard datasets. Attached Figure Description

[0042] Figure 1 This is a flowchart of the method of the present invention;

[0043] Figure 2 This is a flowchart of the in-view detail sampling re-initialization strategy of the present invention;

[0044] Figure 3 For the cross-view of the present invention Figure 1 Flowchart of consistent sampling re-initialization strategy;

[0045] Figure 4 These are comparative experimental images of the Trex scene in the embodiments of the present invention;

[0046] Figure 5 These are comparative experimental images of the room scene in the embodiments of the present invention;

[0047] Figure 6 These are comparative experimental images of the leaves scenario in the embodiments of the present invention;

[0048] Figure 7 These are comparative experimental images of the fortress scene in the embodiments of the present invention;

[0049] Figure 8 These are comparative experimental diagrams of horns scenarios in embodiments of the present invention;

[0050] Figure 9 This is a comparison diagram of the fern scene experiment in the embodiments of the present invention. Detailed Implementation

[0051] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in various embodiments of the present invention can be combined accordingly without mutual conflict.

[0052] This invention proposes a small-sample new view synthesis method based on reinitialization of 3D Gaussian splashing. Its core idea is to rely only on small samples during the training phase and adaptively construct an optimized path through a stepwise reinitialization strategy, thereby achieving high-quality and highly consistent new view synthesis.

[0053] The method of this invention generally includes: sparse point cloud reconstruction, spatially extended hybrid sampling, in-view detail sampling, and cross-view... Figure 1 The process consists of four parts: consistent sampling, final Gaussian optimization, spatial expansion mixture sampling (coarse-grained Gaussian initialization stage), in-view detail sampling, and cross-view sampling. Figure 1 Consistent sampling is a fine-grained Gaussian re-initialization phase, such as... Figure 1 As shown.

[0054] In a preferred embodiment of the present invention, the above-mentioned method for synthesizing small-sample new views based on re-initialized 3D Gaussian splashing includes the following steps S1 to S4. The specific implementation process of each step is described in detail below.

[0055] S1. Use the structure-of-motion algorithm to obtain sparse 3D point cloud, camera extrinsic matrix, and camera intrinsic matrix based on the pre-acquired training view; wherein, the camera extrinsic matrix contains rotation matrix and translation vector, and the camera intrinsic matrix contains focal length and principal point coordinates.

[0056] It should be noted that, in step S1 of this invention, the specific steps for obtaining the sparse 3D point cloud, camera extrinsic matrix, and camera intrinsic matrix using the motion reconstruction structure algorithm include:

[0057] First, the SIFT feature detection algorithm is used to extract feature points from the training view and establish cross-view matching relationships to form matching point pairs. Then, the matching point pairs are filtered using the Random Sample Consensus (RANSAC) algorithm and the five-point method to eliminate erroneous matches. The fundamental matrix is ​​calculated based on the filtered matching point pairs to obtain the initial camera pose of the training view. Using the initial camera pose and the filtered matching point pairs, the initial 3D point coordinates are obtained through triangulation. Bundle adjustment (BA) is then used to optimize all camera poses and 3D point coordinates to minimize reprojection errors. Finally, the camera extrinsic matrix, camera intrinsic matrix, and sparse 3D point cloud for providing geometric constraints are output. These parameters together determine the accurate projection relationship from 3D spatial points to 2D pixels.

[0058] S2. By using a spatially extended hybrid sampling strategy, a set of uniformly and randomly mixed sampling points is generated within the extended bounding box of the sparse 3D point cloud, which is then used to construct and optimize a coarse-grained Gaussian set.

[0059] It should be noted that the specific process of step S2 in this invention is as follows:

[0060] S21: Using the sparse 3D point cloud reconstructed by the structure-of-motion algorithm as the initial sparse point cloud, the bounding box of the initial sparse point cloud is spatially expanded by applying preset expansion coefficients along the X, Y, and Z axes respectively. After scaling, an extended space is formed. .in, And it is a real number.

[0061] In this embodiment, the expansion factor is set to 1.3. Of course, it can also be set by those skilled in the art according to actual needs, and is not limited in this invention.

[0062] S22: Within the extended space, a hybrid sampling method combining uniform grid sampling and random sampling is used to generate sampling points with color attributes, forming a sampling point set. The coordinates of the generated sampling points are used as the initial position parameters of the coarse-grained Gaussian set, and the color attributes of the sampling points are used as the initial color parameters of the coarse-grained Gaussian set. Other parameters of the coarse-grained Gaussian set are obtained based on the generated sampling points using a Gaussian initialization strategy, forming the coarse-grained Gaussian set. This set is then optimized using a training view to form the optimized coarse-grained Gaussian set. The number of sampling points generated by uniform grid sampling... The preset uniform sampling coefficient With the three-axis side length of the extended space , , The product result is obtained by rounding down; the number of sampling points generated by random sampling. The preset random sampling coefficients With the three-axis side length of the extended space , , The product is obtained by rounding down.

[0063] In this embodiment, sampling points generated by uniform grid sampling are used for global geometric structural features of the scene, while sampling points generated by random sampling are used to supplement local geometric detail features of the scene.

[0064] S3. A 3D Gaussian set is formed by initializing the sparse 3D point cloud using a Gaussian initialization strategy. The 3D Gaussian set is then splashed with the camera extrinsic and intrinsic parameters of the trained view to obtain the rendered image. Based on the optimized coarse-grained Gaussian set, an in-view detail sampling re-initialization strategy is used to calculate the composite importance score of each pixel in the rendered image based on depth error and transmittance, thereby selecting key pixels. The selected key pixels are then back-projected into 3D scene points to form a fine-grained Gaussian set and optimized.

[0065] It should be noted that in this invention, for coarse-grained Gaussian sets, an in-view detail sampling reinitialization strategy is used to find pixels in the rendered image that play a key role in the quality of reconstructed details. For example... Figure 2 As shown, the in-view detail sampling reinitialization strategy in step S3 includes:

[0066] S31: For each pixel in the rendered image of each training viewpoint, obtain the importance score of the pixel calculated based on the depth error and the importance score calculated based on the transmittance, and sum the two importance scores by weight to obtain the composite importance score of the pixel.

[0067] In this embodiment, for a pixel Its composite importance score The calculation formula is:

[0068]

[0069] in, This is a preset weighting adjustment coefficient used to balance the contribution ratio of depth information and transmittance information; Represents pixels Importance score calculated based on depth error; Represents pixels Importance score based on transmittance calculation.

[0070] Furthermore, for a pixel in the rendered image, the importance score of that pixel based on the depth error is calculated by comparing the rendered depth map with the pseudo-ground depth map generated by the pre-trained monocular depth estimation network; wherein, the rendered depth map is obtained by splashing a 3D Gaussian set with the camera extrinsic matrix and camera intrinsic matrix of the trained view.

[0071] In this embodiment, for a pixel Its importance score is calculated based on depth error. The calculation formula is:

[0072]

[0073] in, This indicates a normalization operation; Represents pixels in the rendering depth map The depth value; Represents pixels in the pseudo-true depth map The depth value; This represents the L2 norm.

[0074] Furthermore, for a pixel in the rendered image, that pixel has an importance score calculated based on its transmittance. By analyzing the opacity parameter The cumulative transmittance during the mixing process is determined.

[0075] In this embodiment, for a pixel Its importance score is calculated based on transmittance. The formula used to identify areas with insufficient rendering is as follows:

[0076]

[0077]

[0078] in, Indicating the first coarse-grained Gaussian set An opacity parameter of Gaussian along the ray direction; Indicates reaching the th Cumulative transmittance before Gauss; This represents the total number of Gaussians in the coarse-grained Gaussian set; Indicates the first An opacity parameter of Gaussian along the ray direction.

[0079] S32: Based on the calculated composite importance score of each pixel, a preset number of key pixels are selected from the rendered image of each training viewpoint using the importance sampling method, and the key pixels are mapped to three-dimensional scene points.

[0080] It should be noted that in S32 of the present invention, for a key pixel Map it to 3D scene points using the following steps:

[0081] S321: First, calculate the depth value of the key pixel using the following formula. :

[0082]

[0083] in, and Represented as the first and A fixed opacity value preset by Gaussian; Indicates the first Gaussian coordinates; and These represent the rotation matrix and translation vector in the camera extrinsic matrix, respectively. Indicates the center position of the camera.

[0084] S322: Combine the key pixel coordinates with its depth value, back-project the key pixel into a 3D scene point through the camera extrinsic matrix, and obtain the coordinates of the 3D scene point.

[0085] S33: The obtained 3D scene point coordinates are used as the initial position parameters of the fine-grained Gaussian set, and the color value of the key pixel is used as the initial color parameter of the fine-grained Gaussian set. The other parameters of the fine-grained Gaussian set are obtained based on the generated key pixels through the Gaussian initialization strategy, forming a fine-grained Gaussian set. The set is then optimized through the training view to form an optimized fine-grained Gaussian set.

[0086] S4. Adopting cross-view Figure 1 The consistent sampling reinitialization strategy calculates the contribution score of each Gaussian in the optimized fine-grained Gaussian set across all training views. After converting the calculated contribution scores into sampling probabilities, key Gaussian distributions are selected from the optimized fine-grained Gaussian set according to the sampling probabilities to form the final Gaussian set and optimize it, thus completing the synthesis of new views of the scene.

[0087] In this embodiment, cross-viewing is used for the optimized fine-grained Gaussian set. Figure 1 The consistent sampling re-initialization strategy aims to retain only Gaussian points with strong geometric capabilities while removing low-quality Gaussian points, thereby improving the upper limit of reconstruction capability. For example... Figure 3 As shown, cross-view in S4 Figure 1 Consistent sampling re-initialization strategies include:

[0088] S41: For the optimized fine-grained Gaussian set, the first... Gaussian Calculate its contribution score across all training views using the following formula. Then, the probability is normalized to convert the contribution score into a sampling probability. :

[0089]

[0090]

[0091] in, Indicates the coverage discriminant function, when The projection covers the pixels in the training view. hour, Otherwise, it is 0; This represents the set consisting of all training views; This represents a training view; For the optimized fine-grained Gaussian set, the first The contribution score of each Gaussian in all training views; Indicates the first An opacity parameter of Gaussian along the ray direction; Indicates reaching the th Cumulative transmittance before Gauss; This represents the total number of Gaussians in the optimized fine-grained Gaussian set.

[0092] S42: Based on the calculated sampling probability, a preset number of Gaussian distributions are randomly selected from the optimized fine-grained Gaussian set using the importance sampling method. The sampled Gaussian distributions are used as key Gaussian distributions, and the position parameters and color features of the key Gaussian distributions are used as the initialization parameters of the final Gaussian set. Other parameters of the final Gaussian set are obtained based on the generated Gaussian distributions through the Gaussian initialization strategy. The final Gaussian set is then formed and optimized by training the view to form the optimized final Gaussian set, thus completing the synthesis of the new view of the scene.

[0093] The present invention will now demonstrate the application effect of the small sample new view synthesis method based on re-initialized 3D Gaussian splashing described in S1~S4 of the above embodiments on a specific dataset through a specific example, so as to facilitate understanding of the essence of the present invention.

[0094] Example

[0095] The specific implementation process of the small sample new view synthesis method based on re-initialization of three-dimensional Gaussian splashing used in this embodiment is as described above and will not be repeated here.

[0096] This invention, using the LLFF dataset, conducted comparative tests on the novel view synthesis effects of its method with the CoR-GS and DNGaussian methods based on a small sample size. Examples of rendering effects in different scenarios are shown below. Figures 4-9 As shown in Table 1, the indicators are as follows. Figure 4 For Trex scenarios, Figure 5 Corresponding to the room scene, Figure 6 Corresponding to the leaves scenario, Figure 7 Corresponding to the fortress scenario, Figure 8 Corresponding to horns scenarios, Figure 9 For the corresponding Fern scene. In Table 1, PSNR is the peak signal-to-noise ratio, SSIM is the structural similarity, and the higher the value of either indicator, the higher the rendering quality; LPIPS is the perceptual similarity, and the lower the value, the higher the rendering quality.

[0097] Table 1. Test Data Results

[0098]

[0099] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A method for synthesizing new views from small samples based on re-initialized 3D Gaussian splashing, characterized in that, Includes the following steps: S1. Use the structure-of-motion algorithm to obtain sparse 3D point cloud, camera extrinsic matrix, and camera intrinsic matrix based on the pre-acquired training view; wherein, the camera extrinsic matrix contains rotation matrix and translation vector, and the camera intrinsic matrix contains focal length and principal point coordinates; S2. By using a spatial expansion hybrid sampling strategy, a set of uniform and random sampling points is generated within the expanded bounding box of the sparse 3D point cloud, which is used to construct and optimize a coarse-grained Gaussian set. S3. A 3D Gaussian set is formed by initializing the sparse 3D point cloud using a Gaussian initialization strategy. The 3D Gaussian set is then splashed with the camera extrinsic and intrinsic parameters of the trained view to obtain the rendered image. Based on the optimized coarse-grained Gaussian set, an in-view detail sampling re-initialization strategy is used to calculate the composite importance score of each pixel in the rendered image based on depth error and transmittance, thereby selecting key pixels. The selected key pixels are then back-projected into 3D scene points to form a fine-grained Gaussian set and optimized. S4. Adopting a cross-view consistent sampling re-initialization strategy, calculate the contribution score of each Gaussian in all training views in the optimized fine-grained Gaussian set. After converting the calculated contribution scores into sampling probabilities, select key Gaussian distributions from the optimized fine-grained Gaussian set according to the sampling probabilities to form the final Gaussian set and optimize it to complete the synthesis of new views of the scene.

2. The method for synthesizing small-sample new views based on re-initialized 3D Gaussian splashing as described in claim 1, characterized in that, In S1, the specific steps for obtaining sparse 3D point clouds, camera extrinsic matrix, and camera intrinsic matrix using the structure-of-motion (SOG) algorithm include: First, using the SIFT feature detection algorithm to extract feature points from the training view and establishing cross-view matching relationships to form matching point pairs; then, using the random sampling consensus algorithm and the five-point method to filter matching point pairs, eliminating incorrect matches, and calculating the fundamental matrix based on the filtered matching point pairs to obtain the initial camera pose of the training view; using the initial camera pose and the filtered matching point pairs, obtaining the initial 3D point coordinates through triangulation, optimizing all camera poses and 3D point coordinates through bundle adjustment to minimize reprojection errors, and finally outputting the camera extrinsic matrix, camera intrinsic matrix, and sparse 3D point cloud for providing geometric constraints.

3. The method for synthesizing new views from small samples based on re-initialized 3D Gaussian splashing as described in claim 1, characterized in that, The specific process of S2 is as follows: S21: The sparse 3D point cloud reconstructed by the structure-of-motion algorithm is used as the initial sparse point cloud. The bounding box of the initial sparse point cloud is spatially expanded by scaling along the X-axis, Y-axis and Z-axis using preset expansion coefficients to form an expanded space; where the expansion coefficient is a real number greater than 1. S22: Within the extended space, a hybrid sampling method combining uniform grid sampling and random sampling is used to generate sampling points with color attributes, forming a sampling point set. The coordinates of the generated sampling points are used as the initial position parameters of the coarse-grained Gaussian set, and the color attributes of the sampling points are used as the initial color parameters of the coarse-grained Gaussian set. Based on the generated sampling points, other parameters of the coarse-grained Gaussian set are obtained through a Gaussian initialization strategy, forming a coarse-grained Gaussian set. This set is then optimized using a training view to form an optimized coarse-grained Gaussian set. The number of sampling points generated by uniform grid sampling is obtained by rounding down the product of the preset uniform sampling coefficient and the three-axis side length of the extended space. The number of sampling points generated by random sampling is obtained by rounding down the product of the preset random sampling coefficient and the three-axis side length of the extended space.

4. The method for synthesizing small-sample new views based on re-initialized 3D Gaussian splashing as described in claim 1, characterized in that, S3's in-view detail sampling reinitialization strategy includes: S31: For each pixel in the rendered image of each training viewpoint, obtain the importance score of the pixel calculated based on the depth error and the importance score calculated based on the transmittance, and sum the two importance scores by weight to obtain the composite importance score of the pixel. S32: Based on the calculated composite importance score of each pixel, a preset number of key pixels are selected from the rendered image of each training viewpoint using the importance sampling method, and the key pixels are mapped to three-dimensional scene points. S33: The obtained 3D scene point coordinates are used as the initial position parameters of the fine-grained Gaussian set, and the color value of the key pixel is used as the initial color parameter of the fine-grained Gaussian set. The other parameters of the fine-grained Gaussian set are obtained based on the generated key pixels through the Gaussian initialization strategy, forming a fine-grained Gaussian set. The set is then optimized through the training view to form an optimized fine-grained Gaussian set.

5. The method for synthesizing small-sample new views based on re-initialized 3D Gaussian splashing as described in claim 4, characterized in that, In S31, for a pixel in the rendered image, the importance score of that pixel based on the depth error is calculated by comparing the rendered depth map with the pseudo-ground depth map generated by the pre-trained monocular depth estimation network; wherein, the rendered depth map is obtained by splashing a 3D Gaussian set with the camera extrinsic matrix and camera intrinsic matrix of the trained view.

6. The method for synthesizing small-sample new views based on re-initialized 3D Gaussian splashing as described in claim 5, characterized in that, For a pixel Its importance score is calculated based on depth error. The calculation formula is: ; in, This indicates a normalization operation; Represents pixels in the rendering depth map The depth value; Represents pixels in the pseudo-true depth map The depth value; This represents the L2 norm.

7. The method for synthesizing new views from small samples based on re-initialized 3D Gaussian splashing as described in claim 4, characterized in that, In S31, for a pixel in the rendered image, the importance score of that pixel is determined by analyzing the cumulative transmittance of the opacity parameter during the blending process.

8. The method for synthesizing new views from small samples based on re-initialized 3D Gaussian splashing as described in claim 7, characterized in that, For a pixel Its importance score is calculated based on transmittance. The specific calculation formula is as follows: ; ; in, Indicates the first coarse-grained Gaussian set An opacity parameter of Gaussian along the ray direction; Indicates reaching the th Cumulative transmittance before Gaussian ohms; This represents the total number of Gaussians in the coarse-grained Gaussian set; Indicates the first An opacity parameter of Gaussian along the ray direction.

9. The method for synthesizing new views from small samples based on re-initialized 3D Gaussian splashing as described in claim 4, characterized in that, In S32, for a key pixel Map it to 3D scene points using the following steps: S321: First, calculate the depth value of the key pixel using the following formula. : ; in, and Represented as the first and A fixed opacity value preset by Gaussian; Indicates the first Gaussian coordinates; and These represent the rotation matrix and translation vector in the camera extrinsic matrix, respectively. Indicates the center position of the camera; This represents the total number of Gaussians in the coarse-grained Gaussian set; Represents the L2 norm; S322: Combine the key pixel coordinates with its depth value, back-project the key pixel into a 3D scene point through the camera extrinsic matrix, and obtain the coordinates of the 3D scene point.

10. The method for synthesizing small-sample new views based on re-initialized 3D Gaussian splashing as described in claim 1, characterized in that, The cross-view consistency sampling reinitialization strategy in S4 includes: S41: For the optimized fine-grained Gaussian set, the first... Gaussian Calculate its contribution score across all training views using the following formula. Then, the probability is normalized to convert the contribution score into a sampling probability: ; in, Indicates the coverage discriminant function, when The projection covers the pixels in the training view. hour, Otherwise, it is 0; This represents the set consisting of all training views; This represents a training view; Indicates the first An opacity parameter of Gaussian along the ray direction; Indicates reaching the th Cumulative transmittance before Gaussian ohms; S42: Based on the calculated sampling probability, a preset number of Gaussian distributions are randomly selected from the optimized fine-grained Gaussian set using the importance sampling method. The sampled Gaussian distributions are used as key Gaussian distributions, and the position parameters and color features of the key Gaussian distributions are used as the initialization parameters of the final Gaussian set. Other parameters of the final Gaussian set are obtained based on the generated Gaussian distributions through the Gaussian initialization strategy. The final Gaussian set is then formed and optimized by training the view to form the optimized final Gaussian set, thus completing the synthesis of the new view of the scene.

Citation Information

Cited By

  • Multi-model collaborative two-dimensional Gaussian splash three-dimensional reconstruction method

    CN121392157A

  • Physical attribute inversion and three-dimensional reconstruction method based on Gaussian splashing and micro rendering

    CN121482292A