A three-dimensional gaussian sputtering method based on image texture structure guidance
Patent Information
- Application Number
- CN202610576521.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-28
- Publication Date
- 2026-08-07
AI Technical Summary
[0008]为了解决现有3D 高斯溅射技术在稀疏视图新视角合成场景下,因观测信息不全易出现过拟合,进而导致重建结果几何精细结构缺失、边界模糊且产生漂浮伪影的核心问题,同时还需克服现有改进方法的诸多局限,如 FSGS 无差别致密化引入过量基元提升几何误差、DropGaussian 随机丢弃易丢失关键几何基元造成尖锐边界细节缺失,以及传统技术固定分裂尺度无法适配空间变化的几何复杂度、固定丢弃率的 dropout 机制易引发过正则化或欠正则化,损害场景表示紧凑性与表达能力的问题,本发明提出一种基于图像纹理结构指导的三维高斯溅射方法,包括以下步骤:
在合成精度上,于LLFF、Mip-NeRF360和Blender等标准数据集上的广泛实验表明,本方法在多项定量指标上均超越现有主流方法。
Smart Images

Figure CN122530399A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision technology, specifically relating to a three-dimensional Gaussian sputtering method guided by image texture structure. Background Technology
[0002] Existing 3D Gaussian sputtering techniques are prone to overfitting in sparse view synthesis scenarios due to incomplete observation information. This leads to a loss of fine geometric structure, blurred boundary regions, and floating artifacts in the reconstruction results. Current improvement methods for sparse views have significant limitations: for example, the indiscriminate densification strategy used in methods such as FSGS introduces excessive Gaussian primitives, which actually increases the geometric error in structured regions; while random dropout strategies such as DropGaussian can alleviate overfitting, they are prone to erroneously discarding primitives that are crucial to the geometric structure, resulting in the loss of sharp boundary details and a decrease in error map accuracy. In addition, traditional 3D Gaussian sputtering uses a fixed splitting scale, which cannot adapt to the geometric complexity of spatial variations within the scene, making it difficult to achieve a balance between preserving fine structure and maintaining surface continuity. Existing dropout mechanisms also mostly use a fixed dropout rate, failing to dynamically adjust according to the global distribution and local properties of Gaussian primitives, which can easily lead to overregulation or underregulation, impairing the compactness and expressiveness of the scene representation.
[0003] Novel perspective compositing (NVS) is a core task in computer vision and graphics, attracting significant attention driven by immersive technologies such as VR / AR. Early NeRF techniques achieved photorealistic rendering through implicit neural representations, but their reliance on voxel ray stepping resulted in extremely high computational costs. 3D Gaussian sputtering (3DGS), employing explicit point-based representations combined with differentiable rasterization, achieves an excellent balance between real-time rendering efficiency and visual quality, and has become the mainstream NVS technology. However, 3DGS inherently relies on dense multi-view input, making it prone to overfitting to a small number of visible observations under sparse view conditions, leading to incomplete reconstructed geometry and artifacts in insufficiently observed areas.
[0004] To address the challenges of sparse views, recent research has primarily followed two main technical routes: NeRF-based and 3DGS-based. Among NeRF-based methods, Mip-NeRF proposes using frustums for anti-aliasing representations, effectively mitigating aliasing caused by sparsity; DietNeRF implements semantic supervision through CLIP embedding; RegNeRF incorporates depth smoothing regularization; SparseNeRF utilizes monocular depth ranking priors to enhance geometric stability under sparse input; FreeNeRF proposes a frequency control strategy to reduce overfitting; ZeroRF achieves fast sparse view reconstruction without additional pre-training; Pixel-NeRF enriches scene understanding with limited observations by extracting contextual image features; and Reconfusion employs diffusion-based pseudo-view synthesis to supplement missing viewpoints with high quality. Nevertheless, NeRF-based representations are still constrained by low computational efficiency and time-consuming scene-by-scene optimization.
[0005] In improving sparse views based on 3DGS, existing methods can be broadly categorized into scene-by-scene optimization and feedforward networks. Among scene-by-scene optimization methods, FSGS improves coverage by interpolating primitives, but its indiscriminate densification often introduces excessive primitives, leading to increased geometric errors. DropGaussian uses a random drop mechanism to alleviate overfitting, but its randomness may result in the loss of key geometric primitives. Furthermore, methods such as PGDGS, CoR-GS, S2Gaussian, and AD-GS explore adaptive regularization, multi-field collaborative optimization, and self-supervised densification; DNGaussian, LoopSparseGS, CoherentGS, and NexusGS utilize monocular depth, epipolar geometry, or optical flow cues to stabilize geometry and improve scene integrity; Rl3D and GaussianObject employ diffusion-based pseudo-view generation to enhance supervision of sparse regions. However, ensuring strict multi-view consistency among these generated pseudo-views remains a challenge. Figure 1 Consistency remains an ongoing challenge. Feedforward network methods, such as PixelSplat, TranSplat, Mvsplat, and InstantSplat, can achieve extremely fast test-time rendering, but they usually require a lot of computational resources and have limited ability to capture fine geometry under sparse supervision, which can easily lead to overly smoothed or blurry reconstruction results. The cost volume technique that they often rely on also brings high computational costs when there are many input views.
[0006] Existing 3D Gaussian sputtering techniques are prone to overfitting in sparse view synthesis scenarios due to incomplete observation information. This leads to a loss of fine geometric structure, blurred boundary regions, and floating artifacts in the reconstruction results. Current improvement methods for sparse views have significant limitations: for example, the indiscriminate densification strategy used in methods such as FSGS introduces excessive Gaussian primitives, which actually increases the geometric error in structured regions; while random dropout strategies such as DropGaussian can alleviate overfitting, they are prone to erroneously discarding primitives that are crucial to the geometric structure, resulting in the loss of sharp boundary details and a decrease in error map accuracy. In addition, traditional 3D Gaussian sputtering uses a fixed splitting scale, which cannot adapt to the geometric complexity of spatial variations within the scene, making it difficult to achieve a balance between preserving fine structure and maintaining surface continuity. Existing dropout mechanisms also mostly use a fixed dropout rate, failing to dynamically adjust according to the global distribution and local properties of Gaussian primitives, which can easily lead to overregulation or underregulation, impairing the compactness and expressiveness of the scene representation.
[0007] In summary, existing technologies struggle to collaboratively achieve structure-aware primitive compaction, geometry-preserving primitive pruning, and complexity-adaptive primitive splitting in sparse views, which limits the final accuracy and overall robustness of new perspective synthesis. Summary of the Invention
[0008] To address the core issues of existing 3D Gaussian sputtering techniques in sparse view synthesis scenarios, such as overfitting due to incomplete observation information, leading to missing geometric fine structures, blurred boundaries, and floating artifacts in the reconstructed results, and to overcome the limitations of existing improved methods—such as the excessive primitives introduced by FSGS indiscriminate densification increasing geometric errors, the loss of key geometric primitives due to DropGaussian random discarding resulting in sharp boundary details, the inability of traditional fixed splitting scales to adapt to spatially varying geometric complexity, and the tendency of fixed dropout mechanisms to cause over-regularization or under-regularization, impairing the compactness and expressiveness of scene representation—this invention proposes a 3D Gaussian sputtering method guided by image texture structure, comprising the following steps:
[0009] Obtain a set of sparse input images of the same scene; A fine-grained structure-aware Gaussian sputtering model is constructed to train Gaussian points in the same scene, thereby obtaining trained Gaussian points in the same scene. The point cloud of the scene to be rendered is acquired and input into a fine structure-aware Gaussian sputtering model trained in the same scene to generate a high-quality new perspective image.
[0010] Furthermore: the fine-structure sensing Gaussian sputtering framework includes: Calculation module: Used to calculate multi-view gradient maps for a set of sparse input images of the same scene; Fine structure perception resampling module: Based on the multi-view gradient map transmitted by the calculation module, extract the core features representing the fine geometric structure of the scene, and adaptively insert new primitives in the detailed regions with significant gradients; Fine structure-aware splitting module: Based on the gradient map of the inserted new primitives transmitted by the fine structure-aware resampling module, a fine structure-aware splitting strategy is adopted to dynamically adjust the splitting scale of the existing primitives according to the local complexity, thereby generating a more refined new Gaussian representation; Multi-stage global-local adaptive module: Based on the Gaussian representation transmitted by the fine structure perception splitting module, a multi-stage global-local adaptive discarding mechanism is adopted. By analyzing the local opacity and global distribution density of primitives, the discarding probability is dynamically adjusted, thereby achieving an intelligent balance between retaining key structures and pruning redundant primitives.
[0011] Furthermore: the process of extracting core features representing the fine geometric structure of the scene based on the multi-view gradient map transmitted by the computing module, and adaptively inserting new primitives in regions with significant gradients and rich details, is as follows: S21: Candidate Gaussian pairs selected by the k-nearest neighbor criterion, with centers μ_S and μ_T respectively. S22: Connect candidate Gaussian pairs, uniformly sample 15 candidate points along the connecting line segment, and project them onto all K input views using a pinhole camera model; S23: Next, the gradient map of each view is calculated using the Scharr operator to capture high-frequency details, and the geometric detail magnitude of each candidate point in all views is obtained by bilinear interpolation, and then its average detail weight is calculated. S24: Based on the weight distribution, classify the line segment regions: if the maximum weight exceeds the threshold δ_r=0.1, it is considered a region with rich details; otherwise, if the weights are generally low, it is considered a flat region. For regions rich in detail, a discrete probability distribution is constructed based on normalized weights, and sampling is performed through the inverse cumulative distribution function to generate 10 Gaussian pairs, thereby achieving dense sampling in geometrically complex regions. For flat regions, a symmetrical continuous probability distribution is constructed with the endpoints of the line segments as the center, and six points are sampled from it to avoid generating redundant primitives in open areas and maintain the compactness of the scene representation.
[0012] Furthermore: the process of generating a more refined representation based on the gradient map of the inserted new primitive transmitted by the fine structure-aware resampling module, the fine structure-aware splitting strategy, and the dynamic adjustment of the splitting scale of the existing primitives according to the local complexity is as follows: A dynamic adjustment mechanism based on multi-view detail weights is introduced. By using a detail-aware scaling factor, the multi-view gradient weight w of the Gaussian candidate to be split, selected based on the fine structure-aware splitting strategy, is used as a proxy index of local texture complexity. The expression for the detail-aware scaling factor λ is as follows: λ = λ_w + (1 - λ_w)·(1 - w), Where λ_w is a hyperparameter set to 0.6; The scale of the newly generated Gaussian primitive is then updated to s' = λ·s, thus achieving a precise fit between the primitive granularity and the local geometric complexity.
[0013] Furthermore: The Gaussian representation transmitted based on the fine structure perception splitting module employs a multi-stage global-local adaptive discarding mechanism. By analyzing the local opacity and global distribution density of primitives, the discarding probability is dynamically adjusted, thereby achieving an intelligent balance between retaining key structures and pruning redundant primitives. The process is as follows: S31: The mechanism considers the local properties and global distribution of primitives in a coordinated manner to dynamically adjust the discard probability, thereby preserving key structures while suppressing overfitting. At the local level, each primitive is assigned a local discard weight (1 - o_i) that is negatively correlated with its opacity o_i. At the global level, monitor the global average opacity and normalized scene density. = N / V, and thus construct a dynamic global scaling factor β_t = (1 - ō(t))·(1 + (t)), which is updated in each iteration t to respond to the evolution of scene complexity; S32: A three-stage scheduling strategy is adopted for sampling: discarding is completely disabled in the first 600 iterations to allow the scene to initially stabilize and take shape; Between 600 and 2000 iterations, the dropout rate increased linearly; After 2000 iterations, the scheduler is reset and linear promotion begins again, this time by a function. _t is specifically defined; Finally, the probability of the i-th Gaussian being dropped at iteration t is determined by the following formulas: γ_t(g_i) = β_t · (1 - o_i) · _t.
[0014] Where: βt: global scaling factor; 1 oi: Local opacity factor t: scheduling factor during training phase, γt(gi): final output gi: the i-th 3D Gaussian unit.
[0015] Furthermore: the loss function of the fine structure-aware Gaussian sputtering model is defined as: _total = λ_1· _1 + λ_2· _D-SSIM + λ_3· _Corr in: _1 represents pixel-level L1 loss. _D-SSIM is the structural similarity loss. _Corr represents the depth consistency loss of pseudo-views.
[0016] A three-dimensional Gaussian sputtering apparatus guided by image texture structure, comprising: Acquisition module: used to retrieve a set of sparse input images of the same scene; Module: Used to build a fine structure-aware Gaussian sputtering model, and to train Gaussian points in the same scene to obtain trained Gaussian points in the same scene; The generation module is used to acquire and input the point cloud of the scene to be rendered, and input it into a fine structure-aware Gaussian sputtering model trained in the same scene to generate high-quality new perspective images. A readable storage medium that stores a program module, which, when executed in a processor, can implement the method as described in any one of the claims.
[0017] This invention also provides a three-dimensional Gaussian sputtering method based on image texture structure guidance, which brings significant and multifaceted benefits: In terms of synthesis accuracy, extensive experiments on standard datasets such as LLFF, Mip-NeRF360, and Blender show that our method outperforms existing mainstream methods in multiple quantitative metrics.
[0018] On the highly challenging LLFF dataset with a 3-view setting, our method achieves a PSNR of 21.03, which is 2.9% higher than FSGS and 1.3% higher than DropGaussian. On the Mip-NeRF360 dataset with a 12-view setting, the PSNR reaches 19.93, which is 6.0% higher than FSGS. Furthermore, our method also outperforms the current state-of-the-art methods in both SSIM and LPIPS metrics.
[0019] In preserving fine geometric structures, the combined effect of the FSAR and FSAS strategies enables Gaussian primitives to accurately align sharp boundaries and complex textures in the scene, significantly increasing the proportion of small Gaussian kernels with scales between 0.2 and 2.0, fundamentally alleviating the edge blurring and structural loss problems caused by traditional methods. Regarding overfitting suppression, the GLOD-D mechanism adaptively prunes a large number of low-contribution redundant primitives, especially those low-opacity primitives that are prone to causing occlusion artifacts when accumulated.
[0020] Compared to the DropGaussian baseline, this method significantly reduces the spatial density of such redundant primitives, effectively controlling overfitting while perfectly preserving key primitives crucial for geometric integrity. In terms of efficiency and compactness, this framework fully inherits the advantages of 3DGS's efficient differentiable rasterization, converging in just 10,000 iterations on a single NVIDIA RTX 3090 GPU. Furthermore, the resulting Gaussian primitive distribution is more reasonable and compact, avoiding unnecessary computational overhead and achieving a balance between high-quality reconstruction and real-time rendering efficiency. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a framework diagram of the method in this application; Figure 2 Under extremely sparse conditions using only three input views, the method of this invention is compared with the new perspective synthesis visuals of FSGS and DropGaussian baselines. (a) presents the rendering results of the two new perspectives and their corresponding error maps side by side, while (b) shows the overall distribution of all Gaussian metacenters generated by each method in three-dimensional space. Figure 3 It provides a direct comparison of the different sampling behaviors in flat and detailed regions, including (a) the difference in sampling point distribution in detailed regions and (b) the difference in sampling point distribution in flat regions. Figure 4 The effectiveness of the fine structure perception splitting strategy was confirmed by combining quantitative analysis and spatial visualization. Histogram (a) compared the size distribution of all Gaussian kernels in the scene with and without the FSAS strategy, while (b) projected the actual positions of small Gaussian kernels within a specific size range in the 3D scene. Figure 5 It directly compares the state of the 3D Gaussian scene at different iteration stages during the training process with the DropGaussian baseline. Figure 6 It showcases new perspective synthetic rendering results on the LLFF standard real-world scene dataset; Figure 7 This showcases new perspective synthetic rendering results on the Mip-NeRF360 standard real-world scene dataset; Figure 8 Adhering to the core of module ablation in the LLFF dataset 3 perspective, the experimental results of core module ablation under the LLFF dataset 3 perspective setting: comparison of the effects of removing densification, discarding, and segmenting modules. Figure 9 Adhering to the core of component and policy ablation under the LLFF dataset 3 perspective, the experimental results of component and policy ablation under the LLFF dataset 3 perspective setting are as follows: validity verification of FSAR, FSAS and GLOD-D, where (a) basic version: no FSAR, no Dropout (neither the original nor our proposed version), no FSAS (equivalent to setting one parameter); (b) final version: with FSAR, multi-stage our proposed Dropout policy, and FSAS; (c) no FSAR, replacing FSAR with center interpolation; (d) no FSAS, equivalent to setting one parameter; (e) using the original multi-stage drop policy; (f) using the original single-stage drop policy; (g) using the single-stage drop policy our proposed policy; (h) original image. Detailed Implementation
[0023] It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] Figure 1 This is a framework diagram of the method in this application; A three-dimensional Gaussian sputtering method based on image texture structure guidance includes the following steps: Obtain a set of sparse input images of the same scene; A fine-grained structure-aware Gaussian sputtering model is constructed to train Gaussian points in the same scene, thereby obtaining trained Gaussian points in the same scene. The point cloud of the scene to be rendered is acquired and input into a fine structure-aware Gaussian sputtering model trained in the same scene to generate a high-quality new perspective image.
[0026] Furthermore: the fine-structure sensing Gaussian sputtering framework includes: Calculation module: Used to calculate multi-view gradient maps for a set of sparse input images of the same scene; Fine structure perception resampling module: Based on the multi-view gradient map transmitted by the calculation module, extract the core features representing the fine geometric structure of the scene, and adaptively insert new primitives in the detailed regions with significant gradients; Fine structure-aware splitting module: Based on the gradient map of the inserted new primitives transmitted by the fine structure-aware resampling module, a fine structure-aware splitting strategy is adopted to dynamically adjust the splitting scale of the existing primitives according to the local complexity, thereby generating a more refined new Gaussian representation; Multi-stage global-local adaptive module: Based on the Gaussian representation transmitted by the fine structure perception splitting module, a multi-stage global-local adaptive discarding mechanism is adopted. By analyzing the local opacity and global distribution density of primitives, the discarding probability is dynamically adjusted, thereby achieving an intelligent balance between retaining key structures and pruning redundant primitives.
[0027] The core of the model lies in two aspects: First, through fine-structure-aware resampling and fine-structure-aware splitting strategies, geometric detail maps extracted from multi-view images are used to guide Gaussian primitives to adaptively densify and scale in texture-rich regions, enabling primitives to accurately align with the geometric boundaries of the scene. Second, a global-local opacity and density-driven dropout mechanism is introduced, dynamically adjusting the dropout probability based on the local opacity and global distribution density of the primitives. Redundant primitives are adaptively pruned in the later stages of training, effectively suppressing overfitting.
[0028] Furthermore, this fine structure-aware resampling process differs from traditional simple interpolation densification by utilizing geometric cues extracted from multi-view images to achieve adaptive resampling.
[0029] The process of extracting core features representing the fine geometric structure of the scene from the multi-view gradient map transmitted by the computing module, and adaptively inserting new primitives in regions with significant gradients and rich details, is as follows: S21: Candidate Gaussian pairs selected by the k-nearest neighbor criterion, with centers μ_S and μ_T respectively. S22: Connect candidate Gaussian pairs, uniformly sample 15 candidate points along the connecting line segment, and project them onto all K input views using a pinhole camera model; S23: Next, the gradient map of each view is calculated using the Scharr operator to capture high-frequency details, and the geometric detail magnitude of each candidate point in all views is obtained by bilinear interpolation, and then its average detail weight w[i] is calculated. S24: Based on the weight distribution, classify the line segment regions: if the maximum weight exceeds the threshold δ_r=0.1, it is considered a region with rich details; otherwise, if the weights are generally low, it is considered a flat region. For regions rich in detail, a discrete probability distribution is constructed based on normalized weights, and sampling is performed through the inverse cumulative distribution function to generate 10 Gaussian pairs, thereby achieving dense sampling in geometrically complex regions. For flat regions, a symmetrical continuous probability distribution (standard deviation σ=0.1) is constructed concentrated near the endpoints of line segments, and 6 points are sampled from it to avoid generating redundant primitives in open areas and maintain the compactness of the scene representation.
[0030] Furthermore: the process of generating a more refined representation based on the gradient map of the inserted new primitive transmitted by the fine structure-aware resampling module, the fine structure-aware splitting strategy, and the dynamic adjustment of the splitting scale of the existing primitives according to the local complexity is as follows: A dynamic adjustment mechanism based on multi-view detail weights is introduced. By using a detail-aware scaling factor, the multi-view gradient weight w of the Gaussian candidate to be split, selected based on the fine structure-aware splitting strategy, is used as a proxy index of local texture complexity. The expression for the detail-aware scaling factor λ is as follows: λ = λ_w + (1 - λ_w)·(1 - w), Where λ_w is a hyperparameter set to 0.6; this formula ensures that a smaller λ is produced in regions with rich detail (high w), thus generating new primitives with finer scale; while a larger λ is produced in flat regions (low w) to maintain the relative size of the primitives.
[0031] The scale of the newly generated Gaussian primitive is then updated to s' = λ·s, thus achieving a precise fit between the primitive granularity and the local geometric complexity.
[0032] Furthermore: The Gaussian representation transmitted based on the fine structure perception splitting module employs a multi-stage global-local adaptive discarding mechanism. By analyzing the local opacity and global distribution density of primitives, the discarding probability is dynamically adjusted, thereby achieving an intelligent balance between retaining key structures and pruning redundant primitives. The process is as follows: S31: The mechanism considers the local properties and global distribution of primitives in a coordinated manner to dynamically adjust the discard probability, thereby preserving key structures while suppressing overfitting. At the local level, each primitive is assigned a local discard weight (1 - o_i) that is negatively correlated with its opacity o_i. At the global level, to control the excessive growth of primitive numbers and overfitting that may result from interpolation and splitting processes, the global average opacity and normalized scene density are monitored. = N / V, and thus construct a dynamic global scaling factor β_t = (1 - ō(t))·(1 + (t)), which is updated in each iteration t to respond to the evolution of scene complexity; S32: To ensure optimization stability, a three-stage scheduling strategy is adopted for sampling: discarding is completely disabled in the initial 600 iterations to allow the scene to initially stabilize and take shape; Between 600 and 2000 iterations, the dropout rate increased linearly; After 2000 iterations, the scheduler is reset and linear promotion begins again, this time by a function. _t is specifically defined; the function is described. The expression for _t is as follows:
[0033] Finally, the probability of the i-th Gaussian being dropped at iteration t is determined by the following formulas: γ_t(g_i) = β_t · (1 - o_i) · _t.
[0034] Where: βt: global scaling factor; 1 oi: Local opacity factor t: scheduling factor during training phase, γt(gi): final output gi: the i-th 3D Gaussian unit.
[0035] A multi-stage global-local adaptive discarding mechanism considers both the local properties and global distribution of primitives to dynamically adjust the discarding probability, thereby preserving key structures while suppressing overfitting. At the local level, we observed a large number of primitives with extremely low opacity prevalent in the scene. Individually, these primitives contribute very little to the final rendering, but their accumulation leads to unnecessary occlusion and rendering artifacts.
[0036] Furthermore: the loss function of the fine structure-aware Gaussian sputtering model is defined as: _total = λ_1· _1+λ_2· _D-SSIM +λ_3· _Corr in: _1 represents pixel-level L1 loss. _D-SSIM is the structural similarity loss. _Corr represents the pseudoview depth consistency loss. For each generated pseudoview, we render its depth map from the current Gaussian representation and compare it with the depth predicted by a monocular depth estimator. The difference between the two provides a strong geometrically aware constraint, effectively stabilizing scene reconstruction and significantly improving geometric consistency across different unseen viewpoints.
[0037] A three-dimensional Gaussian sputtering apparatus guided by image texture structure, comprising: Acquisition module: used to retrieve a set of sparse input images of the same scene; Module: Used to build a fine structure-aware Gaussian sputtering model, and to train Gaussian points in the same scene to obtain trained Gaussian points in the same scene; The generation module is used to acquire and input the point cloud of the scene to be rendered, and input it into a fine structure-aware Gaussian sputtering model trained in the same scene to generate high-quality new perspective images. A readable storage medium that stores a program module, which, when executed in a processor, can implement the method as described in any one of the claims.
[0038] Example 1 Figure 2 Under extremely sparse conditions using only three input views, the method of this invention is compared with the new perspective synthesis visuals of FSGS and DropGaussian baselines. (a) presents the rendering results of the two new perspectives and their corresponding error maps side by side, while (b) shows the overall distribution of all Gaussian metacenters generated by each method in three-dimensional space. Figure 2The diagram illustrates a comparison of the proposed method with the baselines of FSGS and DropGaussian under extremely sparse conditions using only three input views. (a) juxtaposes the rendering results of the two new perspectives and their corresponding error maps, where the error maps clearly reveal the pixel-level deviations between the outputs of each method and the real images in the form of heatmaps. Key local regions are highlighted with dashed boxes, and the corresponding Gaussian primitive spatial distribution is visualized next to them. Notably, this method additionally provides a geometric detail map extracted from the input image, which serves as the fundamental guide for all subsequent optimization steps. (b) macroscopically displays the overall distribution of all Gaussian primitive centers generated by each method in three-dimensional space. The comparison shows that while the primitive distribution of the FSGS method is extensive, it is redundant; the distribution of DropGaussian is sparse and uneven due to random discarding; while the primitive distribution of this method closely and precisely fits the geometric boundaries of the scene, spatially confirming the basis for the sharper edges and more complete details in the upper part of the rendered image.
[0039] Figure 3 This is a direct comparison of the different sampling behaviors in flat areas and areas with rich details, where (a) shows the difference in the distribution of sampling points in areas with rich details, and (b) shows the difference in the distribution of sampling points in flat areas. Figure 3 This diagram is specifically designed to explain the working principle of the fine structure-aware resampling strategy. It visually compares the different sampling behaviors in flat and detail-rich regions. In flat regions, due to weak image gradient information, the probability distribution of new sampled points is concentrated at both ends of the line segment connecting the two original Gaussian centers, thus avoiding the introduction of uncontributing primitives in empty areas. Conversely, in detail-rich regions, strong image gradients guide the reweighting of the sampling probability distribution, causing newly sampled points to significantly cluster near the true geometric boundaries of the image. This mechanism ensures that additional computational resources are precisely allocated to key locations that characterize high-frequency details of the scene.
[0040] Figure 4 The effectiveness of the fine structure-aware splitting strategy was confirmed through a combination of quantitative analysis and spatial visualization. Histogram (a) compares the size distribution of all Gaussian kernels in the scene with and without the FSAS strategy. The data clearly show that the FSAS strategy significantly increases the proportion of small kernels in the overall distribution. (b) projects the actual positions of small Gaussian kernels within a specific size range in the 3D scene. The visualization results show that these finer primitives are not randomly scattered, but highly concentrated in areas with complex textures and intricate structures, such as the facial features of a statue and the surface of elaborate decorations, thus directly demonstrating the accuracy of the strategy in guiding resource allocation.
[0041] Figure 5 This paper demonstrates the superior performance of the global-local adaptive discarding mechanism in removing redundant primitives by directly comparing the 3D Gaussian scene states at different iteration stages during training with the DropGaussian baseline. The figure uses red dashed boxes to mark typical high-opacity primitive clusters in the scene. The comparison shows that during the optimization process of the DropGaussian baseline, a large number of low-opacity redundant primitives accumulate, forming a fog-like occlusion. In contrast, in the optimization trajectory of this method, the primitive density in the same areas is significantly reduced, resulting in a clearer and more compact scene representation. This visual comparison strongly demonstrates that the discarding mechanism of this method can effectively identify and remove low-importance primitives that contribute little to rendering, thereby significantly improving the efficiency and purity of scene representation.
[0042] Figure 6 It showcases new perspective synthetic rendering results on the LLFF standard real-world scene dataset; Figure 7 This showcases new perspective synthetic rendering results on the Mip-NeRF360 standard real-world scene dataset; These images are not single views, but rather presented in a multi-row, multi-column format, comparing our method with baseline methods under different training view counts (e.g., 3 / 6 / 9 views for LLFF, 12 / 24 views for Mip-NeRF360). Each sub-image is clearly labeled with its corresponding PSNR value. In the images, complex textured, sharp-edged local areas are generally highlighted with red dashed boxes and magnified. The comparison clearly shows that the images generated by our method have sharper edges, richer texture details, and fewer blurs or artifacts. This advantage is particularly evident in high-frequency information areas such as corners, leaves, and sculpted faces. In contrast, the results from methods like FSGS often appear overly smooth or structurally distorted, while DropGaussian may lose crucial details due to random discarding. These direct visual comparisons strongly demonstrate our method's ability to reconstruct and preserve the fine geometric structure of a scene under sparse input, perfectly aligning with the description of improved synthesis accuracy and preservation of fine geometry in the "beneficial effects" section.
[0043] Figure 8 and Figure 9The focus is on the visualization analysis of ablation research, which is key to demonstrating the necessity of each core technology module. The first image, through a side-by-side display of a set of rendering results, visually demonstrates the significant decrease in synthesis quality when Fine Structure Aware Resampling (FSAR), Dropout (GLOD-D), or Splitting (FSAS) modules are removed sequentially from the complete model. For example, removing FSAR results in inaccurate primitive insertion in detailed regions, leading to the loss of local structure; removing GLOD-D causes a large accumulation of low-opacity redundant primitives in the scene, resulting in blurred rendering or cloud-like artifacts, which directly corresponds to overfitting; and removing FSAS, due to the lack of adaptive splitting, makes the primitive granularity unable to match the local complexity, resulting in coarse overall detail. The second image further demonstrates the performance improvement brought by these customized designs compared to the naive approach through more detailed comparative experiments, such as replacing FSAR with simple center interpolation, or setting the scaling factor of FSAS to a fixed value (degenerating to the original splitting). These ablation visualization results irrefutably demonstrate that each of the three strategies—FSAR, FSAS, and GLOD-D—is indispensable, and their combined efforts are what achieve the ultimate superior performance.
[0044] Table 1. Quantitative Comparison Results of Sparse Perspective Synthesis Methods for the LLFF Dataset
[0045] Table 2. Quantitative Comparison Results of Sparse Perspective Synthesis Methods on the Blender Dataset
[0046] Table 3. Quantitative Comparison Results of New Perspective Synthesis Methods for Sparse Perspectives on the Mip-NeRF360 Dataset
[0047] Tables 1, 2, and 3 present a comprehensive performance comparison of our method with a range of state-of-the-art NeRF methods (such as Mip-NeRF and RegNeRF) and 3DGS methods (such as FSGS, DropGaussian, and DNGaussian) on the LLFF, Mip-NeRF360, and Blender datasets, respectively. These tables provide detailed values for PSNR, SSIM, and LPIPS. The data clearly show that, regardless of the sparsity setting, our method achieves best or highly competitive results on most metrics. For example, key conclusions such as a 2.9% improvement in PSNR (compared to FSGS) and a 1.3% improvement in PSNR (compared to DropGaussian) on LLFF 3 views, and a 6.0% improvement on PSNR on Mip-NeRF360 12 views, are derived from these tables. They objectively and statistically confirm our method's superior overall synthesis accuracy.
[0048] Table 4 Ablation experiment results of the stepwise fusion of each core component of FSA-GS
[0049] Table 5 Ablation experiment results of gradient region partitioning and point insertion parameter selection in FSAR strategy
[0050] Table 6 Comparison of geometric detail extraction effects of different frequency thresholds in Fourier transform and the Scharr operator
[0051] Table 7. Ablation experimental results for different scaling factors λ_w in the FSAS strategy.
[0052] Tables 5, 6, and 7 delve into the parameter and module effectiveness analysis within the method. Table 1 provides a core quantitative summary of ablation studies, demonstrating the gradual improvement in PSNR and other metrics as FSAR, FSAS, and GLOD-D modules are added, ultimately achieving the optimal performance of the complete model. The following three tables analyze the impact of different gradient operators (e.g., Scharr vs. Fourier), different parameters in the FSAR strategy (e.g., the number of sampling points m1, m2, m3), and the scaling factor λ_w in the FSAS strategy on performance. For example, Table 2 demonstrates that using the Scharr operator yields a higher PSNR than high-frequency extraction based on Fourier transform, thus validating the rationality of the chosen technical details; Table 4 shows that λ_w=0.6 achieves the best balance and is therefore adopted as the default parameter. These parameter studies demonstrate the rigor of the method design and provide reproducible optimal configurations.
[0053] In the specific implementation of this invention, the following experimental configuration and procedure are followed to verify and reproduce its effects: For experimental setup, three widely used benchmark datasets—LLFF, Mip-NeRF360, and Blender—were selected for comprehensive evaluation. To ensure fair comparison with existing work, the experimental setup strictly followed DropGaussian: LLFF used 3, 6, and 9 training views respectively; Mip-NeRF360 used 12 and 24 views; and Blender used 8 views. The input image resolution was downsampled, with LLFF and Mip-NeRF360 downsampled by 8 times, and Blender downsampled by 2 times. Evaluation employed three authoritative metrics: PSNR, SSIM, and LPIPS, comprehensively measuring the synthesis quality of the new perspectives from the perspectives of pixel error, structural consistency, and perceptual similarity, respectively. Scene initialization involved obtaining initial point clouds from the provided sparse views using Structure for Motion Reconstruction (SfM), and standard 3DGS optimization was used in the first 600 iterations of training to ensure initial stability.
[0054] The training process for the fine structure-aware Gaussian sputtering model is as follows: The FSAR and FSAS policies are activated after the initialization phase (600 iterations) and then executed every 100 iterations thereafter. The FSAR parameters are set to m1=15, m2=10, m3=6, and the detail threshold δ_r=0.1. The detail-aware scaling factor λ_w in the FSAS is fixed at 0.6. The GLOD-D mechanism is initiated at the 600th iteration and restarted as designed at the 2000th iteration. The gradient threshold for the densification operation is set to 5×10^-4, consistent with FSGS to ensure fair comparison. The entire training process consists of 10,000 iterations, and all experiments are performed on a single NVIDIA RTX 3090 GPU using the PyTorch framework.
[0055] The implementation results were fully verified through both quantitative and qualitative methods. Quantitatively, in the LLFF 3-view scenario, this method achieved excellent performance with PSNR of 21.03, SSIM of 0.730, and LPIPS of 0.192, comprehensively surpassing existing technologies. In the 12-view and 24-view settings of the Mip-NeRF360, this method also achieved optimal or highly competitive results.
[0056] Qualitatively, such as Figure 6 , 7 As shown, this method can accurately restore high-frequency details such as edges and textures in the scene, effectively eliminating floating artifacts and blurring issues, and exhibiting stronger robustness even in complex scenes. The system's ablation experiments further confirm the significant contribution of each module—FSAR, FSAS, and GLOD-D—to improving the final reconstruction quality; the absence or substitution of any one component will lead to a significant performance degradation.
[0057] Figure 8 Adhering to the core of module ablation in the LLFF dataset 3 perspective, the experimental results of core module ablation under the LLFF dataset 3 perspective setting: comparison of the effects of removing densification, discarding, and segmenting modules. Figure 9Adhering to the core of component and policy ablation under the LLFF dataset 3 perspective, the experimental results of component and policy ablation under the LLFF dataset 3 perspective setting are as follows: validity verification of FSAR, FSAS and GLOD-D, where (a) basic version: no FSAR, no Dropout (neither the original nor our proposed version), no FSAS (equivalent to setting one parameter); (b) final version: with FSAR, multi-stage our proposed Dropout policy, and FSAS; (c) no FSAR, replacing FSAR with center interpolation; (d) no FSAS, equivalent to setting one parameter; (e) using the original multi-stage drop policy; (f) using the original single-stage drop policy; (g) using the single-stage drop policy our proposed policy; (h) original image.
[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A three-dimensional Gaussian sputtering method based on image texture structure guidance, characterized in that: Includes the following steps: Obtain a set of sparse input images of the same scene; A fine-grained structure-aware Gaussian sputtering model is constructed to train Gaussian points in the same scene, thereby obtaining trained Gaussian points in the same scene. The point cloud of the scene to be rendered is acquired and input into a fine structure-aware Gaussian sputtering model trained in the same scene to generate a high-quality new perspective image.
2. The three-dimensional Gaussian sputtering method based on image texture structure guidance according to claim 1, characterized in that: The fine-structure sensing Gaussian sputtering framework includes: Calculation module: Used to calculate multi-view gradient maps for a set of sparse input images of the same scene; Fine structure perception resampling module: Based on the multi-view gradient map transmitted by the calculation module, extract the core features representing the fine geometric structure of the scene, and adaptively insert new primitives in the detailed regions with significant gradients; Fine structure-aware splitting module: Based on the gradient map of the inserted new primitives transmitted by the fine structure-aware resampling module, a fine structure-aware splitting strategy is adopted to dynamically adjust the splitting scale of the existing primitives according to the local complexity, thereby generating a more refined new Gaussian representation; Multi-stage global-local adaptive module: Based on the Gaussian representation transmitted by the fine structure perception splitting module, a multi-stage global-local adaptive discarding mechanism is adopted. By analyzing the local opacity and global distribution density of primitives, the discarding probability is dynamically adjusted, thereby achieving an intelligent balance between retaining key structures and pruning redundant primitives.
3. The three-dimensional Gaussian sputtering method based on image texture structure guidance according to claim 2, characterized in that: The process of extracting core features representing the fine geometric structure of the scene from the multi-view gradient map transmitted by the computing module, and adaptively inserting new primitives in regions with significant gradients and rich details, is as follows: S21: Candidate Gaussian pairs selected by the k-nearest neighbor criterion, with centers μ_S and μ_T respectively. S22: Connect candidate Gaussian pairs, uniformly sample 15 candidate points along the connecting line segment, and project them onto all K input views using a pinhole camera model; S23: Next, the gradient map of each view is calculated using the Scharr operator to capture high-frequency details, and the geometric detail magnitude of each candidate point in all views is obtained by bilinear interpolation, and then its average detail weight w[i] is calculated. S24: Based on the weight distribution, classify the line segment regions: if the maximum weight exceeds the threshold δ_r=0.1, it is considered a region with rich details; otherwise, if the weights are generally low, it is considered a flat region. For regions rich in detail, a discrete probability distribution is constructed based on normalized weights, and sampling is performed through the inverse cumulative distribution function to generate 10 Gaussian pairs, thereby achieving dense sampling in geometrically complex regions. For flat regions, a symmetrical continuous probability distribution is constructed with the endpoints of the line segments as the center, and six points are sampled from it to avoid generating redundant primitives in open areas and maintain the compactness of the scene representation.
4. The three-dimensional Gaussian sputtering method based on image texture structure guidance according to claim 2, characterized in that: The process of generating a more refined representation based on the gradient map of the newly inserted primitives transmitted by the fine structure-aware resampling module, and the fine structure-aware splitting strategy that dynamically adjusts the splitting scale of existing primitives according to local complexity, is as follows: A dynamic adjustment mechanism based on multi-view detail weights is introduced. By using a detail-aware scaling factor, the multi-view gradient weight w of the Gaussian candidate to be split, selected based on the fine structure-aware splitting strategy, is used as a proxy index of local texture complexity. The expression for the detail-aware scaling factor λ is as follows: λ = λ_w + (1 - λ_w)·(1 - w), Where λ_w is a hyperparameter set to 0.6; The scale of the newly generated Gaussian primitive is then updated to s' = λ·s, thus achieving a precise fit between the primitive granularity and the local geometric complexity.
5. The three-dimensional Gaussian sputtering method based on image texture structure guidance according to claim 2, characterized in that: The Gaussian representation transmitted by the fine-structure perception splitting module employs a multi-stage global-local adaptive discarding mechanism. By analyzing the local opacity and global distribution density of primitives, the discarding probability is dynamically adjusted to achieve an intelligent balance between retaining key structures and pruning redundant primitives. The process is as follows: S31: The mechanism considers the local properties and global distribution of primitives in a coordinated manner to dynamically adjust the discard probability, thereby preserving key structures while suppressing overfitting. At the local level, each primitive is assigned a local discard weight (1 - o_i) that is negatively correlated with its opacity o_i. At the global level, monitor the global average opacity and normalized scene density. = N / V, and thus construct a dynamic global scaling factor β_t = (1 - ō(t))·(1 + (t)), which is updated in each iteration t to respond to the evolution of scene complexity; S32: A three-stage scheduling strategy is adopted for sampling: discarding is completely disabled in the first 600 iterations to allow the scene to initially stabilize and take shape; Between 600 and 2000 iterations, the dropout rate increased linearly; After 2000 iterations, the scheduler is reset and linear promotion begins again, this time by a function. _t is specifically defined; Finally, the probability of the i-th Gaussian being dropped at iteration t is determined by the following formulas: γ_t(g_i) = β_t · (1 - o_i) · _t。 in βt: Global scaling factor; 1 oi: Local opacity factor t: scheduling factor during training phase, γt(gi): final output gi: the i-th 3D Gaussian unit.
6. The three-dimensional Gaussian sputtering method based on image texture structure guidance according to claim 2, characterized in that: The loss function for the fine structure-aware Gaussian sputtering model is defined as: _total = λ_1· _1 + λ_2· _D-SSIM + λ_3· _Corr in: _1 represents pixel-level L1 loss. _D-SSIM is the structural similarity loss. _Corr represents the depth consistency loss of pseudo-views.
7. A three-dimensional Gaussian sputtering apparatus guided by image texture structure, characterized in that: include: Acquisition module: used to retrieve a set of sparse input images of the same scene; Module: Used to build a fine structure-aware Gaussian sputtering model, and to train Gaussian points in the same scene to obtain trained Gaussian points in the same scene; The generation module is used to acquire and input the point cloud of the scene to be rendered, and input it into the fine structure-aware Gaussian sputtering model trained in the same scene to generate a high-quality new perspective image.
8. A readable storage medium storing a program module, characterized in that, The program module, when run in a processor, can implement the method as described in any one of claims 1-6.