A sparse view three-dimensional reconstruction method, device, storage medium and electronic equipment
Patent Information
- Application Number
- CN202611011814.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
AI Technical Summary
这种盲目的采样策略无法有效识别当前重建结果中真正欠缺观测的区域,往往对已充分覆盖的区域生成冗余监督信号,而对亟需补充信息的欠覆盖区域关注不足
本发明通过引入覆盖感知的主动探索策略,能够动态识别并优先探索当前表示中的欠覆盖区域。通过将伪监督生成过程解耦为互补的渲染优化分支与几何引导新视图生成分支,有效克服了单一分支范式在几何一致性与局部细节保真度上不可兼得的困境。设计特征门控融合网络实现双分支信号的空间自适应整合,确保了融合伪监督同时具备高保真局部纹理和全局一致的几何结构,从而提升最终的重建质量。
Smart Images

Figure CN122820979A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional vision reconstruction technology, and more particularly to a sparse vision method. Figure 3 Reconstruction methods, devices, storage media, and electronic equipment. Background Technology
[0002] 3D Gaussian splashing, as a 3D scene representation method based on explicit Gaussian primitives, can achieve real-time and high-fidelity synthesis of new perspectives. However, under conditions of extremely sparse input views, the severe lack of multi-view geometric constraints will lead to overfitting of 3D Gaussian to the limited training perspectives, resulting in a large number of floating artifacts in free space and structural holes in unobserved areas. When rendering from new perspectives with significant deviations from the training pose, these artifact and hole problems will be further exacerbated.
[0003] To alleviate the aforementioned problems, existing methods attempt to introduce diffusion model priors as auxiliary supervision, guiding 3D optimization by generating pseudo-ground images for unobserved regions. However, existing paradigms still have fundamental limitations in their design. One type of method independently applies a diffusion model to each frame for denoising or restoration, using the result as a pseudo-supervision signal. While this frame-by-frame processing approach can achieve effective local fine-grained restoration in mildly degraded scenes, it inherently ignores the geometric consistency across views. In large baseline scenes, the rendered image obtained from sparse reconstruction is often severely damaged by a large number of floating objects, and diffusion-based restoration tends to generate illogical false structures rather than restoring the true underlying geometry. Another type of method formalizes sparse view reconstruction as a video interpolation or completion task, implicitly applying multi-view congruence by utilizing the inter-frame coherence of the video diffusion model. Figure 1 Consistency constraints. However, these models are primarily pre-trained on dynamic video data, rather than on multi-view static scenes. Therefore, they tend to interpret large camera movements as object deformation, leading to geometric drift and motion blur even in static scenes. Furthermore, the inference computation cost of video diffusion models is extremely high, contrasting sharply with the high efficiency of 3D Gaussian splashing.
[0004] Furthermore, both of these paradigms generally rely on predefined fixed trajectories to sample and generate pseudo-supervisory viewpoints. This blind sampling strategy cannot effectively identify regions that are truly lacking in observation in the current reconstruction results. It often generates redundant supervisory signals for areas that are already fully covered, while paying insufficient attention to under-covered areas that urgently need supplementary information.
[0005] It is evident that existing diffusion-prior-guided sparse view reconstruction methods suffer from the following core contradiction: the two mainstream paradigms are complementary in their respective shortcomings, and both lack the ability to actively identify regions that urgently require supervision. Summary of the Invention
[0006] This invention provides a sparse view Figure 3 The invention relates to methods, apparatuses, storage media, and electronic devices for dimensional reconstruction, aiming to solve technical problems in existing technologies such as geometric inconsistencies, loss of detail, blind sampling, and high computational overhead.
[0007] The technical solution adopted in this invention is as follows: Firstly, the present invention provides a sparse view Figure 3 3D reconstruction methods include: Based on the sparse input view and the corresponding camera pose, an initial 3D Gaussian splash representation is constructed. Based on a predefined smooth reference trajectory, high information gain viewpoints are selected through coverage-aware filtering and dynamic backoff mechanisms. For each selected candidate viewpoint, the rendering optimization branch and the new view generation branch are executed in parallel to generate complementary local fine-grained pseudo-supervision and global consistent pseudo-supervision signals; Using a trained feature-gated fusion network, the local fine pseudo-supervision and the global consistent pseudo-supervision signals are spatially adaptively fused to generate a fused pseudo-truth image and its corresponding spatial confidence map. Based on the fused pseudo-true value image and spatial confidence map, the 3D Gaussian splash representation is iteratively optimized until the preset conditions are met.
[0008] Furthermore, the construction of the initial 3D Gaussian splash representation based on the sparse input view and the corresponding camera pose includes: The initial point cloud and camera pose are estimated from the sparse input view using a motion recovery structure algorithm or a pre-trained geometric model; A set of three-dimensional Gaussian elements is initialized based on the initial point cloud. Each Gaussian element is parameterized by mean position, covariance matrix, color, and opacity. The three-dimensional Gaussian primitives are rendered using a differentiable rasterizer to obtain a rendered image corresponding to the sparse input view. The reconstruction loss between the rendered image and the sparse input view is calculated. The parameters of the three-dimensional Gaussian primitives are iteratively optimized through backpropagation to obtain an initial three-dimensional Gaussian splash representation.
[0009] Furthermore, the selection of high information gain viewpoints based on predefined smooth reference trajectories through coverage-aware filtering and dynamic backoff mechanisms includes: The space is divided into multiple spatial bins along a smooth reference trajectory by arc length, and the anchor point camera pose of each bin is extracted. Within each spatial sub-box, rotation-based and translation-based perturbations are applied to the pose of the anchor point camera. Both types of perturbations are normalized using scene scale parameters to generate the candidate viewpoint set. For each candidate viewpoint, render the current 3D Gaussian splash representation to obtain the rendered image and its corresponding transparency map; The transparency map is thresholded and morphologically filtered to obtain a binarized hole mask, and then the hole ratio, which represents the degree of geometric loss of the viewpoint, is calculated. Based on the hole ratio, candidate viewpoints are divided into high-value, medium-value, and low-value levels, and a dynamic backoff mechanism is triggered according to the coverage ratio of high-value candidate viewpoints in each bin: when the coverage ratio is lower than the backoff threshold, the perturbation range is expanded and candidate viewpoints are resampled until the coverage ratio requirement is met or the maximum number of retries is reached. After the final round of candidate sampling is completed, all candidate viewpoints at the high-value and medium-value levels are retained. If the total number of retained candidate viewpoints exceeds the preset global capacity limit, then, provided that each bin has at least one candidate viewpoint, the candidate viewpoint with the lowest void ratio is removed from the bin with the most candidate viewpoints.
[0010] Furthermore, for each selected candidate viewpoint, the rendering optimization branch and the new view generation branch are executed in parallel to generate complementary local fine-grained pseudo-supervision and globally consistent pseudo-supervision signals, including: In the rendering optimization branch, the training view closest to the spatial position of the currently selected viewpoint is retrieved as the reference frame. The reference-guided diffusion model is used to remove floating artifacts and ghost images in the rendered image under the currently selected viewpoint in a single-step denoising manner, generating a local fine pseudo-supervision signal. In the new view generation branch, a geometry-guided multi-view diffusion model is used, which combines sparse input views and corresponding depth geometric priors. By distorting the features of the reference view to the target viewpoint, structural guidance is provided to directly synthesize a new view with cross-view geometric consistency, which serves as a globally consistent pseudo-supervision signal.
[0011] Furthermore, the step of using a trained feature-gated fusion network to perform spatial adaptive fusion of the local fine-grained pseudo-supervision and the globally consistent pseudo-supervision signals to generate a fused pseudo-ground value image and its corresponding spatial confidence map includes: A feature-gated fusion network is constructed, which includes a shared feature encoder, a spatial adaptive fusion gate, and a residual decoder with uncertainty estimation. The shared feature encoder maps the current viewpoint's rendered image, the local fine-grained pseudo-supervised image, and the globally consistent pseudo-supervised image to a unified feature space, respectively, to obtain the corresponding rendered features, fine-grained features, and generated features. The spatial adaptive fusion gate takes the explicit feature residuals between rendered features, fine features, and generated features as input, and fuses them after being processed by parallel local convolutional branches and global compression-activation branches to generate a spatially adaptive gating graph. Based on the gating graph, the fine features and generated features are interpolated element-wise in a unified feature space to obtain the fused features. The residual decoder with uncertainty estimation includes a decoding path consisting of a zero-initialized residual block and a trainable convolutional decoder, as well as an uncertainty estimation path; the decoding path reconstructs a fused pseudo-ground image from the fused features, and the uncertainty estimation path predicts a spatial confidence map from the fused features and constrains the confidence value with a Softplus activation function; The feature-gated fusion network is trained on a self-built dataset, and the training process uses a composite loss function to optimize the network parameters end-to-end.
[0012] Furthermore, the feature-gated fusion network is trained based on a self-built dataset, including: Select m different scenes. For each scene, uniformly select a portion of the view from the scene as the sparse input view for simulation, and optimize a three-dimensional Gaussian splash model based on the sparse input view. From the remaining unobserved cameras in the scene, sample multiple test viewpoints of varying difficulty levels in layers according to the angle between the camera and the training view from near to far. For each test viewpoint, the original 3D Gaussian splash output is rendered and a locally refined pseudo-supervised image is generated through the rendering optimization branch. At the same time, a globally consistent pseudo-supervised image is directly synthesized based on the sparse input view through the new view generation branch. The real image, rendered image, locally refined pseudo-supervised image and globally consistent pseudo-supervised image of the test viewpoint are combined into a training sample quad. Generate n training sample quadruples in the above manner, and divide them into training set and validation set according to the ratio; The training process employs a composite loss function to perform end-to-end optimization of the network parameters, with the loss... include:
[0013] in, Representing pixel-level non-likelihood loss, it is used to jointly optimize the reconstruction error of the fused pseudo-ground image and the uncertainty regularization term of the spatial confidence map; Represents perceptual similarity loss, used to improve the perceptual quality of the fused pseudo-ground image; The gated supervision loss is represented by pseudo-labels constructed based on the mask of the void region and the comparison of errors of each branch, which accelerates the learning convergence of the gated graph. , , , and These are the balance weight coefficients corresponding to each loss term; Gating graph generated for spatial adaptive fusion gate, Spatial confidence plot for path prediction in uncertainty estimation; The total variational smoothing regularization term is applied to the gating graph and the spatial confidence graph, respectively, to promote their spatial continuity.
[0014] Furthermore, the iterative optimization of the 3D Gaussian splash representation based on the fused pseudo-truth image and spatial confidence map until a preset condition is met includes: The pseudo-ground image and spatial confidence map are fused as supervision signals, and the pseudo-supervision loss term is calculated using the following formula. :
[0015] in, The weighting coefficients for the pseudo-supervision loss term from a new perspective. The weighting coefficients are used to balance the pixel-wise L1 loss and the structural similarity loss. Represents the L1 loss term for pixel-wise values. The similarity loss term is represented by a spatial confidence graph. and its global mean The two terms are weighted, and finally the pseudo-supervised loss term and the original supervised loss term on the sparse input view are weighted and summed to form the total loss function, which is used to update the gradient of the three-dimensional Gaussian properties.
[0016] Secondly, the present invention provides a sparse view Figure 3 Reconstruction device, including: An initialization unit is used to construct an initial 3D Gaussian splash representation based on the sparse input view and the corresponding camera pose; The viewpoint selection unit is used to generate a set of candidate viewpoints covering the scene space through spatial binning and pose perturbation strategies, and to perform viewpoint filtering and dynamic backtracking based on the hole ratio of each candidate viewpoint. The dual-branch pseudo-supervised generation unit is used to execute the artifact-aware rendering optimization branch and the geometry-guided new view generation branch in parallel for each selected candidate viewpoint, producing complementary local fine-grained pseudo-supervision and globally consistent pseudo-supervision signals. An adaptive fusion unit is used to perform spatial adaptive fusion of the local fine pseudo-supervision and the global consistent pseudo-supervision signal using a trained feature-gated fusion network, and generate a fused pseudo-truth image and its corresponding spatial confidence map. The iterative optimization unit is used to iteratively optimize the three-dimensional Gaussian splash representation based on the fused pseudo-true image and the spatial confidence map.
[0017] Thirdly, the present invention provides a computer-readable storage medium storing computer instructions that, when executed on a computer, cause the computer to perform the sparse view described above. Figure 3 Dimensional reconstruction method.
[0018] Fourthly, the present invention provides an electronic device, comprising: a memory and a processor; The memory is used to store computer instructions; The processor is configured to invoke computer instructions stored in the memory, causing the electronic device to execute the sparse view described above. Figure 3 Dimensional reconstruction method.
[0019] The beneficial effects of this invention are as follows: This invention introduces an overlay-aware active exploration strategy to dynamically identify and prioritize the exploration of under-covered regions in the current representation. By decoupling the pseudo-supervised generation process into complementary rendering optimization branches and geometry-guided new view generation branches, it effectively overcomes the dilemma of the single-branch paradigm where geometric consistency and local detail fidelity are mutually exclusive. A feature-gated fusion network is designed to achieve spatial adaptive integration of the dual-branch signals, ensuring that the fused pseudo-supervision simultaneously possesses high-fidelity local texture and globally consistent geometric structure, thereby improving the final reconstruction quality. Attached Figure Description
[0020] Figure 1 A sparse view according to one embodiment of this application is shown. Figure 3 A flowchart illustrating the 3D reconstruction method; Figure 2 A sparse view according to one embodiment of this application is shown. Figure 3 A schematic diagram of the 3D reconstruction method; Figure 3 A sparse view according to one embodiment of this application is shown. Figure 3 Block diagram of the reconstruction device; Figure 4 A structural block diagram of an electronic device according to an embodiment of this application is shown; Figure 5 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the application embodiment is shown. Detailed Implementation
[0021] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments are now described. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention; that is, the described embodiments are only a part of the embodiments of the invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0022] This embodiment provides a sparse view Figure 3 The reconstruction methods, devices, storage media, and electronic equipment are described in detail below.
[0023] First, to address the problem that existing diffusion prior paradigms rely on fixed sampling trajectories and cannot actively identify under-covered areas, this embodiment uses a spatial binning and pose perturbation strategy based on predefined trajectories to generate candidate viewpoints. Then, it uses 3D Gaussian splash rendering to calculate the hole ratio for coverage-aware viewpoint selection and dynamic backtracking, thereby improving the information gain of pseudo-supervised viewpoint selection.
[0024] Secondly, to address the issue that single-branch pseudo-supervision generation struggles to balance local detail fidelity with global geometric consistency, this embodiment employs a dual-branch decoupled architecture that parallels the rendering optimization branch and the new view generation branch. This architecture generates locally refined pseudo-supervision signals and globally consistent pseudo-supervision signals, respectively, and then performs spatial adaptive fusion through a feature-gated fusion network, thereby improving the reliability and accuracy of the pseudo-supervision signals.
[0025] like Figure 1 As shown, a sparse view is demonstrated. Figure 3 The dimensional reconstruction method includes steps S100 to S500.
[0026] Step S100: Construct an initial 3D Gaussian splash representation based on the sparse input view and the corresponding camera pose.
[0027] Understandably, the initial 3D Gaussian splash representation provides the basic geometric structure and appearance parameter representation for the entire reconstruction process, and its differentiability supports subsequent gradient-driven parameter optimization. The subsequent active exploration phase relies on this initial model to render candidate viewpoints and evaluate their value, while the pseudo-supervised generation phase relies on this initial model to render new viewpoint images to construct pseudo-ground value supervision signals.
[0028] In some feasible embodiments, based on the foregoing scheme, the step of constructing an initial 3D Gaussian splash representation based on the sparse input view and the corresponding camera pose includes: The initial point cloud and camera pose are estimated from the sparse input view using a motion recovery structure algorithm or a pre-trained geometric model; A set of three-dimensional Gaussian elements is initialized based on the initial point cloud. Each Gaussian element consists of the mean position. Covariance matrix ,color and opacity Parameterized description; The three-dimensional Gaussian primitives are rendered using a differentiable rasterizer to obtain a rendered image corresponding to the sparse input view perspective. Given a camera pose, formula (1) represents the rasterization rendering process: (1) in, The image representing the rendered image. This represents the transparency map corresponding to the rendered image. For pixel coordinates, The contributing Gaussian set is the set of pixels that are ordered along the depth direction for this pixel.
[0029] The initial optimization uses real images of sparse input views as supervision and adopts a weighted combination of L1 loss and structural similarity loss as the optimization objective. After a few iterations, the initial 3D Gaussian splash representation is obtained.
[0030] It should be noted that, due to the extremely limited number of sparse input views, the initial 3D Gaussian splash representation can achieve high rendering quality on the training views, but it still suffers from severe geometric holes and floating object artifacts in new viewpoint directions deviating from the training pose. Subsequent steps S200 to S500 are designed to address these shortcomings.
[0031] Continue to refer to Figure 1 In step S200, based on a predefined smooth reference trajectory, a high information gain viewpoint is selected through coverage perception filtering and dynamic backoff mechanism.
[0032] Understandably, existing methods typically rely on fixed sampling trajectories to select pseudo-supervision viewpoints, failing to address the actual deficiencies in the current reconstruction results. This step introduces a coverage-aware active exploration mechanism, enabling the system to dynamically identify and prioritize regions with the weakest geometric coverage in the current 3D representation, thereby maximizing the potential information gain of the pseudo-supervision signal.
[0033] In some feasible embodiments, based on the aforementioned scheme, the selection of high information gain viewpoints based on a predefined smooth reference trajectory through coverage-aware filtering and dynamic backoff mechanisms includes: Obtain a predefined smooth reference trajectory, and divide it evenly along the reference trajectory according to the arc length. The space is divided into bins, and the anchor point camera pose is extracted from each bin. , This ensures that candidate viewpoints are geometrically uniformly distributed along the reference trajectory in the scene space; Within each spatial bin, a perturbation is applied to the pose of the anchor point camera to generate a... There are 10 candidate poses. Specifically, two complementary pose perturbation strategies are adopted. All perturbation amplitudes are normalized with the scene scale parameter, which is defined as the mean of the pairwise Euclidean distances between the translation components of all anchor camera poses. The calculation method is shown in formula (2): (2) The first disturbance strategy is a rotation-based disturbance: applying interval-based disturbances to the yaw and pitch angles respectively. and interval Large-scale offset of uniform sampling. By perturbing the yaw angle to significantly change the horizontal viewing direction, previously obscured or lateral surfaces at the edge of the view frustum can be exposed; perturbing the pitch angle changes the vertical viewing angle, revealing top and bottom structures that are not visible from the training viewpoint.
[0034] The second perturbation strategy is translation-based: applying a large lateral displacement of the camera. Simultaneously apply from the interval Uniform sampling with longitudinal displacement. By moving the camera position laterally and longitudinally, unobserved scene areas at different depth levels can be revealed while maintaining a roughly constant observation direction.
[0035] Within each compartment The candidate poses are allocated proportionally: among which One disturbance primarily uses rotation, the rest... The perturbation primarily uses translation. This proportion can be adjusted according to the characteristics of the scene; for example, the proportion of rotational perturbation can be increased in scenes with abundant vertical structures.
[0036] For each generated candidate camera pose, a rasterized rendering is performed based on the current 3D Gaussian splash representation, as shown in Formula (1), to obtain the rendered image and the corresponding transparency map. Since relying solely on continuous transparency maps makes it difficult to reliably distinguish between true geometrically missing regions and normal unoccupied background regions, even in distant background regions of the scene where the 3D structure is complete, the transparency value may be low, and if not distinguished, it will be misjudged as a geometric hole. Furthermore, under sparse view conditions, the global floating object problem is severe; densely distributed floating objects may cause the overall transparency map of the area to tend towards saturation, thus masking the true hole distribution. Therefore, this invention applies thresholding and morphological filtering to the transparency map to obtain a stable binary hole mask, calculated as shown in Formula (3): (3) in, The transparency threshold. This represents a morphological opening operation, used to remove isolated noise pixels generated after thresholding, retaining only spatially continuous areas of holes. Based on this hole mask, the hole ratio is defined as the proportion of holed pixels to the total number of pixels in the image, calculated as shown in formula (4): (4) Based on void ratio Each candidate viewpoint is divided into three levels: if The candidate viewpoint was classified as a high-value viewpoint; if The candidate viewpoint was classified as a medium-value viewpoint; if or This candidate viewpoint was classified as a low-value viewpoint. A typical threshold was... .
[0037] To evaluate the overall exploration quality of the current candidate viewpoint set, the coverage ratio is defined as the proportion of spatial bins containing at least one high-value candidate viewpoint to the total number of bins, calculated as shown in formula (5): (5) When coverage ratio Below the preset backoff threshold If the current perturbation range is insufficient to effectively cover under-reconstructed areas in the scene, and most candidate viewpoints fall in extreme areas that are either fully reconstructed or almost completely unreconstructed, a dynamic backoff mechanism is triggered: the angle range of the rotation perturbation and the distance range of the translation perturbation are uniformly expanded by a preset scaling factor, and the candidate pose generation, rendering, and scoring process is re-executed based on the expanded perturbation range. This backoff process iterates a maximum of two times. If the coverage ratio still does not meet the threshold requirement after reaching the maximum number of iterations, the one with the highest coverage ratio in all iterations is selected as the final candidate set. Typical parameters are set as follows: .
[0038] Regardless of whether a rollback is triggered, all candidate viewpoints at high-value and medium-value levels in the selected rounds are ultimately retained to ensure sufficient spatial coverage. To keep the computational cost of subsequent dual-branch pseudo-supervision generation and fusion within a manageable range, a global capacity cap is set, where is the capacity multiplier. Typical value When the total number of retained candidate viewpoints exceeds the upper limit, the three spatial bins containing the most candidate viewpoints are identified, and candidate viewpoints are successively removed from these bins in order of increasing void ratio until the total number meets the capacity constraint.
[0039] It should be noted that the dynamic backoff mechanism is a key safeguard in this step. It ensures that when the initially set perturbation range is insufficient to effectively cover the under-reconstructed areas in the scene, the system can adaptively expand the search range, thereby avoiding the exploration process from getting trapped in local optima.
[0040] Continue to refer to Figure 1 In step S300, for each selected candidate viewpoint, the artifact-aware rendering optimization branch and the geometry-guided new view generation branch are executed in parallel to generate complementary local fine-grained pseudo-supervision signals and globally consistent pseudo-supervision signals.
[0041] Understandably, while frame-by-frame independent processing can preserve local details, it easily disrupts cross-view functionality in existing methods. Figure 1 While video diffusion models can guarantee geometric coherence, they may introduce motion blur and lose local texture. This step decouples the pseudo-supervision generation into two complementary parallel branches, enabling the final fused supervision signal to simultaneously maintain both local texture fidelity and global geometric consistency.
[0042] In some feasible embodiments, based on the foregoing scheme, the parallel execution of the artifact-aware rendering optimization branch and the geometry-guided new view generation branch for each selected candidate viewpoint includes: In the rendering optimization branch, the training view that is closest to the current candidate viewpoint's spatial location is first retrieved as the reference frame using formula (6): (6) Among them, the metric function for measuring the distance between two camera poses can be calculated comprehensively based on the Euclidean distance of the translation component and the angular difference of the rotation component. The rendering optimization branch synchronously inputs the current rendered image under the candidate viewpoint and the retrieved reference frame into the reference-guided single-step diffusion model to obtain the local fine-monitored image, as shown in formula (7): (7) This model employs a cross-attention mechanism to inject reliable structural information and texture cues from the reference frame into the rendered image to be repaired. A single-step denoising process removes artifacts such as floating objects and ghosting in the rendered image. This single-step design makes the computational efficiency of this branch extremely high, requiring only one forward propagation per inference step. The locally refined pseudo-supervised image output by this branch effectively preserves high-frequency texture details in geometrically well-reconstructed regions. However, for regions with large structural holes in the rendered image, the lack of effective pixel information for repair often makes it difficult to generate reasonable missing content.
[0043] In the geometry-guided new view generation branch, pre-computed scene depth information is used as an explicit geometric prior. The depth information can be jointly estimated from the sparse input views using the same geometric prior model (such as DUST3R) as in step S100. This branch performs geometry-guided new view synthesis using a multi-view diffusion model, as shown in Equation (8): (8) in, For a sparse input view collection, and These are the sets of intrinsic and extrinsic parameter matrices for all input views, respectively. The target candidate camera pose is determined. The model then... The features of a sparse reference view are twisted to the target viewpoint based on depth and camera geometry to construct a feature volume with a clear geometric correspondence. Then, the global appearance information is aggregated through a multi-view cross-attention mechanism to finally generate a new view in the target pose.
[0044] Guided by explicit geometric priors, the globally consistent pseudo-supervised image output by this branch exhibits excellent geometric coherence and structural integrity, effectively filling large structural holes. However, because its synthesis process is independent of the rendered image from the current viewpoint, this branch may completely discard reliable high-frequency texture information already present in the rendered image in areas that have been well reconstructed, resulting in texture fidelity in these areas being lower than that of the output from the rendering optimization branch.
[0045] It should be noted that the complementarity of the two branches is reflected in different spatial regions of the same target viewpoint. In geometric regions that have been well reconstructed by the current 3D Gaussian splashing representation, the output of the rendering optimization branch has higher texture fidelity; in regions with large geometric holes, the output of the global generation branch has more reasonable structural completion. This spatially complementary distribution is the fundamental motivation for the subsequent design of the S400 spatial adaptive fusion mechanism.
[0046] Continue to refer to Figure 1 In step S400, the trained feature-gated fusion network is used to perform spatial adaptive fusion of the local fine pseudo-supervision and the global consistent pseudo-supervision signals to generate a fused pseudo-truth image and its corresponding spatial confidence map.
[0047] Understandably, simply blending the outputs of both branches directly in the pixel domain can easily produce artifacts such as ghosting and oversmoothing. This step, by performing gated fusion in the implicit feature space, allows the network to learn the optimal blending ratio of the two branches at each spatial location in a data-driven manner, thereby maximizing the quality of the fused ground truth value.
[0048] In some feasible embodiments, based on the aforementioned scheme, the feature-gated fusion network includes a shared feature encoder, a spatial adaptive fusion gate, and a residual decoder with uncertainty estimation. Its overall processing flow includes: The shared feature encoder is used to input the rendered image of the current candidate viewpoint, the locally refined pseudo-supervised image output by the rendering optimization branch, and the globally consistent pseudo-supervised image output by the new view generation branch into the shared feature encoder, and map them to a unified feature space, as shown in Equation (9): (9) in, This represents a shared feature encoder, employing a pre-trained IFCNN encoder whose weights are frozen during training. The three images share the same encoder's weight parameters, ensuring comparability within the same feature space and giving subsequent feature residual calculations a clear physical meaning.
[0049] In the spatial adaptive fusion gate, the difference between the rendered feature and the fine feature, the difference between the rendered feature and the generated feature, and the channel concatenation of the rendered feature are used as inputs, as shown in formula (10): (10) After processing by parallel local convolutional branches and global squeeze-activation branches, a spatially adaptive gated graph is generated by activation using the Sigmoid function, as shown in Equation (11): (11) Based on the gating graph, the fine features and the generated features are subjected to element-wise weighted linear interpolation in the unified feature space to obtain the fused features, as shown in formula (12): (12) The residual decoder for uncertainty estimation includes a decoding path consisting of a zero-initialized residual block and a trainable convolutional decoder, as well as an uncertainty estimation path. The decoding path reconstructs the fused pseudo-ground image from the fused features, as shown in Equation (13): (13) The uncertainty estimation path predicts the spatial confidence map from the fused features and constrains the confidence value with the Softplus activation function, as shown in Equation (14): (14) To prevent the confidence value from degenerating to zero, the Softplus function is chosen to ensure the non-negativity and smoothness of the confidence value, avoiding the hard truncation problem near zero that occurs with ReLU-like functions. Each pixel value in the spatial confidence map obtained from this path prediction reflects the model's assessment of the reliability of the fusion result at that location.
[0050] Furthermore, based on the aforementioned scheme, the feature-gated fusion network needs to be trained offline before application. The feature-gated fusion network is trained on a self-built large-scale dataset, which is constructed as follows: Specifically, 140 samples (reference value) covering diverse indoor and outdoor scenes from the DL3DV-10K benchmark dataset were selected. For each scene, three views were uniformly selected as simulated sparse input views, and a 3D Gaussian splash model was optimized based on these sparse input views, with 4000 iterations (reference value). From the remaining unobserved cameras in the scene, test viewpoints were sampled hierarchically from near to far according to their camera angles relative to the training views, covering four difficulty levels: easy level (0°–8°), medium level (8°–18°), difficult level (18°–35°), and extreme level (35°–60°). This hierarchical sampling strategy simulates different sparse reconstruction difficulties from slight to large extrapolation, ensuring that the training data covers diverse degradation patterns.
[0051] For each test viewpoint, the original 3D Gaussian splash output is rendered and input into the rendering optimization branch to generate a locally refined pseudo-supervised image. Simultaneously, the sparse input view and its corresponding camera parameters are input into the new view generation branch to directly synthesize a globally consistent pseudo-supervised image. Finally, the real image, rendered image, locally refined pseudo-supervised image, and globally consistent pseudo-supervised image of the test viewpoint are combined into a training sample quadruple. Following this method, a total of 21,665 training sample quadruples are generated from 140 scenes, of which 19,119 are used for training and 2,546 are used for validation (reference value).
[0052] Furthermore, based on the aforementioned scheme, the training process of the feature-gated fusion network employs a composite loss function to perform end-to-end optimization of the network parameters. During training, the weights of the shared feature encoder remain frozen, and only the parameters of the spatial adaptive fusion gate, residual decoder, and uncertainty estimation path are optimized. Specifically, the training uses the Adam optimizer with a batch size of 4, and a two-stage scheduling strategy for the learning rate: initially, training is performed for 10 epochs at a learning rate of , followed by fine-tuning for 40 epochs at a learning rate of . The composite loss function is shown in formula (15): (15) in, Representing pixel-level non-likelihood loss, it is used to jointly optimize the reconstruction error of the fused pseudo-ground image and the uncertainty regularization term of the spatial confidence map; Represents perceptual similarity loss, used to improve the perceptual quality of the fused pseudo-ground image; The gated supervision loss is represented by pseudo-labels constructed based on the mask of the void region and the comparison of errors of each branch, which accelerates the learning convergence of the gated graph. , , , and These are the balance weight coefficients corresponding to each loss term; Gating graph generated for spatial adaptive fusion gate, Spatial confidence plot for path prediction in uncertainty estimation; The total variational smoothing regularization term is applied to the gating graph and the spatial confidence graph, respectively, to promote their spatial continuity.
[0053] It should be noted that after training, the parameters of the feature-gated fusion network remain fixed and can be directly applied to the pseudo-supervised fusion stage in the downstream 3D reconstruction process without the need for additional fine-tuning for each new scene.
[0054] Continue to refer to Figure 1 In step S500, the three-dimensional Gaussian splash representation is iteratively optimized based on the fused pseudo-true image and the spatial confidence map.
[0055] Understandably, steps S200 to S400 provide each candidate viewpoint selected through coverage-aware filtering with a high-quality fused pseudo-ground truth image and a corresponding pixel-by-pixel confidence map. However, these pseudo-supervision signals inevitably contain a certain degree of synthesis error. If pseudo-supervision from all regions is injected into the 3D optimization process with equal weight without differentiation, unreliable gradients in low-quality regions will accumulate and destroy the already well-reconstructed geometry. This step, through a confidence-aware weighted loss function, automatically focuses the optimization process on high-confidence regions while suppressing harmful gradients in low-confidence regions, thereby ensuring optimization stability while fully utilizing the information gain from pseudo-supervision.
[0056] In some feasible embodiments, based on the foregoing scheme, the iterative optimization of the 3D Gaussian splash representation based on the fused pseudo-truth image and the spatial confidence map includes: For each candidate viewpoint selected through coverage perception, a rendered image is obtained using the current 3D Gaussian splash representation. The pseudo-supervised loss function is defined as the confidence-weighted composite photometric loss between the pseudo-ground image and the rendered image, as shown in formula (16): (16) in, The weighting coefficients for the pseudo-supervision loss term from a new perspective. The weighting coefficients are used to balance the pixel-wise L1 loss and the structural similarity loss. Represents the L1 loss term for pixel-wise values. The similarity loss term is represented by a spatial confidence graph. and its global mean The two terms are weighted, and finally the pseudo-supervised loss term and the original supervised loss term on the sparse input view are weighted and summed to form the total loss function, which is used to update the gradient of the three-dimensional Gaussian properties.
[0057] In the overall optimization process, the initial stage is based on standard 3D Gaussian splash optimization. At preset optimization iteration nodes (e.g., the 5000th and 8000th iterations), the coverage-aware active exploration in step S200 is triggered to generate fusion pseudo-true values for newly discovered under-covered viewpoints, and the aforementioned weighted loss function is used for subsequent optimization until the total number of iterations reaches the preset upper limit of 10000.
[0058] It should be noted that the total computation time for the entire process is approximately 20 minutes, which is an order of magnitude more efficient than the video diffusion paradigm, which requires several hours. This efficiency advantage stems from two key designs: first, the rendering optimization branch adopts a single-step diffusion model, requiring only one forward propagation for each inference; second, the coverage-aware active exploration concentrates computational resources on the viewpoint with the highest information gain, avoiding redundant inference of a large number of consecutive frames as required by the video diffusion paradigm.
[0059] Specifically, the following provides an example of implementing the method of this application using a specific module.
[0060] See Figure 2 The diagram illustrates a coverage-aware active exploration implementation using specific modules according to an embodiment of this application.
[0061] like Figure 2As shown, the active viewpoint exploration module first performs spatial binning based on a predefined smooth reference trajectory. After acquiring the anchor point camera pose, it generates candidate viewpoint sets using two strategies: rotation-based and translation-based. Then, based on the current 3D Gaussian rendering, it renders the corresponding transparency maps for each candidate viewpoint. After calculating the hole ratio of the transparency maps, it categorizes them into three types: high-value viewpoints, medium-value viewpoints, and low-value viewpoints. Next, it calculates the coverage ratio of high-value viewpoints. If it falls below a set threshold, a dynamic backoff mechanism is triggered to resample the candidate viewpoints.
[0062] The following describes an embodiment of the apparatus described in this application, which can be used to execute a sparse view in the above embodiments of this application. Figure 3 3D reconstruction method. For details not disclosed in the device embodiments of this application, please refer to the embodiments of the methods described above in this application.
[0063] Reference Figure 3 As shown, a sparse view according to an embodiment of this application Figure 3 The 3D reconstruction device 500 includes an initialization unit 501, an active exploration and viewpoint selection unit 502, a dual-branch pseudo-supervised generation unit 503, an adaptive fusion unit 504, and an iterative optimization unit 505.
[0064] The system comprises the following components: an initialization unit 501, which constructs an initial 3D Gaussian splash representation based on a sparse input view and the corresponding camera pose; a viewpoint selection unit 502, which selects a high information gain viewpoint based on a predefined smooth reference trajectory through overlay-aware filtering and a dynamic backoff mechanism; a dual-branch pseudo-supervised generation unit 503, which executes in parallel an artifact-aware local rendering optimization branch and a geometry-guided global new view generation branch for each selected candidate viewpoint, generating complementary local fine-grained pseudo-supervised signals and globally consistent pseudo-supervised signals; an adaptive fusion unit 504, which uses a trained feature-gated fusion network to perform spatial adaptive fusion of the local fine-grained pseudo-supervised signals and the globally consistent pseudo-supervised signals, generating a fused pseudo-ground truth image and its corresponding spatial confidence map; and an iterative optimization unit 505, which iteratively optimizes the 3D Gaussian splash representation based on the fused pseudo-ground truth image and the spatial confidence map until a preset condition is met.
[0065] like Figure 4 As shown, this application embodiment also provides an electronic device 600, including a memory 610, a processor 620, and a computer program 611 stored in the memory 610 and executable on the processor. When the processor 620 executes the computer program 611, it implements the aforementioned sparse view. Figure 3 The steps of any method in the dimensional reconstruction method.
[0066] Since the electronic device described in this embodiment is a sparse vision implementation of the embodiments of this application... Figure 3 The equipment used in the reconstruction device is based on the methods described in the embodiments of this application. Therefore, those skilled in the art can understand the specific implementation methods and various variations of the electronic equipment in this embodiment. So, how the electronic equipment implements the methods in the embodiments of this application will not be described in detail here. As long as those skilled in the art implement the methods in the embodiments of this application, the equipment used is within the scope of protection of this application.
[0067] In practice, when the computer program 611 is executed by the processor, it can implement any of the embodiments corresponding to the first aspect.
[0068] Figure 5 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.
[0069] It should be noted that, Figure 5 The computer system 700 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0070] like Figure 5 As shown, the computer system 700 includes a Central Processing Unit (CPU) 701, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 702 or programs loaded from storage portion 708 into Random Access Memory (RAM) 703, such as performing the methods described in the above embodiments. The RAM 703 also stores various programs and data required for system operation. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 704. An Input / Output (I / O) interface 705 is also connected to the bus 704.
[0071] The following components are connected to I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to I / O interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 710 as needed so that computer programs read from it can be installed into storage section 708 as needed.
[0072] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is executed by central processing unit (CPU) 701, it performs various functions defined in the system of this application.
[0073] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0074] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0075] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0076] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0077] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this application.
[0078] Other embodiments of this application will readily conceive of by those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. It should be understood that this application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
[0079] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. A sparse view 3D reconstruction method, characterized in that, include: Based on the sparse input view and the corresponding camera pose, an initial 3D Gaussian splash representation is constructed. Based on a predefined smooth reference trajectory, high information gain viewpoints are selected through coverage-aware filtering and dynamic backoff mechanisms. For each selected candidate viewpoint, the rendering optimization branch and the new view generation branch are executed in parallel to generate complementary local fine-grained pseudo-supervision and global consistent pseudo-supervision signals; Using a trained feature-gated fusion network, the local fine pseudo-supervision and the global consistent pseudo-supervision signals are spatially adaptively fused to generate a fused pseudo-truth image and its corresponding spatial confidence map. Based on the fused pseudo-true value image and spatial confidence map, the 3D Gaussian splash representation is iteratively optimized until the preset conditions are met.
2. The sparse view three-dimensional reconstruction method according to claim 1, characterized in that, The initial 3D Gaussian splash representation, constructed based on the sparse input view and corresponding camera pose, includes: The initial point cloud and camera pose are estimated from the sparse input view using a motion recovery structure algorithm or a pre-trained geometric model; A set of three-dimensional Gaussian elements is initialized based on the initial point cloud. Each Gaussian element is parameterized by mean position, covariance matrix, color, and opacity. The three-dimensional Gaussian primitives are rendered using a differentiable rasterizer to obtain a rendered image corresponding to the sparse input view. The reconstruction loss between the rendered image and the sparse input view is calculated. The parameters of the three-dimensional Gaussian primitives are iteratively optimized through backpropagation to obtain an initial three-dimensional Gaussian splash representation.
3. The sparse view three-dimensional reconstruction method according to claim 1, characterized in that, The selection of high information gain viewpoints based on predefined smooth reference trajectories through coverage-aware filtering and dynamic backoff mechanisms includes: The space is divided into multiple spatial bins along a smooth reference trajectory by arc length, and the anchor point camera pose of each bin is extracted. Within each spatial sub-box, rotation-based and translation-based perturbations are applied to the pose of the anchor point camera. Both types of perturbations are normalized using scene scale parameters to generate the candidate viewpoint set. For each candidate viewpoint, render the current 3D Gaussian splash representation to obtain the rendered image and its corresponding transparency map; The transparency map is thresholded and morphologically filtered to obtain a binarized hole mask, and then the hole ratio, which represents the degree of geometric loss of the viewpoint, is calculated. Based on the hole ratio, candidate viewpoints are divided into high-value, medium-value, and low-value levels, and a dynamic backoff mechanism is triggered according to the coverage ratio of high-value candidate viewpoints in each bin: when the coverage ratio is lower than the backoff threshold, the perturbation range is expanded and candidate viewpoints are resampled until the coverage ratio requirement is met or the maximum number of retries is reached. After the final round of candidate sampling is completed, all candidate viewpoints at the high-value and medium-value levels are retained. If the total number of retained candidate viewpoints exceeds the preset global capacity limit, then, provided that each bin has at least one candidate viewpoint, the candidate viewpoint with the lowest void ratio is removed from the bin with the most candidate viewpoints.
4. The sparse view three-dimensional reconstruction method according to claim 1, characterized in that, For each selected candidate viewpoint, the rendering optimization branch and the new view generation branch are executed in parallel to generate complementary local fine-grained pseudo-supervision and globally consistent pseudo-supervision signals, including: In the rendering optimization branch, the training view closest to the spatial position of the currently selected viewpoint is retrieved as the reference frame. The reference-guided diffusion model is used to remove floating artifacts and ghost images in the rendered image under the currently selected viewpoint in a single-step denoising manner, generating a local fine pseudo-supervision signal. In the new view generation branch, a geometry-guided multi-view diffusion model is used, which combines sparse input views and corresponding depth geometric priors. By distorting the features of the reference view to the target viewpoint, structural guidance is provided to directly synthesize a new view with cross-view geometric consistency, which serves as a globally consistent pseudo-supervision signal.
5. The sparse view three-dimensional reconstruction method according to claim 1, characterized in that, The step of using a trained feature-gated fusion network to perform spatial adaptive fusion of the local fine-grained pseudo-supervision and the globally consistent pseudo-supervision signals to generate a fused pseudo-ground value image and its corresponding spatial confidence map includes: A feature-gated fusion network is constructed, which includes a shared feature encoder, a spatial adaptive fusion gate, and a residual decoder with uncertainty estimation. The shared feature encoder maps the current viewpoint's rendered image, the local fine-grained pseudo-supervised image, and the globally consistent pseudo-supervised image to a unified feature space, respectively, to obtain the corresponding rendered features, fine-grained features, and generated features. The spatial adaptive fusion gate takes the explicit feature residuals between rendered features, fine features, and generated features as input, and fuses them after being processed by parallel local convolutional branches and global compression-activation branches to generate a spatially adaptive gating graph. Based on the gating graph, the fine features and generated features are interpolated element-wise in a unified feature space to obtain the fused features. The residual decoder with uncertainty estimation includes a decoding path consisting of a zero-initialized residual block and a trainable convolutional decoder, as well as an uncertainty estimation path; the decoding path reconstructs a fused pseudo-ground image from the fused features, and the uncertainty estimation path predicts a spatial confidence map from the fused features and constrains the confidence value with a Softplus activation function; The feature-gated fusion network is trained on a self-built dataset, and the training process uses a composite loss function to optimize the network parameters end-to-end.
6. The sparse view three-dimensional reconstruction method according to claim 5, characterized in that, The feature-gated fusion network is trained on a self-built dataset, including: Select m different scenes. For each scene, uniformly select a portion of the view from the scene as the sparse input view for simulation, and optimize a three-dimensional Gaussian splash model based on the sparse input view. From the remaining unobserved cameras in the scene, sample multiple test viewpoints of varying difficulty levels in layers according to the angle between the camera and the training view from near to far. For each test viewpoint, the original 3D Gaussian splash output is rendered and a locally refined pseudo-supervised image is generated through the rendering optimization branch. At the same time, a globally consistent pseudo-supervised image is directly synthesized based on the sparse input view through the new view generation branch. The real image, rendered image, locally refined pseudo-supervised image and globally consistent pseudo-supervised image of the test viewpoint are combined into a training sample quad. Generate n training sample quadruples in the above manner, and divide them into training set and validation set according to the ratio; The training process employs a composite loss function to perform end-to-end optimization of the network parameters, with the loss... include: in, Representing pixel-level non-likelihood loss, it is used to jointly optimize the reconstruction error of the fused pseudo-ground image and the uncertainty regularization term of the spatial confidence map; Represents perceptual similarity loss, used to improve the perceptual quality of the fused pseudo-ground image; The gated supervision loss is represented by pseudo-labels constructed based on the mask of the void region and the comparison of errors of each branch, which accelerates the learning convergence of the gated graph. , , , and These are the balance weight coefficients corresponding to each loss term; Gating graph generated for spatial adaptive fusion gate. Spatial confidence plot for path prediction in uncertainty estimation; The total variational smoothing regularization term is applied to the gating graph and the spatial confidence graph, respectively, to promote their spatial continuity.
7. The sparse view three-dimensional reconstruction method according to claim 1, characterized in that, The step involves iteratively optimizing the 3D Gaussian splash representation based on the fused pseudo-ground image and spatial confidence map until preset conditions are met, including: The pseudo-ground image and spatial confidence map are fused as supervision signals, and the pseudo-supervision loss term is calculated using the following formula. : in, The weighting coefficients for the pseudo-supervision loss term from a new perspective. The weighting coefficients are used to balance the pixel-wise L1 loss and the structural similarity loss. Represents the L1 loss term for pixel-wise values. The similarity loss term is represented by a spatial confidence graph. and its global mean The two terms are weighted, and finally the pseudo-supervised loss term and the original supervised loss term on the sparse input view are weighted and summed to form the total loss function, which is used to update the gradient of the three-dimensional Gaussian properties.
8. A sparse view three-dimensional reconstruction device, characterized in that, include: An initialization unit is used to construct an initial 3D Gaussian splash representation based on the sparse input view and the corresponding camera pose; The viewpoint selection unit is used to generate a set of candidate viewpoints covering the scene space through spatial binning and pose perturbation strategies, and to perform viewpoint filtering and dynamic backtracking based on the hole ratio of each candidate viewpoint. The dual-branch pseudo-supervised generation unit is used to execute the artifact-aware rendering optimization branch and the geometry-guided new view generation branch in parallel for each selected candidate viewpoint, producing complementary local fine-grained pseudo-supervision and globally consistent pseudo-supervision signals. An adaptive fusion unit is used to perform spatial adaptive fusion of the local fine pseudo-supervision and the global consistent pseudo-supervision signal using a trained feature-gated fusion network, and generate a fused pseudo-truth image and its corresponding spatial confidence map. The iterative optimization unit is used to iteratively optimize the three-dimensional Gaussian splash representation based on the fused pseudo-true image and the spatial confidence map.
9. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions that, when executed on a computer, cause the computer to perform the sparse view three-dimensional reconstruction method as described in any one of claims 1-7.
10. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer instructions; The processor is configured to invoke computer instructions stored in the memory, causing the electronic device to execute the sparse view three-dimensional reconstruction method as described in any one of claims 1-7.