Three-dimensional gaussian reconstruction method, device, and storage medium
Patent Information
- Application Number
- CN202610926679.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本申请的主要目的在于提供一种三维高斯重建方法、三维高斯重建设备以及存储介质,旨在解决常规稀疏视角下三维高斯泼溅的重建精度低的技术问题
[0019]本申请提出的一个或多个技术方案,至少具有以下技术效果:通过获取由稀疏真实视角图像及对应真实相机位姿初始化并优化得到的当前三维高斯场景,在确定目标伪相机位姿后,在目标伪相机位姿处生成伪视角图像时,同时利用所述当前三维高斯场景的几何信息、所述真实视角图像的风格特征、预设扩散模型的场景记忆能力以及所述当前三维高斯场景在选定的伪相机位姿下的渲染结果,这四重约束共同作用于扩散模型的生成过程。其中,几何信息约束生成图像的深度结构和三维布局与当前已重建几何对齐,风格特征约束使生成内容的色调、光照、纹理与真实视角保持一致,场景记忆能力约束让扩散模型记住当前场景特有的外观细节,而渲染结果作为草稿锚定则防止生成结果偏离当前重建状态。通过上述四重互补约束,将扩散模型原本宽泛的生成自由度压缩至与当前场景几何、风格、细节高度一致的分布内,从而生成与真实场景实际几何结构相一致的可靠伪视角图像。随后,将该可靠伪视角图像及其对应的伪相机位姿,与真实视角图像及其对应的真实相机位姿,共同作为监督信号优化当前三维高斯场景。由于伪视角图像在生成阶段已具备与当前场景一致的几何和外观特性,避免了不可靠内容对梯度更新的干扰,从而使稀疏视角下的三维高斯场景重建精度获得有效提升。
Smart Images

Figure CN122820972A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of scene reconstruction technology, and in particular to a three-dimensional Gaussian reconstruction method, a three-dimensional Gaussian reconstruction device, and a storage medium. Background Technology
[0002] When performing 3D Gaussian reconstruction based on a sparse viewpoint, the number of real cameras is limited and their positions are fixed. This makes it difficult to obtain sufficient geometric constraints and appearance information for spatial regions not covered by real cameras (i.e., unobserved regions), which in turn affects the quality of the 3D structure reconstruction of the complete scene.
[0003] Traditional techniques typically employ pre-trained 2D diffusion models to generate pseudo-viewpoint images, which are then used as additional supervisory signals in the optimization training of 3D Gaussian splashing. However, the camera poses of these pseudo-viewpoints are determined solely through a fixed interpolation method between existing real camera poses. Furthermore, all generated pseudo-viewpoint images and their corresponding camera poses are directly used for scene optimization, regardless of any significant deviations between the rendering results of the pseudo-viewpoint images and the real scene geometry; they participate in scene parameter updates to the same extent as real images. In this context, unreliable content inconsistent with the actual geometry of the real scene (i.e., unreliable pseudo-viewpoint images) can interfere with gradient updates during the optimization process, thus limiting the reconstruction accuracy of 3D Gaussian splashing methods under sparse viewpoints. Summary of the Invention
[0004] The main objective of this application is to provide a three-dimensional Gaussian reconstruction method, a three-dimensional Gaussian reconstruction device, and a storage medium, aiming to solve the technical problem of low reconstruction accuracy of three-dimensional Gaussian splashing under conventional sparse perspective.
[0005] To achieve the above objectives, this application proposes a three-dimensional Gaussian reconstruction method, which includes: Obtain the current 3D Gaussian scene, wherein the current 3D Gaussian scene is determined based on sparse real-view images and their corresponding real camera poses; Determine the target pseudo-camera pose based on the actual camera pose; Based on the geometric information of the current 3D Gaussian scene, the style features of the real-view image, the scene memory capability of the preset diffusion model, and the rendering result of the current 3D Gaussian scene under the pose of the target pseudo-camera, a pseudo-view image is generated at the pose of the target pseudo-camera. Based on the supervision signal, the current 3D Gaussian scene is optimized to obtain the target 3D Gaussian scene, wherein the supervision signal includes: the pseudo-view image, the target pseudo-camera pose, the real view image, and the real camera pose.
[0006] In one embodiment, the step of determining the target pseudo-camera pose based on the real camera pose includes: Based on the existing real camera poses in the current 3D Gaussian scene, determine the candidate pseudo camera poses; Based on the comprehensive evaluation value corresponding to the candidate pseudo camera pose, the target pseudo camera pose is selected from the candidate pseudo camera poses.
[0007] In one embodiment, the step of determining candidate pseudo-camera poses based on existing real camera poses in the current 3D Gaussian scene includes: Based on the existing real camera poses in the current 3D Gaussian scene, generate pose pairs, wherein the pose pairs include two different real camera poses; Spherical linear interpolation is performed on the rotational components of the pose in the pose pair to obtain the first interpolation result; Linear interpolation is performed on the translation components of the pose in the pose pair to obtain a second interpolation result; Based on the first interpolation result and the second interpolation result, an interpolated pseudo-camera pose is generated between the two poses in the pose pair; The rotation and translation components of each pose in the pose pair are extrapolated along both ends of the line connecting the two poses in the pose pair to generate an extrapolated pseudo camera pose located outside the line connecting the two poses in the pose pair. The interpolated pseudo-camera pose and the extrapolated pseudo-camera pose are determined as candidate pseudo-camera poses.
[0008] In one embodiment, the step of selecting the target pseudo-camera pose from the candidate pseudo-camera poses based on the comprehensive evaluation value corresponding to the candidate pseudo-camera poses includes: Obtain the reconstruction information gain index and generation reliability index of each candidate pseudo-camera pose; The comprehensive evaluation value of each candidate pseudo camera pose is calculated based on the preset annealing index, the reconstruction information gain index, and the generation reliability index. The candidate pseudo-camera poses are sorted according to the comprehensive evaluation value to obtain the sorting result; Based on the sorting results, the target pseudo camera pose is determined from each of the candidate pseudo camera poses.
[0009] In one embodiment, the step of determining the target pseudo-camera pose from the candidate pseudo-camera poses based on the sorting result includes: Based on the sorting results, the pose of the current candidate pseudo camera is determined; In the absence of selecting any target pseudo-camera pose, the current candidate pseudo-camera pose is determined as the target pseudo-camera pose; If at least one target pseudo-camera pose has been selected, determine whether the pose distance between the current candidate pseudo-camera pose and the selected target pseudo-camera pose is greater than or equal to a preset distance threshold. If the pose distance between the current candidate pseudo camera pose and each of the selected target pseudo camera poses is greater than or equal to a preset distance threshold, the current candidate pseudo camera pose is determined as the target pseudo camera pose. If the attitude distance between the current candidate pseudo-camera pose and at least one selected target pseudo-camera pose is less than a preset distance threshold, the candidate pseudo-camera pose that is one position after the current candidate pseudo-camera pose in the sorting result is determined as the new current candidate pseudo-camera pose, and the step of determining whether the attitude distance between the current candidate pseudo-camera pose and the selected target pseudo-camera pose is greater than or equal to the preset distance threshold is executed until the number of selected target pseudo-camera poses reaches a preset number and / or all candidate pseudo-camera poses are traversed.
[0010] In one embodiment, the step of generating a pseudo-viewpoint image at the target pseudo-camera pose based on the geometric information of the current 3D Gaussian scene, the style features of the real-view image, the scene memory capability of the preset diffusion model, and the rendering result of the current 3D Gaussian scene at the target pseudo-camera pose includes: Project each Gaussian primitive in the current 3D Gaussian scene onto the image plane where the target pseudo-camera pose is located, and calculate the depth and color of each Gaussian primitive on the image plane to obtain the depth map and rendering map under the target pseudo-camera pose respectively. The depth map is used as a geometric constraint; The style features of the real-view image at a preset position relative to the target pseudo-camera pose are used as style constraints; A preset diffusion model is used as a scene memory constraint, wherein the diffusion model is a model that has been fine-tuned in advance on sparse real-view images; The rendered image is used as the starting image for the diffusion process; Based on the combined constraints of the geometric constraints, the style constraints, the scene memory constraints, and the starting image, a pseudo-viewpoint image is generated at the target pseudo-camera pose.
[0011] In one embodiment, the step of generating a pseudo-viewpoint image at the target pseudo-camera pose based on the common constraints of the geometric constraints, the style constraints, the scene memory constraints, and the starting image further includes: Based on the common constraints of the geometric constraints, the style constraints, the scene memory constraints, and the starting image, an initial pseudo-viewpoint image is generated at the target pseudo-camera pose. If there are void areas in the initial pseudo-viewpoint image, the void areas are repaired using a preset repair model to obtain the pseudo-viewpoint image.
[0012] In one embodiment, before the step of optimizing the current 3D Gaussian scene based on the supervision signal to obtain the target 3D Gaussian scene, the method further includes: Determine the image difference between the pseudo-viewpoint image and the rendered image of the current 3D Gaussian scene under the target pseudo-camera pose, as well as the geometric coverage of the current 3D Gaussian scene under the target pseudo-camera pose; The real-view image, the real camera pose, and the first pseudo-view image and its corresponding first target pseudo-camera pose are determined as supervision signals, wherein the first pseudo-view image is a pseudo-view image whose image difference is less than a preset difference threshold and whose geometric coverage is greater than a preset coverage threshold.
[0013] In one embodiment, before the step of optimizing the current 3D Gaussian scene based on the supervision signal to obtain the target 3D Gaussian scene, the method further includes: Based on the image differences, the pixel-level residual between the pseudo-view image and the rendered image is obtained, and a pixel-level uncertainty map is constructed based on the pixel-level residual. Obtain the overlay map of the current 3D Gaussian scene under the pose of the target pseudo-camera; A confidence weight map is constructed based on the coverage map and the pixel-level uncertainty map; Obtain the extrapolation distance of the target pseudo-camera pose relative to all real camera poses, and determine the exponential decay weight of the pseudo-view image if the extrapolation distance is greater than a preset extrapolation distance threshold. Based on the confidence weight map, the exponential decay weight, and the preset photometric loss function, a weighted pseudo-viewpoint supervision loss is obtained, wherein the weighted pseudo-viewpoint supervision loss includes a first loss component determined based on the pseudo-viewpoint image and its corresponding target pseudo-camera pose. The weighted pseudo-viewpoint supervision loss and the second loss component are determined as supervision signals, wherein the second loss component is determined based on the real viewpoint image and the real camera pose.
[0014] In one embodiment, the three-dimensional Gaussian reconstruction method further includes: Obtain the total number of Gaussian elements in the current 3D Gaussian scene, as well as the current iteration step and the total number of iteration steps in the preset convergence condition; The discard ratio is calculated based on the ratio of the total number of Gaussian cells to a preset density threshold and the ratio of the current iteration step to the total number of iteration steps, wherein the total number of Gaussian cells and the discard ratio are positively correlated. According to the stated discard ratio, a corresponding proportion of Gaussian elements are discarded from the current 3D Gaussian scene.
[0015] Furthermore, to achieve the above objectives, this application also proposes a three-dimensional Gaussian reconstruction device, which includes: The acquisition module is used to acquire the current three-dimensional Gaussian scene, wherein the current three-dimensional Gaussian scene is determined based on sparse real-view images and their corresponding real camera poses; The determination module is used to determine the pose of the target pseudo camera based on the real camera pose. The generation module is used to generate a pseudo-view image at the target pseudo-camera pose based on the geometric information of the current 3D Gaussian scene, the style features of the real view image, the scene memory capability of the preset diffusion model, and the rendering result of the current 3D Gaussian scene under the target pseudo-camera pose. An optimization module is used to optimize the current 3D Gaussian scene based on a supervision signal to obtain a target 3D Gaussian scene, wherein the supervision signal includes: the pseudo-view image, the target pseudo-camera pose, the real view image, and the real camera pose.
[0016] Furthermore, to achieve the above objectives, this application also proposes a three-dimensional Gaussian reconstruction device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the three-dimensional Gaussian reconstruction method as described above.
[0017] Furthermore, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the three-dimensional Gaussian reconstruction method described above.
[0018] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the three-dimensional Gaussian reconstruction method described above.
[0019] The one or more technical solutions proposed in this application have at least the following technical effects: By acquiring and optimizing the current 3D Gaussian scene obtained from sparse real-view images and corresponding real camera poses, and after determining the target pseudo-camera pose, when generating a pseudo-view image at the target pseudo-camera pose, the geometric information of the current 3D Gaussian scene, the style features of the real-view image, the scene memory capability of the preset diffusion model, and the rendering result of the current 3D Gaussian scene under the selected pseudo-camera pose are simultaneously utilized. These four constraints work together in the generation process of the diffusion model. Among them, the geometric information constrains the depth structure and 3D layout of the generated image to align with the currently reconstructed geometry; the style feature constraint ensures that the tone, lighting, and texture of the generated content are consistent with the real viewpoint; the scene memory capability constraint allows the diffusion model to remember the unique appearance details of the current scene; and the rendering result, as a draft anchor, prevents the generated result from deviating from the current reconstruction state. Through the above four complementary constraints, the originally broad degree of freedom of the diffusion model is compressed into a distribution that is highly consistent with the geometry, style, and details of the current scene, thereby generating a reliable pseudo-view image consistent with the actual geometric structure of the real scene. Subsequently, the reliable pseudo-viewpoint image and its corresponding pseudo-camera pose, along with the real viewpoint image and its corresponding real-camera pose, are used together as supervisory signals to optimize the current 3D Gaussian scene. Since the pseudo-viewpoint image already possesses geometric and appearance characteristics consistent with the current scene during the generation stage, interference from unreliable content on gradient updates is avoided, thereby effectively improving the reconstruction accuracy of the 3D Gaussian scene under sparse viewpoints. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating an embodiment of the three-dimensional Gaussian reconstruction method of this application; Figure 2 A schematic diagram of an embodiment of the three-dimensional Gaussian reconstruction method of this application Figure 1 ; Figure 3 A schematic diagram of an embodiment of the three-dimensional Gaussian reconstruction method of this application Figure 2 ; Figure 4 A schematic diagram of an embodiment of the three-dimensional Gaussian reconstruction method of this application Figure 3 ; Figure 5 This is a schematic diagram of the modular structure of the three-dimensional Gaussian reconstruction device of this application; Figure 6 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the three-dimensional Gaussian reconstruction method of this application.
[0023] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0024] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not intended to limit this application.
[0025] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0026] 3D scene reconstruction and novel perspective synthesis are fundamental problems in computer vision and computer graphics, with wide applications in virtual reality, autonomous driving, and digital content creation. 3D Gaussian Splatting (3DGS), as an emerging scene representation method, achieves real-time rendering speed through explicit anisotropic Gaussian primitives and differentiable rasterization rendering, while maintaining visual quality comparable to Neural Radiation Field (NeRF) methods.
[0027] However, high-quality reconstruction of 3D Gaussian splash relies on dense multi-view image input. In practical applications, due to the limited number and fixed positions of real cameras, only a small number (e.g., 3) of sparse viewpoint real images can be obtained. This makes it difficult to obtain sufficient geometric constraints and appearance information for spatial areas not covered by real cameras (i.e., unobserved areas), which in turn affects the quality of 3D structure reconstruction of the complete scene.
[0028] Sparse view reconstruction is essentially an ill-posed inverse problem: limited observations cannot constrain the complete 3D structure and appearance of the real scene, leading to artifacts such as depth estimation errors, floating structures, and inconsistent appearances. At the same time, the adaptive densityization mechanism of 3D Gaussian splashing relies on the gradient signal of the observed view and cannot generate meaningful geometric structures for unobserved regions.
[0029] In conventional techniques, diffusion model-based methods attempt to generate pseudo-viewpoint images using pre-trained two-dimensional diffusion models and use these pseudo-viewpoints as additional supervision signals to participate in the optimization training of three-dimensional Gaussian splashes. For example, the GS-GS method (proposed by Kong et al. in 2025) proposes an alternating optimization framework of three-dimensional Gaussian splashes and diffusion models.
[0030] However, in the aforementioned GS-GS method, the camera pose of the pseudo-viewpoint is determined only through a fixed interpolation method between existing real camera poses, which may produce redundant or uncontrollable pseudo-viewpoints. Furthermore, all generated pseudo-viewpoint images and their corresponding camera poses are directly used for scene optimization. Regardless of whether there is a significant deviation between the rendering result of the pseudo-viewpoint image and the geometry of the real scene, they participate in scene parameter updates to the same extent as real images. Geometrically inconsistent generated content may contaminate the training process. In addition, the above method uses the same Gaussian pixel dropout ratio for all scenes (i.e., complex and simple scenes), failing to adapt to the differences in geometric complexity between different scenes (e.g., the number of Gaussian pixels can differ by more than three times between different scenes).
[0031] In this case, unreliable generated content that is inconsistent with the actual geometry of the real scene will interfere with the gradient update during the optimization process, thus limiting the reconstruction accuracy of the 3D Gaussian splashing method from a sparse perspective.
[0032] This application determines candidate pseudo-camera poses based on existing real camera poses in the current 3D Gaussian scene, and actively selects the target pseudo-camera pose based on the comprehensive evaluation value corresponding to each candidate pseudo-camera pose (integrating reconstruction information gain and generation reliability). This overcomes the shortcomings of traditional methods that rely solely on fixed interpolation between real camera poses, which lack consideration for information gain and generation reliability. Then, at the selected target pseudo-camera pose, it simultaneously utilizes the geometric information of the current 3D Gaussian scene (such as a rendered depth map), the style features of the real viewpoint image, the scene-specific appearance memorized by the pre-set diffusion model after scene fine-tuning, and the rendered image of the current 3D Gaussian scene at that pose as starting anchor points. Through multi-source constraints, it generates pseudo-viewpoint images highly consistent with the scene's geometric structure and visual style, avoiding the uncontrollable content problem caused by unconstrained generation relying on the diffusion model. Furthermore, it calculates the generated pseudo-viewpoint image... The image difference between the viewpoint image and the rendered image of the current 3D Gaussian scene under the same target pseudo-camera pose, as well as the geometric coverage of the current 3D Gaussian scene under the same target pseudo-camera pose, are considered. Only pseudo-viewpoint images with image differences less than a preset difference threshold and geometric coverage greater than a preset coverage threshold, along with their corresponding target pseudo-camera poses, are used as supervision signals to optimize the current 3D Gaussian scene, along with the real viewpoint image and its real camera pose. This eliminates the pollution of the training process by unreliable generated content that is inconsistent with the geometric structure of the real scene, and overcomes the problem of geometric inconsistencies interfering with gradient updates caused by all pseudo-viewpoints participating in optimization without screening. At the same time, during the optimization process, the dropout ratio is dynamically adjusted according to the total number of Gaussian units in the current scene, so that high-density scenes receive stronger regularization and low-density scenes receive weaker regularization, overcoming the problem that a uniform dropout ratio cannot adapt to the differences in geometric complexity of different scenes. Through the aforementioned technical means, this application enables each round of optimization to be based on the pseudo-viewpoint image and its corresponding target pseudo-camera pose, which have been filtered by geometric consistency, as part of the supervision signal. This avoids interference from unreliable content and adaptively adjusts the regularization intensity, thereby significantly improving the reconstruction accuracy of the 3D Gaussian splashing method under sparse viewpoints.
[0033] It should be noted that the execution subject in this embodiment can be a 3D Gaussian reconstruction device, such as a graphics processor or central processing unit in the 3D Gaussian reconstruction device, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or embedded system capable of realizing the above functions. The execution subject in this embodiment can also be a server, workstation, dedicated computing device, or virtual computing environment. This application embodiment does not limit this. The following uses a 3D Gaussian reconstruction device as an example to describe this embodiment and the following embodiments.
[0034] Based on this, the embodiments of this application provide a three-dimensional Gaussian reconstruction method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the three-dimensional Gaussian reconstruction method of this application.
[0035] In one feasible embodiment, the application scenarios of this application can include: scene construction in virtual reality, surround view perception and unknown area completion in autonomous driving, rapid 3D modeling in digital content creation, and environmental exploration and map building in robot navigation. The output of this application is a target 3D Gaussian scene, which is composed of a large number of anisotropic Gaussian primitives. It can be used to render high-quality images from any new perspective in real time, thereby quickly generating a complete, continuous, and visually consistent 3D representation with only a small number of sparse input perspectives.
[0036] In this embodiment, the three-dimensional Gaussian reconstruction method includes steps S1 to S4: Step S1: Obtain the current 3D Gaussian scene, wherein the current 3D Gaussian scene is determined based on sparse real view images and their corresponding real camera poses; In one feasible embodiment, since only a small number (e.g., 3) of photos from different angles are often available in practical applications, directly using the 3D Gaussian Splatting (3DGS) method will result in artifacts such as depth errors and floating structures due to insufficient observation information. Therefore, this step initializes the current 3D Gaussian scene based on these limited number of real-view images and the real camera pose corresponding to each image.
[0037] Alternatively, the real-view image can be a picture taken in the actual scene using a camera or webcam, providing an observation of the scene's appearance.
[0038] Optionally, the true camera pose is the position and orientation parameters of the camera in three-dimensional space when each true viewpoint image is captured. It is usually obtained by calculating the sparse input viewpoint using the Structure from Motion (COLMAP) tool.
[0039] Optionally, the current 3D Gaussian scene is an explicit 3D representation composed of a large number of anisotropic Gaussian primitives. Each Gaussian primitive contains attributes such as position, shape, color, and opacity, and can render images from any new perspective through differentiable rasterization.
[0040] Optionally, the step of obtaining the current 3D Gaussian scene may include: obtaining a sparse 3D point cloud from the real-view image and the real camera pose; initializing the positions of Gaussian primitives in the 3D Gaussian scene to be optimized using the 3D positions of each point in the sparse 3D point cloud to obtain an initial 3D Gaussian scene; and optimizing the initial 3D Gaussian scene using the real-view image to obtain the current 3D Gaussian scene.
[0041] Optionally, initial Gaussian primitives are projected in space using the viewpoint corresponding to each real camera pose, combined with the sparse point cloud output by COLMAP. Each initial Gaussian primitive's appearance color is defined by its center position, covariance matrix, opacity parameter, and spherical harmonic coefficient. All initial Gaussian primitives together constitute the initial optimizable scene representation. The initialization process only utilizes the spatial positions of feature points in the real image and does not use pixel-level photometric information from the real viewpoint image for supervised optimization. That is, the initialization process does not use the pixel-by-pixel color difference (i.e., photometric loss) between the real viewpoint image and the image rendered by the current Gaussian scene under that real viewpoint as a supervisory signal to update the parameters of the Gaussian primitives through backpropagation to obtain the initial 3D Gaussian scene.
[0042] Optionally, the initial 3D Gaussian scene has not undergone photometric supervision optimization of the real-view image, and its geometric structure and appearance information are only in the initial state.
[0043] Optionally, a real-view image is used as a supervision signal to perform a 3D Gaussian splash optimization process on the initial 3D Gaussian scene. That is, a differentiable rasterizer renderer projects Gaussian primitives onto the real camera pose to generate a rendered image. The photometric loss between the rendered image and the corresponding real-view image is calculated, including the weighted sum of L1 loss (absolute error loss) and DSSIM loss (structural dissimilarity loss). Then, the photometric loss is backpropagated and the center position, covariance, opacity and spherical harmonic coefficient of each Gaussian primitive are updated to obtain the current 3D Gaussian scene.
[0044] Understandably, this step, using very limited real photographs and pose information, builds a rough but basic geometric prototype of a 3D scene, which serves as the starting point for subsequent pseudo-viewpoint generation and iterative optimization.
[0045] Step S2: Determine the target pseudo-camera pose based on the real camera pose; In one feasible embodiment, when training using only the real viewpoint, the uncaptured areas are completely unknown, thus requiring the selection of the most valuable virtual camera positions to generate additional training images. This step determines the target pseudo-camera pose for which pseudo-viewpoint images need to be generated based on the existing real camera poses.
[0046] Step S3: Based on the geometric information of the current 3D Gaussian scene, the style features of the real view image, the scene memory capability of the preset diffusion model, and the rendering result of the current 3D Gaussian scene under the pose of the target pseudo camera, generate a pseudo view image at the pose of the target pseudo camera. In one feasible embodiment, directly using an existing diffusion model to generate pseudo-viewpoint images can easily result in geometric distortions or stylistic deviations, which can contaminate the training. Therefore, this step comprehensively utilizes four complementary constraints to generate pseudo-viewpoint images of controllable quality at the selected target pseudo-camera pose.
[0047] Alternatively, geometric information can be obtained by calculating a monocular depth map of the true viewpoint using the Depth-Anything V2 depth estimation model.
[0048] Optionally, the style features are taken from the real training image closest to the target pseudo-camera pose, including appearance attributes such as hue, lighting, and texture. They can be encoded into a 1024-dimensional style feature vector by the IP-Adapter module of the CLIP (Contrastive Language-Image Pre-training) ViT-H / 14 image encoder.
[0049] Optionally, the scene memory capability of the diffusion model refers to inserting a low-rank adaptation module (LoRA) into the attention layer (query, key, value, and output projection) of the UNet (a neural network with an encoder-decoder structure) of the SDXL (Stable Diffusion XL) to enable the model to remember the appearance details unique to the current scene.
[0050] Optionally, the rendering result of the current Gaussian scene under the target pseudo-camera pose includes an RGB (red-green-blue) image (i.e., a rendering map), a depth map, and an alpha (opacity) overlay map, where the RGB image serves as the starting point of the diffusion process (i.e., draft anchoring).
[0051] Optionally, four constraints are applied simultaneously when generating pseudo-viewpoint images: ControlNet (a neural network architecture) depth constraints, IP-Adapter style constraints, LoRA scene memory, and img2img (Image-to-Image) draft anchoring.
[0052] Understandably, this step compresses the free generation space of the diffusion model into a narrow channel that matches the actual geometry and style of the scene through four complementary constraints, thereby obtaining a geometrically correct and style-consistent pseudo-viewpoint image.
[0053] Step S4: Based on the supervision signal, optimize the current 3D Gaussian scene to obtain the target 3D Gaussian scene, wherein the supervision signal includes: the pseudo-view image, the target pseudo-camera pose, the real view image, and the real camera pose.
[0054] In one feasible embodiment, the quality of the generated pseudo-viewpoint images varies, and low-quality pseudo-viewpoints can interfere with training; at the same time, the geometric complexity of different scenes varies greatly, and a uniform regularization intensity cannot be adapted to all situations.
[0055] Optionally, the supervision signal includes four types of data: pseudo-view image and its corresponding target pseudo-camera pose, real view image and its real camera pose; in the alternating optimization framework, the 3D Gaussian reconstruction device samples from the pseudo-view pool with a preset probability and uses it alternately with the real view.
[0056] Optionally, uncertainty-weighted loss can be used for optimization.
[0057] Understandably, this step, while ensuring training stability, makes full use of the additional information provided by the pseudo-viewpoint, and ultimately converges to obtain a high-quality 3D Gaussian scene with geometric correctness and consistent appearance.
[0058] In this embodiment, by acquiring and optimizing the current 3D Gaussian scene obtained from sparse real-view images and corresponding real camera poses, and after determining the target pseudo-camera pose, when generating a pseudo-view image at the target pseudo-camera pose, the geometric information of the current 3D Gaussian scene, the style features of the real-view image, the scene memory capability of the preset diffusion model, and the rendering result of the current 3D Gaussian scene under the selected pseudo-camera pose are all utilized. These four constraints work together in the generation process of the diffusion model. Among them, the geometric information constraint aligns the depth structure and 3D layout of the generated image with the currently reconstructed geometry; the style feature constraint ensures that the tone, lighting, and texture of the generated content are consistent with the real view; the scene memory capability constraint allows the diffusion model to remember the unique appearance details of the current scene; and the rendering result, as a draft anchor, prevents the generated result from deviating from the current reconstruction state. Through the above four complementary constraints, the originally broad degree of freedom of the diffusion model is compressed into a distribution that is highly consistent with the geometry, style, and details of the current scene, thereby generating a reliable pseudo-view image consistent with the actual geometric structure of the real scene. Subsequently, the reliable pseudo-viewpoint image and its corresponding pseudo-camera pose, along with the real viewpoint image and its corresponding real-camera pose, are used together as supervisory signals to optimize the current 3D Gaussian scene. Since the pseudo-viewpoint image already possesses geometric and appearance characteristics consistent with the current scene during the generation stage, interference from unreliable content on gradient updates is avoided, thereby effectively improving the reconstruction accuracy of the 3D Gaussian scene under sparse viewpoints.
[0059] Based on the first embodiment described above, a second embodiment of the three-dimensional Gaussian reconstruction method of this application is proposed. In this embodiment, step S2, the step of determining the target pseudo-camera pose based on the real camera pose, includes: Step A2: Determine candidate pseudo-camera poses based on the existing real camera poses in the current 3D Gaussian scene; In one feasible embodiment, in actual sparse view reconstruction, the few real camera poses are distributed over a limited number of angles in the scene, and it is necessary to find new pseudo-views between and around these poses to supplement the information.
[0060] Optionally, directly selecting points between adjacent real poses ignores the differences in viewpoint information and generation reliability, resulting in either redundant pseudo-camera poses or poses exceeding the generation capabilities of the current diffusion model. Therefore, this step constructs a set of candidate pseudo-camera poses based on existing real camera poses for subsequent selection.
[0061] Optionally, the candidate pseudo-camera pose is a virtual camera position and orientation parameter calculated by interpolation and extrapolation based on the real camera pose.
[0062] Step A3: Select the target pseudo camera pose from the candidate pseudo camera poses based on the comprehensive evaluation value corresponding to the candidate pseudo camera poses.
[0063] In one feasible embodiment, since the candidate pseudo-camera pose pool contains a large number of potentially redundant, uncontrollable, or even exceeding current reconstruction capabilities viewpoints, it is necessary to quantitatively evaluate the value of each candidate pseudo-camera pose in order to select the pseudo-camera pose with the highest information gain and the most reliable generation, thus avoiding low-quality pseudo-viewpoints from contaminating the training. This step calculates the comprehensive evaluation value of each candidate pseudo-camera pose and filters and sorts them according to the comprehensive evaluation value to determine the target pseudo-camera pose used for actually generating pseudo-viewpoint images.
[0064] Optionally, the comprehensive evaluation value is obtained through a joint scoring function. The calculation result can be obtained from the function as follows: ,in Represents the pose of candidate pseudo-cameras. This indicates the reconstructed information gain that can be obtained by acquiring new observations under the candidate pseudo-camera pose. This indicates the reliability of generating a pseudo-viewpoint image using the current diffusion model at this pose. It is the annealing index, used to adjust the relative weight between information gain and reliability.
[0065] Understandably, this step selects the target pseudo-camera pose from the candidate pseudo-camera poses based on the comprehensive evaluation value, which is then used for subsequent pseudo-viewpoint generation and scene optimization.
[0066] In this embodiment, candidate pseudo-camera poses are first constructed by interpolation and extrapolation based on existing real camera poses. Then, each candidate pseudo-camera pose is quantitatively scored according to a comprehensive evaluation value (i.e., the weighted product of reconstruction information gain index and generation reliability index, where the annealing index dynamically adjusts the weights of the two). The candidate pseudo-camera pose with the highest information gain and the most reliable generation is selected as the target pseudo-camera pose based on the score. This method solves the problem of redundant pseudo-camera pose information or exceeding the generation capacity of the diffusion model caused by directly using a fixed interpolation method in sparse view reconstruction, while avoiding low-quality pseudo-views from polluting the training process. By expanding the generation range of candidate pseudo-camera poses from simple neighboring points to multi-level extrapolation regions, and combining joint scoring and ranking selection guided by the annealing index, the active selection and adaptive scheduling of the target pseudo-camera pose are realized. Thus, this embodiment achieves the technical effects of improving the quality of pseudo-view image generation and training stability, and enhancing the geometric integrity and appearance consistency of 3D Gaussian scene reconstruction under sparse view.
[0067] In one feasible implementation, step A2, which involves determining candidate pseudo-camera poses based on the existing real camera poses in the current 3D Gaussian scene, includes: Step A21: Generate a pose pair based on the existing real camera poses in the current 3D Gaussian scene, wherein the pose pair includes two different real camera poses. In one feasible embodiment, in order to generate new virtual camera positions between and around sparse real viewpoints, it is necessary to pair existing real camera poses in pairs for subsequent determination of candidate pseudo camera poses.
[0068] Optionally, a pose pair is a combination of two different poses selected from all existing real camera poses in the current 3D Gaussian scene. Each pose contains a rotation component (representing the camera's orientation) and a translation component (representing the camera's position).
[0069] Optionally, the pose pairs can be generated by iterating through all existing real camera poses and combining any two different poses in sequence to form multiple non-overlapping pose pairs. For example, when there are three real camera poses A, B, and C, the generated pose pairs include (A,B), (A,C), and (B,C).
[0070] Understandably, by pairing real cameras in pairs, a clear geometric reference relationship is established for generating interpolation and extrapolation candidate pseudo-camera poses within the spatial region defined by each pose pair.
[0071] Step A22: Perform spherical linear interpolation on the rotation components of the pose in the pose pair to obtain the first interpolation result; In one feasible embodiment, the rotational change between the poses of the two real cameras is not a simple linear transition, but a smooth change along the shortest path on the sphere, using spherical linear interpolation to calculate the intermediate orientation.
[0072] Alternatively, Spherical Linear Interpolation (SLERP) can smoothly interpolate rotations, ensuring a constant rotational angular velocity and the shortest path during the interpolation process.
[0073] Optionally, for the rotation components of the two poses in the pose pair (typically converted to quaternion form), under given interpolation coefficients... (Within the range of 0 to 1), the intermediate rotation state is calculated, and this intermediate state is the first interpolation result.
[0074] Optionally, The value of determines whether the first interpolation result is closer to the first pose or the second pose. The result is the rotation of the first pose. The result is the rotation of the second pose. The result is the intermediate rotation between the two.
[0075] Understandably, this step generates intermediate rotation values that smoothly transition between the two real camera orientations, used to construct the interpolated pseudo-camera pose.
[0076] Step A23: Perform linear interpolation on the translation components of the pose in the pose pair to obtain a second interpolation result; In one feasible embodiment, the movement of the camera position in three-dimensional space can be approximated by a straight path, and the intermediate position is calculated using linear interpolation.
[0077] Optionally, the same interpolation coefficients as in step A22 are used for the translation components (i.e., position coordinates in three-dimensional space, each component represented as x, y, z values) of the two poses in the pose pair. In this case, follow the formula below: , to perform calculations, where and These are the translation vectors for the two poses, and the calculated three-dimensional coordinates are the second interpolation result.
[0078] Optionally, when The result is the position of the first pose. The result is the position of the second pose. The result is the midpoint of the line connecting the two points.
[0079] Understandably, this step generates the midpoint of the straight path between the two real camera positions through linear interpolation, which, together with the rotation component in the first interpolation result, constitutes the pose parameters of the interpolated pseudo-camera pose.
[0080] Step A24: Based on the first interpolation result and the second interpolation result, generate an interpolated pseudo-camera pose located between the two poses in the pose pair; In one feasible embodiment, the calculated rotation interpolation results and translation interpolation results are combined to obtain candidate pseudo-camera poses located between two real camera poses. These interpolated pseudo-camera poses can fill the observation gap between real viewpoints.
[0081] Optionally, the interpolated pseudo-camera pose is a pseudo-camera pose whose spatial position is located on the line connecting the two real camera poses, determined by simultaneously using rotational interpolation and translational interpolation between the two poses.
[0082] Optionally, in Take multiple values evenly from the range of 0 to 1 (e.g.) =0.2, 0.4, 0.6, 0.8), respectively generating corresponding interpolated pseudo-camera poses, thereby obtaining a large number of smoothly transitioning pseudo-viewpoints between every two real cameras.
[0083] Understandably, this step generates continuous virtual observation positions along the lines connecting real cameras by combining the interpolation results of rotation and translation, in order to cover the unobserved areas between real viewpoints.
[0084] Step A25: Extrapolate the rotation and translation components of each pose in the pose pair along both ends of the line connecting the two poses in the pose pair to generate an extrapolated pseudo camera pose located outside the line connecting the two poses in the pose pair. In one feasible embodiment, interpolation between real cameras can only cover the area inside the existing viewpoint, but cannot reach the unobserved area outside the real cameras. Therefore, it is necessary to extend outward to generate an extrapolated pose located on the extension line connecting the two real cameras in order to explore a wider spatial range.
[0085] Optionally, the extrapolated pseudo-camera pose is a candidate pseudo-camera pose obtained by extending a certain distance outward along the extension line connecting the two real camera poses based on the changing trends of the two real camera poses, and is located outside the line connecting the two real camera poses.
[0086] Optionally, the extrapolated pseudo-camera pose is generated as follows: For the rotation component, the following extrapolation formula is used: ,in, For interpolation coefficients, less than 0 or greater than 1 (e.g., -0.2, +1.2, etc.), calculate the extrapolated rotation matrix; for translation components, use the following extrapolation formula: T ,in, The interpolation coefficient is less than 0 or greater than 1; the extrapolated translation is calculated; the calculated extrapolated rotation and extrapolated translation are combined to obtain an extrapolated pseudo-camera pose.
[0087] Optionally, the interpolation coefficients during extrapolation The value can be set according to the extrapolation level: Level 0 Take [-0.2, 1.2], Level 1 Take [0.5, 1.5], Level 2 The range [-1.0, 2.0] corresponds to different levels of extrapolation.
[0088] Understandably, this step extends the coverage of the virtual camera to the outer regions on both sides of the line connecting the real cameras through extrapolation, providing multiple pseudo-perspectives for exploring unknown geometric structures beyond the actual observation range.
[0089] Step A26: The interpolated pseudo-camera pose and the extrapolated pseudo-camera pose are determined as candidate pseudo-camera poses.
[0090] In one feasible embodiment, all candidate pseudo-camera poses generated by the above two methods are aggregated to form a candidate set, which is used for subsequent screening of target pseudo-camera poses.
[0091] Optionally, candidate pseudo-camera poses include all interpolated pseudo-camera poses and all extrapolated pseudo-camera poses.
[0092] Optionally, the candidate pseudo camera pose can be determined as follows: for each real camera pose pair, generate the corresponding interpolated pseudo camera pose and extrapolated pseudo camera pose respectively, then merge all the interpolated pseudo camera poses and extrapolated pseudo camera poses generated by all pose pairs, and remove duplicate or overly similar poses to obtain the candidate pseudo camera pose.
[0093] Understandably, this step constructs candidate pseudo-camera poses covering the regions between and outside the real cameras by aggregating the pseudo-viewpoints generated by interpolating and extrapolating pseudo-camera poses.
[0094] In this embodiment, if only simple interpolation between real camera poses is relied upon, the generated candidate viewpoints will be limited to the interior of the existing observation area, failing to explore unobserved areas outside the distribution range of real cameras, thus limiting the ability of pseudo-viewpoints to extrapolate the scene. Therefore, this embodiment performs interpolation and extrapolation on each pair of real cameras to cover candidate pseudo-camera poses ranging from conservative interpolation regions to progressively expanding regions.
[0095] In this embodiment, existing real camera poses are paired, and spherical linear interpolation is performed on the rotation component and linear interpolation on the translation component to generate interpolated pseudo-camera poses. Then, rotational exponential mapping extrapolation and translational linear extrapolation are performed outwards along the poses to generate extrapolated pseudo-camera poses. Finally, the interpolation and extrapolation results are merged to form candidate pseudo-camera poses. This method solves the problem that relying solely on simple interpolation between real cameras limits candidate viewpoints to the existing observation area and fails to explore unobserved areas outside the real camera distribution range. Furthermore, by setting multiple extrapolation levels (level 0 to level 2), it achieves spatial coverage of candidate viewpoints from conservative interpolation to progressively expanding extrapolation. Therefore, this embodiment achieves the technical effect of providing rich candidate pseudo-camera poses, including internal filling views and external exploration views, for subsequent active viewpoint selection, and improving the extrapolation capability and geometric integrity in sparse viewpoint 3D reconstruction.
[0096] In one feasible implementation, step A3, which involves selecting the target pseudo-camera pose from the candidate pseudo-camera poses based on the comprehensive evaluation value corresponding to the candidate pseudo-camera poses, includes: Step A31: Obtain the reconstruction information gain index and generation reliability index of the poses of each candidate pseudo camera; In one feasible embodiment, each pseudo-viewpoint in the candidate pseudo-camera pose has different information value and generation quality, and needs to be quantified separately for comprehensive evaluation.
[0097] Optionally, the reconstruction information gain metric is used to measure the information increment that new observations acquired at the candidate pseudo-camera pose can bring to the current 3D Gaussian scene. Specifically, it includes three sub-metrics: alpha sparsity (reflecting the sparsity of Gaussian primitive projection at the candidate pseudo-camera pose), rendering entropy (reflecting the uncertainty of the rendering result at the candidate pseudo-camera pose), and viewpoint difference (reflecting the directional difference between the pseudo viewpoint corresponding to the candidate pseudo-camera pose and the existing real viewpoint).
[0098] Optionally, the generation reliability metric is used to measure the confidence level of generating pseudo-viewpoint images at the candidate pseudo-camera pose using the diffusion model at the current training progress. Specifically, it includes model confidence (the diffusion model's own estimated confidence score for the pose generation result) and historical generation quality (the generation success rate of the region near the candidate pseudo-camera pose in previous diffusion stages).
[0099] Optionally, for each candidate pseudo-camera pose pool, the value of the reconstruction information gain index is calculated based on the parameters of the Gaussian elements in the current 3D Gaussian scene and the existing real camera pose distribution; at the same time, the value of the generation reliability index is calculated based on the model confidence of the diffusion model and the historical generation quality.
[0100] Step A32: Calculate the comprehensive evaluation value of each candidate pseudo-camera pose based on the preset annealing index, the reconstruction information gain index, and the generation reliability index. In one feasible embodiment, poses with high information gain may generate low reliability, while generating reliable poses may result in information redundancy. A dynamic trade-off needs to be made between the two, so an annealing index is introduced to adjust the weights of the two.
[0101] Optionally, the annealing index (denoted as ) ) is an exponential parameter that changes dynamically with the number of training iterations and is used to adjust the influence of the generation reliability index in the overall evaluation.
[0102] Optionally, when calculating the comprehensive evaluation value of each candidate pseudo-camera pose, the value of β is determined according to the current training progress: early training (e.g., the first 5000 iterations). Setting it to 1.5 gives a higher weight to generation reliability, prioritizing poses that the diffusion model is confident in generating; in the later stages of training... The weight is linearly reduced to 0.4, decreasing the weight of generation reliability and giving more weight to poses with higher information gain.
[0103] Understandably, this step dynamically adjusts the proportion of information gain and generation reliability in the overall evaluation value through the annealing index, so that stable poses are conservatively selected in the early stage of training and high-value poses are actively explored in the later stage of training.
[0104] Step A33: Sort the poses of each candidate pseudo camera according to the comprehensive evaluation value to obtain the sorting result; In one feasible embodiment, after calculating the comprehensive evaluation value of all candidate pseudo-camera poses, they can be arranged according to their numerical values so that the optimal poses can be selected from them later.
[0105] Optionally, all poses in the candidate pseudo-camera pose pool can be arranged in descending (or ascending) order of their overall evaluation value, with the poses having the higher overall evaluation value being more valuable in this training phase.
[0106] Step A34: Determine the target pseudo-camera pose from the candidate pseudo-camera poses according to the sorting results.
[0107] In one feasible embodiment, the candidate pseudo-camera poses with the highest comprehensive evaluation values are selected for the actual generation of pseudo-viewpoint images. At the same time, the selected candidate pseudo-camera poses need to be spatially dispersed to avoid the generated pseudo-viewpoints being overly concentrated.
[0108] Optionally, the target pseudo-camera pose is a candidate pseudo-camera pose selected from the candidate pseudo-camera poses and used to actually generate the pseudo-viewpoint image.
[0109] Optionally, the target pseudo-camera pose is determined as follows: Candidate pseudo-camera poses are traversed sequentially from highest to lowest score according to the sorting results. Constraints can also be imposed during selection: First, a minimum pose distance constraint is imposed, meaning the difference in rotation angle and translation distance between the current candidate pseudo-camera pose and any already selected target pseudo-camera pose must both exceed a preset threshold (e.g., a rotation angle difference greater than 15 degrees and a translation distance difference greater than 0.1 times the scene radius). If the difference is less than the threshold, the current candidate pseudo-camera pose is skipped. Second, a fixed number (up to 2) of candidate pseudo-camera poses are selected between each pair of real cameras to prevent oversampling of the region between a pair of cameras. This process is repeated until a fixed total number (up to 16) of target pseudo-camera poses are selected, or all candidate pseudo-camera poses have been traversed.
[0110] In this embodiment, the reconstruction information gain index and generation reliability index of each candidate pseudo-camera pose are first obtained. Then, a comprehensive evaluation value is calculated based on the annealing index, which dynamically changes with the training progress. The pseudo-camera poses are then sorted from largest to smallest. Finally, the sorted results are iterated sequentially, with minimum pose distance constraints and a maximum of a preset number of real cameras selected per pair applied, until a preset number of target pseudo-camera poses are selected or all candidate pseudo-camera poses have been traversed. This method solves the problems of balancing information gain and generation reliability, the inability of fixed thresholds to adapt to changes in the training phase, and viewpoint redundancy caused by spatial clustering of target pseudo-camera poses. By combining dynamic weight scheduling guided by the annealing index with spatial dispersion constraints, this embodiment achieves the technical effect of prioritizing stable and reliable poses in the early stage of training to ensure training stability, actively exploring high information gain poses in the later stage of training to improve reconstruction integrity, and ensuring uniform spatial distribution of selected target pseudo-camera poses, thereby maximizing the pseudo-viewpoint supervision effect.
[0111] In one feasible implementation, step A34, determining the target pseudo-camera pose from the candidate pseudo-camera poses based on the sorting result, includes: Step A341: Determine the pose of the current candidate pseudo camera based on the sorting results; In one feasible embodiment, since all candidate pseudo-camera poses have been sorted according to the comprehensive evaluation value, it is possible to determine in turn whether each candidate pseudo-camera pose is suitable to be selected as the target pseudo-camera pose.
[0112] Optionally, the current candidate pseudo-camera pose is the candidate pseudo-camera pose that is being evaluated and extracted sequentially according to the sorting results.
[0113] Optionally, the current candidate pseudo-camera pose can be determined as follows: starting from the first candidate pseudo-camera pose in the sorting results, where the first candidate pseudo-camera pose can be the candidate pseudo-camera pose with the highest comprehensive evaluation value, and this is taken as the first current candidate pseudo-camera pose; subsequently, each time an evaluation is completed and the next candidate pseudo-camera pose needs to be determined, the candidate pseudo-camera pose that is one position after the current candidate pseudo-camera pose in the sorting results is taken as the new current candidate pseudo-camera pose.
[0114] Optionally, if no target pseudo-camera pose is selected, the current candidate pseudo-camera pose is determined as the target pseudo-camera pose.
[0115] Step A342: Determine whether the pose distance between the current candidate pseudo camera pose and the selected target pseudo camera pose is greater than or equal to a preset distance threshold. In one feasible embodiment, in order to ensure that the selected target pseudo camera poses are evenly distributed in space and do not generate redundant viewpoints, it is necessary to compare the spatial distance between the currently determined candidate pseudo camera poses and the already selected pseudo camera poses used as targets.
[0116] Optionally, the pose distance is the degree of difference between two camera poses in rotational orientation and translational position, which may include the angular difference between rotational components (in degrees) and the Euclidean distance between translational components (in units related to scene size).
[0117] Optionally, the preset distance threshold is a pre-set value used to determine whether the poses of the two cameras are sufficiently dispersed. For example, the threshold for the difference in rotation angle can be set to 15 degrees, and the threshold for the translation distance can be set to 0.1 times the scene radius.
[0118] Optionally, for the current candidate pseudo-camera pose, calculate the rotation angle difference and translation Euclidean distance between the current candidate pseudo-camera pose and each selected target pseudo-camera pose. When the rotation angle difference and translation Euclidean distance are both greater than or equal to their respective thresholds, the pose distance is determined to be greater than or equal to a preset distance threshold; otherwise (i.e., the rotation angle difference is less than 15 degrees, or the translation Euclidean distance is less than 0.1 times the scene radius, or both are less than the corresponding thresholds), it is determined to be less than the preset distance threshold.
[0119] Understandably, this step filters out candidate pseudo-camera poses that are spatially sufficiently dispersed from each of the selected target pseudo-camera poses by quantifying and comparing pose distances.
[0120] Step A343: If the pose distance between the current candidate pseudo camera pose and each of the selected target pseudo camera poses is greater than or equal to a preset distance threshold, the current candidate pseudo camera pose is determined as the target pseudo camera pose. In one feasible embodiment, if the pose distance between the current candidate pseudo camera pose and each of the already selected target pseudo camera poses reaches a preset distance threshold, it means that the current candidate pseudo camera pose can provide new, non-repeating viewpoint information and can be used as the target pseudo camera pose.
[0121] Optionally, each selected target pseudo-camera pose is the pose of all selected target pseudo-cameras.
[0122] Optionally, the current candidate pseudo-camera pose is determined as the target pseudo-camera pose, and the target pseudo-camera pose is subsequently used to generate pseudo-viewpoint images.
[0123] Understandably, this step ensures that the selected target pseudo-camera poses are kept at a sufficient distance to avoid excessive concentration of viewpoints.
[0124] Step A344: If the attitude distance between the current candidate pseudo-camera pose and at least one selected target pseudo-camera pose is less than a preset distance threshold, the candidate pseudo-camera pose that is one position after the current candidate pseudo-camera pose in the sorting result is determined as the new current candidate pseudo-camera pose. Then, the step of determining whether the attitude distance between the current candidate pseudo-camera pose and the selected target pseudo-camera pose is greater than or equal to the preset distance threshold is executed until the number of selected target pseudo-camera poses reaches a preset number and / or all candidate pseudo-camera poses are traversed.
[0125] In one feasible embodiment, if the current candidate pseudo-camera pose is too close to any of the selected target pseudo-camera poses (i.e., the pose distance is less than a preset distance threshold), the current candidate pseudo-camera pose is considered redundant, and the next candidate pseudo-camera pose in the sorting result is determined instead. The above distance judgment and subsequent steps are repeated until the required number of target pseudo-camera poses are selected, and / or all candidate pseudo-camera poses are determined.
[0126] Optionally, determining the candidate pseudo-camera pose that is one position after the current candidate pseudo-camera pose in the sorting results as the new current candidate pseudo-camera pose means moving down one position in the sorting results and taking the next candidate pseudo-camera pose with a slightly lower comprehensive evaluation value as the new judgment object.
[0127] Optionally, performing the judgment step means returning to step A342 and re-comparing the pose distance for the new current candidate pseudo-camera pose.
[0128] Optionally, the preset number is the total number of target pseudo-camera poses to be selected, which can be 16.
[0129] Optionally, traversing all candidate pseudo-camera poses means that all candidate pseudo-camera poses in the sorting results have been determined.
[0130] Optionally, the above process is repeated until one or both of the following conditions are met: the number of selected target pseudo-camera poses has reached a preset number (e.g., 16), or although the preset number has not been reached, all candidate pseudo-camera poses have been traversed (i.e., there are no more candidate pseudo-camera poses available for determination).
[0131] Understandably, this step iteratively skips low-value or redundant candidate pseudo-camera poses and sequentially selects a preset number of target pseudo-camera poses that are evenly distributed and have high comprehensive evaluation values, while ensuring spatial dispersion.
[0132] In this embodiment, candidate pseudo-camera poses are sorted by their comprehensive evaluation values and then sequentially selected as the current candidate pseudo-camera poses. When at least one target pseudo-camera pose is available, the pose distance between the current candidate and all selected poses is checked against a preset distance threshold. Only when all pose distances meet the threshold requirement is the current candidate adopted as the target pseudo-camera pose; otherwise, the current pose is skipped and the next candidate in the sorting results is evaluated. This process is repeated until a preset number of pseudo-camera poses are selected or all candidate pseudo-camera poses have been traversed. This solves the problems of viewpoint spatial clustering, redundancy, and lack of diversity constraints caused by directly selecting poses based solely on their scores. By introducing a pose distance judgment and iterative skipping strategy, this embodiment ensures that the comprehensive evaluation value of the selected target pseudo-camera poses is as high as possible while forcing spatial dispersion among the selected target pseudo-camera poses. This achieves the technical effects of avoiding the repeated generation of pseudo-viewpoint images at similar viewpoints, maximizing the information coverage of new viewpoints, and improving overall 3D reconstruction efficiency and geometric integrity.
[0133] Based on any of the above embodiments, a third embodiment of the 3D Gaussian reconstruction method of this application is proposed. In the previous embodiments, the most valuable target pseudo-camera pose has been determined through active viewpoint selection. However, if a general pre-trained diffusion model is directly used to generate a pseudo-viewpoint image at the target pseudo-camera pose, the generated result often suffers from geometric distortion, style deviation from the current scene, or loss of detail due to insufficient constraints caused by sparse viewpoint input. To this end, this embodiment adopts a multi-source constraint method, using the geometric information rendered in the current 3D Gaussian scene, the style features of the real viewpoint image, the scene memory of the fine-tuned diffusion model, and the current rendered image as a draft anchor. The four constraints work together to compress the generation freedom of the diffusion model to a range close to the real distribution of the scene, thereby generating a pseudo-viewpoint image with good geometric consistency. In this embodiment, step S3, the step of generating a pseudo-viewpoint image at the target pseudo-camera pose based on the geometric information of the current 3D Gaussian scene, the style features of the real viewpoint image, the scene memory capability of the preset diffusion model, and the rendering result of the current 3D Gaussian scene at the target pseudo-camera pose, includes: Step S31: Project each Gaussian primitive in the current 3D Gaussian scene onto the image plane where the target pseudo-camera pose is located, and calculate the depth and color of each Gaussian primitive on the image plane to obtain the depth map and rendering map under the target pseudo-camera pose respectively. In one feasible embodiment, in order to obtain the geometric information and initial appearance of the target pseudo-camera pose, it is necessary to perform differentiable rasterization rendering using an existing 3D Gaussian scene.
[0134] Optionally, based on the rotation and translation parameters of the target pseudo-camera pose, the center coordinates of each Gaussian cell are transformed from the world coordinate system to the camera coordinate system, and then mapped to the pixel position on the two-dimensional image plane where the target pseudo-camera pose is located through perspective projection.
[0135] Optionally, the projected Gaussian elements are sorted on the image plane according to the distance (i.e., depth value) from the Gaussian element to the optical center of the camera. The depth value of the Gaussian element with the smallest depth is taken as the depth of the pixel, and the depth values of all pixels constitute a depth map.
[0136] Optionally, the final RGB color value of each pixel is calculated on the image plane based on the color attributes and opacity of the Gaussian pixels through alpha blending (i.e., weighted superposition from front to back), and the color values of all pixels constitute the rendered image.
[0137] Understandably, this step generates a depth map and a rendering map at the target pseudo-camera pose using 3D Gaussian splashing differentiable rendering, which serve as the geometric conditions (depth map) and initial appearance anchor points (rendering map) in the subsequent quadruple constraints.
[0138] Step S32: Use the depth map as a geometric constraint; In one feasible embodiment, the diffusion model is prone to generating content that is inconsistent with the actual geometry of the scene when generating images. This can be addressed by forcibly constraining the depth distribution of the generated result to match the depth rendered in the current Gaussian scene. Figure 1 This ensures the geometric correctness of the pseudo-viewpoint image.
[0139] Optionally, geometric constraints utilize depth maps as control conditions to guide the generation process of the diffusion model, ensuring that the object shapes, occlusion relationships, and spatial layouts in the generated images remain consistent with the geometric structure of the current 3D Gaussian scene at the target pseudo-camera pose.
[0140] Optionally, one implementation of using the depth map as a geometric constraint can be: normalizing the depth map, scaling its numerical range to the [0,1] interval, and then inputting it into a depth ControlNet model based on SDXL; the ControlNet model injects depth residual features into each layer of the UNet, forcing the latent features of the diffusion model to align with the spatial structure of the depth map during the denoising process.
[0141] Understandably, this step uses the strong geometric guidance of deep ControlNet to anchor the generated content of the diffusion model onto the existing geometric structure of the current Gaussian scene, thus avoiding geometric distortion.
[0142] Step S33: Use the style features of the real viewpoint image at a preset position of the target pseudo-camera pose as a style constraint; In one feasible embodiment, geometric constraints alone can only guarantee the correctness of the structure, but cannot guarantee that the color, lighting and texture style of the generated image are consistent with other perspectives in the real scene. Therefore, style constraints need to be introduced.
[0143] Optionally, the preset position is a pre-defined distance range, which can usually be selected as the real view image closest to the target pseudo camera pose. The closest can be the leftmost or the rightmost target pseudo camera pose, and one or more real view images closest to the target pseudo camera pose can be selected.
[0144] Optionally, style features are high-dimensional feature representations of visual attributes such as hue, lighting, texture, and material of a real-view image.
[0145] Optionally, the specific implementation of using the style features of the real-view image as a style constraint is as follows: use an image encoder (such as CLIP ViT-H / 14) to encode the real-view image closest to the target pseudo-camera pose, and extract a 1024-dimensional style feature vector; input the style feature vector into the IP-Adapter module, and inject it into the generation process of the diffusion model through the cross-attention layer of UNet, so that the style of the generated image is aligned with the style of the real-view image.
[0146] Understandably, this step, through style feature injection, ensures that the pseudo-viewpoint image maintains visual consistency with the real observed scene in appearance, thus avoiding style drift.
[0147] Step S34: Use a preset diffusion model as a scene memory constraint, wherein the diffusion model is a model that has been fine-tuned in advance on sparse real-view images. In one feasible embodiment, the general pre-trained diffusion model does not know the detailed features of the current scene, such as the unique pattern of a texture or the specific shape of an object, so the diffusion model needs to remember the unique appearance features of the current scene.
[0148] Optionally, the preset diffusion model can be an SDXL model that has been fine-tuned for scene adaptation.
[0149] Optionally, the scene memory constraint means that the fine-tuned diffusion model has been trained with a small number of sparse real-view images, storing the appearance details unique to the current scene in the model parameters, and will automatically prefer to generate content that conforms to the details of the current scene during the generation process.
[0150] Optionally, the fine-tuning on sparse real-view images can be done by inserting a low-rank adaptation module into the attention layer (including four linear layers: query, key, value, and output projection) of the SDXL UNet. The rank can be set to 4, and only the inserted LoRA parameters are optimized. After the warm-up phase, the LoRA module is fine-tuned for 50 steps using sparse real-view images so that the diffusion model learns the appearance features of the current scene. After fine-tuning, the number of LoRA parameters accounts for only about 0.06% of the total parameters of the UNet.
[0151] The preset diffusion model, after fine-tuning, can have its parameters stored for later reuse.
[0152] Understandably, this step uses LoRA fine-tuning to solidify the scene memory into the diffusion model, making the generated results more closely match the current scene at the level of detail, rather than a general diffusion model.
[0153] Step S35: Use the rendered image as the starting image for the diffusion process; In one feasible embodiment, if the generation starts entirely from random noise, the diffusion model may produce content that differs too much from the current Gaussian scene rendering result, leading to unstable optimization. Therefore, it is necessary to use the current scene rendering as the starting point to limit the deviation of the generated result.
[0154] Optionally, the starting image refers to the initial input image of the diffusion process, used to replace pure random noise as the starting point for denoising.
[0155] Optionally, the specific implementation of using the rendered image as the starting image of the diffusion process can be as follows: the rendered image obtained in step S31 is used as the initial input of the img2img mode (image-to-image generation mode, a generation mode of the diffusion model), and Gaussian noise with an intensity of 30% is added to the rendered image to obtain the noisy latent representation; then, a partial denoising process is performed from this noisy state, wherein the partial number of steps can usually be set to 30% to 50% of the total number of denoising steps, rather than starting from pure random noise for complete denoising.
[0156] Understandably, this step preserves the rough structure and spatial layout of the rendered image, ensuring that the constraint results do not deviate excessively from the current state of the Gaussian scene during the diffusion generation process.
[0157] Step S36: Based on the common constraints of the geometric constraints, the style constraints, the scene memory constraints, and the starting image, generate a pseudo-viewpoint image at the target pseudo-camera pose.
[0158] In one feasible embodiment, the above four constraints are simultaneously applied to the generation process of the diffusion model, compressing the generation degrees of freedom into a narrow channel that is close to the real distribution of the scene, thereby obtaining a pseudo-viewpoint image that is geometrically consistent, style-matched, and detailed.
[0159] Optionally, the common constraint is to apply the depth map geometric constraints, real-view style constraints, LoRA scene memory constraints, and img2img starting image constraints established in the aforementioned four steps to the same diffusion generation process.
[0160] Optionally, the specific implementation of generating a pseudo-viewpoint image at the target pseudo-camera pose can be as follows: input the normalized depth map after step S32 into the depth ControlNet, input the style feature vector extracted in step S33 into the IP-Adapter, load the LoRA weights after fine-tuning in step S34, and use the noisy rendering image obtained in step S35 as the starting image, and then call the SDXL diffusion model to perform the denoising process; during the denoising process, the ControlNet injects depth conditions at each time step, the IP-Adapter injects style features at each cross-attention layer, the LoRA module adjusts the attention calculation of the UNet, and the starting image constraint makes the denoising trajectory always close to the real image.
[0161] In this embodiment, through the synergistic effect of four constraints, the output of the diffusion model is strictly limited to the range consistent with the geometric structure, visual style, scene memory and rendering anchor point of the current 3D Gaussian scene, and finally a high-quality pseudo-viewpoint image is generated. This solves the technical problem of geometric distortion and style drift caused by directly using the diffusion model under sparse viewpoint, and achieves the technical effect of good geometric consistency, high appearance style matching and complete detail preservation of pseudo-viewpoint image.
[0162] In one feasible implementation, step S36, the step of generating a pseudo-viewpoint image at the target pseudo-camera pose based on the common constraints of the geometric constraints, the style constraints, the scene memory constraints, and the starting image, further includes: Step S361: Based on the common constraints of the geometric constraints, the style constraints, the scene memory constraints, and the starting image, generate an initial pseudo-viewpoint image at the target pseudo-camera pose. In one feasible embodiment, after the quadruple constraints are applied simultaneously to the generation process of the diffusion model, a preliminary pseudo-view image is output. However, since the distribution of Gaussian primitives in the current 3D Gaussian scene may not be adequately covered in some areas, there may be image holes in the preliminary result due to the lack of geometric guidance.
[0163] Optionally, the initial pseudo-viewpoint image refers to the pseudo-viewpoint image to be improved, which is directly generated by the SDXL diffusion model under the combined effect of four constraints (depth map geometric constraints, style feature constraints, LoRA scene memory constraints, and render map initial anchoring).
[0164] Understandably, this step generates an initial pseudo-viewpoint image under the joint guidance of four constraints. However, due to the potential blind spots in the sparse Gaussian overlay, some areas in this initial result may appear as black pixels or holes with no effective content.
[0165] Step S362: If there are void areas in the initial pseudo-view image, the void areas are repaired using a preset repair model to obtain the pseudo-view image.
[0166] In one feasible embodiment, when the Gaussian scene is insufficiently covered at a certain viewpoint, the rendered depth map and alpha coverage map will be missing, resulting in holes in the results generated by the diffusion model and poor quality. These hole areas cannot be used for subsequent training supervision and will pollute the training process, so they need to be filled by repairing the model.
[0167] Optionally, the void region is a continuous or discrete region in the initial pseudo-viewpoint image where the proportion of black pixels exceeds two percent. These regions correspond to viewpoint directions where Gaussian pixels are sparsely or completely absent in the current 3D Gaussian scene.
[0168] Optionally, the preset inpainting model can be the SD2 Inpainting (Stable Diffusion 2 Inpainting) model, which is used to fill content in a specified area of an image.
[0169] Optionally, the specific implementation method for repairing the hole region can be as follows: detect the black pixel region in the initial pseudo-view image, generate a binary mask (pixels with a value of 1 in the mask represent holes that need to be repaired, and pixels with a value of 0 represent normal regions); then dilate the mask (for example, use a 3x3 convolution kernel to extend the hole boundary outward by a few pixels) to smooth the repair boundary; then input the initial pseudo-view image, the dilated mask, and the depth map under the current target pseudo-camera pose into the SD2 Inpainting model, which performs progressive noise reduction and filling of the mask region in the latent space; finally, decode the repaired latent variables back into the RGB image to obtain the final pseudo-view image with the holes properly filled.
[0170] Optionally, during the restoration process, weighted fusion can be performed at the boundary between the restored and unrestored areas to ensure a natural transition in texture and color between the restored content and the surrounding intact image areas.
[0171] Understandably, this step fills in the void areas in the diffusion generation results by repairing the model, solving the problem of incomplete generated images caused by insufficient Gaussian coverage, and ensuring that the pseudo-viewpoint images that finally enter the training pool have complete visual content and consistent boundary transitions.
[0172] Based on any of the above embodiments, a fourth embodiment of the 3D Gaussian reconstruction method of this application is proposed. In the aforementioned embodiments, a batch of pseudo-viewpoint images can be obtained through active viewpoint selection and quadruple constraint generation. However, the quality of different pseudo-viewpoint images varies: some pseudo-viewpoint images may have local areas that differ significantly from the current Gaussian scene rendering result (e.g., edges or high-texture areas). Furthermore, due to the different extrapolation distances of the target pseudo-camera pose, pseudo-viewpoints far from the distribution of the real camera have relatively low overall reliability. If all generated pseudo-viewpoints are included in the supervision signal without distinction, poor-quality or unreliable extrapolation pseudo-viewpoints will introduce incorrect optimization gradients, polluting the optimization process of the 3D Gaussian scene. Therefore, this embodiment filters pseudo-viewpoints based on the image differences between the pseudo-viewpoint images and the rendered images, as well as the geometric coverage of the current Gaussian scene under the target pseudo-camera pose. Only pseudo-viewpoints and their corresponding poses that simultaneously meet the conditions of low difference and high coverage are included in the supervision signal. This effectively filters out unreliable supervision while utilizing pseudo-viewpoints to supplement information, ensuring the stability of the optimization process. In this embodiment, before step S4, which optimizes the current 3D Gaussian scene based on the supervision signal to obtain the target 3D Gaussian scene, the following method is further included: Step S41: Determine the image difference between the pseudo-view image and the rendered image of the current 3D Gaussian scene under the target pseudo-camera pose, as well as the geometric coverage of the current 3D Gaussian scene under the target pseudo-camera pose; In one feasible embodiment, in order to assess the quality of the generated pseudo-viewpoint image, quantitative evaluation can be performed from the dimensions of appearance consistency and geometric support, respectively.
[0173] Optionally, the image difference is the pixel-wise residual between the pseudo-viewpoint image generated by the diffusion model and the rendered image obtained by differentiable rasterization of the current 3D Gaussian scene under the same target pseudo-camera pose. Specifically, it can be calculated as: the absolute difference between corresponding pixels between the generated image and the rendered image, and the average of all pixels is taken as the overall difference. This average value is the mean of the generation-rendering residual.
[0174] Optionally, the geometric coverage is the statistical feature of the Gaussian alpha coverage value of each pixel in the current 3D Gaussian scene under the pose of the target pseudo-camera after alpha mixing. The statistical feature can be quantized by the average value of the alpha coverage of all pixels (i.e., the proportion of non-zero pixels in the alpha coverage map), or by the proportion of pixel-level alpha coverage values greater than a certain threshold.
[0175] Optionally, the specific method for determining the mean of the generation-rendering residual can be as follows: at the target pseudo-camera pose, simultaneously obtain the pseudo-view image generated by the diffusion model and the image rendered by the current Gaussian scene; then calculate the pixel-by-pixel residual map and take its mean to obtain the mean of the generation-rendering residual.
[0176] Optionally, the specific method for determining the geometric coverage can be: to count the proportion of pixels with values greater than 0 in the alpha coverage image to the total number of pixels, and obtain the geometric coverage, which can also be called the alpha coverage rate.
[0177] Step S42: The real viewpoint image, the real camera pose, and the first pseudo viewpoint image and its corresponding first target pseudo camera pose are determined as supervision signals, wherein the first pseudo viewpoint image is a pseudo viewpoint image whose image difference is less than a preset difference threshold and whose geometric coverage is greater than a preset coverage threshold.
[0178] In one feasible embodiment, pseudo-view images that simultaneously meet the conditions of low image difference and high geometric coverage are considered reliable and can be incorporated into the supervision signal for subsequent optimization to prevent pseudo-view images that do not meet the conditions (i.e., excessive residuals or insufficient coverage) from contaminating the training.
[0179] Optionally, the first pseudo-viewpoint image is a pseudo-viewpoint image that has been filtered and whose image difference is less than a preset difference threshold (which can be 0.7) and whose geometric coverage is greater than a preset coverage threshold (which can be 0.25).
[0180] Optionally, the first target pseudo-camera pose is the target pseudo-camera pose used when generating the first pseudo-viewpoint image, which corresponds to the first pseudo-viewpoint image.
[0181] Optionally, the supervision signal is determined by using the real viewpoint image and its real camera pose as the supervision signal, and also by adding the first pseudo viewpoint image that passes the screening criteria and its corresponding first target pseudo camera pose to the supervision signal.
[0182] Understandably, this step, through screening and gating, constructs a supervisory signal composed of both real and high-quality pseudo-viewpoints. This ensures the reliability of the basic supervision while safely utilizing additional pseudo-viewpoint information to enhance the reconstruction of 3D Gaussian scenes under sparse perspectives.
[0183] In this embodiment, by first determining the image differences between the pseudo-view image and the rendered image, as well as the geometric coverage of the current Gaussian scene under the target pseudo-camera pose, the first pseudo-view image that meets the screening criteria and its first target pseudo-camera pose, along with the real view image and the real camera pose, are determined as supervision signals. This solves the problem of inconsistent quality among different pseudo-views and the pollution of the training process by low-quality pseudo-views. This method, through a screening mechanism, only allows pseudo-views with consistent appearance and sufficient geometric support as training data, thereby achieving the technical effect of effectively filtering out unreliable supervision data and ensuring the stability of 3D Gaussian scene optimization while utilizing pseudo-views to supplement information in unobserved areas.
[0184] In one feasible implementation, before step S4, which optimizes the current 3D Gaussian scene based on the supervision signal to obtain the target 3D Gaussian scene, the method further includes: Step E41: Based on the image differences, obtain the pixel-level residual between the pseudo-view image and the rendered image, and construct a pixel-level uncertainty map based on the pixel-level residual; In one feasible embodiment, the overall mean of image differences can reflect the overall quality of the pseudo-viewpoint image. The generation reliability at different pixel locations often varies (e.g., edge regions or high-texture regions may have larger residuals), so it is necessary to refine it down to each pixel.
[0185] Alternatively, pixel-level residuals refer to pseudo-viewpoint images. Rendered image of the current 3D Gaussian scene At the same pixel position The absolute difference is calculated using the following formula: .
[0186] Optionally, a pixel-level uncertainty map, denoted as , A pixel with a larger residual indicates a greater difference between the diffusion generation result and the current Gaussian scene at that location, and its generation reliability is lower.
[0187] Step E42: Obtain the overlay map of the current 3D Gaussian scene under the pose of the target pseudo-camera; In one feasible embodiment, pixel-level residuals alone cannot distinguish whether the difference is due to a diffusion generation error or a lack of geometric support in the Gaussian scene at that location. Therefore, a cover map is needed to identify which pixels have Gaussian support and which pixels are empty regions.
[0188] Optionally, the overlay map is the cumulative opacity value of each pixel obtained by alpha mixing after projecting each Gaussian primitive in the current 3D Gaussian scene under the target pseudo-camera pose, denoted as The value ranges from 0 to 1.
[0189] Step E43: Construct a confidence weight map based on the coverage map and the pixel-level uncertainty map; In one feasible embodiment, the supervision weight of a pixel should simultaneously consider: whether the pixel is covered by Gaussian pixels and the magnitude of the difference between the generated result and the rendered result on the pixel. Only pixels covered by Gaussian pixels can be supervised, and pixel regions with large differences should have their weights reduced.
[0190] Optionally, the confidence weight map is denoted as... , is a matrix of the same size as the image, where the value of each pixel represents the supervision strength of that pixel when calculating the loss.
[0191] Alternatively, the construction method can be: ,in, This is the absolute value of the maximum pixel-by-pixel residual. When This indicates no Gaussian element coverage, with a weight of 0. near The time weight is 0.3 times the original weight.
[0192] Understandably, this step constructs an adaptive pixel-level confidence weight map, which gives high weight to regions with high coverage and low difference, while giving low weight or even zero weight to regions with low coverage or high difference, thereby automatically suppressing the influence of unreliable pixels during the optimization process.
[0193] Step E44: Obtain the extrapolation distance of the target pseudo-camera pose relative to all real camera poses. If the extrapolation distance is greater than a preset extrapolation distance threshold, determine the exponential decay weight of the pseudo-view image. In one feasible embodiment, in addition to pixel-level uncertainty, the credibility of the entire pseudo-view image is also related to its corresponding camera pose: the greater the extrapolation distance, the more the pseudo-view deviates from the distribution of the real camera, and the lower its overall reliability. Therefore, it is necessary to reduce the overall supervision intensity of the pseudo-view.
[0194] Optionally, the extrapolation distance is denoted as , refers to the minimum Euclidean distance between the translation components of the target pseudo-camera pose and the translation components of all real camera poses, that is, the spatial straight-line distance from the target pseudo-camera pose to the nearest real camera position.
[0195] Optionally, the preset extrapolation distance threshold is a pre-set value. When the extrapolation distance is less than the extrapolation distance threshold, the pseudo-viewpoint is considered to be in the safe interpolation region and no additional attenuation is applied. When the extrapolation distance is greater than the threshold, an exponential attenuation weight is applied.
[0196] Optionally, determine the exponential decay weights. The method is as follows: , where 2.0 is the attenuation intensity coefficient.
[0197] Understandably, this step assigns attenuation weights to pseudo-viewpoints at different extrapolation distances, thereby suppressing the contribution of pseudo-viewpoints far from the real camera distribution to the optimization and reducing the risk of extrapolation uncertainty.
[0198] Step E45: Based on the confidence weight map, the exponential decay weight, and the preset photometric loss function, obtain the weighted pseudo-viewpoint supervision loss, wherein the weighted pseudo-viewpoint supervision loss includes a first loss component determined based on the pseudo-viewpoint image and its corresponding target pseudo-camera pose. In one feasible embodiment, the pixel-level confidence weight map is combined with the overall exponential decay weight to weight the photometric loss of the pseudo-viewpoint image, thereby obtaining the final loss value of the pseudo-viewpoint's contribution to the optimization.
[0199] Optionally, the preset photometric loss function includes a combination of L1 loss (pixel-by-pixel absolute error) and DSSIM (Structural Similarity Index Dissimilarity).
[0200] Optionally, the weighted pseudo-viewpoint supervision loss can be calculated as follows: the photometric loss of each pixel is weighted using a confidence weight map to obtain the weighted pixel loss; then the pixel-level weighted loss is summed and divided by the total number of pixels (or the number of effective pixels); finally, it is multiplied by the exponential decay weight of the pseudo-viewpoint.
[0201] Optionally, the first loss component refers to the loss calculated from the pseudo-viewpoint image and its corresponding target pseudo-camera pose.
[0202] Understandably, this step constructs a pseudo-viewpoint supervised loss, enabling adaptive weighting of unreliable pixels and unreliable poses.
[0203] Step E46: The weighted pseudo-viewpoint supervision loss and the second loss component are determined as supervision signals, wherein the second loss component is determined based on the real viewpoint image and the real camera pose.
[0204] In one feasible embodiment, the supervision signal is partly derived from the supervision loss of the real viewpoint image and the real camera pose, and partly derived from the weighted pseudo-viewpoint supervision loss.
[0205] Optionally, the second loss component is a photometric loss calculated using a combination of L1 and DSSIM with real-view images and their real camera poses.
[0206] Optionally, the weighted pseudo-view supervision loss and the second loss component are used together as supervision signals. In the alternating optimization framework, the 3D Gaussian reconstruction device samples from the pseudo-view pool with a preset probability and alternates with the real view. The total optimization loss is the weighted pseudo-view supervision loss and the second loss component.
[0207] In this implementation, the additional supervision provided by the high-quality pseudo-viewpoint is combined with the basic supervision of the real viewpoint to jointly guide the parameter update of the current 3D Gaussian scene, thereby obtaining a more complete and accurate target 3D Gaussian scene under sparse input conditions.
[0208] Based on any of the above embodiments, a sixth embodiment of the 3D Gaussian reconstruction method of this application is proposed. In the conventional approach, in order to suppress overfitting under sparse perspectives, random dropout regularization is usually applied to Gaussian primitives. However, the geometric complexity of different scenes varies greatly: high-density scenes (such as dense vegetation) may require hundreds of thousands of Gaussian primitives, while low-density scenes (such as open indoor spaces) only require a few thousand. If a uniform regularization dropout ratio is applied to all scenes, high-density scenes will produce floating artifacts due to insufficient regularization strength, while low-density scenes will suppress the already fragile detail structure due to excessive regularization, resulting in a decrease in reconstruction quality. Therefore, this embodiment proposes a density-adaptive Gaussian primitive random dropout regularization method, which dynamically adjusts the dropout ratio according to the actual total number of Gaussian primitives in the current scene, so that the regularization strength matches the geometric complexity of the scene, thereby protecting the detail structure while suppressing overfitting. In this embodiment, the 3D Gaussian reconstruction method further includes the following steps: Step D1: Obtain the total number of Gaussian elements in the current 3D Gaussian scene, as well as the current iteration step and the total number of iteration steps in the preset convergence condition; In one feasible embodiment, in order to dynamically adjust the dropout ratio, it is necessary to know the density of the current Gaussian scene and the current training progress in real time.
[0209] Optionally, the total number of Gaussian elements is denoted as... , refers to the number of all anisotropic Gaussian primitives in the current 3D Gaussian scene. Each Gaussian primitive corresponds to an ellipsoidal geometric primitive in the scene. The more primitives there are, the higher the geometric complexity of the scene.
[0210] Optionally, the current iteration step number is denoted as... , refers to the number of 3D Gaussian scene optimization iterations performed from the start of training to the current moment.
[0211] Optionally, the total number of iterations in the preset convergence condition is denoted as... , refers to the total number of optimization iterations that need to be executed in the entire training process, which is set in advance. For example, it can be set to a total number of iterations T = 30,000 steps, and the training ends when t reaches T.
[0212] Step D2: Calculate the discard ratio based on the ratio of the total number of Gaussian cells to a preset density threshold and the ratio of the current iteration step to the total number of iteration steps, wherein the total number of Gaussian cells and the discard ratio are positively correlated. In one feasible embodiment, the dropout ratio should gradually increase as the training process progresses, i.e., the intensity of discontinuation regularization is low in the early stage of training and high in the later stage of training, and should be proportional to the density of the scene: the more Gaussian elements a scene has, the stronger the regularization is needed to suppress overfitting.
[0213] Optionally, a preset density threshold is denoted as... , is a pre-set reference base number used to determine whether the current scene is high-density or low-density, and can be set to 200000.
[0214] Optionally, the discard ratio The calculation formula is: ,in The maximum discard ratio can be set to 0.2. Indicates when Exceed The value is set to 1 if the discard ratio is 1, otherwise the actual ratio is used, ensuring that the upper limit of the discard ratio does not exceed 1 / 3. ; This represents the training progress coefficient, where the dropout rate is lower in the early stages of training and gradually increases in the later stages. It is positively correlated with the discard rate, that is, the larger the total number of Gaussian elements, the higher the discard rate.
[0215] Understandably, this step achieves on-demand allocation of regularization intensity by introducing a coefficient proportional to the total number of Gaussian elements, allowing the dropout ratio to adapt to the actual geometric complexity of the scene.
[0216] Step D3: According to the stated discard ratio, discard a corresponding proportion of Gaussian elements from the current 3D Gaussian scene.
[0217] In one feasible embodiment, after calculating the dynamic drop-off ratio, a random drop-off operation needs to be performed on Gaussian primitives during the actual Gaussian scene optimization process. At the same time, in order to ensure the stability of training, only the opacity gradient should be affected without interfering with the gradient flow of other parameters (such as position, shape, and color).
[0218] Optionally, the opacity parameter of each Gaussian element in the current 3D Gaussian scene can be zeroed out by a discard ratio or a compensation vector can be applied to its gradient. More specifically, this can be achieved by constructing a mask with the same number of Gaussian elements, where each element is assigned a probability. Choose 1, with probability The opacity gradient is set to 0, and then the mask is multiplied by the opacity gradient, so that the opacity gradient of the discarded Gaussian unit is zero, thus preventing it from being updated through backpropagation in this iteration; while the gradient flow of other parameters is unaffected and continues to be updated normally.
[0219] Optionally, discarded Gaussian elements can be removed or merged.
[0220] Understandably, this step, by randomly discarding Gaussian units at a dynamic ratio and only intervening in the opacity gradient, regularizes high-density scenes while avoiding excessive suppression of fragile structures in low-density scenes, thereby improving the generalization ability and detail preservation ability of 3D Gaussian scene reconstruction from a sparse perspective.
[0221] In this embodiment, the dropout ratio is calculated by obtaining the current total number of Gaussian primitives, the current iteration step, and the total number of iteration steps. Then, a corresponding proportion of Gaussian primitives are randomly dropped according to this ratio. This method solves the problem that a uniform dropout ratio cannot adapt to scenes with different geometric complexities. By introducing a density adaptive coefficient that is proportional to the total number of Gaussian primitives, this embodiment achieves the technical effect of obtaining strong regularization to suppress overfitting in high-density scenes and weak regularization to protect detailed structures in low-density scenes. This significantly improves the geometric integrity and visual quality of 3D Gaussian scene reconstruction under sparse viewpoint input.
[0222] In one feasible embodiment, the three-dimensional Gaussian reconstruction method of this application can be: Acquire sparse real view images and the real camera poses corresponding to each real view image. Based on the real view images and the real camera poses, initialize the 3D Gaussian scene to be optimized to obtain the initial 3D Gaussian scene. The initial 3D Gaussian scene is optimized using the real-view image to obtain the current 3D Gaussian scene; Based on the existing real camera poses in the current 3D Gaussian scene, candidate pseudo camera poses are determined, and a target pseudo camera pose is selected from the candidate pseudo camera poses according to the comprehensive evaluation value corresponding to each candidate pseudo camera pose. Based on the multi-source constraint information of the current 3D Gaussian scene, a pseudo-viewpoint image is generated at the pose of the target pseudo-camera. The multi-source constraint information includes the geometric information of the current 3D Gaussian scene, the style features of the real viewpoint image, the scene memory capability of the preset diffusion model, and the rendering result of the current 3D Gaussian scene at the pose of the target pseudo-camera. Determine the image difference between the pseudo-viewpoint image and the rendered image of the current 3D Gaussian scene under the target pseudo-camera pose, as well as the geometric coverage of the current 3D Gaussian scene under the target pseudo-camera pose; The pseudo-viewpoint image with an image difference less than a preset difference threshold and a geometric coverage greater than a preset coverage threshold, along with the corresponding target pseudo-camera pose, and the real-viewpoint image with its corresponding real-camera pose, are determined as supervision signals. The current 3D Gaussian scene is optimized, and the optimized 3D Gaussian scene is taken as the new current 3D Gaussian scene. The process then returns to the step of determining candidate pseudo-camera poses based on the existing real-camera poses in the current 3D Gaussian scene, and subsequent steps, until a preset convergence condition is met. The new current 3D Gaussian scene that meets the preset convergence condition is taken as the target 3D Gaussian scene.
[0223] In one feasible embodiment, the 3D Gaussian reconstruction method of this application employs a training strategy that alternates between 3D Gaussian optimization and diffusion model generation. Specifically, in the warm-up phase, only real-view images are used for optimization to establish an initial scene representation; subsequently, a diffusion generation phase is triggered every 500 iterations. In the diffusion generation phase, the most valuable target pseudo-camera pose is determined through generative active viewpoint selection, and pseudo-view images are generated using four constraints (depth geometry, style features, LoRA scene memory, and render draft anchoring). The pseudo-view images selected through a filtering mechanism (generated-rendered residual mean less than 0.7 and alpha coverage greater than 0.25) are used as subsequent training data. The subsequent 3D Gaussian optimization phase trains by mixing real-view images with the selected pseudo-view images using a preset probability, and this process is repeated until convergence.
[0224] Optionally, during execution, the method of this application can manage multiple large models on a single GPU (Graphics Processing Unit, such as 24GB of video memory) using a time-sharing multiplexing strategy. Specifically, during the 3D Gaussian optimization stage, only the current 3D Gaussian scene (i.e., the scene representation of the Gaussian model) is retained on the GPU, while the depth estimation model, SDXL diffusion model, and SD2 repair model are migrated from the GPU to the CPU (Central Processing Unit) and reside in memory. When the diffusion generation stage is triggered and a pseudo-viewpoint image needs to be generated, the current 3D Gaussian scene is first migrated from the GPU to the CPU while preserving its parameter state. Then, the depth estimation model, SDXL diffusion model, and SD2 repair model are loaded onto the GPU sequentially or in parallel for inference. After the pseudo-viewpoint generation is completed, these diffusion models are migrated back to the CPU, and the current 3D Gaussian scene is reloaded back onto the GPU for further optimization. This on-demand migration ensures that the GPU only accommodates the components required for the current task at any given time, thus running the complete process within the 24GB video memory limit.
[0225] In this approach, the migration of the current 3D Gaussian scene between GPU and CPU employs in-situ data replacement. This means that the stored parameter tensor is directly replaced with a new tensor on the target device using the operation ".data = .data.to(device)", while preserving the parameter's object identity as an nn.Parameter (neural network parameter, representing a tensor requiring differentiation in the PyTorch framework). This ensures that the momentum states stored in the Adam (Adaptive Moment Estimation) optimizer (including first-order moment estimates exp_avg and second-order moment estimates exp_avg_sq) remain associated with the original parameter object, preventing resets or loss, thus ensuring that training continuity is completely unaffected by component migration.
[0226] Optionally, the method in this application was evaluated on the LLFF (Local Light Field Fusion) dataset, which contains eight real-world forward scenes covering various scene types such as vegetation, flowers, buildings, and interiors. Tests were conducted under three sparse settings: 3-view, 6-view, and 9-view, using Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Perceptual Distance (LPIPS) as evaluation metrics.
[0227] Optionally, experimental results show that our proposed method consistently outperforms the baseline 3DGS method under all sparse viewpoint settings, achieving improvements across all three evaluation metrics. The improvement is greatest under the 3-viewpoint setting, indicating that our method yields the most significant benefit under the most information-scarce extreme sparsity conditions.
[0228] Optionally, the effectiveness of each core component proposed in this application was verified through ablation experiments. Specifically, a component was removed from the complete scheme (e.g., density adaptive regularization was "eliminated," the progressive extrapolation hierarchy in active viewpoint selection was "eliminated," or the entire diffusion generation module was "eliminated"), and then the performance difference before and after removal was compared on the same dataset and under the same evaluation metrics. If the performance decreased significantly after removal, it indicates that the component is necessary for the overall scheme; if the performance did not change much after removal, it indicates that the component's contribution is limited.
[0229] In this application, ablation experiments were conducted on the LLFF dataset with 3-view, 6-view, and 9-view sparse settings, and the evaluation metrics used were PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity), and LPIPS (Perceptual Distance). Experimental results show that: First, when density-adaptive regularization was removed from the complete solution (i.e., a fixed dropout ratio was used instead), PSNR and SSIM decreased in most test scenarios, while LPIPS increased (perceived quality deteriorated). The performance degradation was greatest in high-density scenarios (approximately 310,000 Gaussian cells), indicating that the density-adaptive mechanism provided stronger regularization constraints for high-density scenarios, effectively suppressing floating artifacts and overfitting. This validates the necessity of this component and its significant improvement effect in high-density scenarios.
[0230] Second, when the progressive extrapolation hierarchy is removed from the active viewpoint selection (i.e., only a fixed interpolation pose is used, and level 1 and level 2 extrapolation are not enabled), the reconstruction quality in multiple forward scenes (such as flowers and buildings) also decreases significantly, especially the geometric integrity and appearance consistency of large-view extrapolation regions deteriorate. This indicates that the progressive extrapolation hierarchy can automatically expand the spatial range of candidate views based on the historical generation quality, providing richer extrapolation information for unobserved areas, thus verifying the effectiveness of the active viewpoint selection strategy and the additional improvements it brings.
[0231] Third, when the entire diffusion generation module (including pseudo-viewpoint generation, admission, and weighting) is completely removed, and only the real-viewpoint image is used for 3D Gaussian optimization, the PSNR and SSIM under all sparse settings decrease significantly, while LPIPS increases significantly, and the performance degrades to a level close to that of the original sparse 3D Gaussian splashing method. This indicates that pseudo-viewpoint supervision, as the core component of this method, is key to improving the quality of sparse viewpoint reconstruction, verifying its irreplaceable and important role.
[0232] In summary, the ablation experiments not only confirmed the effectiveness of each innovative component in this application but also quantified the performance improvement they brought: density adaptive regularization yielded the greatest benefit in high-density scenarios, the progressive extrapolation level contributed additional improvement in scenarios requiring extrapolation, and the diffusion generation module improved the overall performance of the solution. After multiple experiments to explore hyperparameters, the current parameter configuration (such as a maximum dropout ratio of 0.2, a density threshold of 200,000, an annealing index decaying from 1.5 to 0.4, and three extrapolation levels) was verified as the optimal parameter configuration.
[0233] For example, please refer to Figure 2 , Figure 2 The image shows a test scenario for ferns. Figure 2 (a) is the RGB image of the current 3D Gaussian scene directly rendered in the pose of a certain target pseudo-camera (without being generated by the diffusion model). Figure 2 (b) is a pseudo-view image of the current 3D Gaussian scene under the same target pseudo-camera pose, generated by the diffusion model with four constraints.
[0234] For example, please refer to Figure 3 , Figure 3 (c) is the current 3D Gaussian scene and Figure 2 (b) Depth map directly rendered under the same target pseudo-camera pose (grayscale represents distance, the closer the target, the brighter the map). Figure 3 (d) is the pixel-by-pixel residual map between the pseudo-viewpoint image and the Gaussian rendered image (representing pixel-level uncertainty, with bright areas indicating large differences).
[0235] For example, please refer to Figure 4 , Figure 4 (e) is a real image actually taken from the test viewpoint, used as a reference standard for evaluating the reconstruction quality. Figure 4 (f) is the image of the 3D Gaussian scene finally reconstructed by this method and rendered in the test view using only 3 sparse real viewpoints.
[0236] Through the above Figures 2-4 Visual contrast, the method generates Figure 2 (b) Pseudo-perspective images are similar in geometry and style to Figure 4 (e) Highly consistent with real-world scenarios Figure 3 (d) The pixel-level uncertainty map contains only a few highlighted areas, and Figure 4 (f) Final rendering result and Figure 4 (e) The test results are close to the true values, which verifies the effectiveness of the proposed method in reconstruction under extremely sparse conditions.
[0237] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the three-dimensional Gaussian reconstruction method of this application. Any simple transformations based on this technical concept, such as the interaction and combination of various embodiments, are all within the protection scope of this application.
[0238] This application also provides a three-dimensional Gaussian reconstruction device; please refer to... Figure 5 The three-dimensional Gaussian reconstruction device includes: The acquisition module 10 is used to acquire the current three-dimensional Gaussian scene, wherein the current three-dimensional Gaussian scene is determined based on sparse real view images and their corresponding real camera poses. The determining module 20 is used to determine the target pseudo-camera pose based on the real camera pose; The generation module 30 is used to generate a pseudo-view image at the pose of the target pseudo camera based on the geometric information of the current three-dimensional Gaussian scene, the style features of the real view image, the scene memory capability of the preset diffusion model, and the rendering result of the current three-dimensional Gaussian scene under the pose of the target pseudo camera. The optimization module 40 is used to optimize the current 3D Gaussian scene based on the supervision signal to obtain the target 3D Gaussian scene, wherein the supervision signal includes: the pseudo-view image, the target pseudo-camera pose, the real view image, and the real camera pose.
[0239] The three-dimensional Gaussian reconstruction device provided in this application, employing the three-dimensional Gaussian reconstruction method in the above embodiments, can solve the technical problem of low reconstruction accuracy of three-dimensional Gaussian splashing under conventional sparse viewpoints. Compared with the prior art, the beneficial effects of the three-dimensional Gaussian reconstruction device provided in this application are the same as those of the three-dimensional Gaussian reconstruction method provided in the above embodiments, and other technical features in the three-dimensional Gaussian reconstruction device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0240] This application provides a three-dimensional Gaussian reconstruction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the three-dimensional Gaussian reconstruction method in the first embodiment described above.
[0241] The following is for reference. Figure 6The diagram illustrates a structural schematic of a three-dimensional Gaussian reconstruction device suitable for implementing embodiments of this application. The three-dimensional Gaussian reconstruction device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The illustrated 3D Gaussian reconstruction device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0242] like Figure 6 As shown, the 3D Gaussian reconstruction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the 3D Gaussian reconstruction device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the 3D Gaussian reconstruction device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows a 3D Gaussian reconstruction device with various systems, it should be understood that it is not required to implement or possess all of the systems shown. More or fewer systems may be implemented alternatively.
[0243] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0244] The 3D Gaussian reconstruction device provided in this application, employing the 3D Gaussian reconstruction method described in the above embodiments, can solve the technical problem of low reconstruction accuracy of 3D Gaussian splashing under conventional sparse viewpoints. Compared with the prior art, the beneficial effects of the 3D Gaussian reconstruction device provided in this application are the same as those of the 3D Gaussian reconstruction method provided in the above embodiments, and other technical features of this 3D Gaussian reconstruction device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0245] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0246] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0247] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the three-dimensional Gaussian reconstruction method in the above embodiments.
[0248] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0249] The aforementioned computer-readable storage medium may be included in the 3D Gaussian reconstruction device; or it may exist independently and not be assembled into the 3D Gaussian reconstruction device.
[0250] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0251] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0252] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0253] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described three-dimensional Gaussian reconstruction method, which can solve the technical problem of low reconstruction accuracy of three-dimensional Gaussian splashing under conventional sparse viewpoints. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the three-dimensional Gaussian reconstruction method provided in the above embodiments, and will not be repeated here.
[0254] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the three-dimensional Gaussian reconstruction method as described above.
[0255] The computer program product provided in this application can solve the technical problem of low reconstruction accuracy of 3D Gaussian splashing under conventional sparse viewpoints. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the 3D Gaussian reconstruction method provided in the above embodiments, and will not be repeated here.
[0256] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A three-dimensional Gaussian reconstruction method, characterized in that, The three-dimensional Gaussian reconstruction method includes: Obtain the current 3D Gaussian scene, wherein the current 3D Gaussian scene is determined based on sparse real-view images and their corresponding real camera poses; Determine the target pseudo-camera pose based on the actual camera pose; Based on the geometric information of the current 3D Gaussian scene, the style features of the real-view image, the scene memory capability of the preset diffusion model, and the rendering result of the current 3D Gaussian scene under the pose of the target pseudo-camera, a pseudo-view image is generated at the pose of the target pseudo-camera. Based on the supervision signal, the current 3D Gaussian scene is optimized to obtain the target 3D Gaussian scene, wherein the supervision signal includes: the pseudo-view image, the target pseudo-camera pose, the real view image, and the real camera pose.
2. The method as described in claim 1, characterized in that, The step of determining the target pseudo-camera pose based on the real camera pose includes: Based on the existing real camera poses in the current 3D Gaussian scene, determine the candidate pseudo camera poses; Based on the comprehensive evaluation value corresponding to the candidate pseudo camera pose, the target pseudo camera pose is selected from the candidate pseudo camera poses.
3. The method as described in claim 2, characterized in that, The step of determining candidate pseudo-camera poses based on the existing real camera poses in the current 3D Gaussian scene includes: Based on the existing real camera poses in the current 3D Gaussian scene, generate pose pairs, wherein the pose pairs include two different real camera poses; Spherical linear interpolation is performed on the rotational components of the pose in the pose pair to obtain the first interpolation result; Linear interpolation is performed on the translation components of the pose in the pose pair to obtain a second interpolation result; Based on the first interpolation result and the second interpolation result, an interpolated pseudo-camera pose is generated between the two poses in the pose pair; The rotation and translation components of each pose in the pose pair are extrapolated along both ends of the line connecting the two poses in the pose pair to generate an extrapolated pseudo camera pose located outside the line connecting the two poses in the pose pair. The interpolated pseudo-camera pose and the extrapolated pseudo-camera pose are determined as candidate pseudo-camera poses.
4. The method as described in claim 2, characterized in that, The step of selecting the target pseudo camera pose from the candidate pseudo camera poses based on the comprehensive evaluation value corresponding to the candidate pseudo camera poses includes: Obtain the reconstruction information gain index and generation reliability index of each candidate pseudo-camera pose; The comprehensive evaluation value of each candidate pseudo camera pose is calculated based on the preset annealing index, the reconstruction information gain index, and the generation reliability index. The candidate pseudo-camera poses are sorted according to the comprehensive evaluation value to obtain the sorting result; Based on the sorting results, the target pseudo camera pose is determined from each of the candidate pseudo camera poses.
5. The method as described in claim 4, characterized in that, The step of determining the target pseudo-camera pose from the candidate pseudo-camera poses based on the sorting result includes: Based on the sorting results, the pose of the current candidate pseudo camera is determined; Determine whether the pose distance between the current candidate pseudo camera pose and the selected target pseudo camera pose is greater than or equal to a preset distance threshold; If the pose distance between the current candidate pseudo camera pose and each of the selected target pseudo camera poses is greater than or equal to a preset distance threshold, the current candidate pseudo camera pose is determined as the target pseudo camera pose. If the attitude distance between the current candidate pseudo-camera pose and at least one selected target pseudo-camera pose is less than a preset distance threshold, the candidate pseudo-camera pose that is one position after the current candidate pseudo-camera pose in the sorting result is determined as the new current candidate pseudo-camera pose, and the step of determining whether the attitude distance between the current candidate pseudo-camera pose and the selected target pseudo-camera pose is greater than or equal to the preset distance threshold is executed until the number of selected target pseudo-camera poses reaches a preset number and / or all candidate pseudo-camera poses are traversed.
6. The method as described in claim 1, characterized in that, The step of generating a pseudo-viewpoint image at the target pseudo-camera pose based on the geometric information of the current 3D Gaussian scene, the style features of the real-view image, the scene memory capability of the preset diffusion model, and the rendering result of the current 3D Gaussian scene under the target pseudo-camera pose includes: Project each Gaussian primitive in the current 3D Gaussian scene onto the image plane where the target pseudo-camera pose is located, and calculate the depth and color of each Gaussian primitive on the image plane to obtain the depth map and rendering map under the target pseudo-camera pose respectively. The depth map is used as a geometric constraint; The style features of the real-view image at a preset position relative to the target pseudo-camera pose are used as style constraints; A preset diffusion model is used as a scene memory constraint, wherein the diffusion model is a model that has been fine-tuned in advance on sparse real-view images; The rendered image is used as the starting image for the diffusion process; Based on the combined constraints of the geometric constraints, the style constraints, the scene memory constraints, and the starting image, a pseudo-viewpoint image is generated at the target pseudo-camera pose.
7. The method as described in claim 6, characterized in that, The step of generating a pseudo-viewpoint image at the target pseudo-camera pose based on the common constraints of the geometric constraints, the style constraints, the scene memory constraints, and the starting image further includes: Based on the common constraints of the geometric constraints, the style constraints, the scene memory constraints, and the starting image, an initial pseudo-viewpoint image is generated at the target pseudo-camera pose. If there are void areas in the initial pseudo-viewpoint image, the void areas are repaired using a preset repair model to obtain the pseudo-viewpoint image.
8. The method as described in claim 1, characterized in that, Before the step of optimizing the current 3D Gaussian scene based on the supervision signal to obtain the target 3D Gaussian scene, the method further includes: Determine the image difference between the pseudo-viewpoint image and the rendered image of the current 3D Gaussian scene under the target pseudo-camera pose, as well as the geometric coverage of the current 3D Gaussian scene under the target pseudo-camera pose; The real-view image, the real camera pose, and the first pseudo-view image and its corresponding first target pseudo-camera pose are determined as supervision signals, wherein the first pseudo-view image is a pseudo-view image whose image difference is less than a preset difference threshold and whose geometric coverage is greater than a preset coverage threshold.
9. The method as described in claim 8, characterized in that, Before the step of optimizing the current 3D Gaussian scene based on the supervision signal to obtain the target 3D Gaussian scene, the method further includes: Based on the image differences, the pixel-level residual between the pseudo-view image and the rendered image is obtained, and a pixel-level uncertainty map is constructed based on the pixel-level residual. Obtain the overlay map of the current 3D Gaussian scene under the pose of the target pseudo-camera; A confidence weight map is constructed based on the coverage map and the pixel-level uncertainty map; Obtain the extrapolation distance of the target pseudo-camera pose relative to all real camera poses, and determine the exponential decay weight of the pseudo-view image if the extrapolation distance is greater than a preset extrapolation distance threshold. Based on the confidence weight map, the exponential decay weight, and the preset photometric loss function, a weighted pseudo-viewpoint supervision loss is obtained, wherein the weighted pseudo-viewpoint supervision loss includes a first loss component determined based on the pseudo-viewpoint image and its corresponding target pseudo-camera pose. The weighted pseudo-viewpoint supervision loss and the second loss component are determined as supervision signals, wherein the second loss component is determined based on the real viewpoint image and the real camera pose.
10. The method as described in claim 1, characterized in that, The three-dimensional Gaussian reconstruction method also includes: Obtain the total number of Gaussian elements in the current 3D Gaussian scene, as well as the current iteration step and the total number of iteration steps in the preset convergence condition; The discard ratio is calculated based on the ratio of the total number of Gaussian cells to a preset density threshold and the ratio of the current iteration step to the total number of iteration steps, wherein the total number of Gaussian cells and the discard ratio are positively correlated. According to the stated discard ratio, a corresponding proportion of Gaussian elements are discarded from the current 3D Gaussian scene.
11. A three-dimensional Gaussian reconstruction device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the three-dimensional Gaussian reconstruction method as described in any one of claims 1 to 10.
12. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the three-dimensional Gaussian reconstruction method as described in any one of claims 1 to 10.