3D scene content generation using 2d inpainting diffusion
By training a NeRF model with a 2D inpainting diffusion model as a generative prior, the challenges of generating realistic and consistent 3D content from 2D data are addressed, achieving improved controllability and depth accuracy while reducing computational resources.
Patent Information
- Application Number
- PCT/US2024/057067
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
Current methods for 3D content generation struggle to produce realistic and consistent 3D scenes from 2D data, often requiring extensive multiview data, specific camera pose information, and lacking control over generated content and accurate depth capture.
The approach involves training a Neural Radiance Field (NeRF) model using a 2D inpainting diffusion model as a generative prior, allowing for the generation of realistic 3D content in masked regions without requiring multiview data or specific camera poses, and enabling better control over the generated content and depth information.
This method effectively generates realistic and consistent 3D content in masked regions, improving controllability and depth accuracy, and reducing computational resources needed for training, thus enhancing energy efficiency and practicality for real-world applications.
Smart Images

Figure US2024057067_30052025_PF_FP_ABST
Abstract
Description
3D SCENE CONTENT GENERATION USING 2D INPAINTING DIFFUSION RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of United States Provisional Patent Application Number 63 / 602,044, filed November 22, 2023. United States Provisional Patent Application Number 63 / 602,044 is hereby incorporated by reference in its entirety. FIELD
[0002] The present disclosure relates generally to machine learning. More particularly, the present disclosure relates to a general approach to inpainting 3D content by using a 2D inpainting diffusion model trained on static scenes as a generative prior. BACKGROUND
[0003] The field of three-dimensional (3D) content generation and modeling has seen significant progress with the advent of machine learning techniques. However, generating realistic 3D scenes from two-dimensional (2D) data is still a challenging task due to the high complexity of 3D scenes and the lack of explicit 3D priors in 2D image synthesis models. This presents a technical problem as the generated content can often be inconsistent or blurry when viewed from different angles.
[0004] Traditional methods for 3D content generation typically focus on specific tasks such as object removal, but they struggle to generate realistic content in any masked 3D region while maintaining the complexity of the original scene. Often, these methods require a large amount of multiview data and specific camera pose information, which makes them less practical for real-world applications where such data is not readily available. Moreover, these methods often require retraining of the models with multiview data, which can be computationally expensive and time-consuming.
[0005] Another technical problem is the limited controllability of the generated content, especially when using text-to-image diffusion models. In these models, the generated content is often bottlenecked by the expressivity of the text, limiting the diversity and complexity of the 3D scenes that can be generated.
[0006] Furthermore, the depth information, which is critical for accurate 3D scene generation, is often not adequately captured or utilized in traditional methods. This leads to inaccuracies in the geometry of the generated 3D scene and affects the overall realism of the output.
[0007] Therefore, there is a need for an improved method for 3D content generation that can leverage 2D data, does not require extensive multiview data or specific camera pose information, allows more controllability of the generated content, and accurately captures depth information. SUMMARY
[0008] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or can be learned from the description, or can be learned through practice of the embodiments.
[0009] One example aspect of the present disclosure is directed to a computer- implemented method to train a neural radiance field (NeRF) model to generate novel content for a three-dimensional (3D) scene. The method includes: processing, by a computing system comprising one or more computing devices, data descriptive of a camera pose with the NeRF model to generate a rendered image that depicts the 3D scene from the camera pose; applying, by the computing system, an image mask to the rendered image to generate a masked image; adding, by the computing system, a set of noise to the rendered image to generate a noised image; processing, by the computing system, the noised image with a denoising diffusion model conditioned upon the masked image to generate a denoising prediction; evaluating, by the computing system, a loss function that comprises a distillation loss term that evaluates a difference between the set of noise and the denoising prediction; and modifying, by the computing system, one or more values of one or more parameters of the NeRF model based on the loss function.
[0010] Another example aspect of the present disclosure is directed to one or more non- transitory computer-readable media that collectively store a neural radiance field (NeRF) model that has been trained by performance of the methods described herein.
[0011] Another example aspect of the present disclosure is directed to a computing system comprising one or more processors and one or more non-transitory computer-readable media that collectively store computer-executable instructions for performing operations. The operations include: processing, by the computing system, data descriptive of a camera pose with a differentiable 3D scene representation model to generate a rendered image that depicts the 3D scene from the camera pose; applying, by the computing system, an image mask to the rendered image to generate a masked image; adding, by the computing system, a set of noise to the rendered image to generate a noised image; processing, by the computing system, the noised image with a denoising model conditioned upon the masked image to generate adenoising prediction; evaluating, by the computing system, a loss function that comprises a distillation loss term that evaluates a difference between the set of noise and the denoising prediction; and modifying, by the computing system, one or more values of one or more parameters of the differentiable 3D scene representation model based on the loss function.
[0012] Other aspects of the present disclosure are directed to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.
[0013] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the related principles. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Detailed discussion of embodiments directed to one of ordinary skill in the art is set forth in the specification, which makes reference to the appended figures, in which:
[0015] Figure 1 depicts an example training framework for performing 3D scene content generation using 2D inpainting diffusion according to example embodiments of the present disclosure.
[0016] Figure 2 depicts an example training framework for performing 3D scene content generation using 2D inpainting diffusion according to example embodiments of the present disclosure.
[0017] Figure 3A depicts a block diagram of an example computing system according to example embodiments of the present disclosure.
[0018] Figure 3B depicts a block diagram of an example computing device according to example embodiments of the present disclosure.
[0019] Figure 3C depicts a block diagram of an example computing device according to example embodiments of the present disclosure.
[0020] Reference numerals that are repeated across plural figures are intended to identify the same features in various implementations. DETAILED DESCRIPTION Overview
[0021] The present disclosure provides presents a general approach to inpainting 3D content by using a 2D inpainting diffusion model trained on static scenes as a generativeprior. In particular, while existing methods for 3D inpainting focus on the specific task of object removal, systems and methods of the present disclosure aim to generate realistic content in any masked 3D region while preserving the complexity of the original scene at the object and scene scale. The proposed techniques can be used in various applications such as 3D object removal, 3D inpainting and novel view synthesis.
[0022] More particularly, the present disclosure provides a novel approach to 3D inpainting by distilling a 2D inpainting diffusion model into a learned 3D scene representation, e.g. a Neural Radiance Field (NeRF) model for 3D generation without the need for multiview data or camera poses. Given a set of images of a static scene and some 3D mask (e.g., a volume in world space), one objective of the proposed technique is to synthesize realistic / consistent 3D content in the unknown regions of the scene, while being able to reconstruct visible regions in the images.
[0023] To do so, some example implementations use a frozen 2D inpainting diffusion model that has been pretrained on a set of training data (e.g., domain-specific training data) to guide NeRF optimization. By tailoring the pretrained diffusion model to the downstream tasks and datasets of interest, some example implementations are able to define a distillation loss that encourages the NeRF to synthesize content that resembles samples from the diffusion model in unknown regions. This distillation loss can be optimized in conjunction with reconstruction losses that recover known regions visible in the input images.
[0024] Thus, some example aspects of the present disclosure are directed to a training system and method for training the NeRF model to generate novel content using the diffusion model as a generative prior. The training can be performed over a number of training iterations. In some implementations, each training iteration can include processing data descriptive of a camera pose with the NeRF model to generate a rendered image that depicts the 3D scene from the camera pose.
[0025] The training system can then apply an image mask to the rendered image to generate a masked image. The image mask can be obtained from 3D geometry or object segmentation. This masked image can then be used as a basis for generating new content in the 3D scene. The mask can be applied to any part of the image, allowing for flexibility in the areas that the user wants to modify or enhance.
[0026] The training system can then add a set of noise to the rendered image to generate a noised image. The noised image can then be processed with a denoising diffusion model, which is conditioned upon the masked image, to generate a denoising prediction. For example, the diffusion model can be pretrained on a dataset of training images depictingscenes associated with a particular domain. This can include indoor scenes, outdoor scenes, or any other type of environment.
[0027] The training system can then evaluate a loss function that includes a distillation loss term. This term evaluates the difference between the set of noise added to the rendered image and the denoising prediction generated by the diffusion model. The loss function serves as a measure of how well the model is performing, guiding the system in adjusting the parameters of the NeRF model.
[0028] The training system can modify the values of the parameters of the NeRF model based on the evaluated loss function. For example, this can be done by backpropagating the loss function through the denoising diffusion model while keeping the parameters of the diffusion model fixed. This process helps in optimizing the NeRF model to generate more accurate and realistic 3D content.
[0029] In some applications, the camera pose used in the method can be a novel camera pose unassociated with any ground truth images. This allows for the generation of new views of the 3D scene that have not been previously captured or seen. This feature can be particularly useful in applications such as virtual reality or 3D modeling where novel views of a scene are often required.
[0030] The present disclosure can also incorporate a reconstruction loss term in its loss function. This term compares the unmasked region of a ground truth image with a corresponding region of the rendered image. The ground truth image depicts the 3D scene from the camera pose and includes a masked region and an unmasked region defined by the image mask. The reconstruction loss term helps ensure that the generated 3D content is consistent with the unmasked regions of the original scene.
[0031] In some implementations, the predicted geometry of the 3D scene can be improved by supervising the NeRF depth using a depth estimator. The system generates an inferred depth map from the ground truth image and a rendered depth map by processing data descriptive of the camera pose with the NeRF model. The loss function then includes a depth loss term that evaluates the difference between the inferred depth map and the rendered depth map.
[0032] The method described in the present disclosure can be performed over a number of iterations, each associated with different camera poses. In some implementations, in the initial iterations, the camera pose can be fixed and only the reconstruction loss term is applied. This provides a more stable initialization of the MLP for applying the score distillation sampling. As the optimization progresses, the input view is randomly selected toencourage sample diversity and generate consistent diffusion model outputs. This iterative process ensures a comprehensive and accurate generation of the 3D scene content.
[0033] Notably, some example implementations choose to distill a 2D rather than 3D diffusion model because it enables the use of largescale image datasets of real world scenes. Widely used real world 3D datasets (or simulation data) can be used to train a 3D diffusion model, but are highly object-centric and are far from the complexity captured by image datasets.
[0034] Thus, the present disclosure provides a novel method for 3D scene content generation using 2D inpainting diffusion. This technology leverages the power of pretrained 2D diffusion models to guide 3D synthesis, providing a practical and efficient solution for generating realistic 3D content in masked regions of a scene. With its wide range of potential applications, this technology presents significant advancements in the field of 3D modeling and synthesis.
[0035] The present disclosure provides significant technical effects and benefits in the field of 3D content generation and modeling. One technical problem addressed by the invention is the generation of realistic 3D content in masked regions of a scene using 2D inpainting diffusion. This is a technical problem as it pertains to the field of computer graphics, specifically the generation and manipulation of 3D models based on 2D images. The present disclosure provides a technical solution that includes a novel combination of a Neural Radiance Field (NeRF) model and a diffusion model as a generative prior. This combination is used to guide the 3D synthesis without requiring any diffusion network retraining with multiview data. The proposed techniques have practical applicability in the field of computer graphics, especially in applications such as 3D object removal, 3D inpainting and novel view synthesis. It can also be useful in virtual reality or 3D modeling where novel views of a scene are often required. The proposed techniques therefore enable the generation of realistic 3D content in a more efficient and effective manner.
[0036] Another key technical effect is the increase in energy efficiency during the operation of machine-learned models. This efficiency is achieved by leveraging a pretrained 2D inpainting diffusion model to guide the 3D synthesis without requiring any network retraining with multiview data. This approach reduces the computational resources required for retraining, thereby reducing the energy consumption of the computing system. For instance, in traditional methods, retraining with multiview data is computationally expensive and time-consuming. However, the present disclosure eliminates this need, allowing the system to perform the same task using less energy.
[0037] With reference now to the Figures, example embodiments of the present disclosure will be discussed in further detail. Example 3D Scene Representations
[0038] The geometry and appearance of a scene can be modeled as a NeRF parameterized by a neural network (e.g., multilayer perceptron (MLP)). A NeRF is a learned 3D representation that maps any location in world space to RGB color ^^^^ and density ^^^^. To render a NeRF from a given camera to obtain an RGB image, (differentiable) volumetricraytracing is typically used. For each pixel, a ray defined by ^^^^(^^^^) = ^^^0^ + ^^^^^^^^^^^^ is cast from thecamera aperture ^^^0^ through the pixel center and 3D points are sampled along each ray. Thesepoints ^^^^^^^^, where ^^^^ ∈ {1, … ,^^^^}, are passed to the MLP to obtain output colors ^^^^^^^^ and densities^^^^^^^^. The colors and densities are alpha-composited to obtain the final rendered RGB image with pixels ^^^^(^^^^):
[0039] Similarly, a corresponding depth map ^�^^^(^^^^) can be computed by accumulating density along each ray:
[0040] NeRFs are typically trained by optimizing MLP parameters per-scene given a set of input images ^^^^^^^^and corresponding camera parameters. One example NeRF implementation is from Zip-NeRF [Barron et al.(2023)], which builds on improvements made by MipNerf360 [Barron et al.(2021)] and Instant-NGP [Muller et al.(2022)] to provide anti- aliasing and faster training than the original NeRF implementation. Example Diffusion Models
[0041] Diffusion models can learn a latent-variable generative model through a sequential generation process. Starting with the data distribution ^^^^0~^^^^(^^^^0), a sequence oflatent variables ^^^^1, … , ^^^^^^^^ can be constructed through a forward diffusion process ^^^^(^^^^1:^^^^|^^^^0).The forward process adds increasing amounts of Gaussian noise, where each latent ^^^^^^^^can beconstructed from a datapoint ^^^^0 and random noise ^^^^~^^^^(0, ^^^^):^^^^^^^^(^^^^0, ^^^^) = ^^^^(^^^^)^^^^0 + ^^^^(^^^^)^^^^, (4)where ^^^^(^^^^) and ^^^^2(^^^^) are coefficients chosen to reduce the signal-to-noise ratio of the data as ^^^^ increases while preserving the variance. A diffusion model can be trained to reverse this process and thus learn a generative model that synthesizes samples from noise. Learning the reverse process reduces to learning a function to predict the noise that was added to the data ^^^^^̂^^^:
[0042] Some example implementations of the present disclosure train a 2D inpaintingdiffusion model to predict ^^^^^̂^^^(^^^^^^^^; ^^^^, ^^^^), where ^^^^ is a masked image. An example model can betrained using the RealEstate10k dataset [Zhou et al.(2018)], which contains videos of static indoor and outdoor scenes. Example Score Distillation Sampling
[0043] Score distillation sampling (SDS) [Poole et al.(2022)] uses a pretrained diffusionmodel as a prior for optimizing parameters ^^^^ of a generator function ^^^^ = ^^^^(^^^^) that outputssamples resembling the learned data distribution of the diffusion model. The SDS loss gradients w.r.t. ^^^^ are defined aswhere in practice ^^^^ is randomly sampled at each optimization step and ^^^^(^^^^) is some timestep- dependent weighting function.
[0044] In some example implementations of the present disclosure, a NeRF serves as a differentiable image generator. ^^^^ is a volume-rendered RGB image, ^^^^ is the result of masking ^^^^ with 3D inpainting mask via ray-mask intersection, ^^^^^^^^, is the output of a single forwarddiffusion step applied to ^^^^, and ^^^^^̂^^^(^^^^^^^^;^^^^, ^^^^) is the output of a pre-trained 2D inpaintingdiffusion model. Summary of Example Techniques
[0045] Given a set of multi-view images along a motion path ℐ = {^^^^^^^^}^^^^^^^^=1 and 3D-consistent image masks ℳ = {^^^^^^^^}^^^^^^^^=1 , example implementations of the present disclosureaim to simultaneously synthesize realistic and 3D-consistent content in the masked regions of the reconstruction volume and accurately reconstruct the unmasked regions of the input images.
[0046] More particularly, it is not immediately evident how to use SDS to distill priors from a 2D inpainting diffusion model to guide 3D synthesis. In some possible alternative approaches, a text prompt is taken as input and provides a fixed conditioning signal to the diffusion model during NeRF optimization.
[0047] By contrast, inpainting diffusion models introduce a constraint that the conditioning image ^^^^ should be a masked version of the data sample ^^^^ that is passed through the forward diffusion process (and inpainted). Thus, since example implementations of the present disclosure are inpainting the 3D scene and not the 2D input images, example implementations of the present disclosure use the rendered images as the conditioning signal.
[0048] Example implementations of the present disclosure use a NeRF as the 3D scene representation and differentiable image generator for SDS. The unmasked regions of the 3D scene can be supervised using the RGB values from the input images, which ensures faithfulness to the ground truth scene while maintaining consistent geometry. Priors from the pre-trained 2D inpainting diffusion model can be distilled via SDS to inpaint new content in the masked region. Example Joint Synthesis and Reconstruction
[0049] Example implementations of the present disclosure assume that the input images ℐ have known camera parameters and define a mask function ^^^^^^^^(^^^^) that returns one for each ray in a set ^^^^ that intersects with the masked region and zero otherwise. To reconstruct the unmasked regions of the scene, example implementations of the present disclosure sample random rays ^^^^^^^^across all input images at each optimization step to obtain estimated colors ^^^^^^^^(^^^^^^^^) from the NeRF. A NeRF loss is then applied only to rays intersecting unmasked pixels:
[0050] In the masked regions, some example implementations of the present disclosure can apply SDS as follows. First, some example implementations of the present disclosuresample 256 × 256 patches from the rendered images ^^^^^^^^ that contain some overlap with theinpainting masks. Each rendered patch ^^^^^^^^is masked according to ^^^^^^^^(^^^^^^^^) in order to obtainthe diffusion conditioning signal ^^^^^^^^ = ^^^^^^^^(^^^^^^^^)^^^^^^^^. Per-view diffusion timestep ^^^^^^^^ and noise ^^^^^^^^are randomly sampled, and ^^^^^^^^,^^^^ is computed via ^^^^^^^^,^^^^ = ^^^^(^^^^^^^^)^^^^^^^^ + ^^^^(^^^^^^^^)^^^^^^^^ (as in Equation 4).The diffusion model inputs ^^^^^^^^ ,^^^^, ^^^^^^^^ are passed to the pre-trained inpainting diffusion model toobtain ^^^^^̂^^^(^^^^^^^^;^^^^, ^^^^^^^^), and the SDS loss as defined in Equation 6 is applied via
[0051] Some example implementations of the present disclosure apply additional volume rendering losses to both masked and unmasked rays. These can include distortion and interlevel losses from Mip-Nerf 360, as well as the patch-based depth smoothness loss, ℒDS, from RegNeRF [Niemeyer et al.(2022)] (e.g., modified to include a bilateral weighing term guided by the predicted colors). Thus, one example total loss function iswhere ^^^^1:5are loss weights. In practice, example implementations of the present disclosure constrain ^^^^2to be small relative to the other loss weights. Example Optimization Details
[0052] In some implementations, at the beginning of optimization, only Equation 7 is computed to provide a more stable initialization of the MLP for applying SDS. Some example implementations also fix the input view ^^^^ within a small baseline selected for SDS for a number of steps in order to converge on the inpainted content before handling multi- view consistency. Afterwards, the input views are randomly selected.
[0053] To encourage 3D consistency while still allowing for view-dependent effects, some example implementations can decompose the RGB NeRF prediction into diffuse and sparse components for the final RGB color prediction, where the diffuse component is predicted by a small MLP taking only the features computed from the input 3D point (but not ray direction). At the start of optimization, example implementations optionally use the diffuse component of the RGB image as input to the diffusion model to encourage consistent SDS updates across views, before annealing to both diffuse and specular components. This annealing increases the level of detail in the synthesized images for RealEstate10k scenes, where highly reflective materials are common.
[0054] In order to encourage sample diversity towards the beginning of optimization and generate consistent diffusion model outputs towards the end, some example implementationsof the present disclosure anneal the diffusion timestep such that at any optimization step ^^^^ ∈[0, ^^^^], the value of ^^^^ is sampled from^^^^(^^^^)~^^^^(^^^^ − ^^^^inv, ^^^^), where (10)
[0055] Some example implementations of the present disclosure use ^^^^min = 0.2, ^^^^max =0.95, and ^^^^inv = 0.2. Example implementations of the present disclosure optimize the NeRFover 16 GPUs, where at each step, each GPU takes an optimization step in parallel. Example Implementation Details
[0056] One example implementation uses a variant of Mip-Nerf 360, backed by Instant- NGP hash grids, as the view-synthesis backend. Specifically, example implementations of the present disclosure can use the scene contraction and proposal MLPs from Mip-Nerf 360, hash grids with coarsest and finest resolutions of 16 and 8192, and model exposure and scene lighting variations using GLO codes from NeRF-W [Martin-Brualla et al.(2021)].
[0057] Some example implementations of the present disclosure build on the network architecture from Palette [Saharia et al.(2022)] for the inpainting diffusion model and use aninput resolution of 256 × 256. During training, each image is preprocessed by taking either arandom 256px square crop or the largest center crop resized to 256px, and a random inpainting mask is generated. To generate inpainting masks, example implementations of the present disclosure use a combination of free-form strokes and rectangular masks. The conditioning signal is the input image with the masked region filled with Gaussian noise. Classifier-free guidance [Ho and Salimans (2022)] can optionally be used to enable tuning of sample quality vs. diversity for SDS.
[0058] Some example implementations of the present disclosure use a subset of the RealEstate10k dataset [Zhou et al.(2018)] for training the inpainting diffusion model and use a separate held-out test set for evaluating the 3D inpainting method. In total, the dataset consists of 10 million frames from 10,000 YouTube videos, and includes static indoor andoutdoor scenes. The native resolution of the images is 720 × 1280. The images were resizedto 1080 × 1920 and random 256 × 256 crops were used to train the diffusion model. For 3Dinpainting with the joint synthesis and reconstruction technique, the images weredownsampled to 360 × 640.
[0059] To compare with existing work in object removal (a subset of 3D inpainting scenarios), some example implementations were also evaluated on the SPIn-NeRF dataset [Mirzaei et al.(2023)], which consists of 10 static outdoor scenes captured from 60 views. Each frame also has a corresponding object mask that was manually annotated. Additionally, 40 views of each scene without the object were captured for evaluation. Example 3D Inpainting Masks
[0060] Various example implementations of the present disclosure can use a variety of different inpainting masks, examples of which are as follows.
[0061] Sphere Masks: Some example implementations position a sphere at a fixed distance along the optical axis of the center camera in the input path. Sphere-ray intersections can be computed at each input viewpoint in order to determine the projection of the sphere onto each input image. With these masks, example implementations can inpaint sizable 3D volumes of a scene with new content that is both 3d-consistent and makes semantic sense.
[0062] Object masks: Some example implementations first pre-train a NeRF in order to obtain depth information about the scene. For a single view, example implementations then select points which identify the objects to mask, and provide these as input to a segmentation model such as, for example, the Segment-Anything (SAM) model [Kirillov et al.(2023)]. Some example implementations choose the mask with the highest SAM score and reproject it to the other input views using the estimated NeRF depths. Finally, example implementations dilate the masks in order to fill holes and smooth edges. These masks demonstrate the ability of the proposed approach to be used for tasks such as object removal and replacement.
[0063] Scribble masks: Some example implementations can draw a random path through an input image, and dilate the points along the path with ellipses. As with the object masks, example implementations can first generate the mask in a single input view, then use the estimated depth information from a pre-trained NeRF to project it to other input views. These masks present an interesting set of challenges in that they often span a variety of depths from background to foreground, and partially occlude objects.
[0064] Outpainting masks: These masks can be created by inverting the sphere masks, such that example implementations maintain only a sphere in the center of the image as scene context. The inpainting model then outpaints the majority of the scene.Example Training Methodology Visualizations
[0065] Figure 1 illustrates a graphical representation of a potential embodiment of the method for 3D scene content generation using 2D inpainting diffusion. This figure can illustrate a simplified workflow of the method along with its key elements and steps, potentially aiding in understanding the operation of the described method.
[0066] The process can begin with input images 14 of a static scene and a corresponding 3D inpainting mask 12. The 3D inpainting mask 12 can be a volume in world space that defines the regions of the 3D scene that can need to be filled with newly generated content. The input images 14 can provide a visual representation of the 3D scene from different camera poses. These images can then be masked to isolate the regions of interest, which can be the areas of the scene that are to be inpainted.
[0067] Next, the method can utilize a diffusion model 16, which can be a specialized type of machine learning model trained on a dataset of static scenes. The diffusion model 16 can act as a generative prior for the inpainting process, guiding the synthesis of new content in the masked regions. The diffusion model 16 can be designed to generate realistic, high- quality images that can resemble the original scene's complexity and appearance.
[0068] In parallel, the method can employ a Neural Radiance Field (NeRF) model 18. The NeRF model 18 can be a learned 3D representation that maps any location in world space to color and density, possibly providing a comprehensive 3D representation of the scene. The parameters of the NeRF model 18 can be optimized using a combination of reconstruction losses and score distillation sampling, possibly resulting in an optimized NeRF 24.
[0069] The reconstruction loss 20 can be a measure of the difference between the unmasked regions of the input images 14 and corresponding regions of the rendered images generated by the NeRF model 18. This loss can encourage the NeRF model 18 to accurately reproduce the known parts of the scene, ensuring consistency between the original scene and the generated content.
[0070] The distillation loss 22, on the other hand, can be computed based on the difference between the noise added to the rendered images and the denoising prediction generated by the diffusion model 16. The distillation loss 22 can encourage the synthesized content in the masked regions to resemble samples from the diffusion model 16, leading to potentially more realistic and consistent inpainting results.
[0071] After the optimization process, the optimized NeRF 24 can be used to generate rendered novel views 26 of the scene, including the newly filled regions. In addition to thecolor information, the optimized NeRF 24 can also provide predicted depth maps 28, providing depth information for each pixel in the rendered images. This depth information can enhance the realism of the generated content and can be used in subsequent processing or viewing of the 3D scene.
[0072] Referring now to Figure 2, a diagram depicting an example embodiment of a joint training step using score distillation sampling in masked regions and a reconstruction loss in unmasked regions is shown. This figure illustrates the stages involved in generating an inpainted 3D scene from a 2D image using the methods described in the present disclosure.
[0073] The process begins with the selection of a camera pose 202, which provides a specific viewpoint of the 3D scene. This camera pose information is processed by a Neural Radiance Field (NeRF) model 204, which generates a rendered image 206 that depicts the 3D scene from the chosen camera pose. By processing the camera pose data, the NeRF model 204 can map any location in world space to RGB color and density, thereby generating a 3D representation of the scene as viewed from the camera pose 202.
[0074] Next, an image mask is applied to the rendered image 206, producing a masked image 208. The image mask can correspond to a 2D projection of a 3D mask that defines regions of the 3D scene that are to be inpainted. The masked image 208 therefore contains both known regions (which are unmasked) and unknown regions (which are masked and hence to be inpainted).
[0075] The joint training step further includes adding a set of noise to the rendered image 206 to generate a noised image 210. This noise addition stage perturbs the rendered image 206, creating a noised image 210.
[0076] The noised image 210 is then processed by a denoising diffusion model 212. The diffusion model 212, which has been conditioned upon the masked image 208, generates a denoising prediction 214. This prediction can represent the diffusion model's 212 estimate of what noise should be removed to create a noise-free version of the noised image 210.
[0077] The joint training step also involves evaluating a loss function that includes a distillation loss term 216 and a reconstruction loss term 218. The distillation loss term 216 evaluates the difference between the set of noise added to the rendered image 206 and the denoising prediction 214. The reconstruction loss term 218, on the other hand, compares the unmasked region of a ground truth image (not shown in Figure 2) with a corresponding region of the rendered image 206. These loss terms serve as measures of how well the NeRF model 204 is performing, guiding the system in adjusting the parameters of the NeRF model 204.
[0078] In some embodiments, the joint training step can optionally include additional supervision of NeRF depth using a monocular depth estimator. The method involves generating a rendered depth map 220 and an inferred depth map 222. The rendered depth map 220 is created by processing the camera pose data with the NeRF model 204, while the inferred depth map 222 is generated from the ground truth image. The depth loss term 224, which can optionally form part of the loss function, evaluates the difference between the inferred depth map 222 and the rendered depth map 220.
[0079] The joint training step depicted in Figure 2 therefore provides an approach to generating novel 3D content from 2D images. The use of score distillation sampling in masked regions, reconstruction loss in unmasked regions, and additional supervision of NeRF depth ensures that the generated 3D content is both realistic and consistent with the original scene. This methodology accomplishes this without requiring any network retraining with multiview data. This represents a practical and efficient solution for 3D content generation. Example Devices and Systems
[0080] Figure 3A depicts a block diagram of an example computing system 100 according to example embodiments of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled over a network 180.
[0081] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0082] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 which are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0083] In some implementations, the user computing device 102 can store or include one or more machine-learned models 120. For example, the machine-learned models 120 can be or can otherwise include various machine-learned models such as neural networks (e.g., deepneural networks) or other types of machine-learned models, including non-linear models and / or linear models. Neural networks can include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, neural radiance fields, or other forms of neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models). Example machine-learned models 120 are discussed with reference to Figures 1 and 2.
[0084] In some implementations, the one or more machine-learned models 120 can be received from the server computing system 130 over network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single machine-learned model 120.
[0085] Additionally or alternatively, one or more machine-learned models 140 can be included in or otherwise stored and implemented by the server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine-learned models 140 can be implemented by the server computing system 140 as a portion of a web service (e.g., an inpainting service). Thus, one or more models 120 can be stored and implemented at the user computing device 102 and / or one or more models 140 can be stored and implemented at the server computing system 130.
[0086] The user computing device 102 can also include one or more user input components 122 that receives user input. For example, the user input component 122 can be a touch-sensitive component (e.g., a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component can serve to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.
[0087] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 which areexecuted by the processor 132 to cause the server computing system 130 to perform operations.
[0088] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In instances in which the server computing system 130 includes plural server computing devices, such server computing devices can operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.
[0089] As described above, the server computing system 130 can store or otherwise include one or more machine-learned models 140. For example, the models 140 can be or can otherwise include various machine-learned models. Example machine-learned models include neural networks or other multi-layer non-linear models. Example neural networks include feed forward neural networks, deep neural networks, recurrent neural networks, convolutional neural networks, and / or neural radiance fields. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models). Example models 140 are discussed with reference to Figures 1 and 2.
[0090] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 via interaction with the training computing system 150 that is communicatively coupled over the network 180. The training computing system 150 can be separate from the server computing system 130 or can be a portion of the server computing system 130.
[0091] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 which are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.
[0092] The training computing system 150 can include a model trainer 160 that trains the machine-learned models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, forexample, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.
[0093] In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. The model trainer 160 can perform a number of generalization techniques (e.g., weight decays, dropouts, etc.) to improve the generalization capability of the models being trained.
[0094] In particular, the model trainer 160 can train the machine-learned models 120 and / or 140 based on a set of training data 162. In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some instances, this process can be referred to as personalizing the model.
[0095] The model trainer 160 includes computer logic utilized to provide desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into a memory and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions that are stored in a tangible computer-readable storage medium such as RAM, hard disk, or optical or magnetic media.
[0096] The network 180 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., Internet), or some combination thereof and can include any number of wired or wireless links. In general, communication over the network 180 can be carried via any type of wired and / or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0097] Figure 3A illustrates one example computing system that can be used to implement the present disclosure. Other computing systems can be used as well. For example, in some implementations, the user computing device 102 can include the model trainer 160 and the training dataset 162. In such implementations, the models 120 can be bothtrained and used locally at the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the models 120 based on user-specific data.
[0098] Figure 3B depicts a block diagram of an example computing device 10 that performs according to example embodiments of the present disclosure. The computing device 10 can be a user computing device or a server computing device.
[0099] The computing device 10 includes a number of applications (e.g., applications 1 through N). Each application contains its own machine learning library and machine-learned model(s). For example, each application can include a machine-learned model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.
[0100] As illustrated in Figure 3B, each application can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.
[0101] Figure 3C depicts a block diagram of an example computing device 50 that performs according to example embodiments of the present disclosure. The computing device 50 can be a user computing device or a server computing device.
[0102] The computing device 50 includes a number of applications (e.g., applications 1 through N). Each application is in communication with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and model(s) stored therein) using an API (e.g., a common API across all applications).
[0103] The central intelligence layer includes a number of machine-learned models. For example, as illustrated in Figure 3C, a respective machine-learned model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine-learned model. For example, in some implementations, the central intelligence layer can provide a single model for all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of the computing device 50.
[0104] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As illustrated in Figure 3C, the central device data layer can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API). Additional Disclosure
[0105] The technology discussed herein makes reference to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0106] While the present subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation, not limitation of the disclosure. Those skilled in the art, upon attaining an understanding of the foregoing, can readily produce alterations to, variations of, and equivalents to such embodiments. Accordingly, the subject disclosure does not preclude inclusion of such modifications, variations and / or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. For instance, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the present disclosure cover such alterations, variations, and equivalents.
Claims
WHAT IS CLAIMED IS:
1. A computer-implemented method to train a neural radiance field (NeRF) model to generate novel content for a three-dimensional (3D) scene, the method comprising: processing, by a computing system comprising one or more computing devices, data descriptive of a camera pose with the NeRF model to generate a rendered image that depicts the 3D scene from the camera pose; applying, by the computing system, an image mask to the rendered image to generate a masked image; adding, by the computing system, a set of noise to the rendered image to generate a noised image; processing, by the computing system, the noised image with a denoising diffusion model conditioned upon the masked image to generate a denoising prediction; evaluating, by the computing system, a loss function that comprises a distillation loss term that evaluates a difference between the set of noise and the denoising prediction; and modifying, by the computing system, one or more values of one or more parameters of the NeRF model based on the loss function.
2. The computer-implemented method of claim 1, wherein the denoising diffusion model has been previously trained on a dataset of training images that depict scenes associated with a particular domain.
3. The computer-implemented method of any preceding claim, wherein modifying, by the computing system, the one or more values of the one or more parameters of the NeRF model based on the loss function comprises backpropagating, by the computing system, the loss function through the denoising diffusion model while holding parameters of the denoising diffusion model fixed.
4. The computer-implemented method of any preceding claim, wherein the distillation loss term comprises a score distillation sampling term that evaluates the image mask combined with a weighting function combined with the difference between the set of noise and the denoising prediction combined with the rendered image.
5. The computer-implemented method of any of claims 1-4, wherein the camera pose comprises a novel camera pose unassociated with any ground truth images.
6. The computer-implemented method of any of claim 1-4, further comprising: obtaining, by the computing system, a ground truth image that depicts the 3D scene from the camera pose, wherein the ground truth image comprises a masked region and an unmasked region defined by the image mask; wherein the loss function further comprises a reconstruction loss term that compares the unmasked region of the ground truth image with a corresponding region of the rendered image.
7. The computer-implemented method of claim 6, further comprising: generating, by the computing system, an inferred depth map from the ground truth image; and processing, by the computing system, data descriptive of the camera pose with the NeRF model to generate a rendered depth map; wherein the loss function further comprises a depth loss term that evaluates a difference between the inferred depth map and the rendered depth map.
8. The computer-implemented method of any preceding claim, wherein the method is performed for a plurality of iterations respectively associated with a plurality of different camera poses.
9. The computer-implemented method of claim 8, wherein the camera pose is fixed for a number of initial iterations of the plurality of iterations.
10. The computer-implemented method of claim 8 or 9, wherein only the reconstruction loss term is applied for a number of initial iterations of the plurality of iterations.
11. The computer-implemented method of any preceding claim, wherein processing, by the computing system, the noised image with the denoising diffusion model conditionedupon the masked image to generate the denoising prediction comprises processing, by the computing system, the noised image with the denoising diffusion model conditioned upon the masked image and further conditioned upon a text prompt to generate the denoising prediction.
12. The computer-implemented method of any preceding claim, wherein the 3D scene comprises an indoor scene, and wherein the denoising diffusion model has been previously trained on a dataset of training images that depict indoor scenes.
13. The computer-implemented method of any preceding claim, wherein the image mask comprises a two-dimensional (2D) image mask derived from a 3D mask.
14. One or more non-transitory computer-readable media that collectively store a neural radiance field (NeRF) model that has been trained by performance of the method of any preceding claim.
15. A computing system comprising one or more processors and one or more non- transitory computer-readable media that collectively store computer-executable instructions for performing operations, the operations comprising: processing, by the computing system, data descriptive of a camera pose with a differentiable 3D scene representation model to generate a rendered image that depicts the 3D scene from the camera pose; applying, by the computing system, an image mask to the rendered image to generate a masked image; adding, by the computing system, a set of noise to the rendered image to generate a noised image; processing, by the computing system, the noised image with a denoising model conditioned upon the masked image to generate a denoising prediction; evaluating, by the computing system, a loss function that comprises a distillation loss term that evaluates a difference between the set of noise and the denoising prediction; and modifying, by the computing system, one or more values of one or more parameters of the differentiable 3D scene representation model based on the loss function.
Citation Information
Cited By
3dgs reconstruction noise restoration method, apparatus and device, and medium
CN121147054A
Virtual sonar image generation method and system
CN121414871A
Regularizing neural radiance fields with denoising diffusion models
US12505512B2
Regularizing neural radiance fields with denoising diffusion models
US20240265504A1