Shape and illumination using neural object decomposition via BRDF optimization in-the-wild

The proposed framework addresses the challenge of reconstructing 3D shape and material properties from in-the-wild images by using a hybrid encoding scheme that optimizes over shape, radiance, and pose, resulting in efficient and accurate reconstructions for graphics and AR/VR applications.

WO2025122724A1PCT designated stage expired Publication Date: 2025-06-12GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/058637
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-12-05
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing methods for reconstructing 3D shape and material properties from unconstrained in-the-wild image collections face challenges due to varying lighting, pose, background, and camera intrinsics, leading to inefficient and inaccurate reconstructions.

Method used

A novel end-to-end framework that optimizes over shape, radiance, and pose using a hybrid encoding scheme combining multiresolution hash grids and Fourier feature encodings, enabling efficient and robust reconstruction of 3D assets.

Benefits of technology

The framework achieves fast and accurate reconstruction of 3D shapes and material properties, reducing runtime and improving quality compared to previous methods, making it suitable for applications in AR/VR and graphics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024058637_12062025_PF_FP_ABST
    Figure US2024058637_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an advanced framework designed for the reconstruction of shape, material, and illumination from images captured with varying lighting, pose, and background. This framework addresses the challenge in computer vision and graphics of inverse rendering based on unconstrained image collections by optimizing over shape, radiance, and pose. The proposed framework can utilize a unique implicit shape representation based on a hybrid encoding scheme that includes both a multi-resolution hash encoding and Fourier feature encodings. This hybrid encoding scheme allows for rapid and robust shape reconstruction with joint camera alignment optimization.
Need to check novelty before this filing date? Find Prior Art

Description

SHAPE AND ILLUMINATION USING NEURAL OBJECT DECOMPOSITION VIABRDF OPTIMIZATION IN-THE-WILDRELATED APPLICATIONS

[0001] This application claims priority to and the benefit of United States Provisional Patent Application Number 63 / 606,396, filed December 5, 2023. United States Provisional Patent Application Number 63 / 606,396 is hereby incorporated by reference in its entirety.FIELD

[0002] The present disclosure relates generally to neural object representations. More particularly, the present disclosure relates to an end-to-end framework for reconstruction of shape, material and illumination from images captured with varying lighting, pose, background, camera intrinsics, and / or image resolution.BACKGROUND

[0003] The reconstruction of 3D shape and material properties of objects from unconstrained in-the-wild image collections has been a long-standing challenge in the field of computer vision and graphics. This challenge is exacerbated by the fact that images are often captured in different environments using a variety of devices, resulting in varying background, illumination, and camera intrinsics. Conventional structure-from-motion methods often fail to reconstruct under these challenging circumstances.

[0004] Many existing works on shape and material estimation assume constant camera intrinsics and an initialization of camera poses close to the true poses. However, this assumption may not always hold true, especially when working with in-the-wild image collections. Furthermore, existing methods for material decomposition with camera pose optimization are slow, often running for more than 12 hours on a single object.

[0005] Another issue is that many graphics applications in Augmented Reality (AR), Virtual Reality (VR), games, and movies depend on high-quality 3D assets of real-world objects. The conventional acquisition of these assets involves laborious tasks like 3D modelling, texture painting, camera, and light calibration, or requires controlled setups that are hard to scale.

[0006] Consequently, there is a need for a technique that allows for the efficient and effective reconstruction of 3D assets from unconstrained in-the-wild image collections.SUMMARY

[0007] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or can be learned from the description, or can be learned through practice of the embodiments.

[0008] One example aspect of the present disclosure is directed to a computer- implemented method to synthesize imagery of an object using a machine-learned neural radiance field model. The method is performed for one or more rendering iterations respectively associated with one or more pixels. The method includes obtaining, by a computing system comprising one or more computing devices, a set of camera parameter values. The method includes determining, by the computing system, a set of sampling coordinates based on the set of camera parameter values. The method includes determining, by the computing system, a first set of feature values for each sampling coordinate in the set of sampling coordinates, wherein the first set of feature values are obtained from a multiresolution hash grid. The method includes generating, by the computing system, a second set of feature values for each sampling coordinate in the set of sampling coordinates, wherein generating the second set of feature values for each sampling coordinate comprises performing a Fourier transform on the sampling coordinate. The method includes processing, by the computing system, a combination of the first set of feature values and the second set of feature values with a neural network of the machine-learned neural radiance field model to render a color value for the pixel.

[0009] Example implementations can include any combination of the following features. In some implementations, generating the second set of feature values for each sampling coordinate further comprises processing a set of Fourier values that result from the performance of the Fourier transform on the sampling coordinate with a second neural network to generate the second set of feature values. In some implementations, the combination of the first set of feature values and the second set of feature values comprises a concatenation of the first set of feature values and the second set of feature values. In some implementations, during training of the neural network, an annealing process was performed in which resolution levels were progressively added to the multi-resolution hash grid and frequency bands were progressively added to the Fourier encoding. In some implementations, the set of camera parameter values consists of an eye position, an up rotation angle, and a direction parameter. In some implementations, obtaining, by the computing system, the set of camera parameter values further comprises jittering, by the computing system, the set of camera parameter values to generate a multiplex of cameras; determining, by the computingsystem, the set of sampling coordinates based on the set of camera parameter values further comprises projecting, by the computing system, at least some of the set of sampling coordinates associated with at least some cameras from the multiplex of cameras into a highest ranking camera of the multiplex of cameras; and / or the method further comprises evaluating a multiplex loss that compares colors rendered for projected sampling coordinates with colors rendered for unprojected sampling coordinates. In some implementations, the method further comprises applying, by the computing system, a per- view importance w eighting term during training of the machine-learned neural radiance field model. In some implementations, the method further comprises using, by the computing system, a patch-level loss on rendered color output and alpha density estimate to train the machine-learned neural radiance field model.

[0010] Other aspects of the present disclosure are directed to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices, for example configured to perform and / or storing computer-executable instructions for performing any of the methods described herein.

[0011] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the related principles.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Detailed discussion of embodiments directed to one of ordinary skill in the art is set forth in the specification, which makes reference to the appended figures, in which:

[0013] Figure 1 depicts a graphical representation of an example framework for learning a neural object representation according to example embodiments of the present disclosure.

[0014] Figure 2 depicts a graphical representation of an example constrained camera multiplex according to example embodiments of the present disclosure.

[0015] Figures 3A-C depict graphical representations of an example silhouette-based alignment loss according to example embodiments of the present disclosure.

[0016] Figure 4 depicts a flow chart diagram of an example method to synthesize imagery' of an object according to example embodiments of the present disclosure.

[0017] Figure 5 A depicts a block diagram of an example computing system according to example embodiments of the present disclosure.

[0018] Figure 5B depicts a block diagram of an example computing device according to example embodiments of the present disclosure.

[0019] Figure 5C depicts a block diagram of an example computing device according to example embodiments of the present disclosure.

[0020] Reference numerals that are repeated across plural figures are intended to identify the same features in various implementations.DETAILED DESCRIPTIONOverview

[0021] The present disclosure provides an advanced framework designed for the reconstruction of shape, material, and / or illumination from images captured with varying lighting, pose, and / or background. This framework addresses the challenge in computer vision and graphics of inverse rendering based on unconstrained image collections by optimizing over shape, radiance, and pose. The proposed framework can utilize a unique implicit shape representation based on a hybrid encoding scheme that includes both a multiresolution hash encoding and Fourier feature encodings. This hybrid encoding scheme allows for rapid and robust shape reconstruction with joint camera alignment optimization.

[0022] Further, the framework enables the editing of illumination and object reflectance (material) by jointly optimizing a Bidirectional Reflectance Distribution Function (BRDF) and illumination along with the object's shape. Notably, the proposed framework is classagnostic and works with in-the-wild image collections to produce relightable 3D assets suitable for several use-cases, including AR / VR.

[0023] The proposed framework is designed to handle images captured in various environments using different devices, resulting in varying backgrounds, illuminations, camera intrinsics, and / or resolutions. The images can be obtained via casual acquisition or the use of existing online image collections (e.g. from an image search performed using a websearch engine). Example implementations of the present disclosure also optimize camera parameters and per-image illumination along with a physically plausible decomposition of 3D shape and spatially -varying BRDF material properties for arbitrary camera intrinsics. Furthermore, example implementations integrate multiresolution hash grids into the pipeline and constrain it using a hybrid encoding, thus enabling more rays to be processed in a shorter time during optimization.

[0024] Significant features of the proposed framework include a hybrid encoding with multiresolution hash encoding with level annealing, a modified camera parameterization, acamera multiplex with proj ection constraint, a per-view importance weighting, and patchbased alignment losses. The framework demonstrates improved performance and reduced runtime compared to previous methods, making it a valuable tool for various graphics applications.

[0025] More particularly, one example aspect of the present disclosure is directed to a novel hybrid encoding approach that combines a multi-resolution hash grid and Fourier coordinate mapping to render a color value for each pixel in an image. This approach allows for a more robust and efficient reconstruction of 3D shapes and material properties.

[0026] In one example method disclosed herein, a set of camera parameter values is obtained for each rendering iteration associated with a pixel (e.g.. one set of parameters per input view, updated every rendering iteration). This set of camera parameters can include an eye position, an up rotation angle, a direction parameter, and / or a focal length. The camera parameter values determine the perspective from which the obj ect will be rendered.

[0027] The method further involves determining a set of sampling coordinates based on the camera parameter values. For example, raymarching operations can be performed based on a set of image coordinates importance sampled according to the foreground mask. The sampling coordinates represent the points in 3D space that the method will sample to determine the color value of a pixel. The sampling coordinates can be determined in various ways, including through geometric calculations or machine learning algorithms.

[0028] The present disclosure also introduces a unique way of determining feature values for each sampling coordinate. A first set of feature values is obtained from a multiresolution hash grid, a data structure that allows for efficient storage and retrieval of spatial data. This hash grid can be implemented at multiple resolutions, allowing for a balance between detail and computational efficiency.

[0029] In addition to the hash grid, the method also employs a Fourier transform to generate a second set of feature values for each sampling coordinate. In particular, the Fourier transformation can be applied to the 3D coordinates. The Fourier transform is a mathematical method that decomposes a function into its constituent frequencies, providing a different perspective on the data. This dual approach of using both a hash grid and a Fourier transform enables the method to capture both high-frequency and low-frequency information about the object being rendered. This dual approach also provides an improved gradient flow' for the parameter update.

[0030] The feature values obtained from the hash grid and the Fourier transform are then combined and processed by a neural network. This neural network, part of the machine-learned neural radiance field model, uses the feature values to render a color value for the pixel. The combination of the feature values can be achieved in various ways, such as through concatenation or other mathematical operations.

[0031] The present disclosure also describes an annealing process performed during the training of the neural network. In this process, resolution levels are progressively added to the multi-resolution hash grid. This approach allows the neural network to gradually leam to handle higher-resolution data, improving its performance and robustness.

[0032] Another example aspect of the present disclosure is directed to a camera multiplexing technique, in which multiple cameras are jittered around the initial camera parameters. This technique allows the method to explore different perspectives on the object, reducing the chance of the optimization process getting stuck in local minima. The method also includes a multiplex loss that compares colors rendered for projected sampling coordinates with colors rendered for unprojected sampling coordinates, further improving the robustness of the optimization process.

[0033] Another example aspect of the present disclosure is directed to a per- view importance weighting term and a patch-level loss used in the training of the machine-learned neural radiance field model. The per-view importance weighting term allows the method to give more weight to views that are more useful for optimization, while the patch-level loss aids in camera alignment. These features further enhance the robustness and efficiency of the method.

[0034] The systems and methods of the present disclosure provide a number of technical effects and benefits. As one example, the proposed techniques address the issue of reconstructing the 3D shape, material properties, and / or illumination of objects from unconstrained in-the-wild image collections, an issue that is exacerbated by the varying lighting, pose, and background conditions under which the images are captured. This is a technical problem as it involves the processing and analysis of digital image data to generate a 3D model.

[0035] The disclosed techniques provide a solution to this problem by providing a novel end-to-end framework that optimizes over shape, radiance, and pose to perform the reconstruction. This framework can utilize a unique implicit shape representation based on a hybrid feature representation that leverages both a multi-resolution hash encoding and a Fourier feature encoding, which enables fast and robust shape reconstruction. This is a technical solution as it involves the use of specific data structures and optimization techniques to enhance the performance and accuracy of the 3D reconstruction process.

[0036] With reference now to the Figures, example embodiments of the present disclosure will be discussed in further detail.Example Reconstruction Techniques

[0037] The aim of some example implementations of the present disclosure is to convert 2D image collections into a 3D representation with minimal manual work. The representation includes shape, material parameters and per-view illumination, allowing for view synthesis with relighting.

[0038] Example Problem Setup

[0039] In-the-wild data can include a collection of q images Q G Istx3; j e {1,that show the same object captured with different backgrounds, illuminations and cameras with potentially varying resolutions Sj. In addition, some example implementations assume a rough camera initialization. Some example implementations annotate camera pose quadrants as described in SAMURAI [Boss et al., SAMURAI: Shape And Material from Unconstrained Real -world Arbitrary Image collections. NeurlPS, 2022], Foreground masks can be added if available or automatically generated and might be imperfect at this point. At each point x G IR3in the neural volume V. some example implementations estimate the BRDF parameters for the Cook-Torrance model b G I5(basecolor bcG IR3, metallic bmG IR, roughness brG IR), unit-length surface normal n G IR3and volume density <7 G IR. To enable the decomposition, some example implementations also estimate the latent per-image illumination vectors zj G IR128; j G {1, ... , q}. Furthermore, some example implementations estimate per-image camera poses and intrinsics.

[0040] Description of Related Work

[0041] The next paragraphs provide a brief overview of: NeRF [Mildenhall et al., NeRF: Representing scenes as neural radiance fields for view' synthesis. ECCV, 2020], InstantNGP [Muller et al., Instant neural graphics primitives with a multiresolution hash encoding. ACM TOG, 2022] and SAMURAI.

[0042] Coordinate-based MLPs and NeRF uses a dense neural network to model a continuous function that takes 3D location x G I3and view' direction d G I3and outputs a view'-dependent output color c G IR3and volume density o’ G IR. Mildenhall et al overcome the spectral bias of the MLPs by transforming the input coordinates by a second function; A frequency encoding y that maps from IR to IR2L:y(x) = (sin(2°7ix), COS(2°TTX),... , sin(2L-17rx), COS(2L-1TTX))

[0043] InstantNGP speed up the NeRF optimization drastically by replacing the MLPbased volume representation by a multiresolution voxel hash grid that is tailored to current GPU hardware. For a hash-size T, grid vertices are indexed by a spatial hash function (%) = ( a \ j mod? using large unique prime numbers ny. At each voxel vertex a d-dimensional i=l / embedding is optimized. Instead of the Fourier embedding, the 3D coordinates x are directly used to tri-linearly interpolate between neighboring vertices at each level. The results are concatenated and fed to a MLP to decode the representation. Some example implementations denote the full encoding function including interpolation and concatenation as H(x).

[0044] SAMURAI is a method for joint optimization of 3D shape, BRDF, per-image camera parameters, and illuminations for a given in-the-wild image collection. SAMURAI follows the NeRF idea outlined above but uses the Neural-PIL [Boss et al., Neural-pil: Neural pre-integrated lighting for reflectance decomposition. NeurlPS, 2021] method for physically- based differentiable rendering. It takes 3D locations as input and outputs volume density and BRDF parameters. An additional GLO (generative latent optimization) embedding models the changes in appearances (due to different illuminations) across images. Neural-PIL introduced the use of per-image latent illumination embedding zj and a specialized illumination pre-integration (PIL) network for fast rendering, which some example implementations refer to as ‘PIL rendering’. NeuraLPIL optimizes a per-image embedding to model image-specific illumination. The rendered output color c is equivalent to NeRF’s output c, but due to the explicit BRDF decomposition and illumination modeling, it enables relighting and material editing. To address the unavailability of accurate camera parameters for in-the-wild images. SAMURAI jointly optimizes camera extrinsics and per-view intrinsics from a very coarse initialization. In addition to a coarse-to-fine annealing, this is achieved with a multiplexed optimization scheme where multiple camera proposals per view are kept and weighted according to their performance on the loss over time.

[0045] Example Optimization with Hash Encoding

[0046] Some example implementations of the present disclosure identify misaligned and inconsistent camera poses as the main limiting factor for in-the-wild reconstructions. Joint shape and camera optimization is a severely underdetermined problem. Reconstruction is ty pically slow and often lacks high-frequency detail in textures and shape. Multiresolution hash grids have the potential to speed up the reconstruction while simultaneously allowingfor larger ray counts to be processed and thereby improving visual quality and alignment. However, the naive replacement of the point encoding with Hash grids reduces the reconstruction quality and robustness of the joint camera and shape optimization.

[0047] Hash grids adapt to individual views faster resulting in a noisy shape in the presence of misaligned cameras. As reported previously, multi-resolution hash grids with the default linear interpolation backpropagate noisy and discontinuous gradients with respect to the input position. Additionally, the coarse-to-fine scheme from BARF [Lin et al,. BARF: Bundle- Adjusting Neural Radiance Fields. ICCV, 2021] often used for camera fine-tuning cannot be directly transferred to hash grids. Therefore, some example implementations propose an approach that makes use of a camera multiplex, adds additional geometrical constraints, and a new encoding scheme to be able to improve both reconstruction speed and quality. Next, each of these components is explained in further detail.

[0048] Example Architecture

[0049] A high-level overview of the architecture of some example implementations of the present disclosure is shown in Figure 1. As shown in Figure 1, two resolution annealed encoding branches, the multiresolution hash grid H(x) and the Fourier embedding y(x) can be used to learn a neural volume conditioned on the input coordinates. This enables robust optimization of camera parameters jointly with the shape, material and illumination.

[0050] More particularly, unlike prior works, some example implementations map the input coordinates x using a new hybrid encoding. The combined embedding can be processed by a neural network (e.g., a small MLP) to predict the density cr, and the view and appearance conditioned radiance for a given image patch. Some example implementations also predict a regular direction-dependent radiance c to stabilize the early training stages. The BRDF decoder can operate as in SAMURAI, expanding the feature representation to the BRDF (base color, metallic, roughness). Per sample, some example implementations estimate normal direction from the first order derivative of the density w.r.t. the input positionFrom there the volumetric rendering from NeRF can be performed and the shading for the given pixel coordinate is determined using BRDF, normals and the pre-integrated illumination estimated by the NeuralPIL network. The example architecture shown in Figure is described in further detail below .

[0051] Example Camera Pose Initialization And Parameterization

[0052] Camera pose optimization is a highly non-convex problem and tends to quickly get stuck in local minima. In some settings, initial camera poses are much noisier and featurelarger distances between initial and true poses compared to many related works. To combat this, some example implementations assume a rough initialization in the form of camera pose quadrants in line with SAMURAI. Some example implementations use a ‘lookat + direction’ representation for the camera parameters, storing initial values and offsets for an eye position peyeG IR3, lookat direction Ad^ G IR2. and up rotation angle dupG IR as well as the focal length f G IR per camera. In some settings, this removes the overparameterization regarding the rotation component encoded in eye and center position of the regular ‘lookat’ parameterization.

[0053] Example Hybrid Positional Encoding

[0054] Some example implementations use a hash grid hybrid as coordinate encoding to improve the gradient flow w.r.t the input coordinates x. A Fourier-based coordinate mapping y(x) followed by a neural network (e.g., a small MLP) generates a base embedding that is concatenated with the output of the multiresolution hash grid H(x) resulting in the following formulation of the neural volume F® ((H(x), y(x))). On y, some example implementations apply BARF’s Fourier annealing. Similarly, some example implementations progressively add resolution levels to the hash grid encoding. Starting with only the features from a low7resolution dense grid, some example implementations increase the weights of the higher resolution levels gradually over time.

[0055] Example Camera Multiplexes

[0056] An effective way to reduce the chance of camera pose optimization to be stuck in local minima is the camera multiplex. In some implementations, for each image, m cameras are jittered around the initial camera and simultaneously optimized. Over time the worst performing camera is repeatedly faded out until m = 1. This process is visualized in Figure 2. In particular, as shown in Figure 2, some example implementations can optimize multiple camera proposals per image and weight the contribution to the reconstruction according to a camera’s performance on the loss. Between cameras of a multiplex, some example implementations can add a projection based regularization: Points from all members are projected into the currently best camera and then compared against a new7render to enforce a consistent geometry.

[0057] More particularly, since some example implementations render multiple proposals for a given image anyway, some example implementations can further constrain the optimization using projective geometry. Specifically, some example implementations proj ect the 2D point sets XLrendered by the m — 1 members into the currently highestranking camera Ooof the multiplex using the estimated depth Di from the volumetric rendering. Then some example implementations render the projected coordinates using Ooand compare the rendered color and alpha values a, of all cameras in the multiplex to the ones originally rendered at 6X m-1.where Pi Qis the perspective w arp from image coordinates in camera i to the reference camera. Fvis the rendering function connected to the neural field outputting color c and mask value a. respectively. This regularization comes roughly at the cost of adding a camera to the multiplex. Subsampling of Xtcan decrease the memory footprint if needed. This component may be controlled to be active while there are multiple cameras rendered during the first part of the overall schedule. Used as an additional loss it turns out to be surprisingly effective in constraining the camera optimization and therefore increasing the robustness of the overall optimization. Essentially, some example implementations are enforcing a consistent surface to be generated and smooth the optimization landscape around an initial camera pose.

[0058] Example View Importance Scaling Of Input Images

[0059] Not every input might contribute to the reconstruction in the same way and individual views that are not aligned with the current 3D shape might have a negative impact on the overall optimization progress. To improve high-frequency detail in the reconstruction some example implementations reduce the impact of potentially misaligned cameras while anchoring the optimization using cameras that work well given the loss. Some example implementations keep a circular buffer of around 1000 elements with the recent per-image losses. Like in SAMURAI, this can be used to re-weigh images in the given collection according to:sD,m(3)with the mean / q and standard deviation atof the loss buffer. This limits the influence of badly aligned camera poses on the shape reconstruction. In addition, some example implementations also apply an importance weighting on £camerathat reduces the gradientmagnitude for views that are performing well given the loss history. Specifically, at step t so

[0060] In practice, some example implementations set the hyperparameterpto 0.05.

[0061] Example Losses and Optimization

[0062] Example multiscale patch loss: After a short initial phase of random raysampling, some example implementations render randomly sampled patches of size 16x16 to 32x32. The goal is to constrain the updates and especially the alignment to be consistent on local neighborhoods. Therefore, some example implementations add a multi-scale patch loss on the rendered color c which computes a Charbonnier loss at four different resolution levels, by simple bilinear resampling. Some example implementations weigh each level to compensate for the different pixel counts and enforce the low-resolution version to align first.

[0063] Example mask losses: Some example implementations add a silhouette loss ^silhouette whenever patch-based sampling is active. For example, Figure 3 shows an example silhouette-based alignment loss that can penalize the unaligned pixels given a reference and the rendered gray scale masks. In particular, some example implementations penalize the area between the two silhouettes which can be interpreted as the result of an xor operation on the rendered and input mask. Both masks can be filtered using a Gaussian blur where the radius is heuristically chosen based on the patch size. Figure 3 visualizes how the loss helps with the alignment task. Some example implementations combine this loss with a regular binary-cross-entropy loss on the mask value as well as a loss enforcing a transparent background.

[0064] Example regularization losses: To regularize the hash grid encoding some example implementations apply a normalized weight decay to put a higher penalty on coarser grid levels compared to naive weight decay. Additionally, some example implementations apply regularization to the camera poses and normal output.

[0065] Example Optimization: Some example implementations use three optimizers: e.g., one optimizer (e.g., ADAM) for the networks, hash grid embeddings and cameras, respectively. The learning rate can be decayed exponentially on all optimizers. In addition to the camera representation and constraints mentioned above, some example implementations use ADAM with the / ?1 value reduced to 0.2 to smooth out the noise in the camera updates. In some instances, the learning rate can be tuned between le-3 to 2e-3 depending on scenesize. Render resolution can be continuously increased over the first half of the optimization while the number of active multiplex cameras is reduced. The direct color optimization can be faded to the BRDF optimization and the encoding annealing can be performed over the first third of the optimization. Focal length updates and the view importance weighting can be delayed until an initial shape has been formed.Example Framework Visualization

[0066] Referring again to Figure 1, provided is a graphical representation of an example hybrid encoding system. The system combines a multi-resolution hash grid encoding H(x) 12 and a Fourier feature transformation y(x) 14 to create a robust and efficient encoding scheme. Thus, the hybrid encoding system shown in Figure 1 is represented as two separate branches. The first branch represents the multi-resolution hash grid encoding H(x) 12, while the second branch represents the Fourier feature transformation y(x) 14.

[0067] The branch representing the multi-resolution hash grid encoding H(x) 12 is depicted in the upper branch of Figure 1. The hash grid encoding is a data structure that allows for the efficient storage and retrieval of spatial data. It can be implemented at multiple resolutions, which offers a balance between detail and computational efficiency. This encoding method is particularly advantageous for rendering high-resolution images or objects with intricate details. This encoding method also provides higher speed as the hash grids align well with the architecture of modem GPUs and a neural network 16 (e.g.. a small network such as one or more MLPs) can be used for decoding afterwards.

[0068] The multi -resolution hash grid H(x) 12 can take in 3D coordinates x 18 as input. These coordinates 18 are then used to retrieve features from the hash grid 12 at varying resolutions. This encoding process can involve several steps, including mapping the input coordinates 18 to corresponding cells in the hash grid 12 and retrieving feature values from each cell. The exact process of encoding can vary depending on the specific implementation and the level of detail required. However, some example implementations can perform interpolation of encodings from nearby cells to generate the final output. For example, the grid 12 can be parameterized as a voxel grid and interpolation can be performed using values retrieved from the grid.

[0069] The second branch of the hybrid encoding system, represented as the lower branch in Figure 1, illustrates the Fourier feature transformation y(x) 14. The Fourier transformation 14 is a mathematical method that decomposes a function into its constituent frequencies. This transformation allows for the capture of both high-frequency and low-frequency information about the object being rendered, adding another layer of detail to the encoding process.

[0070] In the Fourier feature transformation 14, the 3D coordinates x 18 are transformed into a higher-dimensional space using a set of sinusoidal basis functions 20. This transformation can be performed using various techniques, including but not limited to the use of sine and cosine functions. The transformed coordinates can optionally then be fed into a neural network 15 (e.g., a small neural network such as one or more MLPs). This neural network 15 processes the transformed coordinates, further refining the data for use in subsequent stages of the method. For example, the coordinate can be transformed into a low7resolution approximation of the common feature space (e.g., shared with hash grid).

[0071] As shown in Figure 1. the feature values obtained from both the hash grid encoding 12 and the Fourier feature transformation 14 can be combined. This combination can be achieved in various ways, such as through concatenation or other mathematical operations. The combined feature values, which contain rich information about the object's shape, material, and illumination, are then used as input to the neural network 16. The neural network 16 can be one network or can be multiple networks.

[0072] The neural network 16 is responsible for rendering BRDF values 22 and density values 24 for each pixel. For example, the neural network 16 can predict BRDF and density per volume sample which is then transformed to a per pixel set of values during the rendering process which comes afterwards. For example, The BRDF values 22 and density values 24 for each pixel can be combined with illumination information 26 from a neural prior to render a color Cj 28 for the pixel.

[0073] The neural network 16 can be trained using various techniques, including but not limited to backpropagation and gradient descent. The neural network 16 can also incorporate various architectural elements, such as hidden layers, activation functions, and dropout layers, to improve its performance.

[0074] Furthermore, Figure 1 illustrates at 30 the process of annealing applied to the hybrid encoding system. In this process, resolution levels are progressively added to the multi-resolution hash grid 12. This approach allows the neural network 16 to gradually learn to handle higher-resolution data, improving its performance and robustness. This progressive addition of resolution levels can be controlled by a set of parameters, which can be adjusted to optimize the performance of the encoding system. This process also improves the cameraparameter optimization as it enables coarse alignment in the beginning with smooth gradients.Example Methods

[0075] Figure 4 illustrates a flow chart diagram of an example method 400 for synthesizing imagery’ of an object using a machine-learned neural radiance field model. The method 400 is executed by a computing system comprising one or more computing devices.

[0076] The method 400 begins at step 402, where the computing system obtains a set of camera parameter values. These camera parameter values can include, for example, eye position, up rotation angle, and direction parameter. These values determine the perspective from which the object is rendered.

[0077] At step 404, the computing system determines a set of sampling coordinates based on the set of camera parameter values obtained in step 402. The sampling coordinates represent points in 3D space that are sampled to determine the color value of a pixel.

[0078] Proceeding to step 406, the computing system determines a first set of feature values for each sampling coordinate in the set of sampling coordinates. This first set of feature values is obtained from a multi-resolution hash grid. A multi-resolution hash grid can be implemented at various resolutions to balance between detail and computational efficiency.

[0079] At step 408, the computing system generates a second set of feature values for each sampling coordinate in the set of sampling coordinates. The generation of the second set of feature values includes performing a Fourier transform on the sampling coordinate. The Fourier transform can be a mathematical method that decomposes a function into its constituent frequencies.

[0080] Finally, at step 410, the computing system processes a combination of the first set of feature values and the second set of feature values with a neural network of the machine- learned neural radiance field model to render a color value for the pixel. The combination of feature values can involve concatenation or other mathematical operations.

[0081] In some implementations of the present disclosure, the method 400 can be performed iteratively to render data for a number of different pixels. The method can be repeated for each pixel in an image or a set of images, thereby enabling the synthesis of imagery' of an object from different perspectives and under vary ing lighting conditions. Each iteration can use different camera parameter values or the same values adjusted slightly toaccount for changes in perspective, thereby enhancing the detail and accuracy of the rendered imagery.Example Devices and Systems

[0082] Figure 5 A depicts a block diagram of an example computing system 100 according to example embodiments of the present disclosure. The system 100 includes a user computing device 102. a server computing system 130. and a training computing system 150 that are communicatively coupled over a network 180.

[0083] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0084] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory' devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 which are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0085] In some implementations, the user computing device 102 can store or include one or more machine-learned models 120. For example, the machine-learned models 120 can be or can otherwise include various machine-learned models such as neural networks (e.g., deep neural networks) or other types of machine-learned models, including non-linear models and / or linear models. Neural networks can include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural netw orks or other forms of neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models).

[0086] In some implementations, the one or more machine-learned models 120 can be received from the server computing system 130 over netw ork 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 can implementmultiple parallel instances of a single machine-learned model 120 (e.g., to perform parallel neural rendering across multiple instances of camera parameters).

[0087] Additionally or alternatively, one or more machine-learned models 140 can be included in or otherwise stored and implemented by the server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine-learned models 140 can be implemented by the server computing system 140 as a portion of a web service (e.g.. a neural rendering service). Thus, one or more models 120 can be stored and implemented at the user computing device 102 and / or one or more models 140 can be stored and implemented at the server computing system 130.

[0088] The user computing device 102 can also include one or more user input components 122 that receives user input. For example, the user input component 122 can be a touch-sensitive component (e.g., a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component can serve to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.

[0089] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory' 134 can store data 136 and instructions 138 which are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0090] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In instances in which the server computing system 130 includes plural server computing devices, such serv er computing devices can operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.

[0091] As described above, the server computing system 130 can store or otherwise include one or more machine-learned models 140. For example, the models 140 can be or can otherwise include various machine-learned models. Example machine-learned models include neural networks or other multi-layer non-linear models. Example neural networksinclude feed forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models).

[0092] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 via interaction with the training computing system 150 that is communicatively coupled over the network 180. The training computing system 150 can be separate from the server computing system 130 or can be a portion of the server computing system 130.

[0093] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory7154 can include one or more non-transi lory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 which are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0094] The training computing system 150 can include a model trainer 160 that trains the machine-learned models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.

[0095] In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. The model trainer 160 can perform a number of generalization techniques (e.g., weight decays, dropouts, etc.) to improve the generalization capability of the models being trained.

[0096] In particular, the model trainer 160 can train the machine-learned models 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, collections of in-the-wild images of one or more objects.

[0097] In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some instances, this process can be referred to as personalizing the model.

[0098] The model trainer 160 includes computer logic utilized to provide desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into a memory and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions that are stored in a tangible computer-readable storage medium such as RAM, hard disk, or optical or magnetic media.

[0099] The network 180 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., Internet), or some combination thereof and can include any number of wired or wireless links. In general, communication over the network 180 can be carried via any type of wired and / or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP, HTTP. SMTP. FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0100] Figure 5 A illustrates one example computing system that can be used to implement the present disclosure. Other computing systems can be used as well. For example, in some implementations, the user computing device 102 can include the model trainer 160 and the training dataset 162. In such implementations, the models 120 can be both trained and used locally at the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the models 120 based on user-specific data.

[0101] Figure 5B depicts a block diagram of an example computing device 10 that performs according to example embodiments of the present disclosure. The computing device 10 can be a user computing device or a server computing device.

[0102] The computing device 10 includes a number of applications (e.g., applications 1 through N). Each application contains its own machine learning library and machine-learned model(s). For example, each application can include a machine-learned model. Exampleapplications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0103] As illustrated in Figure 5B, each application can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0104] Figure 5C depicts a block diagram of an example computing device 50 that performs according to example embodiments of the present disclosure. The computing device 50 can be a user computing device or a server computing device.

[0105] The computing device 50 includes a number of applications (e.g., applications 1 through N). Each application is in communication with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and model(s) stored therein) using an API (e.g., a common API across all applications).

[0106] The central intelligence layer includes a number of machine-learned models. For example, as illustrated in Figure 5C, a respective machine-learned model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine-learned model. For example, in some implementations, the central intelligence layer can provide a single model for all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of the computing device 50.

[0107] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As illustrated in Figure 5C, the central device data layer can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).Additional Disclosure

[0108] The technology discussed herein makes reference to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0109] While the present subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation, not limitation of the disclosure. Those skilled in the art, upon attaining an understanding of the foregoing, can readily produce alterations to, variations of, and equivalents to such embodiments. Accordingly, the subject disclosure does not preclude inclusion of such modifications, variations and / or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. For instance, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the present disclosure cover such alterations, variations, and equivalents.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method to synthesize imagery of an object using a machine-learned neural radiance field model, the method comprising: for one or more rendering iterations respectively associated with one or more pixels: obtaining, by a computing system comprising one or more computing devices, a set of camera parameter values; determining, by the computing system, a set of sampling coordinates based on the set of camera parameter values; determining, by the computing system, a first set of feature values for each sampling coordinate in the set of sampling coordinates, wherein the first set of feature values are obtained from a multi-resolution hash grid; generating, by the computing system, a second set of feature values for each sampling coordinate in the set of sampling coordinates, wherein generating the second set of feature values for each sampling coordinate comprises performing a Fourier transform on the sampling coordinate; and processing, by the computing system, a combination of the first set of feature values and the second set of feature values with a neural network of the machine-learned neural radiance field model to render a color value for the pixel.

2. The computer-implemented method of any preceding claim, wherein generating the second set of feature values for each sampling coordinate further comprises processing a set of Fourier values that result from the performance of the Fourier transform on the sampling coordinate with a second neural network to generate the second set of feature values.

3. The computer-implemented method of any preceding claim, wherein the combination of the first set of feature values and the second set of feature values comprises a concatenation of the first set of feature values and the second set of feature values.

4. The computer-implemented method of any preceding claim, wherein, during training of the neural network, an annealing process was performed in which resolution levelswere progressively added to the multi-resolution hash grid and frequency bands were progressively added to the Fourier encoding.

5. The computer-implemented method of any preceding claim, wherein the set of camera parameter values consists of an eye position, an up rotation angle, and a direction parameter.

6. The computer-implemented method of any preceding claim, wherein: obtaining, by the computing system, the set of camera parameter values further comprises jittering, by the computing system, the set of camera parameter values to generate a multiplex of cameras; determining, by the computing system, the set of sampling coordinates based on the set of camera parameter values further comprises projecting, by the computing system, at least some of the set of sampling coordinates associated with at least some cameras from the multiplex of cameras into a highest ranking camera of the multiplex of cameras; and the method further comprises evaluating a multiplex loss that compares colors rendered for projected sampling coordinates with colors rendered for unprojected sampling coordinates.

7. The computer-implemented method of any preceding claim, further comprising applying, by the computing system, a per-view importance weighting term during training of the machine-learned neural radiance field model.

8. The computer-implemented method of any preceding claim, further comprising using, by the computing system, a patch-level loss on rendered color output and alpha density estimate to train the machine-learned neural radiance field model.

9. One or more non-transitory computer-readable media that collectively store computer-executable instructions for performing the method of any preceding claim.

10. A computing system configured to perform the method of any of claims 1-8.