Parameter interpolation for real-time radiance field rendering

By interpolating spatially-structured parameter values based on the camera's location within a three-dimensional volume and using these values to run a NeRF model, the method addresses the challenges of achieving high visual fidelity and capturing view-dependent effects in real-time radiance field rendering.

WO2025129069A1PCT designated stage expired Publication Date: 2025-06-19GOOGLE LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/US2024/060130
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-12-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current methods for real-time radiance field rendering face challenges in achieving high visual fidelity while managing computational costs and model capacity, particularly in capturing view-dependent effects like specular highlights.

Method used

The method involves obtaining a current camera location within a three-dimensional volume, accessing spatially-structured parameter values associated with the region containing the camera, interpolating these values based on the camera's location, and using the interpolated parameters to run a neural radiance field (NeRF) model for real-time rendering.

Benefits of technology

This approach enhances the model capacity of real-time neural radiance fields, improving the visual fidelity of rendered images with reduced computational costs, and effectively captures view-dependent effects such as specular highlights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024060130_19062025_PF_FP_ABST
    Figure US2024060130_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are systems and methods for real-time radiance field rendering. In particular, example implementations of the present disclosure can learn and use a set of spatially-structured parameter values. For example, the parameter values can be structured so that certain subsets of the parameter values are associated with each of a number of different regions of a three-dimensional volume. At rendering, some of the spatially-structured parameter values can be retrieved based on the location of a virtual camera within the volume. The retrieved parameters values can then be interpolated to generate a set of interpolated parameter values for a neural radiance field (NeRF) model. This model is then run with the interpolated parameters to render one or more pixels depicting the scene from the current camera location.
Need to check novelty before this filing date? Find Prior Art

Description

PARAMETER INTERPOLATION FOR REAL-TIME RADIANCE FIELD RENDERINGRELATED APPLICATIONS

[0001] This application claims priority to and the benefit of United States Provisional Patent Application Number 63 / 609,597, filed December 13, 2023. United States Provisional Patent Application Number 63 / 609,597 is hereby incorporated by reference in its entirety.FIELD

[0002] The present disclosure relates generally to machine learning models known as neural radiance fields. More particularly, the present disclosure relates to parameter interpolation for real-time radiance field rendering.BACKGROUND

[0003] The field of real-time radiance field rendering presents a significant technical challenge, particularly in the context of achieving high visual fidelity while managing computational costs and model capacity. Traditional methods use a single, often computationally expensive model to render an image, which results in a trade-off between rendering speed and the realistic quality of the rendered images. That is, the larger and more computationally demanding the model, the more realistic the rendered images tend to be. However, this approach is not scalable and is often constrained by hardware limitations, which results in a trade-off between visual fidelity and computational efficiency.

[0004] Thus, a particular technical problem in this field is the limitation of the model capacity of a real-time neural radiance field, which negatively impacts the visual fidelity of the rendered images. Current best-performing methods are typically limited to a small amount of per-point and per-view computation, leaving a significant gap in the potential for improving the quality of the rendered images.

[0005] In addition, another technical issue is the inability of traditional methods to efficiently capture and render view-dependent effects, such as specular highlights, in realtime radiance field reconstruction of indoor spaces. Current methods generally struggle to accurately model these effects, which further limits the realistic quality of the rendered images.SUMMARY

[0006] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or can be learned from the description, or can be learned through practice of the embodiments.

[0007] One example aspect of the present disclosure is directed to a computer- implemented method to perform real-time radiance field rendering. The method includes obtaining, by a computing system comprising one or more computing devices, a current camera location within a three-dimensional volume associated with a scene. The method includes accessing, by the computing system, a set of spatially-structured parameter values associated with a region of the three-dimensional volume that contains the current camera location. The method includes interpolating, by the computing system, the set of spatially- structured parameter values based on a location of the current camera location within the region to generate a set of interpolated parameter values for a neural radiance field (NeRF) model. The method includes running, by the computing system, the NeRF model with the set of interpolated parameter values to render one or more pixels depicting the scene from the current camera location.

[0008] Another example aspect of the present disclosure is directed to a computing system for real-time radiance field rendering, the computing system includes one or more processors and one or more non-transitory computer-readable media that collectively store computer-executable instructions for performing operations. The operations include: obtaining, by the computing system comprising one or more computing devices, a training camera location within a three-dimensional volume associated with a scene; accessing, by the computing system, a set of spatially-structured parameter values associated with a region of the three-dimensional volume that contains the training camera location; and interpolating, by the computing system, the set of spatially-structured parameter values based on a location of the current camera location within the region to generate a set of interpolated parameter values for a neural radiance field (NeRF) model; running, by the computing system, the NeRF model with the set of interpolated parameter values to render a predicted pixel color; and modifying, by the computing system, at least some of the set of spatially-structured parameter values associated with the region of the three-dimensional volume that contains the training camera location based at least in part on a loss function that compares the predicted pixel color with a target pixel color.

[0009] Other aspects of the present disclosure are directed to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.

[0010] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the related principles.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Detailed discussion of embodiments directed to one of ordinary skill in the art is set forth in the specification, which makes reference to the appended figures, in which:

[0012] Figure 1 depicts a flow chart diagram of an example method to perform real-time rendering according to example embodiments of the present disclosure.

[0013] Figure 2 depicts a flow chart diagram of an example method to train a neural radiance field model according to example embodiments of the present disclosure.

[0014] Figure 3A depicts a block diagram of an example computing system according to example embodiments of the present disclosure.

[0015] Figure 3B depicts a block diagram of an example computing device according to example embodiments of the present disclosure.

[0016] Figure 3C depicts a block diagram of an example computing device according to example embodiments of the present disclosure.

[0017] Figures 4A-C depict graphical diagrams of example coordinate systems according to example embodiments of the present disclosure.

[0018] Figure 5 depicts a graphical diagram of example teacher supervision according to example embodiments of the present disclosure.

[0019] Figure 6 depicts a graphical diagram of example ray jittering according to example embodiments of the present disclosure.

[0020] Reference numerals that are repeated across plural figures are intended to identify the same features in various implementations.DETAILED DESCRIPTIONOverview

[0021] Generally, the present disclosure is directed to systems and methods for real-time radiance field rendering. In particular, example implementations of the present disclosure can learn and use a set of spatially-structured parameter values. For example, the parameter values can be structured so that certain subsets of the parameter values are associated witheach of a number of different regions of a three-dimensional volume. At rendering, some of the spatially-structured parameter values can be retrieved based on the location of a virtual camera within the volume. The retrieved parameters values can then be interpolated to generate a set of interpolated parameter values for a neural radiance field (NeRF) model. This model is then run with the interpolated parameters to render one or more pixels depicting the scene from the current camera location.

[0022] More particularly, a three-dimensional volume can be partitioned into a number of different regions. For example, the volume can be partitioned into a voxel grid. When performing rendering (e.g., either during training or at test time), parameter values associated with the region (e.g., voxel) that contains the current camera location can be accessed and used by the NeRF model to perform rendering. In particular, the retrieved parameters can be interpolated based on the location of the camera within the region (e.g., voxel). This process can streamline the rendering process by reducing the number of parameters that need to be accessed and interpolated for each rendering operation.

[0023] The interpolation of the spatially-structured parameter values can be performed in a number of different manners. As one example, the computing system can perform trilinear interpolation of the set of spatially-structured parameter values based on the location of the current camera location within the region. This type of interpolation can provide a high degree of accuracy in the rendered image, as it takes into account the position of the camera within the three-dimensional space. Linear interpolation also has the benefit of being differentiable, enabling updates to the interpolated parameters to be scattered back to the underlying spatially-structured parameters.

[0024] In particular, training of the NeRF model in this framework can include modifying at least some of the set of spatially-structured parameter values associated with the region of the three-dimensional volume that contains a training camera. This modification can be based at least in part on a loss function that compares the predicted pixel color with a target pixel color. This training operation can further improve the accuracy and realism of the rendered images.

[0025] In summary, the present disclosure provides a novel and efficient method for real-time radiance field rendering. By interpolating spatially-structured parameter values based on the camera's location within a three-dimensional volume, the method can generate high-quality rendered images with a relatively low computational cost. The method can be implemented in a computer system and can be used in a wide range of applications that require real-time rendering of three-dimensional scenes.

[0026] The proposed techniques provide a technical solution to the problem of real-time radiance field rendering in a three-dimensional (3D) space. This solution introduces a novel method for parameter interpolation that enhances the model capacity of a real-time neural radiance field, thereby improving the visual fidelity of the rendered images. This is achieved by assigning each of a number of 3D regions in volumetric space a set of neural network parameters, which are then used by a model (e.g., a deferred rendering network) to predict a target RGB color.

[0027] As one example, a low-resolution voxel grid covering a 3D volume of interest is established, with neural network parameters assigned to the voxel corners. For a given camera position, trilinear interpolation is applied directly on the network parameters to determine the parameters for the current camera location. This computational operation can be performed once per frame, allowing it to be efficiently executed as part of a real-time rendering pipeline.

[0028] The proposed method also provides a technical solution to the issue of capturing and rendering view-dependent effects, such as specular highlights, in real-time radiance field reconstruction. The use of a voxel grid means that each model is localized to a small volume of space, allowing the model to specialize and achieve higher capacity than can be achieved with a single network. This, in turn, leads to more accurate and realistic rendered images.

[0029] With reference now to the Figures, example embodiments of the present disclosure will be discussed in further detail.Example Methods

[0030] Figure 1 depicts a flow chart diagram of an example method to perform real-time rendering according to example embodiments of the present disclosure.

[0031] In the first step of the process, as shown in block 12 of Figure 1, a computing system obtains a current camera location within a three-dimensional volume associated with a scene. The scene can be any three-dimensional scene that is to be rendered, such as an indoor or outdoor environment, a virtual reality scene, a video game scene, and the like. The camera location can be a virtual camera that is used to capture the scene from a specific viewpoint. The camera location can be specified in any suitable manner, such as by Cartesian coordinates, spherical coordinates, or any other suitable coordinate system.

[0032] In block 14 of Figure 1, the computing system accesses a set of spatially- structured parameter values associated with a region of the three-dimensional volume that contains the current camera location. These spatially-structured parameter values can includeany suitable parameters that are used by the NeRF model to render the scene. The region associated with the current camera location can be a voxel in a three-dimensional voxel grid that covers the entire three-dimensional volume.

[0033] In one example, the voxel grid can be a regular grid, where each voxel has the same size and shape, or an irregular grid, where the voxels can have different sizes and / or shapes. The size and shape of the voxels can be determined based on various factors, such as the complexity of the scene, the desired level of detail in the rendered image, the available computational resources, and the like. In other implementations, the voxel grid can be adaptively refined, where the size and shape of the voxels can change dynamically based on the current camera location, the movement of the camera, the changes in the scene, and the like.

[0034] Thus, in some implementations, the three-dimensional volume is partitioned into a voxel grid, which can encompass the entire scene space. Each voxel can be associated with a set of spatially-structured parameter values. To access these parameter values, the computing system identifies the voxel that contains the current camera location. The system then accesses the set of spatially-structured parameter values associated with this particular voxel. This targeted access approach significantly streamlines the rendering process by reducing the amount of data that needs to be processed, thereby enabling the system to produce high-quality real-time renders.

[0035] In block 16 of Figure 1, the computing system interpolates the set of spatially- structured parameter values based on the location of the current camera location within the region to generate a set of interpolated parameter values for the NeRF model. The interpolation can be performed using any suitable interpolation technique, such as linear interpolation, bilinear interpolation, trilinear interpolation, nearest-neighbor interpolation, and the like. The interpolation can take into account the spatial relationship between the current camera location and the voxel comers, the voxel centers, the voxel faces, or any other suitable points within the voxel.

[0036] In particular, in some implementations, the computing system employs trilinear interpolation of the set of spatially-structured parameter values according to the location of the current camera within the region. Trilinear interpolation is a method of multivariate interpolation on a 3-dimensional regular grid. This technique allows the interpolation of values within a voxel, taking into account the values at each of the eight comers of the voxel. The trilinear interpolation operation is based on the relative position of the current camera location within the voxel. This involves a sequence of linear interpolations first along the x-axis, then the y-axis, and finally the z-axis. Through this trilinear interpolation process, the set of spatially-structured parameter values is interpolated to generate a set of interpolated parameter values for the NeRF model.

[0037] Referring still to Figure 1, in block 18, the computing system runs the NeRF model with the set of interpolated parameter values to render one or more pixels depicting the scene from the current camera location. The NeRF model can be any suitable neural network model that is capable of rendering a three-dimensional radiance field, such as a multilayer perceptron (MLP), and / or the like. The NeRF model can output a set of pixel colors that represent the scene from the current camera location.

[0038] In some implementations, the NeRF model used in the present disclosure comprises a set of predetermined feature values and a deferred Tenderer network. The predetermined feature values can be any suitable features that are used by the NeRF model to render the scene, such as features related to the light source, the material properties of the objects in the scene, the camera properties, and the like. These feature values can be learned during a training phase, where the NeRF model is trained on a dataset of images of the scene with known (or jointly optimized) camera locations. Once the model is trained, these feature values can be used to render the scene from any given camera location. The deferred Tenderer network is an additional part of the NeRF model that is responsible for predicting the final pixel colors based on the interpolated parameter values and the predetermined feature values. This network can be a multilayer perceptron (MLP) or any other suitable type of neural network. Running the NeRF model with the set of interpolated parameter values can include using the interpolated parameter values as the parameters of the Tenderer network to process some of the predetermined feature values to output the final pixel colors that represent the scene from the current camera location.

[0039] In some implementations, the operations of accessing the spatially-structured parameter values and interpolating them based on the current camera location within the region are performed once per image rendered from the current camera location. This introduces a significant computational efficiency in the rendering process. Rather than repeatedly accessing and interpolating the parameter values for each pixel in the image, the system performs these computations once and applies the resultant interpolated parameter values to the entire image. This approach not only reduces the computational load but also helps to maintain consistency across the image, as all pixels in a given rendered image are derived from the same interpolated parameter values. This contributes to the overall visualcoherence and realism of the rendered image, enhancing the quality of the real-time radiance field rendering.

[0040] According to another aspect of the present disclosure, in some implementations, the set of spatially-structured parameter values associated with a region of the three- dimensional volume can have been learned via distillation from a teacher NeRF model. The teacher NeRF model can be a more complex and computationally expensive model that has been trained to render the scene with a high degree of accuracy and realism. The teacher NeRF model can generate a set of teacher parameter values that are used to guide the learning of the spatially-structured parameter values. The distillation process can involve training the NeRF model to mimic the output of the teacher NeRF model by minimizing a loss function that measures the difference between the output of the NeRF model and the output of the teacher NeRF model. This distillation process can enhance the performance of the NeRF model by leveraging the high-quality rendering capabilities of the teacher NeRF model, thus improving the visual fidelity of the rendered images.

[0041] In some implementations, the set of spatially-structured parameter values associated with a region of the three-dimensional volume have been learned via application of a consistency loss that enforces visual consistency between parameter values associated with adjacent regions of the three-dimensional volume. This consistency loss can encourage the NeRF model to produce similar renderings for adjacent regions, thereby reducing potential visual artifacts and discontinuities in the rendered images. The consistency loss can be computed based on the difference between the renderings produced by the NeRF model for adjacent regions, and can be minimized during the training of the NeRF model. This process can be performed for some or all pairs of adjacent regions in the three-dimensional volume, thereby encouraging the NeRF model to produce a consistent and coherent rendering of the entire scene. This consistency loss can provide an additional supervision signal that optionally complements the distillation from the teacher model, leading to improved visual fidelity and realism in the rendered images.

[0042] Referring now to Figure 2, Figure 2 depicts a flow chart diagram of an example method to train a neural radiance field model according to example embodiments of the present disclosure. The method begins at block 202, where the computing system obtains a training camera location within a three-dimensional volume associated with a scene. The training camera location can be any point in the three-dimensional space from which a view of the scene is to be rendered.

[0043] At block 204, the computing system accesses a set of spatially-structured parameter values associated with a region of the three-dimensional volume that contains the training camera location. The region can be any volume of space within the three- dimensional volume that encompasses the training camera location. The spatially-structured parameter values can be numerical values that are associated with the region and are used in the generation of the neural radiance field model.

[0044] In some embodiments, the three-dimensional volume is partitioned into a voxel grid. Each voxel within the grid can correspond to a distinct region of the three-dimensional volume, and each voxel can be associated with a set of spatially-structured parameter values. In such embodiments, accessing the set of spatially-structured parameter values associated with the region can involve identifying the voxel that contains the current camera location and accessing the spatially-structured parameter values associated with that voxel.

[0045] At block 206, the computing system interpolates the set of spatially-structured parameter values based on the location of the current camera location within the region to generate a set of interpolated parameter values for a neural radiance field (NeRF) model. The interpolation process can involve using the location of the current camera location within the region to determine a weighted combination of the spatially-structured parameter values associated with the region. The weights used in the combination can be determined based on the relative distances between the current camera location and points within the region associated with the spatially-structured parameter values. In some embodiments, the interpolation process involves performing trilinear interpolation of the set of spatially- structured parameter values based on the location of the current camera location within the region.

[0046] At block 208, the computing system runs the NeRF model with the set of interpolated parameter values to render a predicted pixel color. The NeRF model can be any computational model that generates a color value for a pixel based on a set of parameter values. In some embodiments, the NeRF model comprises a set of feature values and a deferred Tenderer network. Running the NeRF model with the set of interpolated parameter values can involve using the interpolated parameter values as parameter values of the deferred Tenderer network when processing the feature values to generate the predicted pixel color.

[0047] Finally, at block 210, the computing system modifies at least some of the set of spatially-structured parameter values associated with the region of the three-dimensional volume that contains the training camera location based at least in part on a loss function thatcompares the predicted pixel color with a target pixel color. The loss function can be any function that quantifies the difference between the predicted pixel color and the target pixel color. In some embodiments, the target pixel color comprises a color value predicted by a teacher NeRF model, which can be a pre-existing or previously trained NeRF model that is used as a reference for training the current NeRF model. In other embodiments, the target pixel color can be a color value predicted by a neighboring NeRF model generated by interpolating spatially-structured parameter values associated with a neighboring region of the three-dimensional volume.

[0048] In some implementations, modifying the set of spatially-structured parameter values associated with the region involves scattering a gradient of the loss function to invert the interpolation. This can involve computing the gradient of the loss function with respect to the predicted pixel color, and using this gradient to adjust the spatially-structured parameter values in a direction that reduces the loss function. This process can be repeated iteratively, with the spatially-structured parameter values being updated in each iteration based on the gradient of the loss function, until the loss function reaches a minimum value or some other termination condition is met.

[0049] In this way, the method illustrated in Figure 2 allows for the efficient training of a neural radiance field model for real-time radiance field rendering. By interpolating spatially- structured parameter values based on the location of a camera within a three-dimensional volume, the method can generate high-quality rendered images with a relatively low computational cost. Furthermore, by using a loss function to compare the rendered images with target images, the method can iteratively refine the spatially-structured parameter values to improve the accuracy and realism of the rendered images.Example Devices and Systems

[0050] Figure 3 A depicts a block diagram of an example computing system 100 according to example embodiments of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled over a network 180.

[0051] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0052] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 which are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0053] In some implementations, the user computing device 102 can store or include one or more machine-learned models 120. For example, the machine-learned models 120 can be or can otherwise include various machine-learned models such as neural networks (e.g., deep neural networks) or other types of machine-learned models, including non-linear models and / or linear models. Neural networks can include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks or other forms of neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models). Example machine-learned models 120 are discussed with reference to Figures 1-2.

[0054] In some implementations, the one or more machine-learned models 120 can be received from the server computing system 130 over network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single machine-learned model 120 (e.g., to perform parallel neural rendering across multiple instances of Tenderers).

[0055] Additionally or alternatively, one or more machine-learned models 140 can be included in or otherwise stored and implemented by the server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine-learned models 140 can be implemented by the server computing system 140 as a portion of a web service (e.g., a neural rendering service). Thus, one or more models 120 can be stored and implemented at the user computing device 102 and / or one or more models 140 can be stored and implemented at the server computing system 130.

[0056] The user computing device 102 can also include one or more user input components 122 that receives user input. For example, the user input component 122 can be a touch-sensitive component (e.g., a touch-sensitive display screen or a touch pad) that issensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component can serve to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.

[0057] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 which are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0058] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In instances in which the server computing system 130 includes plural server computing devices, such server computing devices can operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.

[0059] As described above, the server computing system 130 can store or otherwise include one or more machine-learned models 140. For example, the models 140 can be or can otherwise include various machine-learned models. Example machine-learned models include neural networks or other multi-layer non-linear models. Example neural networks include feed forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models). Example models 140 are discussed with reference to Figures 1-2.

[0060] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 via interaction with the training computing system 150 that is communicatively coupled over the network 180. The training computing system 150 can be separate from the server computing system 130 or can be a portion of the server computing system 130.

[0061] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., aprocessor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 which are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0062] The training computing system 150 can include a model trainer 160 that trains the machine-learned models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.

[0063] In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. The model trainer 160 can perform a number of generalization techniques (e.g., weight decays, dropouts, etc.) to improve the generalization capability of the models being trained.

[0064] In particular, the model trainer 160 can train the machine-learned models 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, multiple images of a scene.

[0065] In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some instances, this process can be referred to as personalizing the model.

[0066] The model trainer 160 includes computer logic utilized to provide desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into a memory and executed by one or more processors. In other implementations, the model trainer 160includes one or more sets of computer-executable instructions that are stored in a tangible computer-readable storage medium such as RAM, hard disk, or optical or magnetic media.

[0067] The network 180 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., Internet), or some combination thereof and can include any number of wired or wireless links. In general, communication over the network 180 can be carried via any type of wired and / or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0068] Figure 3 A illustrates one example computing system that can be used to implement the present disclosure. Other computing systems can be used as well. For example, in some implementations, the user computing device 102 can include the model trainer 160 and the training dataset 162. In such implementations, the models 120 can be both trained and used locally at the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the models 120 based on user-specific data.

[0069] Figure 3B depicts a block diagram of an example computing device 10 that performs according to example embodiments of the present disclosure. The computing device 10 can be a user computing device or a server computing device.

[0070] The computing device 10 includes a number of applications (e.g., applications 1 through N). Each application contains its own machine learning library and machine-learned model(s). For example, each application can include a machine-learned model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0071] As illustrated in Figure 3B, each application can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0072] Figure 3C depicts a block diagram of an example computing device 50 that performs according to example embodiments of the present disclosure. The computing device 50 can be a user computing device or a server computing device.

[0073] The computing device 50 includes a number of applications (e.g., applications 1 through N). Each application is in communication with a central intelligence layer. Exampleapplications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and model(s) stored therein) using an API (e.g., a common API across all applications).

[0074] The central intelligence layer includes a number of machine-learned models. For example, as illustrated in Figure 3C, a respective machine-learned model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine-learned model. For example, in some implementations, the central intelligence layer can provide a single model for all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of the computing device 50.

[0075] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As illustrated in Figure 3C, the central device data layer can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).Example Memory-Efficient Radiance Fields

[0076] This section includes a review of MERE (Reiser et al., MERF: Memory-Efficient Radiance Fields for Real-time View Synthesis in Unbounded Scenes, arXiv:2302.12249 (2023)). MERF maps 3D positions x 6 IR3to feature vectors t 6 IR8. MERF parameterizes this mapping using a combination of high-resolution triplanes (Px, Py, PzG IRRXRX8) and a low-resolution sparse voxel grid V G ]R>LXLXLX8query point x is projected onto each of three axis-aligned planes and the underlying 2D grid is queried via bilinear interpolation. Additionally, a trilinear sample is taken from the sparse voxel grid. The resulting four 8- dimensional vectors are then summed:

[0077] This vector is then unpacked into three parts, which are independently rectified to yield a scalar density, a diffuse RGB color, and a feature vector that encodes viewdependence effects:T = exp(tx), c = sigmoid(t2:4), f = sigmoid(t5:8). (2)

[0078] To render a pixel, a ray is cast from that pixel’s center of projection o along the viewing direction d and sampled at a set of distances {t to generate a set of points along that ray x = o + t d. The densities {T are converted into alpha compositing weights {w using the numerical quadrature approximation for volume rendering [?]:where <5, = ti+1— ttis the distance between adjacent samples. After alpha compositing, a deferred shading approach is used to decode the blended diffuse RGB colorswtct and the blended view-dependent color featureswt^t into the final pixel color with the help of the small, deferred rendering MLP h(-; 0):where 6 are the MLP’s parameters.

[0079] In unbounded scenes, far-away content can be modelled coarsely. To achieve a resolution that drops off with the distance from the scene’s focus point, MERF applies a contraction function to each spatial position x before querying the feature field:Example Streamable Memory Efficient Radiance Fields

[0080] Although real-time view-synthesis methods like MERF perform well for a localized environment, they often fail to scale to large multi-room scenes. To this end, some example implementations of the present disclosure leverage a hierarchical architecture. First, some example implementations partition the coordinate space of the entire scene into a series of blocks, where each block is modeled by its own MERF -like representation. Second, some example implementations introduce a grid of spatially-anchored network parameters within each block for modeling view-dependent effects. Finally, some example implementations introduce a gating mechanism for modulating high- and low-resolution contributions to a location’s feature representation. Thus, one example overall architecture can be thought of as a three-level hierarchy: based on the camera origin, (i) some example implementations selectan appropriate submodel, then within a submodel (ii) some example implementations compute the parameters of a deferred appearance network via interpolation, and then within a local voxel neighborhood, (iii) some example implementations determine a location’s feature representation via feature gating.

[0081] This greatly increases the capacity of the proposed model without diminishing rendering speed or increasing memory consumption: even as total storage requirements increase with the number of submodels, only a single submodel is needed to render a given frame. As such, when implemented on a graphics accelerator, the proposed system maintains modest resource requirements comparable to MERF .

[0082] Example Coordinate Space Partitioning Techniques

[0083] While MERF offers sufficient capacity for faithfully representing medium-scale scenes, the use of a single set of triplanes limits its capacity and reduces image quality. In large scenes, numerous surface points project to the same 2D plane location, and the representation therefore struggles to simultaneously represent high-frequency details of multiple surfaces. Although this can be partially ameliorated by increasing the spatial resolution of the underlying representation, doing so significantly impacts memory consumption and is prohibitively expensive in practice.

[0084] Therefore, some example implementations of the present disclosure instead coarsely subdivide the scene into a 3D grid based on camera origin and associate each grid cell with an independent submodel. Each submodel is assigned its own contraction space (Eq. (5)) and is tasked with representing the region of the scene within its grid cell at high detail, while the region outside each submodel’s cell is modeled coarsely. Note that, typically, the entire scene is still represented by each submodel — the submodels differ only in terms of which portions of the scene lies inside or outside of each submodel’s contraction region. As such, rendering a camera only requires a single submodel, implying that only one submodel must be in memory at a time.

[0085] Formally, some example implementations shift and scale all training cameras to lie within a [— K / 2, K / 2]3cube, and then partition this cube into K3identical and tightly packed subvolumes of size [— 1,1]3. Some example implementations can assign training cameras to submodels {<Sfc} by identifying the associated subvolume Rfethat each camera origin o lies within:

[0086] Some example implementations configure the proposed camera-to-submodel assignment procedure to apply to cameras outside of the training set, ensuring its validity during test set rendering. This enables a wide range of features including ray jittering, submodel reassignment, a submodel consensus loss, arbitrary test camera placement, and ping-pong buffers.

[0087] Rather than exhaustively instantiating submodels for all K3subvolumes, some example implementations only consider subvolumes that contain at least one training camera. As most scenes are outdoors or single-story buildings, this reduces the number of submodels from K3to O(K2).

[0088] Example Deferred Appearance Network Partitioning Techniques

[0089] The second level in the proposed partitioning hierarchy concerns the deferred rendering model. Recall that MERF employs a small multi-layer perceptron (MLP) to decode view-dependent colors from blended features as described in Eq. (4). Although the small size of this network is advantageous for fast inference, in some specific cases its capacity is insufficient to accurately reproduce complex view-dependent effects common in larger scenes. Simply increasing the size of this network is not viable as doing so would significantly reduce rendering speed.

[0090] Instead, some example implementations uniformly subdivide the domain of each submodel into a lattice with P vertices along each axis. Some example implementations associate each cell (it, v, w) E {1, . . , P}3with a separate set of network parameters 0UVWand trilinearly interpolate them based on camera origin o:9 = Trilerp o, {9uvw-. (u, v, w) e {l, .., P}3}) (7)

[0091] The use of trilinear interpolation, unlike the nearest-neighbor interpolation used in coordinate space partitioning, helps to prevent aliasing of the view-dependent MLP parameters, which takes the form of conspicuous “popping” artifacts in specular highlights as the camera moves through space. Some example implementations further reduce popping between submodels via regularization. After parameter interpolation, view-dependent colors can be decoded according to Eq. (4).

[0092] Since the size of the view-dependent MLP is negligible compared to total representation size, deferred network partitioning has almost no effect on memory consumption or storage impact. As a result, this technique increases model capacity almost for free. This is in contrast to coordinate space partitioning, which significantly increases storage size. The union of coordinate space and deferred network partitioning effectivelyincreases spatial and view-dependent resolution, respectively. See Figures 4A-C for an illustration of these two forms of partitioning.

[0093] Specifically, Figures 4A-C illustrate coordinate systems in example implementations of the present disclosure for a scene with K3= 33coordinate space partitions and P3= 43deferred appearance network sub-partitions. Each partition is capable of representing the entire scene while allocating the majority of its model capacity to its corresponding partition. Within each partition, some example implementations instantiate a set of spatially-anchored MLP weights {0(J} parameterizing the deferred appearance model, which some example implementations trilinearly interpolate as a function of the camera origin o during rendering. Specifically, Figure 4A illustrates the entire scene 402 in world coordinates with the scene partition 404 and two submodels 406 and 408. Figures 4B and 4C illustrate the same scene 402 from the view of two submodels 406 and 408 in their corresponding contracted coordinate systems. Specifically, Figure 4B visualizes the rendering and parameter interpolation process when the camera origin o 450 lies inside of a submodel’s partition 452, and Figure 4C visualizes the rendering and parameter interpolation process when the camera origin o 460 lies inside of a submodel’s partition 462.

[0094] Example Feature Gating Techniques

[0095] The final level in the proposed hierarchy is at the level of coarse 3D voxel grid V and three high-resolution planes (Px, Py, Pz). In MERF, each 3D position is associated with an 8-dimensional feature vector: the sum of the contributions from these four sources (See Eq. (1)). Although effective, the features generated by this procedure are limited by their naive use of summation to merge high- and low-resolution information, entangling the two together.

[0096] In contrast, some example implementations of the present disclosure instead use low-resolution 3D features to “gate” high-resolution features: if high-resolution features add value for a given 3D coordinate, they should be employed; otherwise, they should be ignored and the smoother, low-resolution features should be used. To this end, some example implementations modify feature aggregation as follows: instead of a naive summation, some example implementations take the last component w(x) = [V(x)]8of the low-resolution voxel grid’s contribution and use it to scale the triplane feature contributions: t(x) = w(x) • (PX(X) + Py(x) + Pz(x) + V(x). (8)

[0097] Some example implementations then build the proposed final feature representation by concatenating the aggregated features t(x) and the voxel grid features V(x): t(x) = t(x) © V(x). (9)

[0098] This incentivizes the model to leverage the low-resolution voxel grid to disable the high-resolution triplanes when rendering low-frequency content such as featureless white walls, and gives the model the freedom to focus on detailed parts of the scene. This change can also be thought of as a sort of “attention”, as a multiplicative interaction is being used to determine when the model should “attend” to the triplane features. This change slightly affects the memory and speed of the proposed model by virtue of increasing the number of rows in the first weight matrix of the proposed MLP, but the practical impact of this is negligible.Example Training Techniques

[0099] Example Radiance Field Distillation Techniques

[0100] NeRF-like models such as MERF are traditionally trained “from scratch” to minimize photometric loss on a set of posed input images. Regularization is very beneficial when training such systems to improve generalization to novel views. One example employs a family of carefully tuned losses in addition to photometric reconstruction error to achieve state-of-the-art performance on large, multi-room scenes. Some example implementations instead adopt “student / teacher” distillation and train the proposed representation to imitate an already -trained, state-of-the-art radiance field model.

[0101] Distillation has several advantages: some example implementations inherit the helpful inductive biases of the teacher model, circumvent the need for onerous hyperparameter tuning for generalization, and enable the recovery of local representations that are also globally consistent. The proposed approach achieves quality comparable to its a larger teacher while being three orders of magnitude faster to render.

[0102] Some example implementations supervise the proposed model by distilling the appearance and geometry of a reference radiance field. Note that the teacher model is frozen during optimization. See Figure 5 for a visualization. Specifically, Figure 5 illustrates an example teacher supervision approach. The student receives photometric supervision via rendered colors and geometric supervision via the rendering weights along camera rays. Both models operate on the same set of ray intervals.

[0103] Appearance'.

[0104] Some example implementations supervise the proposed model by minimizing the photometric difference between patches predicted by the proposed model and a source of ground-truth. Instead of limiting training to a set of photos representing a small subset of a scene’s plenoptic function, the proposed source of “ground-truth” image patches is a teacher model rendered from an arbitrary set of cameras. Specifically, some example implementations distill appearance by penalizing the discrepancy between 3 x 3 patches rendered from student and teacher models. Some example implementations use a weighted combination of the RMSE and DSSIM losses between each student patch C and its corresponding teacher patch C*:£c= 1.5

[0105] Geometry.

[0106] To distill geometry, some example implementations begin by querying the proposed teacher with a given ray origin and direction. This yields a set of weighted intervals along the camera ray {((£ / ,w )}, where each (t£, t£+1) are the metric distances along the ray corresponding to interval i, and each wf is the teacher’s corresponding alpha compositing weight for the same interval as per Eq. (3). The weight of each interval i reflects its contribution to the final predicted radiance. It is this quantity that some example implementations distill into the proposed student. Specifically, some example implementations compute the absolute difference between the teacher and student weights:

[0107] Because these volumetric rendering weights are a function of volumetric density (Eq. (3)), this loss on weights indirectly encourages the student’s and teacher’s density fields to be consistent with each other in visible regions of the scene.

[0108] Example Data Augmentation Techniques

[0109] Because the proposed distillation approach enables the supervision of the proposed student model at any ray in Euclidean space, some example implementations require a procedure for selecting a useful set of rays. While sampling rays uniformly at random throughout the scene is viable, this approach leads to poor reconstruction quality as many such rays originate from inside of objects or walls or are pointed towards unimportant or under-observed parts of the scene. Using camera rays corresponding to pixels in the dataset used to train the teacher model also performs poorly, as those input images represent atiny subset of possible views of the scene. As such, some example implementations adopt a compromise approach by using randomly-perturbed versions of the dataset’s camera rays, which yields a kind of “data augmentation” that improves generalization while focusing the student model’s attention towards the parts of the scene that the photographer deemed relevant.

[0110] To generate a training ray, some example implementations first randomly select a ray from the teacher’s dataset with origin o and direction d. Some example implementations then jitter the origin with isotropic Gaussian noise and draw a uniform sample from an E- neighborhood of the ray’s direction vector to obtain a ray (o, d): o~JV'(o, o-2H)) , (12)

[0111] Some example implementations set o = 0.03K and E = 0.03. Note that o is defined in normalized scene coordinates, where all the input cameras are contained within the [—K / 2, K / 2]3cube; See Figure 6 as an example. Specifically, Figure 6 illustrates an example ray jittering process. To generate training rays for the proposed student model (e.g., shown as the smaller cameras such as camera 602) some example implementations randomly perturb the origins and directions of the camera rays used to supervise the proposed teacher model (e.g., shown as the larger camera 604).

[0112] Example Submodel Consistency Techniques

[0113] Recall that, in spite of employing a single teacher model, coordinate space partitioning means that some example implementations are effectively training multiple, independent student submodels in parallel (in some cases, all submodels are actually trained simultaneously on a single host). This presents a challenge in terms of consistency across submodels — at test time, we seek temporal consistency under smooth camera motion, even when transitioning between submodels. This can be ameliorated by rendering multiple submodels and blending their results, but doing so significantly slows rendering and requires the presence of multiple submodels at once. In contrast, some example implementations aim to render each frame while only querying a single submodel.

[0114] To encourage adjacent submodels to make similar predictions for a given camera ray, some example implementations introduce a photometric consistency loss between submodels. During training some example implementations render each camera ray in the proposed batch twice: once using its “home” submodel s (whichever submodel the ray origin lies within the interior of), and again using a randomly-chosen neighboring submodel s.Some example implementations then impose a straightforward loss between those two rendered colors: s = l |cs(r) — c^(r) | |2. (14)

[0115] Additionally, when constructing batches of training rays, some example implementations take care to assign each ray to a submodel where it will meaningfully improve reconstruction quality. Intuitively, some example implementations expect rays to add the most value to their “home” submodel, but the rays that originate from neighboring submodels may also provide value by providing additional viewing angles of scene content within a submodel’s interior. As such, some example implementations first assign each training ray its “home” submodel, and then randomly re-assign some percentage (e.g., 20%) of rays per batch to a randomly-selected adjacent submodel.Example Rendering Techniques

[0116] Example Baking Techniques

[0117] After training, some example implementations generate precomputed assets for the real-time viewer. Some example implementations broadly follow the “baking” process of MERF with the changes described here and other minor modifications. Recall that MERF extracts a multi-resolution occupancy grid for two purposes: (i) to mark the parts in the scene for baking, and (ii) to skip empty space during rendering. Some example implementations eliminate spurious floaters by post-processing this occupancy grid (at the highest resolution R3= 20483) with a 3x3x3 median filter. As output, some example implementations produce an independent set of baked assets for each submodel, where each asset collection closely mirrors that produced by MERF. These assets include three high-resolution 2D feature maps and a sparse low-resolution 3D feature grid, both represented as quantized byte arrays. The deferred network parameters, in contrast, retain their floating point representation. Unlike MERF, some example implementations store assets as gzip-compressed binary blobs, which are slightly smaller and significantly faster to decode than PNG images.

[0118] Example Live Viewer

[0119] An example viewer proposed herein is based on the MERF volume Tenderer, i.e an OpenGL fragment shader that implements both ray marching and deferred rendering, using texture look-ups to retrieve pointwise feature representations and density. However, the proposed implementation can contain several modifications: support for submodels and parameter interpolation, a distance grid acceleration structure, and / or other low-leveloptimizations. The proposed viewer has modest compute and memory requirements, enabling real-time rendering on smartphones and other resource-constrained devices.

[0120] Submodels'.

[0121] Recall that the proposed method only requires a single submodel to render any given viewpoint in the scene, strongly bounding peak memory usage. To hide network latency, some example implementations nevertheless “ping-pong” two submodels in and out of GPU memory as the user interactively explores the environment. When the camera enters a new subvolume, some example implementations begin loading the corresponding submodel into CPU memory while the previous submodel, which includes an approximate representation of this new adjacent region, continues to be used for rendering. Once the new submodel is available, some example implementations transfer it to GPU memory and immediately use it for rendering. Peak GPU memory usage is thus limited to a maximum of two submodels: one actively being rendered and one being loaded into memory. Some example implementations further limit CPU memory usage by evicting loaded submodels from memory via least-recently-used caching.

[0122] Deferred Appearance Network Interpolation'.

[0123] While trilinearly interpolating deferred appearance network parameters requires additional computation, this feature has a negligible impact on frame rendering time. As all pixels in an image share a common ray origin, some example implementations only need to interpolate parameters once per frame. In practice, some example implementations perform this interpolation on the CPU before executing the fragment shader.Additional Disclosure

[0124] The technology discussed herein makes reference to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0125] While the present subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way ofexplanation, not limitation of the disclosure. Those skilled in the art, upon attaining an understanding of the foregoing, can readily produce alterations to, variations of, and equivalents to such embodiments. Accordingly, the subject disclosure does not preclude inclusion of such modifications, variations and / or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. For instance, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the present disclosure cover such alterations, variations, and equivalents.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method to perform real-time radiance field rendering, the method comprising: obtaining, by a computing system comprising one or more computing devices, a current camera location within a three-dimensional volume associated with a scene; accessing, by the computing system, a set of spatially-structured parameter values associated with a region of the three-dimensional volume that contains the current camera location; interpolating, by the computing system, the set of spatially-structured parameter values based on a location of the current camera location within the region to generate a set of interpolated parameter values for a neural radiance field (NeRF) model; and running, by the computing system, the NeRF model with the set of interpolated parameter values to render one or more pixels depicting the scene from the current camera location.

2. The computer-implemented method of any preceding claim, wherein: the three-dimensional volume is partitioned into a voxel grid; accessing, by the computing system, the set of spatially-structured parameter values associated with the region comprises accessing, by the computing system, the set of spatially- structured parameter values associated with the voxel that contains the current camera location.

3. The computer-implemented method of any preceding claim, wherein interpolating, by the computing system, the set of spatially-structured parameter values based on the location of the current camera location within the region comprises performing, by the computing system, trilinear interpolation of the set of spatially-structured parameter values based on the location of the current camera location within the region.

4. The computer-implemented method of any preceding claim, wherein the NeRF model comprises a set of predetermined feature values and a deferred Tenderer network, and wherein running, by the computing system, the NeRF model with the set of interpolatedparameter values comprises running, by the computing system, the NeRF model with the set of interpolated parameter values used by the deferred Tenderer network.

5. The computer-implemented method of any preceding claim, wherein said steps of accessing and interpolating are performed once per image rendered from the current camera location.

6. The computer-implemented method of any preceding claim, wherein the set of spatially-structured parameter values associated with a region of the three-dimensional volume have been learned via distillation from a teacher NeRF model.

7. The computer-implemented method of any preceding claim, wherein the set of spatially-structured parameter values associated with a region of the three-dimensional volume have been learned via application of a consistency loss that enforces visual consistency between parameter values associated with adjacent regions of the three- dimensional volume.

8. A computing system for real-time radiance field rendering, the computing system comprising one or more processors and one or more non-transitory computer-readable media that collectively store computer-executable instructions for performing operations, the operations comprising: obtaining, by the computing system comprising one or more computing devices, a training camera location within a three-dimensional volume associated with a scene; accessing, by the computing system, a set of spatially-structured parameter values associated with a region of the three-dimensional volume that contains the training camera location; and interpolating, by the computing system, the set of spatially-structured parameter values based on a location of the current camera location within the region to generate a set of interpolated parameter values for a neural radiance field (NeRF) model; running, by the computing system, the NeRF model with the set of interpolated parameter values to render a predicted pixel color; andmodifying, by the computing system, at least some of the set of spatially-structured parameter values associated with the region of the three-dimensional volume that contains the training camera location based at least in part on a loss function that compares the predicted pixel color with a target pixel color.

9. The computing system of claim 8, wherein: the three-dimensional volume is partitioned into a voxel grid; accessing, by the computing system, the set of spatially-structured parameter values associated with the region comprises accessing, by the computing system, the set of spatially- structured parameter values associated with the voxel that contains the current camera location.

10. The computing system of claim 8 or 9, wherein the target pixel color comprises a color value predicted by a teacher NeRF model.

11. The computing system of claim 8 or 9, wherein the target pixel color comprises a color value predicted by a neighboring NeRF model generated by interpolating spatially- structured parameter values associated with a neighboring region of the three-dimensional volume.

12. The computing system of any of claims 8-11, wherein interpolating, by the computing system, the set of spatially-structured parameter values based on the location of the current camera location within the region comprises performing, by the computing system, trilinear interpolation of the set of spatially-structured parameter values based on the location of the current camera location within the region.

13. The computing system of any of claims 8-12, wherein modifying, by the computing system, at least some of the set of spatially-structured parameter values associated with the region of the three-dimensional volume that contains the training camera location comprises scattering, by the computing system, a gradient of the loss function to invert the interpolation.

14. The computing system of any of claims 8-13, wherein the NeRF model comprises a set of feature values and a deferred Tenderer network, and wherein running, by the computing system, the NeRF model with the set of interpolated parameter values comprises running, by the computing system, the NeRF model with the set of interpolated parameter values used by the deferred Tenderer network.

15. One or more non-transitory computer-readable media that collectively store the set of predetermined or spatially-structured parameter values described in any preceding claim.

Citation Information

Cited By

  • Three-dimensional object material mapping generation method based on generative prior

    CN121437759A