Model subdivision for compact real-time radiance fields
By partitioning a 3D volume into NeRF submodels and training them with a submodel consistency loss, the method addresses the challenge of real-time rendering of large-scale 3D spaces, achieving efficient and visually consistent results without requiring powerful hardware.
Patent Information
- Application Number
- PCT/US2024/060152
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Existing neural radiance field (NeRF) models struggle with real-time rendering of large-scale three-dimensional spaces due to high computational requirements and memory constraints, making them unsuitable for applications needing low-latency user interaction.
The approach involves partitioning a three-dimensional volume into regions and associating each region with a neural radiance field (NeRF) submodel. These submodels are trained to render pixels depicting their respective regions, using a submodel consistency loss term to ensure visual consistency across submodels.
This method enables efficient real-time rendering of large-scale 3D spaces without exceeding memory resources, allowing for seamless user interaction and improved visual consistency, while reducing the need for powerful hardware.
Smart Images

Figure US2024060152_19062025_PF_FP_ABST
Abstract
Description
MODEL SUBDIVISION FOR COMPACT REAL-TIME RADIANCE FIELDSRELATED APPLICATIONS
[0001] This application claims priority to and the benefit of United States Provisional Patent Application Number 63 / 609,618, filed December 13, 2023. United States Provisional Patent Application Number 63 / 609,618 is hereby incorporated by reference in its entirety.FIELD
[0002] The present disclosure relates generally to machine learning models known as neural radiance fields. More particularly, the present disclosure relates to efficient real-time rendering of large-scale three-dimensional spaces via the use of a plurality of neural radiance field submodels.BACKGROUND
[0003] The visualization of large-scale three-dimensional (3D) spaces in real time is a beneficial aspect of many technological endeavors, including 3D video gaming, virtual reality (VR), augmented reality (AR), architectural design, and / or geospatial exploration. The ability to render these spaces in a realistic and efficient manner can significantly enhance user experience and interaction.
[0004] However, the existing solutions for rendering these 3D scenes face significant limitations. One of the main challenges lies in the representation of the 3D space. Traditional methods use explicit geometric representations like meshes and point clouds, which are fed into a rendering pipeline to generate 2D images. However, these methods often struggle to capture not only the complexities of natural scenes, but also the subtle interplay of light and matter, both of which are necessary for photorealistic rendering.
[0005] Recently, a new class of models, known as neural radiance fields (NeRF), has shown promise in rendering photorealistic 3D scenes. These models employ a continuous volumetric scene function, which is queried to generate novel views of a 3D scene. Despite their superior performance in rendering high-quality images, NeRF models have to date suffered from significant drawbacks in terms of speed and computational resources. Training a NeRF model can often take hours or even days, and rendering an image can take minutes, making them unsuitable for applications that require real-time interaction.
[0006] Even more critically, these models struggle to scale to large, unbounded scenes without exceeding the finite memory resources of commodity hardware. A complete, detailed3D representation of a large space requires a correspondingly large model. This easily leads to large, monolithic model architectures that require supercomputer-grade hardware to train and serve. This is particularly relevant when the model requires low-latency user interaction, such as when navigating a 3D space on-device.SUMMARY
[0007] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or can be learned from the description, or can be learned through practice of the embodiments.
[0008] One example aspect of the present disclosure is directed to a computer- implemented method to generate compact neural representations with improved visual consistency. The method includes partitioning, by a computing system comprising one or more computing devices, a three-dimensional volume associated with a scene into a plurality of regions. The method includes associating, by the computing system, a plurality of neural radiance field (NeRF) submodels respectively with the plurality of regions. The method includes training, by the computing system, each NeRF submodel to render pixels that depict its corresponding region. In some implementations, training, by the computing system, each NeRF submodel comprises training, by the computing system, each NeRF submodel using a submodel consistency loss term that penalizes a difference between a first prediction made by the NeRF submodel for a pixel and a second prediction made for the pixel by a different NeRF submodel of the plurality of NeRF submodels.
[0009] Another example aspect of the present disclosure is directed to a computing system for efficient neural rendering, the computing system including one or more processors and one or more non-transitory computer-readable media that collectively store: a set of neural radiance field (NeRF) rendering assets, wherein the set of NeRF rendering assets comprise one or more shared feature maps and parameter values for a plurality of NeRF submodels respectively associated with a plurality of regions of a three-dimensional volume associated with a scene, wherein the plurality of NeRF submodels have been trained using a submodel consistency loss that enforces visual consistency between the plurality of NeRF submodels; and instructions that, when executed by the one or more processors, cause the computing system to perform real-time rendering operations. The operations include: determining a current camera location within the three-dimensional volume associated with the scene; loading into memory the NeRF submodel associated with the region of the three-dimensional volume that contains the current camera location; and rendering the scene using the NeRF submodel loaded into memory.
[0010] Other aspects of the present disclosure are directed to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.
[0011] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the related principles.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Detailed discussion of embodiments directed to one of ordinary skill in the art is set forth in the specification, which makes reference to the appended figures, in which:
[0013] Figure 1 depicts a flow chart diagram of an example method to generate compact neural representations with improved visual consistency according to example embodiments of the present disclosure.
[0014] Figure 2 depicts a flow chart diagram of an example method to perform efficient neural rendering according to example embodiments of the present disclosure.
[0015] Figure 3 A depicts a block diagram of an example computing system according to example embodiments of the present disclosure.
[0016] Figure 3B depicts a block diagram of an example computing device according to example embodiments of the present disclosure.
[0017] Figure 3C depicts a block diagram of an example computing device according to example embodiments of the present disclosure.
[0018] Figures 4A-C depict graphical diagrams of example coordinate systems according to example embodiments of the present disclosure.
[0019] Figure 5 depicts a graphical diagram of example teacher supervision according to example embodiments of the present disclosure.
[0020] Figure 6 depicts a graphical diagram of example ray jittering according to example embodiments of the present disclosure.
[0021] Reference numerals that are repeated across plural figures are intended to identify the same features in various implementations.DETAILED DESCRIPTIONOverview
[0022] Example aspects of the present disclosure are directed to generating compact neural representations with improved visual consistency for real-time rendering of large-scale three-dimensional (3D) spaces. The proposed techniques address the problem of creating a detailed 3D representation of a large space, which traditionally requires a correspondingly large model. This often leads to large, monolithic model architectures that require powerful hardware to train and serve. In contrast, the technology proposed herein circumvents this limitation, allowing for real-time radiance fields. In particular, an example method can include partitioning a 3D volume into regions and associating each region with a neural radiance field (NeRF) submodel. These submodels can be trained to render pixels depicting their respective regions, maintaining visual consistency and enabling real-time interaction. This approach allows for the representation of large spaces without exceeding memory resources, overcoming the limitations of existing NeRF models.
[0023] More particularly, some example implementations can begin by partitioning a three-dimensional volume associated with a scene into a plurality of regions. The present disclosure then associates a plurality of neural radiance field (NeRF) submodels respectively with the plurality of regions. Each of these submodels is trained to render pixels that depict its corresponding region. This is a significant departure from traditional models, where one large model is trained to render the entire scene. By using a collection of submodels, the present disclosure allows for more efficient use of computational resources and improved rendering performance. Furthermore, in some implementations, each of the submodels in the present disclosure is capable of rendering the entire scene, with higher fidelity near a corresponding “focus region”. This means that each submodel specializes in a particular region of the scene, but can still render the entire scene if needed.
[0024] According to an aspect of the present disclosure, training each NeRF submodel can include applying a submodel consistency loss term that penalizes a difference between a first prediction made by the NeRF submodel for a pixel and a second prediction made for the pixel by a different NeRF submodel of the plurality of NeRF submodels. This encourages consistency among the submodels, ensuring that they “agree” with each other in terms of geometry and color. This results in a smooth transition when swapping between submodels, without an obvious sudden change in how rendered images appear.
[0025] As one example, the submodel consistency loss can penalize a difference between a first color predicted for the pixel by the NeRF submodel and a second colorpredicted for the pixel by a different NeRF submodel. This ensures that the color rendering is consistent across the different submodels, further enhancing the visual experience for the user.
[0026] Another example aspect of the present disclosure is directed to a novel training strategy in which a shared teacher NeRF model is distilled into the plurality of NeRF submodels. This distillation process allows the submodels to learn from a strong, densely sampled signal, leading to improved training results, more efficient use of computational resources, and improved consistency among the plurality of NeRF submodels.
[0027] Once the submodels have been trained, then at inference time a computing system can choose the most appropriate set of model parameters depending on camera location. This can include a hierarchical architecture where, based on camera origin, a submodel is chosen and within that submodel, deferred MLP parameters are selected (e.g., generated via interpolation). This allows for more accurate rendering of the scene depending on the current camera location.
[0028] Thus, the present disclosure provides a novel approach to real-time 3D rendering, allowing for detailed representations of large spaces without the need for powerful hardware. By partitioning the scene into a collection of independently-renderable submodels, the technology allows for more efficient use of computational resources and improved rendering performance, while still preserving visual consistency among the submodels.
[0029] In particular, the techniques proposed herein address the technical problem of efficiently rendering large-scale, three-dimensional spaces in real-time on standard hardware. This is a technical problem as it includes managing complex computations and resource allocation in a computing system. The solution proposed, which includes partitioning a 3D volume into regions and associating each region with a neural radiance field submodel, provides a tangible effect - the efficient rendering of large-scale 3D spaces in real-time, thereby resulting in faster rendering, using fewer computational resources. Furthermore, the method of training each submodel using a submodel consistency loss term that penalizes a difference between a first prediction made by the submodel for a pixel and a second prediction made for the pixel by a different submodel ensures smooth transitions between submodels and thereby improves the visual consistency of the rendered image.
[0030] With reference now to the Figures, example embodiments of the present disclosure will be discussed in further detail.Example Methods
[0031] Figure 1 illustrates a flowchart of a computer-implemented method to generate compact neural representations with improved visual consistency. The method can be implemented by a computing system that comprises one or more computing devices. The computing system can be a single computing device or a network of computing devices. This could include, for example, a personal computer, a server, a cloud computing system, or any other type of computing device or system that is capable of processing and managing large amounts of data.
[0032] A first step, as indicated at block 12, includes partitioning a three-dimensional volume associated with a scene into a plurality of regions. This partitioning can be performed in any suitable manner and is not limited to any particular method or algorithm. The three- dimensional volume could represent a real-world scene, such as a room or a landscape, or it could represent a virtual or imaginary scene. The regions into which the volume is partitioned could be of any size and shape, and they could be evenly or unevenly distributed throughout the volume. One example partitioning scheme is to subdivide the scene into a coarse 3D grid.
[0033] The next step, as indicated by block 14, includes associating a plurality of neural radiance field (NeRF) submodels respectively with the plurality of regions. Each NeRF submodel can be designed to render pixels that depict its corresponding region. The use of NeRF submodels allows the system to efficiently generate high-quality renderings of the scene with a high level of detail and visual consistency.
[0034] In particular, in some implementations, each submodel is only tasked to represent the scene within the submodel’s grid cell with high detail. The region outside the submodel’s cell is only modelled coarsely. Since the entire scene is represented by each submodel, rendering an image only entails queries to a single submodel, which leads to fast inference.
[0035] In some implementations, at block 14, the computing system associates the plurality of NeRF submodels with the plurality of regions by instantiating an active set of NeRF submodels. These submodels can correspond to only those regions of the three- dimensional volume that contain at least one training camera associated with a training image included in a training dataset. This method allows for efficient utilization of computational resources by focusing only on those regions that are relevant to the training process.
[0036] More specifically, instead of simply instantiating submodel parameters for all possible subvolumes, the system limits the instantiation to an active subset of submodels. This subset is determined by examining the set of camera origins present in the training dataset and marking those subvolumes that contain at least one of these training cameras.
[0037] This approach is based on the intuitive understanding that the model is unlikely to provide accurate or meaningful predictions for regions that are far from the training manifold. Therefore, there is no need to instantiate submodels for these regions, which can result in substantial savings in terms of computational resources.
[0038] For instance, in scenarios where the scenes lie on a 2D surface, such as a singlestory home, this method can significantly reduce the number of active parameters. For example, the number of active parameters can be reduced from a cubic scale (O(KA3), where K is the number of subvolumes along one dimension of the volume) to a quadratic scale (O(KA2)). This reduction can lead to significant improvements in computational efficiency, making the method particularly advantageous for applications that include large and complex three-dimensional volumes.
[0039] The third step, as indicated by block 16, includes training each NeRF submodel. According to an aspect of the present disclosure, the training process can include using a submodel consistency loss term that penalizes a difference between a first prediction made by the NeRF submodel for a pixel and a second prediction made for the pixel by a different NeRF submodel of the plurality of NeRF submodels. This consistency loss term encourages the submodels to make consistent predictions, which helps to ensure that the final renderings of the scene are visually consistent and coherent.
[0040] As an example, in some implementations, the submodel consistency loss can be designed to penalize a difference between a first color predicted for a given pixel by one NeRF submodel and a second color predicted for the same pixel by another NeRF submodel. This penalization process can be performed by rendering each camera ray in a training batch twice, once using its associated or ‘home’ submodel, and once using another randomly chosen neighboring submodel. The system then calculates a loss value based on the difference between the color predictions made by these two submodels. This consistency loss encourages the submodels to make similar color predictions for the same pixels, thereby promoting visual consistency across the scene. As a result of this process, transitions between submodels during navigation of the scene are rendered unnoticeable, eliminating the need for costly blending of frames rendered from adjacent submodels. This further enhances the efficiency and real-time performance of the rendering process.
[0041] According to another aspect, in some implementations, each neural radiance field (NeRF) submodel is configured to apply a contraction function to various spatial positions within the three-dimensional volume associated with the scene. The contraction function enables each NeRF submodel to render the entirety of the three-dimensional volume.Specifically, the contraction function can be applied to each spatial position before querying the feature field. This process allows for modeling of far-away content in an unbounded scene in a coarse manner, thereby achieving a resolution that drops off with the distance from the scene’s focus point.
[0042] An advantage of the contraction function is that it facilitates efficient ray -box intersection tests, which are beneficial for fast empty space skipping. This efficiency in rendering contributes to the high-speed performance and real-time capabilities of the system. The contraction function, therefore, not only aids in the rendering process but also enhances the overall performance and efficacy of the method depicted in Figure 1.
[0043] The contraction function can be used in conjunction with a transformation of world coordinates into the submodel’s coordinate system. With this transformation, all points within the submodel’s cell are mapped to a normalized coordinate space, such as [—1, 1]3. The application of the contraction function to these transformed coordinates results in uniform resolution within the submodel’s grid cell. The resolution then drops off with the distance from the cell’s boundary. Unlike prior work, the present disclosure explicitly focuses the submodel’s capacity on its associated subvolume, thanks to the combination of submodel coordinate transformation and scene contraction.
[0044] According to another aspect of the present disclosure, the training process of each NeRF submodel, as indicated in block 16, can also include the use of a shared teacher NeRF model. This shared teacher model can be a high-quality, state-of-the-art neural radiance field. The training process of the NeRF submodels can be considered as a distillation process where each NeRF submodel learns to imitate the shared teacher model. This distillation process can be implemented by the computing system in a variety of ways, all aiming to replicate the performance of the teacher model in the NeRF submodels.
[0045] The shared teacher model provides a more densely sampled signal than a traditional dataset, thereby increasing the amount of data available for each NeRF submodel. This process serves to smooth the optimization landscape and encourage cross-submodel consistency. The computing system can train each NeRF submodel to imitate the teacher model in terms of both predicted radiance and other quantities.
[0046] An additional benefit of this distillation process is the opportunity for strong supervision. The shared teacher model, being a high-quality representation of the neural radiance field, can provide robust guidance for the training of the NeRF submodels. This supervision can lead to the creation of NeRF submodels that more accurately depict their corresponding regions of the scene.
[0047] In some implementations, the distillation process can be optimized over independently sampled patches of camera rays. This method allows each NeRF submodel to learn from a diverse range of data points, further enhancing the accuracy of the final renderings. The computing system can optimize the distillation process to ensure that the NeRF submodels can accurately depict their corresponding regions while maintaining visual consistency across the entire scene.
[0048] In an additional aspect of the present disclosure, in some implementations, the plurality of NeRF submodels may share at least some of a set of model parameters. These shared model parameters can be distributed in the three-dimensional volume associated with the scene. Further, each of the plurality of NeRF submodels can include a deferred shader model, a type of rendering model that defers certain shading operations to a later stage in the rendering process. This deferred shader model can include parameters that are generated by interpolating at least some of the set of model parameters based on a current camera location within the three-dimensional volume associated with the scene. The interpolation process can take into account various factors such as the distance of the camera location from various points in the scene, the relative orientations of the camera and the scene, and other relevant factors. This allows the system to dynamically adjust the parameters of the deferred shader model as the camera location changes, thereby enabling the system to maintain a high level of detail and visual consistency in the renderings of the scene.
[0049] Referring still to Figure 1, a final step of the method, as indicated by block 18, includes outputting the trained NeRF submodels. This step can include storing rendering assets that can be efficiently used to load and render from different submodels. These stored assets can then be used to generate renderings of the scene from any viewpoint, allowing for a high level of flexibility and user interactivity in the final renderings.
[0050] The method illustrated in Figure 1 can be implemented in a variety of different contexts. For example, it can be used in the field of computer graphics to generate realistic renderings of virtual scenes for video games or virtual reality applications or for enabling a user to explore a three-dimensional space (e.g., a home, restaurant, etc.) via renderings of the space from the trained NeRF models.
[0051] Figure 2 provides a flowchart diagram of an example method for efficient neural rendering according to one embodiment of the present disclosure. The method can be implemented by a computing system that comprises one or more computing devices. The computing system can be a single computing device or a network of computing devices. This could include, for example, a personal computer, a server, a cloud computing system, or anyother type of computing device or system that is capable of processing and managing large amounts of data.
[0052] At block 202, the method includes determining a current camera location within the three-dimensional volume associated with the scene. This step can comprise accessing or detecting the position and / or orientation of a virtual camera within a predefined three- dimensional space. The camera location can be determined based on various factors such as user inputs, sensor data, or predefined paths. This camera location can be represented in any suitable coordinate system relevant to the scene.
[0053] Next, at block 204, the method includes loading into memory the neural radiance field (NeRF) submodel associated with the region of the three-dimensional volume that contains the current camera location. Each NeRF submodel is specialized to render a specific region or “focus region” within the overall scene. This step enables the system to avoid unnecessary computational work and memory usage by only loading the relevant submodel for the current camera location, rather than attempting to process the entire scene at once. The submodel is loaded into an accessible memory of the computing device, such as RAM, from a storage device, which could include local or remote storage, or could be distributed across multiple storage devices.
[0054] In some example implementations of the present disclosure, parameter values for the submodels are distributed in the three-dimensional volume associated with the scene.These parameter values can be the weights and biases of the neural network that forms the NeRF submodel, which are trained to approximate the radiance function of the scene.Therefore, in such implementations, loading into memory the NeRF submodel can include interpolating at least some of the parameter values based on the location of the current camera within the three-dimensional volume associated with the scene.
[0055] This interpolation step serves to compute the parameter values for the submodel that are relevant to the current camera location. This is beneficial as it allows the system to adapt the submodel to the current view of the scene, thus ensuring that the rendered image is accurate and detailed. The interpolation can be performed using various methods, such as linear interpolation or other suitable methods. The interpolation can be based on the spatial proximity of the camera location to the points, regions, or divisions associated with the parameter values, or on other factors such as the orientation or direction of the camera, or the characteristics or features of the scene visible from the camera location.
[0056] The interpolation of parameter values can also contribute to the seamless transition between different NeRF submodels as the camera moves within the three-dimensional volume. By determining the most relevant parameters for the current camera location, the system can smoothly switch between different submodels, thus preventing abrupt changes in the rendered images and maintaining a consistent and high-quality depiction of the scene as the camera navigates through it.
[0057] Referring still to Figure 2, at block 206, the scene is rendered using the NeRF submodel loaded into memory. This rendering step can include generating a two-dimensional image or sequence of images, representing the three-dimensional scene from the perspective of the current camera location. The rendering can utilize ray marching or other graphics rendering techniques to generate realistic images based on the submodel. This could include executing the parameters of the NeRF submodel to render the appearance of surfaces and objects within the scene, including their colors, textures, and the effects of lighting and shadows.
[0058] In some implementations, the method can further optionally include pre-loading into memory one or more of the NeRF submodels that are associated with one or more neighboring regions of the three-dimensional volume. These neighboring regions may be adjacent to the region of the three-dimensional volume that contains the current camera location. The pre-loading operation can be performed in anticipation of the camera moving towards these neighboring regions, thereby ensuring a seamless and efficient rendering process.
[0059] The pre-loading operation can be carried out concurrently with the rendering operation of block 206, leveraging the multi -threading capabilities of modern computing systems. The selection of which neighboring regions to pre-load can be based on various factors such as the current camera direction, speed, or a predicted path based on user inputs or other data. Different pre-loading strategies can be implemented, such as pre-loading the closest neighboring regions, regions in the direction of camera movement, or all neighboring regions.
[0060] In block 208, the method can include detecting that the current camera location has moved from the current region to a second, different region of the three-dimensional volume. In response, in block 210, the method can include a “hot swapping” operation in which the current submodel is replaced with the submodel for the second, different region. Thus, different submodels can be loaded in and out of video memory as the user interactively explores the scene. The “hot swapping” operation is a memory management technique that allows for efficient use of available memory resources. As the user moves around the scene, the submodel corresponding to the current camera location is kept in memory while the othersubmodels are swapped out. This ensures that memory usage is bounded by the size of a single submodel, rather than the cumulative size of all submodels.Example Devices and Systems
[0061] Figure 3 A depicts a block diagram of an example computing system 100 according to example embodiments of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled over a network 180.
[0062] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0063] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 which are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0064] In some implementations, the user computing device 102 can store or include one or more machine-learned models 120. For example, the machine-learned models 120 can be or can otherwise include various machine-learned models such as neural networks (e.g., deep neural networks) or other types of machine-learned models, including non-linear models and / or linear models. Neural networks can include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks or other forms of neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models). Example machine-learned models 120 are discussed with reference to Figures 1-2.
[0065] In some implementations, the one or more machine-learned models 120 can be received from the server computing system 130 over network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 can implementmultiple parallel instances of a single machine-learned model 120 (e.g., to perform parallel neural rendering across multiple instances of Tenderers).
[0066] Additionally or alternatively, one or more machine-learned models 140 can be included in or otherwise stored and implemented by the server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine-learned models 140 can be implemented by the server computing system 140 as a portion of a web service (e.g., a neural rendering service). Thus, one or more models 120 can be stored and implemented at the user computing device 102 and / or one or more models 140 can be stored and implemented at the server computing system 130.
[0067] The user computing device 102 can also include one or more user input components 122 that receives user input. For example, the user input component 122 can be a touch-sensitive component (e.g., a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component can serve to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.
[0068] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 which are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0069] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In instances in which the server computing system 130 includes plural server computing devices, such server computing devices can operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.
[0070] As described above, the server computing system 130 can store or otherwise include one or more machine-learned models 140. For example, the models 140 can be or can otherwise include various machine-learned models. Example machine-learned models include neural networks or other multi-layer non-linear models. Example neural networksinclude feed forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models). Example models 140 are discussed with reference to Figures 1-2.
[0071] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 via interaction with the training computing system 150 that is communicatively coupled over the network 180. The training computing system 150 can be separate from the server computing system 130 or can be a portion of the server computing system 130.
[0072] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 which are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.
[0073] The training computing system 150 can include a model trainer 160 that trains the machine-learned models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.
[0074] In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. The model trainer 160 can perform a number of generalization techniques (e.g., weight decays, dropouts, etc.) to improve the generalization capability of the models being trained.
[0075] In particular, the model trainer 160 can train the machine-learned models 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, multiple images of a scene.
[0076] In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some instances, this process can be referred to as personalizing the model.
[0077] The model trainer 160 includes computer logic utilized to provide desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into a memory and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions that are stored in a tangible computer-readable storage medium such as RAM, hard disk, or optical or magnetic media.
[0078] The network 180 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., Internet), or some combination thereof and can include any number of wired or wireless links. In general, communication over the network 180 can be carried via any type of wired and / or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0079] Figure 3 A illustrates one example computing system that can be used to implement the present disclosure. Other computing systems can be used as well. For example, in some implementations, the user computing device 102 can include the model trainer 160 and the training dataset 162. In such implementations, the models 120 can be both trained and used locally at the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the models 120 based on user-specific data.
[0080] Figure 3B depicts a block diagram of an example computing device 10 that performs according to example embodiments of the present disclosure. The computing device 10 can be a user computing device or a server computing device.
[0081] The computing device 10 includes a number of applications (e.g., applications 1 through N). Each application contains its own machine learning library and machine-learned model(s). For example, each application can include a machine-learned model. Exampleapplications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.
[0082] As illustrated in Figure 3B, each application can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.
[0083] Figure 3C depicts a block diagram of an example computing device 50 that performs according to example embodiments of the present disclosure. The computing device 50 can be a user computing device or a server computing device.
[0084] The computing device 50 includes a number of applications (e.g., applications 1 through N). Each application is in communication with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and model(s) stored therein) using an API (e.g., a common API across all applications).
[0085] The central intelligence layer includes a number of machine-learned models. For example, as illustrated in Figure 3C, a respective machine-learned model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine-learned model. For example, in some implementations, the central intelligence layer can provide a single model for all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of the computing device 50.
[0086] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As illustrated in Figure 3C, the central device data layer can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).Example Memory-Efficient Radiance Fields
[0087] This section includes a review of MERF (Reiser et al., MERF: Memory-Efficient Radiance Fields for Real-time View Synthesis in Unbounded Scenes, arXiv:2302.12249 (2023)). MERF maps 3D positions x 6 IR3to feature vectors t 6 IR8. MERF parameterizes this mapping using a combination of high-resolution triplanes (Px, Py, Pze ^RxRx8) and a low-resolution sparse voxel grid V E ]R>LXLXLX 8query point x is projected onto each of three axis-aligned planes and the underlying 2D grid is queried via bilinear interpolation. Additionally, a trilinear sample is taken from the sparse voxel grid. The resulting four 8- dimensional vectors are then summed:
[0088] This vector is then unpacked into three parts, which are independently rectified to yield a scalar density, a diffuse RGB color, and a feature vector that encodes viewdependence effects:T = exp(tx), c = sigmoid(t2A)> f = sigmoid(t5.Q). (2)
[0089] To render a pixel, a ray is cast from that pixel’s center of projection o along the viewing direction d and sampled at a set of distances {t to generate a set of points along that ray = o + tjd. The densities {T are converted into alpha compositing weights {w using the numerical quadrature approximation for volume rendering [?]:where 6i = ti+1— f is the distance between adjacent samples. After alpha compositing, a deferred shading approach is used to decode the blended diffuse RGB colorswici and the blended view-dependent color featureswif into the final pixel color with the help of the small, deferred rendering MLP hf; Of.where 6 are the MLP’s parameters.
[0090] In unbounded scenes, far-away content can be modelled coarsely. To achieve a resolution that drops off with the distance from the scene’s focus point, MERF applies a contraction function to each spatial position x before querying the feature field:Example Streamable Memory Efficient Radiance Fields
[0091] Although real-time view-synthesis methods like MERF perform well for a localized environment, they often fail to scale to large multi-room scenes. To this end, some example implementations of the present disclosure leverage a hierarchical architecture. First, some example implementations partition the coordinate space of the entire scene into a series of blocks, where each block is modeled by its own MERF -like representation. Second, some example implementations introduce a grid of spatially-anchored network parameters within each block for modeling view-dependent effects. Finally, some example implementations introduce a gating mechanism for modulating high- and low-resolution contributions to a location’s feature representation. Thus, one example overall architecture can be thought of as a three-level hierarchy: based on the camera origin, (i) some example implementations select an appropriate submodel, then within a submodel (ii) some example implementations compute the parameters of a deferred appearance network via interpolation, and then within a local voxel neighborhood, (iii) some example implementations determine a location’s feature representation via feature gating.
[0092] This greatly increases the capacity of the proposed model without diminishing rendering speed or increasing memory consumption: even as total storage requirements increase with the number of submodels, only a single submodel is needed to render a given frame. As such, when implemented on a graphics accelerator, the proposed system maintains modest resource requirements comparable to MERF .
[0093] Example Coordinate Space Partitioning Techniques
[0094] While MERF offers sufficient capacity for faithfully representing medium-scale scenes, the use of a single set of triplanes limits its capacity and reduces image quality. In large scenes, numerous surface points project to the same 2D plane location, and the representation therefore struggles to simultaneously represent high-frequency details of multiple surfaces. Although this can be partially ameliorated by increasing the spatial resolution of the underlying representation, doing so significantly impacts memory consumption and is prohibitively expensive in practice.
[0095] Therefore, some example implementations of the present disclosure instead coarsely subdivide the scene into a 3D grid based on camera origin and associate each grid cell with an independent submodel. Each submodel is assigned its own contraction space (Eq. (5)) and is tasked with representing the region of the scene within its grid cell at high detail, while the region outside each submodel’s cell is modeled coarsely. Note that, typically, the entire scene is still represented by each submodel — the submodels differ only in terms of which portions of the scene lies inside or outside of each submodel’s contraction region. As such, rendering a camera only requires a single submodel, implying that only one submodel must be in memory at a time.
[0096] Formally, some example implementations shift and scale all training cameras to lie within a [— K / 2, K / 2]3cube, and then partition this cube into K3identical and tightly packed subvolumes of size [— 1,1]3. Some example implementations can assign training cameras to submodels {<Sfc} by identifying the associated subvolume Rfethat each camera origin o lies within:
[0097] Some example implementations configure the proposed camera-to-submodel assignment procedure to apply to cameras outside of the training set, ensuring its validity during test set rendering. This enables a wide range of features including ray jittering, submodel reassignment, a submodel consensus loss, arbitrary test camera placement, and ping-pong buffers.
[0098] Rather than exhaustively instantiating submodels for all K3subvolumes, some example implementations only consider subvolumes that contain at least one training camera. As most scenes are outdoors or single-story buildings, this reduces the number of submodels from K3to O(K2).
[0099] Example Deferred Appearance Network Partitioning Techniques
[0100] The second level in the proposed partitioning hierarchy concerns the deferred rendering model. Recall that MERF employs a small multi-layer perceptron (MLP) to decode view-dependent colors from blended features as described in Eq. (4). Although the small size of this network is advantageous for fast inference, in some specific cases its capacity is insufficient to accurately reproduce complex view-dependent effects common in larger scenes. Simply increasing the size of this network is not viable as doing so would significantly reduce rendering speed.
[0101] Instead, some example implementations uniformly subdivide the domain of each submodel into a lattice with P vertices along each axis. Some example implementations associate each cell (it, v, w) E {1, . . , P}3with a separate set of network parameters 0UVWand trilinearly interpolate them based on camera origin o:0 = Trilerp o, {0uvw(u, v, w) e {l, .., P}3}) (7)
[0102] The use of trilinear interpolation, unlike the nearest-neighbor interpolation used in coordinate space partitioning, helps to prevent aliasing of the view-dependent MLP parameters, which takes the form of conspicuous “popping” artifacts in specular highlights as the camera moves through space. Some example implementations further reduce popping between submodels via regularization. After parameter interpolation, view-dependent colors can be decoded according to Eq. (4).
[0103] Since the size of the view-dependent MLP is negligible compared to total representation size, deferred network partitioning has almost no effect on memory consumption or storage impact. As a result, this technique increases model capacity almost for free. This is in contrast to coordinate space partitioning, which significantly increases storage size. The union of coordinate space and deferred network partitioning effectively increases spatial and view-dependent resolution, respectively. See Figures 4A-C for an illustration of these two forms of partitioning.
[0104] Specifically, Figures 4A-C illustrate coordinate systems in example implementations of the present disclosure for a scene with K3= 33coordinate space partitions and P3= 43deferred appearance network sub-partitions. Each partition is capable of representing the entire scene while allocating the majority of its model capacity to its corresponding partition. Within each partition, some example implementations instantiate a set of spatially-anchored MLP weights {0(J} parameterizing the deferred appearance model, which some example implementations trilinearly interpolate as a function of the camera origin o during rendering. Specifically, Figure 4A illustrates the entire scene 402 in world coordinates with the scene partition 404 and two submodels 406 and 408. Figures 4B and 4C illustrate the same scene 402 from the view of two submodels 406 and 408 in their corresponding contracted coordinate systems. Specifically, Figure 4B visualizes the rendering and parameter interpolation process when the camera origin o 450 lies inside of a submodel’s partition 452, and Figure 4C visualizes the rendering and parameter interpolation process when the camera origin o 460 lies inside of a submodel’s partition 462.
[0105] Example Feature Gating Techniques
[0106] The final level in the proposed hierarchy is at the level of coarse 3D voxel grid V and three high-resolution planes (Px, Py, Pz). In MERF, each 3D position is associated with an 8-dimensional feature vector: the sum of the contributions from these four sources (See Eq. (1)). Although effective, the features generated by this procedure are limited by their naive use of summation to merge high- and low-resolution information, entangling the two together.
[0107] In contrast, some example implementations of the present disclosure instead use low-resolution 3D features to “gate” high-resolution features: if high-resolution features add value for a given 3D coordinate, they should be employed; otherwise, they should be ignored and the smoother, low-resolution features should be used. To this end, some example implementations modify feature aggregation as follows: instead of a naive summation, some example implementations take the last component w(x) = [V(x)]8of the low-resolution voxel grid’s contribution and use it to scale the triplane feature contributions: t(x) = w(x) • (PX(X) + Py(x) + Pz(x) + V(x). (8)
[0108] Some example implementations then build the proposed final feature representation by concatenating the aggregated features t(x) and the voxel grid features V(x): t(x) = t(x) ® V(x). (9)
[0109] This incentivizes the model to leverage the low-resolution voxel grid to disable the high-resolution triplanes when rendering low-frequency content such as featureless white walls, and gives the model the freedom to focus on detailed parts of the scene. This change can also be thought of as a sort of “attention”, as a multiplicative interaction is being used to determine when the model should “attend” to the triplane features. This change slightly affects the memory and speed of the proposed model by virtue of increasing the number of rows in the first weight matrix of the proposed MLP, but the practical impact of this is negligible.Example Training Techniques
[0110] Example Radiance Field Distillation Techniques
[0111] NeRF-like models such as MERF are traditionally trained “from scratch” to minimize photometric loss on a set of posed input images. Regularization is very beneficial when training such systems to improve generalization to novel views. One example employs a family of carefully tuned losses in addition to photometric reconstruction error to achievestate-of-the-art performance on large, multi-room scenes. Some example implementations instead adopt “student / teacher” distillation and train the proposed representation to imitate an already -trained, state-of-the-art radiance field model.
[0112] Distillation has several advantages: some example implementations inherit the helpful inductive biases of the teacher model, circumvent the need for onerous hyperparameter tuning for generalization, and enable the recovery of local representations that are also globally consistent. The proposed approach achieves quality comparable to its a larger teacher while being three orders of magnitude faster to render.
[0113] Some example implementations supervise the proposed model by distilling the appearance and geometry of a reference radiance field. Note that the teacher model is frozen during optimization. See Figure 5 for a visualization. Specifically, Figure 5 illustrates an example teacher supervision approach. The student receives photometric supervision via rendered colors and geometric supervision via the rendering weights along camera rays. Both models operate on the same set of ray intervals.
[0114] Appearance :
[0115] Some example implementations supervise the proposed model by minimizing the photometric difference between patches predicted by the proposed model and a source of ground-truth. Instead of limiting training to a set of photos representing a small subset of a scene’s plenoptic function, the proposed source of “ground-truth” image patches is a teacher model rendered from an arbitrary set of cameras. Specifically, some example implementations distill appearance by penalizing the discrepancy between 3 x 3 patches rendered from student and teacher models. Some example implementations use a weighted combination of the RMSE and DSSIM losses between each student patch C and its corresponding teacher patch C*:£c= 1.5
[0116] Geometry.
[0117] To distill geometry, some example implementations begin by querying the proposed teacher with a given ray origin and direction. This yields a set of weighted intervals along the camera ray {((t / >w )}, where each (t£, t£+1) are the metric distances along the ray corresponding to interval i, and each w is the teacher’s corresponding alpha compositing weight for the same interval as per Eq. (3). The weight of each interval i reflects its contribution to the final predicted radiance. It is this quantity that some exampleimplementations distill into the proposed student. Specifically, some example implementations compute the absolute difference between the teacher and student weights:
[0118] Because these volumetric rendering weights are a function of volumetric density (Eq. (3)), this loss on weights indirectly encourages the student’s and teacher’s density fields to be consistent with each other in visible regions of the scene.
[0119] Example Data Augmentation Techniques
[0120] Because the proposed distillation approach enables the supervision of the proposed student model at any ray in Euclidean space, some example implementations require a procedure for selecting a useful set of rays. While sampling rays uniformly at random throughout the scene is viable, this approach leads to poor reconstruction quality as many such rays originate from inside of objects or walls or are pointed towards unimportant or under-observed parts of the scene. Using camera rays corresponding to pixels in the dataset used to train the teacher model also performs poorly, as those input images represent a tiny subset of possible views of the scene. As such, some example implementations adopt a compromise approach by using randomly-perturbed versions of the dataset’s camera rays, which yields a kind of “data augmentation” that improves generalization while focusing the student model’s attention towards the parts of the scene that the photographer deemed relevant.
[0121] To generate a training ray, some example implementations first randomly select a ray from the teacher’s dataset with origin o and direction d. Some example implementations then jitter the origin with isotropic Gaussian noise and draw a uniform sample from an E- neighborhood of the ray’s direction vector to obtain a ray (o, d): o-JVXo, o-2:)) , (12)
[0122] Some example implementations set a = 0.03K and E = 0.03. Note that a is defined in normalized scene coordinates, where all the input cameras are contained within the [—K / 2, K / 2]3cube; See Figure 6 as an example. Specifically, Figure 6 illustrates an example ray jittering process. To generate training rays for the proposed student model (e.g., shown as the smaller cameras such as camera 602) some example implementations randomly perturb the origins and directions of the camera rays used to supervise the proposed teacher model (e.g., shown as the larger camera 604).
[0123] Example Submodel Consistency Techniques
[0124] Recall that, in spite of employing a single teacher model, coordinate space partitioning means that some example implementations are effectively training multiple, independent student submodels in parallel (in some cases, all submodels are actually trained simultaneously on a single host). This presents a challenge in terms of consistency across submodels — at test time, we seek temporal consistency under smooth camera motion, even when transitioning between submodels. This can be ameliorated by rendering multiple submodels and blending their results, but doing so significantly slows rendering and requires the presence of multiple submodels at once. In contrast, some example implementations aim to render each frame while only querying a single submodel.
[0125] To encourage adjacent submodels to make similar predictions for a given camera ray, some example implementations introduce a photometric consistency loss between submodels. During training some example implementations render each camera ray in the proposed batch twice: once using its “home” submodel s (whichever submodel the ray origin lies within the interior of), and again using a randomly-chosen neighboring submodel s. Some example implementations then impose a straightforward loss between those two rendered colors: s = l|cs(r) — c^(r)| |2. (14)
[0126] Additionally, when constructing batches of training rays, some example implementations take care to assign each ray to a submodel where it will meaningfully improve reconstruction quality. Intuitively, some example implementations expect rays to add the most value to their “home” submodel, but the rays that originate from neighboring submodels may also provide value by providing additional viewing angles of scene content within a submodel’s interior. As such, some example implementations first assign each training ray its “home” submodel, and then randomly re-assign some percentage (e.g., 20%) of rays per batch to a randomly-selected adjacent submodel.Example Rendering Techniques
[0127] Example Baking Techniques
[0128] After training, some example implementations generate precomputed assets for the real-time viewer. Some example implementations broadly follow the “baking” process of MERF with the changes described here and other minor modifications. Recall that MERF extracts a multi-resolution occupancy grid for two purposes: (i) to mark the parts in the scenefor baking, and (ii) to skip empty space during rendering. Some example implementations eliminate spurious floaters by post-processing this occupancy grid (at the highest resolution R3= 20483) with a 3x3x3 median filter. As output, some example implementations produce an independent set of baked assets for each submodel, where each asset collection closely mirrors that produced by MERF. These assets include three high-resolution 2D feature maps and a sparse low-resolution 3D feature grid, both represented as quantized byte arrays. The deferred network parameters, in contrast, retain their floating point representation. Unlike MERF, some example implementations store assets as gzip-compressed binary blobs, which are slightly smaller and significantly faster to decode than PNG images.
[0129] Example Live Viewer
[0130] An example viewer proposed herein is based on the MERF volume Tenderer, i.e an OpenGL fragment shader that implements both ray marching and deferred rendering, using texture look-ups to retrieve pointwise feature representations and density. However, the proposed implementation can contain several modifications: support for submodels and parameter interpolation, a distance grid acceleration structure, and / or other low-level optimizations. The proposed viewer has modest compute and memory requirements, enabling real-time rendering on smartphones and other resource-constrained devices.
[0131] Submodels'.
[0132] Recall that the proposed method only requires a single submodel to render any given viewpoint in the scene, strongly bounding peak memory usage. To hide network latency, some example implementations nevertheless “ping-pong” two submodels in and out of GPU memory as the user interactively explores the environment. When the camera enters a new subvolume, some example implementations begin loading the corresponding submodel into CPU memory while the previous submodel, which includes an approximate representation of this new adjacent region, continues to be used for rendering. Once the new submodel is available, some example implementations transfer it to GPU memory and immediately use it for rendering. Peak GPU memory usage is thus limited to a maximum of two submodels: one actively being rendered and one being loaded into memory. Some example implementations further limit CPU memory usage by evicting loaded submodels from memory via least-recently-used caching.
[0133] Deferred Appearance Network Interpolation'.
[0134] While trilinearly interpolating deferred appearance network parameters requires additional computation, this feature has a negligible impact on frame rendering time. As all pixels in an image share a common ray origin, some example implementations only need tointerpolate parameters once per frame. In practice, some example implementations perform this interpolation on the CPU before executing the fragment shader.Additional Disclosure
[0135] The technology discussed herein makes reference to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0136] While the present subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation, not limitation of the disclosure. Those skilled in the art, upon attaining an understanding of the foregoing, can readily produce alterations to, variations of, and equivalents to such embodiments. Accordingly, the subject disclosure does not preclude inclusion of such modifications, variations and / or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. For instance, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the present disclosure cover such alterations, variations, and equivalents.
Claims
WHAT IS CLAIMED IS:
1. A computer-implemented method to generate compact neural representations with improved visual consistency, the method comprising: partitioning, by a computing system comprising one or more computing devices, a three-dimensional volume associated with a scene into a plurality of regions; associating, by the computing system, a plurality of neural radiance field (NeRF) submodels respectively with the plurality of regions; and training, by the computing system, each NeRF submodel to render pixels that depict its corresponding region; wherein training, by the computing system, each NeRF submodel comprises training, by the computing system, each NeRF submodel using a submodel consistency loss term that penalizes a difference between a first prediction made by the NeRF submodel for a pixel and a second prediction made for the pixel by a different NeRF submodel of the plurality of NeRF submodels.
2. The computer-implemented method of any preceding claim, wherein the submodel consistency penalizes a difference between a first color predicted for the pixel by the NeRF submodel and a second color predicted for the pixel by a different NeRF submodel of the plurality of NeRF submodels.
3. The computer-implemented method of any preceding claim, wherein the different NeRF submodel of the plurality of NeRF submodels is randomly selected from a set of neighboring NeRF submodels associated with neighboring regions of the three-dimensional volume associated with the scene.
4. The computer-implemented method of any preceding claim, wherein each NeRF submodel is configured to apply a contraction function to enable each NeRF model to render an entirety of the three-dimensional volume associated with the scene.
5. The computer-implemented method of any preceding claim, wherein training, by the computing system, each NeRF submodel to render pixels that depict its correspondingregion comprises distilling, by the computing system, a shared teacher NeRF model into the plurality of NeRF submodels.
6. The computer-implemented method of any preceding claim, wherein the plurality of NeRF submodels share at least some of a set of model parameters.
7. The computer-implemented method of claim 6, wherein the set of model parameters are distributed in the three-dimensional volume associated with the scene, and wherein each of the plurality of NeRF submodels comprises a deferred shader model with parameters generated by interpolating at least some of the set of model parameters based on a current camera location within the three-dimensional volume associated with the scene.
8. The computer-implemented method of any preceding claim, wherein each of the plurality of NeRF submodels is configured to perform feature gating in which low-resolution voxel grid features are used to disable high-resolution triplanes features when rendering low- frequency content.
9. The computer-implemented method of any preceding claim, wherein associating, by the computing system, the plurality of NeRF submodels respectively with the plurality of regions comprises instantiating, by the computing system, an active set of NeRF submodels that correspond to regions of the three-dimensional volume that contain at least one training camera associated with a training image included in a training dataset.
10. A computing system for efficient neural rendering, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store: a set of neural radiance field (NeRF) rendering assets, wherein the set of NeRF rendering assets comprise one or more shared feature maps and parameter values for a plurality of NeRF submodels respectively associated with a plurality of regions of a three- dimensional volume associated with a scene, wherein the plurality of NeRF submodels havebeen trained using a submodel consistency loss that enforces visual consistency between the plurality of NeRF submodels; and instructions that, when executed by the one or more processors, cause the computing system to perform real-time rendering operations comprising: determining a current camera location within the three-dimensional volume associated with the scene; loading into memory the NeRF submodel associated with the region of the three-dimensional volume that contains the current camera location; and rendering the scene using the NeRF submodel loaded into memory.
11. The computing system of claim 10, wherein the real-time rendering operations further comprise: pre-loading into memory one or more of the NeRF submodels that are associated with one or more neighboring regions of the three-dimensional volume that are adjacent to the region of the three-dimensional volume that contains the current camera location.
12. The computing system of claim 10 or 11, wherein the real-time rendering operations further comprise: in response to the current camera location moving from a first region to a second region of the three-dimensional volume: hotswapping to a second NeRF model associated with the second region.
13. The computing system of any of claims 10-12, wherein all of the plurality of NeRF submodels have been distilled from a shared teacher NeRF model.
14. The computing system of any of claims 10-13, wherein the parameter values are distributed in the three-dimensional volume associated with the scene, and wherein loading into memory the NeRF submodel comprises interpolating at least some of the parameter values based on a location of the current camera location within the three-dimensional volume associated with the scene.
15. One or more non-transitory computer-readable media that collectively store a plurality of neural radiance field (NeRF) submodels as described in any of claims 1-14.
Citation Information
Cited By
4D space-time field construction method and device based on neural field reconstruction
CN121120981A