Real time image rendering for large scenes
Patent Information
- Application Number
- EP2024766152
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-07
- Filing Date
- 2024-03-07
- Publication Date
- 2026-01-14
AI Technical Summary
Achieving both speed and realism in rendering large-scale scenes for virtual worlds has been a long-standing challenge, particularly in providing photorealistic renderings at interactive frame rates for immersive experiences.
A method involving a shading machine learning model that processes a UV feature map and polygonal mesh to generate image renderings with opacity and color values, using a lightweight model to perform real-time rendering by transforming neural features into image renderings, rather than relying on ray casting.
Enables real-time realistic rendering of large scenes, suitable for virtual reality and autonomous system training, with the ability to simulate multiple cameras and diverse scenarios, improving the safety and performance of autonomous systems.
Smart Images

Figure CA2024050286_12092024_PF_FP_ABST
Abstract
Description
REAL TIME IMAGE RENDERING FOR LARGE SCENES CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is a non-provisional application of, and thereby claims benefit to U.S. Patent Application Serial Number 63 / 450,623 filed on March 7, 2023, which is incorporated herein by reference in its entirety. BACKGROUND
[0002] A virtual world is a computer-simulated environment, which interaction in a three-dimensional space as in the real world. In some cases, the virtual world is designed to replicate at least some aspects of the real world. For example, the virtual world may include objects and background reconstructed from the real world. The reconstructing of objects and background from the real world allows the system to replicate aspects of the real world.
[0003] The interaction with the virtual world is generally performed by rendering images for large-scale scenes. The scenes are based on the camera placement of the virtual camera within the virtual world. In many uses of virtual worlds, speed and realism is important. Specifically, the entity interacting with the virtual world should feel as if the entity is interacting in the real world in both relation to time and space. Namely, the scene should appear photorealistic renderings at interactive frame rates for an immersive and seamless experience. Achieving both speed and realism in large-scale scene rendering has been a long-standing challenge. SUMMARY
[0004] In general, in one aspect, one or more embodiments relate to a method that includes identifying a camera location of a camera in a geographic region, and rasterizing, using the camera location and a polygonal mesh, a UV feature map to obtain a feature buffer. The method also includes processing, by a shading machine learning model, the feature buffer using a view direction of the camera to generate an image rendering that includes opacity values andcolor values. The method further includes generating a rendered image from the opacity values and color values.
[0005] In general, in one aspect, one or more embodiments relate to a system that includes memory and a computer processor that includes computer readable program code for performing operations. The operations include identifying a camera location of a camera in a geographic region, and rasterizing, using the camera location and a polygonal mesh, a UV feature map to obtain a feature buffer. The operations also include processing, by a shading machine learning model, the feature buffer using a view direction of the camera to generate an image rendering that includes opacity values and color values. The operations further include generating a rendered image from the opacity values and color values.
[0006] In general, in one aspect, one or more embodiments relate to a non- transitory computer readable medium that includes computer readable program code for performing operations. The operations include identifying a camera location of a camera in a geographic region, and rasterizing, using the camera location and a polygonal mesh, a UV feature map to obtain a feature buffer. The operations also include processing, by a shading machine learning model, the feature buffer using a view direction of the camera to generate an image rendering that includes opacity values and color values. The operations further include generating a rendered image from the opacity values and color values.
[0007] Other aspects of the invention will be apparent from the following description and the appended claims. BRIEF DESCRIPTION OF DRAWINGS
[0008] FIG.1 shows a diagram of an autonomous training and testing system in accordance with one or more embodiments.
[0009] FIG.2 shows a flowchart of the autonomous training and testing system in accordance with one or more embodiments.
[0010] FIG.3 shows a diagram of a rendering system in accordance with one or more embodiments.
[0011] FIG.4 shows a flowchart for generating a rendered image in accordance with one or more embodiments.
[0012] FIG. 5 shows a flowchart for using a UV code map in accordance with one or more embodiments.
[0013] FIG.6 shows a flowchart for processing background in accordance with one or more embodiments.
[0014] FIG.7 shows a flowchart for training the rendering system in accordance with one or more embodiments.
[0015] FIG. 8 shows an example of a rendering system to perform real time rendering of a large scene in accordance with one or more embodiments.
[0016] FIG.9 shows an example of vector quantization to generate a code map in accordance with one or more embodiments.
[0017] FIGs.10A and 10B show a computing system in accordance with one or more embodiments of the invention.
[0018] Like elements in the various figures are denoted by like reference numerals for consistency. DETAILED DESCRIPTION
[0019] In general, embodiments are directed to real time realistic renderings of large scenes. In particular, one or more embodiments define a polygonal mesh for a virtual world and a UV feature map. The UV feature map has neural features for locations in the virtual world. The neural features are learned features that are generated through machine learning. Thus, neural features may not explicitly represent the appearance of objects in the virtual world. The neural features in the UV feature map and the polygonal mesh are used by a rasterization engine that extracts the neural features corresponding to the location of the camera in the virtual world. A shading machine learning modelis configured to transform the neural features into an image rendering having opacity values and color values. The shading machine learning model is a lightweight model that uses the neural features output from the rasterization engine along with the camera direction to generate the image rendering. For example, rather than using ray casting, the shading machine learning model may process feature vectors that include the neural features and the view direction through neural network layers. In one or more embodiments, by using lightweight machine learning model, the rendering may be performed in real time.
[0020] The rendering system may be used with virtual reality systems and with the training or testing of an autonomous system. In virtual reality systems, the user interacts with the virtual world as if the user were in the real world. The virtual world may be a computer-generated environment that may or may not be a virtualized version of the real world. For example, the virtual environment may have objects and background that does not exist in the real world. The rendering system provides the rendered images to another component of an overall system, such as a display, another model, or other component, etc. Further, to simulate multiple cameras, the rendering system may operate in parallel for the different cameras.
[0021] In some embodiments, the processing by system may be used to generate a virtual world that mimics the real world, but with different scenarios implemented. For example, the changed scenarios may be that dynamic and / or static objects are in different locations, the perspective of the player is changed because the player is in a different location than the player was in the real world, or other aspects of the real world are different.
[0022] Embodiments of the invention may be used as part of generating a simulated environment for the training and testing of autonomous systems. An autonomous system is a self-driving mode of transportation that does not require a human pilot or human driver to move and react to the real-world environment. Rather, the autonomous system includes a virtual driver that isthe decision-making portion of the autonomous system. The virtual driver is an artificial intelligence system that learns how to interact in the real world. The autonomous system may be completely autonomous or semi-autonomous. As a mode of transportation, the autonomous system is contained in a housing configured to move through a real-world environment. Examples of autonomous systems include self-driving vehicles (e.g., self-driving trucks and cars), drones, airplanes, robots, etc. The virtual driver is the software that makes decisions and causes the autonomous system to interact with the real-world including moving, signaling, and stopping or maintaining a current state.
[0023] The real-world environment is the portion of the real world through which the autonomous system, when trained, is designed to move. Thus, the real- world environment may include interactions with concrete and land, people, animals, other autonomous systems, human driven systems, construction, and other objects as the autonomous system moves from an origin to a destination. In order to interact with the real-world environment, the autonomous system includes various types of sensors, such as LiDAR sensors amongst other types, which are used to obtain measurements of the real-world environment, and cameras that capture images from the real-world environment.
[0024] The testing and training of the virtual driver of the autonomous systems in the real-world environment is unsafe because of the accidents that an untrained virtual driver can cause. Thus, as shown in FIG.1, a simulator (100) is configured to train and test a virtual driver (102) of an autonomous system. For example, the simulator may be a unified, modular, mixed-reality, closed- loop simulator for autonomous systems. The simulator (100) is a configurable simulation framework that enables not only evaluation of different autonomy components in isolation but also as a complete system in a closed-loop manner. The simulator reconstructs “digital twins” of real-world scenarios automatically, enabling accurate evaluation of the virtual driver at scale. The simulator (100) may also be configured to perform mixed-reality simulation that combines real world data and simulated data to create diverse and realisticevaluation variations to provide insight into the virtual driver’s performance. The mixed reality closed-loop simulation allows the simulator (100) to analyze the virtual driver’s action on counterfactual “what-if” scenarios that did not occur in the real-world. The simulator (100) further includes functionality to simulate and train on rare yet safety-critical scenarios with respect to the entire autonomous system and closed-loop training to enable automatic and scalable improvement of autonomy.
[0025] The simulator (100) creates the simulated environment (104) which is a virtual world. The virtual driver (102) is the player in the virtual world. The simulated environment (104) is a simulation of a real-world environment, which may or may not be in actual existence, in which the autonomous system is designed to move. As such, the simulated environment (104) includes a simulation of the objects (i.e., simulated objects or assets) and background in the real world, including the natural objects, construction, buildings and roads, obstacles, as well as other autonomous and non-autonomous objects. The simulated environment simulates the environmental conditions within which the autonomous system may be deployed. Additionally, the simulated environment (104) may be configured to simulate various weather conditions that may affect the inputs to the autonomous systems. The simulated objects may include both stationary and nonstationary objects. Nonstationary objects are actors in the real-world environment.
[0026] The simulator (100) also includes an evaluator (110). The evaluator (110) is configured to train and test the virtual driver (102) by creating various scenarios in the simulated environment. Each scenario is a configuration of the simulated environment including, but not limited to, static portions, movement of simulated objects, actions of the simulated objects with each other, and reactions to actions taken by the autonomous system and simulated objects. The evaluator (110) is further configured to evaluate the performance of the virtual driver using a variety of metrics.
[0027] The evaluator (110) assesses the performance of the virtual driver throughout the performance of the scenario. Assessing the performance may include applying rules. For example, the rules may be that the automated system does not collide with any other actor, compliance with safety and comfort standards (e.g., passengers not experiencing more than a certain acceleration force within the vehicle), the automated system not deviating from executed trajectory), or other rule. Each rule may be associated with the metric information that relates a degree of breaking the rule with a corresponding score. The evaluator (110) may be implemented as a data-driven neural network that learns to distinguish between good and bad driving behavior. The various metrics of the evaluation system may be leveraged to determine whether the automated system satisfies the requirements of the success criterion for a particular scenario. Further, in addition to system level performance, for modular based virtual drivers, the evaluator may also evaluate individual modules such as segmentation or prediction performance for actors in the scene with respect to the ground truth recorded in the simulator.
[0028] The simulator (100) is configured to operate in multiple phases as selected by the phase selector (108) and modes as selected by a mode selector (106). The phase selector (108) and mode selector (106) may be a graphical user interface or application programming interface component that is configured to receive a selection of phase and mode, respectively. The selected phase and mode define the configuration of the simulator (100). Namely, the selected phase and mode define which system components communicate and the operations of the system components.
[0029] The phase may be selected using a phase selector (108). The phase may be a training phase or a testing phase. In the training phase, the evaluator (110) provides metric information to the virtual driver (102), which uses the metric information to update the virtual driver (102). The evaluator (110) may further use the metric information to further train the virtual driver (102) by generating scenarios for the virtual driver. In the testing phase, the evaluator (110) doesnot provide the metric information to the virtual driver. In the testing phase, the evaluator (110) uses the metric information to assess the virtual driver and to develop scenarios for the virtual driver (102).
[0030] The mode may be selected by the mode selector (106). The mode defines the degree to which real-world data is used, whether noise is injected into simulated data, the degree of perturbations of real-world data, and whether the scenarios are designed to be adversarial. Example modes include open loop simulation mode, closed loop simulation mode, single module closed loop simulation mode, fuzzy mode, and adversarial mode. In an open loop simulation mode, the virtual driver is evaluated with real world data. In a single module closed loop simulation mode, a single module of the virtual driver is tested. An example of a single module closed loop simulation mode is a localizer closed loop simulation mode in which the simulator evaluates how the localizer estimated pose drifts over time as the scenario progresses in simulation. In a training data simulation mode, simulator is used to generate training data. In a closed loop evaluation mode, the virtual driver and simulation system are executed together to evaluate system performance. In the adversarial mode, the actors are modified to perform adversarial. In the fuzzy mode, noise is injected into the scenario (e.g., to replicate signal processing noise and other types of noise). Other modes may exist without departing from the scope of the system.
[0031] The simulator (100) includes the controller (112) which includes functionality to configure the various components of the simulator (100) according to the selected mode and phase. Namely, the controller (112) may modify the configuration of each of the components of the simulator based on the configuration parameters of the simulator (100). Such components include the evaluator (110), the simulated environment (104), an autonomous system model (116), sensor simulation models (114), asset models (117), actor models (118), latency models (120), and a training data generator (122).
[0032] The autonomous system model (116) is a detailed model of the autonomous system in which the virtual driver will execute. The autonomous system model (116) includes model, geometry, physical parameters (e.g., mass distribution, points of significance), engine parameters, sensor locations and type, the firing pattern of the sensors, information about the hardware on which the virtual driver executes (e.g., processor power, amount of memory, and other hardware information), and other information about the autonomous system. The various parameters of the autonomous system model may be configurable by the user or another system.
[0033] For example, if the autonomous system is a motor vehicle, the modeling and dynamics may include the type of vehicle (e.g., car, truck), make and model, geometry, physical parameters such as the mass distribution, axle positions, type and performance of the engine, etc. The vehicle model may also include information about the sensors on the vehicle (e.g., camera, LiDAR, etc.), the sensors’ relative firing synchronization pattern, and the sensors’ calibrated extrinsics (e.g., position and orientation) and intrinsics (e.g., focal length). The vehicle model also defines the onboard computer hardware, sensor drivers, controllers, and the autonomy software release under test.
[0034] The autonomous system model includes an autonomous system dynamic model. The autonomous system dynamic model is used for dynamics simulation that takes the actuation actions of the virtual driver (e.g., steering angle, desired acceleration) and enacts the actuation actions on the autonomous system in the simulated environment to update the simulated environment and the state of the autonomous system. To update the state, a kinematic motion model may be used, or a dynamics motion model that accounts for the forces applied to the vehicle may be used to determine the state. Within the simulator, with access to real log scenarios with ground truth actuations and vehicle states at each time step, embodiments may also optimize analytical vehicle model parameters or learn parameters of a neural network that infers the new state of the autonomous system given the virtual driver outputs.
[0035] In one or more embodiments, the sensor simulation models (114) models, in the simulated environment, active and passive sensor inputs. Passive sensor inputs capture the visual appearance of the simulated environment including stationary and nonstationary simulated objects from the perspective of one or more cameras based on the simulated position of the camera(s) within the simulated environment. Examples of passive sensor inputs include inertial measurement unit (IMU) and thermal. Active sensor inputs are inputs to the virtual driver of the autonomous system from the active sensors, such as LiDAR, RADAR, global positioning system (GPS), ultrasound, etc. Namely, the active sensor inputs include the measurements taken by the sensors, and the measurements being simulated based on the simulated environment based on the simulated position of the sensor(s) within the simulated environment. By way of an example, the active sensor measurements may be measurements that a LiDAR sensor would make of the simulated environment over time and in relation to the movement of the autonomous system. In one or more embodiments, all or a portion of the sensor simulation models (114) may be or include the rendering system (300) shown in FIG. 3. In such a scenario, the rendering system of the sensor simulation models (114) may perform the operations of FIGs.4-6.
[0036] The sensor simulation models (114) are configured to simulate the sensor observations of the surrounding scene in the simulated environment (104) at each time step according to the sensor configuration on the vehicle platform. When the simulated environment directly represents the real-world environment, without modification, the sensor output may be directly fed into the virtual driver. For light-based sensors, the sensor model simulates light as rays that interact with objects in the scene to generate the sensor data. Depending on the asset representation (e.g., of stationary and nonstationary objects), embodiments may use graphics-based rendering for assets with textured meshes, neural rendering, or a combination of multiple rendering schemes. Leveraging multiple rendering schemes enables customizable worldbuilding with improved realism. Because assets are compositional in 3D and support a standard interface of render commands, different asset representations may be composed in a seamless manner to generate the final sensor data. Additionally, for scenarios that replay what happened in the real world and use the same autonomous system as in the real world, the original sensor observations may be replayed at each time step.
[0037] Asset models (117) include multiple models, each model modeling a particular type of individual asset in the real world. The assets may include inanimate objects such as construction barriers or traffic signs, parked cars, and background (e.g., vegetation or sky). Each of the entities in a scenario may correspond to an individual asset. As such, an asset model, or instance of a type of asset model, may exist for each of the objects or assets in the scenario. The assets can be composed together to form the three-dimensional simulated environment. An asset model provides all the information needed by the simulator to simulate the asset. The asset model provides the information used by the simulator to represent and simulate the asset in the simulated environment.
[0038] Closely related to, and possibly considered part of the set of asset models (117) are actor models (118). An actor model represents an actor in a scenario. An actor is a sentient being that has an independent decision-making process. Namely, in the real world, the actor may be animate being (e.g., a person or animal) that makes a decision based on an environment. The actor makes active movement rather than or in addition to passive movement. An actor model, or an instance of an actor model may exist for each actor in a scenario. The actor model is a model of the actor. If the actor is in a mode of transportation, then the actor model includes the model of transportation in which the actor is located. For example, actor models may represent pedestrians, children, vehicles being driven by drivers, pets, bicycles, and other types of actors.
[0039] The actor model leverages the scenario specification and assets to control all actors in the scene and their actions at each time step. The actor’s behavioris modeled in a region of interest centered around the autonomous system. Depending on the scenario specification, the actor simulation will control the actors in the simulation to achieve the desired behavior. Actors can be controlled in various ways. One option is to leverage heuristic actor models, such as an intelligent-driver model (IDM) that try to maintain a certain relative distance or time-to-collision (TTC) from a lead actor or heuristic-derived lane- change actor models. Another is to directly replay actor trajectories from a real log or to control the actor(s) with a data-driven traffic model. Through the configurable design, embodiments may mix and match different subsets of actors to be controlled by different behavior models. For example, far-away actors that initially may not interact with the autonomous system and can follow a real log trajectory, but when near the vicinity of the autonomous system may switch to a data-driven actor model. In another example, actors may be controlled by a heuristic or data-driven actor model that still conforms to the high-level route in a real-log. This mixed-reality simulation provides control and realism.
[0040] Further, actor models may be configured to be in cooperative or adversarial mode. In cooperative mode, the actor model models actors to act rationally in response to the state of the simulated environment. In adversarial mode, the actor model may model actors acting irrationally, such as exhibiting road rage and bad driving.
[0041] In one or more embodiments, the actor models (118), asset models (117), and background may be part of the rendering system (described below with reference to FIG. 3). As another example, the system may be a bifurcated system whereby the operations (e.g., trajectories or positioning) of the assets and actors are defined separately from the appearance, which is part of the rendering system.
[0042] The latency model (120) represents timing latency that occurs when the autonomous system is in a real-world environment. Several sources of timing latency may exist. For example, a latency may exist from the time that an eventoccurs to the sensors detecting the sensor information from the event and sending the sensor information to the virtual driver. Another latency may exist based on the difference between the computing hardware executing the virtual driver in the simulated environment as compared to the computing hardware of the virtual driver. Further, another timing latency may exist between the time that the virtual driver transmits an actuation signal to the autonomous system changing (e.g., direction or speed) based on the actuation signal. The latency model (120) models the various sources of timing latency.
[0043] Stated another way, in the real world, safety-critical decisions in the real world may involve fractions of a second affecting response time. The latency model simulates the exact timings and latency of different components of the onboard system. To enable scalable evaluation without strict requirements on exact hardware, the latencies and timings of the different components of the autonomous system and sensor modules are modeled while running on different computer hardware. The latency model may replay latencies recorded from previously collected real world data or have a data-driven neural network that infers latencies at each time step to match the hardware in a loop simulation setup.
[0044] The training data generator (122) is configured to generate training data. For example, the training data generator (122) may modify real-world scenarios to create new scenarios. The modification of real-world scenarios is referred to as mixed reality. For example, mixed-reality simulation may involve adding in new actors with novel behaviors, changing the behavior of one or more of the actors from the real-world, and modifying the sensor data in that region while keeping the remainder of the sensor data the same as the original log. In some cases, the training data generator (122) converts a benign scenario into a safety-critical scenario.
[0045] The simulator (100) is connected to a data repository (105). The data repository (105) is any type of storage unit or device that is configured to store data. The data repository (105) includes data gathered from the real world. Forexample, the data gathered from the real world include real actor trajectories (126), real sensor data (128), real trajectories of the system capturing the real world (130), and real latencies (132). Each of the real actor trajectories (126), real sensor data (128), real trajectory of the system capturing the real world (130), and real latencies (132) is data captured by or calculated directly from one or more sensors from the real world (e.g., in a real-world log). In other words, the data gathered from the real-world are actual events that happened in real life. For example, in the case that the autonomous system is a vehicle, the real-world data may be captured by a vehicle driving in the real world with sensor equipment.
[0046] Further, the data repository (105) includes functionality to store one or more scenario specifications (140). A scenario specification (140) specifies a scenario and evaluation setting for testing or training the autonomous system. For example, the scenario specification (140) may describe the initial state of the scene, such as the current state of the autonomous system (e.g., the full 6D pose, velocity and acceleration), the map information specifying the road layout, and the scene layout specifying the initial state of all the dynamic actors and objects in the scenario. The scenario specification may also include dynamic actor information describing how the dynamic actors in the scenario should evolve over time which are inputs to the actor models. The dynamic actor information may include route information for the actors, desired behaviors or aggressiveness. The scenario specification (140) may be specified by a user, programmatically generated using a domain-specification- language (DSL), procedurally generated with heuristics from a data-driven algorithm, or adversarial-based generated. The scenario specification (140) can also be conditioned on data collected from a real-world log, such as taking place on a specific real-world map or having a subset of actors defined by their original locations and trajectories.
[0047] The interfaces between the virtual driver and the simulator match the interfaces between the virtual driver and the autonomous system in the realworld. For example, the sensor simulation model (114) and the virtual driver match the virtual driver interacting with the sensors in the real world. The virtual driver is the actual autonomy software that executes on the autonomous system. The simulated sensor data that is output by the sensor simulation model (114) may be in or converted to the exact message format that the virtual driver takes as input as if the virtual driver were in the real world, and the virtual driver can then run as a black box virtual driver with the simulated latencies incorporated for components that run sequentially. The virtual driver then outputs the exact same control representation that it uses to interface with the low-level controller on the real autonomous system. The autonomous system model (116) will then update the state of the autonomous system in the simulated environment. Thus, the various simulation models of the simulator (100) run in parallel asynchronously at their own frequencies to match the real- world setting.
[0048] FIG.2 shows a flow diagram for executing the simulator in a closed loop mode. In Block 201, a digital twin of a real-world scenario is generated as a simulated environment state. Log data from the real world is used to generate an initial virtual world. The log data defines which asset and actor models are used in the initial positioning of assets. For example, using convolutional neural networks on the log data, the various asset types within the real world may be identified. As other examples, offline perception systems and human annotations of log data may be used to identify asset types. Accordingly, corresponding asset and actor modes may be identified based on the asset types and add to the positions of the real actors and assets in the real world. Thus, the asset and actor models create an initial three-dimensional virtual world.
[0049] In Block 203, the sensor simulation model is executed on the simulated environment state to obtain simulated sensor output. The sensor simulation model may use beamforming and other techniques to replicate the view to the sensors of the autonomous system. Each sensor of the autonomous system has a corresponding sensor simulation model and a corresponding system. Thesensor simulation model executes based on the position of the sensor within the virtual environment and generates simulated sensor output. The simulated sensor output is in the same form as would be received from a real sensor by the virtual driver. In one or more embodiments, Block 203 may be performed as shown in FIGs. 4-6 (described below) to generate camera output and lidar sensor output, respectively, for a virtual camera and a virtual lidar sensor, respectively. The operations of FIGs.4-6 may be performed for each camera and lidar sensor on the autonomous system to simulate the output of the corresponding camera and lidar sensor. Location and viewing direction of the sensor with respect to the autonomous vehicle may be used to replicate the originating location of the corresponding virtual sensor on the simulated autonomous system. Thus, the various sensor inputs to the virtual driver match the combination of inputs if the virtual driver were in the real world.
[0050] The simulated sensor output is passed to the virtual driver. In Block 205, the virtual drive executes based on the simulated sensor output to generate actuation actions. The actuation actions define how the virtual driver controls the autonomous system. For example, for an SDV, the actuation actions may be the amount of acceleration, movement of the steering, triggering of a turn signal, etc. From the actuation actions, the autonomous system state in the simulated environment is updated in Block 207. The actuation actions are used as input to the autonomous system model to determine the actual actions of the autonomous system. For example, the autonomous system dynamic model may use the actuation actions in addition to road and weather conditions to represent the resulting movement of the autonomous system. For example, in a wet or snowy environment, the same amount of acceleration action as in a dry environment may cause less acceleration than in the dry environment. As another example, the autonomous system model may account for possibly faulty tires (e.g., tire slippage), mechanical based latency, or other possible imperfections in the autonomous system.
[0051] In Block 209, actors’ actions in the simulated environment are modeled based on the simulated environment state. Concurrently with the virtual driver model, the actor models and asset models are executed on the simulated environment state to determine an update for each of the assets and actors in the simulated environment. Here, the actors’ actions may use the previous output of the evaluator to test the virtual driver. For example, if the actor is adversarial, the evaluator may indicate based on the previous action of the virtual driver, the lowest scoring metric of the virtual driver. Using a mapping of metrics to actions of the actor model, the actor model executes to exploit or test that particular metric.
[0052] Thus, in Block 211, the simulated environment state is updated according to the actors’ actions and the autonomous system state to generate an updated simulated environment state. The updated simulated environment includes the change in positions of the actors and the autonomous system. Because the models execute independently of the real world, the update may reflect a deviation from the real world. Thus, the autonomous system is tested with new scenarios. In Block 213, a determination is made whether to continue. If the determination is made to continue, testing of the autonomous system continues using the updated simulated environment state in Block 203. At each iteration, during training, the evaluator provides feedback to the virtual driver. Thus, the parameters of the virtual driver are updated to improve the performance of the virtual driver in a variety of scenarios. During testing, the evaluator is able to test using a variety of scenarios and patterns including edge cases that may be safety critical. Thus, one or more embodiments improve the virtual driver and increase the safety of the virtual driver in the real world.
[0053] As shown, the virtual driver of the autonomous system acts based on the scenario and the current learned parameters of the virtual driver. The simulator obtains the actions of the autonomous system and provides a reaction in the simulated environment to the virtual driver of the autonomous system. The evaluator evaluates the performance of the virtual driver and creates scenariosbased on the performance. The process may continue as the autonomous system operates in the simulated environment.
[0054] FIG.3 shows a diagram of the rendering system (300) in accordance with one or more embodiments. The rendering system (300) is a system configured to generate a rendered image based on a position of a virtual camera in a geographic region of the virtual world. In one or more embodiments, the geographic region is a subregion of the virtual world. For example, the subregion may be the area up to a threshold distance of from the virtual camera in the virtual world. As another example, the geographic region may be an entire virtual world. In one or more embodiments, the rendering system (300) may inactively render the images as the point of view of the virtual camera changes or as objects move in the virtual world. The rendering system (300) includes a data repository (302) connected to a model framework (304).
[0055] The data repository (302) may include one or more of sensor data (128), polygonal mesh (306), background feature maps (308), a UV feature map (310), and code map (312).
[0056] The sensor data (128) is the sensor data described above with reference to FIG. 1. The sensor data (128) includes actual images (330). Actual images (330) are images captured by one or more cameras of the geographic region. For example, as a sensing vehicle is moving through a geographic region the sensing vehicle may have cameras that gather sensor data from the geographic region. Notably, the sensor data (128) is the time series of data that is captured along the trajectory of the sensing vehicle.
[0057] The polygonal mesh (308) is a mesh structure that maintains the geometry of a geographic region. The polygonal mesh (308) as defined herein corresponds to the standard definition used in the art and may also be referred to as a polygon mesh. The mesh has edges that connect vertices. The vertices have specific locations that map to locations in the geographic region. The combination of edges and vertices define faces. The faces correspond to the surfaces of objects in the geographic region. The faces are polygon faces, suchas triangles, quadrilaterals, or other n-gons. The number of vertices and correspondingly, faces, of the polygonal mesh (308) is configurable based on resolution and processing speed. For example, the polygonal mesh (308) may have fifty thousand vertices or five hundred thousand vertices.
[0058] The UV feature map (310) is a UV map that has neural features in the third dimension. The term UV map corresponds to a standard definition used in the art. The UV map maps locations of a three-dimensional object onto a two-dimensional plane. Thus, each location of the two-dimensional plane has a corresponding location on the three-dimensional object. Similarly, locations on the three-dimensional object each have a corresponding location on the two- dimensional plane. Locations on the UV map may be defined by a horizontal value and a vertical value. The horizontal value may be referred to as the U- axis and the vertical value may be referred to as a V-axis. Thus, the position u,v, where u ^ U and v ^ V, maps to a particular location on the three- dimensional object.
[0059] A UV feature map (310) is a UV map, but with a third dimension being a feature vector. The feature vector has learned neural features that are features of the appearance of a corresponding object at the location corresponding to position u,v. In other words, the third dimension is the feature vector with the neural features for the position defined by the first two dimensions. The neural features are learned from the real sensor data (128) and, as such, may not include direct attributes of color, luminosity, etc., but rather encoded features learned through machine learning.
[0060] In one or more embodiments, at least a portion of the virtual world is represented by a single UV feature map (310) and a single polygonal mesh (308). For example, a single UV feature map and a single polygonal mesh may have all of or several of the stationary objects in a geographic region of the virtual world. In some embodiments, nonstationary objects in the virtual world are represented by individual corresponding UV feature maps (310) andpolygonal meshes (308). Thus, nonstationary objects may move in the virtual world independently of the geographic environment.
[0061] The background feature maps (308) are one or more feature maps defined for the background of the geographic region. The background region is a region that is greater than a threshold distance from a virtual camera. Multiple background feature maps may be defined, whereby each background feature map corresponds to different distance range from locations in the geographic region. For example, a first background feature map may be for objects that would be 400-600 meters away from the camera in the virtual world if the virtual world were a real world, a second background feature map may be for objects 600-2000 meters, and a third background feature map may be for objects, such as the sky, that are more than 2000 meters. The ranges in the example are only for example purposes and not intended to limit the scope of the claims.
[0062] The background feature maps may be a set of neural skyboxes or skydomes, whereby the background is projected onto a cube. The use of the term skybox or skydome corresponds to the standard definition as used in the art of computer graphics. However, in one or more embodiments, rather than a color value, the value at a particular location is a feature vector. The feature vector has neural features that may be similar to the feature vectors of the UV feature map. For example, the neural features of the background feature maps (308) are learned from the real sensor data (128) and, as such, may not include direct attributes of color, luminosity, etc., but rather encoded features learned through machine learning. In one or more embodiments, the neural sky boxes or sky domes represents a scene as a set of cuboid, spherical, or other three- dimensional layers that may be defined on a two-dimensional plane. For cuboid layers, each layer may have six feature maps, where each feature map corresponds to a plane of the cuboid. Each layer may be an individual feature map.
[0063] In some embodiments, a code map (312) is used. In such embodiments, the UV feature map (310) or the background feature map (308) do not store feature vectors directly, but rather codes that map to the feature vectors. In such embodiments, the third dimension of the UV feature map is a code that maps to the neural features. The code map (312) maintains a mapping between codes and feature vectors. Thus, the same stored feature vector may be used by multiple locations of the UV feature map, thereby reducing the size of the UV feature map. In one or more embodiments, the code map (312) is a learned through a machine learning process. The code map (312) is a quantization of an actual UV feature map. Specifically, where multiple feature vectors are close, but not identical, the same code may be used, and a single feature vector may be stored for the multiple feature vectors. Thus, the size of the code map may be further reduced.
[0064] In one or more embodiments, a separate code map may also or alternatively exist for the background feature map(s). In such a scenario, the code map may be specifically trained and generated for one or more of the background feature maps. For example, each background feature map may have a separate code map, or the collection of background feature maps may share the code map. In other embodiments, the code map for the background feature map may be the same as the code map for the UV feature map. Regardless of the configuration, the code map and correspondingly, the background feature map using the code map, may operate in a same or similar manner and be trained in a same or similar manner as the code map of the UV feature map.
[0065] Continuing with FIG. 3, the model framework (304) includes a rasterization engine (314), a shading machine learning model (316), a code map training engine (318), a compositor (320), and a loss function (322). The rasterization engine is software configured to perform rasterization on the geographic region based on the camera. Rasterization is the process of taking the three-dimensional model (e.g., defined by the polygonal mesh) andconverting the three-dimensional model into a raster image which is made up of pixels. The pixel is based on the location of the camera. In one or more embodiments, each pixel has a corresponding feature vector for the pixel. In one or more embodiments, the rasterization engine has multiple distinct processes. A first process may perform the rasterization and a second process may sample the UV feature map or background feature map based on the rasterization.
[0066] The shading machine learning model (316) is a machine learning model that is configured to convert neural features into a color value and an opacity value. The shading machine learning model (316) may be a lightweight machine learning model. For example, the shading machine learning model (316) may be a multilayer perceptron (MLP) model. Generally, an MLP model is a feedforward artificial neural network having at least three layers of nodes. The layers include an input layer, a hidden layer, and an output layer. Each layer has multiple nodes. Each node includes an activation function with learnable parameters. Through training and backpropagation of losses, the parameters are updated and correspondingly, the MLP model improves in making predictions.
[0067] The shading model may include multiple machine learning models (e.g., MLP). The multiple machine learning models may each be an independent MLP. For example, the shading machine learning model may include multiple individual machine learning models, whereby each machine learning model is for at least one of a particular background feature of the plurality of background feature maps, the UV feature map, and the object feature map. In the example, the UV feature map may have an independent machine learning model that is separate from the machine learning model for one or more of the background feature maps. As another example, the background feature maps may have independent machine learning models. As another example, each nonstationary object may have an independent machine learning model for the object. Each of the machine learning models may be independent of the other machinelearning models in that the machine learning models may be at least in part individually trained. Further, the machine learning models may be separate in that at least one or more of the layers are not overlapping with the layers of the other models.
[0068] The code map training engine (318) is configured to train the code map (312). Specifically, the code map training engine (318) is configured to perform a training of which feature vectors should map to which codes in the code map.
[0069] The compositor (320) is configured to composite the image based on the output of the shading machine learning model. Compositing the image combines the different colors based on opacities of the parts of the image into a single image. For example, the compositing combines the foreground as specified by processing the UV map with the background and any nonstationary objects.
[0070] The loss function (322) is a function used to calculate the loss for the system. The loss function (322) uses the various outputs of the model framework (304) to calculate a loss that is used to update, through backpropagation, the model framework (304). During the backpropagation, one or more layers may be frozen to calculate the loss of the other layers.
[0071] FIGs.4-7 show flowcharts in accordance with one or more embodiments. FIG.5 shows a flowchart for generating a rendered image and FIG.5 shows a flowchart for using a code map. FIG. 6 shows a flowchart for rendering a background and compositing a rendered image. While the various steps in these flowcharts are presented and described sequentially, at least some of the steps may be executed in different orders, may be combined or omitted, and at least some of the steps may be executed in parallel. Furthermore, the steps may be performed actively or passively.
[0072] In Block 402, the camera location is identified. In the virtual world, the player may move around. The player may be referred to as an ego system,autonomous system, or viewpoint of the display in the case of a human. Thus, at any point in time, the player has a corresponding position in the virtual world. The player has a corresponding position of at least one virtual camera. In the case of virtual reality, the corresponding position of the virtual camera is the location of the player’s eye(s) in the virtual world. In the case of autonomous systems, the corresponding position of the camera is based on the location of the camera with respect to the autonomous system. Thus, the camera location in the virtual world may be identified. In some embodiments, multiple virtual cameras may exist. If multiple virtual cameras exist, the processing of FIGs. 4-6 is repeated for each camera.
[0073] In Block 404, using the camera location and polygonal mesh, the UV feature map is rasterized to obtain a feature buffer. The rasterization process obtains a u,v coordinate in the UV feature map for each pixel of the rendered image. The UV feature map is sampled to obtain to obtain the feature vector at the UV coordinate. If a code map is used, the sampling may be performed as described in FIG. 5. The process is repeated for each pixel in the rendered image to obtain the feature buffer that has the feature vectors. Each feature vector is related to a corresponding pixel in the rendered image. The rasterization engine may also generate an opacity mask from the polygonal mesh. The opacity mask indicates, for each pixel, whether the pixel is covered by the polygonal mesh.
[0074] In some embodiments, the opacity mask may be used at a compositing stage to composite the image. In other embodiments, the opacity mask may be used to determine whether to obtain the feature vector from the UV feature map.
[0075] In Block 406, a shading machine learning model processes the feature buffer using the view direction of the camera to generate a first set of opacity values and color values. In one or more embodiments, the shading machine learning model processes feature vectors for each pixel independently. For a feature vector, the view direction at the pixel is determined and combined withthe feature vector to generate a combined feature vector. For example, the combination may be a concatenation. Other combinations may be used without departing from the scope of the claims. In one or more embodiments, the view direction indicates the angle of the camera origin to the particular pixel. The shading machine learning model then processes the combined feature vector. For example, the combined feature vector may be processed through one or more layers of the neural network to generate a color value and an opacity value for the pixel. The process is repeated for each pixel to generate an image rendering.
[0076] In Block 408, a rendered image is generated using the first set of opacity values and the color values. In some embodiments, the image rendering is the rendered image and may be used directly. In such a scenario, the rendered image is outputted from the rendering system.
[0077] The process of FIG. 4 may be performed for multiple cameras concurrently. Each camera may have a corresponding view direction and camera location. When performing the processing, the same polygonal mesh and UV feature map may be used for the multiple cameras. Thus, multiple rendered images may be concurrently generated and outputted.
[0078] FIG.5 shows a flowchart for using a code map. In Block 502, the location in the UV feature map is identified. The location is obtained as described in Block 404. In Block 504, the code from the location in the UV feature map is obtained. Instead of obtaining the feature vector from the UV feature map directly, the code is obtained from the UV feature map. The code is used as a lookup into the code map. In Block 506, the feature set mapped to the code from the code map is obtained. In one or more embodiments, the feature set is a feature vector that is mapped to the code map. The feature set may be used to perform the processing described above with reference to FIG.4, Blocks 404 and 406. Because the size(s) of the codes are smaller than the size(s) of the feature vectors, the UV feature map is smaller. The process of FIG. 5 isrepeated for each location of the UV feature map that is identified by the rasterization.
[0079] FIG. 6 shows a flowchart for processing a background. As discussed above, the background may have a different processing than the foreground to save processing time. In other embodiments, a polygonal mesh and a UV feature map may be used as described in FIG.4 to generate the background.
[0080] In Block 602, a camera location in a geographic region is identified. Identifying the camera location may be performed as described above with reference to Block 402 of FIG.4.
[0081] In Block 604, using the camera location, the background feature maps are rasterized to obtain background feature buffers. For each pixel of the image, a camera ray is projected from the camera to the locations in the background feature maps. The intersection location with each of the background feature maps are determined to identify the location in each of the background feature maps. The feature set for the location in each of the background feature maps is obtained and added to the corresponding feature buffer. Thus, in one or more embodiments, each background feature buffer has a feature set for each location that is related to the location.
[0082] In Block 606, the shading machine learning model processes the background feature buffers to generate one or more background renderings. The processing is performed similar to the processing of Block 406 of FIG.4. In one or more embodiments, the processing of Block 606 is performed independently for each feature set for each background pixel, or each pixel, in the image. In one or more embodiments, the feature set is processed through one or more layers of a neural network model of the shading machine learning model. In some embodiments, the view direction is appended to the feature vector. In other embodiments, the view direction is not used when the shading machine learning model processes the feature sets.
[0083] The shading machine learning model may include includes multiple machine learning models. A separate machine learning model may exist for the foreground as for the background. Further, a separate machine learning model may exist for each layer of the background. Accordingly, processing the background feature buffer(s) may include processing, by the trained neural network layers of the corresponding machine learning model, the feature set of each pixel in a layer of the background to obtain an opacity and color value for the pixel in the layer.
[0084] The process of FIG.6 may be extended to include nonstationary objects in the virtual world. As discussed above, nonstationary objects may have corresponding feature maps that are trained and polygonal meshes (i.e., object polygonal meshes) that are trained. Thus, the location of the nonstationary object in the virtual world is identified. Based on the location of the camera with respect to the nonstationary object and the polygonal mesh, the object feature map is rasterized to obtain an object feature buffer. The shading machine learning model then processes the object feature buffer to obtain color values and opacity values for corresponding locations on the object in the image. The process may be repeated for each nonstationary object in the image.
[0085] In Block 608, the one or more background renderings are composited with first set of opacity and color values generated in Block 406 of FIG.4 to generate a rendered image. After the processing of Block 606 and 406 of FIG.4, each pixel may have one or more pairs of opacity value and color value defined for the pixel. If multiple pairs opacity values and color values are defined for the same pixel, then the relative distance to the camera is based on the layer in which the pixel is located is used to order the pair from closest to farthest from the camera. For example, the pair for the foreground are closer than the pairs for the background pixel, which are ordered according to the layers. The pairs are processed in order. For each pair defined for a pixel, the opacity value is used to determine the amount of color that is used for the pixel as compared to the amount of color from farther pairs.
[0086] Stated another way, at the stage of processing of Block 608, the same pixel in the image may have a foreground color value and opacity value pair, a background color value and opacity value pair, and zero, one, or more object color value and opacity value pairs. The pairs of values for the pixel are ordered according to distance from the camera. The opacity value of each current pair is then used, in order of distance, to calculate the amount of color value from the farther pairs as compared to the color value for the current pair. For example, the opacity value of zero on the closest pair indicates that the color value of the closest pair is used. In another example, an opacity value of fifty percent indicates that fifty percent of the color values from the current pair is combined with fifty percent of farther pairs. The combination may be a weighted combination of the current pair and the farther pairs, whereby the weights are determined by the opacity value of the current pair. Further, the combination may be independently performed for the Red, Green, Blue (RGB) portions of the color value. Although RGB is described, other color channels of other color models may be used in a same technique.
[0087] FIG.7 shows a flowchart for training the system in accordance with one or more embodiments. In Block 702, the models are used to generate rendered images. In Block 704, the UV feature maps, background feature maps, code maps, and shading models are trained. Stored camera images are compared to the rendered images to generate a loss. Multiple comparisons may be performed, and the loss may be a combined loss. A first comparison may be a direct comparison of color values of each pixel. Specifically, for each pixel, difference between the color value of the rendered image and the color value of the real camera image is calculated. The differences are combined across the different pixels to obtain the color loss. A second loss may be a perceptual loss. The perceptual loss is the loss determined from the overall image. For the perceptual loss, a trained machine learning model generates a first value from the rendered image and a second value from the real camera image. The difference between the first value and the second value is the perceptual loss. A third loss may bebased on the vector quantization. Namely, the third loss may be based on the amount of quantization of the code map. The vector quantization loss may be calculated as the difference between the original feature and the nearest code in the code map.
[0088] The various losses may be combined to generate a combination loss, such as by performing a weighted averaging. The combination loss may be back propagated through the network. For the UV feature map, the loss updates the neural features. Namely, the system learns the neural network features that are referenced by codes in or directly in the UV feature map.
[0089] In Block 706, a determination is made whether to continue training. If training is to continue, the flow proceeds to Block 702.
[0090] FIG. 8 and FIG. 9 show examples in accordance with one or more embodiments. The discussion of FIG. 8 and FIG. 9 is for example purposes and not intended to limit the scope of the claims.
[0091] One or more embodiments aims to perform real-time rendering of large- scale scenes. Given a set of posed images and a moderate-quality reconstructed mesh, our method generates a scene mesh with neural texture maps and view- dependent fragment shaders. Using the initial mesh, One or more embodiments first generate a UV parameterization to learn neural textures on. One or more embodiments then jointly learn a discrete texture feature codebook and view- dependent lightweight MLPs that can effectively represent scene appearance. Finally, One or more embodiments bake the texture feature codebook and the MLPs into a set of neural texture maps and a fragment shader that can be run in real time with existing graphics pipelines. One or more embodiments now first introduce our approach for representing large-scale scenes (Sec.3.1), then describe how One or more embodiments render and learn the scene (Sec.3.2- 3.3), and finally how One or more embodiments export our model into real time graphics pipelines (Sec.3.4).
[0092] FIG. 8 shows an example of a rendering system to perform real time rendering of a large scene in accordance with one or more embodiments. FIG. 8 shows an example of rendering large-scale outdoor scenes. In order to handle potentially infinite depth ranges (e.g., sky, vegetation, mountain, etc.) as well as nearby regions, a hybrid approach may be used. The entire 3D scene may be partitioned into two regions: an inner cuboid region (foreground) modelled by a polygonal mesh textured with neural features, and an outer cuboid region (background) modelled by neural skyboxes. Such a hybrid scene representation allows the modeling in fine-grained details in both close-by regions and far- away regions, and enables rendering with a remarkable degree of camera movement.
[0093] For the foreground region, one or more embodiments may leverage an explicit geometry mesh scaffold to learn and render neural textures. Various sources of existing polygonal mesh may be used. For example, one or more embodiments can leverage existing neural reconstruction methods or other methods. Initially, the reconstructed mesh may have over tens of millions of triangle faces, which represents the geometry well, but may have self- intersections and duplicate vertices. One or more embodiments may preprocess the obtained polygonal mesh to reduce computational cost and improve UV mapping quality. One or more embodiments may first cluster nearby vertices together and perform quadric mesh decimation to simplify the mesh while preserving structure, and then perform face culling to remove non-visible triangle faces (e.g., source camera views). Finally, a UV map generation tool may be used to unfold the mesh to obtain the UV mappings for each of the polygonal mesh’s vertices.
[0094] In the example, the resulting triangle mesh ^^ ൌ ^ ^^, ^^, ^^^ consists of vertex positions ^^ ∈ ^^ேൈଷ, vertex UV coordinates ^^ ∈ ^^ேൈଶ, and a set of triangle faces ^^. Based on the generated UV mapping, one or more embodiments initialize a learnable UV feature map ^^ ∈ ^^^ൈ^ൈ^to representthe scene appearance covered by the mesh. Using neural features instead of a color texture map enables modelling view-dependent effects during rendering.
[0095] For the background, in the example, a challenge may exist to model the far-away background regions with polygonal mesh because of the complexity and scale of that region. As an alternative, one or more embodiments may use the concept of multi-plane images and multi-sphere images to represent the background region using neural skyboxes. A neural skybox is an example of a background feature map.
[0096] The neural skyboxes may represents the scene as a set of cuboid layers. In the example, each layer ^^^^ , ^^^^^ contains 6 individual feature mapsthat each represents one plane of the cuboid. The background feature maps represent both geometry and view-dependent appearance of the scene in the example. The neural skyboxes may represent a wide range of depths and may be integrated in existing graphics pipelines to enable efficient rendering.
[0097] Turning specifically to the example of FIG. 8, the rendering pipeline is shown. By way of an overview of FIG. 8, the foreground mesh and neural skyboxes are rasterized with neural texture maps to the desired view point, producing a set of image feature buffers. The feature buffers are then processed with MLPs to produce a set of rendering layers, which are composited to synthesize the final RGB image. The process of FIG.8 is described below.
[0098] For the foreground, given a camera pose and intrinsics, one or more embodiments first rasterize (808) the polygonal mesh (802) into screen space, obtaining a UV coordinate ^ ^^, ^^^ for each pixel ^ ^^, ^^^ on the screen. One or more embodiments, then sample the UV feature map ^^ ∈ ^^^ൈ^ൈ^(804) using the rasterized UV coordinates and obtain a feature buffer ^^^∈ ^^ுൈ^ൈ^(810): ^^^^ ^^, ^^^ ൌ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^^ ^^, ^^, ^^^, (1) where ^^ ൈ ^^ is the UV feature map resolution, ^^ ൈ ^^ is the rendering resolution, and ^^ is the feature dimension. In addition to a feature buffer, the rasterization (808) may also generate an opacity mask ^^^∈ ^^ுൈ^(not shown)to indicate whether a pixel is covered by the polygonal mesh (802). To render the RGB image, one or more embodiments concatenate the rendered feature with the view direction ^^^ ^^, ^^^ (814) and pass through a learnable MLPs shader ^^ఏ^ (816):^^^^ ^^, ^^^ ൌ ^^ఏ^^ ^^^^ ^^, ^^^, ^^^ ^^, ^^^^, (2) where ^^்is the MLP parameters, ^^^^ ^^, ^^^ is the rendered RGB color for pixel ^ ^^, ^^^.
[0099] For the background, to render the neural skybox feature layers ^ ^^^^^^ୀ^representing distant background regions, one or more embodiments project camera ray shooting each pixel ^ ^^, ^^^ and compute the ray’s intersection points ^ ^^^^^^ୀ^with layers 1 to ^^, from near to far. Next, one or more embodiments sample the features ^ ^^^^^^ୀ^corresponding to the intersection points on the neural skybox feature map (806) at each layer, generating a set of feature buffers ^^^∈ ^^ுൈ^ൈ^(812), where ^^ ൌ 1,⋯The processing can be efficiently performed with the OpenGL rasterizer (808). One or more embodiments then use a learnable MLPs shader ^^ఏೄ(816) to process the background feature buffers (812) and outputs the opacity map ^^^∈ ^^ுൈ^and RGB color map ^^^∈ ^^ுൈ^ൈଷfor each layer (820): ^^^^ ^^, ^^^, ^^^^ ^^, ^^^ ൌ ^^ఏೄ^^^^^ ^^, ^^^, ^^^ ^^, ^^^^, (3)where ^^ௌrepresents the parameters of the MLPs (816). The MLPs shader first processes the input feature ^^^^ ^^, ^^^, and outputs opacity ^^^^ ^^, ^^^ and an intermediate feature vector. The feature vector is then concatenated with ^^^ ^^, ^^^, the view direction of camera ray, and passed to the last layers that output the view- dependent color ^^^^ ^^, ^^^.
[0100] The foreground and background are then composited. To synthesize the final RGB image, one or more embodiments composite (822) the rendered layers from the foreground mesh ^ ^^^^ ^^, ^^^, ^^^^ ^^, ^^^^ and neural skyboxes^ ^^^^ ^^, ^^^, ^^^^ ^^, ^^^^^^ୀ^by repeatedly compositing the RGB and opacity layers, from near to far: ^^^ ^^, ^^^ ൌ∑^^ୀ^ ^^^^ ^^, ^^^ ⋅ ^^^^ ^^, ^^^ ⋅∏^^ିୀ^^^1 െ ^^^^ ^^, ^^^^. (4)The term ^^^^ ^^, ⋅ ^^, and ∏^^ି^ ^1 െ ^^^^ ^^^^ denotes the fraction of the color that remains afterthe layers in front. The compositing process ensures that the RGB values are correctly blended.
[0101] To encourage the sharing of latent features in visually similar regions such as roads and sky, one or more embodiments apply vector quantization (VQ) to regularize the neural texture maps. The vector quantization allows the features to be supervised from a large range of view directions which improves view- point extrapolation performance. Furthermore, the VQ also compacts the feature representations, reducing the offline storage space. One or more embodiments may maintain two codebooks ^^்and ^^ௌ, each consists of ^^ learnable latent code ^^^∈ ^^^, with ^^ ൌ 1,⋯ , ^^. In the forward pass, one or more embodiments quantize the UV feature map ^^ா∈ ^^^ൈ^ൈ^and neural skyboxes feature maps ^^ா∈ ^^^ൈ^ൈ^ൈ^by mapping each feature to its closest latent code in the codebook: ൌ ^^െ where ^^, ^^ are the spatial coordinates of the feature map, and ^^ is the layer index of the neural skyboxes. Fig. 1 shows the feature map quantization process. One or more embodiments use the quantized features to compute the synthesized image.
[0102] The training may be performed as follows. The feature map ^^, ^^, the codebook ^^், ^^ௌ, and the parameters ^^், ^^ௌof the MLP shaders may be jointly optimized by minimizing the photometric loss and perceptual loss between therendered images and camera observations, as well as the VQ regularizer. The full objective may be: ^^ ൌ ^^^^^^ ^^^^^^^^^^^^^ ^^௩^^^௩^, (6)Where each loss term in more
[0103] Photometric loss may be calculated as follows. ^^^^^measures the ^^ଶdistance between the rendered and the observed images. The photometric loss may be defined as: ^^^^^ ൌ∥ ^^ െ ^^^ ∥ଶ, (7)where ^^ is the rendered image from Eqn.4 and ^^^ is the corresponding observed camera image.
[0104] Perceptual loss may be calculated as follows. One or more embodiments use an additional perceptual loss to enhance the rendered image quality. The perceptual loss measures the “perceptual similarity" that is more consistent with human visual perception:where ^^^denotes the i-th layer with ^^^elements of the pre-trained VGG Network.
[0105] The VQ loss may be calculated as follows. To update the codebook ^^, one or more embodiments may define the VQ loss term as:െଶെଶെ ^^ ∥ଶଶ^ ^^ ∥ ^^ ^^^ ^^ா^ െ ^^ ∥ଶଶ, (9) where ^^ ^^^⋅^ denotes the stop-gradient operator that behaves as the identity map at forward pass and has zero partial derivatives at backward pass. The first two terms form the alignment loss and encourage the codebook latents to follow the feature maps. The last two terms form the commitment loss which stabilizes training by discouraging the features from learning much faster than thecodebook. The quantization step in Eqn.5 is non-differentiable. One or more embodiments approximate the gradient of the feature maps ^^, ^^ using the straight-through estimator, which simply passes the gradient from the quantized feature to the original feature unaltered during back-propagation.
[0106] FIG. 9 shows an example of vector quantization process to generate a code map (902) and a quantized UV feature map (904) in accordance with one or more embodiments. As shown in FIG.9, the original UV feature map (906) having continuous feature values T(u,v) at location u,v is processed through a codebook (i.e., code map) function to generate the code map (902) and quantized feature map (904).
[0107] To enable rendering in real time, one or more embodiments convert our scene representations and multilayer perceptrons (MLPs) to be compatible with the graphics rendering pipeline. The mesh, skyboxes, and texture representations may be directly compatible with the OpenGL, while the learned MLPs ^^ఏ^ and ^^ఏೄ are converted to fragment shaders in OpenGL. During each rendering pass, the triangle mesh ^^ and the skyboxes ^ ^^^^^^ୀ^may be rasterized to the screen as a set of fragments, and each fragment is associated with a feature vector that is bilinearly sampled from the neural texture maps. The fragment shader then maps each fragment’s features to RGB color and opacity. To ensure correct alpha compositing, the scene mesh and cuboids are sorted depth-wise and rendered from back to front, following the procedure outlined in Eqn.4.
[0108] As shown, one or more embodiments use a scaffold mesh as input and incorporates a neural texture field to model view-dependent effects, which can then be exported and rendered in real-time with standard rasterization engines. One or more embodiments may render urban driving scenes at 1920×1080 resolution at over 100 FPS while delivering comparable realism to existing neural rendering approaches. Thus, one or more embodiments may be used for scalable and immersive experiences for self-driving simulation and virtual reality applications.
[0109] Embodiments may be implemented on a computing system specifically designed to achieve an improved technological result. When implemented in a computing system, the features and elements of the disclosure provide a significant technological advancement over computing systems that do not implement the features and elements of the disclosure. Any combination of mobile, desktop, server, router, switch, embedded device, or other types of hardware may be improved by including the features and elements described in the disclosure. For example, as shown in FIG. 10A, the computing system (1000) may include one or more computer processors (1002), non-persistent storage (1004), persistent storage (1006), a communication interface (1012) (e.g., Bluetooth interface, infrared interface, network interface, optical interface, etc.), and numerous other elements and functionalities that implement the features and elements of the disclosure. The computer processor(s) (1002) may be an integrated circuit for processing instructions. The computer processor(s) may be one or more cores or micro-cores of a processor. The computer processor(s) (1002) includes one or more processors. The one or more processors may include a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), combinations thereof, etc.
[0110] The input devices (1010) may include a touchscreen, keyboard, mouse, microphone, touchpad, electronic pen, or any other type of input device. The input devices (1010) may receive inputs from a user that are responsive to data and messages presented by the output devices (1008). The inputs may include text input, audio input, video input, etc., which may be processed and transmitted by the computing system (1000) in accordance with the disclosure. The communication interface (1012) may include an integrated circuit for connecting the computing system (1000) to a network (not shown) (e.g., a local area network (LAN), a wide area network (WAN) such as the Internet, mobile network, or any other type of network) and / or to another device, such as another computing device.
[0111] Further, the output devices (1008) may include a display device, a printer, external storage, or any other output device. One or more of the output devices may be the same or different from the input device(s). The input and output device(s) may be locally or remotely connected to the computer processor(s) (1002). Many different types of computing systems exist, and the aforementioned input and output device(s) may take other forms. The output devices (1008) may display data and messages that are transmitted and received by the computing system (1000). The data and messages may include text, audio, video, etc., and include the data and messages described above in the other figures of the disclosure.
[0112] Software instructions in the form of computer readable program code to perform embodiments may be stored, in whole or in part, temporarily or permanently, on a non-transitory computer readable medium such as a CD, DVD, storage device, a diskette, a tape, flash memory, physical memory, or any other computer readable storage medium. Specifically, the software instructions may correspond to computer readable program code that, when executed by a processor(s), is configured to perform one or more embodiments, which may include transmitting, receiving, presenting, and displaying data and messages described in the other figures of the disclosure.
[0113] The computing system (1000) in FIG.10A may be connected to or be a part of a network. For example, as shown in FIG.10B, the network (1020) may include multiple nodes (e.g., node X (1022), node Y (1024)). Each node may correspond to a computing system, such as the computing system shown in FIG. 10A, or a group of nodes combined may correspond to the computing system shown in FIG. 10A. By way of an example, embodiments may be implemented on a node of a distributed system that is connected to other nodes. By way of another example, embodiments may be implemented on a distributed computing system having multiple nodes, where each portion may be located on a different node within the distributed computing system. Further, one ormore elements of the aforementioned computing system (1000) may be located at a remote location and connected to the other elements over a network.
[0114] The nodes (e.g., node X (1022), node Y (1024)) in the network (1020) may be configured to provide services for a client device (1026), including receiving requests and transmitting responses to the client device (1026). For example, the nodes may be part of a cloud computing system. The client device (1026) may be a computing system, such as the computing system shown in FIG.10A. Further, the client device (1026) may include and / or perform all or a portion of one or more embodiments.
[0115] The computing system of FIG.10A may include functionality to present raw and / or processed data, such as results of comparisons and other processing. For example, presenting data may be accomplished through various presenting methods. Specifically, data may be presented by being displayed in a user interface, transmitted to a different computing system, and stored. The user interface may include a GUI that displays information on a display device. The GUI may include various GUI widgets that organize what data is shown as well as how data is presented to a user. Furthermore, the GUI may present data directly to the user, e.g., data presented as actual data values through text, or rendered by the computing device into a visual representation of the data, such as through visualizing a data model.
[0116] As used herein, the term “connected to” contemplates multiple meanings. A connection may be direct or indirect (e.g., through another component or network). A connection may be wired or wireless. A connection may be temporary, permanent, or semi-permanent communication channel between two entities.
[0117] The various descriptions of the figures may be combined and may include or be included within the features described in the other figures of the application. The various elements, systems, components, and steps shown in the figures may be omitted, repeated, combined, and / or altered as shown in thefigures. Accordingly, the scope of the present disclosure should not be considered limited to the specific arrangements shown in the figures.
[0118] In the application, ordinal numbers (e.g., first, second, third, etc.) may be used as an adjective for an element (i.e., any noun in the application). The use of ordinal numbers is not to imply or create any particular ordering of the elements nor to limit any element to being only a single element unless expressly disclosed, such as by the use of the terms "before", "after", "single", and other such terminology. Rather, the use of ordinal numbers is to distinguish between the elements. By way of an example, a first element is distinct from a second element, and the first element may encompass more than one element and succeed (or precede) the second element in an ordering of elements.
[0119] Further, unless expressly stated otherwise, or is an “inclusive or” and, as such includes “and.” Further, items joined by an or may include any combination of the items with any number of each item unless expressly stated otherwise.
[0120] In the above description, numerous specific details are set forth in order to provide a more thorough understanding of the disclosure. However, it will be apparent to one of ordinary skill in the art that the technology may be practiced without these specific details. In other instances, well-known features have not been described in detail to avoid unnecessarily complicating the description. Further, other embodiments not explicitly described above can be devised which do not depart from the scope of the claims as disclosed herein. Accordingly, the scope should be limited only by the attached claims.
Claims
CLAIMS What is claimed is:
1. A computer-implemented method comprising: identifying a first camera location of a first camera in a geographic region; rasterizing, using the first camera location and a polygonal mesh, a UV feature map to obtain a first feature buffer; processing, by a shading machine learning model, the first feature buffer using a first view direction of the first camera to generate a first image rendering comprising a first set of opacity values and color values; and generating a first rendered image from the first set of opacity values and color values.
2. The computer-implemented method of claim 1, further comprising: rasterizing, using the first camera location, a background feature map to obtain a background feature buffer; processing, by the shading machine learning model, the background feature buffer to generate a background rendering comprising second set of opacity values and color values; and compositing the first set of opacity values and color values with the second set of opacity values and color values.
3. The computer-implemented method of claim 2, wherein the shading machine learning model comprises a first machine learning model that generates the image rendering and a second machine learning model, separate from the first machine learning model, that generates the background rendering.
4. The computer-implemented method of claim 1, further comprising: rasterizing, using the first camera location, a plurality of background feature maps to obtain a plurality of background feature buffers, wherein the plurality of background feature maps each have a corresponding distance from the first camera location;processing, by the shading machine learning model, the plurality of background feature buffers to generate a plurality of background renderings; and compositing the image rendering with the plurality of background renderings.
5. The computer-implemented method of claim 4, wherein the shading machine learning model comprises a plurality of individual machine learning models for at least one of a particular background feature of the plurality of background feature maps and the UV feature map.
6. The computer-implemented method of claim 1, further comprising: rasterizing, using the first camera location and an object polygonal mesh, an object feature map to obtain an object feature buffer; processing, by the shading machine learning model, the object feature buffer to generate an object rendering; and compositing the image rendering with the object rendering.
7. The computer-implemented method of claim 1, further comprising: rasterizing, using the first camera location, a plurality of background feature maps to obtain a plurality of background feature buffers, wherein the plurality of background feature maps each have a corresponding distance from the first camera location; processing, by the shading machine learning model, the plurality of background feature buffers to generate a plurality of background renderings; and rasterizing, using the first camera location and an object polygonal mesh, an object feature map of an object to obtain an object feature buffer, wherein the object is moving in the geographic region; processing, by the shading machine learning model, the object feature buffer to generate an object rendering; and compositing the image rendering with the object rendering and the plurality of background renderings, wherein the shading machine learning model comprises a plurality of individual machine learning models for at least one of a particular backgroundfeature of the plurality of background feature maps, the UV feature map, and the object feature map.
8. The computer-implemented method of claim 1, wherein rasterizing the UV feature map comprises: identifying a location in the UV feature map, obtaining a code from the location in the UV feature map, and obtaining a feature set mapped to the code from a code map.
9. The computer-implemented method of claim 1, further comprising: training the UV feature map to learn a plurality of neural network features from a camera image.
10. The computer-implemented method of claim 1, further comprising: training the UV feature map, a code map, and the shading machine learning model using a plurality of camera images; 11. The computer-implemented method of claim 1, further comprising: identifying a second camera location of a second camera in the geographic region; rasterizing, using the second camera location and the polygonal mesh, the UV feature map to obtain a second feature buffer; processing, by the shading machine learning model, the second feature buffer using a second view direction of the second camera to generate a second image rendering comprising a second set of opacity values and color values; generating a second rendered image from the second set of opacity values and color values; and outputting the first rendered image and the second rendered image concurrently.
12. A system comprising: memory; anda computer processor comprising computer readable program code for performing operations comprising: identifying a camera location of a camera in a geographic region, rasterizing, using the camera location and a polygonal mesh, a UV feature map to obtain a first feature buffer, processing, by a shading machine learning model, the first feature buffer using a view direction of the camera to generate an image rendering comprising a first set of opacity values and color values, and generating a rendered image from the first set of opacity values and color values.
13. The system of claim 12, wherein the operations further comprise: rasterizing, using the camera location, a background feature map to obtain a background feature buffer; processing, by the shading machine learning model, the background feature buffer to generate a background rendering comprising second set of opacity values and color values; and compositing the first set of opacity values and color values with the second set of opacity values and color values.
14. The system of claim 13, wherein the shading machine learning model comprises a first machine learning model that generates the image rendering and a second machine learning model, separate from the first machine learning model, that generates the background rendering.
15. The system of claim 12, wherein the operations further comprise: rasterizing, using the camera location, a plurality of background feature maps to obtain a plurality of background feature buffers, wherein the plurality of background feature maps each have a corresponding distance from the camera location;processing, by the shading machine learning model, the plurality of background feature buffers to generate a plurality of background renderings; and compositing the image rendering with the plurality of background renderings.
16. The system of claim 15, wherein the shading machine learning model comprises a plurality of individual machine learning models for at least one of a particular background feature of the plurality of background feature maps and the UV feature map.
17. The system of claim 12, wherein the operations further comprise: rasterizing, using the camera location and an object polygonal mesh, an object feature map to obtain an object feature buffer; processing, by the shading machine learning model, the object feature buffer to generate an object rendering; and compositing the image rendering with the object rendering.
18. The system of claim 12, wherein the operations further comprise: rasterizing, using the camera location, a plurality of background feature maps to obtain a plurality of background feature buffers, wherein the plurality of background feature maps each have a corresponding distance from the camera location; processing, by the shading machine learning model, the plurality of background feature buffers to generate a plurality of background renderings; and rasterizing, using the camera location and an object polygonal mesh, an object feature map of an object to obtain an object feature buffer, wherein the object is moving in the geographic region; processing, by the shading machine learning model, the object feature buffer to generate an object rendering; and compositing the image rendering with the object rendering and the plurality of background renderings, wherein the shading machine learning model comprises a plurality of individual machine learning models for at least one of a particular backgroundfeature of the plurality of background feature maps, the UV feature map, and the object feature map.
19. A non-transitory computer readable medium comprising computer readable program code for performing operations comprising: identifying a camera location of a camera in a geographic region; rasterizing, using the camera location and a polygonal mesh, a UV feature map to obtain a first feature buffer; processing, by a shading machine learning model, the first feature buffer using a view direction of the camera to generate an image rendering comprising a first set of opacity values and color values; and generating a rendered image from the first set of opacity values and color values.
20. The non-transitory computer readable medium of claim 19, wherein the operations further comprise: rasterizing, using the camera location, a plurality of background feature maps to obtain a plurality of background feature buffers, wherein the plurality of background feature maps each have a corresponding distance from the camera location; processing, by the shading machine learning model, the plurality of background feature buffers to generate a plurality of background renderings; and rasterizing, using the camera location and an object polygonal mesh, an object feature map of an object to obtain an object feature buffer, wherein the object is moving in the geographic region; processing, by the shading machine learning model, the object feature buffer to generate an object rendering; and compositing the image rendering with the object rendering and the plurality of background renderings, wherein the shading machine learning model comprises a plurality of individual machine learning models for at least one of a particular background feature of the plurality of background feature maps, the UV feature map, and the object feature map.