Neural calibration framework

The neural calibration framework addresses the challenges of costly and time-consuming multi-sensor calibration for autonomous systems by using simulated sensor renderings and loss value calculations to update calibration data, resulting in improved accuracy and efficiency.

WO2025102180A1PCT designated stage expired Publication Date: 2025-05-22WAABI INNOVATION INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2024/051524
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-11-18
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Current calibration methods for autonomous systems, such as self-driving vehicles, are costly, time-consuming, and require extensive infrastructure, making it challenging to achieve accurate and efficient multi-sensor calibration.

Method used

A neural calibration framework that uses a feature grid and calibration data to generate simulated sensor renderings, calculates loss values between simulated and real-world renderings, and updates the feature grid and calibration data to improve calibration accuracy.

Benefits of technology

This approach enables automatic, targetless multi-sensor calibration, reducing operational overhead and improving calibration accuracy, while eliminating the need for expensive infrastructure and manual effort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2024051524_22052025_PF_FP_ABST
    Figure CA2024051524_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A method implements a neural calibration framework. The method includes generating a simulated sensor rendering using a feature grid and calibration data. The method further includes generating a loss value between the simulated sensor rendering and a real-world sensor rendering. The method further includes updating the feature grid and the calibration data using the loss value to generate an updated feature grid and updated calibration data to reduce a subsequent loss value. The method further includes calibrating the sensor with the updated calibration data.
Need to check novelty before this filing date? Find Prior Art

Description

NEURAL CALIBRATION FRAMEWORKCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims benefit to US Provisional Application 63 / 600,640, filed November 17, 2023, which is hereby incorporated by reference herein.BACKGROUND

[0002] Autonomous systems, including robotic systems and self-driving vehicles (SDVs), observe the world to perceive and plan safe actions. To observe the world, autonomous systems are equipped with a suite of sensors, which may include light detection and ranging (LiDAR) sensors and camera sensors, that provide depth and appearance information about surroundings. Autonomous systems use accurate calibration of sensor intrinsics and extrinsics, i.e., calibration data that encodes the relative poses between sensors, to interpret and process sensor observations in a shared coordinate frame, such as for multisensor perception or scene reconstruction. A slight shift in the intrinsics and extrinsics estimations may result in an unacceptable misalignment between observations for distant objects, which may result in catastrophic failure.

[0003] Generating calibration data for multi-sensor intrinsics and extrinsics calibration is an arduous process that may utilize large infrastructure, operation costs, and manual effort. The process may involve collecting sensor data in a controlled indoor environment, with fiducials such as checkerboards mounted to fixed locations or held by operators. The autonomous system observes the fiducials from various viewpoints and distances. To enable proper coverage and repeatability, turntables may be employed. Reliance on infrastructure incurs heavy costs, both operationally and in hardware (e.g., vehicle-sized turntables).

[0004] Traditional calibration methods may leverage extracted geometric features and exact checkerboard dimensions to compute correspondences and estimate the relative poses between different sensors. Stage-wise calibrationmay be used to break down the problem: LiDAR-LiDAR, camera-camera, and camera-LiDAR pairs are first obtained, followed by a global optimization stage such as pose graph optimization.

[0005] Additionally, as sensors may shift after long periods of driving, autonomous systems are driven back to the facility for re-calibration during normal operations. The complexity of the process may include multiple sources of error due to hardware, operations, and software that may occur that result in poor calibration, such as warped fiducials, not observing fiducials with enough views that cause ambiguity, or poor convergence of the calibration due to conflicting pose estimates during global optimization. Such a calibration process has high cost and time overhead that makes scaling efficiently challenging. A “drive-and-calibrate” approach (driving the autonomous system outdoors, running an algorithm, and automatically calibrating the multi-sensor intrinsics and extrinsics) suffers from the lack of clear fiducials used by calibration methods.SUMMARY

[0006] In general, in one or more aspects, the disclosure relates to a method of a neural calibration framework. The method includes generating a simulated sensor rendering using a feature grid and calibration data. The method further includes generating a loss value between the simulated sensor rendering and a real-world sensor rendering. The method further includes updating the feature grid and the calibration data using the loss value to generate an updated feature grid and updated calibration data to reduce a subsequent loss value. The method further includes calibrating the sensor with the updated calibration data.

[0007] In general, in one or more aspects, the disclosure relates to a system that includes at least one processor and an application that executes on the at least one processor. Executing the application performs generating a simulated sensor rendering using a feature grid and calibration data. Executing theapplication further performs generating a loss value between the simulated sensor rendering and a real-world sensor rendering. Executing the application further performs updating the feature grid and the calibration data using the loss value to generate an updated feature grid and updated calibration data to reduce a subsequent loss value. Executing the application further performs calibrating the sensor with the updated calibration data.

[0008] In general, in one or more aspects, the disclosure relates to a non- transitory computer readable medium including instructions executable by at least one processor. Executing the instructions performs generating a simulated sensor rendering using a feature grid and calibration data. Executing the instructions further performs generating a loss value between the simulated sensor rendering and a real-world sensor rendering. Executing the instructions further performs updating the feature grid and the calibration data using the loss value to generate an updated feature grid and updated calibration data to reduce a subsequent loss value. Executing the instructions further performs calibrating the sensor with the updated calibration data.

[0009] Other aspects of one or more embodiments may be apparent from the following description and the appended claims.BRIEF DESCRIPTION OF DRAWINGS

[0010] FIG. 1 and FIG. 2 show diagrams of autonomous training and testing systems in accordance with the disclosure.

[0011] FIG. 3 shows a diagram of a system in accordance with the disclosure.

[0012] FIG. 4 shows a flowchart of a method in accordance with the disclosure.

[0013] FIG. 5 and FIG. 6 show example flow diagrams in accordance with the disclosure.

[0014] FIG. 7A and FIG. 7B show computing systems in accordance with the disclosure.

[0015] Similar elements in the various figures are denoted by similar names and reference numerals. The features and elements described in one figure may extend to similarly named features and elements in different figures.DETAILED DESCRIPTION

[0016] One or more embodiments are directed to a neural calibration framework that implements an automatic, targetless (e.g., without specific fiducials), multi-sensor calibration method. The multi-sensor calibration method is based on neural rendering that computes simulated observations using intrinsics and extrinsics for an autonomous system equipped with multiple sensors and multiple types of sensors. Types of sensors may include camera sensors, LiDAR sensors, radio detection and ranging (RADAR) sensors, etc. Replicating sensor observations through the use of neural rendering provides a supervision signal for calibration, without requiring targets or structured scene data. The intrinsics and extrinsics may be learnable parameters of the neural rendering process that are updated during training to accurately correspond to the positions of the sensors. The neural calibration framework may enhance neural rendering specifically for multi-sensor calibration by incorporating a surface-guided alignment loss and a coarse-to-fine sampling approach based on robust feature correspondences, leading to improvements in both calibration and 3D scene reconstruction quality.

[0017] Turning to the Figures, FIG. 1 and FIG. 2 show example diagrams of the autonomous system and virtual driver. Turning to FIG. 1, an autonomous system (116) is a self-driving mode of transportation that does not require a human pilot or human driver to move and react to the real-world environment. The autonomous system (116) may be completely autonomous or semi-autonomous. As a mode of transportation, the autonomous system (116) is contained in a housing configured to move through a real-world environment. Examples of autonomous systems include self-driving vehicles (e.g., self-driving trucks and cars), drones, airplanes, robots, etc.

[0018] The autonomous system (116) includes a virtual driver (102) that is the decision-making portion of the autonomous system (116). The virtual driver (102) is an artificial intelligence system that learns how to interact in the real world and interacts accordingly. The virtual driver (102) is the software executing on a processor that makes decisions and causes the autonomous system (116) to interact with the real world including moving, signaling, and stopping, or maintaining a current state. Specifically, the virtual driver (102) is decisionmaking software that executes on hardware (not shown). The hardware may include a hardware processor, memory or other storage device, and one or more interfaces that form a computing system. A hardware processor is any hardware processing unit that is configured to process computer readable program code and perform the operations set forth in the computer readable program code.

[0019] A real-world environment is the portion of the real world through which the autonomous system (116), when trained, is designed to move. Thus, the real- world environment may include concrete and land, construction, and other objects in a geographic region along with agents. The agents are the other agents in the real-world environment that are capable of moving through the real-world environment. Agents may have independent decision-making functionality. The independent decision-making functionality of the agent may dictate how the agent moves through the environment and may be based on visual or tactile cues from the real -world environment. For example, agents may include other autonomous and non-autonomous transportation systems (e.g., other vehicles, bicyclists, robots), pedestrians, animals, etc.

[0020] In the real world, the geographic region is an actual region within the real world that surrounds the autonomous system. Namely, from the perspective of the virtual driver, the geographic region is the region through which the autonomous system moves. The geographic region includes agents and map elements that are located in the real world. Namely, the agents and map elements each have a physical location in the geographic region that denotes a place in which the corresponding agent or map element is located. The map elements arestationary in the geographic region, whereas the agents may be stationary or nonstationary in the geographic region. The map elements are the elements shown in a map (e.g., road map, traffic map, etc.) or derived from a map of the geographic region.

[0021] The real-world environment changes as the autonomous system (116) moves through the real -world environment. For example, the geographic region may change and the agents may move positions, including new agents being added and existing agents leaving.

[0022] In order to interact with the real-world environment, the autonomous system (116) includes various types of sensors (104), such as LiDAR sensors, radio detection and ranging (RADAR) sensors, and camera sensors, which are used to obtain measurements of the real-world environment. The autonomous system (116) may include other types of sensors as well. The sensors (104) provide input to the virtual driver (102).

[0023] The calibration system (103) generates calibration data that may be used by the autonomous system (116) for the sensors (104). The calibration data may include intrinsics and extrinsics as well as temporal calibration. The intrinsics and extrinsics may be used to align the output of the sensors (104) with the location of the autonomous system (116). Intrinsics may include internal parameters for a sensor. For camera sensors, intrinsics may include focal length (distance between the optical center of a lens and the image sensor when the lens is focused at infinity), principle point (point on the image plane where the optical axis of the lens intersects), distortion (the bending of light passing through a lens of the camera sensor), etc. For LiDAR sensors, intrinsics may include laser wavelength, laser beam angles, pulse repetition rate, field of view, scan pattern, etc. For RADAR sensors, intrinsics may include frequency, pulse width, antenna configuration, modulation scheme, etc. Extrinsics for the sensors (104) may include external parameters that define spatial relationships between the sensors (104) and a reference from of the autonomous system (116). Temporal calibration may align the times a LiDAR sensor shoots beamsto the time the camera sensor is exposed, which may be inaccurate due to inaccurate clocks in the sensors and the autonomous system (116).

[0024] In addition to sensors (104), the autonomous system (116) includes one or more actuators (108). An actuator is hardware and / or software that is configured to control one or more physical parts of the autonomous system based on a control signal from the virtual driver (102). In one or more embodiments, the control signal specifies an action for the autonomous system (e.g., turn on the blinker, apply brakes by a defined amount, apply accelerator by a defined amount, turn the steering wheel or tires by a defined amount, etc. . The actuator(s) (108) are configured to implement the action. In one or more embodiments, the control signal may specify a new state of the autonomous system and the actuator may be configured to implement the new state to cause the autonomous system to be in the new state. For example, the control signal may specify that the autonomous system should turn by a certain amount while accelerating at a predefined rate, while the actuator determines and causes the wheel movements and the amount of acceleration on the accelerator to achieve a certain amount of turn and acceleration rate.

[0025] The testing and training of the virtual driver (102) of the autonomous systems in the real-world environment may be unsafe because of the accidents that an untrained virtual driver can cause. Thus, as shown in FIG. 2, a simulator (200) is configured to train and test a virtual driver (102) of an autonomous system. For example, the simulator may be a unified, modular, mixed-reality, closed-loop simulator for autonomous systems. The simulator (200) is a configurable simulation framework that enables not only evaluation of different autonomy components of the virtual driver (102) in isolation, but also as a complete system in a closed-loop manner. The simulator reconstructs “digital twins” of real-world scenarios automatically, enabling accurate evaluation of the virtual driver at scale. The simulator (200) creates the simulated environment (204) which is a virtual world in which the virtual driver (102) is a player in the virtual world. The simulated environment (204) is a simulation of a real-worldenvironment, which may or may not be in actual existence, in which the autonomous system is designed to move. As such, the simulated environment (204) includes a simulation of the objects (z.e., simulated objects or agents) and background in the real world, including the natural objects, construction, buildings and roads, obstacles, as well as other autonomous and non-autonomous objects. The simulated environment simulates the environmental conditions within which the autonomous system may be deployed. The simulated objects may include both stationary and nonstationary objects. Nonstationary objects are agents in the real-world environment.

[0026] In the simulated environment, the geographic region is a realistic representation of a real-world region that may or may not be in actual existence. Namely, from the perspective of the virtual driver, the geographic region appears the same as if the geographic region were in existence if the geographic region does not actually exist, or the same as the actual geographic region present in the real world. The geographic region in the simulated environment includes virtual agents and virtual map elements that would be actual agents and actual map elements in the real world. Namely, the virtual agents and virtual map elements each have a physical location in the geographic region that denotes an exact spot or place in which the corresponding agent or map element is located. The map elements are stationary in the geographic region, whereas the agents may be stationary or nonstationary in the geographic region. As with the real world, a map exists of the geographic region that specifies the physical locations of the map elements.

[0027] The simulator (200) includes an autonomous system model (216), sensor simulation models (214), and agent models (218). The autonomous system model (216) is a detailed model of the autonomous system in which the virtual driver (102) will execute. The autonomous system model (216) includes model, geometry, physical parameters (e.g., mass distribution, points of significance), engine parameters, sensor locations and type, firing pattern of the sensors, information about the hardware on which the virtual driver executes (e.g.,processor power, amount of memory, and other hardware information), and other information about the autonomous system. The various parameters of the autonomous system model may be configurable by the user or another system.

[0028] The autonomous system model (216) includes an autonomous system dynamic model. The autonomous system dynamic model is used for dynamics simulation that takes the actuation actions of the virtual driver (e.g., steering angle, desired acceleration) and enacts the actuation actions on the autonomous system in the simulated environment to update the simulated environment and the state of the autonomous system. The interface between the virtual driver (102) and the simulator (200) may match the interface between the virtual driver (102) and the autonomous system in the real world. Thus, to the virtual driver (102), the simulator simulates the experience of the virtual driver within the autonomous system in the real world.

[0029] In one or more embodiments, the sensor simulation model (214) models, in the simulated environment, active and passive sensor inputs. The sensor simulation models (114) are configured to simulate the sensor observations of the surrounding scene in the simulated environment (204) at each time step according to the sensor configuration on the vehicle platform. Passive sensor inputs capture the visual appearance of the simulated environment including stationary and nonstationary simulated objects from the perspective of one or more cameras based on the simulated position of the camera(s) within the simulated environment. Examples of passive sensor inputs include inertial measurement unit (IMU) and thermal. Active sensor inputs are inputs to the virtual driver of the autonomous system from the active sensors, such as LiDAR, RADAR, global positioning system (GPS), ultrasound, etc. Namely, the active sensor inputs include the measurements taken by the sensors, and the measurements being simulated based on the simulated environment based on the simulated position of the sensor(s) within the simulated environment.

[0030] Agent models (218) represents an agent in a scenario. An agent is a sentient being that has an independent decision-making process. Namely, in areal world, the agent may be an animate being (e.g., person or animal) that makes a decision based on an environment. The agent makes active movement rather than or in addition to passive movement. An agent model, or an instance of an actor model may exist for each agent in a scenario. The agent model is a model of the agent. If the agent is in a mode of transportation, then the agent model includes the model of transportation in which the agent is located. For example, actor models may represent pedestrians, children, vehicles being driven by drivers, pets, bicycles, and other types of actors.

[0031] The training of the virtual driver may be performed by using log data. Log data from the real world is used to generate an initial virtual world. The log data may be used to define, at least in part, which asset and actor models are used in an initial positioning of assets. The simulator executes a sensor simulation model that may use beamforming and other techniques to replicate the view to the sensors of the autonomous system.

[0032] The simulated sensor output is passed to the virtual driver. The virtual driver executes based on the simulated sensor output to generate actuation actions. The actuation actions define how the virtual driver controls the autonomous system. For example, for a self-driving vehicle, the actuation actions may be the amount of acceleration, movement of the steering, triggering of a turn signal, etc. From the actuation actions, the autonomous system state in the simulated environment is updated. Further, actors’ actions in the simulated environment are modeled based on the simulated environment state. Concurrently with the virtual driver model, the actor models and asset models are executed on the simulated environment state to determine an update for each of the assets and actors in the simulated environment. The actors’ actions may use the previous output of the evaluator to test the virtual driver.

[0033] The updated simulated environment state is updated according to the actors’ actions and the autonomous system state. The updated simulated environment includes the change in positions of the actors and the autonomous system. The process can repeat for a next timestep. Because the models executeindependently of the real world, the update may reflect a deviation from the real world. Thus, the autonomous system is tested with new scenarios.

[0034] One or more embodiments may be used outside of autonomous systems. For example, embodiments may be used in scene reconstruction for gaming environments and mapping environments and simulation. Further, sensors on robots and driving vehicles may be calibrated using embodiments described in Exhibit A. Further, embodiments may be used to estimate poses on and perform reconstruction such as with mobile devices or other mobile sensor platforms.

[0035] Turning to FIG. 3, the calibration system (300) is a collection of hardware and software components that implement a framework to automatically calibrate multi-sensor platforms without specific calibration targets. The calibration system (300) may include at least one processor and memory storing data and programs that execute on the processor. The calibration system (300) processes data from sensors (e.g., the real-world sensor renderings A (353) through N (363)) and optimizes sensor intrinsics and extrinsics (e.g., the calibration data A (305) through N (307)) alongside an underlying scene representation (e.g., the feature grid (303)) as learnable parameters. The intrinsics and extrinsics (corresponding to sensor alignment) and the scene representation are optimized to reduce photometric and geometric consistency errors between simulated sensor renderings and real-world sensor renderings. By concurrently learning the intrinsics, extrinsics, and scene representation, the calibration system (300) may generate accurate calibration data for sensor alignment in unstructured outdoor environments, reducing operational overhead compared to traditional calibration methods.

[0036] The learned parameters (301) are a collection of data that includes learnable elements within the calibration system (300). The learned parameters (301) include the feature grid (303), and the calibration data A (305) through N (307). During calibration, the parameters undergo iterative updates to improve alignment between simulated sensor renderings and real-world sensor data. By using the learned parameters (301), the calibration system (300) adapts tovarious sensor configurations and scenes, enhancing flexibility and generalizability across multi-sensor platforms.

[0037] The feature grid (303) is a collection of data that represents a multiresolution encoding of 3D scene geometry and appearance. The feature grid (303) is a grid of features that may correspond to the points in the 3D scene geometry. The feature grid (303) may have a different resolution than the sensors A (351) through N (361). Each point in the feature grid (303) may include a vector of scalar values that quantify the features at the point, which may then be interpolated for intermediate points. As an example, the feature grid (303) may have a resolution of 10 centimeters (cm), and when a point inside a cell of the grid is queried, the features for the point may be calculated by interpolating the values of the features from the corners of the current cell. The feature grid (303) is the basis for generating simulated sensor renderings through neural rendering techniques. By querying the feature grid (303) at different 3D locations, the calibration system (300) reconstructs the scene from arbitrary viewpoints. For the initial iteration, the feature grid (303) may be filled with random data that is updated with each iteration to better represent the 3D scene. The feature grid (303) undergoes optimization alongside sensor calibration parameters, facilitating learning of an accurate scene representation consistent across multiple sensors and viewpoints.

[0038] The calibration data A (305) through N (307) are collections of data that encode intrinsic and extrinsic parameters for the sensors A (351) through N (361), respectively, within the multi-sensor platform. The intrinsic parameters for an camera sensor may include focal length, principal point, distortion parameters, etc. The intrinsic parameters for a LiDAR sensor may include laser azimuth beam angle, laser elevation beam angle, etc. The extrinsic parameters may include the 6-degree-of-freedom pose (rotation and translation) of the sensor relative to a reference frame on an autonomous system. The calibration data A (305) is used to transform sensor observations into a common coordinate system for comparison, optimization, and further processing by other systems.Refinement of the calibration data A (305) during the calibration process may improve the alignment of the data generated by the sensor A with respect to other sensors and the underlying scene representation. The calibration data A (305) may be for a first type of sensor, which may be a camera sensor, a light detection and ranging (LiDAR) sensor, etc.

[0039] The calibration data N (307) may be similar to calibration data A (305) but for a different sensor in the multi-sensor platform. The calibration data N (307) may be for a different type of sensor than the calibration data A (305). For example, the calibration data A (305) may be for a camera sensor and the calibration data N (307) may be for a LiDAR sensor. Joint optimization of the calibration data A (305) through the calibration data N (307) may align across multiple sensors, even with different modalities or fields of view.

[0040] The sensor models A (311) through N (331) are collection of programs that may simulate the behavior of the sensors A (351) through N (361), respectively. The different sensor models A (311) through N (331) may be for different sensors and be for different sensor types. The sensor models A (311) through N (331) may define sets of rays that are input to the ray sampler (313).

[0041] The ray sampler (313) is a collection of programs that samples (z.e., selects) rays for rendering for the calibration process. Rays may be sampled uniformly or with increased density in regions of interest identified through feature detection. The ray sampler (313) may implement a coarse-to-fine strategy, progressively increasing the probability of selecting rays corresponding to informative areas of the scene. Intelligently selecting rays improves the efficiency and effectiveness of the calibration process. The output of the ray sampler (313) includes the sampled rays A (315) through N (335).

[0042] The sampled rays A (315) through N (335) are rays selected by the ray sampler (313) from the rays defined by the sensor models A (311) through N (331). The sampled rays A (315) through N (335) may respectively correspond to rays defined by the sensor models A (311) through N (331) and respectivelycorrespond to the sensors A (351) through N (361). For example, one of the sampled rays A (315) may correspond to one of the rays output from the sensor model A (311).

[0043] The point sampler (317) is a collection of programs that samples (z.e., selects) points from the sampled rays A (315) through N (335). The point sampler (317) generates the scene points A (319) through N (339) respectively from the sampled rays A (315) through N (335).

[0044] The scene points A (319) through N (339) are collections of data that represent specific 3D locations within the environment being modeled. A scene point is generated by sampling along a selected ray that was defined by a sensor model. By sampling multiple scene points, the calibration system (300) may reconstruct a scene from different viewpoints and sensor perspectives. The scene points A (319) through N (339) may respectively correspond to the sensors A (351) through N (361). As an example, one of the scene points A (319) may be from one of the sampled rays A (315). The scene points A (319) through N (339) may be input to the geometric model (321).

[0045] The geometric model (321) is a collection of programs that process information from the feature grid (303) and the scene points A (319) through N (339) to generate geometric representations of the environment. The geometric model (321) may be implemented as a neural network that processes interpolations of the feature grid (303) based on the scene points A (319) through N (339) to output the feature descriptors A (323) through N (343) and the signed-distance functions A (325) through N (345). The geometric model (313) encodes scene geometry in a differentiable manner that enables end-to- end optimization of both scene representation and sensor calibration using backpropagation, gradient descent, etc.

[0046] The feature descriptors A (323) through N (343) are collections of data generated by the geometric model (313) to encode appearance and geometric information respectively for the scene points A (319) through N (339). Thefeature descriptors A (323) through N (343) may be used by the sensor models A (311) through N (331). The feature descriptors A (323) through N (343) may be respectively input to the Tenderers A (327) through N (347). As an example, the feature descriptors A (323) may be generated from the feature grid (303) and from the scene points A (319) corresponding to the sensor model (311). The feature descriptors N (343) may also be generated from the same underlying scene representation (z.e., from the feature grid (303)) and correspond to the sensor model N (331). To reiterate, the scene points A (319) through N (339) for the sensor models A (311) through N (331) may each utilize the same shared feature grid (z.e., the feature grid (303)) and similar volume rendering procedure so that each of the sensor models A (311) through N (331) uses the same underlying scene representation during optimization to create alignment.

[0047] The signed-distance functions A (325) through N (345) are collections of data output by the geometric model (321) that respectively encode the distance from the scene points A (319) through N (339) to the nearest surface in the environment. As an example, the signed-distance function A (325) may correspond to the scene points A (319), the sensor model A (311), and the sensor model A (351). Positive values for a signed-distance function may indicate points outside of objects, while negative values may indicate points inside of objects. The surface itself may be implicitly defined where the signed- distance functions A (325) through N (345) equal zero. The signed-distance functions A (325) through N (345) provide continuous representations of scene geometry to facilitate differentiable rendering and optimization of the scene representation. The signed-distance functions A (325) through N (345) may be inputs to the Tenderers A (327) through N (347) and to the regularization loss (389).

[0048] The Tenderers A (327) through N (347) render the simulated sensor renderings A (329) through N (349) (also referred to as simulated sensor observations). Different Tenderers may be used for different sensors or differenttypes of sensors. For example, Tenderers for camera sensors may generate RGB colors for simulated images and Tenderers for LiDAR sensors may generate LiDAR intensities and distances for simulated point clouds.

[0049] The simulated sensor renderings A (329) through N (349) are collections of data generated by the Tenderers A (327) through N (347) that represents synthetic sensor observations. The simulated sensor renderings A (329) through N (349) are produced using the current calibration parameters and scene representation (i.e., the current version of the learned parameters (301). Comparison of the simulated sensor renderings A (329) through N (349) with real -world sensor data allows the calibration system (300) to compute loss values to guide optimization of sensor calibration and scene representation.

[0050] The sensors A (351) through N (361) may be hardware components that capture data about the environment, such as a camera or LiDAR sensor. The sensors A (351) through N (361) may include multiple types of sensors. The sensors A (351) through N (361) provide real -world measurements that may be used to calibrate the multi-sensor platform. The data from the sensors A (351) through N (361) may be used as ground truths for comparing against simulated sensor renderings during the calibration process. The sensors A (351) through N (361) may respectively generate the real-world sensor renderings A (353) through N (363).

[0051] The real-world sensor renderings A (353) through N (363) are collections of data captured, respectively, by the sensors A (351) through N (361) in the physical environment. The real -world sensor renderings A (353) through N (363) may include images, point clouds, or other sensor-specific data formats that correspond to the type of sensor that generated the renderings. Comparison of the real-world sensor renderings A (353) through N (363) with the simulated sensor renderings A (329) through N (349) are performed to compute loss values to guide optimization of sensor calibration and scene representation.

[0052] The loss function (371) is a collection of programs that may compute multiple error metrics between simulated sensor renderings and real-world sensor data. The loss function (371) combines multiple loss terms to assess the alignment between sensors and the accuracy of the scene representation. The loss function (371) may generate a scalar value representing the overall discrepancy between the current calibration and scene representation and the observed real-world data. The scalar value output by the loss function (371) may be backpropagated through the differentiable models to refine sensor calibration parameters (e.g., the calibration data A (305) through N (307)) and improve the scene representation (e.g., the feature grid (303)).

[0053] The color rendering loss (373) (which may also be referred to as a camera rendering loss) is a value that quantifies the difference between simulated and real -world color images from camera sensors. As an example, when the sensor A (351) is a camera, the color rendering loss (373) compares pixel values between the simulated sensor rendering A (329) and the real-world sensor rendering A (353). The color rendering loss (373) may be computed as a per- pixel Euclidean distance in RGB space or using perceptual metrics that account for human visual sensitivity. The color rendering loss (373) contributes to the overall loss value (391) to drive improvements in camera calibration and scene appearance modeling through updates to the calibration data A (305) through N (307) and the feature grid (303).

[0054] The intensity rendering loss (375) is a value that measures the discrepancy between simulated and real-world intensity values, which may be from LiDAR sensors. As an example, when the sensor N (361) is a LiDAR sensor, the intensity rendering loss (375) compares intensity values between the simulated sensor rendering N (349) and the real-world sensor rendering N (363). The intensity rendering loss (375) may be computed as a Euclidean distance between corresponding intensity values. The intensity rendering loss (375) aids in refining LiDAR calibration and improving the modeling of surfacereflectance properties in the scene representation through updates to the calibration data A (305) through N (307) and the feature grid (303).

[0055] The photometric consistency loss (377) is a value that combines the color rendering loss (373) and the intensity rendering loss (375) to assess overall photometric alignment across multiple sensors. The photometric consistency loss (377) encourages consistent appearance modeling across different sensor modalities and viewpoints. The photometric consistency loss (377) may weight the contributions of color rendering loss (373) and the intensity rendering loss (375) based on sensor characteristics or calibration priorities. The photometric consistency loss (377) is a component of the overall loss value (391), driving joint optimization of multi-sensor calibration and scene representation through updates to the calibration data A (305) through N (307) and the feature grid (303).

[0056] The surface alignment distance (379) is a value that quantifies the geometric consistency between different sensor observations of the same scene point. The surface alignment distance (379) computes the discrepancy between ray-casted points on the implicit surface representation for corresponding pixels or rays across multiple sensors. The surface alignment distance (379) may be normalized by projecting the 3D points onto image planes to obtain pixel-space distances. The surface alignment distance (379) is a measure of the alignment between sensors based on the current calibration data and scene representation.

[0057] The alignment loss (381) is a value that aggregates multiple surface alignment distances (including the surface alignment distance (379)) across sensor pairs and scene points. The alignment loss (381) computes an overall measure of geometric consistency in the multi-sensor calibration. The alignment loss (381) may weight individual surface alignment distances based on observation confidence to improve accuracy. The alignment loss (381) contributes to the geometric consistency loss (387) and helps drive optimization of sensor intrinsics, extrinsics, and scene geometry to improve cross-sensoralignment through updates to the calibration data A (305) through N (307) and the feature grid (303).

[0058] The depth loss (385) is a value that measures the discrepancy between rendered depth values and observed depth measurements, which may be from LiDAR sensors. The depth loss (385) computes the difference between simulated depth values generated by sensor models and real-world depth measurements for corresponding points of a point cloud. The depth loss (385) may use a distance metric (z.e., a Manhattan or Euclidean distance) to quantify depth errors. The depth loss (385) contributes to the geometric consistency loss (387) to refine both sensor calibration and the geometric accuracy of the scene representation through updates to the calibration data A (305) through N (307) and the feature grid (303).

[0059] The geometric consistency loss (387) is a value that combines multiple loss terms related to the geometric accuracy and consistency of the multi-sensor calibration and scene representation. The geometric consistency loss (387) may include the depth loss (385) and the alignment loss (381) and may weight the individual contributions based on sensor characteristics or calibration priorities. The geometric consistency loss (387) provides a measure of the match between outputs generated from the current calibration and scene representation to the observed 3D structure across multiple sensors. The geometric consistency loss (387) is a component of the loss value (391), driving joint optimization of sensor intrinsics, extrinsics, and scene geometry through updates to the calibration data A (305) through N (307) and the feature grid (303).

[0060] The regularization loss (389) is a value that may measure the accuracy of surface representations. The regularization loss (389) combines multiple terms to encourage desirable properties in the learned scene representation and calibration parameters. The regularization loss (389) includes components that promote smooth surfaces, enforce physical constraints on the signed distance function, and concentrate sample weight distributions around surfaces. By incorporating these regularization terms, the regularization loss (389) helpsprevent overfitting and improves the generalization of the calibrated model. The regularization loss (389) contributes to the overall loss value (391), guiding the optimization process to learn a more robust and physically plausible representation of the scene and sensor configuration through updates to the calibration data A (305) through N (307) and the feature grid (303).

[0061] The loss value (391) is a value that represents the overall discrepancy between the current calibration, scene representation, and observed real-world data. The loss value (391) may be a weighted combination of multiple loss components, including the photometric consistency loss (377), the geometric consistency loss (387), and the regularization loss (389). Each component of the loss value (391) may correspond to different aspects of the calibration quality and scene reconstruction accuracy. The loss value (391) may be an optimization objective for the calibration system (300). By minimizing the loss value (391) through iterative updates to the learned parameters (301), the calibration system (300) improves the alignment between sensors and refines the scene representation. The loss value (391) is backpropagated through the differentiable models to compute gradients and update the calibration data A (305) through N (307) and the feature grid (303).

[0062] Each of the machine learning models utilized within the calibration system (300) may include one or more machine learning models. The machine learning models may include neural networks and may operate using one or more layers of weights that may be sequentially applied to sets of input data, which may be referred to as input vectors. For each layer of a machine learning model, the weights of the layer may be multiplied by the input vector to generate a collection of products, which may then be summed to generate an output for the layer that may be fed, as input data, to a next layer within the machine learning model. Different architectures may be used. The output of the machine learning model may be the output generated from the last layer within the machine learning model. Multiple machine learning models may operate sequentially or in parallel. The output may be a vector or scalar value. Thelayers within the machine learning model may be different and correspond to different types of models. As an example, the layers may include layers for residual neural networks, feature pyramid neural networks, recurrent neural networks, convolutional neural networks, transformer models, attention layers, perceptron models, etc. Perceptron models may include one or more fully connected (also referred to as linear) layers that may convert between the different dimensions used by the inputs and the outputs of a model.

[0063] The machine learning models may be trained by inputting training data to a machine learning model to generate training outputs that are compared to expected outputs. For supervised training, the expected outputs may be labels associated with a given input. For unsupervised learning, the expected outputs may be previous outputs from the machine learning model. The difference between the training output and the expected output may be processed with a loss function to identify updates to the weights of the layers of the model. After training on a batch of inputs, the updates identified by the loss function may be applied to the machine learning model to generate a trained machine learning model. Different algorithms may be used to calculate and apply the updates to the machine learning model, including back propagation, gradient descent, etc.

[0064] FIG. 4 shows a flowchart of a method of a neural calibration framework. The method of FIG. 4 may be implemented using the components of FIG. 1, FIG. 4, and FIG. 3, and one or more of the steps may be performed on, or received at, one or more computer processors. The system may include at least one processor and an application that, when executing on the at least one processor, performs the method. A non-transitory computer readable medium may include instructions that, when executed by one or more processors, perform the method. The outputs from various components (including models, functions, procedures, programs, processors, etc.) for performing the method may be generated by applying a transformation to inputs using the components to create the outputs without using mental processes or human activities.

[0065] Turning to FIG. 4, the process (400) may be part of the application of a calibration system that updates calibration data for sensors of an autonomous vehicle. The process (400) may include multiple steps (e.g., steps (402) through (408)) that may execute on the components described in the other figures, including those of FIG. 1.

[0066] Step (402) includes generating a simulated sensor rendering using a feature grid and calibration data. The simulated sensor rendering may be generated by neural rendering, which may include the use of implicit neural representations (e.g., a feature grid) instead of structured scene geometric data and may use differentiable rendering. Implicit neural representations may use neural networks to encode scenes so that the scenes may be queried for rendering. For example, multilayer perceptrons (MLPs) may learn to map features from a feature grid to color, density, etc., to represent a scene. Differentiable rendering may utilize gradient-based optimization of the scene parameters through backpropagation to generate calibration data. The feature grid used to generate the simulated sensor rendering may represent a multiresolution encoding of 3D scene geometry and appearance, containing features for each point in the 3D scene. The calibration data may encode intrinsic and extrinsic parameters for sensors within a multi-sensor platform. The parameters may include 6-degree-of-freedom pose information relative to a reference frame. A geometric model may process information from the feature grid and scene points to generate geometric representations of the environment. The geometric model may output a signed-distance function encoding surface distances and a feature descriptor encoding appearance information.

[0067] Generating the simulated sensor rendering may include selecting a set of sampled rays by progressively increasing sampling frequency in regions of interest identified in a heat map by blurring the heat map. A ray sampler may implement a coarse-to-fine strategy to select (z.e., sample) rays for rendering. The sampling strategy may start with uniform sampling across the scene and gradually increase the probability of selecting rays corresponding to 1informative areas. Blurring the heat map may help identify broader regions of interest beyond individual feature points. The heat map may be generated by processing the real-world sensor rendering with an interest point detection model to output an image (the heat map) that identifies points of interest from the real-world sensor rendering. The interest point detection model may use neural network (e.g., a convolutional network) trained to identify points of interest. The values for the pixels in the heat map may identify the probability that the ray corresponding to the pixel will be selected. The heat map may be blurred using a Gaussian blur so that at an initial iteration, each value for the pixels has the same probability. For the last iteration, the pixels of interest identified by the interest pixel detector may have a probability approaching or equal to 1 and the remaining pixels may have a probability approaching or equal to 0. The progressive sampling approach may improve efficiency by reducing the number of rays to render, to reduce the computational resources used, and focus on the points of interest, which are relevant to calibration.

[0068] Generating the simulated sensor rendering may include executing a geometric model to generate a signed-distance function and a feature descriptor from the feature grid and a scene point. The geometric model may be implemented as a neural network that processes interpolations of the feature grid based on input scene points. The signed-distance function may encode the distance from a given scene point to the nearest surface, with positive values indicating points outside of objects and negative values representing points inside of objects. The feature descriptor may encode appearance and geometric information used by sensor models to render simulated observations. The geometric model may interpolate the feature grid based on a scene point and then process the interpolation with a neural network (e.g., a multilayer perceptron) to generate the signed-distance function and the feature descriptor.

[0069] Generating the simulated sensor rendering may include generating a first origin and a first direction using a first extrinsic matrix corresponding to the calibration data for the sensor, which may be referred to as a first sensor. Thefirst origin and the first direction may each be one of multiple (e.g., hundreds of thousands) of ray origins and directions for the first sensor according to calibration and sensor resolution. The first sensor may correspond to a camera sensor. Multiple sensors of multiples may be used including camera sensors and LiDAR sensors. The extrinsic matrix may represent the 6-degree-of- freedom transformation between the sensor coordinate system and a reference frame, which may correspond to the autonomous system. The origin may represent the position of the sensor in the reference frame, while the direction may be computed based on the orientation of the sensor and internal parameters. For a camera sensor, the direction may be derived from the focal length and pixel coordinates of the camera sensor.

[0070] Generating the simulated sensor rendering may include executing a first Tenderer corresponding to a first sensor model to generate a portion of the simulated sensor rendering using the first direction and a feature descriptor. The sensor model may simulate the behavior of a specific sensor type, such as a camera sensor or a LiDAR sensor. For a camera model of a camera sensor, the rendering process of the Tenderer may involve projecting 3D points onto a 2D image plane and determining pixel colors based on the feature descriptors. For a LiDAR model of a LiDAR sensor, the process of the Tenderer may involve simulating ray intersections with implicit surfaces from the feature descriptors to compute distance and intensity values.

[0071] Generating the simulated sensor rendering may include generating a second origin and a second direction using a second extrinsic matrix corresponding to second calibration data for a second sensor. The second origin and the second direction may each be one of multiple (e.g., hundreds of thousands) of ray origins and directions for the second sensor according to calibration and sensor resolution. More than two sensors may be used with multiple origins and directions corresponding to each sensor. The second extrinsic matrix may represent the transformation for a different sensor in the multi-sensor platform, which may be a different type of sensor. The secondorigin and direction may be computed similarly to the first, but using the calibration parameters specific to the second sensor.

[0072] Generating the simulated sensor rendering may include executing a second Tenderer corresponding to a second sensor model to generate a portion of a second simulated sensor rendering using the second direction and a feature descriptor. Multiple Tenderers and sensor models may be used, which may correspond to different types of sensors and support an arbitrary number of sensors that are jointly optimized regardless of the number of sensors. The second sensor model may correspond to a different sensor type or configuration than the first. For example, the first model may simulate a camera sensor while the second simulates a LiDAR sensor. The rendering process allows for different characteristics for each sensor type, producing appropriate simulated data such as color images or point clouds with intensity values.

[0073] Step (405) includes generating a loss value between the simulated sensor rendering and a real -world sensor rendering. The loss value quantifies the discrepancies between the current calibration, the scene representation, and the observed real-world data. The loss function computes multiple error metrics to assess the alignment between sensors and the accuracy of the scene representation. The loss value may be calculated as a weighted combination of photometric consistency loss, geometric consistency loss, and regularization loss. Each component of the loss corresponds to different aspects of calibration quality and scene reconstruction accuracy. The loss value is an optimization objective for the calibration system, guiding iterative updates to learned parameters.

[0074] Generating the loss value may include generating a surface alignment distance using the simulated sensor rendering and a second simulated sensor rendering. The surface alignment distance quantifies the geometric consistency between different sensor observations of the same scene point. The distance may be computed by comparing ray-casted points on the implicit surface representation for corresponding pixels or rays across multiple sensors. Tonormalize the measurement, 3D points may be projected onto image planes to obtain pixel-space distances. The surface alignment distance provides a measure of the alignment between sensors based on the current calibration data and scene representation.

[0075] Generating the loss value may include combining the surface alignment distance into a geometric consistency loss to generate the loss value. The geometric consistency loss aggregates multiple loss terms related to the geometric accuracy and consistency of the multi-sensor calibration and scene representation. In addition to the surface alignment distance, the geometric consistency loss may incorporate depth loss and alignment loss. The individual contributions may be weighted based on sensor characteristics or calibration priorities. The geometric consistency loss may measure the match between the current calibration and scene representation with the observed 3D structure across multiple sensors.

[0076] Generating the loss value may include generating the loss value as a weighted combination of a photometric consistency loss, a geometric consistency loss, and a regularization loss. The photometric consistency loss assesses the alignment of appearance information across sensors. The geometric consistency loss evaluates the spatial alignment and structural accuracy. The regularization loss may measure the accuracy of surface representations. Weights for each component may be adjusted to improve the accuracy of the calibration. The combined loss value guides the optimization process to improve both sensor alignment and scene reconstruction simultaneously.

[0077] Generating the loss value may include generating the photometric consistency loss of the loss value as a combination of a first sensor rendering loss and a second sensor rendering loss. The first sensor rendering loss may correspond to color image sensors (e.g., camera sensors), while the second may relate to intensity measurements from other sensor types (e.g., LiDAR sensors). By combining losses from multiple sensor types, the photometric consistency loss measures the consistency of appearance modeling across the multiplesensors of an autonomous system. The relative weighting of the contribution of each sensor may be adjusted based on factors such as sensor resolution, field of view, expected reliability, etc.

[0078] Generating the loss value may include generating a color rendering loss as the first sensor rendering loss over multiple camera sensors. The color rendering loss may quantify the difference between simulated and real-world color images from multiple camera sensors. For each camera in the multi-sensor platform, pixel values may be compared between the simulated sensor rendering and the real-world sensor rendering. The loss may be computed using per-pixel distance metrics in RGB space or more sophisticated perceptual metrics that account for human visual sensitivity. The color rendering loss may be aggregated across multiple cameras for consistent calibration and appearance modeling for an array of camera sensors.

[0079] Generating the loss value may include generating an intensity rendering loss as the second sensor rendering loss over multiple light detection and ranging (LiDAR) sensors. The intensity rendering loss measures the discrepancy between simulated and real -world intensity values from LiDAR sensors. For each LiDAR sensor in the sensor suite of an autonomous vehicle, intensity values may be compared between the simulated sensor rendering and the real-world sensor rendering. The loss may be computed using distance metrics between corresponding intensity values. The intensity rendering loss incorporates intensity information from multiple LiDAR sensors to refine the calibration of the LiDAR array and improve the modeling of surface properties in the scene representation.

[0080] Generating the loss value may include generating a geometric consistency loss of the loss value as a combination of a depth loss and an alignment loss. The geometric consistency loss may be computed by summing the depth loss and the alignment loss with appropriate weighting factors. The depth loss may measure discrepancies between rendered depth values and observed depth measurements from LiDAR sensors. The alignment loss may quantify geometricinconsistencies between different sensor observations of the same scene points. Combining both loss terms may provide an assessment of the geometric accuracy and cross-sensor consistency of the calibration and scene representation.

[0081] Generating the loss value may include generating the depth loss over multiple light detection and ranging (LiDAR) sensors. The depth loss may be calculated for each LiDAR sensor in the multi-sensor platform. For each LiDAR sensor, the difference between simulated depth values from the sensor model and real-world depth measurements may be computed for corresponding points in the point cloud. A distance metric may be used to quantify the depth errors. The individual depth losses from each LiDAR sensor may be aggregated, potentially with sensor-specific weighting factors, to obtain the overall depth loss.

[0082] Generating the loss value may include generating the alignment loss using multiple surface alignment distances between the simulated sensor rendering and a second simulated sensor rendering. Surface alignment distances may be computed for pairs of corresponding pixels or rays across different sensors. The distances may measure the discrepancy between ray-casted points on the implicit surface representation. To normalize the measurements, 3D points may be projected onto image planes to obtain pixel-space distances. The alignment loss may be calculated by aggregating multiple surface alignment distances across various sensor pairs and scene points. Weighting factors may be applied to individual distances based on observation confidence or other criteria.

[0083] Generating the loss value may include generating a regularization loss of the loss value by combining a sample weight distribution around a surface with a signed-distance function. A term of the regularization loss may relate to concentration of the learned sample weight distribution around surfaces. Another term of the regularization loss may enforce constraints on the signed distance function to encourage smooth surfaces or satisfying the Eikonal equation. The sample weight distribution term and the signed-distance function term may be combined with appropriate scaling factors to form the regularization loss.

[0084] Step (410) includes calibrating the sensor with the updated calibration data. The calibration process may involve applying the optimized intrinsic and extrinsic parameters that are used to generate sensor observations in a common coordinate system. The optimized intrinsic and extrinsic parameters may be applied to a sensor by storing the optimized intrinsic and extrinsic parameters in the calibration data for the sensor. The updated calibration data may be used to refine the relative poses between sensors in the multi-sensor platform. Sensor fusion algorithms may utilize the calibrated intrinsic and extrinsic parameters to combine data from multiple sensors. The calibration may improve the alignment of sensor data for downstream perception and planning tasks in the autonomous system. The updated calibration data may be stored in a sensor configuration file or database for future use. Verification steps may be performed to ensure the calibrated sensor data meets accuracy thresholds. The calibration may be applied in real-time to incoming sensor data streams or used to post-process previously collected data.

[0085] Calibrating the sensor may include storing an extrinsic matrix as the calibration data to a sensor profile for the sensor. The extrinsic matrix may encode a 6-degree-of-freedom transformation between the sensor coordinate system and a reference frame for the autonomous system. The sensor profile may contain additional metadata about the sensor, such as intrinsic parameters, model number, installation location, etc. The extrinsic matrix may be stored in a standardized format to facilitate interoperability with different software systems. Version control may be implemented to track changes to the sensor profile over time. Access controls may be applied to restrict modification of the calibration data to authorized personnel or processes. Backup copies of the sensor profile may be created to prevent data loss.

[0086] Calibrating the sensor may include generating calibrated data from the sensor using the extrinsic matrix. The calibrated data may be produced by applying the extrinsic transformation to raw sensor measurements. For camera sensors, pixel coordinates may be projected into 3D space using the calibrationparameters. LiDAR point clouds may be transformed from the sensor frame to the vehicle reference frame. The calibrated data may be time-synchronized with data from other sensors in the multi-sensor platform of the autonomous system. Coordinate system conversions may be applied to express the calibrated data in a desired reference frame, such as a global coordinate system. Quality checks may be performed on the calibrated data to detect potential calibration errors or sensor malfunctions. The calibrated data may be formatted and packaged for consumption by perception, localization, or mapping modules in the autonomous system.

[0087] Turning to FIG. 5, the interface (500) is an example of a display of several views related to sensor data, which may correspond to a data workflow for a neural calibration framework. The interface (500) includes the sensor setup view (510), the collect data view (520), the uncalibrated view (530), and the calibrated view (570).

[0088] The sensor setup view (510) displays information about the sensors of the autonomous system. The sensors include multiple camera sensors and light detection and ranging (LiDAR) sensors. The camera sensors include the camera sensor (511) facing left front, the camera sensor (512) facing to the front of the autonomous system, the camera sensor (513) facing right front, the camera sensor (514) facing to the left rear of the autonomous system, the camera sensor (515) facing to the rear of the autonomous system, and the camera sensor (516) facing to the right rear of the autonomous system. The LiDAR sensors include the LiDAR sensor (518) facing to the front of the autonomous system and the LiDAR sensor (519), which is omnidirectional.

[0089] The collect data view (520) identifies when the system may be collecting data for calibration. The collect data view (520) may display an animation (e.g., moving the lines on the road) during the process of collecting data.

[0090] The uncalibrated view (530) displays real-world sensor renderings (531), (532), (533), (534), (535), (536), (538), and (539) prior to process of calibration.The real-world sensor renderings (531) through (536) are images that correspond respectively to the camera sensors (511) through (516). The real-world sensor renderings (538) and (539) are point clouds that correspond respectively to the LiDAR sensors (518) and (519). The bounding boxes (551), (552), (553), (554), (555), (556), and (558) identify discrepancies due to miscalibration of the sensors. The bounding boxes (551) through (556) identify LiDAR camera alignment issues and the bounding boxes (538) identify LiDAR-LiDAR alignment issues.

[0091] The calibrated view (570) displays real-world sensor renderings (571), (572), (573), (574), (575), (576), (578), and (579) generated after the process of calibration. The real-world sensor renderings (571) through (576) are images that correspond respectively to the camera sensors (511) through (516). The real- world sensor renderings (578) and (579) are point clouds that correspond respectively to the LiDAR sensors (518) and (519). After calibration, the sensors are aligned so the alignment discrepancies are reduced as compared to the discrepancies in the uncalibrated view (530).

[0092] To perform the calibration, the method takes collected data from multisensor robots (such as the autonomous system displayed in the sensor setup view (510)) and automatically calibrates sensor intrinsics and extrinsics. The uncalibrated view (530) shows LiDAR-Camera and LiDAR-LiDAR alignment on collected data with uncalibrated intrinsics and extrinsics. The calibrated view (570) shows sensor alignment with optimized calibration after the performance of calibration.

[0093] Turning to FIG. 6, the neural calibration framework (600) calibrates the sensors (613), (615), (617), and (619) of an autonomous system (611). The sensor (613) is a LiDAR sensor and the sensors (615) through (619) are camera sensors. The calibration data is stored in a sensor calibration graph with nodes that identify the different sensors and edges that may quantify the poses of the sensors (613), (615), (617), and (619). The edges may store the calibration data as an extrinsic matrix for each of the sensors (613), (615), (617), and (619).

[0094] The implicit scene representation (631) is a feature grid that represents the three dimensional scene observed by the autonomous system (611). The implicit scene representations (631) may be initialized as random data and then be updated over multiple iterations.

[0095] A LiDAR model generates a simulated LiDAR rendering (653) for the LiDAR sensor (613) by determining intensity and distance information for a set of selected rays cast into the implicit scene of representation (631). The simulated LiDAR rendering may be stored as a point cloud and approximates the real-world LiDAR rendering (673).

[0096] A camera model generates a simulated camera rendering (657) for the camera sensor (617) by determining color information (e.g., RGB values) for a set of selected rays cast into the implicit scene of representation (631). The simulated camera rendering (657) may be stored as an image and approximates the real-world camera rendering (677).

[0097] The real -world camera renderings (675), (677), and (679) and the real- world LiDAR rendering (673) are compared to the simulated renderings (e.g., the simulated LiDAR rendering (653) and the simulated camera rendering (657)) to calculate multiple loss metrics that are combined into a loss value. The loss value is an objective to be minimized over multiple iterations. The different loss values for the different iterations may be displayed in the graph (681), which shows the reduction of the loss value over multiple iterations. For each iteration, the loss value is used to update the implicit scene representation (631) and the calibration data for the sensors (613), (615), (617), and (619), which are stored in the sensor calibration graph.

[0098] The approach depicted in FIG. 6 jointly optimizes multi-sensor intrinsics and extrinsics and underlying scene representation. The intrinsics and extrinsics and underlying scene representation are optimized within a differentiable framework to minimize photometric and geometric consistency energy terms on collected outdoor data retrospectively.

[0099] The neural calibration framework (600) calibrates the intrinsics and extrinsics of a multi-modal sensor platform by collecting a short trajectory in a general scene, without using calibration targets. The approach builds on a scene representation capable of rendering multi-view geometrically and photometrically consistent sensor observations. Through joint optimization of sensor parameters (e.g., the intrinsics and extrinsics, which may include the extrinsic matrices for the sensors (613) through (619)) and the underlying scene representation (e.g., the implicit scene representation (631)) within a differentiable framework, relationships between sensors are effectively resolved. A differentiable surface alignment loss is introduced to increase geometric consistency across observations and mitigate shape-radiance ambiguity. The method may optimize sensor intrinsics and extrinsics and may use a known 6-DoF vehicle trajectory.

[0100] Implicit neural scene representation may be performed with the neural calibration framework (600). Representing scene geometry and appearance using implicit representations may generate photorealistic results for view synthesis. The neural calibration framework (600) leverages scene representation to calibrate multiple cameras and LiDARs (e.g., the sensors (613) through (619)) mounted on the autonomous system (611). The implicit scene representation (631) may be parameterized using a multi-resolution feature grid with multilayer perceptron (MLP) networks. Given a 3D scene point x G IRx3, the 3D feature grid £Gi=1at each level may be tri-linearly interpolated. The resulting interpolated features may then be concatenated and processed with a neural network (e.g., an MLP) to yield the geometry represented as a signed- distance function s and appearance feature descriptor f. The process may be characterized by a querying function Q'.The feature grid may be optimized using a fixed number of features with a grid index hash function.

[0101] Differentiable sensor models may be used by the neural calibration framework (600). The sensor models may be used for camera and LiDAR sensor calibration. A 6-DoF trajectory for an autonomous system may expressed in an arbitrary world frame, denoted as Pveh(0- Then the state of each sensor at timestamp t is described as:Psensor( =Pveh( EsensorEq.2where E^nsor sensor represents the extrinsic matrix of the i-th sensor, indicating relative pose with respect to the frame of the autonomous system (611).

[0102] A camera model may be used by the neural calibration framework (600) to generate a simulated rendering for a camera sensor. Each pixel u = (u, v)TG IE2is represented in homogeneous coordinates as u = (u, v, l)TG H . Given the intrinsic matrix Kcamand extrinsic matrix Ecam, the viewing ray in 3D world coordinates is expressed as r( / r£) = o +£d, where the ray origin o and ray direction d at timestamp t are given by:

[0103] To render the image from the scene representation, the pixel color Icam(r) is approximated as the weighted sum of colors at sampled points along the camera ray:

[0104] where atG [0, 1] represents opacity, derived from the SDF s£. The feature descriptor f£and SDF stare obtained from the querying function Q(o + h£d) (Eqn. 1), where h£is the i-th sample point along the ray. Dcani(-) serves as the camera decoder, mapping the feature descriptor and view direction to a red, green, blue (RGB) color.

[0105] A LiDAR model may be used by the neural calibration framework (600). The LiDAR sensor (613) emits laser beam pulses and determines the distance from the LiDAR sensor (613) to the reflective surface by measuring the time of flight. With the laser elevation and azimuth angles denoted as w = (0, y) for each emitted ray, the ray direction in sensor coordinates is characterized as (cos 6 cos y, cos 0 sin y, sin 0)T. To model ego-motion during the LiDAR scan, the viewing ray origin and direction in 3D world coordinates at ray firing timestamp t are expressed as:

[0106] Similar to Eq. 4, volume rendering is applied to generate the depth and intensity:where ^LIDARC') denotes the LiDAR decoder, which maps the queried feature descriptor and view direction to the LiDAR intensity value.

[0107] A surface alignment constraint may also be used by the neural calibration framework (600). Recovering both the scene representation and sensor poses presents challenges for unstructured outdoor driving scenes. An unregularized model may learn to render the target observations with incorrect poses and geometry. To address the challenge, the 3D structures inferred from the sensor data may be aligned with the underlying implicit scene surface. For LiDAR data, geometric alignment may be assessed by comparing the rendered depth with the observed depth value. Establishing geometric alignment for camera data is nontrivial due to the absence of direct depth information. A differentiable surface alignment distance provides additional geometric constraints on the camera poses. Sparse correspondences between pairs of camera images may be inferred using multi -view geometry tools. A pair of corresponding pixels between imagepair may be denoted as (u1;u2). The associated camera rays may be computed from Eq. 3 as r^h) = ox+ hd and r2( / r) = o2+ hd2. When calibrated, the rays rxand r2intersect at the scene surface. Inaccuracies in sensor calibration may introduce errors that are measured by computing the distance between the ray-casted points on the scene surface. The continuous nature of the implicit scene representation allows querying depth values for arbitrary rays within the scene using the depth function £>(■) defined in Eq. (6). The ray-casted points are then expressed as:

[0108] The distance of the ray-casted points is normalized by projecting the ray- casted points onto the image plane, and the surface alignment distance is defined as:where TT is the projection operator defined by camera intrinsic and extrinsic parameters.

[0109] Sensor calibration may be learned by the neural framework (600). The optimization jointly minimizes photometric and geometric losses for the scene representation Q and sensor extrinsics E^sor using camera images and LiDAR sweeps. Regularization of the underlying scene geometry ensures smooth surfaces and physical constraints. The learning objective is:T 'Iphotom'fphotom T ' geom'f geom T AegAeg Eq. 9

[0110] The photometric consistency loss may be defined as Tphotom= £rgb+ 2£intand includes camera RGB rendering loss (£rgb) and LiDAR intensity rendering loss (£int). The camera rendering loss calculates an f2loss between observed and rendered images using Eq. 10:where Ncamis the number of camera sensors and E‘ is the i-th camera sensor extrinsic, u] is the sampled pixel from the images captured by the 2-th camera sensor. Icamis the rendered image from Eq. (4) and Icamis the corresponding observed camera image. The LiDAR intensity rendering loss is:where EL;DARis the i-th LiDAR sensor extrinsic and NLis the number of LiDAR sensors, w is the sampled laser beam. IJ DAR (Eq. (6)) and ILIDARarethe rendered and observed (or real-world) LiDAR intensities.

[0111] The geometric consistency loss may be defined as £geom= Tdepth+ / ?£alignand includes a depth rendering loss (Tdepth) and a surface alignment loss (Taiign). The depth rendering loss isterm between rendered and observed LiDAR depth calculated with Eq. 12:where £>(■) is the depth rendering function defined in Eq. 6. For surface alignment loss, an image matching algorithm identifies corresponding pixels between image pairs. Denoting the set of correspondences between camera i and camera j as {uk}k=1and {uk}k=1, where M is the number of correspondences, the surface alignment loss is defined as:where Aurf(') is the surface alignment distance defined in Eq. 8. As image matching and ray-casting results may contain noise, correspondences are filtered out with a re-projection error ||TT(P) — u||2> 8n, or a ray termination probability

[0112] A regularization term may also be used by the neural framework (600) to provide a constraint on the learned surface representation. With the left term of Eq. 14, the learned sample weight distributionis concentrated around a surface. With the right term of Eq. 14, the underlying SDF value is encouraged to satisfy the Eikonal equation. The resulting regularization term (z.e., the regularization loss (£reg)), helps the network learn a smooth zero level set. The regularization loss may be defined as:

[0113] where T£= \ht— D\ represents the distance between the sample pointand its corresponding LiDAR depth observation. The two terms of the regularization loss contribute to learning an accurate surface representation, which yields increased precision of depth values for optimizing the surface alignment loss (Talignof Eq. (13)).

[0114] Ray sampling priors may also be utilized by neural framework (600). Selecting sensor rays to render and supervise affects learning for sensor calibration. Structure-from-motion pipelines identify interest points for establishing correspondences for alignment. Neural rendering may leverage uniform ray sampling for scene reconstruction. A coarse-to-fine sampling strategy may be implemented during training to improve accuracy and reduce the amount of computational resources used for rendering. Initially, sensor rays are uniformly sampled to learn an accurate scene representation. The sampling frequency may be progressively increased in regions of interest to enhance pose registration. The adjustment is made because different sensor rays contribute by different amounts to pose learning. Textureless regions, like the sky and road, may offer insufficient gradients to effectively update sensor pose displacements. Interest points may be detected using an algorithm to generate a corresponding heat map h. Denoting a E [0, 1] as the controllable parameter proportional to theprogress in coarse-to-fine sampling stage, the sampling probability map p(a) is proportional to the Gaussian-blurred version of the heat map: p(a) ocmin+ a ■ GaussBlur( , ka) Eq. 15 where / rmincontrols the minimum score for sampling any rays, ka= akmill+ (1 — a)kmaxis the Gaussian blur kernel.

[0115] Embodiments may be implemented on a computing system specifically designed to achieve an improved technological result. When implemented in a computing system, the features and elements of the disclosure provide a significant technological advancement over computing systems that do not implement the features and elements of the disclosure. Any combination of mobile, desktop, server, router, switch, embedded device, or other types of hardware may be improved by including the features and elements described in the disclosure. For example, as shown in FIG. 7A, the computing system (700) may include one or more computer processors (702), non-persistent storage (704), persistent storage (706), a communication interface (712) (e.g., Bluetooth interface, infrared interface, network interface, optical interface, etc.), and numerous other elements and functionalities that implement the features and elements of the disclosure. The computer processor(s) (702) may be an integrated circuit for processing instructions. The computer processor(s) may be one or more cores or micro-cores of a processor. The computer processor(s) (702) includes one or more processors. The one or more processors may include a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing units (TPU), combinations thereof, etc.

[0116] The input devices (710) may include a touchscreen, keyboard, mouse, microphone, touchpad, electronic pen, or any other type of input device. The input devices (710) may receive inputs from a user that are responsive to data and messages presented by the output devices (708). The inputs may include text input, audio input, video input, etc., which may be processed and transmitted by the computing system (700) in accordance with the disclosure. Thecommunication interface (712) may include an integrated circuit for connecting the computing system (700) to a network (not shown) (e.g., a local area network (LAN), a wide area network (WAN) such as the Internet, mobile network, or any other type of network) and / or to another device, such as another computing device.

[0117] Further, the output devices (708) may include a display device, a printer, external storage, or any other output device. One or more of the output devices may be the same or different from the input device(s). The input and output device(s) may be locally or remotely connected to the computer processor(s) (702). Many different types of computing systems exist, and the aforementioned input and output device(s) may take other forms. The output devices (708) may display data and messages that are transmitted and received by the computing system (700). The data and messages may include text, audio, video, etc., and include the data and messages described above in the other figures of the disclosure.

[0118] Software instructions in the form of computer readable program code to perform embodiments may be stored, in whole or in part, temporarily or permanently, on a non-transitory computer readable medium such as a CD, DVD, storage device, a diskette, a tape, flash memory, physical memory, or any other computer readable storage medium. Specifically, the software instructions may correspond to computer readable program code that, when executed by a processor(s), is configured to perform one or more embodiments, which may include transmitting, receiving, presenting, and displaying data and messages described in the other figures of the disclosure.

[0119] The computing system (700) in FIG. 7 A may be connected to or be a part of a network. For example, as shown in FIG. 7B, the network (720) may include multiple nodes (e.g., node X (722), node Y (724)). Each node may correspond to a computing system, such as the computing system shown in FIG. 7A, or a group of nodes combined may correspond to the computing system shown in FIG. 7A. By way of an example, embodiments may be implemented on a nodeof a distributed system that is connected to other nodes. By way of another example, embodiments may be implemented on a distributed computing system having multiple nodes, where each portion may be located on a different node within the distributed computing system. Further, one or more elements of the aforementioned computing system (700) may be located at a remote location and connected to the other elements over a network.

[0120] The nodes (e.g., node X (722), node Y (724)) in the network (720) may be configured to provide services for a client device (726), including receiving requests and transmitting responses to the client device (726). For example, the nodes may be part of a cloud computing system. The client device (726) may be a computing system, such as the computing system shown in FIG. 7A. Further, the client device (726) may include and / or perform all or a portion of one or more embodiments.

[0121] The computing system of FIG. 7A may include functionality to present raw and / or processed data, such as results of comparisons and other processing. For example, presenting data may be accomplished through various presenting methods. Specifically, data may be presented by being displayed in a user interface, transmitted to a different computing system, and stored. The user interface may include a GUI that displays information on a display device. The GUI may include various GUI widgets that organize what data is shown as well as how data is presented to a user. Furthermore, the GUI may present data directly to the user, e.g., data presented as actual data values through text, or rendered by the computing device into a visual representation of the data, such as through visualizing a data model.

[0122] As used herein, the term “connected to” contemplates multiple meanings. A connection may be direct or indirect (e.g., through another component or network). A connection may be wired or wireless. A connection may be temporary, permanent, or semi-permanent communication channel between two entities.

[0123] The various descriptions of the figures may be combined and may include or be included within the features described in the other figures of the application. The various elements, systems, components, and steps shown in the figures may be omitted, repeated, combined, and / or altered as shown from the figures. Accordingly, the scope of the present disclosure should not be considered limited to the specific arrangements shown in the figures.

[0124] In the application, ordinal numbers (e.g., first, second, third, etc.) may be used as an adjective for an element (z.e., any noun in the application). The use of ordinal numbers is not to imply or create any particular ordering of the elements nor to limit any element to being only a single element unless expressly disclosed, such as by the use of the terms “before”, “after”, “single”, and other such terminology. Rather, the use of ordinal numbers is to distinguish between the elements. By way of an example, a first element is distinct from a second element, and the first element may encompass more than one element and succeed (or precede) the second element in an ordering of elements.

[0125] Further, unless expressly stated otherwise, or is an “inclusive or” and, as such includes “and.” Further, items joined by an or may include any combination of the items with any number of each item unless expressly stated otherwise.

[0126] In the above description, numerous specific details are set forth in order to provide a more thorough understanding of the disclosure. However, it will be apparent to one of ordinary skill in the art that the technology may be practiced without these specific details. In other instances, well-known features have not been described in detail to avoid unnecessarily complicating the description. Further, other embodiments not explicitly described above can be devised which do not depart from the scope of the claims as disclosed herein. Accordingly, the scope should be limited only by the attached claims.

Claims

CLAIMSWhat is claimed is:

1. A method for calibrating a sensor, the method comprising: generating a simulated sensor rendering using a feature grid and calibration data; generating a loss value between the simulated sensor rendering and a real-world sensor rendering; updating the feature grid and the calibration data using the loss value to generate an updated feature grid and updated calibration data to reduce a subsequent loss value; and calibrating the sensor with the updated calibration data.

2. The method of claim 1, wherein generating the loss value comprises: generating a surface alignment distance using the simulated sensor rendering and a second simulated sensor rendering; and combining the surface alignment distance into a geometric consistency loss to generate the loss value.

3. The method of claim 1, wherein generating the simulated sensor rendering comprises: selecting a set of sampled rays by progressively increasing sampling frequency in regions of interest identified in a heat map by blurring the heat map.

4. The method of claim 1, wherein generating the simulated sensor rendering comprises: executing a geometric model to generate a signed-distance function and a feature descriptor from the feature grid and a scene point.

5. The method of claim 1, wherein generating the simulated sensor rendering comprises: generating a first origin and a first direction using a first extrinsic matrix corresponding to the calibration data for the sensor; andexecuting a first Tenderer corresponding to a first sensor model to generate a portion of the simulated sensor rendering using the first direction and a feature descriptor.

6. The method of claim 1, wherein generating the simulated sensor rendering comprises: generating a second origin and a second direction using a second extrinsic matrix corresponding to second calibration data for a second sensor; and executing a second Tenderer corresponding to a second sensor model to generate a portion of a second simulated sensor rendering using the second direction and a feature descriptor.

7. The method of claim 1, wherein generating the loss value comprises: generating the loss value as a weighted combination of a photometric consistency loss, a geometric consistency loss, and a regularization loss; generating the photometric consistency loss of the loss value as a combination of a first sensor rendering loss and a second sensor rendering loss; generating a color rendering loss as the first sensor rendering loss over multiple camera sensors; and generating an intensity rendering loss as the second sensor rendering loss over multiple light detection and ranging (LiDAR) sensors.

8. The method of claim 1, wherein generating the loss value comprises: generating a geometric consistency loss of the loss value as a combination of a depth loss and an alignment loss; generating the depth loss over multiple light detection and ranging (LiDAR) sensors; and generating the alignment loss using multiple surface alignment distances between the simulated sensor rendering and a second simulated sensor rendering.

9. The method of claim 1, wherein generating the loss value comprises: generating a regularization loss of the loss value by combining a sample weight distribution around a surface with a signed-distance function.

10. The method of claim 1, wherein calibrating the sensor comprises: storing an extrinsic matrix as the calibration data to a sensor profile for the sensor; and generating calibrated data from the sensor using the extrinsic matrix.

11. A system comprising: at least one processor; and an application that, when executing on the at least one processor, performs operations comprising: generating a simulated sensor rendering using a feature grid and calibration data, generating a loss value between the simulated sensor rendering and a real- world sensor rendering, updating the feature grid and the calibration data using the loss value to generate an updated feature grid and updated calibration data to reduce a subsequent loss value, and calibrating the sensor with the updated calibration data.

12. The system of claim 11, wherein generating the loss value comprises: generating a surface alignment distance using the simulated sensor rendering and a second simulated sensor rendering; and combining the surface alignment distance into a geometric consistency loss to generate the loss value.

13. The system of claim 11, wherein generating the simulated sensor rendering comprises: selecting a set of sampled rays by progressively increasing sampling frequency in regions of interest identified in a heat map by blurring the heat map.

14. The system of claim 11, wherein generating the simulated sensor rendering comprises: executing a geometric model to generate a signed-distance function and a feature descriptor from the feature grid and a scene point.

15. The system of claim 11, wherein generating the simulated sensor rendering comprises: generating a first origin and a first direction using a first extrinsic matrix corresponding to the calibration data for the sensor; and executing a first Tenderer corresponding to a first sensor model to generate a portion of the simulated sensor rendering using the first direction and a feature descriptor.

16. The system of claim 11, wherein generating the simulated sensor rendering comprises: generating a second origin and a second direction using a second extrinsic matrix corresponding to second calibration data for a second sensor; and executing a second Tenderer corresponding to a second sensor model to generate a portion of a second simulated sensor rendering using the second direction and a feature descriptor.

17. The system of claim 11, wherein generating the loss value comprises: generating the loss value as a weighted combination of a photometric consistency loss, a geometric consistency loss, and a regularization loss; generating the photometric consistency loss of the loss value as a combination of a first sensor rendering loss and a second sensor rendering loss; generating a color rendering loss as the first sensor rendering loss over multiple camera sensors; and generating an intensity rendering loss as the second sensor rendering loss over multiple light detection and ranging (LiDAR) sensors.

18. The system of claim 11, wherein generating the loss value comprises: generating a geometric consistency loss of the loss value as a combination of a depth loss and an alignment loss; generating the depth loss over multiple light detection and ranging (LiDAR) sensors; and generating the alignment loss using multiple surface alignment distances between the simulated sensor rendering and a second simulated sensor rendering.

19. The system of claim 11, wherein generating the loss value comprises: generating a regularization loss of the loss value by combining a sample weight distribution around a surface with a signed-distance function.

20. A non-transitory computer readable medium comprising instructions executable by at least one processor to perform: generating a simulated sensor rendering using a feature grid and calibration data; generating a loss value between the simulated sensor rendering and a real-world sensor rendering; updating the feature grid and the calibration data using the loss value to generate an updated feature grid and updated calibration data to reduce a subsequent loss value; and calibrating the sensor with the updated calibration data.

Citation Information

Patent Citations

  • 3D vision system with automatically calibrated stereo vision sensors and LiDAR sensor

    US11782145B1

  • Determining extrinsic calibration parameters for a sensor

    US20140240690A1

  • Method and system for generating multidimensional maps of a scene using a plurality of sensors of various types

    US20180232947A1

  • Procedural world generation using tertiary data

    US20200050716A1