Calibration of sensor device

Through neural network training methods, NeRF technology is used to calibrate the space and time of sensor equipment, solving the problem of time-consuming and laborious calibration of sensor equipment, and achieving efficient and reliable multimodal sensor system calibration, suitable for real-time calibration and fusion perception on vehicles.

CN120418682APending Publication Date: 2025-08-01YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380089012.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-02-27
Publication Date
2025-08-01

Smart Images

  • Figure CN120418682A_ABST
    Figure CN120418682A_ABST
Patent Text Reader

Abstract

The invention relates to a method of calibrating second sensor devices with respect to a first sensor device, comprising: the first sensor device acquiring first sensor data in a first coordinate system, and each of the second sensor devices acquiring second sensor data in a respective second coordinate system; mapping the first sensor data and the second sensor data into a common coordinate system to obtain mapped first sensor data and second sensor data; training the neural network according to the mapped first and second sensor data by jointly optimizing weights of the neural network for outputting a neural network representation of the environment captured by the first and second sensor devices and a relative transformation of the first coordinate system to the second coordinate system; and spatially calibrating each of the second sensor devices relative to the first sensor device according to the trained neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the calibration of sensor devices, and more particularly to the calibration of LIDAR camera systems, and more specifically to the calibration of LIDAR camera systems of at least partially autonomous vehicles. Background Art

[0002] Sensor systems including multiple sensor devices face the problem of spatial or spatio-temporal calibration of the individual sensor devices relative to each other. Each sensor device acquires sensor data in its local coordinate system according to its respective pose. In order to comprehensively utilize all the acquired sensor data, it is necessary to know the spatial orientation of the sensor devices relative to a common coordinate system, as well as the time offset of the time frame of the sensor data acquired by one sensor device relative to the time frame of the sensor data acquired by another sensor device.

[0003] For example, a Light Detection and Ranging (LIDAR) camera sensing system includes one or more LIDAR devices for acquiring a time series of 3D point cloud data sets of sensed objects and one or more camera devices for capturing a time series of 2D images of these objects, and such systems are applied to various applications. Multimodal sensor calibration is highly demanding because it requires processing different types of data formats (2D images versus 3D point clouds). Specifically, vehicles such as cars, Automated Guided Vehicles (AGVs), and autonomous mobile robots can be equipped with such LIDAR camera sensing systems for navigation, positioning, and obstacle avoidance. In the automotive context, a LIDAR camera sensing system can be composed of an Advanced Driver Assistant System (ADAS).

[0004] Both LIDAR devices and camera devices report information about their local coordinate systems. In order to ensure the normal operation of the LIDAR camera system, precise spatial calibration of one or more LIDAR devices and one or more camera devices relative to each other is required, that is, it is necessary to precisely determine the rotation (tensor) R and translation (tensor) T representing the spatial relationship between the LIDAR device and the camera device. In addition, if the individual sensor devices are not synchronized with each other, for example, through some hardware synchronization, time calibration is also required.

[0005] Spatial calibration is the process of finding the spatial relationship (i.e., transformation matrix) between different sensor devices in a system (such as cameras, radars, LIDARs, inertial measurement units, etc.). Time calibration is the process of finding the time offset between data frames from different sensors.

[0006] These calibration processes pose serious problems. The traditional solution is to conduct experiments after installing the LIDAR camera system. These experiments are based on perceiving specific targets (checkerboards) visible to both sensor devices. Features such as corners and edges can be extracted from the point clouds and images of the known targets (checkerboards), and these features can be used in an optimization process that is used to find the spatial calibration between the two sensor devices, thereby enabling feature matching. However, such experiments are both laborious and time-consuming and require careful operation by professionals. Summary of the Invention

[0007] In view of the above, the basic object of the present application is to provide a technique for accurately spatially and / or spatio-temporally calibrating sensor devices in a sensor system with low cost, high reliability, and high scalability, especially spatial and / or spatio-temporal LIDAR camera calibration.

[0008] The above object and other objects are achieved by the subject matter of the independent claims. Other implementations are apparent from the dependent claims, the description, and the drawings.

[0009] According to a first aspect, there is provided a method for calibrating a set of second sensor devices relative to a first sensor device, wherein the first sensor device is different from the second sensor devices. The method includes the steps of: the first sensor device acquiring first sensor data (first observations) in a first coordinate system, and each of the second sensor devices in the set of second sensor devices acquiring second sensor data (second observations) in a corresponding second coordinate system, wherein the second coordinate systems are different from each other and different from the first coordinate system; mapping the first sensor data and the second sensor data acquired by each of the second sensor devices (from the corresponding local second coordinate system) into a common coordinate system (different from the first and second coordinate systems) to obtain the mapped first and second sensor data. The method further includes: training the neural network based on the mapped first and second sensor data by jointly optimizing (a) the weights of a neural network for outputting a neural network representation of the environment captured by the first and second sensor devices and (b) the relative transformation from the first coordinate system to the second coordinate system (represented in the first coordinate system); spatially calibrating each of the second sensor devices relative to the first sensor device according to the trained neural network (i.e., the obtained optimized relative transformation).

[0010] It goes without saying that the term "neural network" in this article refers to an artificial neural network. The first sensor device can be a camera device or a LIDAR device, and the second sensor device can include (a) at least one camera device, (b) at least one LIDAR device, and (c) at least one of at least one camera device and at least one LIDAR device. The camera device can be a time-of-flight camera, a depth camera, etc., and the LIDAR device can be a Micro-Electro-Mechanical System (MEMS), a LIDAR device, a solid-state LIDAR device, etc.

[0011] Specifically, both the first and second sensor devices can move during the calibration process.

[0012] As is well known, neural networks are trained by optimizing their weights. Optimizing the weights is carried out by minimizing the difference between the output generated by the neural network and the training samples, that is, minimizing the objective function (e.g., loss function) of the neural network. In the method provided in the first aspect, not only are the weights optimized during the training process to obtain a neural network representation of the environment captured by the sensor device (the environment captured by the sensor device is represented by the mapped sensor data), but also while optimizing the weights, the relative transformation from the local first coordinate system of the first sensor device to the local second coordinate systems of the respective second sensor devices is jointly optimized.

[0013] When the training process is completed, the optimized local transformation will be known and can be easily used for the spatial calibration of the sensor devices. The method provided in the first aspect for calibrating a group of second sensor devices relative to the first sensor device can instantaneously calibrate the sensor devices in a reliable and cost-effective manner (e.g., during the driving of a vehicle equipped with sensor devices). Specifically, multiple sensor systems and devices including the sensor devices can be individually calibrated during the training process of the neural network to obtain a synthetic environmental view. There is no need to provide specially designed targets in the environment to perform the calibration process, so the calibration can be completed quickly and effectively. In addition, this method is very suitable for calibrating multi-modal sensor systems that process different data formats (e.g., including a camera device that provides 2D images and a LIDAR device that provides 3D point clouds).

[0014] In the case where the second devices are not synchronized with each other, in addition to spatial calibration, time calibration is also required. Therefore, in one implementation of the method provided according to the first aspect, the first sensor data includes a first time frame, and the second sensor data obtained by each second sensor device in the second sensor devices includes a second time frame; the neural network is trained by: while optimizing the weights of the neural network and the relative transformation from the first coordinate system to the second coordinate system, (c) jointly optimizing the time offset of the second time frame of the second sensor data obtained by at least one second sensor device in the second sensor devices with respect to the first time frame. The method provided by this implementation further includes: performing time calibration on at least one second sensor device in the second sensor devices with respect to the first sensor device according to the trained neural network, that is, the optimized time offset.

[0015] Therefore, at least one second sensor device is spatio-temporally calibrated with respect to the first sensor device according to the trained neural network. If none of the second sensor devices are synchronized with the first sensor device, all the second sensor devices in this group of second sensor devices can be time-calibrated according to this implementation. Different from the prior art, according to this implementation, spatial and time calibration can be jointly performed, and thus can be achieved in a reliable and efficient manner.

[0016] The observations of each sensor device need to be mapped to a common coordinate system. In one implementation of the method provided according to the first aspect, this mapping can be conveniently performed according to the trajectory of the (moving) first sensor device in the common coordinate system (the trajectory interpolation function obtained from discrete position / pose measurements), the prior information of the relative transformation (represented in the first coordinate system), and the prior information of the time offset.

[0017] The prior information represents initial rough estimates. For example, these estimates can be obtained according to some construction plans or computer-aided design (CAD) data.

[0018] In an implementation of the method provided according to the first aspect, the neural network includes a (deep) multi-layer perceptron (MLP) (fully connected feed-forward) neural network. Specifically, the neural network is trained according to the Neural Radiance Field (NeRF) technique, which was proposed in a paper published by B. Mildenhall et al. at the 16th European Conference on "Computer Vision - ECCV 2020" held in Glasgow, UK from August 23rd to 28th, 2020. The title of the paper is "Nerf: Representing scenes as neural radiance fields for view synthesis" and it was published by Springer in Cham in 2020. This technique has now become one of the most popular view synthesis tools. The neural network trained by NeRF provides a neural field that includes color values and spatially related volume density values. This neural field can be queried at multiple positions along a ray for volume rendering (see the detailed description below). According to this implementation, the neural network representation of the environment captured by the sensor device gives the neural field, which is used for subsequent volume rendering. The volume rendering can generate a rendered image or a point cloud, etc. During the training process, the difference between the rendered image or point cloud and the mapped sensor data is minimized.

[0019] According to this implementation, the calibration can be performed according to the neural network trained by NeRF, where the training includes not only optimizing the weights (as described in the paper by Mildenhall et al.), but also optimizing the relative transformation or the relative transformation and time offset, as described above. Since NeRF provides high rendering quality with relatively low computational cost and memory requirements, NeRF can be applied to the method of calibrating a set of second sensor devices relative to the first sensor device provided, especially when NeRF runs on an embedded computing system with limited computational resources.

[0020] According to an implementation, when the neural network is trained according to the neural radiance field technique, the joint optimization of (a) the weights of the neural network for outputting the neural network representation of the environment captured by the first and second sensor devices and (b) the relative transformation from the first coordinate system to the second coordinate system (represented in the first coordinate system) is performed by minimizing the objective function L according to the following equation:

[0021]

[0022] where time(n i) represents associating a timestamp with the corresponding time frame n i The associated function, where θ represents the weight, represents the relative transformation (expressed in the first coordinate system), represents the pose of the first sensor device in the common coordinate system, represents ground truth sensor data (e.g., an image or a point cloud), Z θ represents the neural network-based sensor data (e.g., an image or a point cloud rendered by volume rendering according to the neural field), c0 represents the first sensor device / first coordinate system, C represents the second sensor device (and the second coordinate system) c i the number of, N i represents the number of the second time frames acquired by the i-th second sensor device, ||..|| represents the norm.

[0023] This objective function can be applied to the optimization process in the context of spatial calibration of sensor devices.

[0024] According to another implementation, when the neural network is trained according to the neural radiance field technique and spatio-temporal calibration of sensor devices is to be achieved, the joint optimization (a) the weights of the neural network for outputting the neural network representation of the environment captured by the first and second sensor devices, (b) the relative transformation from the first coordinate system to the second coordinate system, and (c) the time offset of the second time frames acquired by at least one of the second sensor devices in the second sensor device relative to the first time frame is performed by minimizing the objective function L according to the following equation:

[0025]

[0026] where, time(n i ) represents associating a timestamp with the corresponding time frame n i The associated function, where θ represents the weight, represents the relative transformation, dt i represents the time offset of the second time frame relative to the first time frame, represents the trajectory of the first sensor device in the common coordinate system (a trajectory interpolation function obtained from discrete position / pose measurements), represents ground truth sensor data (e.g., an image or a point cloud), Z θ represents the neural network-based sensor data (e.g., an image or a point cloud rendered by volume rendering according to the neural field), c0 represents the first sensor device / first coordinate system, C represents the second sensor device (and the second coordinate system) c i the number of, Ni represents the number of the second time frames acquired by the i-th second sensor device, and ||..|| represents the norm.

[0027] This objective function can be applicable to the optimization process in the context of spatio-temporal calibration of sensor devices.

[0028] The method for calibrating a set of second sensor devices relative to a first sensor device provided by the above first aspect and its implementation manners can be applicable to calibrating a sensor system installed on a vehicle, and the vehicle can be (specifically, fully or partially autonomous) an automobile, an autonomous mobile robot, an AGV, etc. Therefore, according to one implementation manner, the first sensor device and the second sensor device are installed on the vehicle. In this case, the method provided by the first aspect and its implementation manners can be executed during the movement of the vehicle. Reliable calibration based on the method provided by the first aspect and its implementation manners can successfully achieve the fusion perception of numerous (multimodal) sensor devices for tracking and detection verification, ego-vehicle trajectory estimation based on multimodal sensor fusion, precise 3D detection through a multi-camera system using a stereo algorithm, and so on.

[0029] At least one of the training and calibration steps of the method provided by the first aspect and its implementation manners can be executed at a vehicle site or at a remote site provided with sensor data required for calibration using the limited computing resources of an embedded computing system.

[0030] According to a second aspect, there is provided a computer program product including computer-readable instructions. When the computer-readable instructions run on a computer (for example, a computer installed in a vehicle), the computer-readable instructions are used to execute the steps of the method provided by the first aspect or any of its implementation manners.

[0031] According to a third aspect, a calibration system is provided. The calibration system includes: a first sensor device for acquiring first sensor data in a first coordinate system; a set of second sensor devices different from the first sensor device, wherein each second sensor device is for acquiring second sensor data in a corresponding second coordinate system, the second coordinate systems being different from each other and different from the first coordinate system; a neural network; and a processing unit. The processing unit is for mapping the first sensor data and the second sensor data acquired by each second sensor device in the second sensor devices into a common coordinate system (different from the first and second coordinate systems) to obtain the mapped first and second sensor data; training the neural network by jointly optimizing (a) the weights of the neural network for outputting a neural network representation of the environment captured by the first and second sensor devices and (b) the relative transformation from the first coordinate system to the second coordinate system according to the mapped first and second sensor data. Further, the processing unit is for spatially calibrating each second sensor device in the second sensor devices relative to the first sensor device according to the trained neural network.

[0032] The first sensor device may be a camera device or a LIDAR device, and the second sensor devices may include (a) at least one camera device, (b) at least one LIDAR device, and (c) at least one of at least one camera device and at least one LIDAR device. The camera device may be a time-of-flight camera, a depth camera, etc., and the LIDAR device may be a MEMS, a LIDAR device, a solid-state LIDAR device, etc.

[0033] The calibration system provided by the third aspect and its implementation manners provides the same or similar advantages as the method provided by the first aspect and its implementation manners in combination with the above. The calibration system provided by the third aspect and its implementation manners can be used to execute the method provided by the third aspect and its implementation manners.

[0034] According to the implementation method applicable to spatio-temporal calibration, the first sensor data includes a first time frame, and the second sensor data obtained by each second sensor device in the second sensor device includes a second time frame. The processing unit is configured to train the neural network by: while optimizing the weights of the neural network and the relative transformation (represented in the first coordinate system) from the first coordinate system to the second coordinate system, (c) jointly optimizing the time offset of the second time frame obtained by at least one second sensor device in the second sensor device relative to the first time frame; and perform time calibration on at least one second sensor device in the second sensor device relative to the first sensor device according to the trained neural network.

[0035] The processing unit in the calibration system provided by the first aspect and its implementation can be used to map the first sensor data and the second sensor data obtained by each second sensor device in the second sensor device into the common coordinate system according to the trajectory of the first sensor device in the common coordinate system (the trajectory interpolation function obtained from discrete position / pose measurements), the prior information of the relative transformation (represented in the first coordinate system), and the prior information of the time offset.

[0036] According to one implementation, the processing unit in the calibration system is configured to train the neural network according to the neural radiance field technique.

[0037] According to one implementation, the processing unit in the calibration system is configured to train the neural network according to the neural radiance field technique; and jointly optimize (a) the weights of the neural network for outputting the neural network representation of the environment captured by the first and second sensor devices and (b) the relative transformation (represented in the first coordinate system) from the first coordinate system to the second coordinate system by minimizing the objective function L according to the following equation:

[0038]

[0039] where time(n i ) represents a function that associates a timestamp with the corresponding time frame m i , θ represents the weights, represents the relative transformation, represents the pose of the first sensor device in the common coordinate system, represents the ground truth sensor data (e.g., image or point cloud), Z θDenote the neural network-based sensor data (e.g., an image or a point cloud rendered by volume rendering according to the neural field), c0 denote the first sensor device / first coordinate system, C denote the second sensor device (and the second coordinate system) c i quantity of, N i Denote the quantity of the second time frames acquired by the i-th second sensor device, ||..|| denote the norm.

[0040] According to another implementation applicable to spatio-temporal calibration, the processing unit in the calibration system is configured to: train the neural network according to the neural radiance field technique; train the neural network according to the neural radiance field technique; jointly optimize (a) the weights of the neural network for outputting the neural network representation of the environment captured by the first and second sensor devices, (b) the relative transformation from the first coordinate system to the second coordinate system (represented in the first coordinate system), and (c) the time offset of the second time frames acquired by at least one of the second sensor devices in the second sensor device relative to the first time frame by minimizing the objective function L according to the following formula:

[0041]

[0042] where, time(n i ) denote the function that associates the time stamp with the corresponding time frame n i , θ denote the weights, denote the relative transformation, dt i denote the time offset of the time frames in the second time frame relative to the first time frame, denote the trajectory of the first sensor device in the common coordinate system (a trajectory interpolation function obtained according to discrete position / pose measurement values), denote the ground truth sensor data (e.g., an image or a point cloud), Z θ Denote the neural network-based sensor data (e.g., an image or a point cloud rendered by volume rendering according to the neural field), c0 denote the first sensor device / first coordinate system, C denote the second sensor device (and the second coordinate system) c i quantity of, N i Denote the quantity of the second time frames acquired by the i-th second sensor device, ||..|| denote the norm.

[0043] According to a fourth aspect, a vehicle is provided. The vehicle includes the calibration system provided by the third aspect or any one of its implementations. For example, the vehicle can be an automobile, an autonomous mobile robot, or an AGV.

[0044] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0046] Figure 1 The purpose of spatial calibration of the sensor device is shown.

[0047] Figure 2 The purpose of temporal calibration of the sensor device is shown.

[0048] Figure 3 The principle of one embodiment of the method for calibrating a sensor device provided herein is shown.

[0049] Figure 4 and Figure 5 The spatio-temporal calibration of the sensor device provided by one embodiment is shown.

[0050] Figure 6 and Figure 7 The spatial calibration of the sensor device provided by one embodiment is shown.

[0051] Figure 8 is a flowchart of a method for spatially calibrating a sensor device provided by one embodiment.

[0052] Figure 9 is a flowchart of a method for spatially and temporally calibrating a sensor device provided by one embodiment.

[0053] Figure 10 A calibration system provided by one embodiment is shown. DETAILED DESCRIPTION

[0054] A method for calibrating a sensor device (especially for spatially or spatio-temporally calibrating a multi-modal sensor) by jointly optimizing sensor observations and calibration parameters according to a NeRF scene representation is provided herein. Although the following description of the embodiments relates to NeRF technology, other techniques based on explicit representations of environments / scenes using neural fields and image rendering can also be suitably used in alternative embodiments.

[0055] A solution is provided for spatially or spatio-temporally externally calibrating multiple camera and LiDAR sensor setups commonly used in ADAS and autonomous vehicles (e.g., including single or multiple cameras and single or multiple LiDARs) without any prior known targets (i.e., providing target-free calibration). Spatio-temporal calibration includes finding the relative poses of all sensor devices with respect to the master sensor device and the time offsets of the time frames. These poses are represented by the rotation (matrix) R, the translation (matrix) T, and the time offset dt.

[0056] The translation T and the rotation R represent a rigid spatial transformation between the coordinate system centered on the master sensor device and the coordinate systems centered on each of the other sensor devices. The translation can include three translational motions along the three perpendicular axes x, y, and z. The rotation can include three rotational motions around the three perpendicular axes x, y, and z, namely roll, yaw, and pitch. The transformation of coordinates from one coordinate system to another can be achieved through matrix multiplication.

[0057] Taking the calibrated pinhole camera model commonly used in computer vision as an example, the pixel coordinates (u, v) of the projection of a 3D point (whose 3D coordinates are represented in its own coordinate system) are obtained by multiplying the 3D coordinates (subscript K) of the point represented in the camera coordinate system by the camera internal matrix K (where f x 、f y correspond to the focal lengths of the camera, in pixels, f x = f y is applicable to square pixel cameras, and u0, v0 represent the projections of the optical center of the camera on the image plane):

[0058]

[0059] where the subscript L represents the coordinate system of the master sensor device, and R and T represent the rotation and translation of the coordinate system of the master sensor device with respect to the coordinate system of one of the other sensor devices.

[0060] Figure 1 Illustrates the purpose of the spatial calibration of the sensor devices. Vehicle 10 is equipped with one or more LIDAR devices l, at least a first camera device c0 and a second camera device c1. In Figure 1 , the coordinate system (system / frame) associated with vehicle 10 is represented by v, the coordinate system (system / frame) associated with the camera device is represented by c i (i ∈ C), and the coordinate system (system / frame) associated with the LIDAR device is represented by l j (j ∈ L). The external calibration (= spatial calibration) of the sensor devices mounted on the vehicle is represented by a set of relative transformations in the vehicle coordinate system v Given, this set of relative transformations describes the transformation between the vehicle coordinate system and the coordinate systems of the camera / LIDAR device. Figure 1 As shown, the relative transformation of the coordinate system of the vehicle 10 to the coordinate system of the individual sensor devices 1, c0, c1 defines the corresponding extrinsic parameters of translation and rotation.

[0061] The vehicle coordinate system v can be replaced by the coordinate system of one of the sensor devices (the coordinate system of the “master sensor device”), for example, the coordinate system of the LIDAR device l.

[0062] Figure 2 The purpose of temporal calibration of sensor devices is shown. Consider a vehicle V with the pose (pose measurement) of its main sensor device The observations are expressed in a common (global) coordinate system w and collected in time steps (time stamps) k (k = 0, 1, 2, ... n). Based on these observations, by interpolating the measured poses in discrete time steps k, a continuous differentiable trajectory interpolation function of the master sensor device's trajectory in the common coordinate system (including the pose and the coordinate system fixed to the master sensor device) can be obtained. When other sensor devices c0 mounted on vehicle V generate measurements at timestamps k′ (k′≠k) along the vehicle trajectory, the goal of time calibration is to determine a constant time offset dt between the observations (or interpolated observations) produced by the primary sensor of vehicle V and the observations (or interpolated observations) produced by the other sensor devices at timestamps k′=k+dt. In the following, it is assumed that the time offset between the time frames captured by the primary sensor device and the time frames captured by the other sensor devices is constant.

[0063] According to the present invention, a neural network is trained based on data provided by a sensor device to represent the environment and obtain calibration parameters. For example, as shown in the NeRF technology, the neural network outputs a neural field, which can be used for volume rendering. According to a specific implementation, the neural network includes or consists of a (deep) MLP, i.e. a fully connected feedforward neural network, which is trained according to the NeRF technology proposed by B. Mildenhall et al. in a paper published at the 16th European Conference on Computer Vision - ECCV 2020, held in Glasgow, UK from August 23 to 28, 2020, entitled "Nerf: Representing scenes as neural radiance fields for view synthesis" and published by Springer in Cham in 2020. NeRF can obtain a neural network representation of the environment based on color values and spatially correlated volume density values (representing the neural field).

[0064] The input data of the neural network represents the coordinates (x, y, z) of a set of sampled 3D points and the viewing direction corresponding to the 3D points The neural network trained by NeRF outputs view-related color values (e.g., RGB) and volume density values σ. Therefore, the MLP achieves using the optimized weights Θ obtained during training

[0065] Volume rendering is based on rays passing through the scene (projected from all pixels in the image). The volume density σ(x, y, z) can be interpreted as the differential probability that a ray terminates on an infinitesimal particle at (x, y, z). By collecting all volume density values along the ray direction, the cumulative transmittance T(s) along the ray direction can be calculated as follows: The cumulative transmittance T(s) along the ray from the origin 0 to s represents the probability that the ray reaches s along its path without hitting any particles

[0066] The implicit representation (neural field) is queried at multiple positions along the ray, and then the obtained samples are combined into an image. The training objective of the NeRF neural network can be defined as:

[0067]

[0068] where θ represents the trainable parameters of the implicit NeRF representation (i.e., the weights of the neural network), and I n is the training image n (n ∈ N), and I θ (T n , K) is the image (after volume rendering) generated by NeRF at pose T n (the pose of image I n ) using the camera parameters K n (the camera parameters of image U n , e.g., resolution, field of view, etc.) and an appropriate norm ||..||. For example, the L2 norm (L2 loss function) can be adopted

[0069] A similar formula can be used to train NeRF using geometric information (e.g., the depth map of a LIDAR device):

[0070]

[0071] where D n is the training depth map n (n ∈ N), and D θ (T n , K) is the image generated by NeRF at pose T n (the pose of training depth map D n ) using the LIDAR parameters K n (the LIDAR parameters of training depth map Dn A depth map generated from LIDAR parameters (e.g., resolution, field of view, etc.).

[0072] The principle of an embodiment of the method for calibrating a sensor device according to NeRF technology provided in this article is as Figure 3 shown. The NeRF neural network is trained using multi-sensor observations of a scene by a sensor device rigidly mounted on vehicle v. All available sensor observations are mapped to a common coordinate system 33 using the main sensor trajectory (or device / vehicle trajectory; see also the description above) 31 and external prior information (e.g., an initial rough estimate of calibration parameters obtained from some construction plans or CAD data) 32. The implicit representation F of the scene θ is trained by minimizing the difference between multi-modal observations (color or depth) and the predicted modality generated by ray marching based on the implicit representation obtained from a neural network known in the art.

[0073] According to the present invention, while optimizing the weights of the implicit representation, the same optimization objective function for adjusting the neural network weights is used to jointly optimize (simultaneously in the same adjustment / backpropagation step) the rigid external calibration between the sensor devices mounted on the vehicle 34. In the case of spatio-temporal calibration, while optimizing the weights of the implicit representation and the external calibration parameters, the same optimization objective function is also used to jointly optimize the acquisition time offset between the sensor devices. After training convergence is reached, the optimized spatio-temporal calibration 35 of the sensor devices of the vehicle is obtained.

[0074] Figure 4 and Figure 5 shows the spatio-temporal calibration of a sensor device based on NeRF technology provided by an embodiment. For example, the time and space calibration is directly performed on a device (e.g., a vehicle) equipped with one or more camera devices and one or more LIDAR devices. The time and space calibration can be at least partially performed by a remote server provided with the necessary data.

[0075] To train the implicit representation of a scene captured by a multi-modal sensor device, the observations of each sensor device in the sensor device need to be placed in a common coordinate system (global coordinate system, world coordinate system, w). For this purpose, the pose prior information (first estimate) of each sensor device needs to be obtained.

[0076] To calculate the pose prior information of each sensor device in the common coordinate frame, this example requires the following inputs (see the "Input" column in Figure 4 ):

[0077] · The trajectory of the main sensor device (e.g., a LIDAR device or the first camera c0) in the common coordinate system can be easily obtained through technologies such as the Global Navigation Satellite System (GNSS), Simultaneous Localization and Mapping (SLAM), or Structure from Motion (SfM).

[0078] and so on.

[0079] · Prior information on the external calibration of all (second, auxiliary) sensor devices different from the main (first) sensor device, that is, the relative transformation from the coordinate system of the main sensor device to the coordinate systems of each second sensor device (in the coordinate system of the main sensor device).

[0080]

[0081] · Prior information on the time offset dt between the main sensor device and the second sensor device i between them.

[0082] The pose prior information of each observation of each sensor device i in the common coordinate system can be calculated according to these inputs through the following equation (see 51 and 52 in Figure 5 ):

[0083]

[0084] Given the pose prior information of each observation of each sensor device i in the common coordinate system, the observations of each sensor device in sensor device i in the common coordinate system w can be obtained (see 41 in Figure 4 ). The observations of each sensor device in sensor device i in the common coordinate system w are used to train the neural network (see 42 in Figure 4 ). Training the neural network 42 includes training the implicit representation of the corresponding scene, that is, optimizing the weights of the neural network, and jointly optimizing the relative transformation and the time offset dt. After achieving optimization convergence, the optimized actual relative transformation (external spatial calibration) and the time offset dt are obtained, that is, output by the neural network or obtained according to the output of the neural network (see Figure 4 ).

[0085] According to the objective given by Equation 1 above, the optimization problem for C cameras and the observations obtained within N time frames can be defined as:

[0086]

[0087] where, time(n i ) represents a function that associates a timestamp with the corresponding time frame n i and N i represents the number of second time frames obtained by the i-th second sensor device. The training image is the observation (sensor data) of the second sensor device in the common coordinate system. Camera parameters are not considered here because these parameters will remain constant and thus will not affect the target. By using the trajectory interpolation function (see Equation 3), the following can be obtained:

[0088]

[0089] where, is the identity matrix. Due to the differentiability of the image rendering function I θ with respect to the pose, the relative transformation can be added to the optimization parameter θ (weights); due to the differentiability of the trajectory interpolation function, the time offset can also be added to the optimization parameter.

[0090] Other camera devices of the sensor device can also be processed accordingly. For example, the observations of the LIDAR device can be input into the optimization process by replacing the ground truth image U n and the generated (i.e., rendered) image U θ with the ground truth LIDAR ray depth map D n (see Equation 2) and the depth map D θ generated according to the neural field and volume rendering. Due to the implicit representation with weights θ, the camera and LIDAR observations are tightly coupled together in the optimization process.

[0091] Figure 6 and Figure 7 show the spatial calibration of the mutually synchronized sensor devices provided by an embodiment based on NeRF technology. Compared with the embodiment described in combination with Figure 4 and Figure 5 , the calibration is easier and faster when all sensor devices are mutually synchronized (e.g., hardware synchronization).

[0092] To calculate the pose prior information of each sensor device in the common coordinate frame, this example requires the following inputs (see the "Input" column in Figure 6 ):

[0093] · The pose of the main sensor device (e.g., LIDAR device or the first camera c0) that can be easily measured in the common coordinate system

[0094] · Prior information on the external calibration of all (second, auxiliary) sensor devices different from the main (first) sensor device, i.e., the relative transformation from the coordinate system of the main sensor device to the coordinate systems of the respective second sensor devices (in the coordinate system of the main sensor device)

[0095] The pose prior information of each observation of each sensor device i in the common coordinate system can be calculated based on these inputs by the following equation:

[0096]

[0097] Given the pose prior information of each observation of each sensor device i in the common coordinate system, the observations of each sensor device in sensor device i in the common coordinate system w can be obtained (see Figure 6 61 in). The observations of each sensor device in sensor device i in the common coordinate system w are used to train the neural network (see Figure 6 62 in). After the training process is completed, the actual relative transformation (external space calibration) is obtained, that is, output by the neural network or obtained based on the output of the neural network (see Figure 6 ).

[0098] Compared with the embodiment described in combination with Figure 4 and Figure 5 , the goal is simplified to:

[0099]

[0100] The process of jointly optimizing the weights and the external space calibration parameters is as shown in Figure 7 . The observations generated by multiple sensor devices at different poses 71 are used to train the implicit scene representation F θ 72. By updating the scene representation and the observation positions (poses), the difference between the observed scene and the scene representation is minimized 73.

[0101] Figure 8 Fig. shows a method 80 for spatially calibrating a second sensor device relative to a first sensor device provided by an embodiment. The method 80 can be applied to the case where the second sensor device that needs to be spatially calibrated relative to the first sensor device is synchronized with the first sensor device.

[0102] The method 80 includes: the first sensor device (e.g., a camera device or a LIDAR device, e.g., Figure 1The LIDAR device l) shown obtains first sensor data (e.g., images or point clouds of the environment at different time frames) in a first coordinate system (i.e., a coordinate system fixed to and moving with the first sensor device), and each second sensor device in a second sensor device (e.g., a camera device, a LIDAR device, or a combination of a camera device and a LIDAR device, e.g., Figure 1 each second sensor device in the camera devices c0 and c1) shown obtains second sensor data (e.g., images or point clouds of the environment at different time frames) in a corresponding second coordinate system (each second sensor device in the second sensor device has a fixed second coordinate system), where the second coordinate systems are different from each other and different from the first coordinate system.

[0103] Method 80 further includes: mapping S82 the first sensor data and the second sensor data obtained by each second sensor device in the second sensor device to a common (global, world) coordinate system to obtain the mapped first and second sensor data (i.e., mapped to the common coordinate system). The mapping S82 can be performed using the prior information described in the "Input" column above in conjunction with Figure 6 the "Input" column.

[0104] Method 80 further includes: training S83 a neural network by jointly optimizing the weights of the neural network for outputting a neural network representation of the environment captured by the first and second sensor devices and the relative transformation from the first coordinate system to the second coordinate system based on the mapped first and second sensor data. The optimization process can be performed according to the example described above in conjunction with Figure 6 described; see Equation 7 specifically.

[0105] Method 80 further includes: spatially calibrating S84 each second sensor device in the second sensor device relative to the first sensor device, i.e., using the optimized relative transformation from the first coordinate system to the second coordinate system.

[0106] Figure 9 Illustrated is a method 90 for spatially and temporally calibrating a second sensor device relative to a first sensor device provided by an embodiment. This method is applicable to the case where the second sensor device that needs to be calibrated relative to the first sensor device is out of sync with the first sensor device.

[0107] Method 90 includes: a first sensor device (e.g., a camera device or a LIDAR device, e.g., Figure 1The LIDAR device l) shown acquires first sensor data (e.g., images or point clouds of the environment at different time frames) in a first coordinate system (i.e., a coordinate system fixed to and moving with the first sensor device), and each second sensor device in a second sensor device (e.g., a camera device, a LIDAR device, or a combination of a camera device and a LIDAR device, e.g., Figure 1 each second sensor device in the shown camera devices c0 and c1) acquires second sensor data (e.g., images or point clouds of the environment at different time frames) in a corresponding second coordinate system (each second sensor device in the second sensor device has a fixed second coordinate system), where the second coordinate systems are different from each other and different from the first coordinate system. The first sensor data includes a first time frame, and the second sensor data acquired by each second sensor device in the second sensor device includes a second time frame.

[0108] Method 90 further includes: mapping S92 the first sensor data and the second sensor data acquired by each second sensor device in the second sensor device to a common (global, world) coordinate system to obtain the mapped first and second sensor data (i.e., mapped to the common coordinate system). The mapping S92 can be performed using the prior information described in conjunction with Figure 4 the "input" column above.

[0109] Method 90 further includes: training S93 a neural network by jointly optimizing the time offset of the second time frame of the second sensor data acquired by at least one second sensor device in the second sensor device relative to the first time frame while optimizing the weights of the neural network and the relative transformation from the first coordinate system to the second coordinate system. The optimization process can be performed according to the example described in conjunction with Figure 4 above; specifically, see Equation 5.

[0110] Method 90 further includes: spatially and temporally calibrating S94 at least one second sensor device in the second sensor device relative to the first sensor device according to the trained neural network.

[0111] Figure 10 Figure 19 shows a calibration system 100 provided by an embodiment. The calibration system can be installed on a vehicle (e.g., an automobile). The calibration system 100 can be used to perform the method steps of the above method 80 or 90.

[0112] Specifically, Figure 10 the calibration system 100 shown includes: a first sensor device (e.g., a camera device or a LIDAR device, e.g., Figure 1The LIDAR device l) shown is used to obtain first sensor data (e.g., images or point clouds of the environment at different time frames) in a first coordinate system (i.e., a coordinate system fixed to and moving with the first sensor device); a second sensor device (e.g., a camera device, a LIDAR device, or a combination of a camera device and a LIDAR device, e.g., Figure 1 the camera devices c0 and c1 shown) is used to obtain second sensor data (e.g., images or point clouds of the environment at different time frames) in corresponding second coordinate systems (each second sensor device in the second sensor device has a fixed second coordinate system), where the second coordinate systems are different from each other and different from the first coordinate system.

[0113] The calibration system 100 further includes a processing unit 103 and a neural network 104. The neural network 104 can be an MLP neural network. The processing unit 103 is used to map the first sensor data and the second sensor data obtained by each second sensor device in the second sensor device into a common coordinate system to obtain the mapped first and second sensor data. The processing unit 103 can be used to map the first and second sensor data into a common coordinate system using the prior information described in the "Input" column in combination with Figure 4 or Figure 6 above.

[0114] In addition, the processing unit 103 is used to: train the neural network by jointly optimizing (a) the weights of the neural network for outputting the neural network representation of the environment captured by the first and second sensor devices and (b) the relative transformation from the first coordinate system to the second coordinate system according to the mapped first and second sensor data; perform spatial calibration on each second sensor device in the second sensor device relative to the first sensor device according to the trained neural network 104. The optimization process can be performed according to the examples described in combination with Figure 6 above; specifically, see Equation 7.

[0115] Alternatively, the processing unit 103 is used to: train the neural network by jointly optimizing (a) the weights of the neural network for outputting the neural network representation of the environment captured by the first and second sensor devices, (b) the relative transformation from the first coordinate system to the second coordinate system, and (c) the time offset of the second time frame obtained by at least one second sensor device in the second sensor device relative to the first time frame while optimizing the weights of the neural network and the relative transformation from the first coordinate system to the second coordinate system according to the mapped first and second sensor data. The optimization process can be performed according to the examples described in combination with Figure 4 above; specifically, see Equation 5.

[0116] In addition, the processing unit 103 is configured to perform spatial or spatio-temporal calibration on each of the second sensor devices with respect to the first sensor device according to the trained neural network 104.

[0117] Specifically, training the neural network 104 can be performed in the manner described in the embodiments respectively in conjunction with Figure 4 and Figure 6 as shown.

[0118] It should be noted that in the Figures 3 to 10 embodiment shown, training the implicit scene representation and estimating the optimized spatio-temporal or spatial calibration parameters can be performed on the target device (e.g., a vehicle). In the case of limited computing power, the data can be recorded by the device, and the optimization can be performed partially or entirely during the "offline time" of the device. The recorded data can be sent from the device to an external processing unit to perform spatio-temporal or spatial calibration. When the embedded computing resources are powerful enough, the optimized calibration parameters of the host device can be obtained without data transmission.

[0119] Based on the accurately estimated time and spatial calibration parameters, any number of sensor devices, such as camera devices and / or LIDAR devices, can be reliably calibrated at a relatively low cost.

[0120] The embodiments of the above methods and apparatuses can be appropriately integrated in vehicles such as automobiles, AGVs, and autonomous mobile robots for navigation, positioning, and obstacle avoidance. Instant calibration or recalibration can be performed on a series of vehicles equipped with multiple sensors. In one exemplary scenario, an initial sensor calibration is provided to the vehicle (e.g., according to a CAD blueprint). During the first drive, data is collected to apply the spatio-temporal or spatial calibration process provided by the above embodiments. After calibration optimization, the vehicle software can enable automatic functions.

[0121] In the context of automobiles, the embodiments of the above methods and apparatuses can be constituted by ADAS. In the context of autonomous driving / ADAS applications, reliable calibration of sensor devices helps to achieve the fusion perception of numerous sensor devices for tracking and detection verification, ego-vehicle trajectory estimation based on multi-modal sensor fusion, precise 3D detection through a multi-camera system using stereo algorithms, and so on.

[0122] All of the embodiments discussed above are not intended to be limiting, but rather to illustrate the features and advantages of the present invention. It should be understood that the above-mentioned features, in whole or in part, can also be combined in different ways.

Claims

1. A method (80, 90) for calibrating a set of second sensor devices (102) relative to a first sensor device (101), characterized in that, The first sensor device (101) is different from the second sensor device (102), and the method includes the following steps: The first sensor device (101) acquires (S81, S91) first sensor data in a first coordinate system, and each second sensor device in the second sensor device (102) acquires second sensor data in a corresponding second coordinate system, where the second coordinate systems are different from each other and different from the first coordinate system; Map (S82, S92) the first sensor data and the second sensor data acquired by each second sensor device in the second sensor device (102) into a common coordinate system to obtain the mapped first and second sensor data; According to the mapped first and second sensor data, train (S83) a neural network (104) by jointly optimizing the following two: (a) The weights of the neural network (104) for outputting a representation of the environment captured by the first and second sensor devices (102) of the neural network (104) (b) The relative transformation from the first coordinate system to the second coordinate system; According to the trained neural network (104), perform spatial calibration (S84, S94) on each second sensor device in the second sensor device (102) relative to the first sensor device (101).

2. The method (80, 90) according to claim 1, wherein The first sensor data includes a first time frame, and the second sensor data acquired by each second sensor device in the second sensor device (102) includes a second time frame; The neural network (104) is trained (S93) by: while optimizing the weights of the neural network (104) and the relative transformation from the first coordinate system to the second coordinate system, (c) jointly optimizing the time offset of the second time frame of the second sensor data acquired by at least one second sensor device in the second sensor device (102) relative to the first time frame; The method further includes: According to the trained neural network (104), perform time calibration (S94) on the at least one second sensor device in the second sensor device (102) relative to the first sensor device (101).

3. The method according to claim 2, wherein The mapping of the first sensor data and the second sensor data acquired by each second sensor device in the second sensor device (102) into the common coordinate system is performed according to the trajectory of the first sensor device (101) in the common coordinate system, prior information of the relative transformation, and prior information of the time offset.

4. The method (80, 90) according to any one of the preceding claims, characterized in that, The neural network (104) is trained (S83, S93) according to the neural radiance field technique.

5. The method (80, 90) according to claim 1, wherein The neural network (104) is trained (S83) according to the neural radiance field technique; The joint optimization (a) of the weights of the neural network (104) for outputting a neural network (104) representation of the environment captured by the first and second sensor devices (102) and (b) the relative transformation from the first coordinate system to the second coordinate system is performed by minimizing an objective function L according to the following equation: where time(n i ) represents a function that associates a timestamp with the corresponding time frame n i , θ represents the weight, represents the relative transformation, represents the pose of the first sensor device (101) in the common coordinate system, represents the ground truth sensor data, Z θ represents the sensor data representation based on the neural network (104), c0 represents the first sensor device, and C represents the number of the second sensor devices c i (102), N i represents the number of the second time frames obtained by the i-th second sensor device, and ‖..‖ represents the norm.

6. The method (80, 90) according to claim 2 or 3, characterized in that the neural network (104) is trained according to neural radiance field technology (S93); The joint optimization (a) of the weights of the neural network (104) for outputting a neural network (104) representation of the environment captured by the first and second sensor devices (102), (b) the relative transformation from the first coordinate system to the second coordinate system, and (c) the time offset of the second time frame acquired by at least one of the second sensor devices in the second sensor device (102) relative to the first time frame is performed by minimizing an objective function L according to the following equation: where, time(n i ) represents a function that associates a timestamp with the corresponding time frame n i , θ represents the weight, represents the relative transformation, dt i represents the time offset of the second time frame relative to the first time frame, represents the trajectory of the first sensor device (101) in the common coordinate system, represents the ground truth sensor data, Z θ represents the sensor data representation based on the neural network (104), c0 represents the first sensor device, C represents the second sensor device c i (102) quantity, N i represents the quantity of the second time frame acquired by the i-th second sensor device, ‖..‖ represents the norm.

7. The method (80, 90) according to any one of the preceding claims, characterized in that The first sensor device (101) is a camera device or a light detection and ranging LIDAR device, and the second sensor device (102) includes (a) at least one camera device, (b) at least one LIDAR device, and (c) at least one of at least one camera device and at least one LIDAR device.

8. The method (80, 90) according to claim 7, characterized in that, The first sensor device (101) and the second sensor device (102) are mounted on a vehicle.

9. The method (80, 90) according to claim 8, wherein, The method is performed while the vehicle is moving.

10. The method (80, 90) according to claim 8 or 9, characterized in that, At least one of the training (S83, S93) and calibration (S84, S94) steps is performed at the vehicle site or a remote site.

11. A computer program product comprising computer-readable instructions, characterized in that, When the computer-readable instructions are run on a computer, the computer-readable instructions are used to perform the steps of the method (80, 90) according to any one of the above claims.

12. Calibration system (100), characterized in that, Comprising: A first sensor device (101) for acquiring first sensor data in a first coordinate system; A set of second sensor devices (102) different from the first sensor device (101), wherein each second sensor device is used to acquire second sensor data in a corresponding second coordinate system, and the second coordinate systems are different from each other and different from the first coordinate system; A neural network (104); A processing unit (103) for performing the following operations: Mapping the first sensor data and the second sensor data acquired by each second sensor device in the second sensor device (102) into a common coordinate system to obtain the mapped first and second sensor data; According to the mapped first and second sensor data, training the neural network (104) by jointly optimizing the following two: (a) The weights of the neural network (104) for outputting a neural network (104) representation of the environment captured by the first and second sensor devices (102) and (b) The relative transformation from the first coordinate system to the second coordinate system; Perform spatial calibration for each second sensor device in the second sensor device (102) relative to the first sensor device (101) according to the trained neural network (104).

13. The calibration system (100) according to claim 12, wherein The first sensor data includes a first time frame, and the second sensor data acquired by each second sensor device in the second sensor device (102) includes a second time frame; The processing unit (103) is configured to: Train the neural network (104) by: while optimizing the weights of the neural network (104) and the relative transformation from the first coordinate system to the second coordinate system, (c) jointly optimize the time offset of the second time frame acquired by at least one second sensor device in the second sensor device (102) relative to the first time frame; Perform time calibration for the at least one second sensor device in the second sensor device (102) relative to the first sensor device (101) according to the trained neural network (104).

14. The calibration system (100) according to claim 12 or 13, characterized in that, The processing unit (103) is configured to map the first sensor data and the second sensor data acquired by each second sensor device in the second sensor device (102) into the common coordinate system according to the trajectory of the first sensor device (101) in the common coordinate system, the prior information of the relative transformation, and the prior information of the time offset.

15. The calibration system (100) according to any one of claims 12 to 14, characterized in that, The processing unit (103) is configured to train the neural network (104) according to the neural radiance field technique.

16. The calibration system (100) according to claim 12, wherein, The processing unit (103) is configured to: Train the neural network (104) according to the neural radiance field technique; Jointly optimize (a) the weights of the neural network (104) representing the neural network (104) for outputting the environment captured by the first and second sensor devices (102) and (b) the relative transformation from the first coordinate system to the second coordinate system by minimizing the objective function L according to the following equation: where, time(n i ) represents a function that associates a timestamp with the corresponding time frame n i , θ represents the weight, represents the relative transformation, represents the pose of the first sensor device (101) in the common coordinate system, represents the ground truth sensor data, Z θ represents the sensor data representation based on the neural network (104), c0 represents the first sensor device, C represents the second sensor device c i (102) quantity, N i represents the quantity of the second time frame obtained by the i-th second sensor device, ‖..‖ represents the norm.

17. The calibration system (100) according to claim 13 or 14, characterized in that, The processing unit (103) is configured to: Train the neural network (104) according to the neural radiance field technique; Jointly optimize (a) the weights of the neural network (104) representing the neural network (104) for outputting the environment captured by the first and second sensor devices (102), (b) the relative transformation from the first coordinate system to the second coordinate system, and (c) the time offset of the second time frame acquired by at least one second sensor device in the second sensor device (102) relative to the first time frame by minimizing the objective function L according to the following equation: where, time(n i ) represents a function that associates a timestamp with the corresponding time frame n i , θ represents the weight, represents the relative transformation, dt i represents the time offset of the time frame in the second time frame relative to the first time frame, represents the trajectory of the first sensor device (101) in the common coordinate system, represents the ground truth sensor data, Z θ represents the sensor data representation based on the neural network (104), c0 represents the first sensor device, C represents the second sensor device c i (102) quantity, N i represents the quantity of the second time frame obtained by the i-th second sensor device, ‖..‖ represents the norm.

18. The calibration system (100) according to any one of claims 12 to 17, characterized in that, The first sensor device (101) is a camera device or a LIDAR device, and the second sensor device (102) includes (a) at least one camera device, (b) at least one LIDAR device, and (c) at least one of at least one camera device and at least one LIDAR device.

19. A vehicle, characterized in that, Comprising a calibration system (100) according to any one of claims 12 to 18.

Citation Information

Patent Citations

  • Calibration method integrated with three-dimensional point cloud and two-dimensional image based on neural network

    CN112085801A

  • Neural radiation field enhancement method based on joint pose optimization

    CN112613609A

  • Methods and systems for joint pose and shape estimation of objects from sensor data

    US20210150228A1