A target-free space-time calibration method based on staged decoupling optimization
Patent Information
- Application Number
- CN202610778869.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-18
AI Technical Summary
由于事件相机的输出完全依赖于场景的动态变化,固定时间窗口在相机慢速运动或场景纹理稀疏时,截取到的事件数量极少,导致信噪比骤降、对比度方差计算失效;而在极高速运动时,又会导致大量非线性运动被包含在窗口内,破坏了短时匀速运动假设,从而进一步恶化了标定的稳定性和精度
(1)本发明摆脱了传统标定方法对特定标定板和特定实验室光照环境的依赖。通过最大化运动补偿后的事件联合对比度,利用IMU作为刚性运动桥梁,本发明能够直接在自然场景下,实现对完全无重叠视场的多事件相机系统的精确标定,极大提升了机器人和异构传感器平台部署的灵活性。
Smart Images

Figure CN122780409A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision, multi-sensor fusion, and robot perception technology, specifically to a targetless spatiotemporal calibration method based on phased decoupling optimization. Background Technology
[0002] With the development of machine vision technology, event cameras, as a biomimetic sensor inspired by biological vision, have attracted widespread attention in fields such as high-speed motion estimation, robot navigation, and autonomous driving due to their advantages such as extremely high temporal resolution, ultra-high dynamic range, and low power consumption. In practical 3D perception and self-motion estimation tasks, event cameras are usually combined with inertial measurement units to construct a vision-inertial system. To achieve effective fusion of multi-sensor data, obtaining high-precision spatial extrinsic parameters and temporal offsets is an essential prerequisite.
[0003] Currently, calibration methods for event cameras and IMUs are mainly divided into two categories: "target-based" and "targetless." However, both have significant limitations in practical applications. First, traditional target-based calibration methods, such as those relying on checkerboard patterns, AprilTags, or flashing LED arrays, require specific laboratory environments and manual intervention. More critically, many modern robotic platforms, such as drones and multimodal mapping vehicles, typically employ sensor layouts with non-overlapping fields of view to maximize the sensing range. In this hardware topology, target-based methods cannot allow multiple cameras to simultaneously observe the same calibration board, leading to an extremely cumbersome or even completely ineffective calibration process. Second, to eliminate dependence on the calibration board, existing targetless calibration methods calibrate by directly maximizing the sharpness of the motion-compensated event image. However, existing joint optimization frameworks typically incorporate spatial rotation extrinsic parameters and temporal offsets into the same optimizer for simultaneous solution. In real-world natural scenes, this simultaneous optimization can lead to severe parameter coupling or parameter absorption effects. In an attempt to forcibly increase image contrast on an unaligned timeline, optimizers often get trapped in non-physical local extrema, leading to severe angular drift in the output rotation extrinsic parameters. Finally, most existing event camera motion compensation algorithms use fixed time windows to capture the event stream. Since the output of an event camera is entirely dependent on the dynamic changes of the scene, a fixed time window captures very few events when the camera is moving slowly or the scene texture is sparse, resulting in a sharp drop in signal-to-noise ratio and failure in contrast variance calculation. Conversely, at extremely high speeds, a large amount of nonlinear motion is included within the window, violating the assumption of short-term uniform motion and further deteriorating the stability and accuracy of the calibration.
[0004] In summary, existing calibration techniques cannot stably and accurately decouple and calibrate the spatiotemporal parameters of the event camera and IMU in non-overlapping fields of view and natural, target-free scenarios. Therefore, there is an urgent need for a multi-sensor joint calibration method that can adapt to event flow density in natural scenes and effectively overcome spatiotemporal parameter coupling drift. Summary of the Invention In view of the above problems, this invention provides a targetless spatiotemporal calibration method based on phased decoupling optimization. This method can effectively eliminate the coupling drift of spatiotemporal parameters in natural scenes without targets and where the fields of view of multiple event cameras do not overlap. This is achieved through a slicing mechanism based on a fixed number of events and a phased decoupling optimization strategy, enabling high-precision joint calibration of spatial extrinsic parameters, temporal offset, and IMU zero bias between the event camera and the IMU.
[0005] This invention discloses a targetless spatiotemporal calibration method based on phased decoupling optimization, comprising the following steps: acquiring event stream data from synchronously acquired event cameras and angular velocity data from inertial measurement units (IMUs); employing an adaptive slicing strategy based on a fixed number of events to divide the event stream data from multiple event cameras into multiple event data blocks with constant information density, and extracting IMU angular velocity data within the time period corresponding to each event data block to obtain multi-sensor data blocks; constructing a differentiable event motion compensation and joint contrast model; based on the joint contrast model, employing a phased decoupling optimization strategy to jointly optimize the spatiotemporal parameters of the targetless event cameras and IMUs to obtain calibration results; wherein, the spatiotemporal parameters include the rotational extrinsic parameters from each event camera to the IMU, time offset, and IMU zero offset; The specific steps of the phased decoupling optimization strategy include: Synchronize and align the data from all event cameras with the IMU data on the time axis; freeze the rotation extrinsic parameters from each event camera to the IMU, and perform a stage to optimize the time offset and IMU zero bias; The stage of freezing the time offset and optimizing the rotating extrinsic parameters with prior constraints; The stage of global joint fine-tuning of all spatiotemporal parameters. Optionally, the specific steps of dividing the event stream data of multiple event cameras into multiple event data blocks with constant information density and extracting IMU angular velocity data within the time period corresponding to the event data block to obtain multi-sensor data blocks include: setting a fixed event threshold; when the number of events accumulated by the event camera designated as the main camera reaches the fixed event threshold, recording the corresponding time interval; dividing the continuous event stream data into multiple event data blocks according to the time interval, and extracting IMU angular velocity data within the time period corresponding to each event data block. Optionally, the construction of a differentiable event motion compensation and joint contrast model includes: using B-splines combined with differentiable spherical linear interpolation to construct a continuous time trajectory model based on the angular velocity after IMU zero bias compensation; correcting the event timestamp according to the time offset, and using the continuous time trajectory model and the rotation extrinsic parameters to be optimized to map the event points to the reference time to obtain the compensated sub-pixel coordinates; generating event compensation images based on the compensated sub-pixel coordinates of all event cameras, and calculating the negative value of the sum of the image variances of all event compensation images as the contrast loss. Optionally, in the phased decoupling optimization strategy: in the stage of optimizing the time offset and IMU zero bias, the weight of the prior constraint term is set to zero; in the stage of optimizing the rotational extrinsic parameters with prior constraints, the prior penalty term is activated, which is used to constrain the fine-tuning amplitude of the rotational extrinsic parameters near the initial value; in the global joint fine-tuning stage, all spatiotemporal parameters are jointly optimized with a learning rate lower than that of the first two stages. Optionally, it also includes a targetless spatial extrinsic parameter coarse initialization step: screening multi-sensor data blocks that meet the conditions of violent motion; for each event camera, finding the local visual angular velocity that maximizes the contrast of the event accumulated image through an optimization algorithm; using modulus consistency constraints and variance gain constraints to screen effective angular velocity observation vector pairs, and obtaining the initial rotational extrinsic parameters from each event camera to the IMU by solving the Wahba problem. Optionally, the total loss function of the phased decoupling optimization strategy includes contrast loss, IMU zero bias regularization term, and prior constraint term of rotational extrinsic parameters; wherein the calculation of contrast loss depends on the time offset and rotational extrinsic parameters. Optionally, the method further includes: recording the updated spatiotemporal parameter values after processing each event data block during the last optimization cycle, and calculating the arithmetic mean of all spatiotemporal parameters during the optimization cycle as the final calibration result. Optionally, the event camera data consists of data acquired by at least two event cameras with non-overlapping fields of view; wherein one event camera is defined as the master camera, and the others are slave cameras.
[0006] Compared with the prior art, the present invention has at least the following beneficial effects: (1) This invention eliminates the dependence of traditional calibration methods on specific calibration boards and specific laboratory lighting environments. By maximizing the joint contrast of events after motion compensation and using the IMU as a rigid motion bridge, this invention can directly achieve accurate calibration of multi-event camera systems with completely non-overlapping fields of view in natural scenes, greatly improving the flexibility of robot and heterogeneous sensor platform deployment.
[0007] (2) To address the parameter absorption effect that joint optimization can easily lead to in low-texture scenes, this invention innovatively introduces a three-stage optimization mechanism of "time-first, prior extrinsic parameter refinement, and global fine-tuning". Combined with prior regularization constraints on the rotation matrix, it effectively cuts off the erroneous coupling path between spatiotemporal parameters, reduces the angle error of blind calibration from several degrees to sub-degree level in conventional methods, and significantly improves the physical correctness and robustness of calibration.
[0008] (3) This invention abandons the traditional fixed-time-window truncation method and adopts an adaptive slicing strategy based on a fixed number of events. This mechanism ensures the accumulation of sufficient edge features in slow-motion or sparse-texture scenes, while automatically shortening the time window during extremely high-speed rotation to strictly satisfy the mathematical assumption of uniform motion. This enables the invention to maintain a constant and high-quality contrast gradient signal in an extremely high dynamic range, significantly improving the convergence success rate of the objective function. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the embodiments will be briefly introduced below. The features and advantages of the present invention can be more clearly understood by referring to the accompanying drawings. The accompanying drawings are schematic and should not be construed as limiting the present invention in any way. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a flowchart of the targetless spatiotemporal calibration method based on phased decoupling optimization according to the present invention. Detailed Implementation
[0011] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.
[0012] A specific embodiment of the present invention, such as Figure 1 This paper presents a targetless spatiotemporal calibration method based on staged decoupling optimization, with the following specific steps: Step S1: Adaptively slice the multi-sensor data, including event camera data and IMU data, acquired synchronously to obtain multi-sensor data blocks.
[0013] Specifically, event stream data is data used to record dynamic visual information. The form of event stream data is an event sequence, which contains multiple events arranged chronologically. Each event contains a parameter vector. x , y , p , t ],in, x and y For the event x axis pixel coordinates and y The axis pixel coordinates indicate the location in the image where the event occurred. p The polarity of the event represents the change in brightness. p A value of 1 indicates an increase in brightness (positive polarity). p A value of 0 indicates reduced brightness (negative polarity). t This is the timestamp of the event, indicating when the event occurred.
[0014] Since the event camera outputs a non-stationary asynchronous event stream, the traditional fixed-time-window slicing method results in an extremely low signal-to-noise ratio during slow motion and violates the short-time uniform motion assumption during fast motion. Therefore, this embodiment employs an adaptive slicing mechanism based on a fixed number of events.
[0015] Step S1-1: Acquire event stream data and IMU angular velocity data from multiple event cameras. Define one event camera as the master camera and the rest as slave cameras.
[0016] Step S1-2: Set fixed event thresholds When the number of events accumulated by the main camera reaches Record fixed event thresholds at that time. Internal corresponding start timestamp and end timestamp The event data stream within the time interval.
[0017] Step S1-3: Based on time interval The continuous event stream data is divided into multiple event data blocks containing the same information density, and the IMU angular velocity data within the time period corresponding to the event data block is extracted to obtain a multi-sensor data block.
[0018] The mechanism of this invention ensures that the image information density used for subsequent optimization remains constant at any motion speed.
[0019] Step S2: Coarse initialization of extrinsic parameters in the targetless space.
[0020] To prevent subsequent non-convex optimization from getting trapped in local minima, we first obtain initial rotational extrinsic parameters that have been filtered through double constraints, without any prior time offset. .
[0021] Step S2-1: Perform intense motion screening and calculate the average modulus of the IMU angular velocity data within the current multi-sensor data block. ;like If the data is less than a preset motion threshold, skip this multi-sensor data block; otherwise, proceed to the next step. This represents the angular velocity of the IMU.
[0022] Step S2-2: For each event camera, use a grid search combined with the Nelder-Mead optimization algorithm to find the local visual angular velocity that maximizes the cumulative image contrast of the current multi-sensor data block events. .
[0023] Furthermore, to accelerate the search, the IMU angular velocity modulus is employed. As a priori guide for the sampling step size.
[0024] Step S2-3: Introduce dual constraints: (1) Module length consistency constraint: Judgment (1) Whether it is within the allowable error range; (2) Variance gain constraint: Determine whether the variance of the image after motion compensation is significantly improved compared to the uncompensated version. Angular velocity observation vectors that satisfy the constraints As valid vector pairs, the random sample consensus algorithm is used to eliminate outliers caused by directional ambiguity, and the Wahba problem is solved through singular value decomposition to obtain the initial rotation extrinsic parameters from the camera to the IMU for each event. .
[0025] Step S3: Construct a differentiable event motion compensation and joint contrast model.
[0026] Specifically, in this step, an end-to-end differentiable mathematical model is established, so that the objective function, with the joint contrast of the event images as its core, can adapt to the time offset. IMU zero bias and rotation extrinsic matrix Calculate the partial derivatives to perform joint gradient optimization.
[0027] Step S3-1: Zero bias of the IMU gyroscope to be optimized The compensation is added to the original angular velocity of the IMU. A continuous-time trajectory model is constructed using B-splines combined with differentiable spherical linear interpolation. This allows it to query the IMU's attitude at any floating-point time.
[0028] Step S3-2: Construct differentiable event motion compensation: Offset correction of event timestamps. For in The coordinates of the k-th event point triggered at time t. Map it to the reference time To achieve motion compensation for events and obtain the compensated sub-pixel coordinates, the compensation formula is as follows:
[0029]
[0030] in, The equivalent rotation matrix in the camera coordinate system. Let be the intrinsic parameter matrix of the event camera, and k represent the k-th event in the current event data block. The sub-pixel coordinates after compensation for the k-th event; Indicates the event alignment time after time offset correction; Indicates the time offset; This represents the rotation extrinsic matrix from the IMU coordinate system to the event camera coordinate system to be optimized; This represents the attitude matrix of the IMU relative to the world coordinate system at a specific moment, obtained through the integration of IMU inertial data.
[0031] Step S3-3: Use bilinear interpolation to calculate the compensated sub-pixel coordinates. The event compensation image is generated by summing the pixel grids of the corresponding event camera. The sum of the image variances of all event compensation images from all event cameras is calculated, and the negative value is taken as the contrast loss. .
[0032] Step S4: Based on the contrast loss obtained from the event motion compensation and joint contrast model in step S2, perform multi-stage decoupled spatiotemporal joint optimization to obtain the calibration result; To address the parameter coupling and parameter absorption effects that occur in direct joint optimization under low-texture or non-overlapping field of view, for example, the optimizer compensates for time errors. However, the rotation extrinsic parameter matrix is incorrectly distorted. This leads to a physical deviation of more than ten degrees. In this embodiment, a three-stage decoupling optimization strategy is innovatively adopted.
[0033] Furthermore, the total loss function in the multi-stage decoupling joint optimization process is defined as:
[0034] in, For zero-partial regularization terms, This represents the zero-biased regularization weight coefficient; Indicates IMU zero bias; These are prior constraints on external parameters. Indicates the prior penalty term; , Indicates prior weights, Represents the rotation extrinsic parameter matrix of the camera; This represents the initial rotational extrinsic parameters from each event camera to the IMU.
[0035] It should be noted that the aforementioned contrast loss The calculation process directly depends on the time offset. With rotation extrinsic matrix At this stage, by minimizing the above total loss function This allows for iterative updates of spatiotemporal parameters in the active state using the gradient descent algorithm.
[0036] Step S4-1: Synchronize and align the data from all event cameras with the IMU data on the time axis; freeze the rotation extrinsic parameters of all event cameras. This locks it to the initial rotational extrinsic parameters of each event camera to the IMU. Prior weights Set to 0. Only set the time offset. and IMU zero bias We use learnable parameters for gradient descent optimization to minimize the total loss function. The forced system first uses high-frequency motion edges to align the time axis, preventing time errors from contaminating spatial parameters, thereby providing a precise time reference for subsequent external parameter refinement.
[0037] Step S4-2: Refine with prior extrinsic parameters. Freeze time bias shift. and IMU zero bias Release the rotation extrinsic matrix of the event camera. These are learnable parameters. The prior penalty term is then activated. It allows for precise sub-degree adjustments near the initial value, but severely punishes large drifts that cross physical limits.
[0038] Step S4-3: Global Joint Fine-Tuning. Release all learnable parameters. With a learning rate lower than that of the first two stages, the total loss function is minimized. Perform joint gradient descent optimization.
[0039] Specifically, within each optimization cycle, event data blocks are processed one by one, the gradient of the total loss function with respect to each learnable parameter is calculated, and the rotation extrinsic parameters are updated. Time offset and IMU zero bias The value is calculated until all event data blocks are processed; then the process returns to execute the next optimization cycle, until the preset maximum number of optimization cycles is reached.
[0040] Furthermore, during the final optimization cycle, the updated values of the spatiotemporal parameters to be optimized (i.e., rotational extrinsic parameters) are recorded after the model processes each event data block. Time offset IMU zero bias The arithmetic mean of the above parameters over the entire optimization period is calculated, and the final calibration result is used to eliminate the instantaneous oscillation of parameters caused by local extreme noise, so as to obtain the globally optimal and physically realistic spatiotemporal extrinsic parameters.
[0041] All the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of the present invention, and will not be described in detail here.
[0042] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A targetless spatiotemporal calibration method based on phased decoupling optimization, characterized in that, Includes the following steps: Acquire event stream data from synchronously acquired event cameras and angular velocity data from inertial measurement units (IMUs); adopt an adaptive slicing strategy based on a fixed number of events to divide the event stream data from multiple event cameras into multiple event data blocks with constant information density, and extract the IMU angular velocity data within the time period corresponding to the event data block to obtain multi-sensor data blocks; A differentiable event motion compensation and joint contrast model is constructed. Based on the joint contrast model, a phased decoupling optimization strategy is adopted to jointly optimize the spatiotemporal parameters of the targetless event camera and the IMU to obtain the calibration results. The spatiotemporal parameters include the rotation extrinsic parameters from each event camera to the IMU, the time offset, and the IMU zero offset. The specific steps of the phased decoupling optimization strategy include: Synchronize and align the data from all event cameras with the IMU data on the time axis; freeze the rotation extrinsic parameters from each event camera to the IMU, and perform a stage to optimize the time offset and IMU zero bias; The stage of freezing the time offset and optimizing the rotating extrinsic parameters with prior constraints; The stage of global joint fine-tuning of all spatiotemporal parameters.
2. The method according to claim 1, characterized in that, The specific steps for dividing the event stream data from multiple event cameras into multiple event data blocks with constant information density and extracting IMU angular velocity data within the time period corresponding to each event data block to obtain multi-sensor data blocks include: setting a fixed event threshold; when the number of events accumulated by the event camera designated as the main camera reaches the fixed event threshold, recording the corresponding time interval; dividing the continuous event stream data into multiple event data blocks according to the time interval, and extracting IMU angular velocity data within the time period corresponding to each event data block.
3. The method according to claim 1, characterized in that, The construction of the differentiable event motion compensation and joint contrast model includes: using B-splines combined with differentiable spherical linear interpolation to construct a continuous time trajectory model based on the angular velocity after IMU zero bias compensation; correcting the event timestamps according to the time offset, and using the continuous time trajectory model and the rotation extrinsic parameters to be optimized, mapping the event points to the reference time to obtain the compensated sub-pixel coordinates; generating event-compensated images based on the compensated sub-pixel coordinates of all event cameras, and calculating the negative value of the sum of the image variances of all event-compensated images as the contrast loss.
4. The method according to claim 1, characterized in that, In the phased decoupling optimization strategy: in the phase of optimizing the time offset and IMU zero bias, the weight of the prior constraint term is set to zero; in the phase of optimizing the rotation extrinsic parameters with prior constraints, the prior penalty term is activated, which is used to constrain the fine-tuning amplitude of the rotation extrinsic parameters near the initial value; in the global joint fine-tuning phase, all spatiotemporal parameters are jointly optimized with a learning rate lower than that of the first two phases.
5. The method according to claim 1, characterized in that, It also includes a targetless spatial extrinsic coarse initialization step: screening multi-sensor data blocks that meet the conditions of violent motion; for each event camera, finding the local visual angular velocity that maximizes the contrast of the event's accumulated image through an optimization algorithm; using modulus consistency constraints and variance gain constraints to screen effective angular velocity observation vector pairs, and obtaining the initial rotational extrinsic parameters from each event camera to the IMU by solving the Wahba problem.
6. The method according to claim 1, characterized in that, The total loss function of the phased decoupling optimization strategy includes contrast loss, IMU zero-bias regularization term, and prior constraint term of rotation extrinsic parameters; wherein, the calculation of contrast loss depends on time offset and rotation extrinsic parameters.
7. The method according to claim 1, characterized in that, The method further includes: in the last optimization cycle, recording the updated spatiotemporal parameter values after processing each event data block, and calculating the arithmetic mean of all spatiotemporal parameters in the optimization cycle as the final calibration result.
8. The method according to claim 1, characterized in that, The event camera data consists of data collected by at least two event cameras with non-overlapping fields of view; wherein one event camera is defined as the master camera and the others are slave cameras.