Method and system for calibrating any relative pose of laser radar and camera, and storage medium

By using a lightweight cylindrical calibration target and a spatiotemporal decoupling reconstruction strategy, combined with visual SFM and laser SLAM algorithms, the problem of high-precision calibration of LiDAR and camera without initial common field of view was solved, realizing a flexible and fast calibration process and improving calibration accuracy and robustness.

CN121767461APending Publication Date: 2026-03-31HANGZHOU NORMAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve high-precision relative pose calibration between lidar and camera without initial common field of view and hardware synchronization constraints. Furthermore, existing online calibration methods without calibration boards suffer from low accuracy and high environmental dependence.

Method used

Using a lightweight cylinder as the calibration target, and employing asynchronous data acquisition and spatiotemporal decoupling reconstruction strategies, combined with visual SFM and laser SLAM algorithms, images and point cloud data are independently reconstructed. Nonlinear optimization is performed by considering reprojection error and geometric consistency error to remove the constraints of hardware synchronization and fixed spatial pose.

Benefits of technology

It achieves high-precision calibration of lidar and camera under asynchronous conditions, reduces system complexity, and performs joint optimization in the global 3D scene, improving calibration accuracy and robustness, and is suitable for various environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767461A_ABST
    Figure CN121767461A_ABST
Patent Text Reader

Abstract

The invention discloses a laser radar and camera arbitrary relative pose calibration method and system and a storage medium. The method is characterized in that a calibration process is decoupled into two independent stages, namely a multi-sensor independent reconstruction stage and a map-level joint optimization stage. According to the method, a laser sequence and an image sequence are independently reconstructed to generate respective environment maps, and then a constraint relation is established on a three-dimensional map level, so that the dependence on time-space synchronization is fundamentally eliminated. The method can be applied to autonomous inspection, anomaly detection and intelligent early warning requirements proposed by development and application of the intelligent safety inspection robot for hazardous chemical substance storage safety, and tasks such as storage environment map construction, hazardous chemical substance leakage detection and storage environment risk level evaluation necessary for the inspection robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a method, system and storage medium for arbitrary relative pose calibration of LiDAR and camera. Background Technology

[0002] With the rapid development of autonomous driving and mobile robotics technologies, multi-sensor fusion has become a key technology for achieving environmental perception and navigation / localization. LiDAR provides precise 3D spatial information, while cameras capture rich texture and color information. Effectively fusing this information can fully leverage their respective advantages and enhance the system's perception capabilities. Accurate relative pose calibration between LiDAR and cameras is fundamental to multi-sensor fusion.

[0003] Traditional methods rely on hardware time synchronization and initial shared field of view (FLP) of sensors, which are difficult to meet in practical applications. Mobile robot sensor configuration often prioritizes perception coverage, resulting in the lack of an initial shared field of view; adding hardware synchronization increases cost and complexity. Existing online calibration methods without calibration boards suffer from low accuracy and high environmental dependence. Therefore, there is an urgent need for a practical calibration technique that can guarantee accuracy while overcoming spatiotemporal constraints. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method, system, and storage medium for calibrating the arbitrary relative pose of a lidar and a camera.

[0005] In a first aspect, the present invention provides a method for calibrating the arbitrary relative pose of a lidar and a camera, the method comprising the following steps:

[0006] Step S1: Use a lightweight cylinder with diffuse surface reflection as a calibration target and pre-calibrate the camera's intrinsic parameters;

[0007] Step S2: Keep the relative pose of the lidar and the camera fixed. Under the condition of asynchronous acquisition, move the lightweight cylinder so that it appears in both the camera's field of view and the lidar's scanning range at the same time, and acquire no less than fifteen sets of images and point cloud data in different spatial poses.

[0008] Step S3: Perform spatiotemporally decoupled independent reconstruction of the obtained image sequence and point cloud sequence: use the visual SFM algorithm to generate sparse 3D point cloud and recover the camera pose of each frame image, and use the laser SLAM algorithm to construct a laser global 3D map and obtain the laser pose of each frame point cloud.

[0009] Step S4: Extract the 2D projected straight line of the lightweight cylinder from the reconstructed image, and extract the 3D axis of the lightweight cylinder from the reconstructed laser point cloud.

[0010] Step S5: Using the 2D projection line and 3D axis as constraints, construct the point cloud reprojection error and geometric consistency error, and solve the external parameter rotation matrix and translation vector between the lidar and the camera through joint nonlinear optimization to complete the arbitrary relative pose calibration.

[0011] In a second aspect, the present invention provides a system for calibrating arbitrary relative poses of a lidar and a camera, comprising a memory and a processor, wherein the memory stores a program for a method of calibrating arbitrary relative poses of a lidar and a camera, and the method of calibrating arbitrary relative poses of a lidar and a camera, when executed by the processor, implements the steps described in the first aspect of the present invention.

[0012] In a third aspect, the present invention provides a computer-readable storage medium storing a program for calibrating an arbitrary relative pose of a lidar and a camera, wherein when the program is executed by a processor, it implements the steps of a method for calibrating an arbitrary relative pose of a lidar and a camera.

[0013] Compared with existing technologies, the beneficial effects of this invention are as follows: It completely removes the rigid constraints of hardware time synchronization and fixed sensor spatial pose, allowing the LiDAR and camera to acquire data under asynchronous and freely changing relative pose conditions. This not only significantly reduces system complexity but also makes it possible to perform rapid and flexible calibration of online and in-service systems. Furthermore, by introducing a spatiotemporal decoupling reconstruction strategy for SFM and LiDAR SLAM, joint optimization is achieved within a globally consistent 3D scene model, effectively overcoming the noise and viewpoint limitations of single-frame data, ultimately achieving a significant improvement in calibration accuracy and overall robustness. Another outstanding advantage lies in its strong versatility; it can achieve extremely high accuracy using dedicated calibration targets and can also adapt to natural features in the scene for calibration, expanding the application boundaries of the technology. Attached Figure Description

[0014] Figure 1 This application provides a schematic flowchart of a laser camera calibration technique based on spatiotemporal decoupling for arbitrary relative poses.

[0015] Figure 2 This application provides a schematic diagram of the apparatus in an embodiment. Detailed Implementation

[0016] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0017] This application provides a method for arbitrary relative pose calibration of a laser camera based on spatiotemporal decoupling. This method can be applied to the autonomous inspection, anomaly detection, and intelligent early warning requirements of the development and application of intelligent safety inspection robots for hazardous chemical storage. It addresses the essential tasks of inspection robots, such as building storage environment maps, detecting hazardous chemical leaks, and assessing storage environment risk levels.

[0018] The core of this application lies in decoupling the calibration process into two independent stages: a multi-sensor independent reconstruction stage and a map-level joint optimization stage. Laser sequences and image sequences are reconstructed independently to generate their respective environmental maps. Constraints are then established at the 3D map level, fundamentally eliminating the reliance on spatiotemporal synchronization.

[0019] The method includes the following steps:

[0020] Step 1: Use a lightweight cylinder as the calibration target. Select a lightweight cylinder with good surface diffuse reflectivity. The diameter and height of the cylinder need to be precisely known. For ease of image recognition, it can be placed in front of a monochrome background or a background with high contrast.

[0021] Choose a feature-rich, well-lit indoor or semi-outdoor environment. Avoid direct sunlight or excessive darkness. Place one or more lightweight cylinders in the scene as calibration targets. The cylinder surfaces should be made of a diffuse reflective material to facilitate laser reflection and image feature extraction. Ensure there is sufficient space around the calibration targets to allow the equipment to observe them from different angles and distances.

[0022] Step 2: Fix the lidar and camera, ensuring that their relative positions are strictly fixed during the calibration process. The camera's focal length, principal point, distortion coefficient, and other intrinsic parameters need to be pre-calibrated to complete the sensor system initialization.

[0023] The lidar and camera are fixedly mounted on the work platform. There is no need to precisely calibrate their relative pose to record the camera's intrinsic parameters. First, traditional calibration is performed using a checkerboard pattern to check that the sensor is working properly and set appropriate parameters.

[0024] Step 3: In this step, with asynchronous acquisition by the camera and LiDAR, the cylinder is held or moved so that it simultaneously appears within the camera's field of view and the LiDAR's scanning range. Multiple sets of data, including image data and point cloud data, are acquired at different positions and orientations, and the acquired data are independently reconstructed under spatiotemporal decoupling.

[0025] For each set of data, it is necessary to simultaneously acquire one image from the camera and one point cloud frame from the laser camera. It is necessary to ensure accuracy by acquiring 15-30 poses. The cylinders should be distributed as evenly as possible within the field of view, covering the entire field of view and different depths, and their poses should also be diverse.

[0026] like Figure 2 As shown, the data acquisition process requires an ideal motion trajectory, which should ensure that:

[0027] Orbital observation: Orbit the cylinder at least 180 degrees around the calibration object and observe it from different angles.

[0028] Variable distance observation: While circling, the distance between the platform and the calibration object is changed, such as from near to far, and then from far to near.

[0029] Multiple perspectives: Try having the sensor observe the calibration object from different heights (such as eye level, top view, and bottom view).

[0030] Environmental coverage: During movement, ensure that the sensor can also see the environmental features around the calibration object (such as walls, corners, other stationary objects), which will provide key scene constraints for SLAM and SFM reconstruction.

[0031] The movement should be smooth and slow, avoiding violent shaking or high-speed movement to prevent image blurring and laser point cloud distortion.

[0032] Throughout the acquisition process, it is essential to ensure that the target object remains simultaneously within the field of view of both the lidar and the camera. This is crucial for establishing a correlation between the two; brief obstructions are tolerable, but prolonged loss of the target will negatively impact reconstruction.

[0033] The independent reconstruction process under spatiotemporal decoupling is as follows:

[0034] The Lego-LOAM algorithm in laser SLAM is used to perform motion estimation and map construction on the point cloud sequence to generate a high-precision 3D scene model and obtain the pose of each frame of point cloud in the laser coordinate system. Visual SFM is used to extract features, match and perform motion recovery structure processing on the image sequence to generate a sparse 3D point cloud model and recover the camera pose of each frame of image.

[0035] Step 4: Perform data processing and feature extraction on the reconstructed image. Process the image to extract the 2D projection lines of the cylinder, and process the point cloud to extract the 3D axis of the cylinder.

[0036] Furthermore, step four involves processing the image to extract the 2D projection lines of the cylinder, including the following steps:

[0037] 1. Image preprocessing improves image quality, suppresses noise, enhances target features, and lays the foundation for edge detection.

[0038] 2. Use the Canny edge detector to find pixels in the image with drastic brightness changes.

[0039] 3. Outline search and cylinder outline filtering: Find the outlines that are most likely to belong to the left and right edges of the cylinder from the edge map.

[0040] 4. By using RANSAC line fitting, the selected set of left and right contour pixels is fitted into an accurate straight line equation.

[0041] 5. Calculate the 2D projection midline using the midpoint fitting method.

[0042] Furthermore, step four involves processing the point cloud to extract the 3D axis of the cylinder, including the following steps:

[0043] 1. Preprocess the point cloud data by voxel grid downsampling and statistical outlier removal to improve data quality and reduce the burden on subsequent processing.

[0044] 2. The cylindrical point cloud is separated from the background (such as the ground, walls, and other objects) using Euclidean clustering. First, a KD-Tree is constructed for the point cloud to accelerate nearest neighbor search. Then, each point is traversed, and its distance is compared with its neighbors. If the distance is less than a threshold, they are merged into the same cluster. Finally, iterative growth is performed to ensure that every point is processed. The threshold is a clustering distance threshold, which should be greater than the maximum gap between points on the cylindrical surface but less than the minimum distance between the cylinder and the background objects.

[0045] 3. RANSAC coarse fitting is used to robustly estimate the initial values ​​of the cylinder parameters, providing a good starting point for subsequent fine fitting. Then, nonlinear least squares method is used for fine fitting.

[0046] In this embodiment, the Euclidean clustering segmentation core adopts the KD-Tree data structure, which constructs a dataset from point clouds. A KD-Tree data structure is constructed for the input point cloud. By recursively partitioning the space, an efficient tree-like index structure is established, which can significantly accelerate nearest neighbor search and radius search operations.

[0047] The point cloud axis extraction process in this embodiment, through a standardized pipeline of "preprocessing-segmentation-fitting" and the core strategy of "RANSAC robust initialization + nonlinear optimization refinement", ensures that the three-dimensional geometric information of the calibration object can be obtained stably and with high accuracy even in complex real-world environments, laying a reliable data foundation for the entire spatiotemporal decoupling calibration system.

[0048] Step 5: Define reprojection error and geometric consistency error, and perform nonlinear optimization on the final result.

[0049] The reprojection error is defined based on constraints of the camera imaging model to ensure alignment between the laser point cloud and the image texture. This error measures the distance between a 3D point in the LiDAR coordinate system, projected onto the 2D image plane using the currently estimated extrinsic parameters and known camera intrinsic parameters, and the corresponding feature point actually observed in the image.

[0050] One advantage of calculating this error is that it forces the projection of the laser point cloud to align with the rich texture features of the image, which is the fundamental goal of data fusion. Another advantage of calculating this error is that each 3D point reconstructed by SFM can provide such an error term, contributing a massive number of constraints to the optimization problem.

[0051] Point cloud reprojection error is based on a 3D point in the laser point cloud. Camera intrinsic matrix parameter K, extrinsic rotation matrix to be optimized With translation vector The two-dimensional feature points corresponding to the three-dimensional point in the image are obtained by reconstructing the three-dimensional point using the structure for motion reconstruction (SFM). These four types of data can be calculated, where the subscript L indicates that the 3D point is in the lidar coordinate system. yes The coordinates of this point in the lidar coordinate system for The coordinates of this point in the image pixel coordinate system.

[0052] In a preferred embodiment, the point cloud reprojection error calculation process is as follows:

[0053] 1. Transformation from laser coordinate system to camera coordinate system

[0054] .

[0055] 2. Projected onto the image normalization plane

[0056]

[0057] 3. Obtain the pixel coordinates of the reprojection.

[0058]

[0059] Calculated

[0060] 4. Calculation error

[0061]

[0062] Furthermore, the geometric consistency error is calculated, which is the degree of geometric fit between the calibration lightweight cylinder fitted from the laser point cloud and the 3D point cloud corresponding to the same calibration object reconstructed from the image sequence via SFM.

[0063] One advantage of calculating this error is that it does not depend on specific pixels, but rather directly constrains the overall geometry. Even if there are minor errors in the absolute positions of the SFM reconstructed points, the overall shape must still conform to the laser model.

[0064] Another advantage of calculating this error is that it is insensitive to the feature extraction error of the calibration object: even if a feature point extracted from the image is deviated, as long as the overall shape of the point cloud reconstructed by SFM is correct, this error can still provide the correct constraint direction.

[0065] Another advantage of calculating this error is its very strong constraint: a cylindrical model (5 parameters) can simultaneously impose constraints on hundreds or thousands of SFM point clouds, greatly enhancing the convexity of the optimization problem and helping the algorithm escape local optima.

[0066] Geometric consistency error is based on the parameters of the cylindrical model fitted from the laser map: a point on the axis Unit vector along the axis ,radius A set of three-dimensional points belonging to the surface of the cylinder obtained from the SFM map. The extrinsic rotation matrix to be optimized With translation vector These three sets of data can be calculated.

[0067] In a preferred embodiment, the geometric consistency error calculation process is as follows:

[0068] 1. Reconstruct the points from SFM The coordinates are transformed from the camera coordinate system to the laser coordinate system for comparison with the laser cylindrical model.

[0069]

[0070] 2. For each point Calculate its distance from the laser cylinder axis.

[0071] vector

[0072] The vector is in the direction of the axis. The projection scalar on is

[0073] The foot of the perpendicular from the point to the axis is

[0074] The distance from the point to the axis is

[0075] 3. Finally, the error was calculated to be...

[0076] Finally, the above errors are combined with nonlinear optimization, and the overall objective function is as follows:

[0077]

[0078] in: For reprojection error, For geometric consistency error, M is the number of laser points available for reprojection, N is the number of calibration point clouds reconstructed by SFM, and λ is a weighting factor used to balance the magnitudes of the two error terms to ensure they contribute equally to the optimization.

[0079] The joint optimization in this application has the following advantages:

[0080] Complementarity: Reprojection error has a strong constraint in areas with rich texture, but may fail in areas with weak texture. Geometric consistency error only cares about the macroscopic shape of the object and is independent of texture. The combination of the two ensures that a strong constraint is in effect in any scenario.

[0081] Robustness: If some SFM reconstruction points are inaccurate, their reprojection error will be large, but it may not affect the overall geometric consistency error. The optimizer integrates all information and automatically reduces the impact of unreliable data, thus obtaining more stable results.

[0082] This application embodiment also provides a lidar and camera arbitrary relative pose calibration system, including a memory and a processor. The memory stores a lidar and camera arbitrary relative pose calibration method program. When the lidar and camera arbitrary relative pose calibration method is executed by the processor, it implements the above method steps.

[0083] This application embodiment also provides a computer-readable storage medium storing a program for calibrating an arbitrary relative pose of a lidar and a camera. When the program is executed by a processor, it implements the steps of the method.

[0084] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0085] Alternatively, if this application is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0086] It should be understood that the specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art can make various interpretations, including modifications, combinations, and substitutions, based on actual design considerations and other factors. All such modifications within the spirit and principles of this application fall within the scope of protection claimed in this application.

Claims

1. A method for calibrating arbitrary relative poses between a lidar and a camera, characterized in that The method comprises the following steps: Step S1, a light-weight cylinder with a diffusely reflective surface is used as a calibration target, and the intrinsic parameters of the camera are calibrated in advance; Step S2, the relative poses of the laser radar and the camera are fixed, under the condition of asynchronous acquisition of the two, the light-weight cylinder is moved so as to simultaneously appear in the field of view of the camera and the scanning range of the laser radar, and at least fifteen groups of images and point cloud data are acquired at different spatial poses; Step S3, the obtained image sequence and point cloud sequence are independently reconstructed by space-time decoupling: a sparse three-dimensional point cloud is generated by using a visual SFM algorithm, and the camera poses of each frame of image are recovered, and a laser global three-dimensional map is constructed by using a laser SLAM algorithm, and the laser poses of each frame of point cloud are acquired; Step S4, a 2D projection straight line of the light-weight cylinder is extracted from the reconstructed image, and a 3D axis of the light-weight cylinder is extracted from the reconstructed laser point cloud; Step S5, the 2D projection straight line and the 3D axis are used as constraints to construct a point cloud re-projection error and a geometric consistency error, and a nonlinear optimization is used to solve the extrinsic rotation matrix and the translation vector between the laser radar and the camera, and the relative pose calibration is completed.

2. The method of claim 1, wherein: The diameter and height of the light-weight cylinder are accurately known in advance, and the surface is a uniform diffusely reflective material.

3. The method of claim 1, wherein: In step S2, the light-weight cylinder is distributed in the entire field of view and at different depths during the acquisition process, and the attitude angle changes by not less than ±30°.

4. The method of claim 3, wherein: When the images and the point cloud are acquired, the sensor platform completes a ≥180° surrounding trajectory around the light-weight cylinder, and simultaneously realizes multi-view variable-distance observation from near to far and from far to near, so as to ensure that the reconstructed point cloud covers the cylinder in all directions.

5. The method of any one of claims 1 to 4, wherein: The laser SLAM algorithm in step S3 is a Lego-LOAM algorithm, and the visual SFM algorithm is an incremental motion structure recovery algorithm.

6. The method of claim 5, wherein: The point cloud re-projection error is defined as follows: the three-dimensional points in the laser coordinate system are transformed into the camera coordinate system through the to-be-optimized extrinsic parameters, projected onto the image plane through the camera intrinsic parameters, and the Euclidean distance between the corresponding two-dimensional feature points obtained by the SFM reconstruction is calculated.

7. The method of claim 6, wherein: The geometric consistency error is defined as follows: the three-dimensional points of the cylinder surface reconstructed by the SFM are transformed into the laser coordinate system, the perpendicular distance of each point to the laser-fitted cylinder axis is calculated, and the square sum of the difference between the perpendicular distance and the cylinder radius is calculated.

8. The method of claim 7, wherein: The objective function of the joint nonlinear optimization is as follows: where: is the point cloud reprojection error, is the geometric consistency error, M is the number of laser points available for reprojection, N is the number of object point cloud points reconstructed by SFM, is a weight factor to balance the magnitude of the two error terms, ensuring that they contribute equally in the optimization. 9.A system for calibrating arbitrary relative poses between a lidar and a camera, the system comprising: A memory and a processor are included, and a laser radar and camera relative pose calibration method program is stored in the memory, and the laser radar and camera relative pose calibration method is executed by the processor to realize the following steps: Step S1, a light-weight cylinder with a diffusely reflective surface is used as a calibration target, and the intrinsic parameters of the camera are calibrated in advance; Step S2, the relative poses of the laser radar and the camera are fixed, under the condition of asynchronous acquisition of the two, the light-weight cylinder is moved so as to simultaneously appear in the field of view of the camera and the scanning range of the laser radar, and at least fifteen groups of images and point cloud data are acquired at different spatial poses; Step S3, independently reconstructing the obtained image sequence and point cloud sequence respectively by spatio-temporal decoupling: generating a sparse three-dimensional point cloud and restoring the camera pose of each frame of image by using a visual SFM algorithm, and constructing a laser global three-dimensional map and obtaining the laser pose of each frame of point cloud by using a laser SLAM algorithm; Step S4, extracting the 2D projection straight line of the lightweight cylinder in the reconstructed image and the 3D axis of the lightweight cylinder in the reconstructed laser point cloud; Step S5, taking the 2D projection straight line and the 3D axis as constraints, constructing a point cloud re-projection error and a geometric consistency error, jointly solving the extrinsic rotation matrix and the translation vector between the laser radar and the camera by nonlinear optimization, and completing the arbitrary relative pose calibration.

10. A computer-readable storage medium, characterized in that, The medium stores a laser radar and camera arbitrary relative pose calibration method program, and when the laser radar and camera arbitrary relative pose calibration method program is executed by the processor, the steps of the laser radar and camera arbitrary relative pose calibration method in any one of claims 1 to 8 are realized.