External parameter calibration method, device and equipment for laser radar and camera, and medium
By dividing the calibration space into sub-space regions, using lidar and cameras to acquire features and perform joint optimization, the problems of limited field of view coverage and inconsistent calibration accuracy in the depth direction of single target calibration are solved, achieving high-precision extrinsic parameter calibration, which is suitable for autonomous driving and robot navigation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG GAOYU TECHNOLOGY CO LTD
- Filing Date
- 2025-12-23
- Publication Date
- 2026-05-01
AI Technical Summary
Existing LiDAR and camera extrinsic parameter calibration methods have limited field-of-view coverage in single-target calibration and inconsistent calibration accuracy in the depth direction, making it difficult to meet the high-precision calibration requirements in complex scenarios.
Multiple planar reference objects at different distances are set in the calibration space, which are divided into multiple sub-space regions. Point cloud features are obtained by LiDAR and image features are obtained by camera. Combined with the initial extrinsic parameters, joint optimization is performed to obtain high-precision target extrinsic parameters.
It significantly improves calibration accuracy and robustness, making it suitable for applications requiring precise multi-sensor fusion, such as autonomous driving and robot navigation.
Smart Images

Figure CN121962282A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-sensor fusion technology, and in particular to a method, apparatus, device, and medium for calibrating the extrinsic parameters of a lidar and camera. Background Technology
[0002] In fields such as intelligent driving, mobile robots, and aircraft, multi-sensor collaborative operation has become a core technological approach for achieving environmental perception and autonomous decision-making. LiDAR and cameras are core sensors for perceiving the environment, and they are often used in combination to achieve high-precision environmental modeling and target recognition.
[0003] The extrinsic parameters (coordinate transformation relationship) of lidar and camera are the core factors that determine the fusion perception accuracy of lidar and camera. Current extrinsic parameter calibration methods generally rely on manually designed calibration boards or specific scene features, and are mostly limited to parameter solving within a single distance range. The calibration field of view coverage of a single target is limited, and the calibration accuracy in the depth direction is inconsistent, making it difficult to meet the high-precision calibration requirements in complex scenes. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and medium for calibrating the extrinsic parameters of a lidar and a camera, in order to solve the problems of limited calibration field of view coverage and inconsistent calibration accuracy in the depth direction for a single target.
[0005] A method for extrinsic parameter calibration of a lidar and a camera, wherein the lidar and the camera are jointly positioned in a calibration space at a calibration location for extrinsic parameter calibration, and planar reference objects are respectively set at multiple different distances from the calibration location in the calibration space. The extrinsic parameter calibration method includes the following steps: dividing the calibration space into multiple sub-space regions corresponding to different calibration distance intervals according to the calibration location; acquiring point cloud features of the planar reference objects in each sub-space region through the lidar, and acquiring image features of the planar reference objects in each sub-space region through the camera; and obtaining target extrinsic parameters between the lidar and the camera based on preset initial extrinsic parameters and combining the point cloud features and the image features.
[0006] An extrinsic parameter calibration device for a lidar and a camera, wherein the lidar and the camera are jointly disposed in a calibration space at a calibration position for extrinsic parameter calibration, and planar reference objects are respectively disposed at multiple different distances from the calibration position in the calibration space. The extrinsic parameter calibration device includes: a spatial partitioning module, used to divide the calibration space into multiple sub-space regions corresponding to different calibration distance intervals according to the calibration position; a feature acquisition module, used to acquire point cloud features of the planar reference objects in each of the sub-space regions through the lidar and to acquire image features of the planar reference objects in each of the sub-space regions through the camera; and an extrinsic parameter acquisition module, used to acquire target extrinsic parameters between the lidar and the camera based on preset initial extrinsic parameters and combining the point cloud features and the image features.
[0007] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned extrinsic parameter calibration method for a lidar and a camera.
[0008] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for extrinsic parameter calibration of a lidar and camera.
[0009] In the aforementioned technical solution for extrinsic parameter calibration of LiDAR and camera, the LiDAR and camera are jointly positioned in a calibration space at a calibration location for extrinsic parameter calibration. Planar reference objects are positioned at multiple different distances from the calibration location within the calibration space. The extrinsic parameter calibration method includes the following steps: dividing the calibration space into multiple sub-space regions corresponding to different calibration distance intervals based on the calibration location; acquiring point cloud features of the planar reference objects in each sub-space region using the LiDAR, and acquiring image features of the planar reference objects in each sub-space region using the camera; and obtaining the target extrinsic parameters between the LiDAR and camera based on preset initial extrinsic parameters, combined with the point cloud features and image features. This method can obtain high-precision target extrinsic parameters, significantly improving calibration accuracy and robustness, and is particularly suitable for applications requiring precise multi-sensor fusion, such as autonomous driving and robot navigation. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1This is a flowchart of an external parameter calibration method for a lidar and a camera according to an embodiment of the present invention; Figure 2 This is a calibration scenario diagram of the external parameter calibration method for lidar and camera in one embodiment of the present invention; Figure 3 This is a detailed flowchart of step S1 in the external parameter calibration method for lidar and camera in one embodiment of the present invention; Figure 4 This is a result diagram of step S11 in the external parameter calibration method for lidar and camera in one embodiment of the present invention; Figure 5 This is a detailed flowchart of step S4 in the external parameter calibration method for lidar and camera in one embodiment of the present invention; Figure 6 This is a schematic diagram of a reference point cloud obtained in step S21 of the external parameter calibration method for lidar and camera in one embodiment of the present invention; Figure 7 This is a schematic diagram of a point cloud feature obtained in step S22 of the external parameter calibration method for lidar and camera in one embodiment of the present invention; Figure 8 This is a detailed flowchart of step S3 in the external parameter calibration method for lidar and camera in one embodiment of the present invention; Figure 9 This is a detailed flowchart of step S32 in the external parameter calibration method for lidar and camera in one embodiment of the present invention; Figure 10 This is a schematic diagram illustrating the principle of calibration error evaluation in the external parameter calibration method for lidar and camera according to an embodiment of the present invention; Figure 11 This is a schematic diagram of an external parameter calibration device for a lidar and camera according to an embodiment of the present invention; Figure 12 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] This invention provides a method, apparatus, device, and medium for extrinsic parameter calibration of a lidar and a camera. The lidar and camera are jointly positioned in a calibration space at a calibration location for extrinsic parameter calibration. Planar reference objects are positioned at multiple different distances from the calibration location within the calibration space. The extrinsic parameter calibration method includes: dividing the calibration space into multiple sub-space regions corresponding to different calibration distance intervals based on the calibration location; further dividing the calibration space into multiple sub-space regions corresponding to different calibration distance intervals based on the calibration location; acquiring point cloud features of the planar reference objects in each sub-space region using the lidar, and acquiring image features of the planar reference objects in each sub-space region using the camera; and obtaining the target extrinsic parameters between the lidar and the camera based on preset initial extrinsic parameters, combined with the point cloud features and image features. This method, by acquiring the correspondence between point cloud and image features in multiple sub-space regions and iteratively solving using a joint optimization algorithm, can obtain high-precision target extrinsic parameters, significantly improving calibration accuracy and robustness. It is particularly suitable for applications requiring precise multi-sensor fusion, such as autonomous driving and robot navigation.
[0014] First, it's important to clarify that the coordinate system used by the LiDAR is the LiDAR coordinate system. Specifically, it's a three-dimensional coordinate system established with the LiDAR's optical transceiver as the origin and the LiDAR's optical axis as the z-axis; it's also a right-handed coordinate system. The camera's coordinate system is the camera coordinate system. Specifically, it's a three-dimensional coordinate system established with the camera's optical center as the origin and the camera's optical axis as the z-axis; it's also a right-handed coordinate system. Extrinsic parameter calibration involves solving for the rotation matrix and translation vector between the two coordinate systems, ensuring that the LiDAR point cloud data can be accurately projected onto the camera's imaging plane.
[0015] In one embodiment, such as Figure 1 As shown, a method for calibrating the extrinsic parameters of a lidar and a camera is provided, including the following steps: Step S1: Based on the calibration location, divide the calibration space into multiple sub-space regions corresponding to different calibration distance intervals.
[0016] It should be noted that the calibration location refers to the common coordinate origin or installation reference point of the lidar and camera. Each planar reference object is a target with a known geometric structure and high reflectivity, distributed across near, medium, and far distances. The calibration distance refers to the vertical distance from the calibration location to each planar reference object. Based on this distance, the calibration space is divided into multiple sub-regions, each corresponding to a different calibration distance interval, to achieve refined regional calibration and segmented modeling and optimization of extrinsic parameter errors at different calibration distances. Each sub-region contains at least one planar reference object to ensure the integrity of multimodal data acquisition and the accuracy of spatial correspondence.
[0017] like Figure 2As shown, a specific calibration scenario of this embodiment is illustrated. The lidar and camera are jointly set in the calibration space at calibration position 1 for external parameter calibration. Planar reference objects 10 are respectively set at multiple different distances from calibration position 1 in the calibration space. The planar reference objects 10 are symmetrically distributed to enhance calibration stability.
[0018] In this embodiment, by acquiring a global image of the calibration space, dividing the global image into multiple local images, and based on each local image, dividing the calibration space into multiple sub-space regions, with each local image corresponding to a sub-space region, the accurate identification and matching of planar reference objects within different distance ranges can be achieved.
[0019] Specifically, such as Figure 3 As shown, step S1 includes the following sub-steps: Step S11: Obtain the global image of the calibration space and divide the global image to obtain multiple local images corresponding to different calibration distance intervals.
[0020] It should be noted that the global image is a complete field-of-view image covering the entire calibration space, synchronously acquired by a camera or global image acquisition device at the calibration location. It contains imaging information of each planar reference object at different distances. The division of local images is based on the spatial distribution and depth information correspondence of each planar reference object in the global image. The division boundaries of local images match the calibration distance range, ensuring that each local image only covers planar reference objects within a specific distance range, avoiding matching ambiguities caused by overlapping of near and far targets.
[0021] In this application, the global image is first divided into grid cells to obtain multiple grid units. Then, based on the correspondence between the depth information of each planar reference object and the calibration distance interval, each grid unit is assigned to the local image corresponding to the calibration distance interval, ensuring that each local image only contains the imaging area of the planar reference object within the corresponding distance range. This partitioning method can effectively reduce cross-distance interference, improve feature matching accuracy, and provide clear spatial constraints for subsequent region-based optimization.
[0022] like Figure 4As shown, using the image coordinate system (with the camera optical center as a reference), the global image is first divided into a 3×3 grid using two horizontal and two vertical dividing lines, resulting in nine equally divided local sub-images (nine equally sized rectangular sub-regions in the image). The horizontal coverage angle of the grid is ≥120°, and the vertical coverage angle is ≥60°, ensuring that each local sub-image covers the image center and the four corners of the image, achieving "full-area coverage without blind spots." Then, each local sub-image is mapped to its corresponding calibration distance range based on the depth information of the planar reference objects it contains. The local sub-images are then merged according to the calibration distance range to obtain three local images (obtained by dividing the global image using two horizontal dividing lines). Each local image contains the visual feature information of the planar reference objects within the corresponding distance range in the calibration space. From bottom to top, the three local images are a near-range local image, a medium-range local image, and a far-range local image, corresponding to the three calibration distance intervals of near, medium, and far distances, respectively. The sub-space region covered by each local image expands as the calibration distance increases to accommodate the differences in the imaging resolution of the planar reference object at different calibration distances.
[0023] In this application, a nine-grid arrangement is used to cover the entire image, eliminating calibration errors at the field of view edges. Experiments show that far-field calibration errors can be reduced by 30%-50%. Furthermore, the standardized nine-grid arrangement eliminates the need for precise control of the relative positions between targets, reducing implementation difficulty and cost. Through this image segmentation strategy, the calibration extrinsic parameters between the camera and LiDAR can be specifically optimized for each calibration distance interval, improving the accuracy of cross-modal data fusion. This is particularly effective in significantly enhancing the alignment between point clouds and images in distant, weakly reflective regions, providing high-quality input for subsequent feature extraction and joint calibration.
[0024] Step S12: Divide the calibration space into regions based on each local image to obtain the subspace regions corresponding to each local image.
[0025] In this embodiment, three local images are mapped to a calibration space to obtain three sub-space regions with calibration distance ranges of near, medium, and far. The near-range local image corresponds to a calibration distance range of 1-3 meters, with a straight-line distance of 1-3 meters between the planar reference object and the camera's optical center, providing high-precision point cloud and image features. The medium-range local image corresponds to a calibration distance range of 4-7 meters, with a straight-line distance of 4-7 meters between the planar reference object and the camera's optical center, providing clear feature imaging with moderate coverage, balancing point cloud density and field of view coverage. The far-range local image corresponds to a calibration distance range of 8-15 meters, with a straight-line distance of 8-15 meters between the planar reference object and the camera's optical center, effectively capturing sparse point cloud distribution and low-resolution image features at long distances, optimizing long-distance projection accuracy.
[0026] like Figure 2 As shown, each subspace region expands in a gradient along the depth direction. The near-range subspace region 11 has the smallest volume but the highest feature density; the mid-range subspace region 12 has a moderate volume and balanced coverage; and the far-range subspace region 13 has the largest volume (lateral span) but sparse point cloud and low image resolution. Within each subspace region, three planar reference objects are placed near the center and left and right boundaries to ensure coverage of the field of view within that distance range. Each planar reference object has known geometric dimensions and high-contrast texture, facilitating image feature extraction and point cloud matching. Furthermore, through multi-view observation and joint optimization algorithms, the camera-LiDAR extrinsic parameters of the near-range, mid-range, and far-range subspace regions are independently calibrated, improving the alignment accuracy of each calibration distance range, especially achieving smooth transitions across distances, and enhancing the overall calibration stability and robustness. The transition zone between each subspace region achieves continuous interpolation of extrinsic parameters through a weighted fusion strategy, effectively suppressing boundary interruptions that may be introduced by segmented calibration, ensuring calibration consistency in the depth direction, and solving the problem of calibration accuracy decaying with distance in traditional methods.
[0027] Furthermore, during actual scene setup, laser rangefinders can be used to precisely calibrate the positions of each planar reference object, ensuring that it is located within the corresponding sub-space region, thereby guaranteeing the accuracy of matching local images with the calibrated distance range. For planar reference objects, calibration boards with regular geometric patterns such as high-contrast checkerboards or QR codes can be selected, as their corner or edge features are easy to extract and have high positioning accuracy, which can effectively improve the consistency of image and point cloud matching.
[0028] In this application, the aforementioned partitioning method not only achieves ordered zoning of the calibration space but also enhances the accuracy and matching efficiency of feature extraction from planar reference objects at different calibration distances, providing a refined spatial distribution basis for subsequent cross-modal data alignment. This distance-driven region partitioning mechanism essentially reflects the adaptive logic of the perception system in a multi-scale environment, revealing a cognitive evolution process from global to local and from coarse to fine.
[0029] In other embodiments, the partitioning of the calibration space can be extended to more distance levels to meet the dynamic range requirements of practical application scenarios, thereby improving adaptability in complex environments. The division of the calibration space can also be dynamically adjusted based on the geometric distribution density of the planar reference object, that is, a finer-grained subspace region division is used in high-density areas, while adjacent subspace regions are merged in low-density areas to reduce redundant calculations.
[0030] In other embodiments, such as for long-range flight missions of aircraft, the subspace region can be extended to ultra-long-range intervals exceeding 50 meters from the calibration distance, such as 50-100 meters, 100-150 meters, and 150-200 meters, to meet the needs of capturing sparse features in a large airspace. In this case, the division of local images is no longer confined to a fixed grid, but rather adaptively cropped based on the aircraft's attitude and trajectory to ensure that each subspace region always covers the key reference object. The size and material of the planar reference object can also be differentiated. For size differentiation, the size gradually increases from near to far (e.g., from 2m×2m to 6m×6m; or for every 50m increase, the size of the planar reference object increases by 50%) to enhance recognizability under long-range imaging. Regarding materials, high reflectivity or materials with strong texture contrast are selected based on the ambient lighting characteristics, such as high-contrast diffuse reflection coatings and high-reflectivity stripes, to ensure that the sensor can still stably extract features at long distances. The above design not only improves the robustness of feature recognition across distance scales, but also enhances the system's adaptability in complex lighting and dynamic environments.
[0031] Step S2: Obtain the point cloud features of the planar reference objects in each sub-space region using LiDAR, and obtain the image features of the planar reference objects in each sub-space region using a camera.
[0032] It should be noted that point cloud features are the three-dimensional spatial distribution information of the reference point cloud corresponding to a planar reference object, including its fitted plane equation and geometric attributes such as the coordinates of three-dimensional feature points. Image features, on the other hand, correspond to the two-dimensional projection information of the planar reference object in the reference image, including the coordinates of two-dimensional feature points, edge contours, and texture features. Although their representations differ in their respective modes, they both originate from the geometric and appearance attributes of the same physical entity. After the LiDAR and camera are fixed in their calibration positions, it is necessary to ensure that they synchronously acquire point cloud data and image data to avoid spatial registration errors caused by time asynchrony and to ensure the spatiotemporal consistency of data acquisition.
[0033] like Figure 5 As shown, step S2 includes the following sub-steps: Step S21: Align the lidar and camera with each subspace region in turn, and simultaneously acquire the reference point cloud and reference image of the planar reference object in each subspace region.
[0034] In this embodiment, by aligning with each sub-space region, the LiDAR scans the surface of the planar reference object at high frequency to acquire dense and accurate reference point clouds, focusing on capturing its geometric contours and spatial pose information; the camera simultaneously captures high-resolution texture images, recording the visual features and illumination response characteristics of the planar reference object. To ensure the alignment accuracy of multimodal data, nanosecond-level time synchronization is achieved based on a hardware triggering mechanism, and preliminary spatial registration is performed in conjunction with the pose prior of the calibration position. After each sub-space region independently completes data acquisition, it forms a one-to-one corresponding reference point cloud and reference image, which serve as the basic input for subsequent cross-modal matching and joint optimization. During the acquisition process, the joint calibration of the LiDAR point cloud and camera image is further refined through ICP (Iterative Closest Point) algorithm and SIFT feature matching to eliminate minor pose deviations. To address the issues of point cloud sparsity and image resolution degradation at long distances, super-resolution reconstruction and point cloud completion network preprocessing data are introduced to improve feature integrity. The edge contours and corners of the reference object are emphasized to support the accuracy of subsequent cross-modal matching. Data from all subspace regions are stored in a normalized manner according to a unified spatiotemporal benchmark, forming a structured multimodal feature library that provides highly consistent input for the construction of a global calibration model.
[0035] like Figure 6 As shown, the image displays point cloud data of a subspace region. The red point cloud in the image represents the 3D point cloud outline of a rectangular planar reference object. Its surface is uniformly covered with a high-reflectivity coating, exhibiting a dense and continuous geometric structure under LiDAR scanning. The edge contours are sharp, and the corner points are evenly distributed, effectively reflecting the excellent capture capability of LiDAR under high-reflectivity materials.
[0036] Furthermore, during the acquisition process, the scanning resolution of the lidar can be dynamically adjusted according to the distance level of the subspace region. A high-resolution mode is used in the near-distance region to capture detailed features, while the resolution is appropriately reduced in the far-distance region to improve the uniformity of point cloud density.
[0037] Step S22: Perform feature extraction on each reference point cloud and each reference image to obtain the point cloud features and image features of each planar reference object.
[0038] In this embodiment, the point cloud features include the corresponding fitted plane and three-dimensional corner points extracted based on the reference point cloud of the corresponding planar reference object, while the image features include the two-dimensional corner points detected in the reference image.
[0039] For each planar reference object, the fitting plane is constructed by first randomly selecting at least three non-collinear points from each reference point cloud to build the corresponding reference plane. Then, the distances from all points in each reference point cloud to the corresponding reference plane are calculated, and points that meet the distance threshold are identified as inliers. Next, it is determined whether all combinations of points in each reference 3D point cloud have been traversed. If not, at least three non-collinear points are randomly selected from each reference point cloud to construct the corresponding reference plane, and this process is repeated. If all combinations have been traversed, the reference plane with the most inliers is selected as the fitting plane for each reference point cloud, ensuring optimal spatial consistency and geometric representativeness of the fitting result. The equation of the corresponding fitting plane can be expressed as: Ax + By + Cz + D = 0, where A, B, and C are the plane normal vector components, and D is the offset of the plane from the origin. This fitting plane preserves the geometric features of the original point cloud to the greatest extent, effectively suppresses noise interference, and provides basic geometric constraints for subsequent corner detection and pose calculation.
[0040] For the 3D corner points corresponding to each planar reference object, the k nearest neighbors of each point cloud point must first be found. The normal direction between each point cloud point and its k nearest neighbors is calculated using the cross product of vectors, and the normal vector is normalized. Then, the outer product of the normals of the k nearest neighbors of each point cloud point is summed to generate a 2×2 structure tensor (or a 3×3 structure tensor, adjusted according to the scene). By calculating the determinant and tensor trace of the structure tensor, the corner detection formula is applied to extract the 3D corner response value. The larger the 3D corner response value, the more likely the point cloud point is to be a 3D corner point (due to drastic changes in local geometry). Subsequently, non-maximum suppression is performed to filter weak corner points, and point cloud points with 3D corner response values greater than the 3D corner response threshold are retained for planar adaptation. The 3D corner point is generated by setting the z-axis coordinate of the corresponding point cloud point to all zeros, ensuring that the normal direction is perpendicular to the fitted plane. The corner detection formula is as follows: R = det(A) - α·(trace(A)) 2 ; Where R represents the corner response value; det(A) represents the determinant of the structure tensor; α is an empirical parameter, usually taken as 0.04-0.06; and trace(A) is the tensor trace.
[0041] like Figure 7 As shown, the figure displays the point cloud distribution of a rectangular planar reference object in three-dimensional space, as well as the results of its fitting plane and three-dimensional corner detection. The four circles in the figure mark the four three-dimensional corners of the corresponding fitting plane.
[0042] For two-dimensional corner points in image features, the reference image is first preprocessed by using Gaussian filtering to suppress noise and converting it to grayscale to reduce the impact of illumination changes. Then, Sobel or Canny edge detection is used to extract image gradient information and calculate the gradient magnitude and direction of each pixel in the grayscale image. Next, non-maximum suppression is performed along the gradient direction, and linear interpolation is applied to retain local maxima points for dual-threshold detection. By setting high and low gradient thresholds, points with gradient magnitudes higher than the high threshold are retained as strong edge points, while points with gradient magnitudes between the low and high thresholds in their neighborhood are considered weak edge points. Connecting weak edge points forms a continuous edge structure. Finally, two-dimensional corner point detection is performed on the strong edge points. Using the determinant and tensor trace of the structure tensor corresponding to the reference image, the two-dimensional corner response value of each strong edge point is calculated according to the aforementioned corner point detection formula. Non-maximum suppression is then applied to filter weak corner points, and strong edge points with two-dimensional corner response values greater than the set two-dimensional corner response threshold are selected as two-dimensional corner points.
[0043] Furthermore, a correspondence can be established between two-dimensional and three-dimensional corner points through projection relationships. By using the camera's intrinsic parameters to calibrate extrinsic parameters, the three-dimensional corner points are reprojected onto the reference image. The pixel-level deviation between the projected points of the three-dimensional corner points and the detected two-dimensional corner points is calculated to optimize the calibration parameters and minimize the reprojection error.
[0044] Step S3: Based on the preset initial extrinsic parameters, and combined with the features of each point cloud and each image, obtain the target extrinsic parameters between the LiDAR and the camera.
[0045] It should be noted that the initial extrinsic parameters are the starting point of the calibration process and are usually obtained through coarse registration methods, such as hand-eye calibration based on geometric constraints or pose estimation using known reference objects.
[0046] In this embodiment, the initial extrinsic parameters are determined by the relative pose and structural features of the LiDAR and camera. A three-dimensional calibration model of the LiDAR and camera is established, and the initial rotation matrix and initial translation vector between the optical centers of the LiDAR and camera are calculated. That is, the initial extrinsic parameters consist of the initial rotation matrix and the initial translation vector, used to map the point cloud features in the LiDAR coordinate system to the camera coordinate system. Subsequently, through a joint optimization strategy, using geometric consistency constraints between the point cloud features and image features within each calibration distance interval, the initial extrinsic parameters are refined and iteratively optimized to obtain the target extrinsic parameters, ensuring that the point cloud features and image features achieve optimal alignment under multi-view geometric constraints.
[0047] Specifically, such as Figure 8 As shown, step S3 includes the following sub-steps: Step S31: Obtain the initial extrinsic parameters between the lidar and the camera based on their relative pose and structural features.
[0048] It should be noted that relative pose refers to the fixed geometric relationship between the lidar and camera during installation, including the translation and rotation relationships reflected by the installation position and orientation of the lidar and camera. This can be obtained through mechanical design drawings or physical measurements and serves as an important basis for calibration initialization. Structural features refer to the physical structural parameters of the lidar and camera, including structural composition and dimensions. Through the coupled analysis of structural features and relative pose, the initial rotation matrix and initial translation vector between the lidar and camera can be derived, thereby constructing the rigid body transformation matrix from the lidar coordinate system to the camera coordinate system.
[0049] In this embodiment, a three-dimensional calibration model of the lidar and camera is first established, including parameters such as the installation position, installation direction, structural composition, and dimensions of the lidar and camera. Then, using this three-dimensional calibration model, the initial rotation matrix and initial translation vector between the lidar and camera are calculated, thus completing the construction of the initial extrinsic parameters.
[0050] First, in the 3D calibration model, the vector between the lidar and the camera's optical center is measured as the initial translation vector. Then, the initial rotation matrix is determined by selecting three non-parallel camera planes in the 3D model corresponding to the camera and measuring the normal vectors of these three camera planes. n cx , n cy and n cz Simultaneously, three non-parallel radar planes were selected in the 3D model of the lidar, and the normal vectors of these three radar planes were measured. n lx , n ly and n lz ,calculate n cx and n lx The first angle between , n cy and n ly The second included angle between ,as well as n cz and n lz The third angle between Then, based on the first included angle The second included angle and the third angle Calculate the initial rotation matrix to describe the orientation relationship between the lidar coordinate system and the camera coordinate system. The initial rotation matrix is: .
[0051] It should be noted that for both LiDAR and camera, three non-parallel end faces can be found as the radar plane or camera plane. The selection of the three normal vectors should cover the main sensing directions of the LiDAR and camera to ensure the geometric representativeness of the rotation matrix.
[0052] In this embodiment, the calculation process of the initial rotation matrix follows the Euler angle transformation principle, by rotating the first included angle around the coordinate axis in sequence. The second included angle and the third angle The coordinate system is aligned using angles to ensure accurate modeling of the spatial orientation relationship between the LiDAR and the camera. Combined with the previously measured initial translation vector, this forms the complete initial extrinsic parameters. These initial extrinsic parameters serve as the starting values for subsequent joint optimization, providing a high-precision spatial alignment foundation for the deep fusion of point cloud and image features, effectively reducing the convergence difficulty in the subsequent nonlinear optimization process.
[0053] Step S32: Based on the initial extrinsic parameters, combine the features of each point cloud and each image to obtain the target extrinsic parameters.
[0054] In this embodiment, by constructing a joint optimization model, the point cloud features and image features are aligned in a common space. Using initial extrinsic parameters as initial values, the initial rotation matrix and initial translation vector are iteratively optimized to maximize the matching between the edge features of the LiDAR point cloud projected onto the camera view and the image gradient direction.
[0055] Specifically, such as Figure 9 As shown, step S32 includes the following sub-steps: Step S321: Project each 3D feature point onto the reference image corresponding to each 2D feature point through the initial extrinsic parameters to obtain the corresponding point cloud projection points.
[0056] In this embodiment, each 3D corner point is transformed to the camera coordinate system using initial extrinsic parameters and then projected onto the corresponding reference image to obtain its point cloud projection point position in the 2D reference image. Subsequently, the geometric distance between each point cloud projection point and its corresponding 2D corner point in the reference image is calculated, yielding the deviation between the point cloud projection point and its corresponding 2D corner point, which is used to quantitatively evaluate the alignment accuracy of the initial extrinsic parameters. This deviation serves as an input term to the optimization objective function, guiding the adjustment direction of the initial rotation matrix and initial translation vector in subsequent iterations, ensuring that the point cloud edge structure and image gradient distribution tend to be consistent, thereby improving the accuracy and robustness of cross-modal feature matching.
[0057] Step S322: Based on the deviation between each point cloud projection point and each two-dimensional feature point, obtain the pose offset between the lidar and the camera.
[0058] In this embodiment, an error optimization function is constructed based on the deviation between each point cloud projection point and its corresponding 2D corner point. The deviation is used as the residual input to iteratively solve the pose offset between the lidar and the camera. This pose offset includes rotation correction terms and translation correction terms, which are used to update the initial rotation matrix and initial translation vector in the initial extrinsic parameters, gradually approximating the optimal target extrinsic parameters.
[0059] Step S323: Adjust the initial extrinsic parameters according to the pose offset to obtain the target extrinsic parameters.
[0060] Specifically, the updated rotation and translation correction terms are applied to the initial extrinsic parameters to iteratively adjust the relative pose of the LiDAR and camera until the error function converges to a preset error threshold, ultimately obtaining high-precision target extrinsic parameters. At this point, after multiple iterations of optimization, the initial rotation matrix and initial translation vector gradually converge, and the point cloud projection edges and image gradient directions tend to be consistent.
[0061] Furthermore, the reprojection error (the deviation between the projected point cloud points and the 2D feature points) is minimized using an iterative algorithm (such as Levenberg-Marquardt), and the initial rotation matrix and initial translation vector are finely adjusted to obtain the optimal target extrinsic parameters. In each iteration, the derivative matrix J (such as the Jacobian matrix) of the reprojection error with respect to the optimization variables (initial rotation matrix and initial translation vector) is calculated. A linear equation is then constructed using this derivative matrix: (J...) T J+λI)δ=-J Te, where e is the current reprojection error vector, λ is the damping factor, and δ is the pose update amount to be solved, which includes the update step sizes δ_R and δ_T of the initial rotation matrix and initial translation vector, respectively. The pose update amount δ is obtained by solving this linear equation, and then the parameters are updated using the formula R_new = R_old + δ_R, T_new = T_old + δ_T, where R_old and T_old represent the initial rotation matrix and initial translation vector, respectively, and R_new and T_new represent the updated initial rotation matrix and initial translation vector, respectively. Finally, the above iterative process is repeated until the pose update amount δ is less than a preset update threshold or the number of iterations reaches the upper limit, completing the accurate calibration of the target extrinsic parameters. The final target extrinsic parameters include the optimized target rotation matrix and target translation vector, which can accurately describe the spatial geometric relationship between the LiDAR and the camera, achieving high-precision alignment of point cloud data and image data in a unified coordinate system. The calibration results can be directly used for subsequent multimodal perception tasks, such as 3D object detection, semantic segmentation and environment mapping, significantly improving the system's perception capability and positioning accuracy in complex scenes.
[0062] It should be noted that the robustness of the calibration results can be further improved through joint optimization using multi-frame data, effectively suppressing abnormal noise interference from single frames. By selecting multi-frame point cloud and image data at different distances in a static scene, a joint error function is constructed, which weights and fuses the reprojection errors of the multi-frame data, enhancing the calibration process's tolerance to noise and local errors. Furthermore, the global consistency of the calibration results can be further optimized through steps such as random sampling, model estimation, model validation, iterative repetition, and final optimization.
[0063] In some embodiments, the method further includes the step of: evaluating the calibration error of the target extrinsic parameters to obtain the calibration error value corresponding to each planar reference object. Specifically, each reference point cloud is projected onto a corresponding reference image using the target extrinsic parameters, and the projected point cloud points are drawn on the reference image. The projection process can be expressed as: Z·p=K·(R´·P+T´), where P is the three-dimensional feature point of the reference point cloud, R´ and T´ are the target rotation matrix and target translation vector from the laser radar to the camera, K is the camera intrinsic parameter, Z is the scale coefficient, and p is the point cloud projection point projected onto the reference image. Then, the projection error between the farthest point cloud projection point and the corresponding real two-dimensional feature point in the reference image is calculated; this is the calibration error, used to quantify the calibration accuracy.
[0064] like Figure 10As shown, the dots in the figure represent point cloud projection points that extend beyond the range of the planar reference object 10. Point cloud projection points extending beyond the range of the planar reference object 10 are located around it. The point cloud projection point furthest from the planar reference object 10 is identified as the target projection point 3. The length of the telescopic rod 2 is adjusted so that its two ends are aligned with the edge of the planar reference object 10 and the target projection point 3, respectively. Then, the telescopic rod 2 is placed on the support rod 4 of the bracket below the planar reference object 10, aligned with the scale on the support rod 4. The length of the telescopic rod is read from the scale on the support rod 4, which is the calibration error. By introducing physical measurement to assist verification, the projection error is transformed into an intuitive geometric length reading, significantly improving the reliability and interpretability of the calibration results.
[0065] Furthermore, for long-distance calibration scenarios, the error threshold is 2cm if the calibration distance is within 50m; 4cm if the calibration distance is between 50m and 150m; and 6cm if the calibration distance exceeds 150m. When the calibration error is lower than the corresponding error threshold at the distance, the external parameter calibration result is deemed to meet the accuracy requirements and can be used in practical applications; otherwise, data needs to be re-collected and calibration parameters optimized. The reprojection error is gradually widened from ≤1.5 pixels at 50-100 meters to ≤3.5 pixels at 200-300 meters, and the three-dimensional deviation is widened from ≤0.5 meters to ≤1.5 meters, conforming to the geometric law of perspective projection where "error increases with distance." Compared to a single threshold, this system can avoid "misjudgment of long-distance errors," making the acceptance criteria more scientific, and improving the accuracy of accuracy judgment across the entire distance range by 80%. This method balances algorithm optimization with engineering testing, effectively ensuring the collaborative perception performance of LiDAR and cameras in long-distance, large-scale scenarios, and is suitable for high-reliability scenarios such as autonomous driving and intelligent transportation.
[0066] This method combines real-time reference images with dynamic feedback from point cloud projections to achieve visualized monitoring of calibration errors. It can also directly measure the distance between the projected points on the reference image and the actual two-dimensional feature points, converting this distance into actual spatial error based on the proportional relationship between the reference image and the actual object, thus enabling quantitative assessment of calibration accuracy. This method is not only applicable to static calibration scenarios but can also be extended to online calibration updates in dynamic environments. By combining multi-frame perception data from vehicle or aircraft operation, it continuously optimizes extrinsic parameters, improving the stability and reliability of LiDAR and camera systems during long-term operation.
[0067] Furthermore, the target extrinsic parameters can be optimized based on each calibration error value to obtain the optimized final extrinsic parameters, ensuring its generalization ability at different distances and viewpoints, and further improving calibration accuracy and stability.
[0068] In summary, this method provides an extrinsic parameter calibration approach for LiDAR and cameras. By dividing the calibration space into multiple different calibration distance intervals and performing simultaneous calibration and error evaluation for each interval, a multi-scale calibration optimization framework is constructed, achieving collaborative optimization of accuracy across the entire range from near to far distances. Simultaneously, by dynamically adjusting the reprojection error and 3D deviation threshold, the parameter drift problem caused by error accumulation at long distances is effectively suppressed. Combined with real-time alignment feedback of point cloud projection and image features, online verification and iterative optimization of calibration results are supported, significantly enhancing the adaptability of LiDAR and camera extrinsic parameter calibration in complex environments. This method has been validated on various vehicle-mounted and airborne platforms, achieving a calibration success rate of over 95%, meeting the reliability requirements of high-precision perception systems in practical applications. Furthermore, this method supports universal deployment across sensor configurations, adaptable to combinations of cameras with different resolutions and multi-beam LiDARs, demonstrating excellent performance in both calibration efficiency and accuracy consistency.
[0069] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0070] In one embodiment, an extrinsic parameter calibration device for a lidar and camera is provided, which corresponds one-to-one with the extrinsic parameter calibration method for the lidar and camera described in the above embodiments. For example... Figure 11 As shown, the extrinsic parameter calibration device for the lidar and camera includes a spatial partitioning module 101, a feature acquisition module 102, and an extrinsic parameter acquisition module 103. Detailed descriptions of each functional module are as follows: The spatial partitioning module 101 is used to divide the calibration space into multiple sub-space regions corresponding to different calibration distance intervals according to the calibration position.
[0071] The feature acquisition module 102 is used to acquire point cloud features of planar reference objects in each subspace region through LiDAR, and to acquire image features of planar reference objects in each subspace region through a camera.
[0072] The extrinsic parameter acquisition module 103 is used to acquire the target extrinsic parameters between the lidar and the camera based on the preset initial extrinsic parameters and the features of each point cloud and each image.
[0073] Specific limitations regarding the extrinsic parameter calibration device for LiDAR and cameras can be found in the limitations of the extrinsic parameter calibration method for LiDAR and cameras mentioned above, and will not be repeated here. Each module in the aforementioned extrinsic parameter calibration device for LiDAR and cameras can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0074] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 12 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a method for extrinsic parameter calibration of a lidar and camera.
[0075] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the extrinsic parameter calibration method for the lidar and camera described in the above embodiments, for example... Figure 1 S1-S3, as shown, will not be described again here to avoid repetition. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the external parameter calibration device for LiDAR and camera, for example... Figure 11 The functions of the spatial partitioning module 101, feature acquisition module 102, and extrinsic parameter acquisition module 103 shown are not described again here to avoid repetition.
[0076] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the extrinsic parameter calibration method for the lidar and camera described in the above embodiment, for example... Figure 1 S1-S3, as shown, will not be described again here to avoid repetition. Alternatively, when this computer program is executed by the processor, it implements the functions of each module / unit in the above embodiment of the external parameter calibration device for lidar and camera, for example... Figure 11 The functions of the spatial partitioning module 101, feature acquisition module 102, and extrinsic parameter acquisition module 103 shown are not described again here to avoid repetition. The computer-readable storage medium can be non-volatile or volatile.
[0077] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0078] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0079] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for calibrating the extrinsic parameters of a lidar and camera, characterized in that, The lidar and the camera are jointly positioned in the calibration space at a calibration location for extrinsic parameter calibration. Planar reference objects are respectively positioned at multiple different distances from the calibration location within the calibration space. The extrinsic parameter calibration method includes the following steps: Based on the calibration position, the calibration space is divided into multiple sub-space regions corresponding to different calibration distance intervals; The point cloud features of the planar reference object in each of the sub-space regions are obtained by the lidar, and the image features of the planar reference object in each of the sub-space regions are obtained by the camera; Based on the preset initial extrinsic parameters, and combined with the point cloud features and the image features, the target extrinsic parameters between the lidar and the camera are obtained.
2. The external parameter calibration method according to claim 1, characterized in that, The step of dividing the calibration space into multiple sub-space regions corresponding to different calibration distance intervals based on the calibration position includes: The global image of the calibration space is acquired, and the global image is divided to obtain multiple local images corresponding to different calibration distance intervals; The calibration space is partitioned according to each of the local images to obtain the subspace region corresponding to each of the local images.
3. The external parameter calibration method according to claim 2, characterized in that, The process of dividing the global image to obtain multiple local images corresponding to different calibration distances includes: The global image is divided into a nine-grid structure to obtain nine equally divided local sub-images; The local sub-images are merged to obtain three local images; The step of partitioning the calibration space according to each of the local images to obtain a subspace region corresponding to each of the local images includes: The three local images are mapped to the calibration space to obtain three sub-space regions with calibration distance intervals of near distance, medium distance, and far distance, respectively.
4. The external parameter calibration method according to claim 1, characterized in that, The step of obtaining target extrinsic parameters between the lidar and the camera based on preset initial extrinsic parameters, combined with the point cloud features and the image features, includes: Based on the relative pose and structural features of the lidar and the camera, the initial extrinsic parameters between the lidar and the camera are obtained; Based on the initial extrinsic parameters, the target extrinsic parameters are obtained by combining the point cloud features and the image features.
5. The external parameter calibration method according to claim 4, characterized in that, The point cloud features include three-dimensional feature points, and the image features include two-dimensional feature points; obtaining the target extrinsic parameters based on the initial extrinsic parameters, by combining each of the point cloud features and each of the image features, includes: Each of the three-dimensional feature points is projected onto the reference image corresponding to each of the two-dimensional feature points through the initial extrinsic parameters to obtain the corresponding point cloud projection points; Based on the deviation between each of the point cloud projection points and each of the two-dimensional feature points, the pose offset between the lidar and the camera is obtained; The initial extrinsic parameters are adjusted based on the pose offset to obtain the target extrinsic parameters.
6. The external parameter calibration method according to claim 1, characterized in that, Each of the aforementioned sub-space regions contains at least one of the aforementioned planar reference objects. The step of acquiring point cloud features of the planar reference objects within each of the aforementioned sub-space regions using the lidar, and acquiring image features of the planar reference objects within each of the aforementioned sub-space regions using the camera, includes: The lidar and the camera are sequentially aligned with each of the sub-space regions, and reference point clouds and reference images of the planar reference objects within each of the sub-space regions are acquired synchronously. Feature extraction is performed on each of the reference point clouds and each of the reference images to obtain the point cloud features and image features of each of the planar reference objects.
7. The external parameter calibration method according to claim 1, characterized in that, The external parameter calibration method also includes: The calibration error of the target extrinsic parameters is evaluated to obtain the calibration error value corresponding to each of the planar reference objects.
8. A device for calibrating the external parameters of a lidar and camera, characterized in that, The lidar and the camera are jointly positioned in the calibration space at a calibration location for extrinsic parameter calibration. Planar reference objects are respectively positioned at multiple different distances from the calibration location within the calibration space. The extrinsic parameter calibration device includes: The spatial partitioning module is used to divide the calibration space into multiple sub-space regions corresponding to different calibration distance intervals according to the calibration position; The feature acquisition module is used to acquire point cloud features of the planar reference object in each of the sub-space regions through the lidar, and to acquire image features of the planar reference object in each of the sub-space regions through the camera; The extrinsic parameter acquisition module is used to acquire the target extrinsic parameters between the lidar and the camera based on preset initial extrinsic parameters, combined with the point cloud features and the image features.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the external parameter calibration method for the lidar and camera as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the external parameter calibration method for the lidar and camera as described in any one of claims 1 to 7.