A method and apparatus for joint calibration of intrinsic and extrinsic parameters of a camera and a lidar.
By introducing the pose sequence of the robotic arm end effector during the camera and lidar calibration process, and constructing an error objective function for optimization, the problem of error accumulation in the existing technology is solved, and a joint calibration with high accuracy and high efficiency is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN LIUXING TECHNOLOGY LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies suffer from error accumulation during the calibration process of cameras and lidar, resulting in low accuracy of calibration results and difficulty in achieving joint calibration, thus failing to meet the needs of automated and batch calibration.
By acquiring image sequences and point cloud sequences of the robotic arm under different calibration poses, and combining the installation poses of the visual target and the reflective target, objective functions for reprojection error and point-to-surface distance error are constructed. The robotic arm end-effector pose sequence is used for optimization and solution to determine the intrinsic and extrinsic parameters of the camera and lidar, thereby achieving joint calibration.
It significantly improves the accuracy and consistency of calibration results, realizes the automation and structuring of the calibration process, and meets the needs of batch calibration on the production line.
Smart Images

Figure CN122492839A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sensor calibration technology, and more particularly to the field of joint calibration of intrinsic and extrinsic parameters, specifically to a method and apparatus for joint calibration of intrinsic and extrinsic parameters of a camera and a lidar. Background Technology
[0002] In fields such as autonomous driving, robot navigation, and 3D reconstruction, the combined use of cameras and LiDAR has become the mainstream solution. To achieve accurate fusion of multi-sensor data, it is necessary to ensure the accurate calibration of camera intrinsic parameters to avoid geometric distortion of the image; at the same time, it is necessary to calibrate the intrinsic parameters of the LiDAR to reduce point cloud distortion; in addition, it is also necessary to accurately calibrate the extrinsic parameters between the camera and the LiDAR to achieve accurate registration between the point cloud and the image.
[0003] Existing technologies typically employ two calibration methods: The first is the traditional manual calibration method, which relies on operators holding a chessboard grid and placing it at multiple angles and distances. However, this method suffers from poor pose repeatability, low consistency in calibration results, and uneven sampling distribution, leading to insufficient coverage of the field of view edges and extreme angle areas. Furthermore, manual operation is inefficient and time-consuming, making it difficult to meet the batch calibration needs of production lines, and the true pose value of the calibration board usually relies on external measurements, resulting in limited accuracy. The second method is automated calibration, such as using guide rails or turntables. Guide rails only provide one-dimensional motion, limiting the degree of freedom in pose adjustment, while turntables cannot simultaneously change the position and orientation of both the visual and reflective targets, thus limiting spatial sampling capabilities. In addition, existing solutions are mostly designed for single sensors, making it difficult to achieve joint calibration of LiDAR and cameras.
[0004] Therefore, existing technologies are prone to error accumulation during calibration, resulting in low accuracy of calibration results. Summary of the Invention
[0005] This application provides a method and apparatus for joint calibration of intrinsic and extrinsic parameters of a camera and a lidar, thereby improving the accuracy of the calibration results.
[0006] According to one aspect of this application, a method for joint calibration of intrinsic and extrinsic parameters of a camera and a lidar is provided, comprising: The image sequence, point cloud sequence, and end-effector pose sequence of the robotic arm at each calibration pose are obtained from the pre-generated calibration pose sequence; wherein, the image sequence and the point cloud sequence are obtained by acquiring visual targets and reflective targets fixed at the end of the robotic arm by a camera and a lidar fixed at the calibration station, respectively. Based on the image sequence, the point cloud sequence, the robotic arm end-effector pose sequence, the installation pose of the visual target, the installation pose of the reflective target, the camera parameters to be solved, and the lidar parameters to be solved, the objective function for the reprojection error of the camera and the objective function for the point-to-surface distance error of the lidar are constructed respectively, and optimized to obtain the camera calibration results and lidar calibration results. Based on the image sequence, the point cloud sequence, the camera calibration results, and the lidar calibration results, the extrinsic parameter transformation matrix between the camera and the lidar is determined.
[0007] According to another aspect of this application, a joint calibration device for intrinsic and extrinsic parameters of a camera and a lidar is provided, comprising: The multi-source data synchronous acquisition module is used to acquire image sequences, point cloud sequences, and end-effector pose sequences of the robotic arm at each calibration pose from a pre-generated calibration pose sequence; wherein, the image sequences and the point cloud sequences are acquired by a camera and a lidar fixed at the calibration station to acquire visual targets and reflective targets fixed at the end of the robotic arm, respectively. The camera and lidar intrinsic parameter calibration module is used to construct the reprojection error objective function of the camera and the point-to-surface distance error objective function of the lidar based on the image sequence, the point cloud sequence, the pose sequence of the robotic arm end effector, the installation pose of the visual target, the installation pose of the reflective target, the camera parameters to be solved, and the lidar parameters to be solved, and then optimize and solve them to obtain the camera calibration results and lidar calibration results. The camera and lidar extrinsic parameter calibration module is used to determine the extrinsic parameter transformation matrix between the camera and lidar based on the image sequence, the point cloud sequence, the camera calibration result, and the lidar calibration result.
[0008] According to another aspect of this application, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the camera and lidar joint calibration method for intrinsic and extrinsic parameters as described in any embodiment of this application.
[0009] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the camera and lidar internal and external parameter joint calibration method described in any embodiment of this application.
[0010] According to another aspect of this application, a computer program product is provided, comprising a computer program that, when executed by a processor, implements any of the camera and lidar internal and external parameter joint calibration methods provided in the embodiments of this application.
[0011] The technical solution of this application embodiment acquires image sequences, point cloud sequences, and robotic arm end-effector pose sequences. Based on the image sequences, point cloud sequences, robotic arm end-effector pose sequences, the mounting poses of the visual target and the reflective target, the camera parameters to be solved, and the lidar parameters to be solved, a reprojection error objective function for the camera and a point-to-surface distance error objective function for the lidar are constructed respectively, and optimized to obtain camera calibration results and lidar calibration results. Based on the image sequences, point cloud sequences, camera calibration results, and lidar calibration results, the extrinsic parameter transformation matrix between the camera and lidar is determined. This solution introduces the robotic arm end-effector pose sequence, allowing image sequences and point cloud sequences under different calibration poses to participate in the camera and lidar parameter solving process under unified constraints. This avoids error accumulation caused by independent calibration or step-by-step estimation, significantly improving the accuracy and consistency of the calibration results. At the same time, it realizes the structuring and automation of the calibration process, improves calibration efficiency, and can meet the batch calibration needs of production lines.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0013] Figure 1 This is a flowchart of a method for joint calibration of intrinsic and extrinsic parameters of a camera and a lidar according to Embodiment 1 of this application; Figure 2 This is a flowchart of another method for joint calibration of camera and lidar intrinsic and extrinsic parameters according to Embodiment 2 of this application; Figure 3 This is a schematic diagram of the structure of a camera and lidar internal and external parameter joint calibration device according to Embodiment 3 of this application; Figure 4 This is a schematic diagram of the structure of an electronic device that implements the joint calibration method for the intrinsic and extrinsic parameters of the camera and lidar according to the embodiments of this application. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] Example 1 Figure 1 This is a flowchart illustrating a method for joint calibration of intrinsic and extrinsic parameters of a camera and a LiDAR according to Embodiment 1 of this application. This embodiment is applicable to joint calibration of intrinsic and extrinsic parameters of a camera and a LiDAR in multi-sensor fusion application scenarios. This method can be executed by a joint calibration device for the intrinsic and extrinsic parameters of a camera and a LiDAR, which can be implemented in hardware and / or software and can be configured in a computer device. Figure 1 As shown, the method includes: S101. Obtain the image sequence, point cloud sequence and end-effector pose sequence of the robotic arm in each calibration pose from the pre-generated calibration pose sequence.
[0017] The image sequence and the point cloud sequence are obtained by acquiring visual targets and reflective targets fixed at the end of the robotic arm by a camera and a lidar fixed at the calibration station, respectively.
[0018] In this embodiment, the calibration pose sequence refers to a pre-generated set of robotic arm poses, used to guide the robotic arm to move sequentially to different spatial positions and attitudes during the calibration process, thereby achieving coverage of the sensor observation space. The sensors include a camera and a LiDAR. The robotic arm end-effector pose refers to the spatial pose information of the robotic arm end-effector relative to the robotic arm base coordinate system, including position parameters and attitude parameters. The visual target is a calibration pattern recognizable by the camera; the reflective target is a highly reflective target suitable for LiDAR detection, used to provide effective observation objects for both the camera and LiDAR.
[0019] For example, visual targets are used for camera calibration and can employ high-contrast black and white patterns such as checkerboard or dot arrays, with corner point or center-to-center spacing accuracy better than ±0.01mm to meet the accuracy requirements of camera distortion parameter calibration. Reflective targets are used for lidar calibration and can be planar or polyhedral targets with known geometric dimensions, coated with a high diffuse reflectance material with a reflectivity of not less than 90% to ensure the lidar can obtain a stable and effective echo signal. The visual and reflective targets are fixedly mounted at the end of the robotic arm, their relative positions determined by their respective installation positions and remaining stable during calibration. The camera and lidar are fixedly mounted at the calibration station, and their relative positions remain unchanged during calibration. The camera can use global exposure to eliminate rolling shutter artifacts generated during robotic arm movement, thereby improving the accuracy of image feature point extraction under high-speed continuous acquisition conditions.
[0020] Specifically, during the data acquisition phase, the robotic arm is controlled to move sequentially according to a pre-generated sequence of calibration poses, maintaining a preset time interval at each calibration pose to ensure stability. While the robotic arm is in each calibration pose, the camera and LiDAR are simultaneously triggered to acquire data. The camera acquires images of the visual target positioned at the end of the robotic arm, and the LiDAR acquires point cloud data of the reflective target positioned at the end of the robotic arm, simultaneously obtaining the corresponding end-effector pose data. This process is repeated for each calibration pose in the sequence to obtain the corresponding image sequence, point cloud sequence, and end-effector pose sequence. This method allows for the acquisition of camera and LiDAR observation data under multiple pose conditions, as well as end-effector pose data, providing a reliable data foundation for subsequent joint calibration of internal and external parameters of the camera and LiDAR.
[0021] S102. Based on the image sequence, the point cloud sequence, the robotic arm end-effector pose sequence, the installation pose of the visual target, the installation pose of the reflective target, the camera parameters to be solved, and the lidar parameters to be solved, construct the objective function of the camera reprojection error and the objective function of the lidar point-to-surface distance error respectively, and optimize and solve them to obtain the camera calibration results and lidar calibration results.
[0022] In this embodiment, the mounting pose of the visual target refers to the fixed mounting relationship of the visual target relative to the end effector of the robotic arm, used to describe the position and orientation of the visual target in the coordinate system of the end effector of the robotic arm; the mounting pose of the reflective target refers to the fixed mounting relationship of the reflective target relative to the end effector of the robotic arm, used to describe the position and orientation of the reflective target in the coordinate system of the end effector of the robotic arm. The camera parameters to be solved include camera intrinsic parameters and camera extrinsic parameters. The camera intrinsic parameters include the camera intrinsic parameter matrix and camera distortion parameters, used to characterize the imaging model parameters of the camera. The camera extrinsic parameters are used to describe the spatial positional relationship of the camera relative to the reference coordinate system. The lidar parameters to be solved include lidar intrinsic parameter compensation parameters and lidar extrinsic parameters. The intrinsic parameter compensation parameters are used to correct system errors during lidar ranging and scanning, and the lidar extrinsic parameters are used to describe the spatial positional relationship of the lidar relative to the reference coordinate system. When describing the camera and lidar extrinsic parameters, the reference coordinate system is the coordinate system of the robotic arm base.
[0023] The reprojection error objective function is an error function constructed based on the camera imaging model, calculating the deviation between the projected image feature points in 3D space onto the image plane and the actually detected image feature points. The point-to-plane distance error objective function is an error function constructed based on the relationship between point cloud data and a plane with known geometric structures, calculating the distance from each point in the point cloud to the plane. Camera calibration results refer to the set of camera parameters obtained through optimization, including the determined camera intrinsic matrix, camera distortion parameters, and camera extrinsic parameters. LiDAR calibration results refer to the set of LiDAR parameters obtained through optimization, including the determined LiDAR intrinsic compensation parameters and LiDAR extrinsic parameters.
[0024] Specifically, for each calibration pose in the calibration pose sequence, the corresponding image, point cloud, and robotic arm end-effector pose are acquired. Based on the robotic arm end-effector pose sequence and the installation poses of the visual and reflective targets, the data acquired under each calibration pose can be processed in conjunction with the corresponding robotic arm end-effector pose during parameter solving. Furthermore, for camera calibration, the observed positions of visual target feature points in the image are processed in conjunction with the robotic arm end-effector pose and the installation pose of the visual target. The difference between the processed result and the predicted position obtained based on the camera parameters to be solved is used to construct a reprojection error objective function. For lidar calibration, the data related to the reflective target in the point cloud are processed in conjunction with the robotic arm end-effector pose and the installation pose of the reflective target. A point-to-surface distance error objective function is constructed based on the distance relationship between the processed point cloud data and the reflective target plane. Thus, the data corresponding to each calibration pose participates in the construction of the corresponding error objective function under the condition of incorporating the robotic arm end-effector pose, and is optimized and solved under the camera and lidar parameters to be solved. This method introduces the pose sequence of the robotic arm end effector, enabling data acquired under different calibration poses to be uniformly processed according to the corresponding robotic arm end effector pose during the parameter solution process. This allows the information corresponding to each pose in the calibration pose sequence to be comprehensively utilized in the same solution process, avoiding the accumulation of errors caused by independent or distributed estimation of each calibration pose, and improving the accuracy and consistency of the calibration results.
[0025] S103. Determine the extrinsic parameter transformation matrix between the camera and the lidar based on the image sequence, the point cloud sequence, the camera calibration result, and the lidar calibration result.
[0026] In this embodiment, the extrinsic parameter transformation matrix between the camera and the lidar refers to a parameter matrix used to describe the relative spatial relationship between the camera coordinate system and the lidar coordinate system, and is used to realize the spatial correspondence and transformation between image data and point cloud data. This extrinsic parameter transformation matrix differs from the camera extrinsic parameters and lidar extrinsic parameters. The camera extrinsic parameters and lidar extrinsic parameters are used to describe the spatial pose of the camera and lidar relative to the reference coordinate system, respectively, while the extrinsic parameter transformation matrix is used to directly characterize the relative position and attitude relationship between the camera and the lidar.
[0027] Specifically, after obtaining the camera calibration results and LiDAR calibration results, based on the image sequence and point cloud sequence, and combining the camera extrinsic parameters and LiDAR extrinsic parameters from both calibration results, the correspondence between the spatial representations of the camera and LiDAR is determined under the same reference coordinate system. This yields the extrinsic parameter transformation matrix used to describe the relative position and attitude relationship between the camera and LiDAR. In this way, the camera calibration results and LiDAR calibration results participate in the extrinsic parameter determination process under a unified spatial reference, avoiding spatial inconsistencies caused by separate independent descriptions, thereby improving the accuracy of extrinsic parameter determination between the camera and LiDAR.
[0028] The technical solution of this application embodiment acquires image sequences, point cloud sequences, and robotic arm end-effector pose sequences. Based on the image sequences, point cloud sequences, robotic arm end-effector pose sequences, the mounting poses of the visual target and the reflective target, the camera parameters to be solved, and the lidar parameters to be solved, a reprojection error objective function for the camera and a point-to-surface distance error objective function for the lidar are constructed respectively, and optimized to obtain camera calibration results and lidar calibration results. Based on the image sequences, point cloud sequences, camera calibration results, and lidar calibration results, the extrinsic parameter transformation matrix between the camera and lidar is determined. This solution introduces the robotic arm end-effector pose sequence, allowing image sequences and point cloud sequences under different calibration poses to participate in the camera and lidar parameter solving process under unified constraints. This avoids error accumulation caused by independent calibration or step-by-step estimation, significantly improving the accuracy and consistency of the calibration results. At the same time, it realizes the structuring and automation of the calibration process, improves calibration efficiency, and can meet the batch calibration needs of production lines.
[0029] Example 2 Figure 2 This is a flowchart of another method for joint calibration of camera and lidar intrinsic and extrinsic parameters according to Embodiment 2 of this application. The technical solution of this embodiment further refines the application of the robotic arm end-effector pose sequence in the calibration process based on the technical solutions of the above embodiments. For example... Figure 2 As shown, the method includes: S201. Obtain the image sequence, point cloud sequence and end-effector pose sequence of the robotic arm corresponding to each calibration pose from the pre-generated calibration pose sequence.
[0030] The image sequence and the point cloud sequence are obtained by acquiring visual targets and reflective targets fixed at the end of the robotic arm by a camera and a lidar fixed at the calibration station, respectively.
[0031] S202. Based on the end-effector pose sequence and the installation pose of the visual target, calculate the actual pose sequence of the visual target in the coordinate system of the robot arm base. Based on the actual pose sequence of the visual target, the image sequence, and the camera parameters to be solved, construct the reprojection error objective function and optimize it to obtain the camera calibration result.
[0032] S203. Based on the end-effector pose sequence and the mounting pose of the reflective target, calculate the actual pose sequence of the reflective target in the coordinate system of the robot arm base. Based on the actual pose sequence of the reflective target, the point cloud sequence, and the parameters of the lidar to be solved, construct the objective function of the point-to-surface distance error from the point cloud to the reflective target plane, and optimize the solution to obtain the lidar calibration result.
[0033] In this embodiment, the actual pose sequence of the visual target refers to the set of spatial poses of the visual target in the coordinate system of the robotic arm base, determined based on the pose sequence of the robotic arm end effector and the installation pose of the visual target, corresponding to the changes in pose of each target. This set is used to characterize the true spatial pose of the visual target. The actual pose sequence of the reflective target refers to the set of spatial poses of the reflective target in the coordinate system of the robotic arm base, determined based on the pose sequence of the robotic arm end effector and the installation pose of the reflective target. This set is used to characterize the true spatial pose of the reflective target.
[0034] Specifically, based on the end-effector pose sequence and the installation pose of the visual target, the actual pose sequence of the visual target is determined. Combined with the image sequences under each calibration pose and the camera parameters to be solved, a reprojection error objective function is constructed. This transforms the frame-by-frame camera extrinsic variables, which originally varied with the calibration pose, into unified camera extrinsic variables. Camera calibration is optimized only in a fixed parameter space composed of the camera intrinsic matrix, camera distortion parameters, and camera extrinsic parameters. This transforms the optimization variables from multi-pose local variables to a centralized optimization of unified camera parameters. Simultaneously, the reprojection error objective function corresponding to each calibration pose accumulates in the same parameter space to form an overall constraint, improving the stability of the solution process. Based on the end-effector pose... The system generates the actual pose sequence of the reflective target by using the sequence of poses and the installation pose of the reflective target. Combined with the point cloud sequence under each calibration pose and the parameters of the lidar to be solved, a point-to-plane distance error objective function is constructed. By using the pose sequence of the robotic arm end effector and the installation pose of the reflective target, the planar geometric relationship of the reflective target in the coordinate system of the robotic arm base is directly determined. This eliminates the reliance on the planar model obtained by fitting noisy point clouds during lidar calibration, thus avoiding the error propagation problem introduced by fitting a plane based on point clouds. At the same time, since the planar constraints are directly determined by the pose of the robotic arm end effector, the error calculation from point cloud to plane is based on a deterministic geometric benchmark, thereby improving the stability and reliability of lidar intrinsic parameter solution.
[0035] S204. Based on the image sequence, the point cloud sequence, the camera calibration result, and the lidar calibration result, determine the extrinsic parameter transformation matrix between the camera and the lidar.
[0036] The technical solution of this embodiment further refines the application of the robotic arm end-effector pose sequence in the calibration process. During camera calibration, the robotic arm end-effector pose sequence and the mounting pose of the visual target are used to determine the actual pose sequence of the visual target in the robotic arm base coordinate system. Based on this, image sequences under each calibration pose jointly participate in the construction and optimization of the reprojection error objective function, enabling the camera parameter solution to be completed under unified parameter constraints. During lidar calibration, the robotic arm end-effector pose sequence and the mounting pose of the reflective target are used to determine the actual pose sequence of the reflective target in the robotic arm base coordinate system. Based on this deterministic geometric relationship, a point-to-surface distance error objective function from the point cloud to the reflective target plane is constructed, allowing the lidar parameter solution to be optimized based on a stable geometric reference, thereby avoiding the uncertainty caused by relying on the fitting results of the point cloud itself. This method enables camera calibration and lidar calibration to complete parameter solving under the unified spatial constraints provided by the end-effector pose of the robotic arm. During camera calibration, image observation errors at different robotic arm acquisition positions jointly affect the optimization process of the same camera parameter model. During lidar calibration, error calculation and parameter solving are performed based on deterministic plane constraints, thereby completing the optimization solution of lidar parameters.
[0037] In one optional implementation, the camera parameters to be solved include a camera intrinsic matrix, camera distortion coefficients, and camera extrinsic parameters; the camera extrinsic parameters are the camera pose parameters in the coordinate system of the robotic arm base; the step of constructing a reprojection error objective function based on the actual pose sequence of the visual target, the image sequence, and the camera parameters to be solved, and optimizing the solution to obtain the camera calibration result, includes: using the camera extrinsic parameters to be solved, transforming the actual pose sequence of the visual target to the camera coordinate system to obtain a theoretical pose sequence; projecting preset visual target feature points on the visual target onto the camera imaging plane based on the theoretical pose sequence, the camera intrinsic matrix to be solved, and the camera distortion coefficients to obtain predicted pixel positions; extracting the actual pixel positions of the visual target feature points from the image sequence; constructing a reprojection error objective function based on the difference between the predicted pixel positions and the actual pixel positions; and minimizing the reprojection error objective function to jointly optimize the solution of the camera parameters to be solved to obtain the camera calibration result.
[0038] In this embodiment, the theoretical pose sequence refers to the set of spatial poses of the visual target in the camera coordinate system obtained by transforming the actual pose sequence of the visual target into the camera coordinate system through the camera extrinsic parameters to be solved. The predicted pixel position refers to the two-dimensional pixel coordinate position obtained by projecting the visual target feature points onto the camera imaging plane based on the camera intrinsic parameter matrix, camera distortion coefficients, and the theoretical pose sequence. The actual pixel position refers to the actual observed pixel coordinate position of the visual target feature points extracted from the image sequence in the imaging plane. Here, the visual target feature points refer to geometric feature points set on the visual target, used for stable extraction and coordinate localization in the image sequence, such as the corner points of a checkerboard pattern or the center points of a dot array pattern.
[0039] Specifically, the actual pose sequence of the visual target is transformed using the camera extrinsic parameters to be solved, so that it is uniformly expressed in the camera coordinate system. Then, combined with the camera intrinsic parameter matrix and camera distortion coefficients to be solved, the visual target feature points are mapped from three-dimensional space to the two-dimensional imaging plane, obtaining the predicted pixel positions corresponding to the image observations. At the same time, the actual pixel positions of the corresponding visual target feature points are extracted from the image sequence, and the reprojection error objective function is constructed based on the deviation between the predicted pixel positions and the actual pixel positions. By minimizing this reprojection error objective function, the joint optimization of the camera intrinsic parameter matrix, camera distortion parameters, and camera extrinsic parameters is achieved, thereby obtaining the camera calibration results. In this way, image observation information from different robotic arm acquisition positions is transformed into error constraints on the same camera parameter model, allowing all observation data to participate in parameter solving within a unified optimization framework. This improves the overall constraint strength and optimization stability of the parameter solving process, reduces the error accumulation caused by observation dispersion during parameter estimation, and improves the accuracy of the camera calibration results.
[0040] In one optional embodiment, the lidar parameters to be solved include lidar intrinsic parameter compensation parameters and lidar extrinsic parameters; the lidar extrinsic parameters are the pose parameters of the lidar in the coordinate system of the robotic arm base; the step of constructing a point-to-surface distance error objective function from the point cloud to the plane of the reflecting target based on the actual pose sequence of the reflecting target, the point cloud sequence, and the lidar parameters to be solved, and optimizing the solution to obtain the lidar calibration result, includes: determining the plane of the reflecting target in the coordinate system of the robotic arm base based on the actual pose sequence of the reflecting target and the geometry of the reflecting target. The theoretical plane model is obtained by transforming the real plane equation into the lidar coordinate system using the lidar extrinsic parameters to be solved. Point cloud data belonging to the reflecting target are extracted from the point cloud sequence. The point cloud data is then corrected according to the lidar intrinsic parameter compensation parameters to be solved, resulting in corrected point cloud data. A point-to-surface distance error objective function is constructed based on the distance from points in the corrected point cloud data to the theoretical plane model. By minimizing the point-to-surface distance error objective function, the lidar parameters to be solved are jointly optimized to obtain the lidar calibration result.
[0041] In this embodiment, the geometry of the reflective target is predetermined, providing a stable geometric reference for the point cloud sequence acquired by the lidar. This allows the point cloud data acquired by the lidar to correspond to a known spatial geometric structure, thereby transforming the point cloud data processing into an error calculation process based on a predetermined geometric model. For example, when the reflective target has a planar structure, it can be used as a standard reference plane to calculate the distance deviation of each point in the point cloud from this plane.
[0042] The true plane equation refers to the analytical expression of the plane corresponding to the reflective target, obtained by mathematically representing the spatial plane of the reflective target in the coordinate system of the robotic arm base, based on the actual pose sequence and geometry of the reflective target. It describes the true spatial position and attitude of the reflective target in the coordinate system of the robotic arm base. The theoretical plane model refers to the plane expression obtained by transforming the true plane equation to the coordinate system of the LiDAR through the extrinsic parameters to be solved. It is used as a reference geometric model in the calculation of LiDAR point cloud errors.
[0043] Specifically, based on the actual pose sequence and geometry of the reflective target, the true plane equation of the reflective target plane in the coordinate system of the robotic arm base is determined to achieve a unified mathematical expression of the spatial geometric relationship of the reflective target. Based on the extrinsic parameters of the lidar to be solved, the true plane equation is transformed to the lidar coordinate system to obtain a theoretical plane model, providing a unified reference benchmark for the error calculation of the point cloud data. Point cloud data belonging to the reflective target is extracted from the point cloud sequence, and the point cloud data is corrected by combining the intrinsic compensation parameters of the lidar to be solved, resulting in corrected point cloud data. A point-to-surface distance error objective function is constructed based on the distance from each point in the corrected point cloud data to the theoretical plane model. By minimizing this objective function, the joint optimization solution of the lidar intrinsic compensation parameters and the lidar extrinsic parameters is achieved, thereby obtaining the lidar calibration result. This approach transforms the original method of plane fitting based on point cloud itself into a deterministic plane constraint based on the pose determination of the robotic arm. This eliminates the influence of point cloud noise and sparsity on the plane reference, thereby avoiding the parameter instability problem introduced by plane estimation error in traditional methods and improving the accuracy of lidar calibration results.
[0044] In one optional implementation, determining the extrinsic transformation matrix between the camera and the lidar based on the image sequence, the point cloud sequence, the camera calibration results, and the lidar calibration results includes: calculating the initial extrinsic transformation matrix from the lidar to the camera using a coordinate transformation chain based on the camera extrinsic parameters in the camera calibration results and the lidar extrinsic parameters in the lidar calibration results; transforming the point cloud sequence to the camera coordinate system based on the initial extrinsic transformation matrix to obtain a transformed point cloud sequence; and projecting the transformed point cloud sequence onto the camera image plane using the camera intrinsic parameter matrix and camera distortion coefficients from the camera calibration results to obtain a projected point set sequence; and calculating the spatial distance between the points in the projected point set sequence and the edges of the visual target and the reflective target in the image sequence. A cross-modal alignment error objective function is constructed. The reprojection error objective function, the point-to-surface distance error objective function, and the cross-modal alignment error objective function are weighted and summed to construct a joint error objective function. In this joint error objective function, the camera intrinsic parameter matrix and camera distortion coefficients in the reprojection error objective function are determined as known constants based on the camera intrinsic parameter matrix and camera distortion coefficients in the camera calibration results. Similarly, the lidar intrinsic parameter compensation parameters in the point-to-surface distance error objective function are determined as known constants based on the lidar intrinsic parameter compensation parameters in the lidar calibration results. Only the initial extrinsic parameter transformation matrix is optimized. By minimizing the joint error objective function, the initial extrinsic parameter transformation matrix is optimized to obtain the final extrinsic parameter transformation matrix.
[0045] In this embodiment, the initial extrinsic transformation matrix is used to characterize the initial spatial pose relationship between the lidar coordinate system and the camera coordinate system. The transformed point cloud sequence represents the point cloud data sequence transformed to the camera coordinate system by the initial extrinsic transformation matrix. The projected point set sequence represents the two-dimensional point set sequence obtained by projecting the transformed point cloud sequence onto the image plane. The visual target and reflective target edges are used to characterize the boundary regions of the target's outer contour structure in the image. Their determination methods include, but are not limited to: extracting the target's outer contour boundary points through an edge detection algorithm; constructing the boundary contour based on the outermost feature points of a checkerboard or dot array structure; or obtaining the edge positions by geometrically fitting the corresponding region in the image using a pre-known target geometric model, thereby characterizing the spatial distribution range of the target in the image plane.
[0046] The cross-modal alignment error objective function measures the spatial consistency deviation between the point cloud projection result and the edges of the visual and reflective targets in the image. Essentially, it calculates the distance deviation between the projected points and the target structure boundaries in the image after converting the point cloud data to the camera's imaging plane, thus characterizing the spatial consistency between the observation results of the two sensors. The joint error objective function is a comprehensive optimization objective composed of the reprojection error objective function, the point-to-surface distance error objective function, and the cross-modal alignment error objective function, combined with preset weights. The reprojection error objective function and the point-to-surface distance error objective function serve as constraints to maintain the stability and rationality of their respective calibration results, while the cross-modal alignment error objective function serves as the core optimization term to constrain the spatial consistency between the camera and the lidar.
[0047] Specifically, based on the camera extrinsic parameters from the camera calibration results and the lidar extrinsic parameters from the lidar calibration results, an initial extrinsic parameter transformation matrix is obtained through coordinate transformation chain calculation. This matrix is used to characterize the initial spatial pose relationship between the lidar coordinate system and the camera coordinate system. Based on the initial extrinsic parameter transformation matrix, the point cloud sequence is transformed to the camera coordinate system to obtain a transformed point cloud sequence, ensuring that the point cloud data and image data are in a unified coordinate representation system. Combining the camera intrinsic parameter matrix and camera distortion coefficients from the camera calibration results, the transformed point cloud sequence is projected onto the image plane to obtain a projected point set sequence, which is used to analyze the spatial correspondence between the projected point set sequence and the edges of the visual and reflective targets in the image sequence. A cross-modal alignment error objective function is constructed by calculating the spatial distance between the projected point set sequence and the edges of the visual and reflective targets in the image sequence. Simultaneously, the reprojection error objective function, the point-to-surface distance error objective function, and the cross-modal alignment error objective function are weighted and combined to construct a joint error objective function. The initial extrinsic parameter transformation matrix is optimized by minimizing the joint error objective function to obtain the final extrinsic parameter transformation matrix. This approach transforms the spatial relationship between the camera and the lidar from a coarse alignment based on initial chain calculations into a refined adjustment process based on the joint optimization of multiple constraint errors. Cross-modal alignment errors provide spatial consistency constraints between the image and the point cloud, while reprojection errors and point-to-surface distance errors maintain the stability of their respective sensor calibration results. Thus, under the combined effect of multiple constraints, the accuracy and stability of extrinsic parameter solutions are improved, and the cumulative bias caused by initial extrinsic parameter errors is reduced.
[0048] In one optional implementation, the calibration pose sequence is determined as follows: According to a preset hierarchical sampling strategy, spatial positions are sampled hierarchically within the calibration space to obtain a set of spatial positions; for each spatial position in the set of spatial positions, the attitude parameters of the visual target and the reflection target are discretely sampled and combined to obtain an initial calibration pose sequence; wherein, the attitude parameters include pitch angle, yaw angle, and roll angle; based on the robotic arm motion constraints, the initial calibration pose sequence is subjected to reachability verification and collision detection, and unreachable or collision-risk calibration poses are eliminated to obtain the final calibration pose sequence.
[0049] In this embodiment, the calibration space refers to the spatial region jointly defined by the effective sensing range of the camera and lidar, and the reachable range of the robotic arm's movement during the calibration process. It serves to constrain the spatial boundaries for calibration pose generation and data acquisition. A preset hierarchical sampling strategy is used to perform discrete sampling in the calibration space according to spatial position and attitude parameters, and then combine the sampling results to generate candidate calibration poses. The spatial position set refers to the set of spatial position points obtained after discrete sampling within the calibration space according to a preset spatial step size. The initial calibration pose sequence refers to the set of robotic arm calibration poses obtained by combining the spatial position set and the discrete sampling results of the attitude parameters.
[0050] Robotic arm motion constraints are used to limit whether the end-effector pose of the robotic arm meets the actual executable conditions. These constraints include limitations on the workspace range and joint range of motion of the robotic arm. They are used as constraints in the accessibility verification and collision detection process. Accessibility verification is used to determine whether the calibration pose meets the kinematic feasibility conditions of the robotic arm. Collision detection is used to determine whether there is a risk of collision between the robotic arm and its own structure or the external environment during execution, thereby eliminating calibration poses with motion conflict risks.
[0051] Specifically, within the calibration space, spatial positions are discretely sampled according to a preset hierarchical sampling strategy to obtain a set of spatial positions, which is used to characterize the distribution of selectable spatial sampling points during the calibration process. For each spatial position, discrete sampling is combined with the attitude parameters of the visual target and the reflection target to generate an initial calibration pose sequence, which is used to characterize the set of candidate end-effector poses formed by the combination of spatial position and attitude angle. The initial calibration pose sequence is screened based on the motion constraints of the robotic arm. Reachability verification is used to determine whether each calibration pose in the initial calibration pose sequence meets the kinematic realizability conditions of the robotic arm, and collision detection is used to determine whether each calibration pose in the initial calibration pose sequence collides with the robotic arm's own structure or the external environment during execution, thereby eliminating unexecutable calibration poses and obtaining the final calibration pose sequence. This approach transforms pose generation during calibration from an unconstrained random or empirically selected method to an automatic generation method based on spatial range constraints and the executableness constraints of the robotic arm's motion. While ensuring the coverage of the sampling space, it improves the effectiveness of calibration poses, reduces redundant poses that are unexecutable or pose a risk of collision, thereby enhancing the stability of calibration data acquisition and the execution efficiency of the calibration process.
[0052] In one optional implementation, the method further includes: for verification poses that have not participated in calibration optimization, collecting image verification sequences, point cloud verification sequences, and robotic arm end-effector pose verification sequences corresponding to each verification pose; based on the camera calibration results and the lidar calibration results, calculating the camera reprojection error, the distance error from the lidar point cloud to the target plane, and the reprojection error from the lidar point cloud to the image under each verification pose; if any error exceeds a preset error threshold, increasing the number of calibration poses in the calibration pose sequence and re-executing the calibration method.
[0053] In this embodiment, the verification pose refers to the pre-reserved calibration pose from the calibration pose sequence that has not participated in the parameter optimization solution of the calibration process, and is used to verify the accuracy of the calibration results. The image verification sequence, point cloud verification sequence, and robotic arm end-effector pose verification sequence are image data sequences, point cloud data sequences, and robotic arm end-effector spatial position and attitude data sequences respectively acquired by the camera, LiDAR, and robotic arm under each verification pose.
[0054] Camera reprojection error, based on camera calibration results, is the deviation between the predicted pixel position obtained by projecting the 3D coordinates of pre-defined visual feature points on the visual target in the target coordinate system onto the image plane, and the actual pixel position of the corresponding visual feature points extracted from the image verification sequence. The distance error from the LiDAR point cloud to the target plane refers to the spatial distance deviation between each point in the point cloud data belonging to the reflective target extracted in the point cloud verification sequence and the target plane determined based on the actual pose and geometric model of the reflective target. The reprojection error from the LiDAR point cloud to the image refers to the spatial position deviation between the projected points obtained by projecting the point cloud data in the point cloud verification sequence onto the image plane based on camera calibration results, and the corresponding target structure position in the image verification sequence.
[0055] Specifically, for verification poses that did not participate in calibration optimization, image verification sequences, point cloud verification sequences, and robotic arm end-effector pose verification sequences corresponding to each verification pose are collected to construct a verification dataset independent of the calibration optimization process. Based on camera calibration results and LiDAR calibration results, errors are calculated for the data under each verification pose. Specifically, the consistency of the camera calibration results in image space is verified by the camera reprojection error, the consistency between the point cloud and the known geometric model is verified by the distance error from the LiDAR point cloud to the target plane, and the cross-modal space alignment effect is verified by the reprojection error from the LiDAR point cloud to the image. The validity of the calibration results is judged by comparing each error with a preset error threshold. If any error exceeds the preset error threshold, the number of calibration poses in the calibration pose sequence is increased and the calibration method is re-executed to improve the coverage density of the calibration data, thereby increasing the accuracy of the calibration results. This method enables multi-dimensional error verification of calibration results under the verification pose, thereby achieving external consistency verification of calibration accuracy. By driving adaptive adjustment of the number of calibration poses through error feedback, the spatial coverage and constraint strength of calibration data are improved, thus enhancing the reliability of the overall calibration results.
[0056] For example, the acquired image sequence is The robotic arm end-effector pose sequence is as follows: ,in, For the end effector of the robotic arm at the 1st The pose of the target relative to the coordinate system of the robotic arm base under the positioning pose is in the form of a 4×4 homogeneous transformation matrix; the mounting pose of the visual target is... For image sequences Each frame in Sub-pixel corner detection algorithms (such as Harris corner detection combined with sub-pixel refinement, or Stone-Tommasi feature extraction algorithm) are used to extract the set of two-dimensional pixel coordinates of visual feature points in the visual target. ,in, For the first A set of two-dimensional pixel coordinates of visual feature points extracted from a frame image. For the first The first frame of the image Two-dimensional pixel coordinates of a visual feature point This represents the total number of visual feature points detected in each frame of the image. Simultaneously, the three-dimensional coordinates of these visual feature points in the target's local coordinate system are known. (Visual target plane, Z=0). Using the forward kinematics of the robotic arm, the actual pose sequence of the visual target in the robotic arm base coordinate system for each pose is calculated as follows: ;in, The position and pose of the robotic arm's end effector are directly provided by the robotic arm. This refers to the mounting pose of the visual target, which directly transforms the end effector pose of the robotic arm into the true spatial pose of the visual target. Let the extrinsic parameters of the camera in the robotic arm's base coordinate system be... The theoretical pose sequence of the visual target in the camera coordinate system can be expressed as: ; through the end effector pose of the robotic arm Mounting pose of visual target First, obtain the pose of the visual target in the base coordinate system. Then, it is transformed to the camera coordinate system. In this way, the three-dimensional coordinates of the visual target in the camera coordinate system at each pose can be accurately expressed, and thus used to construct the reprojection error. This chain transformation combines the observation data from different poses with the camera extrinsic parameters to be solved. Direct correlation forms a global optimization constraint, avoiding error accumulation. Construct the reprojection error objective function: ,in, This is the sum of squared reprojection errors for each visual feature point in each image of the image sequence. For the first The first frame of the image Two-dimensional pixel coordinates of each corner point The projection function of the camera depends on the camera intrinsic parameter matrix. Camera distortion parameters The pose of the visual target in the camera coordinate system and the three-dimensional coordinates of visual feature points in the visual target coordinate system A nonlinear least squares optimization algorithm is used to optimize the camera intrinsic parameter matrix. Camera distortion parameters and camera external parameters A joint optimization solution is performed. In the traditional Zhang Zhengyou calibration method, the camera extrinsic parameters at each pose are... All are treated as independent unknowns for optimization. A total of 100 images were generated Each external parameter has a degree of freedom. However, in the method proposed in this application, due to... ,and robotic arm end-effector pose and Since the pose of the visual target is known, the only unknown is the globally unique camera extrinsic parameter. Only 6 degrees of freedom need to be solved. For example, in the traditional Zhang Zhengyou calibration method, the total number of parameters to be optimized is 9 + 6N (5 intrinsic parameters + 4 distortion parameters + 6N extrinsic parameters). When acquiring 40 frames of images, the number of parameters reaches 249. The high-dimensional space leads to strong coupling between the focal length f and translation Z, making the algorithm prone to getting trapped in local optima. The method proposed in this application, after introducing the pose of the robotic arm's end effector, reduces the optimization degrees of freedom to 15 (independent of N), bringing the following benefits: complete decoupling of intrinsic and extrinsic parameters: the coupling between focal length and translation Z is broken, allowing for accurate searching in a low-dimensional subspace; a significant improvement in the observation constraint ratio: from approximately 23.0 (N=40, M=70) in the traditional method to 373.3, enhancing statistical reliability; elimination of local optima: the objective function surface is smoother, and the number of algorithm iterations is reduced by more than 50%. Wherein, the observation constraint ratio = total number of observation equations / total number of unknown parameters to be optimized, where the total number of observation equations is provided by each corner point in each frame of the image. and There are two coordinate observations, therefore the total is 2×M×N, where M is the number of corner points in each image frame (e.g., the number of corner points within a chessboard grid), and N is the number of image frames acquired. The total number of unknown parameters to be optimized is 9+6N in the traditional method, and 9+6=15 in this method.
[0057] For example, the collected point cloud sequence is The end effector pose sequence of the robotic arm is The mounting posture of the reflective target is Based on reflectivity thresholds and spatial connectivity, from each frame of point cloud... Extract point cloud data belonging to high reflectivity targets. ,in, For the first The number of points in the target point cloud within the frame. Optional segmentation strategies include: filtering by reflection intensity to retain points with a reflectivity of at least 90%; using Euclidean clustering to remove outliers; and combining the robot arm's end-effector pose information to define a spatial region of interest, further filtering out point cloud data belonging to the reflective target. The spatial region of interest is determined by pre-calculating the theoretical position range of the reflective target in three-dimensional space based on the current pose of the robot arm's end-effector and the target's mounting pose at the end-effector, thus defining a three-dimensional spatial region in the point cloud. The robot arm's end-effector pose is then utilized. Mounting pose of the reflective target on the end effector of the robotic arm Calculate the first The actual pose of the reflective target plane in the coordinate system of the robot arm base under the specified positioning posture. Therefore, the unit normal vector of the reflecting target plane in the coordinate system of the robotic arm base can be directly obtained. and planar distance parameters Thus, the equations of the real plane are established: The accuracy of this true plane equation is entirely determined by the pose of the robotic arm's end effector, rather than relying on the fitting results of the point cloud data, thus ensuring the reliability of the true plane equation. Let the extrinsic parameters of the lidar coordinate system relative to the robotic arm's base coordinate system be... The actual plane equations in the robot arm base coordinate system are transformed to the lidar coordinate system to obtain the theoretical plane model. The original point cloud measurement value of the lidar is the distance. Horizontal angle vertical angle An intrinsic parameter error model is introduced to correct the original point cloud measurements: corrected distance. Correcting the horizontal angle Correcting the vertical angle Convert the corrected polar coordinates to Cartesian coordinates. ,in Let these be the intrinsic parameters of the lidar to be solved. Construct the objective function for the point-to-surface distance error from the point to the reflecting target plane: ,in, This represents the Euclidean distance from a point to a plane. A nonlinear least squares optimization algorithm is used to jointly solve for the intrinsic parameters of the lidar for compensation. With lidar extrinsic parameters Traditional targetless or static target-based methods rely on fitting a plane from a noisy point cloud to calibrate the intrinsic parameters of the lidar. This method suffers from self-explicit circular dependence, meaning that a plane is fitted using a point cloud with intrinsic parameter errors, and then the same fitted plane is used to calibrate the intrinsic parameters. High noise levels lead to unreliable calibration results. The method proposed in this application directly generates the true plane equation of the reflecting target using the pose of the robotic arm's end effector. The accuracy is determined by the positioning accuracy of the robotic arm's end effector, breaking the self-explicit circular dependence. It transforms the point cloud fitting problem into a projection problem from a point to a defined plane, significantly improving the calibration accuracy of the lidar's intrinsic parameter compensation parameters and extrinsic parameters.
[0058] For example, the camera's extrinsic parameters in the robot arm's base coordinate system are obtained separately. and lidar extrinsic parameters The initial extrinsic transformation matrix from lidar to the camera can be directly calculated using the coordinate system chain rule: The robotic arm base serves as a fixed global reference frame, avoiding the singularity problem caused by parallel rotation axes in traditional hand-eye calibration AX=XB. Using this initial extrinsic parameter transformation matrix as the initial value, a joint error objective function is constructed: ;in, , and These are the weight coefficients corresponding to the reprojection error objective function, the point-to-surface distance error objective function, and the cross-modal alignment error objective function, respectively. This is the point cloud sequence corresponding to the visual target and the reflection target. The edges of visual and reflective targets in the image sequence are represented. A nonlinear least squares optimization algorithm is used to optimize the joint error objective function, yielding the final extrinsic parameter transformation matrix. Traditional extrinsic parameter calibration methods involve three steps: perspective n-point positioning and pose estimation, random sampling consistency plane fitting, and point-to-surface registration, which suffers from a progressively amplified three-level error. The method proposed in this application uses the robotic arm base as the reference coordinate system. Based on the camera extrinsic parameters from the camera calibration results and the lidar extrinsic parameters from the lidar calibration results, an initial extrinsic parameter transformation matrix from the lidar to the camera is calculated through a coordinate transformation chain, obtaining a mathematically closed initial solution. Its accuracy depends only on the independent calibration results, avoiding error propagation. The reprojection error objective function, the point-to-surface distance error objective function, and the cross-modal alignment error objective function are weighted and summed to construct a joint error objective function. Cross-modal physical alignment is introduced to achieve sub-pixel level projection accuracy.
[0059] Example 3 Figure 3 This is a schematic diagram of a camera and LiDAR internal and external parameter joint calibration device according to Embodiment 3 of this application. This embodiment is applicable to the joint calibration of camera and LiDAR internal and external parameters in multi-sensor fusion application scenarios. This camera and LiDAR internal and external parameter joint calibration device can be implemented in hardware and / or software, and can be configured in a computer device. Figure 3 As shown, the camera and lidar internal and external parameter joint calibration device 300 includes: The multi-source data synchronous acquisition module 310 is used to acquire the image sequence, point cloud sequence and end-effector pose sequence of the robotic arm in each calibration pose from the pre-generated calibration pose sequence; wherein, the image sequence and the point cloud sequence are acquired by a camera and a lidar fixed at the calibration station to acquire the visual target and the reflective target fixed at the end of the robotic arm, respectively. The camera and lidar intrinsic parameter calibration module 320 is used to construct the reprojection error objective function of the camera and the point-to-surface distance error objective function of the lidar based on the image sequence, the point cloud sequence, the pose sequence of the robotic arm end effector, the installation pose of the visual target, the installation pose of the reflective target, the camera parameters to be solved, and the lidar parameters to be solved, and then optimize and solve them to obtain the camera calibration results and lidar calibration results. The camera and lidar extrinsic parameter calibration module 330 is used to determine the extrinsic parameter transformation matrix between the camera and lidar based on the image sequence, the point cloud sequence, the camera calibration result, and the lidar calibration result.
[0060] In one optional embodiment, the camera and lidar intrinsic parameter calibration module 320 is specifically used for: Based on the end-effector pose sequence and the mounting pose of the visual target, the actual pose sequence of the visual target in the coordinate system of the robot arm base is calculated. Based on the actual pose sequence of the visual target, the image sequence, and the camera parameters to be solved, a reprojection error objective function is constructed and optimized to obtain the camera calibration result. Based on the end-effector pose sequence and the mounting pose of the reflective target, the actual pose sequence of the reflective target in the coordinate system of the robot arm base is calculated. Based on the actual pose sequence of the reflective target, the point cloud sequence, and the parameters of the lidar to be solved, the objective function of the point-to-surface distance error from the point cloud to the reflective target plane is constructed and optimized to obtain the lidar calibration result.
[0061] In one optional embodiment, the camera and lidar intrinsic parameter calibration module 320 is specifically used for: Using the camera extrinsic parameters to be solved, the actual pose sequence of the visual target is transformed to the camera coordinate system to obtain the theoretical pose sequence; Based on the theoretical pose sequence, the camera intrinsic parameter matrix to be solved, and the camera distortion coefficient, the preset visual target feature points on the visual target are projected onto the camera imaging plane to obtain the predicted pixel positions. Extract the actual pixel positions of visual target feature points from the image sequence; Based on the difference between the predicted pixel position and the actual pixel position, a reprojection error objective function is constructed; By minimizing the reprojection error objective function, the camera parameters to be solved are jointly optimized to obtain the camera calibration results.
[0062] In one optional embodiment, the camera and lidar intrinsic parameter calibration module 320 is specifically used for: Based on the actual pose sequence of the reflective target and the geometry of the reflective target, determine the true plane equation of the reflective target plane in the coordinate system of the robotic arm base; Using the extrinsic parameters of the lidar to be solved, the real plane equation is transformed to the lidar coordinate system to obtain the theoretical plane model; Extract point cloud data belonging to the reflection target from the point cloud sequence; Based on the lidar intrinsic parameter compensation parameters to be solved, the point cloud data is corrected to obtain corrected point cloud data; Based on the distance from the points in the corrected point cloud data to the theoretical plane model, a point-to-surface distance error objective function is constructed. By minimizing the objective function of the point-to-surface distance error, the parameters of the lidar to be solved are jointly optimized to obtain the lidar calibration result.
[0063] In one optional implementation, the camera and lidar extrinsic parameter calibration module 330 is specifically used for: Based on the camera extrinsic parameters in the camera calibration results and the lidar extrinsic parameters in the lidar calibration results, the initial extrinsic parameter transformation matrix from lidar to camera is calculated through a coordinate transformation chain. Based on the initial extrinsic transformation matrix, the point cloud sequence is transformed to the camera coordinate system to obtain the transformed point cloud sequence. Then, combined with the camera intrinsic matrix and camera distortion coefficients in the camera calibration results, the transformed point cloud sequence is projected onto the camera image plane to obtain the projected point set sequence. Calculate the spatial distance between points in the projection point set sequence and the edges of visual and reflective targets in the image sequence, and construct a cross-modal alignment error objective function; The reprojection error objective function, the point-to-surface distance error objective function, and the cross-modal alignment error objective function are weighted and summed to construct a joint error objective function. In this joint error objective function, the camera intrinsic parameter matrix and camera distortion coefficients in the reprojection error objective function are known constants based on the camera intrinsic parameter matrix and camera distortion coefficients in the camera calibration results. Similarly, the lidar intrinsic parameter compensation parameters in the point-to-surface distance error objective function are known constants based on the lidar intrinsic parameter compensation parameters in the lidar calibration results. Only the initial extrinsic parameter transformation matrix is optimized. The initial extrinsic transformation matrix is optimized by minimizing the joint error objective function to obtain the final extrinsic transformation matrix.
[0064] In one optional embodiment, the camera and lidar internal and external parameter joint calibration device 300 further includes a calibration pose sequence determination module, which is specifically used for: According to the preset hierarchical sampling strategy, the spatial location is sampled hierarchically within the calibration space to obtain a set of spatial locations; For each spatial location in the set of spatial locations, the attitude parameters of the visual target and the reflected target are discretely sampled and combined to obtain an initial target attitude sequence; wherein, the attitude parameters include pitch angle, yaw angle and roll angle; Based on the robotic arm motion constraints, the initial calibration pose sequence is subjected to reachability verification and collision detection. Unreachable or collision-risk calibration poses are eliminated to obtain the final calibration pose sequence.
[0065] In one optional embodiment, the camera and lidar internal and external parameter joint calibration device 300 further includes a calibration result verification module, which is specifically used for: For verification poses that were not included in the calibration and optimization, image verification sequences, point cloud verification sequences, and robotic arm end-effector pose verification sequences were collected for each verification pose. Based on the camera calibration results and the lidar calibration results, the camera reprojection error, the distance error from the lidar point cloud to the target plane, and the reprojection error from the lidar point cloud to the image are calculated for each verification pose. If any error exceeds a preset error threshold, the number of calibration poses in the calibration pose sequence is increased and the calibration method is re-executed.
[0066] The camera and lidar internal and external parameter joint calibration device provided in this application embodiment can execute the camera and lidar internal and external parameter joint calibration method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the execution method.
[0067] This application also provides an electronic device, a readable storage medium, and a computer program product. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for joint calibration of intrinsic and extrinsic parameters of any camera and lidar according to this application.
[0068] Example 4 Figure 4 This is a schematic diagram of the structure of an electronic device that implements the joint calibration method for the intrinsic and extrinsic parameters of the camera and lidar according to the embodiments of this application. Figure 4 A schematic diagram of an electronic device 410 that can be used to implement embodiments of this application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0069] like Figure 4As shown, the electronic device 410 includes at least one processor 411 and a memory, such as a read-only memory (ROM) 412 or a random access memory (RAM) 413, communicatively connected to the at least one processor 411. The memory stores computer programs executable by the at least one processor. The processor 411 can perform various appropriate actions and processes based on the computer program stored in the ROM 412 or loaded from storage unit 418 into the RAM 413. The RAM 413 may also store various programs and data required for the operation of the electronic device 410. The processor 411, ROM 412, and RAM 413 are interconnected via a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.
[0070] Multiple components in electronic device 410 are connected to I / O interface 415, including: input unit 416, such as keyboard, mouse, etc.; output unit 417, such as various types of displays, speakers, etc.; storage unit 418, such as disk, optical disk, etc.; and communication unit 419, such as network card, modem, wireless transceiver, etc. Communication unit 419 allows electronic device 410 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0071] Processor 411 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 411 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 411 performs the various methods and processes described above, such as the joint calibration method of intrinsic and extrinsic parameters of a camera and LiDAR.
[0072] In some embodiments, the camera and LiDAR combined intrinsic and extrinsic parameter calibration method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 418. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 410 via ROM 412 and / or communication unit 419. When the computer program is loaded into RAM 413 and executed by processor 411, one or more steps of the camera and LiDAR combined intrinsic and extrinsic parameter calibration method described above can be performed. Alternatively, in other embodiments, processor 411 can be configured to perform the camera and LiDAR combined intrinsic and extrinsic parameter calibration method by any other suitable means (e.g., by means of firmware).
[0073] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0074] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0075] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0076] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0077] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0078] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0079] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0080] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for joint calibration of intrinsic and extrinsic parameters of a camera and a lidar, characterized in that, include: The image sequence, point cloud sequence, and end-effector pose sequence of the robotic arm at each calibration pose are obtained from the pre-generated calibration pose sequence; wherein, the image sequence and the point cloud sequence are obtained by acquiring visual targets and reflective targets fixed at the end of the robotic arm by a camera and a lidar fixed at the calibration station, respectively. Based on the image sequence, the point cloud sequence, the robotic arm end-effector pose sequence, the installation pose of the visual target, the installation pose of the reflective target, the camera parameters to be solved, and the lidar parameters to be solved, the objective function for the reprojection error of the camera and the objective function for the point-to-surface distance error of the lidar are constructed respectively, and optimized to obtain the camera calibration results and lidar calibration results. Based on the image sequence, the point cloud sequence, the camera calibration results, and the lidar calibration results, the extrinsic parameter transformation matrix between the camera and the lidar is determined.
2. The method according to claim 1, characterized in that, The process involves constructing objective functions for camera reprojection error and LiDAR point-to-surface distance error, respectively, based on the image sequence, point cloud sequence, robotic arm end-effector pose sequence, installation pose of the visual target, installation pose of the reflective target, camera parameters to be solved, and LiDAR parameters to be solved. These functions are then optimized to obtain camera calibration results and LiDAR calibration results. Based on the end-effector pose sequence and the mounting pose of the visual target, the actual pose sequence of the visual target in the coordinate system of the robot arm base is calculated. Based on the actual pose sequence of the visual target, the image sequence, and the camera parameters to be solved, a reprojection error objective function is constructed and optimized to obtain the camera calibration result. Based on the end-effector pose sequence and the mounting pose of the reflective target, the actual pose sequence of the reflective target in the coordinate system of the robot arm base is calculated. Based on the actual pose sequence of the reflective target, the point cloud sequence, and the parameters of the lidar to be solved, the objective function of the point-to-surface distance error from the point cloud to the reflective target plane is constructed and optimized to obtain the lidar calibration result.
3. The method according to claim 2, characterized in that, The camera parameters to be solved include the camera intrinsic matrix, camera distortion coefficients, and camera extrinsic parameters; the camera extrinsic parameters are the camera pose parameters in the coordinate system of the robotic arm base; the process of constructing a reprojection error objective function based on the actual pose sequence of the visual target, the image sequence, and the camera parameters to be solved, and optimizing the solution to obtain the camera calibration result, includes: Using the camera extrinsic parameters to be solved, the actual pose sequence of the visual target is transformed to the camera coordinate system to obtain the theoretical pose sequence; Based on the theoretical pose sequence, the camera intrinsic parameter matrix to be solved, and the camera distortion coefficient, the preset visual target feature points on the visual target are projected onto the camera imaging plane to obtain the predicted pixel positions. Extract the actual pixel positions of visual target feature points from the image sequence; Based on the difference between the predicted pixel position and the actual pixel position, a reprojection error objective function is constructed; By minimizing the reprojection error objective function, the camera parameters to be solved are jointly optimized to obtain the camera calibration results.
4. The method according to claim 2, characterized in that, The lidar parameters to be solved include lidar intrinsic compensation parameters and lidar extrinsic parameters; the lidar extrinsic parameters are the pose parameters of the lidar in the coordinate system of the robotic arm base; the objective function of the point-to-surface distance error from the point cloud to the plane of the reflecting target is constructed based on the actual pose sequence of the reflecting target, the point cloud sequence, and the lidar parameters to be solved, and optimized to obtain the lidar calibration result, including: Based on the actual pose sequence of the reflective target and the geometry of the reflective target, determine the true plane equation of the reflective target plane in the coordinate system of the robotic arm base; Using the extrinsic parameters of the lidar to be solved, the real plane equation is transformed to the lidar coordinate system to obtain the theoretical plane model; Extract point cloud data belonging to the reflection target from the point cloud sequence; Based on the lidar intrinsic parameter compensation parameters to be solved, the point cloud data is corrected to obtain corrected point cloud data; Based on the distance from the points in the corrected point cloud data to the theoretical plane model, a point-to-surface distance error objective function is constructed. By minimizing the objective function of the point-to-surface distance error, the parameters of the lidar to be solved are jointly optimized to obtain the lidar calibration result.
5. The method according to claim 1, characterized in that, The step of determining the extrinsic parameter transformation matrix between the camera and the lidar based on the image sequence, the point cloud sequence, the camera calibration result, and the lidar calibration result includes: Based on the camera extrinsic parameters in the camera calibration results and the lidar extrinsic parameters in the lidar calibration results, the initial extrinsic parameter transformation matrix from lidar to camera is calculated through a coordinate transformation chain. Based on the initial extrinsic transformation matrix, the point cloud sequence is transformed to the camera coordinate system to obtain the transformed point cloud sequence. Then, combined with the camera intrinsic matrix and camera distortion coefficients in the camera calibration results, the transformed point cloud sequence is projected onto the camera image plane to obtain the projected point set sequence. Calculate the spatial distance between points in the projection point set sequence and the edges of visual and reflective targets in the image sequence, and construct a cross-modal alignment error objective function; The reprojection error objective function, the point-to-surface distance error objective function, and the cross-modal alignment error objective function are weighted and summed to construct a joint error objective function. In this joint error objective function, the camera intrinsic parameter matrix and camera distortion coefficients in the reprojection error objective function are known constants based on the camera intrinsic parameter matrix and camera distortion coefficients in the camera calibration results. Similarly, the lidar intrinsic parameter compensation parameters in the point-to-surface distance error objective function are known constants based on the lidar intrinsic parameter compensation parameters in the lidar calibration results. Only the initial extrinsic parameter transformation matrix is optimized. The initial extrinsic transformation matrix is optimized by minimizing the joint error objective function to obtain the final extrinsic transformation matrix.
6. The method according to claim 1, characterized in that, The calibration pose sequence is determined as follows: According to the preset hierarchical sampling strategy, the spatial location is sampled hierarchically within the calibration space to obtain a set of spatial locations; For each spatial location in the set of spatial locations, the attitude parameters of the visual target and the reflected target are discretely sampled and combined to obtain an initial target attitude sequence; wherein, the attitude parameters include pitch angle, yaw angle and roll angle; Based on the robotic arm motion constraints, the initial calibration pose sequence is subjected to reachability verification and collision detection. Unreachable or collision-risk calibration poses are eliminated to obtain the final calibration pose sequence.
7. The method according to claim 1, characterized in that, The method further includes: For verification poses that were not included in the calibration and optimization, image verification sequences, point cloud verification sequences, and robotic arm end-effector pose verification sequences were collected for each verification pose. Based on the camera calibration results and the lidar calibration results, the camera reprojection error, the distance error from the lidar point cloud to the target plane, and the reprojection error from the lidar point cloud to the image are calculated for each verification pose. If any error exceeds a preset error threshold, the number of calibration poses in the calibration pose sequence is increased and the calibration method is re-executed.
8. A device for joint calibration of intrinsic and extrinsic parameters of a camera and a lidar, characterized in that, include: The multi-source data synchronous acquisition module is used to acquire image sequences, point cloud sequences, and end-effector pose sequences of the robotic arm at each calibration pose from a pre-generated calibration pose sequence; wherein, the image sequences and the point cloud sequences are acquired by a camera and a lidar fixed at the calibration station to acquire visual targets and reflective targets fixed at the end of the robotic arm, respectively. The camera and lidar intrinsic parameter calibration module is used to construct the reprojection error objective function of the camera and the point-to-surface distance error objective function of the lidar based on the image sequence, the point cloud sequence, the pose sequence of the robotic arm end effector, the installation pose of the visual target, the installation pose of the reflective target, the camera parameters to be solved, and the lidar parameters to be solved, and then optimize and solve them to obtain the camera calibration results and lidar calibration results. The camera and lidar extrinsic parameter calibration module is used to determine the extrinsic parameter transformation matrix between the camera and lidar based on the image sequence, the point cloud sequence, the camera calibration result, and the lidar calibration result.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the camera and lidar joint calibration method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the camera and lidar joint calibration method for intrinsic and extrinsic parameters as described in any one of claims 1-7.