Multi-camera multi-laser radar joint calibration method and device and electronic equipment
By using a joint calibration method involving multiple cameras and multiple lidars, target scene data is acquired, pose is generated, and a joint optimization problem is constructed. This solves the problem of low automation in production line calibration methods, achieves high-precision calibration parameter calculation, and is suitable for complex scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-04-03
AI Technical Summary
Existing production line calibration methods have low automation levels, cannot automatically adapt to target movement and calibration site changes, and have low calibration parameter accuracy, posing challenges, especially in sensor layout schemes with no common field of view or a small common field of view.
A joint calibration method using multiple cameras and multiple lidar is adopted. By acquiring target scene images and point cloud data from cameras and lidar, the pose of the target and the pose of the reference frame camera coordinate system are generated, the initial calibration parameters are calculated, and a joint optimization problem is constructed to solve the target calibration parameters. Reprojection error and point cloud matching error constraints are added, which is suitable for scenarios where the target changes frequently and the calibration site changes.
It significantly improves the accuracy of calibration parameters between multiple cameras and multiple lidars, realizes automated calibration without the need to build target maps in advance, adapts to frequent target changes and calibration site changes, and improves calibration efficiency and accuracy.
Smart Images

Figure CN121788620A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of production line calibration technology, and in particular to a joint calibration method, apparatus and electronic equipment for multiple cameras and multiple lidar. Background Technology
[0002] With the rapid development of autonomous driving technology and advanced driver-assistance systems (ADAS), onboard sensors, as core components of environmental perception, directly determine the safety and reliability of the system through their performance and collaborative capabilities.
[0003] Currently, cameras and LiDAR are the two most commonly used types of sensors in autonomous driving systems: cameras provide rich texture and color information, suitable for visual tasks such as target recognition and lane detection; LiDAR acquires high-precision 3D point cloud data by actively emitting laser beams, possessing strong ranging and spatial perception capabilities. By fusing data from multiple sensors, a more comprehensive and accurate perception of the surrounding environment can be achieved, thereby improving the robustness and safety of autonomous driving decisions.
[0004] In practical applications, autonomous vehicles are typically equipped with multiple sensors, including cameras, LiDAR, millimeter-wave radar, and inertial measurement units. These sensors collect environmental information from different angles and positions to achieve 360-degree coverage around the vehicle without blind spots. However, due to vehicle structural limitations and field-of-view layout requirements, the overlapping areas of the observation fields of the various sensors are small or even completely nonexistent, posing a significant challenge to data fusion between multiple sensors. To achieve cross-sensor information alignment and fusion, the extrinsic parameters (i.e., relative installation position and attitude) of each sensor must be accurately calibrated beforehand to ensure that observation data from different coordinate systems can be unified into a single reference coordinate system for processing.
[0005] During the vehicle production line phase, the actual installation position and angle of sensors often have manufacturing tolerances, thus requiring factory calibration. Traditional production line calibration methods typically rely on placing high-precision targets at fixed locations and using pre-built target maps to assist in sensor positioning and extrinsic parameter calculation. However, these production line calibration methods are not highly automated, inefficient, and cannot automatically adapt to target movement or changes in the calibration site, nor can they achieve complete automation of the calibration process.
[0006] Therefore, how to solve the problems of low automation in existing production line calibration methods, inability to automatically adapt to target movement and calibration site changes, and low accuracy of calibration parameters is an important issue that urgently needs to be addressed in the field of production line calibration. Summary of the Invention
[0007] This invention provides a joint calibration method, apparatus, and electronic device for multiple cameras and multiple lidars, which overcomes the shortcomings of existing production line calibration methods, such as low automation, inability to automatically adapt to target movement and calibration site changes, and low calibration parameter accuracy. It eliminates the need to build a target map in advance and is applicable not only to sensor layout schemes with no common field of view or a small common field of view, but also to scenarios with frequent target changes and calibration site changes, significantly improving the calibration parameter accuracy between multiple cameras and multiple lidars.
[0008] On one hand, the present invention provides a joint calibration method for multiple cameras and multiple lidars, comprising: acquiring target scene images collected by multiple cameras and target scene point clouds collected by multiple lidars; acquiring target detection results of the target scene images collected by each camera, and generating camera pose and target pose in a reference frame camera coordinate system based on the target detection results; acquiring target point cloud clusters of the target scene point clouds collected by each lidar; calculating initial calibration parameters of the multiple cameras and the multiple lidars based on the camera pose and target pose of the multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple lidars; constructing and solving a joint optimization problem based on the initial calibration parameters of the multiple cameras and the multiple lidars to obtain target calibration parameters of the multiple cameras and the multiple lidars; wherein the joint optimization problem uses camera pose, target pose, and initial calibration parameters of each camera and lidar as optimization variables, and adds reprojection error constraints and point cloud matching error constraints.
[0009] Furthermore, the step of obtaining the target detection result of the target scene image acquired by each camera includes: performing target detection on the target scene image acquired by each camera to obtain the target detection result; wherein, the target detection result includes the pose of the target center and the three-dimensional coordinates of the target vertex in the camera coordinate system where the camera is located, as well as the target identifier.
[0010] Further, generating the camera pose and target pose in the reference frame camera coordinate system based on the target detection results includes: optimizing the pixel projection error of the same target vertex in different frame target scene images using perspective n-point calculation to obtain the camera poses of the multiple cameras in different frame target scene images; constructing a camera pose relationship diagram between the cameras in different frame target scene images based on the relative pose changes of each camera in different frame target scene images; determining the initial camera pose of the multiple cameras in the reference frame camera coordinate system based on the camera pose relationship diagram corresponding to the multiple cameras; calculating the initial target pose of each target in the reference frame camera coordinate system based on the target detection results and the initial camera poses of the multiple cameras in the reference frame camera coordinate system; constructing and solving a pose optimization problem based on the initial camera pose, the initial target pose, and the three-dimensional coordinates of the target vertex to obtain the camera pose and target pose in the reference frame camera coordinate system; wherein the pose optimization problem uses the initial camera pose and the initial target pose as optimization variables and has reprojection error constraints.
[0011] Further, the step of acquiring the target point cloud clusters of the target scene point cloud collected by each lidar includes: converting the target scene point cloud collected by each lidar into a two-dimensional depth map; performing a breadth-first search on the two-dimensional depth map based on the depth continuity of the two-dimensional depth map to obtain point cloud segmentation results, the point cloud segmentation results including multiple point cloud segmentation clusters; determining the center point and distribution covariance value of each point cloud segmentation cluster, and performing principal component analysis on the distribution covariance value to obtain the eigenvalues corresponding to each point cloud segmentation cluster; and selecting target point cloud clusters with target shapes from the point cloud segmentation results based on the eigenvalues corresponding to each point cloud segmentation cluster.
[0012] Further, the step of calculating the initial calibration parameters of the multiple cameras and the multiple lidars based on the camera poses and target poses of the multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple lidars, includes: determining the main camera based on the number of targets observed by each of the multiple cameras, wherein the main camera is one of the multiple cameras, and the coordinate system of the main camera is the main camera coordinate system; using the degree of overlap between the target pose observed by the non-main camera and the target pose in the main camera coordinate system as the target cost function to optimize the camera pose of the non-main camera, thereby obtaining the initial calibration parameters of the non-main camera relative to the main camera; wherein, the non-main camera refers to the multiple cameras other than the main camera.
[0013] Further, the step of calculating the initial calibration parameters of the multiple cameras and the multiple lidars based on the camera poses and target poses of the multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple lidars, includes: calculating the three-dimensional coordinates of all target vertices in the main camera coordinate system based on the initial calibration parameters of the non-main camera relative to the main camera, to obtain the target vertex point cloud in the main camera coordinate system; determining the center point of the target point cloud cluster observed by each lidar, and optimizing the degree of matching between it and the center point of the corresponding target vertex point cloud in the main camera coordinate system as the target cost function, to obtain the initial calibration parameters of the multiple lidars relative to the main camera.
[0014] Secondly, the present invention also provides a joint calibration device for multiple cameras and multiple lidars, comprising: a target scene image and point cloud acquisition module, used to acquire target scene images collected by multiple cameras and target scene point clouds collected by multiple lidars; a camera pose and target pose generation module, used to acquire target detection results of target scene images acquired by each camera, and generate camera pose and target pose in a reference frame camera coordinate system based on the target detection results; a target point cloud cluster acquisition module, used to acquire target point cloud clusters of target scene point clouds acquired by each lidar; and an initial calibration parameter calculation module, used to calculate the initial calibration parameters based on the multiple camera poses and target poses of each lidar. The system calculates the initial calibration parameters of the multiple cameras and multiple lidars based on the camera pose and target pose in the reference frame camera coordinate system, as well as the target point cloud clusters corresponding to the multiple lidars. A multi-camera and lidar joint calibration module is used to construct and solve a joint optimization problem based on the initial calibration parameters of the multiple cameras and multiple lidars to obtain the target calibration parameters of the multiple cameras and multiple lidars. The joint optimization problem uses the camera pose, target pose, and the initial calibration parameters of each camera and lidar as optimization variables, and includes reprojection error constraints and point cloud matching error constraints.
[0015] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the joint calibration method for multiple cameras and multiple lidar as described above.
[0016] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the joint calibration method for multiple cameras and multiple lidar as described above.
[0017] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the joint calibration method for multiple cameras and multiple lidar as described above.
[0018] The present invention provides a joint calibration method for multiple cameras and multiple lidars, which acquires target scene images collected by multiple cameras and target scene point clouds collected by multiple lidars; acquires the target detection results of the target scene images collected by each camera, and generates camera pose and target pose in the camera coordinate system of the reference frame based on the target detection results; acquires the target point cloud clusters of the target scene point clouds collected by each lidar; calculates the initial calibration parameters of multiple cameras and multiple lidars based on the camera poses and target poses of multiple cameras in the camera coordinate system of the reference frame, and the target point cloud clusters corresponding to multiple lidars; and constructs and solves a joint optimization problem based on the initial calibration parameters of multiple cameras and multiple lidars to obtain the target calibration parameters of multiple cameras and multiple lidars. The joint optimization problem uses camera pose, target pose, and the initial calibration parameters of each camera and lidar as optimization variables, and adds reprojection error constraints and point cloud matching error constraints. This method eliminates the need to build a target map in advance. It is applicable not only to sensor deployment schemes with no common field of view or a small common field of view, but also to scenarios where the target changes frequently and the calibration site changes, significantly improving the accuracy of calibration parameters among multiple cameras and multiple lidars. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the joint calibration method for multiple cameras and multiple lidar provided in an embodiment of the present invention.
[0021] Figure 2 This is a schematic diagram of the structure of the multi-camera and multi-lidar joint calibration device provided in an embodiment of the present invention.
[0022] Figure 3 This is a schematic diagram of the physical structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0024] It should be noted that existing production line calibration methods, in order to address the problems caused by the lack of common vision or small common vision between sensors, mostly rely on fixed targets and pre-built target maps, using accurately estimated target coordinates to assist in sensor positioning. This production line calibration method suffers from low automation, low efficiency, inability to automatically adapt to target movement and changes in calibration site, and inability to achieve fully automated calibration.
[0025] To address this, this invention proposes a highly automated joint calibration method for multi-camera and multi-LiDAR systems that eliminates the need for pre-constructed target maps and is suitable for situations involving frequent changes in production line environments and targets, and where there is little or no common field of view. Specifically, Figure 1 A flowchart illustrating the joint calibration method for multiple cameras and multiple lidar provided in an embodiment of the present invention is shown.
[0026] like Figure 1 As shown, the method includes: S110, acquiring target scene images collected by multiple cameras and target scene point clouds collected by multiple lidars; S120, acquiring target detection results of the target scene images collected by each camera, and generating camera pose and target pose in the reference frame camera coordinate system based on the target detection results; S130, acquiring target point cloud clusters of the target scene point clouds collected by each lidar; S140, calculating initial calibration parameters of the multiple cameras and the multiple lidars based on the camera pose and target pose of the multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple lidars; S150, constructing and solving a joint optimization problem based on the initial calibration parameters of the multiple cameras and the multiple lidars to obtain target calibration parameters of the multiple cameras and the multiple lidars; wherein, the joint optimization problem uses camera pose, target pose, and initial calibration parameters of each camera and lidar as optimization variables, and adds reprojection error constraints and point cloud matching error constraints.
[0027] The following will provide a detailed description of steps S110-S150 and related steps.
[0028] S110 acquires target scene images from multiple cameras and target scene point clouds from multiple lidar sensors.
[0029] It is easy to understand that when implementing the multi-camera and multi-lidar joint calibration method, it is first necessary to obtain the observation data of the multi-camera and multi-lidar in the actual working scenario.
[0030] Specifically, within the calibration area surrounding the vehicle, multiple fixed targets marked with AprilTags are evenly distributed within the field of view of each camera and LiDAR. These targets are securely mounted on the ground or walls, with stable positions and a reasonable distribution, ensuring they can be observed simultaneously or in shifts by the various cameras and LiDARs around the vehicle. The AprilTag is a visual marker similar to a two-dimensional barcode, possessing excellent geometric features and a unique encoding, facilitating accurate identification and positioning in images. Each AprilTag has a unique ID, allowing it to be used to distinguish different markers or locations.
[0031] Subsequently, the vehicle is controlled to move slowly back and forth in a straight line within the calibration area at a preset speed (e.g., a low speed of about 3 km / h) for a certain duration (e.g., about 30 seconds). This movement process allows each camera and lidar to observe one or more targets from different angles and distances, thereby increasing the diversity and redundancy of the observation data, which is beneficial to improving the accuracy and robustness of subsequent calibration.
[0032] During vehicle movement, sensor data from all cameras and LiDAR are simultaneously acquired. Specifically, each camera continuously captures scene images containing the AprilTag target at a certain frame rate, recording the two-dimensional visual information of the target in different poses; simultaneously, each LiDAR continuously scans the surrounding environment, acquiring three-dimensional point cloud data including the target's reflection area. All sensor data from the cameras and LiDAR are precisely timestamped to ensure cross-modal data synchronization.
[0033] Through the above process, multiple sets of time-aligned sensor data can be obtained, namely, target scene images acquired by multiple cameras at the same frame rate and target scene (3D) point clouds acquired synchronously by multiple lidars.
[0034] It should be noted that the number of cameras and LiDAR in this step can be set according to actual needs, and no specific limit is made here.
[0035] Based on the target scene images acquired by multiple cameras and the target scene point clouds acquired by multiple lidar sensors in step S110, steps S120-S130 are further executed. Steps S120 and S130 are two independent steps, which can be executed in parallel or sequentially, without specific limitations here.
[0036] S120: Obtain the target detection result of the target scene image captured by each camera, and generate the camera pose and target pose in the reference frame camera coordinate system based on the target detection result.
[0037] After acquiring the target scene images, it is necessary to perform target detection operation frame by frame on the target scene images (sequences) acquired by each camera in order to extract the visual feature information contained therein.
[0038] Specifically, the target scene image can be preprocessed first, including but not limited to grayscale conversion and filtering / denoising, to improve the accuracy of subsequent target detection. Then, edge- or contour-based detection algorithms or an AprilTag target detection module can be used to locate all possible AprilTag regions in the target scene image.
[0039] For successfully detected AprilTags, their IDs are recorded, and the 3D coordinates of their four target corner points in the camera coordinate system and the pose of the target center in the camera coordinate system are extracted. The target corner points are the vertices where black and white squares intersect, possessing clear geometric features that can be used for subsequent pose estimation. Since each target carries a known geometric structure and a uniquely encoded AprilTag, target scene images can be stably detected even under conditions of uneven lighting, tilted viewpoints, or partial occlusion.
[0040] By iterating through all target scene images acquired by each camera, the target detection operation is completed one by one, ultimately obtaining the complete observation results of the target by each camera at different times and from different perspectives, i.e., the target detection results. The results include the target's ID identifier, the pose of the target center, and the three-dimensional coordinates of the target vertices (target corners) in each target scene image.
[0041] Furthermore, based on the target detection results of multiple frames and multiple viewpoints of the target scene images from each camera, the SFM (Structure From Motion) algorithm is used to recover the camera's motion trajectory (composed of the camera's pose in multiple consecutive frames of target scene images) and the target's pose in three-dimensional space (i.e., position and orientation), thus obtaining the camera pose and target pose in the reference frame camera coordinate system. The specific process here will be elaborated in detail in the following embodiments.
[0042] SFM is a technique that uses a series of two-dimensional image sequences to reconstruct the three-dimensional structure of a scene. It is used to recover the three-dimensional structure of a scene and the position and orientation of the camera from a set of images taken by a moving camera.
[0043] The reference frame camera coordinate system can be the initial frame camera coordinate system or other frame camera coordinate systems; no specific limitation is made here. The initial frame camera coordinate system is the camera coordinate system in which the first frame of the target scene image captured by each camera is located.
[0044] S130: Obtain the target point cloud clusters from the target scene point clouds collected by each lidar.
[0045] After completing the target scene point cloud acquisition, it is necessary to perform point cloud segmentation and detection processing on a time frame basis for the target scene point cloud acquired by each lidar, so as to obtain the point cloud corresponding to the target area in the target scene point cloud of each time frame, i.e., the target point cloud cluster.
[0046] Specifically, each frame of the target scene point cloud is first converted into a depth map. Based on the criterion of depth continuity, a breadth-first search is performed on the depth map to obtain multiple point cloud segmentation clusters. Then, geometric feature analysis is performed on these multiple point cloud segmentation clusters to determine whether each cluster corresponds to the actual target. Finally, only the point cloud clusters corresponding to the actual target are retained. This process will be detailed in the following embodiments.
[0047] By traversing all time frames of the target scene point cloud of all lidars and performing point cloud segmentation and detection processing operations one by one, the target point cloud clusters (sets) collected by each lidar at different observation times can be obtained, thus obtaining the target point cloud clusters (sets) of the target scene point cloud collected by each lidar.
[0048] In step S120, the target detection results of the target scene images acquired by each camera are obtained, and the camera pose and target pose in the reference frame camera coordinate system are generated based on the target detection results. In step S130, based on the target point cloud clusters of the target scene point clouds acquired by each lidar, step S140 is further executed.
[0049] S140, calculate the initial calibration parameters of the multiple cameras and the multiple lidars based on the camera pose and target pose of the multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple lidars.
[0050] Specifically, based on the camera pose and target pose of each camera in the reference frame camera coordinate system, and the target point cloud clusters observed by each lidar, the initial calibration parameters of each camera relative to the reference camera, and the initial calibration parameters of each lidar relative to the reference camera, are calculated using the ICP (Iterative Closest Point) algorithm. This process will be explained in detail in the following embodiments.
[0051] Among them, the ICP algorithm is a point cloud registration method based on iterative optimization. By continuously matching the nearest points, it finds the optimal rigid body transformation (including rotation and translation) between two point clouds so that the two point clouds are aligned as much as possible in space.
[0052] The reference camera can be any one of multiple cameras, or it can be a camera selected from multiple cameras according to preset rules; no specific limitation is made here. For example, preferably, the reference camera is the main camera, which is the camera that observes the most targets.
[0053] In step S140, based on the camera poses and target poses of multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to multiple lidars, the initial calibration parameters of multiple cameras and multiple lidars are calculated, and then step S150 is further executed.
[0054] S150, based on the initial calibration parameters of the multiple cameras and the multiple lidars, a joint optimization problem is constructed and solved to obtain the target calibration parameters of the multiple cameras and the multiple lidars; wherein, the joint optimization problem uses the camera pose, the target pose, and the initial calibration parameters of each camera and lidar as optimization variables, and adds reprojection error constraints and point cloud matching error constraints.
[0055] After obtaining the initial calibration parameters of each camera and each lidar relative to the reference camera, a joint optimization method is used to minimize the reprojection error of all target corner points in all frames of camera images (target scene images) and minimize the distance between the center point of all targets and the center point of the target point cloud cluster observed by the lidar. This allows for the calculation of more accurate calibration parameters, target pose, and camera pose trajectory.
[0056] Specifically, a time-aligned approach is used to establish the correspondence between all target scene images, all target detection results, and all target scene point clouds. Based on this correspondence, a joint optimization problem is constructed using the camera pose and target pose determined in step S120, and the initial calibration parameters of each camera and LiDAR determined in step S130, as optimization variables. Furthermore, reprojection error constraints and point cloud matching error constraints are added to the joint optimization problem.
[0057] Among them, the reprojection error constraint requires that the reprojection error of all targets in the target scene image be minimized, and the point cloud matching error constraint requires that the distance between the center point of all targets and the center point of the corresponding target point cloud cluster observed by the lidar be minimized.
[0058] By iteratively optimizing the constructed joint optimization problem until the optimization variables become relatively small, the joint optimization problem converges, yielding target calibration parameters for multiple cameras and multiple LiDARs. Furthermore, solving the joint optimization problem can also yield more accurate target poses and camera pose trajectories.
[0059] In this embodiment, target scene images acquired by multiple cameras and target scene point clouds acquired by multiple lidars are obtained; target detection results of the target scene images acquired by each camera are obtained, and camera pose and target pose in the reference frame camera coordinate system are generated based on the target detection results; target point cloud clusters of the target scene point clouds acquired by each lidar are obtained; initial calibration parameters of multiple cameras and multiple lidars are calculated based on the camera pose and target pose of multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to multiple lidars; based on the initial calibration parameters of multiple cameras and multiple lidars, a joint optimization problem is constructed and solved to obtain the target calibration parameters of multiple cameras and multiple lidars; wherein, the joint optimization problem uses camera pose, target pose, and initial calibration parameters of each camera and lidar as optimization variables, and adds reprojection error constraints and point cloud matching error constraints. This method eliminates the need to build a target map in advance. It is applicable not only to sensor deployment schemes with no common field of view or a small common field of view, but also to scenarios where the target changes frequently and the calibration site changes, significantly improving the accuracy of calibration parameters among multiple cameras and multiple lidars.
[0060] Based on the above embodiments, the following will further describe in detail the process of obtaining the target detection results of the target scene images captured by each camera in step S120.
[0061] Obtaining the target detection results of the target scene images acquired by each camera includes: inputting the target scene images acquired by each camera into the target detection module, performing target detection on the target scene images acquired by each camera, and obtaining the corresponding target detection results; wherein, the target detection results include the pose of the target center and the three-dimensional coordinates of the target vertex in the camera coordinate system where the camera is located, the target identifier ID, and the two-dimensional pixel coordinates of the target vertex in the image.
[0062] It is easy to understand that for all target scene images captured by each camera, each frame of the target scene image is input into the target detection module. This module is specifically used to identify AprilTag targets with specific coding structures in the target scene images and extract their key geometric and semantic information.
[0063] Specifically, the target detection module first preprocesses the input target scene image, performing processes such as grayscale conversion, image enhancement, and noise filtering to improve detection robustness under complex lighting or low-resolution conditions. Then, using edge detection and contour analysis algorithms, it locates potential target marker regions in the target scene image and records them as candidate target regions. For each candidate target region, perspective correction and binarization are further performed, and its internal AprilTag is decoded to determine whether the candidate target region is a valid AprilTag target. The unique code of the AprilTag, i.e., the target identifier, is then read. This identifier is used to distinguish targets in different locations or of different types.
[0064] For a valid AprilTag target, the target detection module further extracts the pixel coordinates of the four target corners / vertices in the corresponding target scene image. These corners are vertices where black and white squares intersect, possessing high contrast and well-defined geometric features, making them suitable as reference points for high-precision positioning.
[0065] Furthermore, by combining the camera's intrinsic parameters (such as focal length) and the target's physical dimensions, the Perspective N-Point (PnP) method can be used to solve for the target's 6-DOF pose relative to the current camera, i.e., the pose of the target center, or the target pose when the local coordinate system has the target's geometric center as its origin. The target center pose represents the rigid body transformation from the target's local coordinate system to the current camera coordinate system, including rotation and translation parameters.
[0066] At the same time, based on the known target geometry, such as the target's side length and the theoretical three-dimensional coordinates of each corner point in the target's local coordinate system, the three-dimensional coordinates of each target vertex in the current coordinate system can be calculated in reverse according to the pose of the target center.
[0067] Therefore, for each frame of the target scene image, the corresponding target detection results can be obtained, including the detected target identifier, the center pose of the target in the current camera coordinate system (the pose of the target center), and the three-dimensional coordinates of each target vertex in the current camera coordinate system. In addition, the target detection results may also include the two-dimensional pixel coordinates of each target vertex.
[0068] By traversing all frames of target scene images from all cameras, the target detection process is completed one by one, and the complete observation results of each camera on each target at different times are obtained.
[0069] Based on the above embodiments, the following will further describe in detail the process of generating the camera pose and target pose in the reference frame camera coordinate system according to the target detection results in step S120.
[0070] Based on the target detection results, the camera pose and target pose in the reference frame camera coordinate system are generated, including: optimizing the pixel projection error of the same target vertex in different frame target scene images using perspective n-point solution, obtaining the camera poses of multiple cameras in different frame target scene images (i.e., the camera pose of each camera in different frame target scene images); constructing a camera pose relationship graph between different frame target scene images based on the relative pose changes of each camera in different frame target scene images; and determining the multiple cameras in the reference frame camera coordinate system based on the camera pose relationship graph corresponding to multiple cameras. The initial camera pose in the reference frame camera coordinate system (i.e., the camera pose trajectory / camera motion trajectory of each camera in the reference frame camera coordinate system); based on the target detection results and the initial camera poses of the multiple cameras in the reference frame camera coordinate system, the initial target pose of each target in the reference frame camera coordinate system is calculated; based on the initial camera pose, the initial target pose, and the three-dimensional coordinates of the target vertices, a pose optimization problem is constructed and solved to obtain the camera pose and target pose in the reference frame camera coordinate system; wherein, the pose optimization problem uses the initial camera pose and the initial target pose as optimization variables and has reprojection error constraints.
[0071] First, for each target scene image detected in a frame, the PnP solution method is used to calculate the pose of the current frame camera coordinate system relative to other frames, i.e., the camera pose in the current frame camera coordinate system, based on the two-dimensional pixel coordinates of the target vertex in the target scene image and the theoretical three-dimensional coordinates of the target vertex in the target local coordinate system.
[0072] To improve estimation accuracy, a nonlinear optimization strategy is introduced to minimize the pixel projection error of the target vertex after it is projected from 3D space onto the target scene image plane, i.e., the Euclidean distance between the actual detected target vertex and the reprojected corner point. Through this optimization process, the camera pose corresponding to each frame of the target scene image is obtained.
[0073] Because of the temporal sequence relationship between multiple frames of target scene images, adjacent frames often observe the same target. Based on this, the relative change in camera pose between adjacent frames of target scene images is calculated, i.e., the rotational transformation domain translational transformation of the current frame's camera coordinate system relative to the previous frame's camera coordinate system. Treating each frame of the target scene image as a point and the relative pose changes between different frames as edges, a camera pose relationship graph can be constructed. This graph describes the continuous evolution of the pose of a single camera during motion.
[0074] For multiple cameras, construct their respective camera pose relationship graphs. Since multiple cameras may observe the same target at different times, cross-camera observation relationships can be established using target identifiers and timestamps. A reference frame is selected as the global reference baseline, and its corresponding camera coordinate system is defined as the reference frame camera coordinate system (usually the camera coordinate system of the first / initial frame target scene image is chosen). Starting from the reference frame, align the camera poses of all frames other than the reference frame according to the relative pose relationship graph, and uniformly transform them to the reference frame camera coordinate system. This yields the initial camera poses of all cameras in their respective reference frame coordinate systems, serving as the initial values for subsequent optimization.
[0075] After obtaining the initial camera poses for each camera, for each target, observation data of the target scene in multiple frames from multiple cameras is collected. Based on the camera poses of each frame of the target scene image and the 3D coordinates of the detected target vertices, the initial target pose in the reference frame camera coordinate system is calculated. The initial target pose represents the position and orientation of the target relative to the reference frame camera coordinate system, serving as the initial value for subsequent optimization.
[0076] To further improve accuracy, the initial camera poses of all cameras and the initial target poses of all targets are used as optimization variables, and the overall reprojection error is minimized to construct a pose optimization problem. Specifically, for each observed target vertex, its theoretical 3D coordinates in the target's local coordinate system are transformed to the reference frame camera coordinate system using the initial target pose, and then projected onto the corresponding target scene image plane using the initial camera pose. The deviation between this deviation and the actual detected 2D pixel coordinates of the target vertex is calculated; this is the reprojection error. The sum of the squares of all such deviations constitutes the objective function, which is the overall reprojection error. Minimizing the objective function is the reprojection error constraint.
[0077] The pose optimization problem is a BA (Bundle Adjustment) optimization problem. After constructing the pose optimization problem, nonlinear optimization algorithms such as the LM (Levenberg-Marquardt) algorithm can be used to iteratively optimize the reprojection error of each target vertex until the changes in the optimization variables tend to stabilize and the pose optimization problem converges, thereby calculating a more accurate camera pose and target pose in the reference frame camera coordinate system.
[0078] It is worth mentioning that introducing reprojection error constraints into the pose optimization problem can ensure that the optimized camera pose and target pose can most consistently interpret the observation data in all target scene images.
[0079] Based on the above, we can finally obtain high-precision camera motion trajectories (camera pose sequences) and target poses (sets) in the reference frame camera coordinate system.
[0080] Based on the above embodiments, the process of obtaining the target point cloud cluster of the target scene point cloud collected by each lidar in step S130 will be described in detail below.
[0081] The process of acquiring target point cloud clusters from the target scene point cloud collected by each lidar includes: converting the target scene point cloud collected by each lidar into a two-dimensional depth map; performing a breadth-first search on the two-dimensional depth map based on its depth continuity to obtain point cloud segmentation results, wherein the point cloud segmentation results include multiple point cloud segmentation clusters; determining the center point and distribution covariance value of each point cloud segmentation cluster, and performing principal component analysis on the distribution covariance value to obtain the eigenvalues corresponding to each point cloud segmentation cluster; and selecting target point cloud clusters with target shapes from the point cloud segmentation results based on the eigenvalues corresponding to each point cloud segmentation cluster.
[0082] It's easy to understand that for each point cloud of the target scene acquired by the LiDAR, it first needs to be converted into a two-dimensional depth map in time frames. Specifically, the horizontal and vertical angles of each point cloud of the target scene are mapped onto a two-dimensional table to form a two-dimensional depth map. Each pixel value in the two-dimensional depth map represents the distance value returned by the LiDAR in that direction, i.e., the depth of the point.
[0083] After obtaining the 2D depth map, point cloud segmentation is performed based on the depth continuity between pixels in the depth map. A breadth-first search algorithm can be used to traverse the entire 2D depth map, searching for regions with small depth variations between adjacent pixels as potential point cloud segmentation clusters. Specifically, starting with any unvisited pixel, its depth difference with its neighbors is checked. If the depth difference is less than a preset threshold, the two pixels are considered to belong to the same object, and the search continues until no new neighbors satisfying the condition are found. This search generates multiple independent point cloud segmentation clusters, each representing a possible physical object or a part thereof. The preset threshold can be set according to actual needs and is not specifically limited here.
[0084] For each segmented point cloud cluster, calculate its corresponding center point location and distribution covariance. The center point location is the geometric center of the point cloud cluster, which can be obtained by averaging the coordinates of all points within the cluster. The distribution covariance reflects the distribution of points within the cluster. When calculating the distribution covariance, the center point location is removed; the mean is first calculated for each dimension, and then the cross product of the deviations of each point relative to the mean is calculated and averaged. Here, "dimension" refers to the length, width, and height dimensions of the object represented by the point cloud cluster.
[0085] Principal Component Analysis (PCA) is performed on the calculated distribution covariance values to obtain the eigenvalues corresponding to each point cloud segmentation cluster. The magnitude of the eigenvalues characterizes the dispersion of data along the corresponding direction; larger eigenvalues indicate higher variability of data along that direction. Here, "direction" refers to the length, width, and height directions of the object represented by the point cloud segmentation cluster.
[0086] Finally, based on the feature values corresponding to each point cloud segmentation cluster, the point cloud segmentation cluster corresponding to the actual target is identified. Specifically, the target has a specific geometric shape (e.g., a rectangle), and its corresponding point cloud segmentation cluster will exhibit some unique characteristics. For example, if the target is a flat surface, the feature values of its corresponding point cloud segmentation cluster will show a larger feature value in one direction and a relatively smaller feature value in the other two orthogonal directions. Based on this pattern, point cloud clusters whose feature values match the expected characteristics can be selected from all point cloud segmentation clusters, thus obtaining the target point cloud cluster where the actual target is located.
[0087] Based on the above embodiments, the process of calculating the initial calibration parameters of multiple cameras and multiple lidars in step S130 will be described in detail below.
[0088] Based on the camera poses and target poses of multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to multiple lidars, the initial calibration parameters of multiple cameras and multiple lidars are calculated, including: determining the main camera based on the number of targets observed by each of the multiple cameras, the main camera being one of the multiple cameras, and the coordinate system of the main camera being the main camera coordinate system; using the degree of overlap between the target poses observed by non-main cameras and the target poses in the main camera coordinate system as the target cost function to optimize the camera poses of non-main cameras, thereby obtaining the initial calibration parameters of non-main cameras relative to the main camera; where non-main cameras refer to the multiple cameras other than the main camera.
[0089] Based on the initial calibration parameters of the non-main camera relative to the main camera, the three-dimensional coordinates of all target vertices in the main camera coordinate system are calculated to obtain the target vertex point cloud in the main camera coordinate system. The center point of the target point cloud cluster observed by each lidar is determined, and the matching degree between it and the center point of the corresponding target vertex point cloud in the main camera coordinate system is used as the target cost function for optimization to obtain the initial calibration parameters of multiple lidars relative to the main camera.
[0090] The straightforward approach is to calculate the total number of valid targets observed by each camera in its target scene image based on the target identification results. The camera with the highest total number of valid targets is then selected as the reference camera, also known as the main camera. The coordinate system of the main camera is denoted as the main camera coordinate system, and the other cameras are denoted as non-main cameras.
[0091] It should be noted that the camera with the largest number of effective targets observed was selected as the main camera because the main camera has the richest space observation information and can provide a more stable and reliable geometric reference.
[0092] For the same target (determined based on the target identifier), the target pose in the reference frame camera coordinate system observed by the non-main camera is transformed to the main camera coordinate system, and spatially aligned with the already determined target pose in the main camera coordinate system (i.e., the target pose observed by the main camera). Specifically, an ICP optimization problem is constructed, using the degree of spatial overlap between the target center or target vertex between the two as the objective cost function. By minimizing the objective cost function, the pose transformation relationship between the non-main camera and the main camera is optimized, and the initial calibration parameters of the non-main camera relative to the main camera are obtained.
[0093] After obtaining the initial calibration parameters of all non-main cameras relative to the main camera, the three-dimensional coordinates of the target vertices observed by each non-main camera are further transformed to the main camera coordinate system based on these initial calibration parameters, resulting in a target vertex point cloud (set) in the main camera coordinate system. The target vertex point cloud describes the geometric distribution of all targets in the unified reference coordinate system.
[0094] For each lidar, calculate the center point (the centroid of all points in the target point cloud cluster) of each target point cloud cluster observed by it, and use it as the position representative of the corresponding target in the lidar coordinate system.
[0095] For the same target, the center point of the target point cloud cluster observed by each lidar is matched and aligned with the center point of the target vertex point cloud in the main camera coordinate system. Specifically, another optimization problem is constructed, with the spatial distance matching error between the center point of the target point cloud cluster observed by the lidar and the center point of the corresponding target vertex point cloud in the main camera coordinate system as the objective cost function (i.e., point cloud matching error). By minimizing the point cloud matching error, the pose transformation parameters of the lidar relative to the main camera are adjusted to make the two sets of center points as aligned as possible in space, and finally the initial calibration parameters of each lidar relative to the main camera are obtained. The alignment object here can also be the degree of matching between the center point of the target point cloud cluster observed by the lidar and the corresponding target pose in the main camera coordinate system.
[0096] Based on the above, the initial calibration parameters of all non-main cameras and all LiDARs relative to the main camera can be obtained, which provides high-quality initial values for subsequent joint optimization.
[0097] During joint optimization, the correspondence between all target scene images, all target detection results, and all target scene point clouds is first established based on the target identifier (ID) and timestamp. Then, the correspondence between target point cloud clusters observed by multiple lidars, target detection results observed by multiple cameras, and target identifiers is established at the same time.
[0098] Subsequently, a joint optimization problem is constructed, which is a BA optimization problem. The initial calibration parameters of each camera and LiDAR relative to the main camera, all target poses, and the main camera's camera pose are used as optimization variables. The reprojection error of the target vertices and the point cloud matching error are minimized through the LM optimization algorithm until the degree of change of the optimization variables is small. The joint optimization problem converges, and the final calibration parameters are obtained, namely the target calibration parameters of multiple cameras and multiple LiDARs.
[0099] The multi-camera and multi-LiDAR joint calibration method provided in this embodiment has the following advantages: 1) The calibration parameters of multiple cameras and multiple LiDARs are more accurate; 2) The pose of each target in the main camera coordinate system is more reliable; 3) Drift and cumulative errors in SFM are eliminated, resulting in a smoother and more consistent camera motion trajectory (camera pose); 4) The reprojection of camera image pixels and the geometry of LiDAR point cloud are adjusted collaboratively under the same framework, realizing global optimization under the dual constraints of "vision + LiDAR".
[0100] Corresponding to the joint calibration method for multi-camera and multi-lidar described in the above embodiments, the present invention also proposes a joint calibration device for multi-camera and multi-lidar.
[0101] Specifically, Figure 2 A schematic diagram of the structure of the multi-camera and multi-lidar joint calibration device provided in an embodiment of the present invention is shown.
[0102] like Figure 2As shown, the device includes: a target scene image and point cloud acquisition module 210, used to acquire target scene images collected by multiple cameras and target scene point clouds collected by multiple lidars; a camera pose and target pose generation module 220, used to acquire target detection results of the target scene images acquired by each camera, and generate camera pose and target pose in the reference frame camera coordinate system based on the target detection results; a target point cloud cluster acquisition module 230, used to acquire target point cloud clusters of the target scene point clouds collected by each lidar; and an initial calibration parameter calculation module 240, used to calculate the initial calibration parameters based on the multiple cameras in the reference frame camera coordinate system. The system calculates the initial calibration parameters of the multiple cameras and multiple lidars based on the camera pose and target pose in the coordinate system, as well as the target point cloud clusters corresponding to the multiple lidars. The multi-camera and lidar joint calibration module 250 is used to construct and solve a joint optimization problem based on the initial calibration parameters of the multiple cameras and multiple lidars to obtain the target calibration parameters of the multiple cameras and multiple lidars. The joint optimization problem uses the camera pose, target pose, and the initial calibration parameters of each camera and lidar as optimization variables, and adds reprojection error constraints and point cloud matching error constraints.
[0103] In this embodiment, the target scene image and point cloud acquisition module 210 acquires target scene images collected by multiple cameras and target scene point clouds collected by multiple lidars; the camera pose and target pose generation module 220 acquires the target detection results of the target scene images acquired by each camera, and generates the camera pose and target pose in the reference frame camera coordinate system based on the target detection results; the target point cloud cluster acquisition module 230 acquires the target point cloud clusters of the target scene point clouds collected by each lidar; and the initial calibration parameter calculation module 240 calculates the target point cloud clusters of the target scene point clouds acquired by the multiple cameras in the reference frame. The system calculates initial calibration parameters for multiple cameras and lidars based on the camera pose and target pose in the camera coordinate system, as well as the target point cloud clusters corresponding to multiple lidars. A multi-camera and lidar joint calibration module 250 constructs and solves a joint optimization problem based on these initial calibration parameters to obtain the target calibration parameters for the multiple cameras and lidars. The joint optimization problem uses camera pose, target pose, and the initial calibration parameters of each camera and lidar as optimization variables, and incorporates reprojection error constraints and point cloud matching error constraints. This device eliminates the need for pre-constructing a target map, making it suitable not only for sensor deployment schemes with no or small common field of view, but also for scenarios with frequent target changes and altered calibration sites, significantly improving the accuracy of calibration parameters among multiple cameras and lidars.
[0104] It should be noted that the multi-camera and multi-LiDAR joint calibration device provided in the embodiments of the present invention can be referred to in correspondence with the multi-camera and multi-LiDAR joint calibration methods described in the above embodiments, and will not be repeated here.
[0105] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communications bus 340. The processor 310 can call logic instructions in the memory 330 to execute a joint calibration method for multiple cameras and multiple LiDARs. This method includes: acquiring target scene images collected by multiple cameras and target scene point clouds collected by multiple LiDARs; acquiring target detection results for the target scene images collected by each camera; generating camera pose and target pose in a reference frame camera coordinate system based on the target detection results; acquiring target point cloud clusters in the target scene point clouds collected by each LiDAR; calculating initial calibration parameters for the multiple cameras and the multiple LiDARs based on the camera poses and target poses in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple LiDARs; constructing and solving a joint optimization problem based on the initial calibration parameters of the multiple cameras and the multiple LiDARs to obtain target calibration parameters for the multiple cameras and the multiple LiDARs; wherein the joint optimization problem uses camera pose, target pose, and the initial calibration parameters of each camera and LiDAR as optimization variables, and adds reprojection error constraints and point cloud matching error constraints.
[0106] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0107] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the joint calibration method for multiple cameras and multiple lidars provided by the above methods. The method includes: acquiring target scene images collected by multiple cameras and target scene point clouds collected by multiple lidars; acquiring target detection results of the target scene images collected by each camera; generating camera pose and target pose in a reference frame camera coordinate system based on the target detection results; and acquiring the target pose of the target scene images collected by each lidar. A target point cloud cluster is formed by collecting the target scene point cloud. Based on the camera poses and target poses of the multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple lidars, initial calibration parameters for the multiple cameras and lidars are calculated. Based on the initial calibration parameters of the multiple cameras and lidars, a joint optimization problem is constructed and solved to obtain the target calibration parameters for the multiple cameras and lidars. The joint optimization problem uses camera poses, target poses, and the initial calibration parameters of each camera and lidar as optimization variables, and includes reprojection error constraints and point cloud matching error constraints.
[0108] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a joint calibration method for multiple cameras and multiple lidars provided by the above methods. The method includes: acquiring target scene images collected by multiple cameras and target scene point clouds collected by multiple lidars; acquiring target detection results of the target scene images collected by each camera; generating camera pose and target pose in a reference frame camera coordinate system based on the target detection results; and acquiring target point cloud clusters of the target scene point clouds collected by each lidar. Based on the camera poses and target poses of the multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple lidars, the initial calibration parameters of the multiple cameras and the multiple lidars are calculated; based on the initial calibration parameters of the multiple cameras and the multiple lidars, a joint optimization problem is constructed and solved to obtain the target calibration parameters of the multiple cameras and the multiple lidars; wherein, the joint optimization problem uses the camera pose, target pose, and the initial calibration parameters of each camera and lidar as optimization variables, and adds reprojection error constraints and point cloud matching error constraints.
[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0110] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A joint calibration method for multiple cameras and multiple lidar, characterized in that, include: Acquire target scene images from multiple cameras, as well as target scene point clouds from multiple lidar sensors; Obtain the target detection results of the target scene images captured by each camera, and generate the camera pose and target pose in the reference frame camera coordinate system based on the target detection results; Obtain the target point cloud clusters from the target scene point clouds collected by each lidar. Based on the camera pose and target pose of the multiple cameras in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple lidars, calculate the initial calibration parameters of the multiple cameras and the multiple lidars. Based on the initial calibration parameters of the multiple cameras and multiple lidars, a joint optimization problem is constructed and solved to obtain the target calibration parameters of the multiple cameras and multiple lidars; The joint optimization problem uses camera pose, target pose, and initial calibration parameters of each camera and lidar as optimization variables, and adds reprojection error constraints and point cloud matching error constraints.
2. The joint calibration method for multiple cameras and multiple lidar as described in claim 1, characterized in that, The acquisition of target detection results from the target scene images captured by each camera includes: Target detection is performed on the target scene images captured by each camera to obtain the target detection results; The target detection results include the pose of the target center and the three-dimensional coordinates of the target vertex in the camera coordinate system where the camera is located, as well as the target identifier.
3. The joint calibration method for multiple cameras and multiple lidar as described in claim 2, characterized in that, The step of generating the camera pose and target pose in the reference frame camera coordinate system based on the target detection results includes: The pixel projection error of the same target vertex in different frame target scene images is optimized by using perspective n-point solution, so as to obtain the camera pose of the multiple cameras in different frame target scene images; Based on the relative pose changes of each camera in different frame target scene images, a camera pose relationship diagram is constructed between different frame target scene images. Based on the camera pose relationship diagram corresponding to the multiple cameras, the initial camera pose of the multiple cameras in the reference frame camera coordinate system is determined; Based on the target detection results and the initial camera poses of the multiple cameras in the reference frame camera coordinate system, calculate the initial target pose of each target in the reference frame camera coordinate system. Based on the initial camera pose, the initial target pose, and the three-dimensional coordinates of the target vertex, a pose optimization problem is constructed and solved to obtain the camera pose and target pose in the reference frame camera coordinate system. The pose optimization problem uses the initial camera pose and the initial target pose as optimization variables, and has reprojection error constraints.
4. The joint calibration method for multiple cameras and multiple lidar as described in claim 1, characterized in that, The acquisition of the target point cloud clusters from the target scene point clouds collected by each lidar includes: Convert the point cloud of the target scene collected by each lidar into a two-dimensional depth map; Based on the depth continuity of the two-dimensional depth map, a breadth-first search is performed on the two-dimensional depth map to obtain point cloud segmentation results, which include multiple point cloud segmentation clusters. The center point and distribution covariance value of each point cloud segmentation cluster are determined, and principal component analysis is performed on the distribution covariance value to obtain the eigenvalues corresponding to each point cloud segmentation cluster. Based on the feature value corresponding to each point cloud segmentation cluster, target point cloud clusters with target shapes are selected from the point cloud segmentation results.
5. The joint calibration method for multiple cameras and multiple lidar as described in claim 1, characterized in that, The step of calculating the initial calibration parameters of the multiple cameras and the multiple lidars based on the camera poses and target poses in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple lidars, includes: The main camera is determined based on the number of targets observed by each of the multiple cameras. The main camera is one of the multiple cameras, and the coordinate system of the main camera is the main camera coordinate system. The degree of overlap between the target pose observed by the non-main camera and the target pose in the main camera coordinate system is used as the target cost function to optimize the camera pose of the non-main camera, thereby obtaining the initial calibration parameters of the non-main camera relative to the main camera. The non-main camera refers to multiple cameras other than the main camera.
6. The joint calibration method for multiple cameras and multiple lidar as described in claim 5, characterized in that, The step of calculating the initial calibration parameters of the multiple cameras and the multiple lidars based on the camera poses and target poses in the reference frame camera coordinate system, and the target point cloud clusters corresponding to the multiple lidars, includes: Based on the initial calibration parameters of the non-main camera relative to the main camera, the three-dimensional coordinates of all target vertices in the main camera coordinate system are calculated to obtain the target vertex point cloud in the main camera coordinate system. The center point of the target point cloud cluster observed by each lidar is determined, and the matching degree between the lidar and the center point of the corresponding target vertex point cloud in the main camera coordinate system is used as the target cost function for optimization, thereby obtaining the initial calibration parameters of multiple lidars relative to the main camera.
7. A joint calibration device for multiple cameras and multiple lidar, characterized in that, include: The target scene image and point cloud acquisition module is used to acquire target scene images collected by multiple cameras and target scene point clouds collected by multiple lidar sensors. The camera pose and target pose generation module is used to obtain the target detection results of the target scene images captured by each camera, and generate the camera pose and target pose in the reference frame camera coordinate system based on the target detection results. The target point cloud cluster acquisition module is used to acquire the target point cloud cluster of the target scene point cloud collected by each lidar. The initial calibration parameter calculation module is used to calculate the initial calibration parameters of the multiple cameras and the multiple lidars based on the camera pose and target pose of the multiple cameras in the reference frame camera coordinate system, as well as the target point cloud clusters corresponding to the multiple lidars. A multi-camera and lidar joint calibration module is used to construct and solve a joint optimization problem based on the initial calibration parameters of the multiple cameras and the multiple lidars to obtain the target calibration parameters of the multiple cameras and the multiple lidars. The joint optimization problem uses camera pose, target pose, and initial calibration parameters of each camera and lidar as optimization variables, and adds reprojection error constraints and point cloud matching error constraints.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the joint calibration method for multiple cameras and multiple lidar as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the joint calibration method for multiple cameras and multiple lidar as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the joint calibration method for multiple cameras and multiple lidar as described in any one of claims 1 to 6.