Lidar and camera joint calibration method, electronic device and storage medium
By using a joint calibration method for lidar and camera, motion distortion compensation and conversion of point cloud data into three-dimensional Gaussian voxels are performed. Combined with color consistency constraints and rendering error terms, the accuracy and robustness of extrinsic parameter calibration are improved, solving the calibration problem under dynamically acquired data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 江淮前沿技术协同创新中心
- Filing Date
- 2025-10-14
- Publication Date
- 2026-08-04
AI Technical Summary
When the structural features of the scene are not obvious, the calibration accuracy of the external parameters of the lidar and camera decreases, and it is difficult to perform high-precision calibration based on dynamically acquired data.
By acquiring point cloud data from the lidar and image data from the camera, motion distortion compensation is performed on the point cloud data to generate a global point cloud map. This map is then converted into three-dimensional Gaussian voxels with color and covariance attributes. By jointly optimizing the attributes of the three-dimensional Gaussian voxels and the camera extrinsic parameters, the first loss function is minimized to complete the extrinsic parameter calibration.
It improves the robustness and accuracy of extrinsic parameter calibration in complex dynamic scenarios, solves the problem of decreased calibration accuracy in the absence of obvious structural features in the absence of calibration objects, and realizes high-precision extrinsic parameter calibration for dynamically acquired data.
Smart Images

Figure CN121304805B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of computer vision, and particularly to a method for joint calibration of lidar and camera, electronic devices and storage media. Background Technology
[0002] In fields such as computer vision, multimodal fusion of LiDAR and cameras is fundamental to environmental perception, while accurate extrinsic parameter calibration is essential for effective data fusion. Existing technologies primarily rely on calibration objects for extrinsic parameter calibration. For example, using artificial targets like checkerboard patterns, extrinsic parameters are calculated by detecting common features of these targets in data from different sensors. These methods are cumbersome, have low automation, and require manual intervention. The second type is calibration methods without targets, which directly match common natural structural features in scenes collected by different sensors.
[0003] However, the aforementioned calibration methods without calibration objects are highly dependent on scene conditions; when structural features are not obvious in the scene, the accuracy drops significantly. Furthermore, most existing calibration methods require the sensor system to remain stationary during data acquisition, making it difficult to perform high-precision extrinsic parameter calibration based on dynamically acquired data. Summary of the Invention
[0004] The purpose of this invention is to provide a joint calibration method for lidar and camera, an electronic device, and a storage medium, which can solve the problem that the accuracy of related extrinsic parameter calibration methods decreases significantly when the structural features of the scene are not obvious, and it is difficult to perform high-precision extrinsic parameter calibration based on dynamically acquired data.
[0005] To address the aforementioned technical problems, embodiments of the present invention provide a joint calibration method for lidar and camera, comprising: acquiring point cloud data from multiple lidars and image data from a camera; performing motion distortion compensation on the point cloud data to generate a global point cloud map; converting the global point cloud map into multiple three-dimensional Gaussian voxels with color attributes and covariance attributes; and minimizing a first loss function by jointly optimizing the attributes of the three-dimensional Gaussian voxels and the extrinsic parameters of the camera to complete the extrinsic parameter calibration of the camera; wherein the first loss function includes a color consistency constraint term and a rendering error term, the color consistency constraint term being used to constrain the color attributes of a single three-dimensional Gaussian voxel to remain consistent across multiple time points and multiple viewpoints, and the rendering error term being the pixel color error between the projected image generated by the three-dimensional Gaussian voxel and the image data from the camera.
[0006] Embodiments of the present invention also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the lidar and camera joint calibration method as described above.
[0007] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the lidar and camera joint calibration method as described above.
[0008] In this embodiment of the invention, firstly, the scheme compensates for motion distortion in the point cloud data collected during the dynamic process, transforming the deformed raw data into an accurate and distortion-free global point cloud map. This lays the foundation for extrinsic parameter calibration under dynamically acquired data. Then, the global point cloud map is converted into three-dimensional Gaussian voxels with color and covariance attributes, thereby constructing a continuous scene model containing structural features of three-dimensional Gaussian voxels. This makes the structural features upon which the extrinsic parameter calibration between the LiDAR and the camera depends more apparent, thus improving the extrinsic parameter calibration effect. Furthermore, a color consistency constraint term is designed into the first loss function. By forcing the same three-dimensional Gaussian voxel to maintain a stable color when observed at different times and angles, this provides a strong global constraint for the optimization process. Combined with the rendering error term, this significantly improves the robustness and accuracy of extrinsic parameter calibration in complex dynamic scenes.
[0009] Furthermore, the motion distortion compensation for the point cloud data includes: firstly, obtaining pose data through point cloud registration; then, performing distortion processing on the point cloud data based on the pose data; iteratively performing point cloud registration and distortion processing until preset constraints are met, thus completing the motion distortion compensation; wherein, the preset constraints are that the pose change, objective function convergence, and global pose consistency simultaneously satisfy the target constraints. By iteratively performing point cloud registration and distortion processing until the preset constraints are met, it is possible to compensate for point cloud distortion caused by line-by-line scanning of LiDAR, achieving distortion-free point clouds under a unified time reference, thereby improving the map accuracy constructed from dynamically acquired point cloud data.
[0010] Furthermore, the method includes: when the number of lidars is greater than or equal to two, sequentially performing motion distortion compensation on the point cloud data of multiple lidars to obtain point cloud maps of multiple lidars respectively; selecting one of the point cloud maps of multiple lidars as a reference point cloud map, and sequentially matching the remaining point cloud maps to the reference point cloud map to complete the extrinsic parameter calibration among multiple lidars, thereby generating a global point cloud map. Through the above method, motion distortion compensation can be performed on the point cloud data collected by multiple lidars, and extrinsic parameter calibration among multiple lidars can be further achieved. Attached Figure Description
[0011] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0012] Figure 1 This is a flowchart of a joint calibration method for lidar and camera provided according to an embodiment of the present invention; Figure 2 This is a flowchart of the external parameter calibration between lidars in a lidar and camera joint calibration method provided by an embodiment of the present invention; Figure 3 This is a structural diagram of an electronic device used in a laser radar and camera joint calibration method provided by an embodiment of the present invention. Detailed Implementation
[0013] In fields such as computer vision, multimodal fusion of LiDAR and cameras is a key technology for achieving robust environmental perception, while accurate extrinsic parameter calibration is the foundation for effective data fusion. To achieve extrinsic parameter calibration, existing technologies mainly face the following challenges: The first category is traditional calibration methods based on specific calibration objects. For example, this method typically uses targets such as checkerboard patterns that require precise manual placement, and calculates extrinsic parameters by detecting common calibration object features in data from different sensors. Although this type of method can guarantee a certain level of accuracy in a controlled environment, its process is cumbersome, has a low degree of automation, requires a lot of manual intervention, is not suitable for large-scale deployment, and is difficult to quickly recalibrate in complex field environments.
[0014] The second category is calibration methods without calibration objects, designed to improve automation. These methods typically rely on direct matching based on common natural structural features (such as edges, planes, or corners) in scenes captured by different sensors. However, these methods are highly dependent on scene conditions; when obvious common structural features are lacking (e.g., in the field or in scenes with repetitive features), their calibration accuracy drops significantly. Furthermore, when the overlap area of the fields of view of multiple sensors is small, the principle of direct feature matching makes it difficult to find sufficient constraints, thus failing to obtain high-precision calibration results.
[0015] Furthermore, all the aforementioned existing methods share a common fundamental limitation: their design is based on the premise that the sensor system must remain static during data acquisition. However, in real-world applications such as autonomous driving, sensors are precisely mounted on moving platforms. When the LiDAR scans in motion, the sequential acquisition of scan points results in severe motion distortion (e.g., "ghosting" or "distortion") in the generated single-frame point cloud, making it impossible for the point cloud to accurately reflect the real-world scene geometry. Consequently, dynamic extrinsic parameter calibration between the LiDAR and the camera is also difficult.
[0016] Therefore, there is an urgent need for a joint calibration method for lidar and cameras that can solve the problem that the accuracy of relevant extrinsic parameter calibration methods decreases significantly when the structural features of the scene are not obvious, and that it is difficult to perform high-precision extrinsic parameter calibration based on dynamically acquired data.
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details are presented in the various embodiments of the present invention to facilitate a better understanding of this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the various embodiments below is for ease of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined with and referenced by each other without contradiction.
[0018] One embodiment of the present invention relates to a joint calibration method for lidar and camera, which can be applied to workstations or vehicle-mounted computing units. For example, the executing entity can be an embedded computing unit integrated into an autonomous vehicle, intelligent engineering machinery, or mobile robot, including an on-board computing platform, an autonomous driving domain controller, or a robot controller, as well as a host computer with data processing capabilities. The embodiment of the present invention includes: acquiring point cloud data from multiple lidars and image data from a camera; performing motion distortion compensation on the point cloud data to generate a global point cloud map; converting the global point cloud map into multiple three-dimensional Gaussian voxels with color attributes and covariance attributes; and minimizing a first loss function by jointly optimizing the attributes of the three-dimensional Gaussian voxels and the extrinsic parameters of the camera to complete the extrinsic parameter calibration of the camera; wherein the first loss function includes a color consistency constraint term and a rendering error term, the color consistency constraint term constraining the color attributes of a single three-dimensional Gaussian voxel observed at multiple times and from multiple viewpoints to remain consistent, and the rendering error term being the pixel color error between the projected image generated by the three-dimensional Gaussian voxel and the image data from the camera. In this embodiment of the invention, firstly, the scheme compensates for motion distortion in the point cloud data collected during the dynamic process, transforming the deformed raw data into an accurate and distortion-free global point cloud map. This lays the foundation for extrinsic parameter calibration under dynamically acquired data. Then, the global point cloud map is converted into three-dimensional Gaussian voxels with color and covariance attributes, thereby constructing a continuous scene model containing structural features of three-dimensional Gaussian voxels. This makes the structural features upon which the extrinsic parameter calibration between the LiDAR and the camera depends more apparent, thus improving the extrinsic parameter calibration effect. Furthermore, a color consistency constraint term is designed into the first loss function. By forcing the same three-dimensional Gaussian voxel to maintain a stable color when observed at different times and angles, this provides a strong global constraint for the optimization process. Combined with the rendering error term, this significantly improves the robustness and accuracy of extrinsic parameter calibration in complex dynamic scenes.
[0019] The following is a detailed description of the implementation details of the lidar and camera joint calibration method according to an embodiment of the present invention. The following content is only for the convenience of understanding and is not necessary for implementing this solution.
[0020] like Figure 1 As shown, the method of this embodiment of the invention includes steps 110 to 140.
[0021] In step 110, point cloud data from the LiDAR and image data from the camera are acquired. The number of LiDAR and camera must be at least one.
[0022] In step 120, motion distortion compensation is performed on the point cloud data to generate a global point cloud map.
[0023] In an optional embodiment, the number of LiDARs is greater than or equal to two. In this case, the steps for generating a global point cloud map are as follows: Figure 2 As shown, the method includes steps 121 and 122: Step 121, motion distortion compensation is performed sequentially on the point cloud data of multiple lidars to obtain point cloud maps for each lidar; Step 122, one of the point cloud maps from the multiple lidars is selected as the reference point cloud map, and the remaining point cloud maps are sequentially matched to the reference point cloud map to complete the extrinsic parameter calibration between the multiple lidars, thereby generating a global point cloud map. Through the above method, motion distortion compensation can be performed on the point cloud data collected by multiple lidars, and extrinsic parameter calibration between multiple lidars can be further achieved.
[0024] In a specific example, the motion distortion compensation in step 121 above is achieved through the following steps: For point cloud data, pose data is first obtained through point cloud registration. Then, based on the pose data, distortion processing is performed on the point cloud data. This process of point cloud registration and distortion processing is iteratively repeated until preset constraints are met, thus completing the motion distortion compensation. The preset constraints are that the pose change, objective function convergence, and global pose consistency simultaneously satisfy the target constraints. By iteratively performing point cloud registration and distortion processing until the preset constraints are met, it is possible to compensate for point cloud distortion caused by line-by-line scanning by LiDAR, achieving distortion-free point clouds under a unified time reference, thereby improving the accuracy of maps constructed from dynamically acquired point cloud data.
[0025] It should be noted that the aforementioned objective constraints include: the rotational change and translational update of the pose before and after point cloud registration are less than a preset pose change threshold; the change of the objective function value before and after point cloud registration is less than a preset objective function value change threshold; and the deviation of the pose after point cloud registration in the global range is less than a preset global deviation threshold.
[0026] Optionally, the point cloud registration method mentioned in step 121 above is a point cloud registration method based on probability distribution modeling. The specific details of the point cloud registration method based on probability distribution modeling are as follows: First, one point cloud is selected as the reference point cloud, and another point cloud is used for registration onto the reference point cloud. To transform the discrete reference point cloud into a continuous mathematical representation that is more conducive to matching, the reference point cloud is first voxelized. The reference point cloud is divided into a spatial voxel grid, and the point set within each voxel cell approximates a Gaussian distribution. : , For any voxel cell, calculate all 3D points within that cell. The statistical characteristics. Here, the subscript k represents the Kth point in a voxel cell of the reference point cloud. This refers to the k-th point within one of the cubes. Therefore, k ranges from 1 to the total number of points N within this small cube. Furthermore, the mean vector of the point set is calculated using the following formula. Covariance Matrix : , Among them, the mean vector Physically, it represents the geometric center of the local point cloud, while the covariance matrix... This characterizes the distribution pattern and orientation of the local point cloud in three-dimensional space, effectively describing the geometric structure of the local region, such as whether the region is approximately a plane, a straight line, or a disordered cluster of points. Then, during the registration process, for any point in the point cloud to be registered... It is necessary to evaluate its behavior after undergoing a rigid body transformation. (This transformation is composed of rotation and translation parameters) (Definition) After the action, the degree of matching with the reference point cloud model. Therefore, for the points in the point cloud to be registered... After rigid body transformation Then, its probability density representing the degree of matching is: , Where the subscript j represents the j-th point in the entire point cloud to be registered. This refers to the j-th point in the new frame of scan data. The range of j is from 1 to the total number of points in the entire point cloud frame. The formula above calculates the transformed point. Falling into a local Gaussian distribution in the reference point cloud The possibility of [something]. It is worth noting that the exponent term in this formula is essentially the point of transformation. With the center of Gaussian distribution The square of the Mahalanobis distance between them is obtained by passing through the inverse of the covariance matrix. Distance weighting was applied. The advantage of this weighting method is that it adaptively adjusts the error penalty based on the local geometry of the reference point cloud. Specifically, it allows larger deviations in directions with gentle geometric changes (e.g., within a plane) and imposes larger penalties in directions with drastic geometric changes (e.g., perpendicular to the plane), thus achieving more robust and accurate alignment. To evaluate the transformation... A global log-likelihood function was constructed to assess the alignment effect of the entire point cloud to be registered. : , This function will select all points in the point cloud to be registered. The logarithmic probability densities are accumulated to form a probability density function relating to the transformation parameters. The global objective function. The higher the value, the more significant the current transformation. This aims to make the point cloud to be registered as closely as possible to the reference point cloud as a whole. Using logarithmic form is a common technique in optimization problems; it transforms the probability product into a logarithmic sum without changing the extreme points of the function, thus simplifying computation and improving numerical stability. Therefore, the point cloud registration problem is formalized as an optimization problem, namely, finding a set of optimal transformation parameters. This makes the overall log-likelihood function above... The maximum value is reached. The registration problem is represented as: , The optimal transformation parameters The corresponding rigid body transformation This refers to the optimal pose transformation that minimizes the alignment error between two point cloud frames. Finally, based on the point cloud registration method described above, numerical iterative optimization (such as Newton's method or quasi-Newton's method) yields the pose estimation of a single lidar sensor at a series of discrete time points, denoted as the pose sequence. : , in, The subscript 'i' represents the frame number, such as... It is the pose of the point cloud in the first frame. This is the pose of the second frame, and so on. Pose It contains the sensor's 3D rotation and 3D translation information at time i, used to describe the sensor's precise position and orientation in 3D space. Each pose... All of these are rigid body transformation matrices belonging to the SE(3) space. These poses roughly describe the position and attitude of the sensor at different times and are the basic data points for subsequent construction of continuous trajectories.
[0027] Optionally, the distortion processing method mentioned in step 121 above is a distortion processing method based on B-spline curves. The specific distortion processing method based on B-spline curves is as follows: First, since the scanning process of a lidar is continuous, discrete poses alone are insufficient to accurately correct distortions caused by sensor motion within a single frame scan. To model the continuous trajectory of the sensor during the scanning time, this embodiment of the invention introduces a B-spline curve B(t). This B-spline curve can mathematically and accurately describe a smooth and continuous motion trajectory: , in, For k-th order B-spline basis functions, These are the control points. Specifically, this B-spline curve... It consists of a set of control points and k-th order B-spline basis functions The control points are defined by a weighted sum. These are a series of parameters that determine the shape of the pose curve, while the basis functions... As a weight, its value at time t determines the degree of influence of each control point on the current pose. To determine the optimal control point... This embodiment uses an optimization process to find a set of control points that minimizes the fitting error: , Here, j represents the index of the discrete pose sequence, which loops from 1 to n, traversing all known discrete poses. It is the j-th discrete pose. It is the timestamp corresponding to the j-th discrete pose. It is the constructed continuous B-spline curve at time point The interpolation result on the B-spline curve is its "predicted pose" at that time. Therefore, the fitting error is defined as the result of the B-spline curve at discrete times... interpolated pose B Discrete pose obtained by registration with point cloud The sum of the differences between them. By minimizing this error, a set of optimal control points can be obtained. This allows us to determine a smooth, continuous pose curve that most accurately approximates the actual discrete pose observation points. .
[0028] Line-by-line scanning by lidar results in different points in the same frame of point cloud having different sampling timestamps. Let the points in the i-th frame of point cloud be... Its timestamp is The interpolation value of the pose curve at that moment is: , Because during the scanning process of a frame of data, each laser point within a lidar... They are all at different timestamps The data is collected, and therefore its coordinates are based on the sensor's own coordinate system at that moment. To eliminate this deformation introduced by motion, this embodiment first uses the timestamp of each point... In continuous pose curves Interpolation is performed to obtain the precise pose of the sensor when the point was acquired. Subsequently, this precise pose was used. Point Coordinate transformation can distort point clouds into... : .
[0029] in, The subscript 'i' also represents the i-th frame of the point cloud, and the subscript 'j' represents the j-th laser point within the point cloud of the i-th frame. By performing this operation on all points within a frame, all points can be uniformly transformed to a common reference coordinate system (e.g., the coordinate system at the start of the frame), thereby effectively eliminating motion distortion and obtaining a geometrically accurate, distortion-free point cloud with a unified time reference. .
[0030] It is important to note that the pose estimation obtained after a single point cloud registration and distortion processing in the above specific case is not accurate. To ensure the accuracy and robustness of the final output pose sequence and 3D map, this embodiment of the invention employs a joint convergence criterion. This criterion is not based on a single indicator, but rather comprehensively judges from three dimensions: local changes, global optimization progress, and consistency of the final result. The iterative process can only terminate when the constraints of all dimensions are simultaneously satisfied. The constraints are: , The iteration will only stop when the following four conditions are met simultaneously: (1) the local pose update is small enough; (2) the objective function no longer shows significant improvement; and (3) the global pose sequence is smooth and consistent. This multi-dimensional and rigorous joint judgment criterion can effectively avoid premature termination due to the satisfaction of a single criterion, ensure the full optimization of the iteration process, and ultimately obtain a high-quality solution that is locally stable and globally consistent.
[0031] The constraint (1) above is used to monitor whether the estimated sensor pose has stabilized between two adjacent iterations. Specifically, it measures the amount of pose update in the k-th iteration, including the rotation component. Translation section If both update amounts are small enough, less than the preset rotation update amount convergence threshold. Convergence threshold for translation update This indicates that, from a local perspective, the pose estimation has converged, and subsequent iterations are unlikely to produce significant optimization.
[0032] The above constraint (2) is used to determine whether the optimization process of point cloud registration has reached its limit. This is the objective function in the point cloud registration process during the k-th iteration, and its value represents the degree of matching between the two frames of point clouds at the current pose. This formula calculates the relative rate of change of the objective function value between two adjacent iterations. If this rate of change is less than a preset convergence threshold for the relative rate of change of the objective function... This means that even if the pose is still being fine-tuned, the improvement effect on the overall matching quality is negligible, indicating that the optimization has entered a flat region or reached near the optimal solution.
[0033] The above constraint (3) is used to evaluate the quality and consistency of the overall result of the current iteration output. It calculates the final discrete pose sequence. The smooth, continuous motion trajectory fitted from this sequence The overall deviation between them. If the calculated result is less than the preset convergence threshold of the overall deviation. This indicates that the discrete pose points are closely distributed on a smooth motion trajectory, the entire pose sequence is smooth and consistent globally, and the solution has high stability.
[0034] In a specific example, the extrinsic parameter calibration in step 122 above is achieved through the following steps: by iterating the extrinsic parameters among multiple lidars, the second loss function between the remaining point cloud map and the reference point cloud map is minimized, thus completing the extrinsic parameter calibration among multiple lidars. The second loss function is obtained by weighted calculation of the error vector and the covariance matrix. The error vector is calculated from the values of corresponding point clouds in the remaining and reference point cloud maps, and the covariance matrix is calculated from the local covariance of corresponding point clouds in the remaining and reference point cloud maps.
[0035] The extrinsic parameter calibration among multiple lidars in step 122 above is implemented as follows: Each lidar independently runs the motion distortion processing algorithm in step 121 above to generate a dense point cloud map. Then, the map from one of the lidars is selected as the reference, i.e., the target point cloud. And use the map of another lidar to be calibrated as the source point cloud. To achieve accurate registration using the geometric features of the point clouds, the system constructs a local covariance matrix for each point in both the source and target point clouds by analyzing the distribution of its neighborhood points. This matrix is denoted as follows: and The covariance matrix effectively describes the local surface geometry at the point. Next, the system employs an iterative matching method based on local covariance weighting, solving for the optimal rigid body transformation between the two lidars by minimizing a weighted squared error function. The optimal rotation matrix is obtained. Translation vector They are combined into a 4x4 homogeneous extrinsic parameter matrix. By repeating this calibration process for all non-reference lidars, the extrinsic parameter matrix of each lidar relative to the reference lidar can be obtained. Finally, using these extrinsic parameter matrices, all independent point cloud maps can be transformed to a unified global coordinate system based on the reference lidar, aligned, and fused into a more complete and denser global point cloud map.
[0036] Those skilled in the art will recognize that step 120 primarily achieves two objectives: first, motion distortion processing is applied to the dynamically acquired point cloud data from the LiDAR, enabling the construction of a point cloud map based on the dynamic point cloud data; second, extrinsic parameter calibration and global map construction are performed on the point cloud maps constructed from two or more LiDARs, enabling both extrinsic parameter construction and providing a more complete calibration base map for subsequent extrinsic parameter calibration between the LiDAR and the camera, thereby improving the accuracy of subsequent camera extrinsic parameter calibration.
[0037] In step 130, the global point cloud map is converted into a three-dimensional Gaussian voxel with color and covariance attributes.
[0038] In a specific example, to establish a differentiable mapping relationship between the LiDAR point cloud and the camera image, this embodiment of the invention converts the previously generated global point cloud map into a continuous scene representation composed of multiple three-dimensional Gaussian voxels with color and covariance attributes. Specifically, the global point cloud map is first regarded as a whole three-dimensional point set stitched together from point clouds at multiple time points. , where i is the i-th frame. Subsequently, these discrete points are spatially aggregated to generate M three-dimensional Gaussian spheres (or Gaussian voxels). Each Gaussian sphere k in three-dimensional space is defined by a set of properties, including: center position. Covariance matrix ,color and transparency Among them, the covariance matrix The size and orientation of the Gaussian sphere in three-dimensional space are controlled, enabling it to adaptively cover local regions of the point cloud, thus transforming the discrete point cloud into a continuous volumetric representation. In three-dimensional space, the contribution of the k-th Gaussian sphere to any point x can be expressed by the following formula: , Where x represents the coordinates of any point in three-dimensional space. Next, to associate this three-dimensional Gaussian voxel with the camera image, the three-dimensional Gaussian voxel needs to be projected onto the plane of the camera image. For each camera c in the system, its image data is... Where c is the camera index, t is the timestamp index, and (u,v) are the pixel coordinates. Its inherent camera intrinsic parameter matrix... for: , in, and These are the focal lengths of the camera along the x and y axes, respectively. and The coordinates are the principal point coordinates. The camera extrinsic parameters to be calibrated are determined by the rotation matrix. Translation vector The rendering process includes the following steps: 1. Coordinate System Transformation: First, transform each 3D Gaussian sphere k in the world coordinate system using the currently estimated camera extrinsic parameters ( The coordinates are transformed to the coordinate system of camera c. The transformed center position is... and the transformed covariance matrix Follow the formula below: , 2. Projection onto the pixel plane: Projection function via pinhole camera By projecting a three-dimensional Gaussian sphere in the camera coordinate system onto a two-dimensional pixel plane, a two-dimensional Gaussian distribution is obtained. This two-dimensional distribution is centered at the two-dimensional projection center. and projection covariance To describe.
[0039] 3. Calculating Rendering Weights and Pixel Values: For any pixel (u, v) in the image, its final rendered color is a weighted mixture of the colors projected onto that pixel region using a two-dimensional Gaussian distribution. The weight of a single Gaussian sphere to pixel (u, v) is:
[0040] The weight of the individual Gaussian sphere mentioned above is determined by its transparency. The final rendered color value of this pixel is determined by the probability density of its center point under its two-dimensional Gaussian distribution, and normalization is applied to all influencing Gaussian spheres. Here, j in the denominator is the summation index of all M Gaussian spheres. These are all the colors of the Gaussian sphere. Weighted sum:
[0041] Through the above steps, the system constructs a rendering process from 3D point cloud, Gaussian voxel attributes, and camera extrinsic parameters to 2D image color. This lays the foundation for subsequent joint optimization of 3D Gaussian voxels and camera extrinsic parameters using gradient descent.
[0042] As those skilled in the art will understand, step 130 lays the foundation for subsequent high-precision calibration by constructing a three-dimensional Gaussian voxel from the global point cloud map. Specifically, the above steps can construct a continuous scene model composed of multiple three-dimensional Gaussian voxels with richer structural features from the inherently discrete and sparse point cloud data. Each three-dimensional Gaussian voxel contains not only center position and color information, but also a covariance matrix. This covariance matrix can characterize the geometry and orientation of the local point cloud region covered by the voxel, thus forming a continuous volume representation. In this way, the structural features on which the extrinsic parameter calibration between the LiDAR and the camera depends become more obvious and robust. This can effectively solve a core problem in related uncalibrated extrinsic parameter calibration technology, namely, the problem that calibration accuracy will significantly decrease when there are no obvious common structural features such as edges, planes, or corners in the scene. Therefore, by constructing such a continuous scene representation containing rich geometric information, this invention significantly improves the adaptability and accuracy of the calibration method to scene conditions.
[0043] In step 140, the first loss function is minimized by jointly optimizing the properties of the three-dimensional Gaussian voxels and the camera's extrinsic parameters to complete the camera's extrinsic parameter calibration.
[0044] It should be noted that the first loss function mentioned above includes a color consistency constraint term and a rendering error term. The color consistency constraint term is used to ensure that the color attributes of a single 3D Gaussian voxel remain consistent across multiple time points and viewpoints. The rendering error term is the pixel color error between the projected image generated by the 3D Gaussian voxel and the image data from the camera.
[0045] In a specific embodiment, step 140 includes: projecting multiple three-dimensional Gaussian voxels onto the image plane of the camera based on initial extrinsic parameters to obtain a projected image; constructing a first loss function based on the projected image and the image data of the camera, and calculating the gradient; iteratively updating the attributes of the three-dimensional Gaussian voxels, the extrinsic parameters of the camera, and the color and covariance attributes of the three-dimensional Gaussian voxels according to the gradient, until the first loss function converges to the minimum value.
[0046] In step 140, the system performs joint calibration of the LiDAR and cameras, the core of which is the joint optimization of the properties of the 3D Gaussian voxels and the extrinsic parameters of each camera. Specifically, during the optimization process, the center position of the 3D Gaussian sphere is determined. Covariance ,color and transparency And external parameters of each camera (rotation) Peaceful relocation Both can be updated simultaneously using gradient descent. The size and orientation of the Gaussian sphere are determined by the covariance matrix. The decision is made to effectively cover local point cloud regions to ensure rendering continuity. Meanwhile, transparency and color are used to match the visual features of the real-world images captured by the camera. The optimization objective of this step is to minimize a first loss function, which aims to minimize the pixel color error between the rendered 3D Gaussian voxel image and the real camera image, and to ensure color consistency within the same spatial region under multi-viewpoint and multi-time conditions. This first loss function is defined as: , in, This represents the average color of the same Gaussian voxel k across multiple viewpoints and time points. Let represent the color observed by Gaussian voxel k at time t. The first loss function consists of the following two key parts: 1. Rendering error term: The first part of the formula This is the rendering error term. This term calculates the projected image generated from 3D Gaussian voxels. Compared with real image data captured by the camera The sum of squared color errors across all pixels (u,v), all cameras c, and all times t. Minimizing this error term drives the optimization process to adjust various properties of the camera extrinsic parameters and Gaussian voxels, making the 2D image rendered from the 3D point cloud as visually consistent as possible with the real photograph, thus achieving precise geometric and visual alignment.
[0047] 2. Color consistency constraint: Part 2 of the formula This is a color consistency constraint term used to ensure that the color attribute of a single 3D Gaussian voxel remains consistent across different times and viewpoints. It is based on the physical fact that the intrinsic color of the same point in space (represented by a Gaussian voxel k) is fixed. This term penalizes the color of a Gaussian voxel observed at different times t. Its average color across all observations The differences between them provide a strong global constraint for the optimization process.
[0048] As those skilled in the art will recognize, in step 140 above, a color consistency constraint term is designed into the first loss function. This is because the point cloud data is dynamically acquired, with point cloud data collected at different times and angles. Therefore, the color consistency constraint term ensures that the color of the same 3D Gaussian voxel remains stable when observed at different times and angles, thereby avoiding accuracy errors caused by extrinsic parameter calibration of data collected by moving lidar and cameras. Furthermore, a rendering error term is also set to constrain the pixel rendering values in the projection image of the 3D Gaussian voxel to be as consistent as possible with the pixel rendering values in the camera image data.
[0049] Those skilled in the art will recognize that the method provided in the above embodiments is applicable to extrinsic parameter calibration of different numbers of LiDARs and cameras. Specifically, the process of this method can be adaptively adjusted according to the number of LiDARs: when there is only one LiDAR, the system will directly perform motion distortion compensation on the point cloud data collected by this single LiDAR to generate a global point cloud map, which is then used for extrinsic parameter calibration with the camera. When there is more than one LiDAR, the system will first independently generate its local, distortion-free dense point cloud map for each LiDAR, and then complete the extrinsic parameter calibration between multiple LiDARs by matching these local maps, finally fusing them into a unified global point cloud map, which is then used for extrinsic parameter calibration with the camera. Furthermore, this method is also applicable to one or more cameras. Regardless of the number of cameras, during the optimization phase, the system will jointly calibrate all cameras with the 3D Gaussian voxels generated from the point cloud map by minimizing a first loss function.
[0050] Those skilled in the art will recognize that the method provided in the above embodiments can perform extrinsic parameter calibration on data collected by a sensor during motion. This method is specifically designed for the characteristics of dynamic data acquisition, and its key features are: First, the system performs motion distortion compensation on the point cloud data of the LiDAR. This method first obtains the sensor's pose data through point cloud registration, and then corrects the distortion of the point cloud based on this pose change information, thereby constructing an accurate global point cloud map. Second, in the joint optimization process of the LiDAR and camera, a color consistency constraint term is added to the first loss function. The core function of this constraint term is to ensure that the color attributes of a single three-dimensional Gaussian voxel constructed from dynamically acquired data (i.e., observations at multiple times and from multiple perspectives) remain consistent. This design provides a strong global constraint for the optimization process, thereby effectively guaranteeing the final accuracy and robustness of the extrinsic parameter calibration with camera image data.
[0051] In this embodiment of the invention, firstly, the scheme compensates for motion distortion in the point cloud data collected during the dynamic process, transforming the deformed raw data into an accurate and distortion-free global point cloud map. This lays the foundation for extrinsic parameter calibration under dynamically acquired data. Then, the global point cloud map is converted into three-dimensional Gaussian voxels with color and covariance attributes, thereby constructing a continuous scene model containing structural features of three-dimensional Gaussian voxels. This makes the structural features upon which the extrinsic parameter calibration between the LiDAR and the camera depends more apparent, thus improving the extrinsic parameter calibration effect. Furthermore, a color consistency constraint term is designed into the first loss function. By forcing the same three-dimensional Gaussian voxel to maintain a stable color when observed at different times and angles, this provides a strong global constraint for the optimization process. Combined with the rendering error term, this significantly improves the robustness and accuracy of extrinsic parameter calibration in complex dynamic scenes.
[0052] The steps described above are for clarity only. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of this invention. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the protection scope of this invention.
[0053] Furthermore, the examples mentioned in the above embodiments can be freely combined, and any combination can be understood as an embodiment. The terms "embodiment" or "example" appearing in various locations in the specification do not necessarily refer to the same embodiment, nor are they independent or alternative embodiments mutually exclusive with other embodiments. Those skilled in the art will understand that the embodiments described herein can be combined with other embodiments.
[0054] Furthermore, in the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0055] Another embodiment of the present invention relates to an electronic device, such as Figure 3 As shown, it includes at least one processor 210; and a memory 220 communicatively connected to at least one processor 210; wherein the memory 220 stores instructions executable by at least one processor 210, the instructions being executed by at least one processor 210 to enable at least one processor 210 to perform the lidar and camera joint calibration method as described above.
[0056] The memory 220 and processor 210 are connected via a bus, which may include any number of interconnecting buses and bridges, connecting various circuits of one or more processors 210 and memory 220 together. The bus may also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides an interface between the bus and the transceiver. The transceiver may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 210 is transmitted over a wireless medium via an antenna, which further receives data and transmits it to processor 210.
[0057] Processor 210 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory 220 can be used to store data used by processor 210 during operation.
[0058] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the method embodiments described above.
[0059] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0060] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes in form and detail may be made without departing from the spirit and scope of the present invention.
Claims
1. A method for joint calibration of lidar and camera, characterized in that, include: Acquire point cloud data from a lidar and image data from a camera; wherein the number of lidar and camera is at least one. Motion distortion compensation is performed on the point cloud data to generate a global point cloud map; The global point cloud map is converted into a series of three-dimensional Gaussian voxels with color and covariance attributes. By jointly optimizing the properties of the three-dimensional Gaussian voxels and the extrinsic parameters of the camera, the first loss function is minimized to complete the extrinsic parameter calibration of the camera; The first loss function includes a color consistency constraint term and a rendering error term. The color consistency constraint term is used to constrain the color attributes of a single 3D Gaussian voxel to remain consistent across multiple time points and viewpoints. The color consistency constraint term is calculated from the deviation between the color of a single 3D Gaussian voxel observed at each time point and the average color of the 3D Gaussian voxel, where the average color is the average value of the color of the 3D Gaussian voxel observed across multiple time points and viewpoints. The rendering error term is the pixel color error between the projected image generated by the 3D Gaussian voxel and the image data from the camera.
2. The laser radar and camera joint calibration method according to claim 1, characterized in that, The step of jointly optimizing the properties of the three-dimensional Gaussian voxels and the extrinsic parameters of the camera to minimize the first loss function, thereby completing the extrinsic parameter calibration of the camera, includes: Based on the initial extrinsic parameters, multiple three-dimensional Gaussian voxels are projected onto the image plane of the camera to obtain a projected image. A first loss function is constructed based on the projected image and the image data from the camera, and the gradient is calculated; Based on the gradient, the properties of the 3D Gaussian voxel, the extrinsic parameters of the camera, and the color and covariance properties of the 3D Gaussian voxel are iteratively updated until the first loss function converges to its minimum value.
3. The laser radar and camera joint calibration method according to claim 1, characterized in that, The method includes: When the number of lidars is greater than or equal to two, motion distortion compensation is performed on the point cloud data of multiple lidars in sequence to obtain point cloud maps of multiple lidars respectively. One of the point cloud maps from multiple lidars is selected as the reference point cloud map, and the remaining point cloud maps are sequentially matched to the reference point cloud map to complete the extrinsic parameter calibration among the multiple lidars, so as to generate a global point cloud map.
4. The laser radar and camera joint calibration method according to claim 3, characterized in that, The step of sequentially matching the remaining point cloud maps to the reference point cloud map to complete the extrinsic parameter calibration among the multiple lidars includes: By iterating through the extrinsic parameters among multiple lidars, the second loss function between the remaining point cloud map and the reference point cloud map is minimized, thus completing the extrinsic parameter calibration among multiple lidars. The second loss function is obtained by weighted calculation of the error vector and the covariance matrix. The error vector is calculated from the values of the corresponding point clouds in the remaining point cloud map and the reference point cloud map. The covariance matrix is calculated from the local covariance of the corresponding point clouds in the remaining point cloud map and the reference point cloud map.
5. The laser radar and camera joint calibration method according to any one of claims 1 to 4, characterized in that, The motion distortion compensation of the point cloud data includes: For the point cloud data, pose data is first obtained through point cloud registration, and then the point cloud data is distorted based on the pose data. The point cloud registration and distortion processing are iteratively performed until the preset constraint conditions are met, and then the motion distortion compensation is completed. The preset constraints are that the pose change, the convergence of the objective function, and the global consistency of the pose simultaneously satisfy the objective constraints.
6. The laser radar and camera joint calibration method according to claim 5, characterized in that, The pose change, objective function convergence, and global pose consistency simultaneously satisfy the objective constraints, including: The rotational change and translational update of the pose before and after point cloud registration are less than a preset pose change threshold, and the change of the objective function value before and after point cloud registration is less than a preset objective function value change threshold, and the deviation of the pose after point cloud registration in the global range is less than a preset global deviation threshold.
7. The laser radar and camera joint calibration method according to claim 5, characterized in that, The distortion processing method is a distortion processing method based on B-spline curves.
8. The laser radar and camera joint calibration method according to claim 5, characterized in that, The point cloud registration method is a point cloud registration method based on probability distribution modeling.
9. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the lidar and camera joint calibration method as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the joint calibration method for lidar and camera as described in any one of claims 1 to 8.