Automatic calibration method and calibration device for laser radar-camera external parameters

By synchronously collecting data from vehicle-mounted LiDAR and cameras and combining semantic segmentation and geometric attributes of vectorized maps, a target consistency function is constructed. This solves the problem of insufficient calibration accuracy of LiDAR-camera extrinsic parameters in existing technologies, achieving high-precision and robust automatic calibration suitable for complex environments.

CN120931733BActive Publication Date: 2026-05-12WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2025-07-15
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, lidar-camera extrinsic parameter calibration methods rely on a large amount of labeled data or specific environments, have weak generalization ability, and are difficult to effectively combine with topological structure information in high-precision maps, resulting in insufficient calibration accuracy and failing to meet the requirements of autonomous driving systems for high-precision environmental perception and fusion.

Method used

By simultaneously collecting point cloud data and image data from vehicle-mounted LiDAR and cameras, semantic segmentation is used to generate a set of semantic masks. Combined with the geometric attributes of vectorized maps, a target consistency function is constructed, and the extrinsic parameter matrix is ​​adjusted to achieve automatic calibration, reducing reliance on manual intervention and prior samples, and improving calibration accuracy and robustness.

Benefits of technology

It achieves high-precision and robust LiDAR-camera calibration in complex traffic scenarios, enhances the system's perception fusion capability, is suitable for dynamic environments and areas with weak textures, and reduces calibration costs and the intensity of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931733B_ABST
    Figure CN120931733B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of multi-sensor fusion, in particular to a laser radar-camera external parameter automatic calibration method and a calibration device, wherein the method comprises the following steps: based on a vehicle-mounted laser radar and a camera, point cloud data and image data of a road scene are synchronously collected; image data is subjected to semantic segmentation to generate a semantic mask set of each image; based on the point cloud data, the geometric properties of the road scene are extracted from a vectorized map, and the laser radar coordinate system of the vehicle-mounted laser radar and points in the vectorized map are projected into the image in combination with the semantic mask set, a target consistency function is constructed; based on the target consistency function, an external parameter matrix is adjusted until the target consistency function reaches a maximum value, and automatic calibration is completed. Therefore, the problems that related technologies rely on a large amount of labeling data or specific environments, have poor generalization ability, cannot combine high-precision maps and road prior information, and result in insufficient calibration precision and difficulty in meeting the high-precision perception requirements of automatic driving are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-sensor fusion technology, and in particular to an automatic calibration method and calibration device for lidar-camera extrinsic parameters. Background Technology

[0002] In related technologies, to achieve efficient fusion of multi-sensor data, it is necessary to accurately calibrate the extrinsic parameters (i.e., relative pose relationship) between LiDAR and camera. LiDAR-camera extrinsic parameter calibration methods can achieve precise registration of point cloud data and image data in the same spatial coordinate system by estimating the rotation matrix and translation vector between the two. This improves the consistency and complementarity of multi-source sensing information, providing a high-precision fusion foundation for downstream tasks such as target detection, environmental perception, and 3D reconstruction.

[0003] However, related technologies typically rely on large amounts of labeled data or specific environmental conditions, have weak generalization capabilities, are difficult to adapt to complex and ever-changing real-world road scenarios, and cannot be effectively combined with topological information in high-precision maps. They also lack full utilization of prior constraints such as roads and lane lines, resulting in calibration accuracy that is difficult to reach the centimeter level. These problems make it difficult to meet the stringent requirements of autonomous driving systems for high-precision environmental perception and fusion, and urgently need to be solved. Summary of the Invention

[0004] This invention provides an automatic calibration method and device for lidar-camera extrinsic parameters, addressing the technical problem that in related technologies, lidar-camera extrinsic parameter calibration methods rely on a large amount of labeled data or specific environmental conditions, have weak generalization ability, cannot be effectively combined with topological structure information in high-precision maps, and lack full utilization of prior constraints such as roads and lane lines, resulting in calibration accuracy that is difficult to reach the centimeter level and cannot meet the stringent requirements of autonomous driving systems for high-precision environmental perception and fusion.

[0005] The first aspect of this invention provides an automatic calibration method for the extrinsic parameters of a LiDAR-camera system, comprising the following steps: simultaneously acquiring point cloud data and image data of a road scene based on a vehicle-mounted LiDAR and a camera; performing semantic segmentation on the image data to generate a semantic mask set for each image; extracting the geometric attributes of the road scene from a vectorized map based on the point cloud data, and combining the semantic mask set to project the LiDAR coordinate system of the vehicle-mounted LiDAR and the points in the vectorized map onto the image to construct a target consistency function; adjusting the extrinsic parameter matrix based on the target consistency function until the target consistency function reaches its maximum value, thereby completing the automatic calibration.

[0006] By employing the above technical means, semantic segmentation is performed on the acquired images, and combined with the geometric attributes of the vectorized map, a target consistency function is constructed to achieve automatic extrinsic parameter calibration between the LiDAR and the camera. This fully utilizes the spatial correspondence between image semantics and map geometry, improves the accuracy and robustness of feature matching, effectively enhances the automation level and environmental adaptability of multimodal sensor calibration, reduces reliance on manual intervention and prior samples, and thus enhances the system's perception fusion capability in complex traffic scenarios.

[0007] Optionally, in one embodiment of the present invention, the step of projecting the lidar coordinate system of the vehicle-mounted lidar and the points in the vectorized map onto the image includes: projecting the points of the lidar coordinate system onto the image to obtain a first set of points falling on each mask in the semantic mask set; and projecting the points in the vectorized map onto the image to obtain a second set of points falling on each mask in the semantic mask set.

[0008] By using the above techniques, points in the LiDAR coordinate system and points in the vectorized map are projected onto the image to obtain a set of points corresponding to the image semantic mask. This enables semantic alignment and spatial association of different data sources in a unified image coordinate system, thereby enhancing the multimodal fusion capability between images, point clouds, and maps. It can effectively improve the accuracy and consistency of target matching during automatic calibration, reduce registration errors caused by differences in perspective or data noise, and ultimately achieve a more accurate and robust LiDAR-camera calibration effect.

[0009] Optionally, in one embodiment of the present invention, constructing the target consistency function includes: based on the normal vector consistency, intensity consistency, and category consistency of points in the first point set, and obtaining map normal vector consistency and map category consistency jointly by the first point set and the second point set; determining the target consistency function based on the map normal vector consistency and the map category consistency.

[0010] By using the above technical means, the target consistency function is determined based on the consistency of map normal vectors and map category. This function can comprehensively reflect the degree of matching of projection points in geometric structure and semantic attributes, enabling accurate identification and association of the same target in multi-source data. This improves the fusion accuracy and automatic calibration reliability between lidar point clouds and images and maps, and enhances the system's adaptability and practicality in complex road environments.

[0011] Optionally, in one embodiment of the present invention, the expression of the target consistency function can be expressed as:

[0012] ,

[0013] in, and These represent the consistency of normal vectors, intensity, and category of points in the first set of points, as well as the consistency of map normal vectors and map category obtained jointly from the first set of points and the second set of points; All represent hyperparameters; N, I, C, M, c These represent the normal vector, intensity, and category of points in the first set of points, as well as the map normal vector and map category obtained jointly from the first set of points and the second set of points.

[0014] By using the above technical means, the target consistency function is determined based on parameters such as normal vector consistency and intensity consistency in the first set of points. This can comprehensively measure the matching degree of point cloud features in spatial structure and reflection characteristics, thereby enhancing the correspondence accuracy of different data sources on the same target, improving the fusion accuracy between point cloud and image / map data, further enhancing the stability and robustness of the external parameter calibration process, and providing reliable multimodal alignment support for high-precision perception systems.

[0015] Optionally, in one embodiment of the present invention, the adjustment formula for the extrinsic parameter matrix can be expressed as:

[0016] ,

[0017] in, T Represents the extrinsic parameter matrix; num Indicates the total selection num Zhang picture; i Indicates the first i Zhang picture; Indicates the first i The target consistency function for the images.

[0018] By using the above technical means, the external parameter matrix can be optimized based on the consistency function after weighted fusion of multiple frames of images. This can effectively alleviate the uncertainty caused by occlusion, noise or local matching error in a single frame of image, improve the stability and accuracy of external parameter estimation, and thus achieve more robust and high-precision automatic calibration between LiDAR and camera. This is applicable to complex perception scenarios in dynamic environments or weak texture areas.

[0019] A second aspect of the present invention provides an automatic calibration device for the extrinsic parameters of a LiDAR-camera system, comprising: an acquisition module for simultaneously acquiring point cloud data and image data of a road scene based on a vehicle-mounted LiDAR and a camera; a generation module for performing semantic segmentation on the image data to generate a semantic mask set for each image; a construction module for extracting the geometric attributes of the road scene from a vectorized map based on the point cloud data, and projecting the LiDAR coordinate system of the vehicle-mounted LiDAR and the points in the vectorized map onto the image in combination with the semantic mask set to construct a target consistency function; and a calibration module for adjusting the extrinsic parameter matrix based on the target consistency function until the target consistency function reaches its maximum value, thereby completing the automatic calibration.

[0020] By employing the above technical means, semantic segmentation is performed on the acquired images, and combined with the geometric attributes of the vectorized map, a target consistency function is constructed to achieve automatic extrinsic parameter calibration between the LiDAR and the camera. This fully utilizes the spatial correspondence between image semantics and map geometry, improves the accuracy and robustness of feature matching, effectively enhances the automation level and environmental adaptability of multimodal sensor calibration, reduces reliance on manual intervention and prior samples, and thus enhances the system's perception fusion capability in complex traffic scenarios.

[0021] Optionally, in one embodiment of the present invention, the construction module includes: a first projection unit, used to project points of the lidar coordinate system onto the image to obtain a first set of points falling on each mask in the semantic mask set; and a second projection unit, used to project points in the vectorized map onto the image to obtain a second set of points falling on each mask in the semantic mask set.

[0022] By using the above techniques, points in the LiDAR coordinate system and points in the vectorized map are projected onto the image to obtain a set of points corresponding to the image semantic mask. This enables semantic alignment and spatial association of different data sources in a unified image coordinate system, thereby enhancing the multimodal fusion capability between images, point clouds, and maps. It can effectively improve the accuracy and consistency of target matching during automatic calibration, reduce registration errors caused by differences in perspective or data noise, and ultimately achieve a more accurate and robust LiDAR-camera calibration effect.

[0023] Optionally, in one embodiment of the present invention, the construction module includes: an acquisition unit, configured to acquire map normal vector consistency and map category consistency based on the normal vector consistency, intensity consistency, and category consistency of points in the first point set, as well as the map normal vector consistency and map category consistency jointly acquired by the first point set and the second point set; and a determination unit, configured to determine the target consistency function based on the map normal vector consistency and the map category consistency.

[0024] By using the above technical means, the target consistency function is determined based on the consistency of map normal vectors and map category. This function can comprehensively reflect the degree of matching of projection points in geometric structure and semantic attributes, enabling accurate identification and association of the same target in multi-source data. This improves the fusion accuracy and automatic calibration reliability between lidar point clouds and images and maps, and enhances the system's adaptability and practicality in complex road environments.

[0025] Optionally, in one embodiment of the present invention, the expression of the target consistency function can be expressed as:

[0026] ,

[0027] in, and These represent the consistency of normal vectors, intensity, and category of points in the first set of points, as well as the consistency of map normal vectors and map category obtained jointly from the first set of points and the second set of points; All represent hyperparameters; N, I, C, M, c These represent the normal vector, intensity, and category of points in the first set of points, as well as the map normal vector and map category obtained jointly from the first set of points and the second set of points.

[0028] By using the above technical means, the target consistency function is determined based on parameters such as normal vector consistency and intensity consistency in the first set of points. This can comprehensively measure the matching degree of point cloud features in spatial structure and reflection characteristics, thereby enhancing the correspondence accuracy of different data sources on the same target, improving the fusion accuracy between point cloud and image / map data, further enhancing the stability and robustness of the external parameter calibration process, and providing reliable multimodal alignment support for high-precision perception systems.

[0029] Optionally, in one embodiment of the present invention, the adjustment formula for the extrinsic parameter matrix can be expressed as:

[0030] ,

[0031] in, T Representing the extrinsic parameter matrix num Indicates the total selection num Zhang picture; i Indicates the first i Zhang picture; Indicates the first i The target consistency function for the images.

[0032] By using the above technical means, the external parameter matrix can be optimized based on the consistency function after weighted fusion of multiple frames of images. This can effectively alleviate the uncertainty caused by occlusion, noise or local matching error in a single frame of image, improve the stability and accuracy of external parameter estimation, and thus achieve more robust and high-precision automatic calibration between LiDAR and camera. This is applicable to complex perception scenarios in dynamic environments or weak texture areas.

[0033] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to achieve automatic calibration of lidar-camera extrinsic parameters as described in the above embodiments.

[0034] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described automatic calibration method for lidar-camera extrinsic parameters.

[0035] A fifth aspect of the present invention provides a computer program product, including a computer program that, when executed, is used to implement the above-described automatic calibration method for lidar-camera extrinsic parameters.

[0036] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0037] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0038] Figure 1 This is a flowchart illustrating an automatic calibration method for lidar-camera extrinsic parameters according to an embodiment of the present invention.

[0039] Figure 2 This is a schematic diagram of the automatic calibration process of lidar-camera extrinsic parameters according to an embodiment of the present invention;

[0040] Figure 3 This is a block diagram of an automatic calibration device for lidar-camera extrinsic parameters according to an embodiment of the present invention;

[0041] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention.

[0042] Figure label:

[0043] 10-Automatic calibration device for lidar-camera extrinsic parameters; 100-Acquisition module, 200-Generation module, 300-Construction module, 400-Calibration module; 401-Memory, 402-Processor, 403-Communication interface. Detailed Implementation

[0044] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0045] The automatic calibration method and device for lidar-camera extrinsic parameters according to embodiments of the present invention are described below with reference to the accompanying drawings. Addressing the technical problems mentioned in the background section regarding lidar-camera extrinsic parameter calibration methods, which rely on large amounts of labeled data or specific environments, have poor generalization ability, fail to incorporate high-precision map topology information and road prior constraints, resulting in insufficient calibration accuracy and difficulty in meeting the high-precision perception and fusion requirements of autonomous driving, the present invention provides an automatic calibration method for lidar-camera extrinsic parameters. In this method, point cloud data and image data of road scenes are simultaneously collected by vehicle-mounted lidar and camera. A SAM (Segment Anything Model) model, along with the normal vectors and category attributes of the vectorized map, is introduced to construct a target consistency function, achieving automatic extrinsic parameter calibration of the lidar and camera. This method eliminates the need for training samples, significantly reducing calibration costs and manual intervention intensity, and exhibits significant autonomous adaptability and deployment flexibility. Furthermore, by incorporating the topological structure information of high-precision maps, the accuracy and robustness of the calibration results in different scenarios are further improved, enhancing the system's practical application capabilities in complex environments. This solves the technical problem in related technologies where calibration methods rely on a large amount of labeled data or specific environments, have poor generalization ability, cannot combine high-precision map topology information and road prior constraints, resulting in insufficient calibration accuracy and difficulty in meeting the requirements of autonomous driving for high-precision perception and fusion.

[0046] Specifically, Figure 1 This is a flowchart illustrating an automatic calibration method for extrinsic parameters of a lidar-camera system provided in an embodiment of the present invention.

[0047] like Figure 1 As shown, the automatic calibration method for the extrinsic parameters of the lidar-camera system includes the following steps:

[0048] In step S101, point cloud data and image data of the road scene are collected simultaneously based on the vehicle-mounted LiDAR and camera.

[0049] Understandably, vehicle-mounted perception systems typically include LiDAR and cameras. LiDAR is used to acquire spatial structural information about the vehicle's surroundings, while cameras collect image data to identify semantic targets such as traffic signs and traffic lights. Fusion processing of both technologies can improve the accuracy and robustness of environmental perception.

[0050] It should be noted that synchronous acquisition refers to the process in which multiple sensors simultaneously collect data at the same or comparable time points based on a unified time reference. This can ensure the consistency of various types of sensing data in the time dimension, and facilitate data fusion and spatiotemporal alignment.

[0051] In embodiments of the present invention, point cloud data may include the spatial coordinates and intensity information (the intensity of the reflected signal recorded by the lidar when measuring the target point) of each lidar point; image data may include RGB information of the road scene.

[0052] In step S102, the image data is semantically segmented to generate a set of semantic masks for each image.

[0053] Semantic segmentation divides images into semantically meaningful categories, such as vehicles, pedestrians, and roads, which can refine environmental perception and scene understanding, and assist autonomous vehicles in making behavioral decisions.

[0054] Specifically, embodiments of the present invention can utilize SAM for semantic segmentation of images. The SAM model, through deep learning algorithms, can accurately identify various objects in an image and generate a semantic mask set for each image, which can be represented as:

[0055] M = ,

[0056] in, It can represent the first i One category; i Indicates the position index. .

[0057] Specifically, in embodiments of the present invention, the mask is labeled as Each mask is a binary matrix of the same size as the image, representing a region of an object in the image, which can be further subdivided into multiple categories, such as lane lines, buildings, pedestrians, etc.

[0058] It should be noted that, Represents pixels Does it belong to the first i A mask, among which... Represents pixels Belongs to the One example, Represents pixels It does not belong to the i-th instance.

[0059] In step S103, based on point cloud data, the geometric attributes of the road scene are extracted from the vectorized map, and combined with the semantic mask set, the LiDAR coordinate system of the vehicle-mounted LiDAR and the points in the vectorized map are projected onto the image to construct a target consistency function.

[0060] In one embodiment of the present invention, the geometric properties may include, but are not limited to, the road surface normal vector. Category of objects in the scene wait.

[0061] The embodiments of the present invention extract geometric attributes from vector maps, which can provide structured prior information for spatial alignment of images and point clouds, assist in constructing spatial correspondences between multimodal data, thereby improving the accuracy and stability of feature matching.

[0062] Optionally, in one embodiment of the present invention, projecting the LiDAR coordinate system of the vehicle-mounted LiDAR and the points in the vectorized map into the image includes: projecting the points of the LiDAR coordinate system into the image to obtain a first set of points falling on each mask in the semantic mask set; and projecting the points in the vectorized map into the image to obtain a second set of points falling on each mask in the semantic mask set.

[0063] As one possible implementation, embodiments of the present invention can project points in the lidar coordinate system onto an image, as shown in the following formula:

[0064] ,

[0065] in, K This represents the intrinsic parameter matrix of the camera; T This represents the extrinsic parameter matrix of the camera; P Represents a point in the lidar coordinate system; This indicates the depth of a point in the camera coordinate system; These represent the x-coordinate and y-coordinate of a point in the pixel coordinate system, respectively. These represent the x-coordinate, y-coordinate, and vertical coordinate of a point in the lidar coordinate system, respectively.

[0066] It should be noted that, in the embodiments of the present invention, it is assumed that the camera's intrinsic parameters are known and remain unchanged; Represents the coordinates in the image coordinate system; ( , , )This represents the coordinates of a point in the lidar coordinate system.

[0067] Furthermore, for each mask We can obtain a set of points that fall on it, which can be denoted as S :

[0068] ,

[0069] The following is a detailed explanation of how the present invention projects points from a vectorized map onto an image, using a specific example.

[0070] This invention can transform points on a vectorized map. Q( , , ) Using the transformation matrix Transforming to the lidar coordinate system, the formula can be shown below:

[0071] ,

[0072] in, Represents a point in the lidar coordinate system; Represents the transformation matrix; Represents points on a vectorized map.

[0073] It should be noted that the transformation matrix It can be obtained in real time through the positioning system.

[0074] Furthermore, embodiments of the present invention can employ the same projection method to project corresponding points. Projected onto the image, we get q(u1, v1) .

[0075] Furthermore, for each mask We can obtain a set of points that fall on it, which can be denoted as :

[0076] S1 ,

[0077] Understandable, It is a collection of LiDAR projection points. It is a set of map projection points.

[0078] Optionally, in one embodiment of the present invention, constructing a target consistency function includes: based on the normal vector consistency, intensity consistency, and category consistency of points in the first point set, and the map normal vector consistency and map category consistency obtained jointly by the first point set and the second point set; determining the target consistency function based on the map normal vector consistency and map category consistency.

[0079] As one possible implementation method, embodiments of the present invention may employ... This represents the number of points in set S, and at the same time S1 The same selection process was also conducted. This point makes this Points and sets S In N Each point in the graph has a one-to-one correspondence.

[0080] Furthermore, in order to measure sets S Consistency of midpoints, and sets S With sets S1 To ensure consistency at corresponding points, this invention proposes five consistency functions. , respectively represent in the set S Consistency of normal vectors at midpoints, consistency of intensity, consistency of category, and consistency of set S and set S1 The consistency of the obtained map normal vectors and map categories.

[0081] The following examples illustrate the five consistency functions.

[0082] (1) For the normal vector, the function It can be the average of the pairwise dot products of all vectors in set S, thus allowing the construction of a matrix. A and B .

[0083] Specifically, it can be expressed as:

[0084] , ,

[0085] in, Represents a set S The normal vector of the midpoint.

[0086] Furthermore, The formula can be as follows:

[0087] ,

[0088] Where N represents the number of points in set S, i represents the i-th normal vector in A, jLet A represent the j-th normal vector in A.

[0089] (2) For intensity, the function Through sets S The formula for calculating the variance of all intensity values ​​is as follows:

[0090] ,

[0091] in, express S Points in The corresponding strength.

[0092] (3) For segmentation categories, first calculate the set S The points for each category can be represented as: ,

[0093] in, Clustering i The number of points in the middle, C indicates The number of clusters.

[0094] In an embodiment of the present invention, the array can be sorted from smallest to largest, as follows:

[0095] ,

[0096] As one possible approach, the category consistency score can be calculated as follows:

[0097] ,

[0098] ,

[0099] in, k This is an empirical value, usually taken as 0.5.

[0100] Understandably, if the categories are more concentrated, then the consistency score will be higher.

[0101] (4) For map normal vector consistency The formula can be as follows:

[0102] ,

[0103] in, Let P be the normal vector of the lidar point P. Represents the normal vector of a map point; Indicates mask The number of valid point pairs; P represents the point in the image mask; Q represents the corresponding point in the high-precision map.

[0104] (5) For map category consistency In each mask Below, the projection points of the lidar can be counted. S semantic categories and corresponding point sets in vector maps Category tags Match rate:

[0105] ,

[0106] in, This indicates an indicator function, which is 1 when the categories match and 0 otherwise. Let P be the semantic category; Indicates mask The number of valid pairs within the region; Let P be the category of point Q in the map; P represents the point in the image mask, and Q represents the corresponding point in the high-precision map.

[0107] In actual execution, for each frame of the image, a score for each type of consistency can be obtained, and the formula can be as follows:

[0108] ,

[0109] in, , indicating the type of consistency function; m This indicates that there are a total of [number] images in a frame. m Seed mask; ; Indicates weight; This represents a compensation function for the sparsity of point clouds.

[0110] Specifically, the weight of each mask It is the ratio of the number of LiDAR points that fall on the corresponding mask positions in the image to the total number of points, which can be expressed as:

[0111] ,

[0112] It should be noted that in the embodiments of the present invention, distant or occluded areas are projected onto a smaller point cloud. Because even under incorrect extrinsic parameters, fewer points tend to have higher consistency and therefore lower reliability, a penalty is applied to the mask with fewer points. The function is a monotonically increasing function of the number of points and can be expressed as:

[0113] ,

[0114] in, and It's a hyperparameter; Indicates mask The number of valid pairs of points within the area.

[0115] Ultimately, the objective function obtained in this embodiment of the invention for a single frame is a weighted fusion of the above five consistency functions.

[0116] Optionally, in one embodiment of the present invention, the expression for the target consistency function can be:

[0117] ,

[0118] in, and These represent the consistency of normal vectors, intensity, and category of points in the first set of points, as well as the consistency of map normal vectors and map category obtained jointly from the first and second sets of points. Both represent hyperparameters.

[0119] In step S104, the extrinsic parameter matrix is ​​adjusted based on the target consistency function until the target consistency function reaches its maximum value, thus completing the automatic calibration.

[0120] Optionally, in one embodiment of the present invention, the adjustment formula for the extrinsic parameter matrix can be expressed as:

[0121] ,

[0122] in, T Represents the extrinsic parameter matrix; num Indicates the total selection num Zhang picture; i Indicates the first i Zhang picture; Indicates the first i The target consistency function for the images.

[0123] In some cases, in order to obtain stronger constraints, embodiments of the present invention select multiple frames of images for optimization to obtain the final extrinsic parameter matrix T.

[0124] As a concrete example, the overall process of automatic calibration of LiDAR-camera extrinsics can be as follows: Figure 2 As shown. Embodiments of this invention can achieve accurate semantic segmentation without training samples using the SAM model, improving target extraction capabilities in unknown environments; they can also optimize the extrinsic parameter matrix using multiple frames of images, enhancing the stability and accuracy of LiDAR and camera extrinsic parameter calibration. This approach is applicable to areas with weak texture and complex road environments, demonstrating good versatility and practical application value.

[0125] The automatic calibration method for lidar-camera extrinsic parameters proposed in this embodiment of the invention integrates lidar point cloud, image semantic segmentation, and geometric attributes of vectorized maps to achieve automatic calibration in complex scenarios. It can achieve automatic calibration without relying on training samples, and has significant autonomous adaptability and deployment flexibility. It can significantly reduce calibration costs and the intensity of manual intervention. At the same time, it can combine the topological structure information of high-precision maps to further improve the accuracy and robustness of calibration results in different scenarios, and enhance the system's practical application capability in complex environments.

[0126] Next, with reference to the accompanying drawings, an automatic calibration device for the extrinsic parameters of a lidar-camera system according to an embodiment of the present invention is described.

[0127] Figure 3 This is a block diagram of an automatic calibration device for lidar-camera extrinsic parameters according to an embodiment of the present invention.

[0128] like Figure 3 As shown, the automatic calibration device 10 for the external parameters of the lidar-camera includes: an acquisition module 100, a generation module 200, a construction module 300, and a calibration module 400.

[0129] The acquisition module 100 is used to simultaneously acquire point cloud data and image data of the road scene based on the vehicle-mounted LiDAR and camera.

[0130] The generation module 200 is used to perform semantic segmentation on the image data to generate a set of semantic masks for each image.

[0131] Module 300 is used to extract the geometric attributes of the road scene from the vectorized map based on point cloud data, and combine the semantic mask set to project the LiDAR coordinate system of the vehicle-mounted LiDAR and the points in the vectorized map onto the image to construct the target consistency function.

[0132] The calibration module 400 is used to adjust the extrinsic parameter matrix based on the target consistency function until the target consistency function reaches its maximum value, thus completing the automatic calibration.

[0133] Optionally, in one embodiment of the present invention, the construction module 300 includes: a first projection unit and a second projection unit.

[0134] The first projection unit is used to project points from the lidar coordinate system onto the image to obtain the first set of points falling on each mask in the semantic mask set.

[0135] The second projection unit is used to project points in the vectorized map onto the image to obtain a second set of points that fall on each mask in the semantic mask set.

[0136] Optionally, in one embodiment of the present invention, the construction module includes: an acquisition unit and a determination unit.

[0137] The acquisition unit is used to acquire map normal vector consistency and map category consistency based on the normal vector consistency, intensity consistency, and category consistency of points in the first point set, as well as the map normal vector consistency and map category consistency obtained jointly by the first point set and the second point set.

[0138] The unit is determined based on the consistency of map normals and map categories to determine the target consistency function.

[0139] Optionally, in one embodiment of the present invention, the expression for the target consistency function can be expressed as:

[0140] ,

[0141] in, and These represent the consistency of normal vectors, intensity, and category of points in the first set of points, as well as the consistency of map normal vectors and map category obtained jointly from the first and second sets of points. All represent hyperparameters; N, I, C, M, c These represent the normal vector, intensity, and category of points in the first set of points, as well as the map normal vector and map category obtained jointly from the first set of points and the second set of points.

[0142] Optionally, in one embodiment of the present invention, the adjustment formula for the extrinsic parameter matrix can be expressed as:

[0143] ,

[0144] in, T Represents the extrinsic parameter matrix; num Indicates the total selection num Zhang picture; i Indicates the first i Zhang picture; Indicates the first i The target consistency function for the images.

[0145] It should be noted that the explanation of the above-described embodiment of the automatic calibration method for lidar-camera extrinsic parameters also applies to the automatic calibration device for lidar-camera extrinsic parameters in this embodiment, and will not be repeated here.

[0146] The automatic calibration device for lidar-camera extrinsics proposed in this embodiment of the invention integrates lidar point cloud, image semantic segmentation, and geometric attributes of vectorized maps to achieve automatic calibration in complex scenarios. It can achieve automatic calibration without relying on training samples, and has significant autonomous adaptability and deployment flexibility. It can significantly reduce calibration costs and the intensity of manual intervention. At the same time, it can combine the topological structure information of high-precision maps to further improve the accuracy and robustness of calibration results in different scenarios, and enhance the system's practical application capability in complex environments.

[0147] Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. The electronic device may include:

[0148] The memory 401, the processor 402, and the computer program stored on the memory 401 and capable of running on the processor 402.

[0149] When the processor 402 executes the program, it implements the automatic calibration method for the extrinsic parameters of the lidar-camera provided in the above embodiments.

[0150] Furthermore, electronic devices also include:

[0151] Communication interface 403 is used for communication between memory 401 and processor 402.

[0152] The memory 401 is used to store computer programs that can run on the processor 402.

[0153] Memory 401 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0154] If the memory 401, processor 402, and communication interface 403 are implemented independently, then the communication interface 403, memory 401, and processor 402 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0155] Optionally, in a specific implementation, if the memory 401, processor 402, and communication interface 403 are integrated on a single chip, then the memory 401, processor 402, and communication interface 403 can communicate with each other through an internal interface.

[0156] Processor 402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0157] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described automatic calibration method for lidar-camera extrinsic parameters.

[0158] This invention also provides a computer program product, including a computer program that can run computer instructions. When these computer instructions are executed by a processor, they implement the automatic calibration method for lidar-camera extrinsic parameters provided in this invention.

[0159] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0160] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0161] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0162] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0163] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0164] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0165] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0166] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. An automatic calibration method for the extrinsic parameters of a lidar-camera system, characterized in that, Includes the following steps: Based on vehicle-mounted LiDAR and cameras, point cloud data and image data of road scenes are collected simultaneously; The image data is semantically segmented to generate a set of semantic masks for each image; Based on the point cloud data, the geometric attributes of the road scene are extracted from the vectorized map, and combined with the semantic mask set, the lidar coordinate system of the vehicle-mounted lidar and the points in the vectorized map are projected onto the image to construct a target consistency function. The extrinsic parameter matrix is ​​adjusted based on the target consistency function until the target consistency function reaches its maximum value, thus completing the automatic calibration. The step of projecting the LiDAR coordinate system of the vehicle-mounted LiDAR and the points in the vectorized map onto the image to construct a target consistency function includes: projecting the points of the LiDAR coordinate system onto the image to obtain a first set of points falling on each mask in the semantic mask set; projecting the points in the vectorized map onto the image to obtain a second set of points falling on each mask in the semantic mask set; obtaining map normal vector consistency and map category consistency based on the normal vector consistency, intensity consistency, and category consistency of the points in the first set of points, as well as the combined map normal vector consistency and map category consistency of the first set of points and the second set of points; and determining the target consistency function based on the map normal vector consistency and the map category consistency. For map normal vector consistency The formula is as follows: , in, Let P be the normal vector of the lidar point P. Represents the normal vector of a map point; Indicates mask The number of valid point pairs; P represents the point in the image mask; Q represents the corresponding point in the high-precision map; For map category consistency In each mask Below, statistical lidar projection points S semantic categories and corresponding point sets in vector maps Category tags The matching rate is calculated using the following formula: , in, This indicates an indicator function, which is 1 when the categories match and 0 otherwise. Let P be the semantic category.

2. The automatic calibration method for lidar-camera extrinsic parameters according to claim 1, characterized in that, The expression for the target consistency function is: , in, and These represent the consistency of normal vectors, intensity, and category of points in the first set of points, as well as the consistency of map normal vectors and map category obtained jointly from the first set of points and the second set of points; All represent hyperparameters; N, I, C, M, c These represent the normal vector, intensity, and category of points in the first set of points, as well as the map normal vector and map category obtained jointly from the first set of points and the second set of points.

3. The automatic calibration method for lidar-camera extrinsic parameters according to claim 2, characterized in that, The adjustment formula for the extrinsic parameter matrix is: , in, T Represents the extrinsic parameter matrix; num Indicates the total selection num Zhang picture; i Indicates the first i Zhang picture; Indicates the first i The target consistency function for the images.

4. An automatic calibration device for the extrinsic parameters of a lidar-camera system, characterized in that, include The acquisition module is used to simultaneously acquire point cloud data and image data of road scenes based on vehicle-mounted LiDAR and cameras; The generation module is used to perform semantic segmentation on the image data to generate a set of semantic masks for each image; The construction module is used to extract the geometric attributes of the road scene from the vectorized map based on the point cloud data, and combine the semantic mask set to project the lidar coordinate system of the vehicle-mounted lidar and the points in the vectorized map onto the image to construct a target consistency function. The calibration module is used to adjust the extrinsic parameter matrix based on the target consistency function until the target consistency function reaches its maximum value, thereby completing the automatic calibration. The construction module includes: a first projection unit for projecting points from the lidar coordinate system onto the image to obtain a first set of points falling on each mask in the semantic mask set; a second projection unit for projecting points from the vectorized map onto the image to obtain a second set of points falling on each mask in the semantic mask set; an acquisition unit for acquiring map normal vector consistency and map category consistency based on the normal vector consistency, intensity consistency, and category consistency of points in the first set of points, as well as the combined map normal vector consistency and map category consistency of the first set of points and the second set of points; and a determination unit for determining the target consistency function based on the map normal vector consistency and the map category consistency. For map normal vector consistency The formula is as follows: , in, Let P be the normal vector of the lidar point P. Represents the normal vector of a map point; Indicates mask The number of valid point pairs; P represents the point in the image mask; Q represents the corresponding point in the high-precision map; For map category consistency In each mask Below, statistical lidar projection points S semantic categories and corresponding point sets in vector maps Category tags The matching rate is calculated using the following formula: , in, This indicates an indicator function, which is 1 when the categories match and 0 otherwise. Let P be the semantic category.

5. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the automatic calibration method for lidar-camera extrinsic parameters as described in any one of claims 1-3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the automatic calibration method for lidar-camera extrinsic parameters as described in any one of claims 1-3.

7. A computer program product, comprising a computer program, characterized in that, The computer program is executed to implement the automatic calibration method for lidar-camera extrinsic parameters as described in any one of claims 1-3.