Method for camera and lidar joint extrinsic calibration based on a single calibration board

EP4690118A1Pending Publication Date: 2026-02-11HARMAN INT IND INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023721587
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2026-02-11

Smart Images

  • Figure CN2023086915_10102024_PF_FP_ABST
    Figure CN2023086915_10102024_PF_FP_ABST
Patent Text Reader

Abstract

A method for camera and lidar joint extrinsic calibration based on a single calibration board is provided. The method applies novel algorithms to jointly estimate extrinsic of both at least one camera and at least one lidar through a single calibration board. More specifically, the proposed novel calibration board is compact and has distinct calibration targets for both the camera and lidar sensors. The proposed algorithms can efficiently and accurately detect calibration targets from the cameras and the lidars. Consequently, based on the detected calibration targets, extrinsic of the at least one camera and the at least one lidar can be jointly estimated. Due to the novel design of calibration board and algorithms, the calibration setups are simple, and fast, meanwhile, the calibration ability are robust, and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR CAMERA AND LIDAR JOINT EXTRINSIC CALIBRATION BASED ON A SINGLE CALIBRATION BOARDTECHNICAL FIELD

[0001] The present inventive subject matter generally relates to driver assistance and image / video processing. More particularly, the present inventive subject matter relates to a method for camera and lidar joint extrinsic calibration based on a single calibration board.BACKGROUND

[0002] Advanced Driver Assistance Systems (ADAS) and Autonomous Driving (AD) are systems developed to automate / adapt / enhance vehicles for safe and convenient driving. The ADAS / AD can monitor the internal / external environment of vehicles through various sensors, such as cameras, lidars, radars and ultrasound. The collected sensor data can be processed by the onboard processer (s) to extract useful information and take appropriate actions. These actions could be as simple as turning on / off a warning light or as complex as taking control of the throttle, the braking, and the steering of a vehicle.

[0003] For various ADAS / AD systems, to take appropriate actions on vehicles, it is necessary to extract accurate perceptions of the surrounding environment, which may include: 1) static objects / scenes, such as buildings, roads, lanes, and traffic signs; 2) dynamic objects, such as vehicles and pedestrians. The recognition task includes various information about an object, such as its categories, attributes, positions, distances, postures, and speeds. Among all tasks, for 3D related sensing tasks (position, distance, etc. ) , the sensor extrinsic calibration can be one of the most important issues.

[0004] With the development of ADAS / AD systems, due to the requirements for accuracy and robustness, the design of these systems is facing topologies with multiple sensors for implementing multi-modality perception and fusion. In order to represent and fuse perception results in a common reference frame (i.e., ego-vehicle coordinate) , the rigid transformations (3D  rotation and translation) between all sensors and ego-vehicle should be known. Consequently, this kind of fusion raises a further calibration issue, namely, the joint calibration of multiple sensors.

[0005] Due to the ability to provide abundant appearance information, camera sensors have been widely applied in various ADAS / AD systems. On the other hand, lidar sensors have high accuracy of range measurement, which thus have been widely applied. Since these two kinds of sensors can provide complementary information, it is suitable to apply them in a same ADAS / AD system. In this kind of design, to effectively combine data / results of both sensors for implementing multi-modality perception and fusion, the relative extrinsic parameters between cameras and lidars should be jointly determined first. However, existing joint calibration methods suffer from several problems: 1) complex setups; 2) lack of generalization ability. Therefore, the accuracy of calibration is strongly dependent on the carefully designed settings and the structured environments.

[0006] There is a need to provide a method for camera and lidar joint extrinsic calibration, which may apply algorithms to jointly estimate extrinsic of both camera and lidar sensors through a single calibration board.

[0007] SUMMARY OF THE INVENTIVE SUBJECT MATTER

[0008] A method provided in the present inventive subject matter may be used for the joint extrinsic calibration between at least one camera and at least one lidar. The method may comprise positioning first plurality of center points from 2D image comprising the single calibration board captured by the at least one camera, and positioning second plurality of center points from 3D point cloud containing points reflected from the single calibration board returned by the at least one lidar. The method may further comprise estimating an extrinsic transformation between the at least one camera and the at least one lidar. In the method, the first plurality of center points belong to rectangular grids on the single calibration board, respectively, and the second plurality of center points belong to circular holes in the single calibration board, respectively. It is essential that the first plurality of center points and the second plurality of center points are completely coincident in position.

[0009] A non-transitory computer readable medium storing instruction that can be processed by one or more processor to realize that positioning first plurality of center points from 2D image comprising the single calibration board captured by the at least one camera, and positioning second plurality of center points from 3D point cloud containing points reflected from the single calibration board returned by the at least one lidar. The non-transitory computer readable medium storing instruction that can be processed by one or more processor to further realize estimating an extrinsic transformation between the at least one camera and the at least one lidar. The first plurality of center points belong to rectangular grids on the single calibration board, respectively, and the second plurality of center points belong to circular holes in the single calibration board, respectively. It is essential that the first plurality of center points and the second plurality of center points are completely coincident in position.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The present inventive subject matter may be better understood from reading the following description of non-limiting embodiments, with reference to the attached drawings. In the figures, like reference numeral designates corresponding parts, wherein below:

[0011] FIG. 1 shows an example diagram illustrating a scene of joint external calibration of a camera and a lidar based on a single calibration board, in accordance with one or more embodiments of the inventive subject matter;

[0012] FIG. 2A shows an example diagram illustrating a front view of a joint calibration pattern, based on which the method for camera and lidar joint extrinsic calibration can be performed, in accordance with one or more embodiments of the inventive subject matter;

[0013] FIG. 2B and 2C each shows an alternative example of FIG. 2A, in accordance with one or more embodiments of the inventive subject matter;

[0014] FIG. 2D shows an expanded example of FIG. 2C, in accordance with one or more embodiments of the inventive subject matter;

[0015] FIG. 3 shows an example flowchart illustrating the method for camera and lidar joint extrinsic calibration, in accordance with one or more embodiments of the inventive subject matter;

[0016] FIG. 4A-4D shows an example processing for positioning the four center points from a 2D camera image of the single calibration board captured by the camera, wherein the single calibration target of FIG. 2C is contained, in accordance with one or more embodiments of the inventive subject matter; and

[0017] FIG. 5A-5D shows an example processing for positioning the four center points from 3D point cloud scanned and returned by the lidar, wherein the single calibration target of FIG. 2C is contained, and wherein the X and Y coordinate axes are shown as a reference, in accordance with one or more embodiments of the inventive subject matter.

[0018] DETAILED DESCRIPTION OF THE INVENTIVE SUBJECT MATTER

[0019] The detailed description of the one or more embodiments of the present inventive subject matter is disclosed hereinafter; however, it is understood that the disclosed embodiments are merely exemplary of the inventive subject matter that may be embodied in various and alternative forms. The figures are not necessarily to scale; some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and function details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ the present inventive subject matter.

[0020] Among the numerous types of sensor data fusion methods, the combination of at least one lidar and at least one camera is one of the most commonly used pairs of sensors for driving environment perception. The lidar may be able to provide 3D point cloud data, which include accurate depth and reflection intensity information, while the camera captures rich semantic information of the scene. The combination of the at least one camera and the at least one lidar provides the feasibility to overcome the flaws of each of the sensors. The main challenge in fusing these two heterogeneous sensors is to find the rigid body transformation between the sensor coordinate systems by performing the joint extrinsic calibration.

[0021] In one aspect of the present inventive subject matter, a method for at least one camera and at least one lidar joint extrinsic calibration based on a single calibration board is provided.

[0022] FIG. 1 shows an example diagram illustrating a scene 100 of joint external calibration of a camera 120 and a lidar 130 based on a single calibration board 110, in accordance with one or  more embodiment of the inventive subject matter. As shown in the scene 100 of FIG. 1, the camera 120 and the lidar 130 are fixed to each other and both installed on the roof of an autonomous driving vehicle 140, while let the targets on the board completely appear in the overlapping area of the fields of vision (FOVs) of both the camera 120 and the lidar 130, and then simultaneously enable the camera 120 and the lidar 130 to detect their respective calibration features from the board.

[0023] In target-based extrinsic calibration methods, the setting of calibration targets is particularly important and critical. For both the camera and lidar sensors, edges and corners are the features that can be detected accurately and robustly. For example, for the camera sensors, a chessboard with rectangular grids may be the most straightforward calibration board design, and many existing algorithms can be conveniently applied for detecting these calibration patterns with high accuracy. However, for lidar sensors, since the data are much sparse, nearly horizontal edges of the rectangular chessboard might not intersect with any of the lidar scan beams. Therefore, it is hard to detect rectangular shaped objects from the lidar data. Moreover, the rectangular patterned chessboard can only provide limited information.

[0024] On the other hand, a board with circular holes may be a good design for lidar sensors (also for stereo cameras) as the holes can be accurately detected when intersecting with few lidar beams. However, due to the lack of distinct semantic appearance, it is much harder for a mono camera to accurately detect the calibration patterns.

[0025] In the aspect, a novel-designed calibration pattern with detecting targets that can be detected by both at least one camera and at least one lidar to be calibrated has been introduced herein, from which all the sensors to be calibrated can detect their corresponding distinct features, respectively.

[0026] In order to present a calibration pattern that can be detectable in all relevant sensors of the at least one camera and the at least one lidar to be joint calibrated, and observable features extracted for localization, the calibration pattern may be designed with detecting targets including rectangular grids with circular holes each concentrically arranged inside. All the vertices of the rectangular grids have each been arranged as one corner point, which generally can be defined as the point where black (or white) rectangles intersect, that is, such as the common vertex of the opposite corners of the two black rectangles diagonally adjacent to each other. By detecting such calibration targets, the cameras may determine the corresponding  rectangular grids by detecting each four corner points thereof, respectively, while the lidars may obtain the corresponding edges of the circular holes. Thereby the at least one camera and the at least one lidar to be joint calibrated may further work out their respective center points from the rectangular grids and the circular holes at each of the completely coincident locations, respectively.

[0027] FIG. 2A shows an example diagram illustrating a front view of a calibration pattern 200, based on which the method for camera and lidar joint extrinsic calibration can be performed, in accordance with one or more embodiment of the inventive subject matter.

[0028] As shown in FIG. 2A, the calibration pattern 200 includes planar targets of four rectangular grids 210, 212, 214, 216 and four corresponding circular holes 220, 222, 224, 226 concentrically arranged. The four rectangular grids 210, 212, 214, 216 may be identical in size and shape, and symmetrically disposed on a two-dimensional plane as 2×2 (in 2 rows and 2 columns) grids, inside of which the four circular holes 220, 222, 224, 226 may be disposed, correspondingly. It is easy to make the side length of each rectangular grid at least being larger than the diameter of the corresponding circular hole inside thereof, so as to set the corner points 230a-230i at each of the vertices of the four intersect rectangular grids 210, 212, 214, 216. The corner points and be each set by a tiled shape of the diagonal intersection of two black rectangles, such as the two diagonal intersect black rectangles at the top left and bottom right and the two white at the bottom left and top right as shown in FIG. 2A, or vice versa in black and white. In the joint calibration method herein, it is essential to arrange each rectangular grid 210, 212, 214, 216 and its circular hole 220, 222, 224, 226 therein to have a completely coincident position for the center point P0, P1, P2 or P3, correspondingly.

[0029] According to the example shown in FIG. 2A, there may be nine corner points 230a-230i in the calibration pattern 200 desired to be detected by the camera, in turn the four rectangular grids 210, 212, 214, 216 can be determined, and thus the positions of the four center points P0, P1, P2, and P3 can be worked out. While for the lidar, the edges of the four circular holes 220, 222, 224, 226 can be detected to determine the positions of their four center points P0, P1, P2 and P3, respectively, which completely coincide with those of the rectangular in position.

[0030] FIG. 2B shows an alternative example of FIG. 2A, in accordance with one or more embodiments of the inventive subject matter, wherein the tiled black and white rectangles used to define the four corner points 230b’、 230d’、 230f’、 230h’ of the nine are color inversed from  those 230b、 230d、 230f、 230h in FIG. 2A. Those skilled in the art may understand that such differences will not change the position detection of the camera and the lidar on their respective calibration targets.

[0031] In one and more embodiments, as shown in FIG. 2A and FIG. 2B, the four rectangular grids 210, 212, 214, 216 are represented in 2×2 grids and are arranged all in a light color, e.g., in white. In this way, it can be seen that most areas of the calibration pattern 200 are in light color. As the camera requires a good lighting environment, such most area of the entire calibration pattern formed in light color can be more easily segmented in a scene with insufficient light, such as in underground garages. Alternatively, it is easy to detect a relatively bright area and perform rough positioning of the calibration board area, so that all subsequent corner detection can be performed near this rough positioning area.

[0032] In one and more embodiments, it is also possible to set some outer edges with the inverse colors from the 2×2 rectangular grids, respectively, which may be easier for the grids to be detected by the camera. In an example as shown in FIG. 2C, there are two of the four rectangular grids each has been arranged with inverse-colored elongate borders 240, 242, thereby forming a semi-wrapped structure for the four rectangular grids 210, 212, 214, 216. The corner areas 244, 246 of a light background color near the black wrap can be detected easier. If these two corners can be successfully detected, the area between the vertices of these two corners (i.e., the two corner points 230c, 230g) may be detected as the approximate location of the calibration board, which brings advantageous for the camera to be joint calibrated herein.

[0033] In one or more alternative embodiments, setting a square gridded pattern outside of each corresponding circular hole, instead of the rectangular grid, may be an optimal choice, not only because it is easy to detect itself, but also because it may simplify the algorithm and save the computation on the single calibration board, that is, the combination of the concentric circular and square shapes may be occupying smallest area. Compared to other conventional calibration targets used for multi-sensors joint calibrations, for the calibration boards being made of the same size, the information collectable per frame (such as the number of detection points) can be multiply increased by the calibration pattern of the inventive subject matter, resulting in correspondingly improvements in calibration accuracy. Therefore, the joint calibration pattern with the rectangular-grid arrangement may preferably be square.

[0034] In one or more embodiments, in order for all relevant sensors of the at least one camera and  the at least one lidar to be able to detect the graphical features set in the joint calibration pattern, the detecting targets should be processed and manufactured on a single calibration board, for example, with pre-set dimensions. The single calibration board then can be as compact as possible, easy to detect, and easy for calibration process performing.

[0035] The single calibration board provided in the inventive subject matter can provide sufficient information with very compact arrangement, and it may have variable scalable implementations when needed. The number of combinations of the rectangular grids with the corresponding concentrically circular holes arranged therein can be expanded to such arrangement of M×N, where M≥2 and N≥2 in integer. In this way, the amount of information obtained per detection can be further increased to improve the accuracy or speed of calibration. In one or more embodiments of the calibration target according to the inventive subject matter, the combined arrangement with the rectangular plus circular hole therein may be, for example, as 2×3, 3×3, 2×4, 3×4 , …, or M×N, where M≥2 and N≥2 in integer, respectively. As long as it may satisfy the requirements of the at least one camera and the at least one camera to be calibrated for the clarity and plenty of the information extraction for their joint calibration target detection. The more combinations, the more detective information, and thus the more accurate calibration achieved.

[0036] FIG. 2D shows an expanded example FIG. 2C, in accordance with one or more embodiments of the inventive subject matter. As shown in FIG. 2D, the arrangement of the rectangular grids with the concentric circular holes are expanded to 4×2, i.e., there are eight grid rectangular grids and corresponding circular holes contained in the joint calibration pattern. Therefore, it can be expected that, compared to the joint calibration pattern in FIG. 2C, there may be eight center points P0-P7 desired to be positioned by the camera and the lidar to be joint calibrated, respectively, doubling the amount of information for the joint calibration.

[0037] It can be noted that, the novel joint calibration pattern provided in the inventive subject matter have significant technical advantages, which is compact and has distinct calibration targets and / or features for both the at least on camera and the at least one lidar. As the robust calibration targets, the rectangular grids and the circular holes can be accurately detected from the lidar and camera sensors, respectively. Meanwhile, the proposed joint calibration pattern can be easily manufactured in a compact board, so that can be easily applied in most cases. Moreover, because of the compact design, the joint calibration board can be easily extended by adding more  rectangular grids and circular holes, to provide further accuracy and robustness.

[0038] FIG. 3 shows an example flowchart 300 illustrating the method for at least one camera and at least one lidar joint extrinsic calibration based on a single calibration board, in accordance with one or more embodiments of the inventive subject matter. the method comprises the steps S310-S370, as will be described below. In the raised one and more embodiments, these steps will be described in conjunction with the scene as shown in FIG. 1 and the calibration patterns as shown in FIG. 2C. However, those skilled in the art may appreciate that other scenes or calibration patterns may be also applicable.

[0039] In step S310, the designed calibration targets can be instantiated and fabricated on a single calibration board with known pre-set dimensions. Then, the fabricated single calibration board may be placed to ensure it appears in a place of a common field of view (FOV) of the at least one camera and the at least one lidar to be joint calibrated.

[0040] As shown in FIG. 1, the single joint calibration board 110 containing the calibration pattern of FIG. 2C is suggested to horizontally placed in front of the camera 120 and the lidar 130 to be joint calibrated, and is generally oriented as facing towards both the relevant sensors. The calibration board should appear in the common FOV of the camera 120 and the lidar130. Then, the camera 120 may be able to capture 2D image frames containing the single joint calibration board 110, and the lidar 130 may be able to continuously scan the joint calibration board 110 with its lidar beams and return 3D point clouds. The calibration targets in turn may be contemporaneously detected by the camera 120 and the lidar 130, and the position of each center points P0, P1, P2 and P3 can be expected to be worked out therefrom, by each of the camera 120 and the lidar 130, respectively.

[0041] On one hand, in step 310, the at least one camera may capture 2D camera images of the calibration scene as described in FIG. 1, and the joint calibration board should be in each frame of the captured images.

[0042] FIG. 4A shows an example frame of 2D camera image of the joint calibration scene captured by the camera, wherein the single calibration board fabricated with the calibration pattern of FIG. 2C is contained. As can be seen from the image, the real calibration scene is in an underground garage with insufficient lighting. The calibration board is located in an area within the image, the nine corner points 410a-410i on the calibration board are expected to be detected by the camera. There also can be some noise comes from the backgrounds or generated during the  image processing. A large number of the environmental noises detected from the image may be due to, for example, insufficient light in the environment (such as the points 440) , possible light spots generated by lighting lamps (such as the points 450, 460) , and objects with contrasting color boundaries in the background (such as points 470, 472, 474, 476) . Therefore, a series of image processing shall be required in order to accurately detect the desired corner points from the calibration board.

[0043] In step S322, due to more or less distortion inherent in the camera lens, the raw image captured by the camera needs to be un-distorted, firstly.

[0044] In the method of the inventive subject matter, it is a prerequisite for the implementation of the provided joint calibration method that the factory calibrations for the camera and the lidar have been completed, respectively. Therefore, it is possible to obtain the un-distorted image based on the known camera intrinsic matrix Kraw and the camera distortion coefficients draw, the specific steps are as followings:

[0045] a. For the distorting original image of the camera, its coordinates may be represented as:

[0046] wherein the camera intrinsic matrix Kraw and the camera distortion coefficients draw are known by the camera calibration completed, as noted;

[0047] b. For a new introduced un-distorting image of the camera, its new coordinates can be:

[0048] where the new camera intrinsic coefficients “Knew” can be set for free as there is no any distortion, nor any new distortion coefficients d;

[0049] c. According to the resolution of the new image, query the pixel position of each of the pixels on the original image, the relationship can be described as:

[0050] Then, the distorted coordinate can be found by the know draw, and

[0051] Accordingly, the 2D un-distorted new image of the camera can be derived base on the above corresponding relationships as described above. The following described image processing steps shall be performed to detect the nine corner points on the basis of the new 2D un-distorted image.

[0052] In the next step S324, some image pre-processing used to the 2D un-distorted image may be performed to suppress the noise that comes from the environment. In one or more embodiments, the 2D un-distorted image of the underground garage scene may be firstly smoothed and then binarized into a black and white image by, specifically, selecting a suitable mediate pixel value, traversing all the pixels of the image and judging the pixels with the pixel value less than this selected mediate value as black, and the pixels with the pixel value greater than this selected mediate value as white. Updating all the pixel value for the black pixels to be such as biased to 0, and the pixel value for the white pixels to be such as biased to 255. Then, after further processing the image with the image erosion and the image expansion, the binarized image may results as shown in FIG. 4B.

[0053] In step S326, to further limit the range searching for the corner points, in one or more embodiments, a 2D filter with a created kernel may be employed for filtering the whole image. The kernel for the 2D filter, for example, can be used to weight the 2D images. For example, in the binarized image in FIG. 4B, the corner points expected to be detected are the intersections of the diagonal black and white rectangles with known dimensions on the custom-fabricated calibration board, so the kernel to be established is also of the diagonal type. By way of example, as an exemplary reference kernel corresponding to the tiled-shape area of the diagonal intersection of the two black rectangles at the top left and bottom right, and the two white at the bottom left and top right, as indicated by the dashed block 480 in FIG. 4B, the basic settings for this kernel can be created as the following format:

[0054] Where the weight is set to 1 for the white areas and -1 for the black areas, respectively, and A is the normalization coefficient, which renders the maximum response value of 2D filtering = 1. It can be appreciated that the shape weighted with such kernel can be as denoted with the dashed  block 480, as well as 482, 484, 486, 488, which are all set for the definition of the corner points on the joint calibration board. If the 2D filter configured with the kernel is used to filter the whole image, the response value near the corner points 410a, 410c, 410e, 410g and 410i that conforms to the kernel setting shall be relatively much larger. Only the areas with such black and white diagonal-intersect shape 480 can maximize the response value. Therefore, as shown in FIG 4C, there are five strong response spots 480a, 480c, 480e, 480g, and 480i appear around the positions of the center points 410a, 410c, 410e, 410g, and 410i, respectively.

[0055] It can be appreciated by those skilled in the art, the size of the kernel can be set and adjusted based on the size of the black and white rectangles. In one example, the size of the kernel can be set to the following format based on the size of the black and white rectangles in the calibration board image. In an ideal situation, setting all the white part = 255 and the black part = 0, for performing full image filtering with the 2D filter configured with the kernel in size of 8×8, the corner point position each gets the largest response is 255*16+255*16. That is, all black areas are 0, and once there is a value other than 0, the response intensity will be decrease.

[0056] Those skilled in the art can appreciate that the size of the kernel may further be related to the camera shooting distance and the camera intrinsic coefficients. For example, for the example of FIG. 4A, the size of the kernel is set of 16×16 in size, accounting in the overall normalization coefficient A, so that the maximum response value = 1. Similarly, with properly configuring the kernel, such as exchange the locations of element 1 and -1 in the matrix, the spots for other four corner points 410b, 410d, 410f, 410h may be located, respectively.

[0057] The above is the simplest example of setting the kernel, where all white areas are set to 1 and all black areas are set to –1. Those skilled in the art can imagine that the specific settings of the 2D filter kernel in image filtering need to be adjusted based on real scenes. Therefore, with adjusting the settings of the kernel and performing image filtering Iteratively, the approximate regions of all the nine corner points 410a through 410i can be determined in the image.

[0058] Next, In the step S328, Non-Maximum Compression (NMS) target detection may be applied to the calibrated board image data after the 2D image filtering, by which non-maximum points in each respective region of the nine spots may be removed, the unnecessary points are reduced, and the points adjacent to each other being too close or tight within each region may get sparser. Therefore, redundant points are further eliminated, and better candidates for the object detection are located. The processing in this step greatly reduces the computational  complexity of the subsequent steps, and accelerates the detection. However, after the above multi-step image processing process, there may be still many individual points in the respective regions covering around the nine corner points, respectively, in which still include some noise points.

[0059] Afterwards, in the step S330, in order to further accurately detected the nine desired center points, clustering may be simply performed on those individual points remained. According to geometric principles, it is known that such nine corner points shall meet the constraints, that is, each three points divided into a category can be linked by a straight line, forming a 3-point collinear tuple. Thus, complied with the clustering, the expected nine corner points can be finally determined with the 3-point collinear constraint, respectively.

[0060] FIG 4D shows the partially enlarged image of the joint calibration board in FIG. 4A. Based on the design of the joint calibration board, it can be assumed that a total of 8 straight lines can be formed by connecting each three (and only three) of the nine expected corner points, namely, the horizontal three lines (410a-410b-410c; 410d-410i-410h and 410e-410f-410g) , the vertical three lines (410a-410h-410g; 410b-410i-410f and 410c-410d-410e) , and the diagonal two lines (410a-410i-410e; 410c-410i-410g) , as shown by the light-colored lines, respectively. This process can find out the final corner points from each of the nine regions that satisfy the constrains by traversing to logically determine the distances and the relative angles among the individual points remained in the image.

[0061] In one or more embodiments, the clustering process comprises, firstly, calculating the distances between any two points of the remained individual points and the angle between the lines linking the any two points and the X-coordinate axis, then continuing to calculating the two lines from the set of lines from the previous step in two-by-two, to find the set of line for all three points on a straight line by determining that the absolute angular difference between the angles formed by the respective two straight line and the x-coordinate axis should be less than a small threshold Δθthreshold in angular degree. In the example as shown in FIG. 4D, taking any three points from the three regions of corner points 410a, 410i and 410e as an example to illustrate the algorithmic constraints of clustering, for the angle θ1 between the line 410a-410i and the x-coordinate axis, and the angle θ2 formed between the line 410i-410e and the x-coordinate axis:

[0062] when |θ1-θ2| < Δθthreshold                 (6)

[0063] the current three points 410a, 410i and 410e can be determined to be on a straight line, and the three points are determined to be a 3-point collinear tuple of 410a-410i-410e. In this way, the step for any two of these individual points may be repeated for traversing to obtain a set of candidate lines including all three points on a line. In one and more embodiments, the small threshold value Δθthreshold can be adjusted, preferably Δθthreshold≤ 5°. All points that cannot be categorized into any of the 3-point collinear tuples may be eliminated, thereby the candidate corner points are further reduced for the final clustering. It is noted that the nine corner points finally determined in the image must be complied with the 3-point collinear constraint, and each corner points may be in more than one lines.

[0064] Since most of the noise has been filtered out and all remaining points have been traversed. Using any of the remained candidate points as the point at the center location of the calibration target, i.e., the corner point 410i, whether or not there are 8 surrounding points that can satisfy the eight of 3-point collinear can be determined.

[0065] Due to the fact that the method of the inventive subject matter has undergone various optimizations as previously described and most of the noise points have been eliminated, the selection of the nine corner points can be running through a general-purpose processor, such as in the processor (s) of a laptop, in real-time, for example.

[0066] So far, the nine desired corner points in the calibration pattern has been detected. Each of the detected corner points in the image may then be subpixeled to improve its accuracy, that is, the integer pixel points of these nine discrete corner points can be converted into floating point numbers. Therefore, higher resolution is obtained, thereby improving the accuracy of the position of each point.

[0067] Finally, the coordinates of the four center points in the 2D camera image can be calculated by some known geometric algorithms. The position of ach of the four center points P0, P1, P2, P3, which are un-distorted, can be formulated as:

[0068] The above positioned four center points P0-P3 can be represented in the homogeneous coordinate as:

[0069] In the next step S332, projecting the four center points in the undistorted image to the  camera physical coordinate system to obtain the camera coordinates thereof. The newly set camera intrinsic matrix Knew for the un-distorting situation can be found by the relationship described in (3) as noted above, and thus the four center points in the homogeneous coordinate of camera physical coordinate can be derived as:

[0070] On the other hand, for the lidar, the aim is to detect the four center points that belong to the four circular holes, respectively. In particular, the lidar can scan with its lidar beams and return 3D point cloud data. The information returned is 3D point cloud representing the range measurements.

[0071] In step S340, in the joint calibration scene 100 in connecting with FIG. 1 and FIG. 4A, the lidar 130 scans and returns 3D point cloud. Firstly, the lidar receives the 3D point cloud returned and automatically invokes un-distortion processing to adjust the returned 3D point cloud to 3D lidar point cloud in step S342. Then, an image of the 3D lidar point cloud can be generated.

[0072] FIG. 5A shows an example frame of image generated from the 3D point cloud returned from the lidar. Because one single scanned frame of the returned lidar point cloud may be very sparse, in order to increase the density of the lidar point cloud, a number of frames of lidar point cloud may be collected for generating each frame of the image. Therefore, the frame of image as shown in FIG 5A can be generated from the scanning results of multiple lidar point clouds or from the scanning results of multiple selected lidar point clouds. In one or more embodiments, every ten frames of lidar point clouds may be collected for generating each single frame of the image.

[0073] Referring to FIG. 5A, the frame of image generated with ten frames of the 3D lidar point cloud 500, wherein the point cloud scanned by the lidar beam around the joint calibration board is selected to generate the frame of image which is similar to FIG. 4A in its FOV, for example, and wherein the edges of the four circular holes 520, 522, 524, 526 of the joint calibration board are the targets that are expected to be detected by the lidar. Then, a plane point selection may be performed, in step S344, by properly running the depth estimation, which applies predefined distances to filter the lidar point cloud, to select the desired points distinguishable with the background points. In the example, the points with distances much greater than the predefined  distances may be filtered, resulting in the desired points selected out from background points, and such desired points may be considered as formed from the echo reflected from the calibration board.

[0074] Depth discontinuity points detection in step S346 is performed to find out the potential points belong to the circular holes by the depth difference of adjacent points. In this step, each point in the plane of the calibration board is assigned a magnitude representing the depth difference with respect to their neighbors. For the lidar point pi with a given range measurement belongs to the plane of the calibration board, performing with that:

[0075] where for its scanned neighbor points are pi+1 and pi-1 . All the points from the plane of the calibration board can be filtered out with a discontinuity value pΔ<δdiscount, plane, resulting in the point cloud selected, as shown in FIG. 5B. The shown points in FIG. 5B are not belonging to the circular holes and can be found and removed out therefrom.

[0076] Next, in step S348, the four circle on the calibration board are intend to be segmented. The points not belonging to the circles have been removed by previously performed the filtering processes, and only discontinuities are present in the clouds at this stage. Accordingly, outliers to be removed are, mostly, from the outer boundaries of the calibration target. The lidar point cloud can be processed to keep only the rings formed with a number of points compatible with the presence of a circle and subsequently to remove the outer point in these rings. The filtered clouds are as shown in FIG. 5C.

[0077] Afterwards, in step S350, projecting the circle points to a plane, such as to the XY plane of the lidar coordinate system, and performing a 2D space circle fitting, resulting in four 2D circles representing the edges of the corresponding circular holes on the calibration board, respectively, as shown in FIG. 5D. At this point, based on the four fitted circles, and since the dimensions of the calibration board are pre-set and known, it is easy to determine the four circle centers by taking into account of the radius of the circular hole as a constraint through some geometric algorithms. Finally, transforming the positioned four center points back to the 3D space, such four center points will be used as the correspondence key points between the camera and the lidar.

[0078] Following the steps above, the four center points P0, P1, P2, P3 can have been detected  from each calibration cloud-point of lidar, and each of the four center points can be formulated as: And the homogeneous coordinate of the center points P0, P1, P2, P3 in the 3D lidar coordinate can be represented as:

[0079] In one and more embodiments, to improve the accuracy of the joint calibration, as an additional or alternative step S360, the single calibration board may be moved to several different places within the common FOV of the camera and lidar, and the processing steps S320-S350 may be repeated for the several time. For these repetitions, by recording the positioning results of the corresponding several sets of center points for the camera and the lidar, respectively, the joint extrinsic estimation can be performed for the several times, as will be described below.

[0080] So far, the four center points P0, P1, P2, P3 in the 2D camera coordinate system of  and in the 3D lidar coordinate system with the 3D coordinate of both are represented in the homogeneous coordinates, correspondingly. Therefore, finding the transformation relationship between the 2D camera coordinate system and the 3D lidar coordinate system translates into solving the transformation between 2D and 3D points, which can be formulated as a least square problem that is also focused by, for example, PnP algorithms.

[0081] Accordingly, in the final step S370, it is possible to perform the joint extrinsic estimation between the lidar coordinate system and the camera coordinate system. The extrinsic transformation, included the extrinsic (the rotation matrix R and the translation vector t) between the camera and the lidar, can be solved, respectively.

[0082] In one and more embodiments, the PnP algorithm can be, such as, the EPnP and / or the Levenberg-Marquardt optimization. In the example using the single calibration as shown in FIG. 2, n = 4 for the four center points, P0, P1, P2, P3, to be detected. In the example using the single calibration as shown in FIG. 3, n = 8 for the eight center points, P0, P1, …, P7, to be detected.

[0083] The algorithms provided can efficiently and accurately detect the calibration target from the camera and lidar sensors. Consequently, based on the detected calibration target, extrinsic of camera and lidar can be jointly estimated. Due to the novel design of calibration board and algorithms, the calibration setups are simple, and efficiency, meanwhile, the calibration ability  are robust, and accuracy.

[0084] Any combination of one or more computer readable medium (s) may be utilized to perform the method in the present disclosure. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0085] The inventive subject matter comprises, but not limited to, the items listed hereinafter.

[0086] Item 1: a method for at least one camera and at least one lidar joint extrinsic calibration based on a single calibration board, comprising following steps of:

[0087] positioning first plurality of center points from 2D image comprising the single calibration board captured by the at least one camera;

[0088] positioning second plurality of center points from 3D point cloud containing points reflected from the single calibration board returned by the at least one lidar; and

[0089] estimating an extrinsic transformation between the at least one camera and the at least one lidar,

[0090] wherein the first plurality of center points belong to rectangular grids on the single calibration board, respectively;

[0091] wherein the second plurality of center points belong to circular holes in the single calibration board, respectively;

[0092] wherein the first plurality of center points and the second plurality of center points are completely coincident in position.

[0093] Item 2: the method of item 1, wherein the four vertices of each of the rectangular grids are  arranged as corner points that can be detected by the at least one camera.

[0094] Item 3: the method of any of items 1or 2, wherein the rectangular grids, in which the corresponding circular holes are coincidentally arranged, are disposed in M×N grids, M≥2 and N≥2 in integer, on the single calibration board.

[0095] Item 4: the method of any of items 1-3, wherein detecting the corner points by the at least one camera comprises filtering the 2D image by using a 2D filter with a kernel created based on the calibration board.

[0096] Item 5: the method of any of items 1-4, wherein the corner points are each defined by the diagonal intersection of two black rectangles, and the kernel is created based on the shape of the two black rectangles intersecting diagonally.

[0097] Item 6: the method of any of items 1-5, wherein detecting the corner points by the at least one camera comprises accurately detecting the corner points by determining that each three of the corner points are on one straight line.

[0098] Item 7: the method of any of items 1-6, the second plurality of center points can be positioned on a 2D space in which 2D circles are fit from points of edges of the circular holes on the single calibration board, respectively.

[0099] Item 8: the method of any of items 1-7, wherein the edges of the circular holes can be detected from the single calibration board, respectively, using a depth discontinuity points detection.

[0100] Item 9: the method of any of items 1-8, wherein the 2D image comprising the single calibration board captured by the at least one camera is applied with un-distorting processing before positioning the first plurality of center points therefrom, wherein the un-distorting processing is based on an intrinsic matrix and intrinsic distortion coefficients.

[0101] Item 10: the method of any of items 1-9, wherein estimating the extrinsic transformation of the at least one camera and the at least one lidar comprises working out a rotation matrix and a translation vector, respectively, between the at least one camera and the at least one lidar.

[0102] Item 11. a non-transitory computer readable medium storing instruction that can be processed by one or more processor to realize following steps. comprising:

[0103] positioning first plurality of center points from 2D image comprising the single calibration board captured by the at least one camera;

[0104] positioning second plurality of center points from 3D point cloud containing points reflected  from the single calibration board returned by the at least one lidar; and

[0105] estimating an extrinsic transformation between the at least one camera and the at least one lidar,

[0106] wherein the first plurality of center points belong to rectangular grids on the single calibration board, respectively;

[0107] wherein the second plurality of center points belong to circular holes in the single calibration board, respectively;

[0108] wherein the first plurality of center points and the second plurality of center points are completely coincident in position.

[0109] Item 12. The non-transitory computer readable medium according to item 11, wherein the four vertices of each of the rectangular grids are arranged as corner points that can be detected by the at least one camera.

[0110] Item 13. The non-transitory computer readable medium according to item 11 or 12, wherein the rectangular grids, in which the corresponding circular holes are coincidentally arranged, are disposed in M×N grids, M≥2 and N≥2 in integer, on the single calibration board.

[0111] Item 14. The non-transitory computer readable medium according to any of items 11-13, wherein detecting the corner points by the at least one camera comprises filtering the 2D image by using a 2D filter with a kernel created based on the calibration board.

[0112] Item 15. The non-transitory computer readable medium according to any of items 11-14, wherein the corner points are each defined by the diagonal intersection of two black rectangles, and the kernel is created based on the shape of the two black rectangles intersecting diagonally.

[0113] Item 16. The non-transitory computer readable medium according to any of items 11-15, wherein detecting the corner points by the at least one camera comprises accurately detecting the corner points by determining that each three of the corner points are on one straight line.

[0114] Item 17. The non-transitory computer readable medium according to any of items 11-16, wherein the second plurality of center points can be positioned on a 2D space in which 2D circles are fit from points of edges of the circular holes on the single calibration board, respectively.

[0115] Item 18. The non-transitory computer readable medium according to any of items 11-17, wherein the edges of the circular holes can be detected from the single calibration board, respectively, using a depth discontinuity points detection.

[0116] Item 19. The non-transitory computer readable medium according to any of items 11-18,  wherein the 2D image comprising the single calibration board captured by the at least one camera is applied with un-distorting processing before positioning the first plurality of center points therefrom, wherein the un-distorting processing is based on an intrinsic matrix and intrinsic distortion coefficients.

[0117] Item 20. The non-transitory computer readable medium according to any of items 11-19, wherein estimating the extrinsic transformation of the at least one camera and the at least one lidar comprises working out a rotation matrix and a translation vector, respectively, between the at least one camera and the at least one lidar.

[0118] As used in this application, an element or step recited in the singular and proceeded with the word “a” or “an” should be understood as not excluding plural of said elements or steps, unless such exclusion is stated. Furthermore, references to “one embodiment” or “one example” of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features. The terms “first, ” “second, ” and “third, ” etc. are used merely as labels, and are not intended to impose numerical requirements or a particular positional order on their objects.

[0119] While exemplary embodiments are described above, it is not intended that these embodiments describe all possible forms of the present disclosure. Rather, the words used in the specification are words of description rather than limitation, and it is understood that various changes may be made without departing from the spirit and scope of the present disclosure. Additionally, the features of various implementing embodiments may be combined to form further embodiments of the inventive subject matter.

Claims

1.A method for at least one camera and at least one lidar joint extrinsic calibration based on a single calibration board, comprising following steps of:positioning first plurality of center points from 2D image comprising the single calibration board captured by the at least one camera;positioning second plurality of center points from 3D point cloud containing points reflected from the single calibration board returned by the at least one lidar; andestimating an extrinsic transformation between the at least one camera and the at least one lidar,wherein the first plurality of center points belong to rectangular grids on the single calibration board, respectively;wherein the second plurality of center points belong to circular holes in the single calibration board, respectively;wherein the first plurality of center points and the second plurality of center points are completely coincident in position.2.The method according to claim 1, wherein the four vertices of each of the rectangular grids are arranged as corner points that can be detected by the at least one camera.3.The method according to claim 1, whereinthe rectangular grids, in which the corresponding circular holes are coincidentally arranged, are disposed in M×N grids, M≥2 and N≥2 in integer, on the single calibration board.4.The method according to claim 2, whereindetecting the corner points by the at least one camera comprises filtering the 2D image by using a 2D filter with a kernel created based on the calibration board.5.The method according to claim 4, whereinthe corner points are each defined by the diagonal intersection of two black rectangles, and the kernel is created based on the shape of the two black rectangles intersecting diagonally.6.The method according to claim 2, whereindetecting the corner points by the at least one camera comprises accurately detecting the corner points by determining that each three of the corner points are on one straight line.7.The method according to claim 2, wherein the second plurality of center points can be positioned on a 2D space in which 2D circles are fit from points of edges of the circular holes on the single calibration board, respectively.8.The method according to claim 7, wherein the edges of the circular holes can be detected from the single calibration board, respectively, using a depth discontinuity points detection.9.The method according to claim 1, whereinthe 2D image comprising the single calibration board captured by the at least one camera is applied with un-distorting processing before positioning the first plurality of center points therefrom, wherein the un-distorting processing is based on an intrinsic matrix and intrinsic distortion coefficients.10.The method according to claim 1, wherein estimating the extrinsic transformation of the at least one camera and the at least one lidar comprises working out a rotation matrix and a translation vector, respectively, between the at least one camera and the at least one lidar.11.A non-transitory computer readable medium storing instruction that can be processed by one or more processor to realize following steps. comprising:positioning first plurality of center points from 2D image comprising the single calibration board captured by the at least one camera;positioning second plurality of center points from 3D point cloud containing points reflected from the single calibration board returned by the at least one lidar; andestimating an extrinsic transformation between the at least one camera and the at least one lidar,wherein the first plurality of center points belong to rectangular grids on the single calibration board, respectively;wherein the second plurality of center points belong to circular holes in the single calibration board, respectively;wherein the first plurality of center points and the second plurality of center points are completely coincident in position.12.The non-transitory computer readable medium according to claim 11, wherein the four vertices of each of the rectangular grids are arranged as corner points that can be detected by the at least one camera.13.The non-transitory computer readable medium according to claim 11, whereinthe rectangular grids, in which the corresponding circular holes are coincidentally arranged, are disposed in M×N grids, M≥2 and N≥2 in integer, on the single calibration board.14.The non-transitory computer readable medium according to claim 12, whereindetecting the corner points by the at least one camera comprises filtering the 2D image by using a 2D filter with a kernel created based on the calibration board.15.The non-transitory computer readable medium according to claim 14, whereinthe corner points are each defined by the diagonal intersection of two black rectangles, and the kernel is created based on the shape of the two black rectangles intersecting diagonally.16.The non-transitory computer readable medium according to claim 12, whereindetecting the corner points by the at least one camera comprises accurately detecting the corner points by determining that each three of the corner points are on one straight line.17.The non-transitory computer readable medium according to claim 12, wherein the second plurality of center points can be positioned on a 2D space in which 2D circles are fit from points of edges of the circular holes on the single calibration board, respectively.18.The non-transitory computer readable medium according to claim 17, wherein the edges of the circular holes can be detected from the single calibration board, respectively, using a depth discontinuity points detection.19.The non-transitory computer readable medium according to claim 11, whereinthe 2D image comprising the single calibration board captured by the at least one camera is applied with un-distorting processing before positioning the first plurality of center points therefrom, wherein the un-distorting processing is based on an intrinsic matrix and intrinsic distortion coefficients.20.The non-transitory computer readable medium according to claim 19, wherein estimating the extrinsic transformation of the at least one camera and the at least one lidar comprises working out a rotation matrix and a translation vector, respectively, between the at least one camera and the at least one lidar.