Camera and laser radar combined external parameter calibration method based on single calibration board
By designing a calibration pattern that combines rectangular grids and concentric circular holes, and using a camera and LiDAR to detect feature points respectively, the problems of complexity and lack of generalization ability in existing technologies are solved, achieving efficient and accurate sensor data fusion and improving the robustness and accuracy of calibration.
Patent Information
- Application Number
- CN202380095779.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-07
- Publication Date
- 2025-10-17
Smart Images

Figure CN120813974A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present subject matter relates generally to driver assistance and image / video processing. More specifically, the present subject matter relates to a method for camera and lidar joint extrinsic calibration based on a single calibration board. BACKGROUND
[0002] Advanced Driver Assistance Systems (ADAS) and Autonomous Driving (AD) are systems designed to automate / regulate / enhance the safe and convenient driving of vehicles. ADAS / AD can monitor the internal / external environment of a vehicle through various sensors such as cameras, lidars, radars, and ultrasonic waves. The collected sensor data can be processed by an on-board processor to extract useful information and take appropriate actions. These actions can be as simple as turning on / off warning lights or as complex as controlling the throttle, brakes, and steering of the vehicle.
[0003] For various ADAS / AD systems, in order to take appropriate actions on the vehicle, accurate perception of the surrounding environment must be extracted, which can include: 1) static objects / scenes, such as buildings, roads, lanes, and traffic signs; 2) dynamic objects, such as vehicles and pedestrians. The identification task includes various information about the object, such as the category, attributes, location, distance, pose, and speed of the object. Among all tasks, for 3D related sensing tasks (location, distance, etc.), sensor extrinsic calibration can be one of the most important issues.
[0004] With the development of ADAS / AD systems, due to the requirements for accuracy and robustness, the design of these systems faces a topology with multiple sensors for realizing multi-modal perception and fusion. In order to represent and fuse the perception results in a common reference frame (i.e., ego-vehicle coordinate system), the rigid transformation (3D rotation and translation) between all sensors and the ego-vehicle should be known. Therefore, such fusion raises a further calibration problem, i.e., joint calibration of multiple sensors.
[0005] Due to the ability to provide rich appearance information, camera sensors have been widely used in various ADAS / AD systems. On the other hand, lidar sensors have high ranging accuracy and are therefore widely used. Since these two sensors can provide complementary information, they are suitable for application in the same ADAS / AD system. In such a design, in order to effectively combine the data / results of the two sensors to realize multi-modal perception and fusion, the relative extrinsic parameters between the camera and the lidar should first be determined jointly. However, existing joint calibration methods have several problems: 1) complex setup; 2) lack of generalization ability. Therefore, the accuracy of the calibration depends largely on the carefully designed setup and structured environment.
[0006] There is a need to provide a method for camera and lidar joint extrinsic calibration, which can apply an algorithm to jointly estimate the extrinsic parameters of both camera and lidar sensors through a single calibration board. SUMMARY
[0007] The method provided in the inventive subject matter can be used for joint extrinsic calibration between at least one camera and at least one lidar. The method can include locating a first plurality of center points from a 2D image captured by the at least one camera including the single calibration board, and locating a second plurality of center points from a 3D point cloud returned by the at least one lidar containing points reflected from the single calibration board. The method can further include estimating an extrinsic transformation between the at least one camera and the at least one lidar. In the method, the first plurality of center points respectively belong to a rectangular grid on the single calibration board, while the second plurality of center points respectively belong to circular holes in the single calibration board. It is crucial that the first plurality of center points and the second plurality of center points are perfectly coincident in position.
[0008] A non-transitory computer-readable medium stores instructions processable by one or more processors to locate a first plurality of center points from a 2D image captured by the at least one camera including the single calibration board, and locate a second plurality of center points from a 3D point cloud returned by the at least one lidar containing points reflected from the single calibration board. The non-transitory computer-readable medium stores instructions processable by one or more processors to further estimate an extrinsic transformation between the at least one camera and the at least one lidar. The first plurality of center points respectively belong to a rectangular grid on the single calibration board, while the second plurality of center points respectively belong to circular holes in the single calibration board. It is crucial that the first plurality of center points and the second plurality of center points are perfectly coincident in position. BRIEF DESCRIPTION OF DRAWINGS
[0009] The inventive subject matter can be better understood with reference to the following description of non-limiting embodiments and with reference to the accompanying drawings, in which like reference numerals refer to corresponding parts in the several views. In the drawings:
[0010] Figure 1 An example diagram showing a scenario of single calibration board based joint extrinsic calibration of camera and lidar according to one or more embodiments of the inventive subject matter is shown;
[0011] Figure 2A An example diagram showing a front view of a joint calibration pattern according to one or more embodiments of the inventive subject matter, based on which a method for camera and lidar joint extrinsic calibration can be performed, is shown;
[0012] Figure 2B And Figure 2CEach shows one or more embodiments according to the present subject matter Figure 2A Alternative examples of;
[0013] Figure 2D One or more embodiments of the present invention are shown. Figure 2C An extended example of ;
[0014] Figure 3 An example flow chart illustrating a method for joint extrinsic calibration of a camera and a lidar according to one or more embodiments of the present subject matter is shown;
[0015] Figures 4A-4D An example process for locating four center points from a 2D camera image of a single calibration plate captured by a camera is shown, including Figure 2C Single calibration target; and
[0016] Figures 5A-5D An example process for locating four center points from a 3D point cloud scanned and returned by a LiDAR is shown, according to one or more embodiments of the present subject matter. Figure 2C A single calibration target, with the X-axis and Y-axis shown as references. DETAILED DESCRIPTION
[0017] The following discloses a detailed description of one or more embodiments of the subject matter of the present invention; however, it should be understood that the disclosed embodiments are merely exemplary of the subject matter of the present invention, which may be embodied in various and alternative forms. The drawings are not necessarily drawn to scale; some features may be exaggerated or minimized to illustrate details of particular components. Therefore, the specific structural and functional details disclosed herein should not be interpreted as limiting, but merely represent a basis for teaching one skilled in the art to variously apply the subject matter of the present invention.
[0018] Among numerous sensor data fusion methods, the combination of at least one lidar and at least one camera is one of the most commonly used sensor pairs for driving environment perception. Lidar can provide 3D point cloud data, including accurate depth and reflection intensity information, while cameras capture rich semantic information about the scene. The combination of at least one camera and at least one lidar offers the potential to overcome the limitations of each sensor. The primary challenge in fusing these two heterogeneous sensors is finding a rigid-body transformation between the sensor coordinate systems by performing a joint extrinsic calibration.
[0019] In one aspect of the present subject matter, a method for joint extrinsic calibration of at least one camera and at least one lidar based on a single calibration plate is provided.
[0020] Figure 1 An example diagram showing a scene 100 illustrating joint extrinsic calibration of camera 120 and lidar 130 based on single calibration board 110 according to one or more embodiments of the inventive subject matter is shown. As shown in scene 100, Figure 1 As shown in scene 100, camera 120 and lidar 130 are fixed to each other and both mounted on the roof of an autonomous driving vehicle 140, while letting the targets on the board fully appear within the overlapping region of the field of view (FOV) of both camera 120 and lidar 130, and then enabling camera 120 and lidar 130 to detect their respective calibration features from the board simultaneously.
[0021] In target-based extrinsic calibration methods, the setup of the calibration target is particularly important and critical. For both camera sensors and lidar sensors, features whose edges and corners can be accurately and robustly detected. For example, for camera sensors, a checkerboard with a rectangular grid can be the most straightforward calibration board design, and many existing algorithms can be conveniently applied to detect these calibration patterns with high accuracy. However, for lidar sensors, due to the very sparse data, the nearly horizontal edges of the rectangular checkerboard can not intersect with any of the lidar scanning beams. Therefore, it is difficult to detect objects of rectangular shape from lidar data. Moreover, the checkerboard of rectangular patterns can only provide limited information.
[0022] On the other hand, a board with circular holes can be a good design for lidar sensors (also for stereo cameras), as these holes can be accurately detected when intersected with a small number of lidar beams. However, it is much more difficult for monochrome cameras to accurately detect the calibration patterns due to the lack of unique semantic appearance.
[0023] In this regard, a novel designed calibration pattern is introduced herein, which has detection targets that can be detected by both at least one camera and at least one lidar to be calibrated, and all the sensors to be calibrated can respectively detect their corresponding different features from the calibration pattern.
[0024] To present a calibration pattern that can be detected in all relevant sensors of the at least one camera and the at least one lidar to be jointly calibrated, as well as observable features that can be extracted for localization, the calibration pattern can be designed to have detection targets including a rectangular grid with respective circular holes arranged concentrically inside. All vertices of the rectangular grid have been arranged to be one corner point each, which can be generally defined as a point of intersection of black (or white) rectangles, that is, a common vertex of opposite corners of two black rectangles diagonally adjacent to each other, for example. By detecting such calibration targets, the camera can determine the corresponding rectangular grid by detecting every four corner points of the corresponding rectangular grid, respectively, while the lidar can obtain the corresponding edges of the circular holes. Thereby, the at least one camera and the at least one lidar to be jointly calibrated can further calculate their respective center points from the rectangular grid and the circular holes in each of the fully coinciding positions, respectively.
[0025] Figure 2A An example diagram showing a front view of a calibration pattern 200 according to one or more embodiments of the inventive subject matter is shown, based on which a method for camera-lidar joint extrinsic calibration can be performed.
[0026] As Figure 2A shown, the calibration pattern 200 includes planar targets of four rectangular grids 210, 212, 214, 216 and four corresponding circular holes 220, 222, 224, 226 arranged concentrically. The four rectangular grids 210, 212, 214, 216 can have the same size and shape and be symmetrically arranged in 2 2 (2 rows and 2 columns) grids in a two-dimensional plane, inside which four circular holes 220, 222, 224, 226 can be correspondingly arranged. It is easy to make the side length of each rectangular grid at least greater than the diameter of the corresponding circular hole inside it, so as to arrange a corner point 230a-230i at each of the vertices of the four intersecting rectangular grids 210, 212, 214, 216. The corner point can be arranged by two black rectangles intersecting at opposite corners each, such as shown in Figure 2A As shown, two black rectangles intersecting at opposite corners are located at the upper left corner and the lower right corner, while two white rectangles are located at the lower left corner and the upper right corner, or black and white are reversed. In the joint calibration method herein, it is required to arrange each rectangular grid 210, 212, 214, 216 and the circular hole 220, 222, 224, 226 therein to be fully coincided in the positions of the center points P0, P1, P2, or P3, respectively.
[0027] According to Figure 2AIn the illustrated example, there can be nine corner points 230a-230i in the calibration pattern 200 that are expected to be detected by the camera, which in turn can determine four rectangular grids 210, 212, 214, 216, and thus can calculate the positions of four center points P0, P1, P2, and P3. As for the lidar, the edges of the four circular holes 220, 222, 224, 226 can be detected to determine the positions of their four center points P0, P1, P2, and P3, respectively, which are in position exactly coinciding with those of the rectangles.
[0028] Figure 2B An alternative example of the calibration pattern 200 according to one or more embodiments of the inventive subject matter is shown in FIG. 2B, in which the tiling black-and-white rectangles used to define four of the nine corner points 230b', 230d', 230f', 230h' are colored differently from those used to define the other five corner points 230a, 230c, 230e, 230g, 230i. Figure 2A In this example, the tiling black-and-white rectangles used to define four of the nine corner points 230b', 230d', 230f', 230h' are colored differently from those used to define the other five corner points 230a, 230c, 230e, 230g, 230i. Figure 2A In this example, the tiling black-and-white rectangles used to define four of the nine corner points 230b', 230d', 230f', 230h' are colored differently from those used to define the other five corner points 230a, 230c, 230e, 230g, 230i.
[0029] In one or more embodiments, as shown in FIG. 2C, the four rectangular grids 210, 212, 214, 216 are represented in a 2 Figure 2A 2 grid representation, and all are arranged in light color (e.g., in white). In this way, it can be seen that a large portion of the calibration pattern 200 is in light color. Since the camera requires a good lighting environment, such a large portion of the calibration pattern in light color can be more easily segmented in a light deficient scene, such as in a basement garage. Alternatively, it is easier to detect the relatively brighter area and perform a coarse localization on the calibration board area, so that all subsequent corner detections can be performed near this coarse localization area. Figure 2B In one or more embodiments, some of the outer edges can also be arranged with colors opposite to the 2x2 rectangular grids, which can make the grids easier to be detected by the camera. In the example shown in FIG. 2D, two of the four rectangular grids are each arranged with an elongated border 240, 242 in opposite color, thereby forming a half-wrapped structure of the four rectangular grids 210, 212, 214, 216. The corner areas 244, 246 near the black wrap in light background color can be more easily detected. If these two corners can be successfully detected, the area between the top points of these two corners (i.e., the two corner points 230c, 230g) can be detected as the approximate location of the calibration board, which brings advantages to the camera to be jointly calibrated herein.
[0030] In one or more embodiments, some of the outer edges can also be arranged with colors opposite to the 2x2 rectangular grids, which can make the grids easier to be detected by the camera. In the example shown in FIG. 2D, two of the four rectangular grids are each arranged with an elongated border 240, 242 in opposite color, thereby forming a half-wrapped structure of the four rectangular grids 210, 212, 214, 216. The corner areas 244, 246 near the black wrap in light background color can be more easily detected. If these two corners can be successfully detected, the area between the top points of these two corners (i.e., the two corner points 230c, 230g) can be detected as the approximate location of the calibration board, which brings advantages to the camera to be jointly calibrated herein. Figure 2C
[0031] In one or more alternative embodiments, a square grid pattern, rather than a rectangular grid, may be the best choice outside each corresponding circular hole, not only because it is easier to detect, but also because it can simplify the algorithm and save computation on a single calibration plate. That is, the combination of concentric circles and squares can occupy the smallest area. Compared to other conventional calibration targets used for multi-sensor joint calibration, for calibration plates composed of the same size, the calibration pattern of the subject matter of the present invention can multiply the information that can be collected per frame (such as the number of detection points), thereby correspondingly improving calibration accuracy. Therefore, the joint calibration pattern with a rectangular grid arrangement can preferably be square.
[0032] In one or more embodiments, in order for all relevant sensors of at least one camera and at least one lidar to detect the graphical features set in the joint calibration pattern, the detection target should be processed and manufactured on a single calibration plate, for example, with predetermined dimensions. The single calibration plate can then be as compact as possible, easy to detect, and easy to perform calibration processes.
[0033] The single calibration plate provided in the subject matter of the present invention can provide enough information in a very compact arrangement and can have a variable and scalable implementation when necessary. The number of combinations of rectangular grids and corresponding concentric circular holes arranged therein can be extended to such arrangements of M×N, where M≥2 and N≥2 are integers. In this way, the amount of information obtained in each detection can be further increased to improve the accuracy or speed of calibration. In one or more embodiments of the calibration target according to the subject matter of the present invention, the combination arrangement of the rectangle plus the circular holes therein can be, for example, 2×3, 3×3, 2×4, 3×4, ..., or M×N, where M≥2 and N≥2 are integers respectively. As long as the requirements for clarity and richness of information extracted by at least one camera to be calibrated and at least one camera for its joint calibration target detection can be met. The more combinations, the more detection information, and therefore a more accurate calibration is achieved.
[0034] Figure 2D One or more embodiments of the present invention are shown. Figure 2C An example of an extension of . Figure 2D As shown in , the arrangement of rectangular grids and concentric circular holes is expanded to 4×2, that is, the joint calibration pattern contains eight rectangular grids and corresponding circular holes. Figure 2C Compared with the joint calibration pattern in , there may be eight center points P0-P7 expected to be located by the camera and lidar to be jointly calibrated, respectively, thereby doubling the amount of information used for joint calibration.
[0035] It can be noted that the novel joint calibration pattern provided in the present subject matter has significant technical advantages, being compact and having different calibration targets and / or features for both at least one camera and at least one lidar. As robust calibration targets, the rectangular grids and the circular holes can be accurately detected from the lidar and camera sensors, respectively. At the same time, the proposed joint calibration pattern can be easily manufactured in a compact board, making it easy to apply in most cases. Moreover, due to the compact design, the joint calibration board can be easily extended by adding more rectangular grids and circular holes to provide further accuracy and robustness.
[0036] Figure 3 An example flowchart 300 showing a method for joint extrinsic calibration of at least one camera and at least one lidar based on a single calibration board according to one or more embodiments of the present subject matter is shown. The method comprises steps S310-S370, which will be described as follows. In the proposed one or more embodiments, these steps will be described in connection with the scenario shown in Figure 1 and the calibration pattern shown in Figure 2C However, the skilled person can understand that other scenarios or calibration patterns are also applicable.
[0037] In step S310, the designed calibration target instances can be instantiated and manufactured on a single calibration board with known pre-set dimensions. The manufactured single calibration board can then be placed to ensure that it appears in a location of a common field of view (FOV) of at least one camera and at least one lidar to be jointly calibrated.
[0038] As shown in Figure 1 A single joint calibration board 110 containing the calibration pattern of Figure 2C is suggested to be placed horizontally in front of a camera 120 and a lidar 130 to be jointly calibrated, and is generally oriented to face both related sensors. The calibration board should appear in the common FOV of the camera 120 and the lidar 130. The camera 120 can then be able to capture 2D image frames containing the single joint calibration board 110, and the lidar 130 can be able to continuously scan the joint calibration board 110 using its lidar beams and return a 3D point cloud. The calibration targets can in turn be detected by the camera 120 and the lidar 130 simultaneously, and the position of each center point P0, P1, P2, and P3 can be expected to be computed from it by each of the camera 120 and the lidar 130, respectively.
[0039] In one aspect, in step 310, at least one camera can capture 2D camera images of the calibration scene as described in Figure 1 and the joint calibration board should be located in each frame of the captured images.
[0040] Figure 4AAn example frame of a 2D camera image of a joint calibration scene captured by a camera is shown, containing a single calibration board made with Figure 2C calibration pattern. From the image, it can be seen that the real calibration scene is in an underground garage with insufficient light. The calibration board is located within the area inside the image, and nine corner points 410a-410i on the calibration board are expected to be detected by the camera. There can also be some noise from the background or generated during image processing. The large amount of environmental noise detected from the image can be due to, for example, insufficient light in the environment (such as point 440), possible light points generated by the lighting lamps (such as points 450, 460), and objects in the background with contrasting color boundaries (such as points 470, 472, 474, 476). Therefore, a series of image processing should be required in order to accurately detect the expected corner points from the calibration board.
[0041] In step S322, due to the more or less distortion inherent to the camera lens, it is necessary to first de-distort the original image captured by the camera.
[0042] In the method of the present subject matter, the premise for implementing the provided joint calibration method is that the camera and the lidar have respectively completed factory calibration. Therefore, based on the known camera intrinsic matrix K 原始 and camera distortion coefficients d 原始 , a de-distorted image can be obtained, with the specific steps as follows:
[0043] a. For the distorted original image of the camera, its coordinates can be represented as:
[0044] (1)
[0045] where the camera intrinsic matrix K 原始 and the camera distortion coefficients d 原始 are known through the completed camera calibration, as described;
[0046] b. For the newly introduced de-distorted image of the camera, its new coordinates can be:
[0047] (2)
[0048] where the new camera intrinsic coefficient “K 新 ” can be freely set, as there is no distortion and there is no new distortion coefficient d;
[0049] c. According to the resolution of the new image, the pixel position of each of the pixel points on the original image is queried, and the relationship can be described as:
[0050] (3)
[0051] Then, the known d原始 finding distorted coordinates and
[0052] (4)
[0053] Thus, a new 2D de-distorted image of the camera can be derived based on the above correspondence as described above. The following image processing steps should be performed to detect the nine corner points on the basis of the new 2D de-distorted image.
[0054] In the next step S324, some image pre-processing for the 2D de-distorted image can be performed to suppress noise from the environment. In one or more embodiments, the 2D de-distorted image of the underground garage scene can be first smoothed, and then binarized into a black and white image, in particular by selecting a suitable middle pixel value, traversing all pixels of the image, and deciding pixels with pixel values smaller than the selected middle value as black and pixels with pixel values larger than the selected middle value as white. All pixel values of the black pixels are updated to, e.g., biased to 0, and all pixel values of the white pixels are updated to, e.g., biased to 255. Then, after further processing the image with image erosion and image dilation, a binarized image can be generated as shown in Figure 4B .
[0055] In step S326, to further limit the range search for the corner points, in one or more embodiments, a 2D filter with the created kernel can be employed to filter the whole image. For example, the kernel of the 2D filter can be used to weight the 2D image. For example, in the binarized image in Figure 4B , the expected detected corner points are intersections of diagonally black and white rectangles of known sizes on the customized calibration board, so the kernel to be established is also of diagonal type. By way of example, as an exemplary reference kernel corresponding to a tiled shape area intersecting diagonally with two black rectangles of the top-left corner and the bottom-right corner and intersecting diagonally with two white rectangles of the bottom-left corner and the top-right corner, as indicated by the dashed box 480 in Figure 4B , the basic setting of this kernel can be created as the following format:
[0056] 2D filter kernel = A · (5)
[0057] wherein respectively, the weight of the white region is set to 1, while the weight of the black region is set to -1, and A is a normalization coefficient, which makes the maximum response value of the 2D filter = 1. It can be understood that the shape of the kernel weighting can be represented by the dashed boxes 480 and 482, 484, 486, 488, which are all set to define the corner points on the joint calibration board. If the entire image is filtered using the 2D filter configured with the kernel, the response value near the corner points 410a, 410c, 410e, 410g and 410i that meets the kernel setting should be much larger. Only the region with the shape 480 of the black and white diagonal intersection can maximize the response value. Therefore, as shown in FIG. 4B, five strong response points 480a, 480c, 480e, 480g and 480i appear around the positions of the center points 410a, 410c, 410e, 410g and 410i, respectively. Figure 4C
[0058] It can be understood by those skilled in the art that the size of the kernel can be set and adjusted based on the size of the black and white rectangle. In one example, the size of the kernel can be set to the following format based on the size of the black and white rectangle in the calibration board image. In an ideal case, let all white parts = 255, and black parts = 0, to perform the entire image filtering with the 2D filter configured with the kernel size of 8 8, the corner point position gets the maximum response of 255*16+255*16 each. That is, all black regions are 0, and once there is a value other than 0, the response intensity will decrease.
[0059] It can be understood by those skilled in the art that the size of the kernel can be further related to the camera shooting distance and the camera intrinsic coefficient. For example, for the example of Figure 4A , the size of the kernel is set to 16 16, considering the overall normalization coefficient A, so that the maximum response value = 1. Similarly, by appropriately configuring the kernel, such as swapping the positions of elements 1 and -1 in the matrix, the spots of the other four corner points 410b, 410d, 410f, 410h can be located respectively.
[0060] The above is the simplest example of setting the kernel, in which all white regions are set to 1, and all black regions are set to -1. Those skilled in the art can imagine that the specific setting of the 2D filter kernel in image filtering needs to be adjusted based on the real scene. Therefore, by iteratively adjusting the setting of the kernel and performing image filtering, the approximate regions of all nine corner points 410a to 410i in the image can be determined.
[0061] Next, in step S328, a non-maximum suppression (NMS) object detection can be applied to the calibration board image data after 2D image filtering, whereby non-maximum points in each respective region of the nine bright spots can be removed, reducing unnecessary points, and points that are too close or too tight to each other within each region can be made sparser. Thus, further redundant points are eliminated, and better candidates for object detection are located. The processing in this step greatly reduces the computational complexity of the subsequent steps, and speeds up the detection. However, after the above multi-step image processing process, there can still be many individual points in the respective regions surrounding the nine corner points, including some noise points.
[0062] Subsequently, in step S330, to further accurately detect the nine desired center points, a clustering can be simply performed on those remaining individual points. According to geometric principles, it is known that the nine corner points should satisfy constraints, that is, each three points grouped into a class can be connected by a straight line, forming a 3-point collinear tuple. Thus, according to clustering, the expected nine corner points can finally be determined by 3-point collinear constraints, respectively.
[0063] Figure 4D A local zoomed-in image of the joint calibration board in Figure 4A is shown. Based on the design of the joint calibration board, it can be assumed that a total of 8 straight lines can be formed by connecting each three (and only three) of the nine expected corner points, that is, three horizontal lines (410a-410b-410c; 410d-410i-410h and 410e-410f-410g), three vertical lines (410a-410h-410g; 410b-410i-410f and 410c-410d-410e), and two diagonal lines (410a-410i-410e; 410c-410i-410g), respectively as shown by the light-colored lines. This process can logically determine the distance and relative angle between the remaining individual points in the image by traversing from each of the nine regions that satisfy the constraints to find the final corner points.
[0064] In one or more embodiments, the clustering process includes: first, calculating the distance between any two points of the remaining individual points and the angle of the line connecting the two points with the x-coordinate axis; then, continuing to calculate two lines from the line set in the previous step, to find the line set of all three points in a straight line by determining that the absolute angle difference between the angles formed by the corresponding two straight lines and the x-coordinate axis should be less than a small threshold angle. As shown in the example in Figure 4D , taking any three points in the three regions of corner points 410a, 410i and 410e as an example to show the algorithmic constraints of clustering, for the angle of line 410a-410i with the x-coordinate axis and the angle between line 410i-410e and the x-coordinate axis :
[0065] When (6)
[0066] It can be determined that the current three points 410a, 410i and 410e are on a straight line, and the three points are determined as a 3-point collinear tuple of 410a-410i-410e. In this way, this step can be repeated for any two of these individual points to be traversed to obtain a set of candidate lines including all three points on a line. In one or more embodiments, the small threshold , preferably ≤ 5°. All points that cannot be classified into any 3-point collinear tuple can be eliminated, thereby further reducing the candidate corner points for final clustering. Note that the final nine corner points determined in the image must follow the 3-point collinear constraint, and each corner point can be located on more than one line.
[0067] Since most of the noise has been filtered out, and all the remaining points have been traversed. Using any of the remaining candidate points as a point of the center position of the calibration target, that is, the corner point 410i, it can be determined whether there are 8 surrounding points that can satisfy 8 groups of 3-point collinear.
[0068] Due to the fact that the method of the present subject matter has been optimized as described previously and most of the noise points have been eliminated, the selection of the nine corner points can be run in real time, for example, by a general-purpose processor (in a processor such as a laptop computer).
[0069] So far, the nine expected corner points in the calibration pattern have all been detected. Each of the detected corner points in the image can then be sub-pixelized to improve its accuracy, that is, the integer pixel points of the nine discrete corner points can be converted to floating-point numbers. Therefore, a higher resolution is obtained, thereby improving the accuracy of the position of each point.
[0070] Finally, the coordinates of the four center points in the 2D camera image can be calculated by some known geometric algorithm. The position of each of the four center points P0, P1, P2, P3 after de-warping can be expressed as: .
[0071] The above-mentioned positioned four center points P0-P3 can be represented in homogeneous coordinates as:
[0072] (7)
[0073] In the next step S332, the four center points in the de-distorted image are projected to the camera physical coordinate system to obtain their camera coordinates. The camera intrinsic matrix K is newly set for the de-distorted case 新 Can be found from the relationship described in (3) above, and thus the four center points in the homogeneous coordinates of the camera physical coordinates can be derived as:
[0074] (8)
[0075] On the other hand, for the laser radar, the purpose is to detect the four center points respectively belonging to the four circular holes. Specifically, the laser radar can scan with its laser radar beams and return 3D point cloud data. The returned information is a 3D point cloud representing range measurements.
[0076] In step S340, in the joint calibration scenario 100 combining Figure 1 and Figure 4A , the laser radar 130 scans and returns 3D point clouds. First, the laser radar receives the returned 3D point clouds, and in step S342, a de-distortion process is automatically invoked to adjust the returned 3D point clouds to 3D laser radar point clouds. Then, an image of the 3D laser radar point clouds can be generated.
[0077] Figure 5A An example image frame generated from 3D point clouds returned from the laser radar is shown. Because one single frame of returned laser radar point clouds can be very sparse, in order to increase the density of the laser radar point clouds, multiple frames of laser radar point clouds can be collected to generate each frame of the image. Thus, as Figure 5A shown, the image frame can be generated from the scan results of multiple laser radar point clouds, or from the scan results of multiple selected laser radar point clouds. In one or more embodiments, every ten frames of laser radar point clouds can be collected to generate each single frame of the image.
[0078] Referring to Figure 5A , an image frame is generated from ten frames of 3D laser radar point clouds 500, where the point clouds scanned by the laser radar beams around the joint calibration board are selected to generate an image frame similar to Figure 4A on the FOV, and where the edges of the four circular holes 520, 522, 524, 526 of the joint calibration board are the targets expected to be detected by the laser radar. Then, plane point selection can be performed in step S344 by running a depth estimation appropriately, which applies a predefined distance to filter the laser radar point clouds to select the desired points distinguishable from the background points. In this example, the points with a distance much larger than the predefined distance can be filtered out, thus the desired points are selected from the background points, and such desired points can be considered to be formed from the echoes reflected from the calibration board.
[0079] A depth discontinuity point detection is performed in step S346 to find potential points belonging to the circular holes by the depth difference of neighboring points. In this step, each point in the plane of the calibration plate is assigned an amplitude which represents the depth difference with respect to its neighboring points. For each point belonging to the plane of the calibration plate with a given distance measurement of the LIDAR points the following is performed:
[0080] (9)
[0081] where the neighboring points for which the scan is performed are and . The discontinuity value is used to filter out all points on the plane of the calibration plate, resulting in the selection of the point cloud as shown in Figure 5B . The points shown in Figure 5B do not belong to the circular holes and can be found and removed from them.
[0082] Next, in step S348, the four circles on the calibration plate are intended to be segmented. The points that do not belong to the circles have been removed by the filtering process performed previously and only discontinuities exist in the cloud in this stage. Thus, the outliers to be removed mostly come from the outer border of the calibration target. The LIDAR point cloud can be processed to keep only the rings formed by a number of points compatible with the existence of a circle and then remove the outer points in these rings. The filtered cloud is shown in Figure 5C .
[0083] Thereafter, in step S350, the circle points are projected onto a plane, such as onto the XY plane of the LIDAR coordinate system, and a 2D space circle fitting is performed, resulting in four 2D circles representing the edges of the corresponding circular holes on the calibration plate, as shown in Figure 5D . At this time, based on the four fitted circles, and since the dimensions of the calibration plate are preset and known, it is easy to determine the four circle centers by considering the radius of the circular holes as a constraint by some geometric algorithm. Finally, the four located center points, which will be used as the corresponding key points between the camera and the LIDAR, are transformed back to the 3D space.
[0084] Following the above steps, four center points P0, P1, P2, P3 can be detected from each calibration cloud point of the LIDAR, and each of the four center points can be expressed as: . And the homogeneous coordinates of the center points P0, P1, P2, P3 in the 3D LIDAR coordinates can be expressed as:
[0085] (10)
[0086] In one or more embodiments, to improve the accuracy of the joint calibration, as an additional or alternative step S360, the single calibration board can be moved to several different positions within the common FOV of the camera and the lidar, and the processing steps S320-S350 can be repeated several times. For these repetitions, the joint extrinsic parameter estimation can be performed several times by recording the positioning results of the corresponding sets of center points of the camera and the lidar, respectively, as will be described below.
[0087] The coordinates in the 2D camera coordinate system are thus The 3D coordinates in the 3D lidar coordinate system are thus The four center points P0, P1, P2, P3 are all represented in homogeneous coordinates, respectively. Thus, finding the transformation between the 2D camera coordinate system and the 3D lidar coordinate system translates into solving the transformation between 2D points and 3D points, which can be formulated as a least squares problem, as also concerned by the PnP algorithm.
[0088] Thus, in a final step S370, the joint extrinsic parameter estimation between the lidar coordinate system and the camera coordinate system can be performed. The extrinsic parameter transformation, including the extrinsic parameters (rotation matrix R and translation vector t) between the camera and the lidar, can be solved, respectively.
[0089] In one or more embodiments, the PnP algorithm can be such as EPnP and / or Levenberg-Marquardt optimization. In the example of single calibration shown in Fig. 2, for the four center points P0, P1, P2, P3 to be detected, n = 4. In the example of single calibration shown in Fig. 3, for the eight center points P0, P1,..., P7 to be detected, n = 8. Figure 3
[0090] The provided algorithm can efficiently and accurately detect the calibration target from the camera and lidar sensors. Thus, based on the detected calibration target, the extrinsic parameters of the camera and the lidar can be jointly estimated. Due to the novel design of the calibration board and the algorithm, the calibration setup is simple and efficient, while the calibration capability is robust and accurate.
[0091] The methods in this disclosure can be performed using any combination of one or more computer readable media. The computer readable media can be computer readable signal media or computer readable storage media. Computer readable storage media can be, for example but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0092] The inventive subject matter includes, but is not limited to, the following listed items.
[0093] Item 1: A method for joint extrinsic calibration of at least one camera and at least one lidar based on a single calibration board, comprising the steps of:
[0094] locating a first plurality of center points from a 2D image captured by the at least one camera including the single calibration board;
[0095] locating a second plurality of center points from a 3D point cloud returned by the at least one lidar containing points reflected from the single calibration board; and estimating an extrinsic transformation between the at least one camera and the at least one lidar, wherein the first plurality of center points respectively belong to a rectangular grid on the single calibration board;
[0096] wherein the second plurality of center points respectively belong to circular holes in the single calibration board;
[0097] wherein the first plurality of center points and the second plurality of center points fully coincide in position.
[0098] Item 2: The method according to item 1, wherein four vertices of each of the rectangular grids are arranged as corner points detectable by the at least one camera.
[0099] Item 3: The method according to any one of items 1 or 2, wherein the rectangular grids are arranged on the single calibration board as M N grids, M and N are integers, and corresponding circular holes are arranged coincidently in the rectangular grids.
[0100] Item 4: The method according to any one of items 1-3, wherein detecting the corner points by the at least one camera comprises filtering the 2D image using a 2D filter having a kernel created based on the calibration plate.
[0101] Item 5: The method according to any one of items 1-4, wherein the corner points are each defined by the intersection of the diagonals of two black rectangles, and the kernel is created based on the shape of the intersection of the diagonals of the two black rectangles.
[0102] Item 6: The method according to any one of Items 1-5, wherein detecting the corner points by the at least one camera comprises accurately detecting the corner points by determining that every three of the corner points are on a straight line.
[0103] Item 7: According to the method of any one of Items 1 to 6, the second plurality of center points may be located in a 2D space, wherein 2D circles are respectively fitted by points on the edge of the circular hole on the single calibration plate.
[0104] Item 8: The method according to any one of Items 1 to 7, wherein the edge of the circular hole can be detected from the single calibration plate using depth discontinuity detection.
[0105] Item 9: A method according to any one of Items 1-8, wherein the 2D image captured by the at least one camera including the single calibration plate is subjected to dedistortion processing before the first plurality of center points are located therefrom, wherein the dedistortion processing is based on an intrinsic parameter matrix and intrinsic parameter distortion coefficients.
[0106] Item 10: A method according to any one of Items 1-9, wherein estimating the extrinsic parameter transformation of the at least one camera and the at least one lidar comprises respectively calculating the rotation matrix and translation vector between the at least one camera and the at least one lidar.
[0107] Item 11. A non-transitory computer-readable medium storing instructions, the instructions being processable by one or more processors to implement the following steps, the steps comprising:
[0108] locating a first plurality of center points from a 2D image captured by the at least one camera and including the single calibration plate;
[0109] locating a second plurality of center points from a 3D point cloud returned by the at least one lidar including points reflected from the single calibration plate; and estimating an extrinsic transformation between the at least one camera and the at least one lidar, wherein the first plurality of center points respectively belong to a rectangular grid on the single calibration plate;
[0110] wherein the second plurality of center points respectively belong to the circular holes in the single calibration plate;
[0111] wherein the first plurality of center points and the second plurality of center points are completely coinciding in position.
[0112] Item 12. The non-transitory computer-readable medium of item 11, wherein the four vertices of each of the rectangular grids are arranged as corner points detectable by the at least one camera.
[0113] Item 13. The non-transitory computer-readable medium of item 11 or 12, wherein the rectangular grids are arranged on the single calibration plate as M N grids, M and N are integers, and corresponding circular holes are arranged coinciding in the rectangular grids.
[0114] Item 14. The non-transitory computer-readable medium of any one of items 11-13, wherein detecting the corner points by the at least one camera comprises filtering the 2D image by using a 2D filter with a kernel created based on the calibration plate.
[0115] Item 15. The non-transitory computer-readable medium of any one of items 11-14, wherein the corner points are each defined by an intersection of diagonals of two black rectangles, and the kernel is created based on a shape of the intersection of diagonals of the two black rectangles.
[0116] Item 16. The non-transitory computer-readable medium of any one of items 11-15, wherein detecting the corner points by the at least one camera comprises accurately detecting the corner points by determining that each three of the corner points are on a straight line.
[0117] Item 17. The non-transitory computer-readable medium of any one of items 11-16, wherein the second plurality of center points are locatable on a 2D space, wherein 2D circles are respectively fitted by points of edges of the circular holes on the single calibration plate.
[0118] Item 18. The non-transitory computer-readable medium of any one of items 11-17, wherein the edges of the circular holes are respectively detected from the single calibration plate using depth discontinuity point detection.
[0119] Item 19. The non-transitory computer-readable medium of any one of items 11-18, wherein the 2D image including the single calibration plate captured by the at least one camera is applied a de-warping process before locating the first plurality of center points therefrom, wherein the de-warping process is based on an intrinsic matrix and intrinsic distortion coefficients.
[0120] Item 20. A non-transitory computer-readable medium according to any one of Items 11-19, wherein estimating the extrinsic parameter transformation of the at least one camera and the at least one lidar comprises respectively calculating a rotation matrix and a translation vector between the at least one camera and the at least one lidar.
[0121] As used in this application, an element or step recited in the singular and preceded by the word "a(n)" should be understood as not excluding plural of said elements or steps, unless such exclusion is indicated. In addition, reference to "one embodiment" or "an example" of the present disclosure is not intended to be interpreted as excluding the existence of additional embodiments that also include the recited features. The terms "first," "second," and "third," etc. are used merely as labels and are not intended to impose numerical requirements or a particular positional order on their objects.
[0122] Although exemplary embodiments have been described above, these embodiments are not intended to describe all possible forms of the present disclosure. On the contrary, the words used in this specification are descriptive, rather than limiting, and it should be understood that various changes can be made without departing from the spirit and scope of the present disclosure. In addition, the features of the various implemented embodiments can be combined to form further embodiments of the present invention.
Claims
1. A method for joint extrinsic calibration of at least one camera and at least one lidar based on a single calibration plate, comprising the following steps: locating a first plurality of center points from a 2D image captured by the at least one camera and including the single calibration plate; locating a second plurality of center points from a 3D point cloud returned by the at least one lidar including points reflected from the single calibration plate; and estimating an extrinsic transformation between the at least one camera and the at least one lidar, wherein the first plurality of center points respectively belong to a rectangular grid on the single calibration plate; wherein the second plurality of center points respectively belong to the circular holes in the single calibration plate; The first plurality of center points and the second plurality of center points completely overlap in position. 2 . The method according to claim 1 , wherein four vertices of each of the rectangular meshes are arranged as corner points that can be detected by the at least one camera.
3. The method according to claim 1, wherein the rectangular grid is arranged as M on the single calibration plate. N grids, M and N is an integer, and corresponding circular holes are arranged in a coincident manner in the rectangular grid.
4. The method according to claim 2, wherein Detecting the corner points by the at least one camera includes filtering the 2D image using a 2D filter having a kernel created based on the calibration plate. 5 . The method of claim 4 , wherein the corner points are each defined by the intersection of diagonals of two black rectangles, and the kernel is created based on the shape of the intersection of the diagonals of the two black rectangles.
6. The method according to claim 2, wherein Detecting the corner points by the at least one camera includes accurately detecting the corner points by determining that every three of the corner points are on a straight line. 7 . The method according to claim 2 , wherein the second plurality of center points can be located in a 2D space, wherein 2D circles are respectively fitted by points of the edges of the circular holes on the single calibration plate. 8 . The method according to claim 7 , wherein the edges of the circular holes can be detected separately from the single calibration plate using depth discontinuity detection.
9. The method according to claim 1, wherein The 2D image captured by the at least one camera and including the single calibration plate is subjected to a dedistortion process before the first plurality of center points are located therefrom, wherein the dedistortion process is based on an intrinsic parameter matrix and intrinsic distortion coefficients.
10. The method according to claim 1, wherein estimating the extrinsic parameter transformation of the at least one camera and the at least one lidar comprises respectively calculating a rotation matrix and a translation vector between the at least one camera and the at least one lidar.
11. A non-transitory computer-readable medium storing instructions capable of being processed by one or more processors to implement the following steps, the steps comprising: locating a first plurality of center points from a 2D image captured by the at least one camera and including the single calibration plate; locating a second plurality of center points from a 3D point cloud returned by the at least one lidar including points reflected from the single calibration plate; and estimating an extrinsic transformation between the at least one camera and the at least one lidar, wherein the first plurality of center points respectively belong to a rectangular grid on the single calibration plate; wherein the second plurality of center points respectively belong to the circular holes in the single calibration plate; The first plurality of center points and the second plurality of center points completely overlap in position. 12 . The non-transitory computer-readable medium of claim 11 , wherein four vertices of each of the rectangular meshes are arranged as corner points detectable by the at least one camera.
13. The non-transitory computer readable medium according to claim 11, wherein the rectangular grid is arranged as M on the single calibration plate. N grids, M and N is an integer, and corresponding circular holes are arranged in a coincident manner in the rectangular grid.
14. The non-transitory computer-readable medium of claim 12, wherein Detecting the corner points by the at least one camera includes filtering the 2D image using a 2D filter having a kernel created based on the calibration plate. 15 . The non-transitory computer-readable medium of claim 14 , wherein the corner points are each defined by the intersection of diagonals of two black rectangles, and the kernel is created based on the shape of the intersection of the diagonals of the two black rectangles.
16. The non-transitory computer readable medium of claim 12, wherein Detecting the corner points by the at least one camera includes accurately detecting the corner points by determining that every three of the corner points are on a straight line. 17 . The non-transitory computer-readable medium of claim 12 , wherein the second plurality of center points can be located on a 2D space, wherein 2D circles are respectively fitted by points of edges of the circular holes on the single calibration plate. 18 . The non-transitory computer-readable medium of claim 17 , wherein the edges of the circular holes can be detected separately from the single calibration plate using depth discontinuity detection.
19. The non-transitory computer-readable medium of claim 11, wherein The 2D image captured by the at least one camera and including the single calibration plate is subjected to a dedistortion process before the first plurality of center points are located therefrom, wherein the dedistortion process is based on an intrinsic parameter matrix and intrinsic distortion coefficients.
20. The non-transitory computer-readable medium of claim 19, wherein estimating the extrinsic transformation of the at least one camera and the at least one lidar comprises respectively calculating a rotation matrix and a translation vector between the at least one camera and the at least one lidar.