A camera and radar joint calibration method and device for columnar object semantic segmentation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI JIANGHUAI AUTOMOBILE GRP CORP LTD
- Filing Date
- 2026-04-22
- Publication Date
- 2026-08-04
AI Technical Summary
一是受限激光反射强度和激光点云数据量大、稀疏且无结构情况,鲁棒性较差且计算效率低下;
[0007]The beneficial effects of this invention are as follows: By acquiring the cylindrical point cloud of the vehicle in the current vehicle body coordinate system, and using the widely existing cylindrical objects in natural scenes as calibration references, there is no need to arrange specific artificial calibration objects, thus overcoming the environmental limitations of the calibration field; based on the initial extrinsic parameters characterizing the spatial transformation relationship between the lidar coordinate system and the camera coordinate system, the cylindrical point cloud is projected onto the image space to obtain a two-dimensional projected point cloud image. Then, by binarizing the two-dimensional projected point cloud image and extracting the edge points of the cylindrical objects through morphological operations, the edge point cloud of the lidar cylindrical objects is obtained. This transforms the edge extraction problem of the three-dimensional point cloud into a morphological processing problem in the two-dimensional image space, avoiding the high computational complexity caused by directly performing curvature calculations or boundary detection in three-dimensional space, and improving the computational efficiency and robustness of point cloud edge extraction; by acquiring vehicle visual image data and using the SURF feature combined with Hough transform to extract the straight line edge features in the image to obtain a visual edge feature point set, the SURF feature... Scale invariance and the robust detection capability of Hough transform for linear structures are utilized to stably extract the linear edge features corresponding to columnar objects in images. The laser columnar object edge point cloud is projected onto the image space and matched with the visual edge feature point set. Based on the matching result, the reprojection error value, which characterizes the geometric position error between the pixel points after the laser columnar object edge point cloud is projected onto the image plane and the visual edge feature points, is determined, establishing a cross-modal geometric constraint relationship between the laser point cloud and the image. The relative pose of the camera and the lidar is iteratively optimized using a nonlinear least squares method until the reprojection error value is minimized. The relative pose corresponding to the minimum reprojection error value is taken as the optimal joint calibration result, realizing the automatic refinement of the calibration parameters. This avoids the problem of step search in traditional methods potentially getting trapped in local optima and having limited accuracy. Therefore, it does not rely on specific manual calibration objects and overcomes the problems of poor robustness and low computational efficiency caused by the large, sparse, and unstructured laser point cloud data, which are limited by the laser reflection intensity.
Smart Images

Figure CN122510286A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and more specifically, to a camera and radar joint calibration method and apparatus for columnar semantic segmentation. Background Technology
[0002] The prior art, patent application CN112819903A, discloses a method for joint calibration of a camera and a lidar based on an L-shaped calibration board. The method includes: mounting the camera and lidar on the device to be calibrated; placing the L-shaped calibration board on the ground within the field of view of the camera and lidar; activating the camera and lidar to acquire data, obtaining image data including the L-shaped calibration board from the camera and point cloud data including the L-shaped calibration board from the lidar; performing corner detection on the acquired image data to obtain the coordinates of the checkerboard corner points on the two planes of the L-shaped calibration board in the pixel coordinate system; performing plane segmentation and fitting on the acquired point cloud data to obtain the equations of the two planes of the L-shaped calibration board, and using geometric information to obtain the coordinates of the checkerboard corner points on the two planes of the L-shaped calibration board in the lidar coordinate system; and calculating the pose changes of the camera and lidar based on the coordinates of the corner points of the L-shaped calibration board in the pixel coordinate system and the coordinates of the corner points of the L-shaped calibration board in the lidar coordinate system.
[0003] The above-mentioned prior art has the following technical defects: First, the limitations of laser reflection intensity and the large, sparse, and unstructured nature of laser point cloud data result in poor robustness and low computational efficiency. Secondly, while it can achieve high accuracy under specific calibration conditions, the operation process is cumbersome, requires a high degree of human intervention, and is limited by the geometric characteristics of the calibration object, making it difficult to meet the actual needs of online calibration, automated calibration, or large-scale fleet calibration.
[0004] In summary, there is an urgent need for a camera and radar joint calibration method that does not rely on specific calibration objects, can utilize common structural features in the natural environment, and is suitable for online automated calibration. This would overcome the technical shortcomings of traditional methods, such as being limited by sparse and unstructured laser point clouds, cumbersome calibration processes, high degree of human intervention, and difficulty in meeting the needs of large-scale vehicle fleets or real-time calibration. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a camera and radar joint calibration method and apparatus for semantic segmentation of columnar objects, which aims to solve at least one of the above-mentioned technical problems.
[0006] In a first aspect, the technical solution of the present invention to solve the above-mentioned technical problems is as follows: a camera and radar joint calibration method for semantic segmentation of columnar objects, the method comprising: Obtain the cylindrical point cloud of the vehicle in the vehicle body coordinate system at the current moment; Based on the initial extrinsic parameters of the camera and lidar, the point cloud of the columnar object is projected onto the image space to obtain a two-dimensional projected point cloud image. The initial extrinsic parameters characterize the spatial transformation relationship between the lidar coordinate system and the camera coordinate system. The two-dimensional projected point cloud image is binarized, and the edge points of the column are extracted through morphological operations to obtain the edge point cloud of the laser column. The vehicle visual image is acquired, and the straight line edge features in the vehicle visual image are extracted by combining SURF features with Hough transform to obtain the visual edge feature point set. The laser column edge point cloud is projected onto the image space and matched with the visual edge feature point set. The reprojection error value is determined based on the matching result. The reprojection error value characterizes the geometric position error between the pixel point after the laser column edge point cloud is projected onto the image plane and the visual edge feature point in the visual edge feature point set. The relative pose of the camera and the lidar is iteratively optimized using a nonlinear least squares method until the reprojection error is minimized. The relative pose corresponding to the minimum reprojection error is then taken as the optimal joint calibration result.
[0007] The beneficial effects of this invention are as follows: By acquiring the cylindrical point cloud of the vehicle in the current vehicle body coordinate system, and using the widely existing cylindrical objects in natural scenes as calibration references, there is no need to arrange specific artificial calibration objects, thus overcoming the environmental limitations of the calibration field; based on the initial extrinsic parameters characterizing the spatial transformation relationship between the lidar coordinate system and the camera coordinate system, the cylindrical point cloud is projected onto the image space to obtain a two-dimensional projected point cloud image. Then, by binarizing the two-dimensional projected point cloud image and extracting the edge points of the cylindrical objects through morphological operations, the edge point cloud of the lidar cylindrical objects is obtained. This transforms the edge extraction problem of the three-dimensional point cloud into a morphological processing problem in the two-dimensional image space, avoiding the high computational complexity caused by directly performing curvature calculations or boundary detection in three-dimensional space, and improving the computational efficiency and robustness of point cloud edge extraction; by acquiring vehicle visual image data and using the SURF feature combined with Hough transform to extract the straight line edge features in the image to obtain a visual edge feature point set, the SURF feature... Scale invariance and the robust detection capability of Hough transform for linear structures are utilized to stably extract the linear edge features corresponding to columnar objects in images. The laser columnar object edge point cloud is projected onto the image space and matched with the visual edge feature point set. Based on the matching result, the reprojection error value, which characterizes the geometric position error between the pixel points after the laser columnar object edge point cloud is projected onto the image plane and the visual edge feature points, is determined, establishing a cross-modal geometric constraint relationship between the laser point cloud and the image. The relative pose of the camera and the lidar is iteratively optimized using a nonlinear least squares method until the reprojection error value is minimized. The relative pose corresponding to the minimum reprojection error value is taken as the optimal joint calibration result, realizing the automatic refinement of the calibration parameters. This avoids the problem of step search in traditional methods potentially getting trapped in local optima and having limited accuracy. Therefore, it does not rely on specific manual calibration objects and overcomes the problems of poor robustness and low computational efficiency caused by the large, sparse, and unstructured laser point cloud data, which are limited by the laser reflection intensity.
[0008] Based on the above technical solution, the present invention can be further improved as follows.
[0009] Furthermore, the acquisition of the cylindrical point cloud of the vehicle in the vehicle body coordinate system at the current moment includes: The LiDAR point cloud of the vehicle is acquired. Each point in the LiDAR point cloud corresponds to a semantic value. For each semantic value, the semantic value represents the category of the corresponding point. Based on the semantic values of all points in the LiDAR point cloud, all points in the LiDAR point cloud are filtered out, and the point cloud corresponding to the semantic value as columnar is retained, so as to obtain the columnar point cloud of the vehicle in the vehicle body coordinate system at the current time.
[0010] Furthermore, the above-mentioned binarization processing of the two-dimensional projected point cloud image and the extraction of the edge points of the columnar object through morphological operations to obtain the laser columnar object edge point cloud include: Binarize the two-dimensional projected point cloud image to generate a binary image; Dilation and erosion operations are performed sequentially on the binary image. Edge points are obtained by the difference between the binary image and the eroded image. Based on the laser point cloud corresponding to the edge points, the edge point cloud of the laser column is obtained.
[0011] Furthermore, the above-mentioned dilation and erosion operations are performed sequentially on the binary image, and the edge points are obtained by the difference between the binary image and the eroded image, including: A rectangular sliding window with a size of 3 pixels × 3 pixels is used; The binary image is traversed using a rectangular sliding window to obtain the pixel corresponding to the maximum value among all pixels within the rectangular sliding window, resulting in the dilated image. The dilated image is then traversed using the same rectangular sliding window to obtain the pixel corresponding to the minimum value among all pixels within the rectangular sliding window, resulting in the eroded image. The grayscale values of pixels at the same location in the binary image and the eroded image are subtracted to obtain the difference image. Pixels with grayscale values greater than 0 in the difference image are extracted as the edge points of the columnar structures.
[0012] Furthermore, based on the initial extrinsic parameters of the camera and LiDAR, the columnar point cloud is projected onto the image space to obtain a two-dimensional projected point cloud image, including: Based on the initial extrinsic parameters, each point in the columnar point cloud is projected into the image space to obtain a two-dimensional projected point cloud image.
[0013] Furthermore, the above-mentioned method, which combines SURF features with Hough transform, extracts straight line edge features from vehicle visual images to obtain a set of visual edge feature points, including: The vehicle visual image is converted to grayscale to obtain a grayscale image; Perform Canny edge detection on the grayscale image to obtain a binary edge image; The Hough transform is used to detect straight lines in a binary edge image, and the edge points along the detected lines are extracted as the Hough edge feature point set. A scale space is constructed for the grayscale image, and extreme points are detected in the scale space using the Hessian matrix as SURF feature points, so as to generate a SURF feature point set based on all SURF feature points. The Hough edge feature point set and the SURF feature point set are fused, and the SURF feature points that are simultaneously located on the detected lines are retained as the visual edge feature point set.
[0014] Furthermore, the above-mentioned projection of the laser columnar object edge point cloud onto the image space and matching it with the visual edge feature point set, and the determination of the reprojection error value based on the matching result, includes: Project each point in the laser columnar object edge point cloud onto the image plane to obtain the laser projection point set; calculate the pixel distance between each point in the laser projection point set and each point in the visual edge feature point set, and determine the closest pair of points as a matching point pair; The RANSAC algorithm is used to filter all matching point pairs, outlier matching point pairs are removed, and interior point matching point pairs that satisfy the same projection transformation relationship are retained as valid matching point pairs. The sum of squared pixel distances between each laser projection point and the corresponding matched visual edge feature point in the effective matching point pair is calculated as the reprojection error value.
[0015] Secondly, to solve the above-mentioned technical problems, the present invention also provides a camera and radar joint calibration device for columnar semantic segmentation, the device comprising: The acquisition module is used to acquire the cylindrical point cloud of the vehicle in the vehicle body coordinate system at the current moment; The two-dimensional projection module is used to project the cylindrical point cloud onto the image space based on the initial extrinsic parameters of the camera and the lidar to obtain a two-dimensional projected point cloud image. The initial extrinsic parameters characterize the spatial transformation relationship between the lidar coordinate system and the camera coordinate system. The edge point cloud determination module is used to binarize the two-dimensional projected point cloud image and extract the edge points of the columnar object through morphological operations to obtain the edge point cloud of the laser columnar object. The visual edge feature point set determination module is used to acquire vehicle visual images. It uses the SURF feature combined with Hough transform to extract straight line edge features in the vehicle visual images and obtain the visual edge feature point set. The reprojection error value determination module is used to project the laser column edge point cloud onto the image space, match it with the visual edge feature point set, and determine the reprojection error value based on the matching result. The reprojection error value characterizes the geometric position error between the pixel point after the laser column edge point cloud is projected onto the image plane and the visual edge feature point in the visual edge feature point set. The joint calibration result determination module is used to iteratively optimize the relative pose of the camera and the lidar using a nonlinear least squares method until the reprojection error value is minimized. The relative pose corresponding to the minimum reprojection error value is then taken as the optimal joint calibration result.
[0016] Thirdly, in order to solve the above-mentioned technical problems, the present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the camera and radar joint calibration method for columnar semantic segmentation of the present application.
[0017] Fourthly, in order to solve the above-mentioned technical problems, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the camera and radar joint calibration method for columnar semantic segmentation of the present application.
[0018] Additional aspects and advantages of this application will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below.
[0020] Figure 1 This is a flowchart illustrating a camera and radar joint calibration method for semantic segmentation of columnar objects, provided in one embodiment of the present invention. Figure 2 A schematic flowchart illustrating another method for joint camera and radar calibration of semantic segmentation of columnar objects, provided in an embodiment of the present invention; Figure 3 A schematic diagram of a camera and radar joint calibration device for columnar semantic segmentation provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. Detailed Implementation
[0021] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0022] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.
[0023] The data acquisition process involved in this invention follows the principles of legality, legitimacy, and necessity. Based on obtaining the explicit authorization and consent of the user, only the minimum necessary information required to achieve the purpose is collected, and data security protection obligations are fulfilled in accordance with the law.
[0024] The solution provided in this invention can be applied to any application scenario that requires joint radar calibration.
[0025] This invention provides a possible implementation, such as... Figure 1 The diagram shows a flowchart of a camera and radar joint calibration method for semantic segmentation of columnar objects. This method can be executed by any electronic device, such as a terminal device, or jointly executed by a terminal device and a server. For ease of description, the method provided in this embodiment will be described below using a terminal device as the execution subject as an example. Figure 1 The flowchart shown indicates that the method may include the following steps: S10, Obtain the cylindrical point cloud of the vehicle in the vehicle body coordinate system at the current moment; S20, based on the initial extrinsic parameters of the camera and lidar, projects the columnar point cloud onto the image space to obtain a two-dimensional projected point cloud image. The initial extrinsic parameters characterize the spatial transformation relationship between the lidar coordinate system and the camera coordinate system. S30 performs binarization processing on the two-dimensional projected point cloud image and extracts the edge points of the columnar object through morphological operations to obtain the laser columnar object edge point cloud. S40: Acquire a vehicle visual image. Use the SURF feature combined with Hough transform to extract the straight line edge features in the vehicle visual image and obtain a set of visual edge feature points. S50: Project the laser column edge point cloud onto the image space and match it with the visual edge feature point set. Based on the matching result, determine the reprojection error value. The reprojection error value characterizes the geometric position error between the pixel points after the laser column edge point cloud is projected onto the image plane and the visual edge feature points in the visual edge feature point set. S60 iteratively optimizes the relative pose of the camera and the lidar using a nonlinear least squares method until the reprojection error is minimized. The relative pose corresponding to the minimum reprojection error is then taken as the optimal joint calibration result.
[0026] The method of this invention acquires a cylindrical point cloud of the vehicle in the current vehicle body coordinate system, using widely existing cylindrical objects in natural scenes as calibration references, eliminating the need for specific artificial calibration objects and overcoming the environmental limitations of the calibration field. Based on initial extrinsic parameters characterizing the spatial transformation relationship between the lidar coordinate system and the camera coordinate system, the cylindrical point cloud is projected into the image space to obtain a two-dimensional projected point cloud image. By binarizing the two-dimensional projected point cloud image and extracting the edge points of the cylindrical objects through morphological operations, the edge point cloud of the lidar cylindrical objects is obtained. This transforms the edge extraction problem of three-dimensional point clouds into a morphological processing problem in two-dimensional image space, avoiding the high computational complexity caused by directly performing curvature calculations or boundary detection in three-dimensional space, thus improving the computational efficiency and robustness of point cloud edge extraction. Furthermore, by acquiring vehicle visual image data and using SURF features combined with Hough transform to extract straight line edge features from the image to obtain a visual edge feature point set, the method utilizes the scale of SURF features... The method leverages degree invariance and the robust detection capability of Hough transform for linear structures to stably extract linear edge features corresponding to columnar objects in images. It projects the laser columnar object edge point cloud onto the image space and matches it with the visual edge feature point set. Based on the matching result, it determines the reprojection error value, which characterizes the geometric position error between the pixels projected onto the image plane and the visual edge feature points, establishing a cross-modal geometric constraint relationship between the laser point cloud and the image. It iteratively optimizes the relative pose of the camera and lidar using a nonlinear least squares method until the reprojection error value is minimized. The relative pose corresponding to the minimum reprojection error value is taken as the optimal joint calibration result, achieving automatic refinement of calibration parameters. This avoids the problem of step search potentially getting trapped in local optima and limited accuracy in traditional methods, thus eliminating the need for specific manual calibration objects. It overcomes the limitations of traditional methods, which are constrained by laser reflection intensity and the large, sparse, and unstructured nature of laser point cloud data, resulting in poor robustness and low computational efficiency.
[0027] The following specific embodiments further illustrate the solution of the present invention, and the technical problems to be solved by the present invention include: 1. A scheme for feature extraction using the semantic segmentation results of columnar objects in real-time vehicle-mounted lidar perception is proposed to solve the problem of poor robustness and low computational efficiency of feature extraction caused by the large, sparse and unstructured laser point cloud data due to limited laser reflection intensity. 2. An extraction scheme is used to extract the edge points of the columnar object as key features and perform a fine matching with the edge feature points extracted from the image. This solves the problem that the geometric features of the calibration object are limited and cannot meet the requirements of online calibration, automated calibration or large-scale fleet calibration.
[0028] Based on this, this embodiment provides a joint camera and radar calibration method for columnar semantic segmentation, which can be applied to autonomous vehicles equipped with LiDAR and cameras. Figure 1 As shown, the method includes the following steps S10 to S60, specifically: S10, Obtain the cylindrical point cloud of the vehicle in the vehicle body coordinate system at the current moment; One implementation of S10 above is as follows: S101, acquire the LiDAR point cloud of the vehicle. Each point in the LiDAR point cloud corresponds to a semantic value, and for each semantic value, the semantic value represents the category of the corresponding point. In this embodiment, the semantic value of the LiDAR point cloud can be obtained by reasoning from the original LiDAR point cloud through a pre-trained semantic segmentation model. The semantic value of each point is used to identify the object category to which the point belongs, such as columnar objects (e.g., street lamp poles, traffic sign poles, tree trunks), ground, vehicles, pedestrians, etc.
[0029] S102, based on the semantic values of all points in the LiDAR point cloud, filter out all points in the LiDAR point cloud, retaining only the point cloud whose semantic value corresponds to the columnar object category, to obtain the columnar object point cloud of the vehicle in the vehicle body coordinate system at the current moment. Specifically, traverse the semantic value corresponding to each point in the LiDAR point cloud, determine whether the semantic value indicates that the point is a columnar object, retain only the points whose semantic category is columnar object, and filter out points of other categories.
[0030] After the above filtering, the remaining point cloud is the cylindrical object point cloud in the vehicle coordinate system at the current moment, denoted as . Each point in the columnar point cloud has three-dimensional coordinates in the vehicle body coordinate system. .
[0031] Between S101 and S102, the method further includes: inspecting the acquired lidar point cloud and removing outliers. Specifically, each point in the lidar point cloud is traversed, and points with values of NaN (non-numerical) or a distance of more than 200 meters are removed to avoid the impact of abnormal data on calibration accuracy.
[0032] S20, based on the initial extrinsic parameters of the camera and lidar, projects the columnar point cloud onto the image space to obtain a two-dimensional projected point cloud image. The initial extrinsic parameters characterize the spatial transformation relationship between the lidar coordinate system and the camera coordinate system. One specific implementation of S20 is as follows: based on the initial extrinsic parameters, each point in the columnar point cloud is projected onto the image space to obtain a two-dimensional projected point cloud image.
[0033] In this embodiment, the aforementioned initial external parameters can be obtained through vehicle design and installation parameters or coarse calibration. They include a rotation matrix R and a translation vector t, which are used to describe the spatial transformation relationship from the lidar coordinate system to the camera coordinate system.
[0034] In this process, projecting each point in the columnar point cloud onto the image space based on the initial extrinsic parameters can be achieved using a camera imaging geometry model. Specifically, the columnar point cloud obtained in step S10... Each point in the array (this point is a three-dimensional point) The corresponding pixel coordinates are obtained by mapping the camera projection model onto the two-dimensional image plane. This leads to the acquisition of a two-dimensional projected point cloud image.
[0035] The projection relationship can be expressed as: in, is the projection function used to convert homogeneous coordinates into two-dimensional pixel coordinates; K is the camera intrinsic parameter matrix, which can be obtained in advance through camera calibration; R and t are the rotation matrix and translation vector of the LiDAR relative to the camera, respectively, i.e., the initial extrinsic parameters.
[0036] All points in the columnar point cloud are projected to form a two-dimensional projected point cloud image. I .
[0037] S30 performs binarization processing on the two-dimensional projected point cloud image and extracts the edge points of the columnar object through morphological operations to obtain the laser columnar object edge point cloud. Alternatively, one implementation of S30 above is as follows: S301, perform binarization processing on the two-dimensional projected point cloud image to generate a binary image. Specifically, each pixel in the two-dimensional projected point cloud image I obtained in step S20 is binarized: for pixel positions with projection points, their grayscale value is set to 255 (foreground); for pixel positions without projection points, their grayscale value is set to 0 (background), thereby generating a binary image. .
[0038] S302, perform dilation and erosion operations sequentially on the binary image, obtain edge points by the difference between the binary image and the eroded image, and obtain the laser point cloud of the laser columnar object based on the laser point cloud corresponding to the edge points.
[0039] In step S302 above, dilation and erosion operations are performed sequentially on the binary image. One way to obtain edge points by the difference between the binary image and the eroded image is as follows: S3021 uses a rectangular sliding window with a size of 3 pixels × 3 pixels; S3022, the binary image is traversed using a rectangular sliding window to obtain the pixel corresponding to the maximum value among all pixels within the rectangular sliding window, and the dilated image is obtained; this operation expands the coverage of foreground pixels in the binary image.
[0040] Each slide yields a pixel corresponding to a maximum value, which is the maximum grayscale value among all pixels within the sliding window.
[0041] The mathematical expression for S3022 above is: Where B represents a 3-pixel × 3-pixel rectangular sliding window, and ⊕ is the dilation operator. This represents the dilated image. In step S3023, a rectangular sliding window is used to traverse the dilated image, and the pixel with the minimum value among all pixels within the rectangular sliding window is used to obtain the eroded image; this operation shrinks the boundary of the foreground region. The combined dilation and erosion operation fills the small holes in the foreground region while maintaining the overall outline and area of the foreground region essentially unchanged.
[0042] Each time the slider is moved, a pixel corresponding to a minimum value is obtained. This minimum value is the minimum grayscale value among all pixels within the sliding window.
[0043] The mathematical expression for S3023 above is: Where ⊖ is the erosion operator, This represents the eroded image; the erosion operation can shrink the expanded region. S3024: Subtract the grayscale values of pixels at the same location in the binary image from those in the eroded image to obtain the difference image. The mathematical expression for S3024 above is: in, This represents the difference image. S3025: Extract pixels with grayscale values greater than 0 from the difference image as the edge points of the columnar structure. Specifically, iterate through the difference image. For each pixel in the graph, if the grayscale value of the pixel is greater than 0, then the pixel is determined to correspond to the edge point of the column.
[0044] Finally, based on the laser point cloud corresponding to the edge points, the edge point cloud of the laser column is obtained. Specifically, the point cloud indices corresponding to the edge points in the difference image are retained, and the 3D point clouds corresponding to these point cloud indices are extracted to form the final edge point cloud of the laser column, denoted as . .
[0045] Each 3D point in the laser point cloud has a unique identifier, which can be a point cloud index (such as a position number or ID in the point cloud array). When the 3D point is projected onto the image plane and is finally determined to be an edge point, it is necessary to trace back to the original point cloud based on this unique identifier in order to extract the corresponding 3D point coordinates.
[0046] By extracting edge points through the above morphological operations, the amount of data in the laser point cloud can be greatly reduced, improving the computational efficiency of subsequent matching and optimization.
[0047] S40: Acquire a vehicle visual image. Use the SURF feature combined with Hough transform to extract the straight line edge features in the vehicle visual image and obtain a set of visual edge feature points. In this embodiment, a camera installed on the vehicle can capture real-time visual images of the surrounding environment.
[0048] Alternatively, one implementation of the above S40 is as follows: S401, perform grayscale processing on the vehicle visual image to obtain a grayscale image; specifically, converting the original color image into a grayscale image can reduce computational complexity.
[0049] S402, perform Canny edge detection on the grayscale image to obtain a binary edge image; wherein, Canny edge detection extracts stable edge contours in the image by calculating the gradient magnitude and direction of the image, and using double threshold detection and edge connection.
[0050] S403 uses the Hough transform to detect straight lines in a binary edge image, extracting the edge points traversed by the detected lines as the Hough edge feature point set. Specifically, the Hough transform is used to detect straight lines in the binary edge image, outputting the parametric equation of each line or the coordinate set of all edge points on the line, i.e., the Hough edge feature point set, denoted as... S404. Construct a scale space for the grayscale image, and detect extreme points in the scale space using the Hessian matrix as SURF feature points to generate a SURF feature point set based on all SURF feature points. Specifically, SURF (Speed-Up Robust Features) feature extraction includes the following steps: First, construct a scale space for the grayscale image; then, calculate the Hessian matrix for each pixel in the scale space, and locate candidate feature points by detecting the extreme values of the determinant of the Hessian matrix; finally, perform principal direction assignment and feature descriptor generation to obtain a SURF feature point set with scale invariance and rotation invariance, denoted as . S405, the Hough edge feature point set and the SURF feature point set are fused, retaining the SURF feature points that are simultaneously located on the detected straight line as the visual edge feature point set. Specifically, the SURF feature point set is traversed. For each feature point in the dataset, it is determined whether the point lies on any straight line detected by the Hough transform. If it does, the point is retained; otherwise, it is discarded. The fused visual edge feature point set combines the geometric constraints of the straight-line edges of the Hough transform with the robustness of the SURF features, denoted as . .
[0051] S50: Project the laser column edge point cloud onto the image space and match it with the visual edge feature point set. Based on the matching result, determine the reprojection error value. The reprojection error value characterizes the geometric position error between the pixel points after the laser column edge point cloud is projected onto the image plane and the visual edge feature points in the visual edge feature point set. In the above S50, one way to project the point cloud of the laser columnar object's edge onto the image space, match it with the visual edge feature point set, and determine the reprojection error value based on the matching result is as follows: S501, Project each point in the laser columnar object edge point cloud onto the image plane to obtain the laser projection point set; specifically, the rotation matrix R and translation vector t in the initial extrinsic parameters can be used to transform the laser columnar object edge point cloud obtained in step S30. Each point in The laser projection point set is obtained by mapping the camera projection model onto the image plane. ,in: Where R and t are the rotation matrix and translation vector in the initial extrinsic parameters, used to describe the transformation relationship from the lidar coordinate system to the camera coordinate system.
[0052] S502, calculate the pixel distance between each point in the laser projection point set and each point in the visual edge feature point set, and determine the closest pair of points as a matching point pair; specifically, for the laser projection point set... For each point A in the dataset, calculate its relationship with the set of visual edge feature points. The Euclidean distance (pixel distance) between each point in the graph is used to select the point with the smallest distance as the candidate matching point for point A, thus forming a matching point pair.
[0053] S503, the RANSAC algorithm is used to filter all matching point pairs, removing outlier matching point pairs (in the matching relationship established between the laser projection point and the visual edge feature point, those erroneous matches whose geometric position deviation is significantly greater than other matching point pairs and do not conform to the true projection transformation relationship), and retaining interior point matching point pairs that satisfy the same projection transformation relationship as valid matching point pairs; the specific process of the RANSAC (Random Sample Consensus) algorithm is as follows: randomly select a minimum subset (e.g., 3 or 4 pairs of matching point pairs) from all candidate matching point pairs, use this subset to estimate a projection transformation model; then count the number of matching point pairs that conform to the transformation model (i.e., the number of interior points) among all matching point pairs; repeat the above random selection and counting process multiple times, select the matching point pair with the largest number of interior points as the valid matching point pair, and at the same time remove outlier matching point pairs that do not satisfy the transformation model.
[0054] S504, calculate the sum of squared pixel distances between each laser projection point and the corresponding matched visual edge feature point in the effective matching point pair, as the reprojection error value.
[0055] Suppose that after RANSAC filtering, n pairs of valid matching points are obtained. For the i-th pair of matching points, the points in the laser projection point set are... The points in the set of matched visual edge feature points are The reprojection error value for: in, Let be the pixel coordinates of the i-th point in the laser projection point set. Let n be the pixel coordinates of the point that successfully matches the i-th point, and n represent the number of matching point pairs in the valid matching point pairs. Reprojection error value. This describes the cumulative pixel distance between pixels projected from the laser-column edge point cloud onto the image plane and visual edge feature points in the visual edge feature point set. It reflects the geometric positional error between these pixels and the visual edge feature points. A smaller J value indicates better geometric alignment between the two types of edge features and higher calibration accuracy.
[0056] S60 iteratively optimizes the relative pose of the camera and the lidar using a nonlinear least squares method until the reprojection error is minimized. The relative pose corresponding to the minimum reprojection error is then taken as the optimal joint calibration result.
[0057] Specifically, the reprojection error value J constructed in step S50 is used as the cost function, which is a function of the rotation matrix R and the translation vector t. In this embodiment, a nonlinear least squares method (such as the Levenberg-Marquardt algorithm) is used to iteratively optimize and solve this cost function.
[0058] The optimization process uses the initial extrinsic parameters (rotation matrix R and translation vector t) from step S20 as the initial values for iteration. In each iteration, the correction values of the rotation matrix and translation vector are calculated based on the current reprojection error value, and R and t are corrected. Then, step S50 (i.e., reprojection, matching, and calculation of the reprojection error value) is re-executed. The above iterative process is repeated until the reprojection error value converges to below a preset threshold or the preset maximum number of iterations is reached.
[0059] When the iteration terminates, the rotation matrix R corresponding to the minimum reprojection error value is... ∗ Translation vector t ∗ The optimal joint calibration result is output. This optimal extrinsic parameter describes the final accurate spatial transformation relationship between the LiDAR coordinate system and the camera coordinate system, which can be used for subsequent autonomous driving tasks such as sensor data fusion, target detection, and tracking.
[0060] To better illustrate and understand the principle of the method provided by this invention, the following description uses an optional specific embodiment to illustrate the solution of this invention. It should be noted that the specific implementation of each step in this specific embodiment should not be construed as a limitation of the solution of this invention. Other implementations that can be conceived by those skilled in the art based on the principle of the solution provided by this invention should also be considered within the scope of protection of this invention.
[0061] In this embodiment, see Figure 2 The flowchart shown illustrates the method, which includes the following steps: Step 1: Input the laser columnar point cloud: Obtain the columnar point cloud of the vehicle in the vehicle body coordinate system at the current moment.
[0062] Specifically, the LiDAR point cloud of the vehicle is acquired, with each point corresponding to a semantic value that characterizes the point's category. After removing outliers, non-columnar points are filtered out based on their semantic values, retaining only the point cloud with the semantic category of columnar.
[0063] Step 2: Point cloud back-projection to image space: Based on the initial extrinsic parameters of the camera and LiDAR, the point cloud of the columnar object is projected to the image space to obtain a two-dimensional projected point cloud image.
[0064] The initial extrinsic parameters include a rotation matrix and a translation vector, which describe the initial spatial transformation relationship from the lidar coordinate system to the camera coordinate system. Specifically, using the camera imaging geometry model, each point in the columnar point cloud is projected onto the image plane to obtain the corresponding pixel coordinates, and all projected points form a two-dimensional projected point cloud image.
[0065] Step 3: Laser point cloud edge feature extraction: The two-dimensional projected point cloud image is binarized, and the edge points of the column are extracted through morphological operations to obtain the laser column edge point cloud.
[0066] Specifically, the two-dimensional projected point cloud image is binarized to generate a binary image. A rectangular sliding window of size 3 pixels × 3 pixels is used to traverse the binary image, and the maximum value of all pixels within the window is used as the output value of the center pixel to obtain the dilated image. Then, the dilated image is traversed, and the minimum value of all pixels within the window is used as the output value of the center pixel to obtain the eroded image. The gray values of pixels at the same position in the binary image and the eroded image are subtracted to obtain the difference image. Pixels with gray values greater than 0 in the difference image are extracted as edge points, and their corresponding laser point cloud indices are retained to obtain the edge point cloud of the laser columnar object.
[0067] Step 4: Visual image input, acquire vehicle visual image; Step 5: Visual line element feature extraction: Using the SURF feature extraction method combined with Hough transform, the straight line edge features in the image are extracted to obtain the visual edge feature point set.
[0068] Specifically, the vehicle visual image is converted to grayscale to obtain a grayscale image. Canny edge detection is then performed on the grayscale image to obtain a binary edge image. Hough transform is used to detect straight lines in the image, and the edge points along the lines are extracted as the Hough edge feature point set. Simultaneously, a scale space is constructed for the grayscale image, and extreme points are detected using the Hessian matrix as SURF feature points, generating a SURF feature point set. The Hough edge feature point set and the SURF feature point set are then fused, retaining the SURF feature points that simultaneously lie on the detected straight lines, as the visual edge feature point set.
[0069] Step 6: Edge matching and reprojection error calculation: Project the edge point cloud of the laser columnar object onto the image space, match it with the visual edge feature point set, and determine the reprojection error value based on the matching result.
[0070] Specifically, each point in the point cloud of the laser columnar object's edge is projected onto the image plane to obtain a laser projection point set. The pixel distance between each point in the laser projection point set and each point in the visual edge feature point set is calculated, and the closest pair of points is identified as a matching point pair. The RANSAC algorithm is used to filter all matching point pairs, removing outlier pairs and retaining inlier matching point pairs that satisfy the same projection transformation relationship as valid matching point pairs. The sum of the squared pixel distances between each laser projection point and the corresponding matched visual edge feature point in the valid matching point pair is calculated as the reprojection error value.
[0071] Step 7: Improved joint edge calibration: Iteratively optimize the relative pose of the camera and LiDAR using the nonlinear least squares method until the reprojection error is minimized. The relative pose corresponding to the minimum reprojection error is taken as the optimal joint calibration result.
[0072] Specifically, the reprojection error value is used as the cost function, and the initial extrinsic parameters are used as the initial values for iteration. A nonlinear least squares method is employed to iteratively optimize the rotation matrix and translation vector. After each iteration, the projection, matching, and error calculation in step 5 are re-executed until the reprojection error value converges or the maximum number of iterations is reached. The rotation matrix and translation vector corresponding to the minimum reprojection error value are output as the optimal joint calibration result.
[0073] The solution of the present invention has the following beneficial effects: 1. Using the semantic segmentation results of columnar objects in real-time vehicle-mounted lidar perception for point cloud feature extraction can effectively solve the problems of poor feature extraction robustness and low computational efficiency caused by limited laser reflection intensity and large, sparse and unstructured laser point cloud data. 2. A refined matching scheme using the edge points of columnar objects as key features and the edge feature points extracted from the image can effectively solve the problem that the geometric features of the calibration object are limited and cannot meet the requirements of online calibration, automated calibration or large-scale fleet calibration.
[0074] Based on and Figure 1 Using the same principle as the method shown, this embodiment of the invention also provides a camera and radar joint calibration device 20 for columnar semantic segmentation, such as... Figure 3 As shown, the camera and radar joint calibration device 20 for semantic segmentation of columnar objects may include an acquisition module 210, a two-dimensional projection module 220, an edge point cloud determination module 230, a visual edge feature point set determination module 240, a reprojection error value determination module 250, and a joint calibration result determination module 260, wherein: The acquisition module 210 is used to acquire the cylindrical point cloud of the vehicle in the vehicle body coordinate system at the current moment; The two-dimensional projection module 220 is used to project the cylindrical object point cloud onto the image space based on the initial extrinsic parameters of the camera and the lidar to obtain a two-dimensional projected point cloud image. The initial extrinsic parameters characterize the spatial transformation relationship between the lidar coordinate system and the camera coordinate system. The edge point cloud determination module 230 is used to perform binarization processing on the two-dimensional projected point cloud image and extract the edge points of the columnar object through morphological operations to obtain the edge point cloud of the laser columnar object. The visual edge feature point set determination module 240 is used to acquire vehicle visual images. It uses the SURF feature combined with Hough transform to extract straight line edge features in the vehicle visual images and obtain the visual edge feature point set. The reprojection error value determination module 250 is used to project the laser column edge point cloud onto the image space, match it with the visual edge feature point set, and determine the reprojection error value based on the matching result. The reprojection error value characterizes the geometric position error between the pixel point after the laser column edge point cloud is projected onto the image plane and the visual edge feature point in the visual edge feature point set. The joint calibration result determination module 260 is used to iteratively optimize the relative pose of the camera and the lidar using a nonlinear least squares method until the reprojection error value is minimized. The relative pose corresponding to the minimum reprojection error value is then taken as the optimal joint calibration result.
[0075] Optionally, when acquiring the cylindrical point cloud of the vehicle in the vehicle body coordinate system at the current moment, the acquisition module 210 is specifically used for: The LiDAR point cloud of the vehicle is acquired. Each point in the LiDAR point cloud corresponds to a semantic value. For each semantic value, the semantic value represents the category of the corresponding point. Based on the semantic values of all points in the LiDAR point cloud, all points in the LiDAR point cloud are filtered out, and the point cloud corresponding to the semantic value as columnar is retained, so as to obtain the columnar point cloud of the vehicle in the vehicle body coordinate system at the current time.
[0076] Optionally, when the edge point cloud determination module 230 performs binarization processing on the two-dimensional projected point cloud image and extracts the edge points of the columnar object through morphological operations to obtain the edge point cloud of the laser columnar object, it is specifically used for: Binarize the two-dimensional projected point cloud image to generate a binary image; Dilation and erosion operations are performed sequentially on the binary image. Edge points are obtained by the difference between the binary image and the eroded image. Based on the laser point cloud corresponding to the edge points, the edge point cloud of the laser column is obtained.
[0077] Optionally, when the edge point cloud determination module 230 performs dilation and erosion operations sequentially on the binary image and obtains edge points through the difference between the binary image and the eroded image, it is specifically used for: A rectangular sliding window with a size of 3 pixels × 3 pixels is used; The binary image is traversed using a rectangular sliding window to obtain the pixel corresponding to the maximum value among all pixels within the rectangular sliding window, resulting in the dilated image. The dilated image is then traversed using the same rectangular sliding window to obtain the pixel corresponding to the minimum value among all pixels within the rectangular sliding window, resulting in the eroded image. The grayscale values of pixels at the same location in the binary image and the eroded image are subtracted to obtain the difference image. Pixels with grayscale values greater than 0 in the difference image are extracted as the edge points of the columnar structures.
[0078] Optionally, when the aforementioned two-dimensional projection module 220 projects the columnar object point cloud onto the image space based on the initial extrinsic parameters of the camera and lidar to obtain a two-dimensional projected point cloud image, it is specifically used for: Based on the initial extrinsic parameters, each point in the columnar point cloud is projected into the image space to obtain a two-dimensional projected point cloud image.
[0079] Optionally, when the aforementioned visual edge feature point set determination module 240 extracts straight line edge features from the vehicle visual image using the SURF feature combined with Hough transform method to obtain the visual edge feature point set, it is specifically used for: The vehicle visual image is converted to grayscale to obtain a grayscale image; Perform Canny edge detection on the grayscale image to obtain a binary edge image; The Hough transform is used to detect straight lines in a binary edge image, and the edge points along the detected lines are extracted as the Hough edge feature point set. A scale space is constructed for the grayscale image, and extreme points are detected in the scale space using the Hessian matrix as SURF feature points, so as to generate a SURF feature point set based on all SURF feature points. The Hough edge feature point set and the SURF feature point set are fused, and the SURF feature points that are simultaneously located on the detected lines are retained as the visual edge feature point set.
[0080] Optionally, when the reprojection error value determination module 250 projects the laser columnar object edge point cloud onto the image space, matches it with the visual edge feature point set, and determines the reprojection error value based on the matching result, it is specifically used for: Project each point in the laser columnar object edge point cloud onto the image plane to obtain the laser projection point set; calculate the pixel distance between each point in the laser projection point set and each point in the visual edge feature point set, and determine the closest pair of points as a matching point pair; The RANSAC algorithm is used to filter all matching point pairs, outlier matching point pairs are removed, and interior point matching point pairs that satisfy the same projection transformation relationship are retained as valid matching point pairs. The sum of squared pixel distances between each laser projection point and the corresponding matched visual edge feature point in the effective matching point pair is calculated as the reprojection error value.
[0081] The camera and radar joint calibration device for columnar semantic segmentation in this embodiment of the invention can execute the camera and radar joint calibration method for columnar semantic segmentation provided in this embodiment of the invention. The implementation principle is similar. The actions performed by each module and unit in the camera and radar joint calibration device for columnar semantic segmentation in each embodiment of the invention correspond to the steps in the camera and radar joint calibration method for columnar semantic segmentation in each embodiment of the invention. For detailed functional descriptions of each module of the camera and radar joint calibration device for columnar semantic segmentation, please refer to the descriptions in the corresponding camera and radar joint calibration methods for columnar semantic segmentation shown above, which will not be repeated here.
[0082] The aforementioned camera and radar joint calibration device for columnar semantic segmentation can be a computer program (including program code) running on a computer device. For example, the camera and radar joint calibration device for columnar semantic segmentation is an application software. The device can be used to execute the corresponding steps in the method provided in the embodiments of the present invention.
[0083] In some embodiments, the camera and radar joint calibration device for columnar semantic segmentation provided in this invention can be implemented using a combination of hardware and software. As an example, the camera and radar joint calibration device for columnar semantic segmentation provided in this invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the camera and radar joint calibration method for columnar semantic segmentation provided in this invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0084] In other embodiments, the camera and radar joint calibration device for columnar semantic segmentation provided in this invention can be implemented in software. Figure 3A camera and radar joint calibration device for columnar semantic segmentation stored in memory is shown. It can be software in the form of programs and plug-ins, and includes a series of modules, including an acquisition module 210, a two-dimensional projection module 220, an edge point cloud determination module 230, a visual edge feature point set determination module 240, a reprojection error value determination module 250, and a joint calibration result determination module 260, for implementing the camera and radar joint calibration method for columnar semantic segmentation provided in the embodiments of the present invention.
[0085] The modules described in the embodiments of the present invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.
[0086] Based on the same principles as the methods shown in the embodiments of the present invention, the embodiments of the present invention also provide an electronic device, which may include, but is not limited to: a processor and a memory; the memory for storing computer programs; and the processor for executing the methods shown in any embodiment of the present invention by invoking the computer programs.
[0087] In one alternative embodiment, an electronic device is provided, such as Figure 4 As shown, Figure 4 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.
[0088] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0089] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0090] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0091] The memory 4003 stores application code (computer program) for executing the present invention, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0092] Among these, electronic devices can also be terminal devices. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention.
[0093] This invention provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.
[0094] According to another aspect of the present invention, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.
[0095] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0096] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0097] The computer-readable storage medium provided in this invention can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0098] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.
[0099] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.
Claims
1. A joint camera and radar calibration method for semantic segmentation of columnar objects, characterized in that, include: Obtain the cylindrical point cloud of the vehicle in the vehicle body coordinate system at the current moment; Based on the initial extrinsic parameters of the camera and lidar, the columnar point cloud is projected onto the image space to obtain a two-dimensional projected point cloud image. The initial extrinsic parameters characterize the spatial transformation relationship between the lidar coordinate system and the camera coordinate system. The two-dimensional projected point cloud image is binarized, and the edge points of the column are extracted by morphological operations to obtain the laser column edge point cloud. A vehicle visual image is acquired, and the straight line edge features in the vehicle visual image are extracted by using SURF features combined with Hough transform to obtain a set of visual edge feature points. The laser column edge point cloud is projected onto the image space and matched with the visual edge feature point set. Based on the matching result, the reprojection error value is determined. The reprojection error value characterizes the geometric position error between the pixel point after the laser column edge point cloud is projected onto the image plane and the visual edge feature point in the visual edge feature point set. The relative pose of the camera and the lidar is iteratively optimized using a nonlinear least squares method until the reprojection error value is minimized. The relative pose corresponding to the minimum reprojection error value is then taken as the optimal joint calibration result.
2. The method according to claim 1, characterized in that, The acquisition of the cylindrical point cloud of the vehicle in the vehicle body coordinate system at the current moment includes: Acquire the LiDAR point cloud of the vehicle, where each point in the LiDAR point cloud corresponds to a semantic value, and for each semantic value, the semantic value represents the category of the corresponding point; Based on the semantic values of all points in the lidar point cloud, all points in the lidar point cloud are filtered out, and the point cloud with the category of columnar object corresponding to the semantic value is retained to obtain the columnar object point cloud of the vehicle in the vehicle body coordinate system at the current time.
3. The method according to claim 1, characterized in that, The step of binarizing the two-dimensional projected point cloud image and extracting the edge points of the columnar object through morphological operations to obtain the laser columnar object edge point cloud includes: Binarize the two-dimensional projected point cloud image to generate a binary image; Dilation and erosion operations are performed sequentially on the binary image. Edge points are obtained by the difference between the binary image and the eroded image. Based on the laser point cloud corresponding to the edge points, the edge point cloud of the laser column is obtained.
4. The method according to claim 3, characterized in that, The step of sequentially performing dilation and erosion operations on the binary image, and obtaining edge points through the difference between the binary image and the eroded image, includes: A rectangular sliding window with a size of 3 pixels × 3 pixels is used; The binary image is traversed using the rectangular sliding window to obtain the pixel corresponding to the maximum value among all pixels within the rectangular sliding window, resulting in a dilated image. The dilated image is then traversed using the rectangular sliding window to obtain the pixel corresponding to the minimum value among all pixels within the rectangular sliding window, resulting in an eroded image. The grayscale values of pixels at the same position in the binary image and the eroded image are subtracted to obtain a difference image. Pixels with grayscale values greater than 0 in the difference image are extracted as edge points of the columnar structure.
5. The method according to any one of claims 1 to 4, characterized in that, The process of projecting the columnar point cloud onto the image space based on the initial extrinsic parameters of the camera and lidar to obtain a two-dimensional projected point cloud image includes: Based on the initial extrinsic parameters, each point in the columnar point cloud is projected into the image space to obtain a two-dimensional projected point cloud image.
6. The method according to any one of claims 1 to 4, characterized in that, The method employing SURF features combined with Hough transform extracts straight line edge features from the vehicle visual image, obtaining a set of visual edge feature points, including: The vehicle visual image is converted to grayscale to obtain a grayscale image; Perform Canny edge detection on the grayscale image to obtain a binary edge image; The Hough transform is used to detect straight lines in the binary edge image, and the edge points along the detected straight lines are extracted as the Hough edge feature point set. A scale space is constructed for the grayscale image, and extreme points in the scale space are detected by the Hessian matrix as SURF feature points, so as to generate a SURF feature point set based on all SURF feature points. The Hough edge feature point set and the SURF feature point set are fused, and the SURF feature points that are simultaneously located on the detected straight lines are retained as the visual edge feature point set.
7. The method according to any one of claims 1 to 4, characterized in that, The step of projecting the point cloud of the laser columnar object's edge onto the image space, matching it with the visual edge feature point set, and determining the reprojection error value based on the matching result includes: Project each point in the laser columnar object edge point cloud onto the image plane to obtain a laser projection point set; calculate the pixel distance between each point in the laser projection point set and each point in the visual edge feature point set, and determine the closest pair of points as a matching point pair; The RANSAC algorithm is used to filter all matching point pairs, outlier matching point pairs are removed, and interior point matching point pairs that satisfy the same projection transformation relationship are retained as valid matching point pairs. The sum of squared pixel distances between each laser projection point and the corresponding matched visual edge feature point in the effective matching point pair is calculated as the reprojection error value.
8. A camera and radar joint calibration device for columnar semantic segmentation, characterized in that, include: The acquisition module is used to acquire the cylindrical point cloud of the vehicle in the vehicle body coordinate system at the current moment; The two-dimensional projection module is used to project the columnar point cloud onto the image space based on the initial extrinsic parameters of the camera and the lidar, so as to obtain a two-dimensional projected point cloud image. The initial extrinsic parameters characterize the spatial transformation relationship between the lidar coordinate system and the camera coordinate system. The edge point cloud determination module is used to perform binarization processing on the two-dimensional projected point cloud image and extract the edge points of the columnar object through morphological operations to obtain the edge point cloud of the laser columnar object. The visual edge feature point set determination module is used to acquire vehicle visual images and extract straight line edge features from the vehicle visual images by using SURF features combined with Hough transform to obtain the visual edge feature point set. The reprojection error value determination module is used to project the laser column edge point cloud onto the image space, match it with the visual edge feature point set, and determine the reprojection error value based on the matching result. The reprojection error value characterizes the geometric position error between the pixel points after the laser column edge point cloud is projected onto the image plane and the visual edge feature points in the visual edge feature point set. The joint calibration result determination module is used to iteratively optimize the relative pose of the camera and the lidar using a nonlinear least squares method until the reprojection error value is minimized. The relative pose corresponding to the minimum reprojection error value is then taken as the optimal joint calibration result.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-7.