A Detection Method and Device for the Fusion of LiDAR and Camera
By using time-synchronized data and calibration parameters in lidar and camera fusion detection, combined with point cloud detection and image detection, the deep fusion of lidar and camera is achieved, solving the problem of large errors in the detection results in the prior art and improving detection accuracy.
Patent Information
- Application Number
- CN202210709758.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-06-22
AI Technical Summary
The existing lidar and camera fusion detection methods have shortcomings in data fusion and accuracy of detection results, resulting in large errors in fusion detection results.
By obtaining time-synchronized lidar point cloud and camera images, as well as calibration parameters from lidar coordinate system to camera image pixel coordinate system, point cloud detection and image detection are performed, and matching and correction are combined with 2D bounding boxes to achieve deep fusion between lidar and camera.
The accuracy of the fusion detection results of lidar and cameras is improved, and the ability to obtain three-dimensional information for road targets and traffic targets is enhanced.
Smart Images

Figure CN114937255B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent driving perception, and more specifically, to a detection method and device for the fusion of lidar and camera. Background Art
[0002] When an intelligent driving vehicle performs environmental perception, since the target detection effect of a single sensor is limited, the detection results of multiple sensors are usually integrated to obtain more reliable perception information. Currently, the combination of lidar and camera is relatively common.
[0003] There are mainly two existing detection methods for the fusion of lidar and camera: one is pre-fusion, where the data information of lidar and camera is directly fused at the raw data level, and then the fused data is processed by a perception algorithm. However, the current pre-fusion neural networks (such as MV3D, AVOD, F-PointNet, etc.) have low detection accuracy and it is difficult to extract all the required information from the fused data. The other is post-fusion, where lidar and camera are detected separately, and then the detection results are fused. Since lidar and camera are detected as independent single sensors during their respective detection processes and there is no data interaction between them, there are certain errors in the respective detection results of lidar and camera, resulting in a large error in the final fused detection result. Summary of the Invention
[0004] In view of this, the present invention discloses a detection method and device for the fusion of lidar and camera to achieve the deep fusion of lidar and camera and improve the accuracy of the fused detection results of lidar and camera.
[0005] A detection method for the fusion of lidar and camera includes:
[0006] Obtain time-synchronized lidar point cloud and camera image, as well as the calibration parameters from the lidar coordinate system to the camera image pixel coordinate system, where the camera image pixel coordinate system is the pixel coordinate system where the camera image is located;
[0007] Perform point cloud detection on the lidar point cloud to obtain a preliminary detection result of road targets including category and heading angle;
[0008] Project the preliminary detection result of road targets into the camera image pixel coordinate system through the calibration parameters to obtain the 2D bounding box of the radar detection target in the camera image pixel coordinate system;
[0009] Perform target detection on the camera image in combination with the 2D bounding box of the radar detection target to obtain the 2D bounding box of the camera detection target in the camera image pixel coordinate system;
[0010] Using the camera to detect the 2D bounding box of the target, matching it with the 2D bounding box detected by the radar and correcting the category to obtain the corrected 3D detected road target of the lidar point cloud;
[0011] Based on the depth map corresponding to the lidar point cloud, back-projecting the depth points of the traffic target within the 2D bounding box detected by the camera to obtain the 3D information of the traffic target;
[0012] Combining the 3D detected road target of the lidar point cloud and the 3D information of the traffic target to obtain the target detection result of the lidar and camera fusion.
[0013] Optionally, the lidar point cloud is subjected to point cloud detection to obtain a preliminary detection result of the road target including the category and heading angle, including:
[0014] Performing ground point segmentation on the lidar point cloud to obtain non-ground point cloud;
[0015] Clustering the non-ground point cloud to obtain the coordinate positions and size information of each cluster of point cloud;
[0016] Downsampling the coordinate positions and the size information of each cluster of point cloud and inputting them into the PointNet neural network to obtain the preliminary detection result of the road target.
[0017] Optionally, the preliminary detection result of the road target is projected into the camera pixel coordinate system through the calibration parameters to obtain the 2D bounding box of the radar-detected target in the camera image pixel coordinate system, including:
[0018] Projecting the preliminary detection result of the road target into the camera image pixel coordinate system through the calibration parameters to obtain each initial 2D bounding box of the radar-detected target;
[0019] Calculating the average depth of each of the initial 2D bounding boxes;
[0020] Calculating the intersection over union of any two of all the initial 2D bounding boxes;
[0021] If there is a target intersection over union greater than the set threshold among all the intersections over union, then filtering out the initial 2D bounding box with the larger average depth among the two initial 2D bounding boxes corresponding to the target intersection over union to obtain the non-occluded 2D bounding box in the camera image pixel coordinate system, and determining the non-occluded 2D bounding box as the 2D bounding box of the radar-detected target.
[0022] Optionally, performing object detection on the camera image in combination with the 2D bounding box of the radar-detected object to obtain the 2D bounding box of the camera-detected object in the pixel coordinate system of the camera image, including:
[0023] Performing object detection on the camera image using a region proposal network in combination with the 2D bounding box of the radar-detected object to obtain a camera image detection proposal box;
[0024] Inputting the camera image detection proposal box and the non-occluded 2D bounding box into a region of interest network for classification and regression to obtain the 2D bounding box of the camera-detected object.
[0025] Optionally, matching and correcting the category of the 2D bounding box of the radar-detected object using the 2D bounding box of the camera-detected object to obtain a corrected 3D detected road object of the lidar point cloud, including:
[0026] Calculating the intersection over union of the 2D bounding box of the camera-detected object and the 2D bounding box of the radar-detected object;
[0027] Based on the intersection over union, performing optimal matching using the Hungarian matching algorithm, performing probability fusion on the road object category for the matched 3D detected point cloud objects using a probability fusion algorithm, and increasing the probability of the obstacle category for the unmatched 3D detected point cloud objects to obtain the corrected 3D detected road object of the lidar point cloud.
[0028] Optionally, based on the depth map corresponding to the lidar point cloud, back-projecting the depth points of the traffic object within the 2D bounding box of the camera-detected object to obtain the 3D information of the traffic object, including:
[0029] Based on the depth map corresponding to the lidar point cloud and the calibration parameters, back-projecting the depth points of the traffic object within the 2D bounding box of the camera-detected object to obtain the original 3D information of the traffic object;
[0030] Judging whether there are object 3D points in the original traffic 3D information that do not meet the preset position requirements;
[0031] If so, filtering out the object 3D points and obtaining the 3D information of the traffic object according to the remaining 3D points.
[0032] Optionally, the determination process of the depth map corresponding to the lidar point cloud includes:
[0033] Projecting the lidar point cloud into the coordinate system of the camera image through the calibration parameters to obtain an original depth map;
[0034] Complement and fill the original depth map in ascending order of holes to obtain an intermediate depth map;
[0035] Reduce the output noise and smooth the local plane of the intermediate depth map to obtain the depth map corresponding to the lidar point cloud.
[0036] A detection device for lidar and camera fusion, comprising:
[0037] An acquisition unit for acquiring time-synchronized lidar point clouds and camera images, as well as calibration parameters from the lidar coordinate system to the camera image pixel coordinate system, where the camera image pixel coordinate system is the pixel coordinate system where the camera image is located;
[0038] A first detection unit for performing point cloud detection on the lidar point cloud to obtain a preliminary detection result of road targets including categories and heading angles;
[0039] A projection unit for projecting the preliminary detection result of the road target into the camera image pixel coordinate system through the calibration parameters to obtain a 2D bounding box of the radar detection target in the camera image pixel coordinate system;
[0040] A second detection unit for performing target detection on the camera image in combination with the 2D bounding box of the radar detection target to obtain a 2D bounding box of the camera detection target in the camera image pixel coordinate system;
[0041] A correction unit for matching and correcting the categories of the 2D bounding box of the radar detection target by using the 2D bounding box of the camera detection target to obtain a corrected 3D detection road target of the lidar point cloud;
[0042] A back-projection unit for back-projecting the depth points of traffic targets within the 2D bounding box of the camera detection target based on the depth map corresponding to the lidar point cloud to obtain 3D information of traffic targets;
[0043] A result fusion unit for merging the 3D detection road target of the lidar point cloud and the 3D information of traffic targets to obtain a target detection result of lidar and camera fusion.
[0044] Optionally, the first detection unit includes:
[0045] A segmentation sub-unit for segmenting ground points from the lidar point cloud to obtain non-ground point clouds;
[0046] A clustering sub-unit for clustering the non-ground point clouds to obtain the coordinate positions and size information of each cluster of point clouds;
[0047] The target detection subunit is configured to downsample the coordinate positions and the dimension information of each cluster of point clouds and input them into a PointNet neural network to obtain a preliminary detection result of the road target.
[0048] Optionally, the projection unit includes:
[0049] The projection subunit is configured to project the preliminary detection result of the road target into the camera image pixel coordinate system through the calibration parameters to obtain each initial 2D bounding box of the radar detection target;
[0050] The depth calculation subunit is configured to calculate the average depth of each of the initial 2D bounding boxes;
[0051] The first intersection over union (IoU) calculation subunit is configured to calculate the IoU between any two of all the initial 2D bounding boxes;
[0052] The bounding box determination subunit is configured to, if there is a target IoU greater than a set threshold among all the IoUs, filter out the initial 2D bounding box with a larger average depth among the two initial 2D bounding boxes corresponding to the target IoU to obtain an unoccluded 2D bounding box in the camera image pixel coordinate system, and determine the unoccluded 2D bounding box as the 2D bounding box of the radar detection target.
[0053] As can be seen from the above technical solutions, the present invention discloses a detection method and device for lidar-camera fusion, which acquires time-synchronized lidar point clouds and camera images, as well as calibration parameters from the lidar coordinate system to the camera image pixel coordinate system, performs point cloud detection on the lidar point clouds to obtain a preliminary detection result of the road target including the category and heading angle, projects the preliminary detection result of the road target into the camera image pixel coordinate system through the calibration parameters to obtain the 2D bounding box of the radar detection target in the camera image pixel coordinate system, performs target detection on the camera image in combination with the 2D bounding box of the radar detection target to obtain the 2D bounding box of the camera detection target in the camera image pixel coordinate system, matches and corrects the category of the 2D bounding box of the radar detection target by using the 2D bounding box of the camera detection target to obtain a corrected 3D detection road target of the lidar point clouds, projects the depth points of the traffic target within the 2D bounding box of the camera detection target based on the depth map corresponding to the lidar point clouds to obtain the 3D information of the traffic target, and combines the 3D detection road target of the lidar point clouds and the 3D information of the traffic target to obtain a lidar-camera fusion target detection result. The present invention not only fuses the lidar point clouds and camera images at the raw data level, but also fuses the camera image target detection result and the lidar point cloud detection result, thereby realizing the deep fusion of the lidar and the camera, and greatly improving the accuracy of the lidar-camera fusion detection result. Brief Description of the Drawings
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the disclosed drawings without creative efforts.
[0055] Figure 1 It is a flowchart of a detection method for the fusion of a lidar and a camera disclosed in an embodiment of the present invention;
[0056] Figure 2 It is a flowchart of a method for performing point cloud detection on lidar point cloud to obtain a preliminary detection result of a road target including category and heading angle disclosed in an embodiment of the present invention;
[0057] Figure 3 It is a detection flowchart of a PointNet neural network disclosed in an embodiment of the present invention;
[0058] Figure 4 It is a flowchart of a method for determining a 2D bounding box of a radar detection target in the pixel coordinate system of a camera image disclosed in an embodiment of the present invention;
[0059] Figure 5 It is a schematic diagram of the working principle of a traditional Faster RCNN network;
[0060] Figure 6 It is a flowchart of a method for determining three-dimensional information of a traffic target disclosed in an embodiment of the present invention;
[0061] Figure 7 It is a schematic structural diagram of a detection device for the fusion of a lidar and a camera disclosed in an embodiment of the present invention. Detailed Embodiments
[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0063] An embodiment of the present invention discloses a detection method and device for lidar-camera fusion, which not only fuses lidar point clouds and camera images at the raw data level, but also fuses the camera image object detection results and lidar point cloud detection results, thereby achieving deep fusion of lidar and camera, and greatly improving the accuracy of the fusion detection results of lidar and camera. In addition, by fusing lidar and camera at the raw data level, the present invention can respectively improve the object detection accuracy of lidar point clouds and camera images, and at the same time, the three-dimensional information of the camera image detection target can be obtained from the depth map corresponding to the lidar point cloud, especially the detection of traffic targets can be realized.
[0064] See Figure 1 , a flowchart of a detection method for lidar-camera fusion disclosed in an embodiment of the present invention, the method includes:
[0065] Step S101, obtain time-synchronized lidar point clouds and camera images, and calibration parameters from the lidar coordinate system to the camera image pixel coordinate system;
[0066] The camera image pixel coordinate system in this embodiment is the pixel coordinate system where the camera image is located;
[0067] Among them, the time synchronization of lidar point clouds and camera images means that the time difference between lidar point clouds and camera images is within a small range.
[0068] The determination process of the calibration parameters from the lidar coordinate system to the camera image pixel coordinate system can refer to existing mature solutions and will not be elaborated here.
[0069] Step S102, perform point cloud detection on the lidar point cloud to obtain a preliminary detection result of road targets including category and heading angle;
[0070] In practical applications, lidar detection methods such as PointNet neural network and PointPillar can be used to perform point cloud detection on the lidar point cloud to obtain a preliminary detection result of road targets including category and heading angle. The category is the type of object that lidar can detect, such as road targets (hereinafter simply referred to as targets when not necessary) such as cars, buses, bicycles, pedestrians, and tricycles.
[0071] Step S103, project the preliminary detection result of road targets into the camera image pixel coordinate system through the calibration parameters to obtain a 2D bounding box of the radar detection target in the camera image pixel coordinate system;
[0072] Among them, the radar detection target is a road target.
[0073] In practical applications, the eight vertices of the preliminary detection result of the road target can be projected into the pixel coordinate system of the camera image through calibration parameters, and the 2D bounding box of the radar detection target in the pixel coordinate system of the camera image can be obtained according to the positions of the eight projection points.
[0074] Step S104: Perform object detection on the camera image in combination with the 2D bounding box of the radar detection target to obtain the 2D bounding box of the camera detection target in the pixel coordinate system of the camera image;
[0075] Among them, the camera detection target includes a road target and a traffic target.
[0076] In practical applications, an improved Faster RCNN network can be used to perform object detection on the camera image. The improved Faster RCNN network is divided into two stages. In the first stage, a detection proposal box of the camera image where a target may exist is obtained. In the second stage, the detection proposal box of the camera image is detected to obtain the final 2D bounding box of the camera detection target. In practical applications, the category of the camera detection target can also be obtained.
[0077] Step S105: Match the 2D bounding box of the radar detection target with the 2D bounding box of the camera detection target and correct the category to obtain the corrected 3D detected road target of the lidar point cloud;
[0078] In this embodiment, the 2D bounding box of the camera detection target can be used as the object detection result of the camera image, and the 2D bounding box of the radar detection target can be used as the preliminary detection result of the lidar point cloud. Matching the 2D bounding box of the radar detection target with the 2D bounding box of the camera detection target and correcting the category is actually using the object detection result of the camera image to correct the category probability of the lidar point cloud road target detection result to obtain the corrected 3D detected road target of the lidar point cloud.
[0079] Step S106: Based on the depth map corresponding to the lidar point cloud, back-project the depth points of the traffic target within the 2D bounding box of the camera detection target to obtain the 3D information of the traffic target;
[0080] In this embodiment, the depth points are pixel coordinate points of the traffic target in the depth map with depth information.
[0081] Among them, the traffic target includes: traffic lights and traffic signs.
[0082] The specific steps for back-projecting to obtain the 3D information of the traffic target in the present invention are as follows:
[0083] Multiply the pixel coordinate points of the traffic target on the depth map by the inverse matrix of the calibration parameters to obtain the original 3D point cloud of the traffic target;
[0084] Determine whether there are target three-dimensional points in the original traffic target three-dimensional point cloud that do not meet the preset position requirements;
[0085] If so, filter out the target three-dimensional points, and obtain the three-dimensional information of the traffic target according to the remaining three-dimensional point cloud.
[0086] Step S107: Merge the three-dimensional detection road target of the lidar point cloud and the three-dimensional information of the traffic target to obtain the target detection result of the lidar and camera fusion.
[0087] Among them, merging the three-dimensional detection road target of the lidar point cloud and the three-dimensional information of the traffic target means adding the three-dimensional information of the traffic target to the three-dimensional detection road target of the lidar point cloud, so as to obtain the target detection result of the lidar and camera fusion.
[0088] In summary, the present invention discloses a detection method for lidar and camera fusion, which obtains time-synchronized lidar point cloud and camera image, as well as calibration parameters from the lidar coordinate system to the camera image pixel coordinate system. Perform point cloud detection on the lidar point cloud to obtain a preliminary detection result of the road target including the category and heading angle, project the preliminary detection result of the road target into the camera image pixel coordinate system through the calibration parameters to obtain the 2D bounding box of the radar detection target in the camera image pixel coordinate system, perform target detection on the camera image in combination with the 2D bounding box of the radar detection target to obtain the 2D bounding box of the camera detection target in the camera image pixel coordinate system, use the 2D bounding box of the camera detection target to match and correct the category of the 2D bounding box of the radar detection target to obtain the corrected three-dimensional detection road target of the lidar point cloud, based on the depth map corresponding to the lidar point cloud, back-project the depth points of the traffic target within the 2D bounding box of the camera detection target to obtain the three-dimensional information of the traffic target, and merge the three-dimensional detection road target of the lidar point cloud and the three-dimensional information of the traffic target to obtain the target detection result of the lidar and camera fusion. The present invention not only fuses the lidar point cloud and the camera image at the original data level, but also fuses the camera image target detection result and the lidar point cloud detection result, thus realizing the deep fusion of the lidar and the camera, and greatly improving the accuracy of the fusion detection result of the lidar and the camera.
[0089] In addition, the present invention fuses the lidar and the camera at the original data level, which can respectively improve the target detection accuracy of the lidar point cloud and the camera image, and at the same time, the three-dimensional information of the camera image detection target can be obtained from the depth map corresponding to the lidar point cloud, especially the detection of traffic targets can be realized.
[0090] To further optimize the above embodiments, see Figure 2, A method flowchart for detecting lidar point cloud to obtain a preliminary detection result of road targets including categories and heading angles, disclosed in an embodiment of the present invention, the method includes:
[0091] Step S201, perform ground point segmentation on the lidar point cloud to obtain non-ground point cloud;
[0092] In the present invention, ground point cloud and non-ground point cloud can be obtained by performing ground segmentation on the lidar point cloud. In this embodiment, the ground point cloud will be filtered out, and the non-ground point cloud will be clustered.
[0093] Step S202, cluster the non-ground point cloud to obtain the coordinate positions and size information of each cluster of point cloud;
[0094] Step S203, downsample the coordinate positions and size information of each cluster of point cloud and input them into the PointNet neural network to obtain a preliminary detection result of road targets.
[0095] Among them, the detection flowchart of the PointNet neural network is as Figure 3 shown. First, centralize the coordinate of each cluster of point cloud and rotate it to the front of the vehicle to make the heading angle more evenly distributed, then perform point cloud feature extraction. The extracted point cloud features pass through the pooling layer to obtain comprehensive features, and then identify the category and heading angle information to obtain a preliminary detection result of road targets.
[0096] To further optimize the above embodiment, refer to Figure 4 A method flowchart for determining the 2D bounding box of radar detection targets in the pixel coordinate system of camera images, disclosed in an embodiment of the present invention, that is, step S103 includes:
[0097] Step S301, project the preliminary detection result of road targets into the pixel coordinate system of camera images through calibration parameters to obtain each initial 2D bounding box of radar detection targets;
[0098] Specifically, project the eight vertices of the preliminary detection result of road targets into the pixel coordinate system of camera images, and obtain each initial 2D bounding box of radar detection targets according to the positions of the eight projected points.
[0099] Step S302, calculate the average depth of each initial 2D bounding box;
[0100] Among them, the calculation process of the average depth of the initial 2D bounding box can refer to existing mature solutions and will not be elaborated here.
[0101] Step S303, calculate the intersection over union (IoU) of any two initial 2D bounding boxes among all the initial 2D bounding boxes;
[0102] Intersection-over-Union (IoU) is a concept used in object detection. It is the overlap rate between the generated candidate bounding boxes and the original ground truth bounding boxes, that is, the ratio of their intersection to their union. In this embodiment, the IoU of any two initial 2D bounding boxes is calculated.
[0103] Step S304: If there is a target IoU greater than the set threshold among all the IoUs, filter out the initial 2D bounding box with a larger average depth among the two initial 2D bounding boxes corresponding to the target IoU, obtain the non-occluded 2D bounding box in the pixel coordinate system of the camera image, and determine the non-occluded 2D bounding box as the 2D bounding box of the radar detection target.
[0104] It should be noted that the initial 2D bounding box with a larger average depth among the two initial 2D bounding boxes corresponding to the target IoU is the occluded 2D bounding box. In this embodiment, the occluded 2D bounding box is found and filtered out from all the initial 2D bounding boxes based on the IoU, so as to obtain the non-occluded 2D bounding box, that is, the 2D bounding box of the radar detection target.
[0105] To further optimize the above embodiment, step S104 may specifically include:
[0106] Use the Region Proposal Network to perform object detection on the camera image in combination with the 2D bounding box of the radar detection target to obtain the detection proposal boxes of the camera image;
[0107] Input the detection proposal boxes of the camera image and the non-occluded 2D bounding box into the Region of Interest Network for classification and regression, and then obtain the 2D bounding box of the camera detection target.
[0108] In this embodiment, the improved Faster RCNN network is used to perform object detection on the camera image.
[0109] For ease of understanding the working principle of the improved Faster RCNN network, refer to Figure 5 the schematic diagram of the working principle of the traditional Faster RCNN network shown in. First, use the Feature Pyramid Network (FPN) to extract features. FPN is a general feature extractor, and through the top-down process and the lateral connection structure (see Figure 5High-level semantic features of various sizes are constructed using C1 to C5 and P2 to P5, which can better handle the multi-scale variation problem in object detection. At each pixel point on the feature map extracted by the FPN network, anchor boxes of different sizes are set. The RPN (Region Proposal Network) is used for further feature extraction, and the foreground and background classification and size and position regression of the anchor boxes are performed in the first stage. The anchor boxes with the classification result of background are deleted, and the foreground anchor boxes after regression are filtered by IOU to obtain the proposed boxes where objects may exist. Then, the proposed boxes are sent to the ROI (Region of Interest) network for the second-stage classification and regression. In the ROI network, the ROI Align layer integrates the feature maps of proposed boxes of different sizes into the same size through bilinear interpolation, and then the features are combined through the fully connected layer (FC) to predict the classification and position and size information of the target.
[0110] The improved Faster RCNN network in the present invention is divided into two stages. In the first stage, the proposed boxes for camera image detection where the target may exist are obtained. In the second stage, the proposed boxes for camera image detection are detected to obtain the 2D bounding box and category of the final camera detection target. Compared with the traditional Faster RCNN network, on the basis of obtaining the proposed boxes for camera image detection using the RPN (Region Proposal Network) in the first stage, the improved Faster RCNN network inputs them together with the non-occluded 2D bounding boxes into the ROI (Region of Interest) network for the second-stage classification and regression to increase the accuracy of the input to the ROI network, thereby improving the accuracy of the camera image target detection result.
[0111] To further optimize the above embodiment, step S105 may specifically include:
[0112] Calculate the intersection over union of the 2D bounding box of the camera detection target and the 2D bounding box of the radar detection target;
[0113] Based on the intersection over union, use the Hungarian matching algorithm for optimal matching. For the point cloud three-dimensional detection targets that are matched, use the probability fusion algorithm to perform probability fusion on the road target categories, and increase the probability of the obstacle category for the point cloud three-dimensional detection targets that are not matched to obtain the corrected lidar point cloud three-dimensional detection road targets.
[0114] Among them, the probability fusion algorithm can be the evidence theory, etc.
[0115] This embodiment mainly uses the camera target detection results (i.e., the 2D bounding boxes of the camera-detected targets) to correct the class probabilities of the initial lidar point cloud detection results (i.e., the 2D bounding boxes of the lidar-detected targets).
[0116] To further optimize the above embodiment, refer to Figure 6 , the flowchart of a method for determining three-dimensional information of traffic targets disclosed in an embodiment of the present invention, that is, step S106 includes:
[0117] Step S401: Based on the depth map corresponding to the lidar point cloud and the calibration parameters, back-project the depth points of the traffic targets within the 2D bounding boxes of the camera-detected targets to obtain the original three-dimensional information of the traffic targets;
[0118] Step S402: Determine whether there are target three-dimensional points in the original traffic three-dimensional information that do not meet the preset position requirements. If so, execute step S403;
[0119] Among them, when all the three-dimensional points in the original traffic three-dimensional information meet the preset position requirements, the original traffic three-dimensional information is directly determined as the traffic target three-dimensional information.
[0120] Step S403: Filter out the target three-dimensional points and obtain the traffic target three-dimensional information based on the remaining three-dimensional points.
[0121] It should be noted that the camera image has rich pixel information and can detect traffic targets such as traffic lights and traffic signs that cannot be detected in the lidar point cloud. Since the 2D bounding boxes of the camera-detected targets are obtained in step 104, only the 2D information of the traffic targets can be obtained. In this embodiment, the lidar point cloud and the calibration parameters from the lidar coordinate system to the camera image pixel coordinate system are used to back-project and obtain the traffic target three-dimensional information.
[0122] Among them, the determination process of the depth map corresponding to the lidar point cloud includes:
[0123] Project the lidar point cloud into the camera image pixel coordinate system through the calibration parameters to obtain the original depth map;
[0124] Fill in and complete the original depth map in the order of the holes from small to large to obtain the intermediate depth map;
[0125] Reduce the output noise and smooth the local plane of the intermediate depth map to obtain the depth map corresponding to the lidar point cloud.
[0126] It should be noted that when completing the original depth map, image processing methods or machine learning methods can be used. In this embodiment, a classic image fast completion method based on OpenCV is used. Utilizing the idea that null values around valid depths may have similar values, the order of filling small holes first and then large holes is adopted. Finally, the completed depth information is obtained by reducing output noise and smoothing local planes.
[0127] Corresponding to the above method embodiment, the present invention also discloses a detection device for lidar and camera fusion.
[0128] See Figure 7 , a schematic structural diagram of a detection device for lidar and camera fusion disclosed in an embodiment of the present invention. The device includes:
[0129] An acquisition unit 501, configured to acquire lidar point clouds and camera images with time synchronization, as well as calibration parameters from the lidar coordinate system to the camera image pixel coordinate system;
[0130] Among them, the time synchronization of the lidar point clouds and the camera images means that the time difference between the lidar point clouds and the camera images is within a small range.
[0131] A first detection unit 502, configured to perform point cloud detection on the lidar point clouds to obtain a preliminary detection result of road targets including categories and heading angles;
[0132] In practical applications, lidar detection methods such as PointNet neural network, PointPillar, etc. can be used to perform point cloud detection on the lidar point clouds to obtain a preliminary detection result of road targets including categories and heading angles.
[0133] A projection unit 503, configured to project the preliminary detection result of the road targets into the camera image pixel coordinate system through the calibration parameters to obtain a 2D bounding box of the radar detection target in the camera image pixel coordinate system;
[0134] In practical applications, the eight vertices of the preliminary detection result of the road targets can be projected into the camera image pixel coordinate system through the calibration parameters, and a 2D bounding box of the radar detection target in the camera image pixel coordinate system can be obtained according to the positions of the eight projection points.
[0135] A second detection unit 504, configured to perform target detection on the camera image in combination with the 2D bounding box of the radar detection target to obtain a 2D bounding box of the camera detection target in the camera image pixel coordinate system;
[0136] In practical applications, an improved Faster RCNN network can be used for object detection of camera images. The improved Faster RCNN network is divided into two stages. In the first stage, detection proposal boxes of camera images where objects may exist are obtained. In the second stage, the detection proposal boxes of camera images are detected to obtain the 2D bounding boxes of the final camera detection objects. In practical applications, the categories of camera detection objects can also be obtained.
[0137] A correction unit 505 is configured to match and correct the categories of the 2D bounding boxes of the lidar detection objects by using the 2D bounding boxes of the camera detection objects, so as to obtain corrected 3D detected road objects of the lidar point cloud.
[0138] In this embodiment, the 2D bounding boxes of camera detection objects can be used as the object detection results of camera images, and the 2D bounding boxes of lidar detection objects can be used as the preliminary detection results of lidar point clouds. Matching and correcting the categories of the 2D bounding boxes of lidar detection objects by using the 2D bounding boxes of camera detection objects is actually using the object detection results of camera images to correct the category probabilities of the preliminary detection results of lidar point clouds, so as to obtain corrected 3D detected road objects of the lidar point cloud.
[0139] A back-projection unit 506 is configured to back-project the depth points of traffic objects within the 2D bounding boxes of the camera detection objects based on the depth map corresponding to the lidar point cloud, so as to obtain 3D information of traffic objects.
[0140] A result fusion unit 507 is configured to merge the 3D detected road objects of the lidar point cloud and the 3D information of traffic objects to obtain the object detection results of lidar and camera fusion.
[0141] Among them, merging the 3D detected road objects of the lidar point cloud and the 3D information of traffic objects means adding the 3D information of traffic objects to the 3D detected road objects of the lidar point cloud, so as to obtain the object detection results of lidar and camera fusion.
[0142] In summary, the present invention discloses a detection device for fusing lidar and camera, which acquires lidar point clouds and camera images with time synchronization, as well as calibration parameters from the lidar coordinate system to the camera image pixel coordinate system. It performs point cloud detection on the lidar point clouds to obtain a preliminary detection result of road targets including categories and heading angles, projects the preliminary detection result of road targets into the camera image pixel coordinate system through the calibration parameters to obtain a 2D bounding box of the radar detection target in the camera image pixel coordinate system, performs target detection on the camera image in combination with the 2D bounding box of the radar detection target to obtain a 2D bounding box of the camera detection target in the camera image pixel coordinate system, matches and corrects the categories of the 2D bounding box of the radar detection target by using the 2D bounding box of the camera detection target to obtain a corrected 3D detection of road targets from the lidar point clouds. Based on the depth map corresponding to the lidar point clouds, the depth points of traffic targets within the 2D bounding box of the camera detection target are back-projected to obtain 3D information of traffic targets, and the 3D detection of road targets from the lidar point clouds and the 3D information of traffic targets are merged to obtain a target detection result of lidar and camera fusion. The present invention not only fuses lidar point clouds and camera images at the raw data level, but also fuses the camera image target detection result and the lidar point cloud detection result, thus realizing the deep fusion of lidar and camera and greatly improving the accuracy of the fusion detection result of lidar and camera.
[0143] In addition, the present invention fuses lidar and camera at the raw data level, which can respectively improve the target detection accuracy of lidar point clouds and camera images, and at the same time, 3D information of camera image detection targets can be obtained from the depth map corresponding to the lidar point clouds, especially the 3D detection of traffic targets can be realized.
[0144] To further optimize the above embodiment, the first detection unit 502 may include:
[0145] A segmentation sub-unit for segmenting ground points from the lidar point clouds to obtain non-ground point clouds;
[0146] A clustering sub-unit for clustering the non-ground point clouds to obtain the coordinate positions and size information of each cluster of point clouds;
[0147] A target detection sub-unit for downsampling the coordinate positions and the size information of each cluster of point clouds and inputting them into the PointNet neural network to obtain the preliminary detection result of the road targets.
[0148] To further optimize the above embodiment, the projection unit 503 may include:
[0149] A projection subunit, configured to project the preliminary detection result of the road target into the camera image according to the calibration parameters, so as to obtain respective initial 2D bounding boxes of the radar detection target;
[0150] A depth calculation subunit, configured to calculate the average depth of each of the initial 2D bounding boxes;
[0151] A first intersection over union (IoU) calculation subunit, configured to calculate the IoU between any two of all the initial 2D bounding boxes;
[0152] A bounding box determination subunit, configured to, if there is a target IoU greater than a set threshold among all the IoUs, filter out the initial 2D bounding box with a larger average depth among the two initial 2D bounding boxes corresponding to the target IoU, obtain an unoccluded 2D bounding box in the pixel coordinate system of the camera image, and determine the unoccluded 2D bounding box as the 2D bounding box of the radar detection target.
[0153] To further optimize the above embodiment, the second detection unit 504 may include:
[0154] A target detection subunit, configured to perform target detection on the camera image by using a region proposal network in combination with the 2D bounding box of the radar detection target, so as to obtain a camera image detection proposal box;
[0155] An input subunit, configured to input the camera image detection proposal box and the unoccluded 2D bounding box into a region of interest network for classification and regression, so as to obtain a 2D bounding box of the camera detection target.
[0156] To further optimize the above embodiment, the correction unit 505 may include:
[0157] A second IoU calculation subunit, configured to calculate the IoU between the 2D bounding box of the camera detection target and the 2D bounding box of the radar detection target;
[0158] A correction subunit, configured to perform optimal matching based on the IoU by using the Hungarian matching algorithm, perform probability fusion on the road target category for the point cloud three-dimensional detection target that is matched by using a probability fusion algorithm, and increase the probability of the obstacle category for the point cloud three-dimensional detection target that is not matched, so as to obtain the corrected lidar point cloud three-dimensional detection road target.
[0159] To further optimize the above embodiment, the back-projection unit 506 may include:
[0160] A back-projection subunit, configured to perform back-projection on the depth points of the traffic target within the 2D bounding box of the camera detection target based on the depth map corresponding to the lidar point cloud and the calibration parameters, so as to obtain the original traffic target three-dimensional information;
[0161] A judgment subunit, configured to judge whether there are target three-dimensional points in the original traffic three-dimensional information that do not meet the preset position requirements;
[0162] A filtering subunit, configured to, when the judgment subunit determines that it is the case, filter out the target three-dimensional points, and obtain the traffic target three-dimensional information according to the remaining three-dimensional points.
[0163] The detection device for lidar and camera fusion may further include: a depth map determination unit.
[0164] The depth map determination unit may specifically be configured to:
[0165] Project the lidar point cloud into the camera image through calibration parameters to obtain an original depth map;
[0166] Complement and fill the original depth map in the order of holes from small to large to obtain an intermediate depth map;
[0167] Reduce the output noise and smooth the local plane of the intermediate depth map to obtain the depth map corresponding to the lidar point cloud.
[0168] It should be particularly noted that for the specific working principles of the components in the device embodiments, please refer to the corresponding parts of the method embodiments, which will not be elaborated here.
[0169] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0170] The various embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0171] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A detection method for the fusion of lidar and camera, characterized in that, it includes: Obtain the lidar point cloud and camera image synchronized in time, as well as the calibration parameters from the lidar coordinate system to the camera image pixel coordinate system, where the camera image pixel coordinate system is the pixel coordinate system where the camera image is located; Perform point cloud detection on the lidar point cloud to obtain a preliminary detection result of road targets including category and heading angle; Project the preliminary detection result of the road target into the camera image pixel coordinate system through the calibration parameters to obtain the 2D bounding box of the radar detection target in the camera image pixel coordinate system; Combine the 2D bounding box of the radar detection target to perform target detection on the camera image to obtain the 2D bounding box of the camera detection target in the camera image pixel coordinate system; Calculate the intersection over union of the 2D bounding box of the camera detection target and the 2D bounding box of the radar detection target; Based on the intersection over union, use the Hungarian matching algorithm for optimal matching, perform probability fusion on the road target category for the matched 3D detection targets of the point cloud using the probability fusion algorithm, and increase the probability of the obstacle category for the unmatched 3D detection targets of the point cloud to obtain the corrected 3D detection road targets of the lidar point cloud; Based on the depth map corresponding to the lidar point cloud, back-project the depth points of the traffic target within the 2D bounding box of the camera detection target to obtain the 3D information of the traffic target; Merge the 3D detection road targets of the lidar point cloud and the 3D information of the traffic target to obtain the target detection result of the fusion of lidar and camera.
2. The detection method according to claim 1, characterized in that, the performing point cloud detection on the lidar point cloud to obtain a preliminary detection result of road targets including category and heading angle includes: Perform ground point segmentation on the lidar point cloud to obtain non-ground point cloud; Cluster the non-ground point cloud to obtain the coordinate position and size information of each cluster of point cloud; Downsample the coordinate position and the size information of each cluster of point cloud and input them into the PointNet neural network to obtain the preliminary detection result of the road target.
3. The detection method according to claim 1, characterized in that, the projecting the preliminary detection result of the road target into the camera image pixel coordinate system through the calibration parameters to obtain the 2D bounding box of the radar detection target in the camera image pixel coordinate system includes: Project the preliminary detection result of the road target into the camera image pixel coordinate system through the calibration parameters to obtain each initial 2D bounding box of the radar detection target; Calculate the average depth of each of the initial 2D bounding boxes; Calculate the intersection over union of any two of all the initial 2D bounding boxes; If there is a target intersection over union greater than a set threshold among all the intersection over unions, filter out the initial 2D bounding box with a larger average depth among the two initial 2D bounding boxes corresponding to the target intersection over union, obtain the non-occluded 2D bounding box in the camera image pixel coordinate system, and determine the non-occluded 2D bounding box as the 2D bounding box of the radar detection target.
4. The detection method according to claim 3, wherein, performing target detection on the camera image in combination with the 2D bounding box of the radar detection target to obtain the 2D bounding box of the camera detection target in the camera image pixel coordinate system, including: performing target detection on the camera image using a region proposal network in combination with the 2D bounding box of the radar detection target to obtain camera image detection proposal boxes; inputting the camera image detection proposal boxes and the non-occluded 2D bounding boxes into a region of interest network for classification and regression to obtain the 2D bounding box of the camera detection target.
5. The detection method according to claim 1, wherein, performing back-projection on the depth points of the traffic target within the 2D bounding box of the camera detection target based on the depth map corresponding to the lidar point cloud to obtain three-dimensional information of the traffic target, including: performing back-projection on the depth points of the traffic target within the 2D bounding box of the camera detection target based on the depth map corresponding to the lidar point cloud and the calibration parameters to obtain the original three-dimensional information of the traffic target; judging whether there are target three-dimensional points that do not meet the preset position requirements in the original traffic three-dimensional information; if so, filtering out the target three-dimensional points and obtaining the three-dimensional information of the traffic target according to the remaining three-dimensional points.
6. The detection method according to claim 1, wherein, the determination process of the depth map corresponding to the lidar point cloud includes: projecting the lidar point cloud into the camera image coordinate system through the calibration parameters to obtain an original depth map; performing filling and complementing on the original depth map in the order of holes from small to large to obtain an intermediate depth map; reducing the output noise and smoothing the local plane of the intermediate depth map to obtain the depth map corresponding to the lidar point cloud.
7. A detection device for lidar and camera fusion, wherein, comprising: an acquisition unit, configured to acquire time-synchronized lidar point cloud and camera image, and calibration parameters from the lidar coordinate system to the camera image pixel coordinate system, where the camera image pixel coordinate system is the pixel coordinate system where the camera image is located; a first detection unit, configured to perform point cloud detection on the lidar point cloud to obtain a preliminary detection result of road targets including categories and heading angles; a projection unit, configured to project the preliminary detection result of the road targets into the camera image pixel coordinate system through the calibration parameters to obtain the 2D bounding box of the radar detection target in the camera image pixel coordinate system; a second detection unit, configured to perform target detection on the camera image in combination with the 2D bounding box of the radar detection target to obtain the 2D bounding box of the camera detection target in the camera image pixel coordinate system; A correction unit for calculating the intersection over union (IoU) between the 2D bounding box of the target detected by the camera and the 2D bounding box of the target detected by the radar; Based on the IoU, use the Hungarian matching algorithm for optimal matching. For the point cloud three-dimensional detection targets that are matched, use the probability fusion algorithm to perform probability fusion on the road target categories. For the point cloud three-dimensional detection targets that are not matched, increase the probability of the obstacle category to obtain the corrected lidar point cloud three-dimensional detection road targets; A back-projection unit for back-projecting the depth points of the traffic target within the 2D bounding box of the target detected by the camera based on the depth map corresponding to the lidar point cloud to obtain the three-dimensional information of the traffic target; A result fusion unit for combining the lidar point cloud three-dimensional detection road targets and the three-dimensional information of the traffic target to obtain the target detection result of the fusion of the lidar and the camera.
8. The detection device according to claim 7, wherein, the first detection unit includes: a segmentation sub-unit for segmenting the ground points from the lidar point cloud to obtain non-ground point clouds; a clustering sub-unit for clustering the non-ground point clouds to obtain the coordinate positions and size information of each cluster of point clouds; a target detection sub-unit for downsampling the coordinate positions and the size information of each cluster of point clouds and inputting them into the PointNet neural network to obtain the preliminary detection result of the road target.
9. The detection device according to claim 7, wherein, the projection unit includes: a projection sub-unit for projecting the preliminary detection result of the road target into the camera image pixel coordinate system through the calibration parameters to obtain each initial 2D bounding box of the radar detection target; a depth calculation sub-unit for calculating the average depth of each of the initial 2D bounding boxes; a first IoU calculation sub-unit for calculating the IoU between any two of all the initial 2D bounding boxes; a bounding box determination sub-unit for, if there is a target IoU greater than a set threshold among all the IoUs, filtering out the initial 2D bounding box with the larger average depth among the two initial 2D bounding boxes corresponding to the target IoU to obtain the non-occluded 2D bounding box in the camera image pixel coordinate system, and determining the non-occluded 2D bounding box as the 2D bounding box of the radar detection target.
Citation Information
Patent Citations
Target detection method based on laser radar and image pre-fusion
CN110363820A
Unmanned driving platform real-time target 3D detection method based on camera and laser radar
CN110879401A