A multi-source heterogeneous sensor fusion method and device of a camera and a laser radar

By calculating the optimal matching results of image information and point cloud information from cameras and LiDAR, and combining the confidence scores of the sensors to select the main sensor, the fusion of multi-source heterogeneous sensors was achieved. This solved the detection problem of intelligent driving vehicles in complex traffic scenarios under all weather and all time conditions, and improved detection accuracy and adaptability.

CN114488181BActive Publication Date: 2026-01-16BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210016754.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-07
Publication Date
2026-01-16
Estimated Expiration
2042-01-07

AI Technical Summary

Technical Problem

Existing single-sensor detection methods cannot meet the detection requirements of complex traffic scenarios in intelligent driving vehicles, which are conducted around the clock and in all weather conditions. Furthermore, existing sensor fusion strategies have low adaptability and cannot effectively fuse the detection results of LiDAR and cameras.

Method used

By calculating the optimal matching result between image information and point cloud information acquired by camera and lidar, the fusion target is determined, and the master sensor is selected according to the confidence level of the sensor, realizing the fusion of multi-source heterogeneous sensors. The fusion target sequence includes 2D and 3D information of image target and point cloud target.

Benefits of technology

It improves detection accuracy and adaptability in complex environments, reduces target omissions, and meets the detection needs of all weather and all time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114488181B_ABST
    Figure CN114488181B_ABST
Patent Text Reader

Abstract

The application discloses a multi-source heterogeneous sensor fusion method and equipment of a camera and a laser radar, and is used for solving the technical problem that an existing sensor fusion strategy cannot meet the sensing demand of an intelligent automobile in an all-weather and all-working-hour working condition for a complex traffic scene. Wherein, a sequence of image target detection corresponding to the camera and a sequence of point cloud three-dimensional target detection corresponding to the laser radar are determined; a first point cloud target matched with the image target successfully and a second point cloud target not matched with the image target successfully are determined respectively; the first point cloud target and the image target matched with the first point cloud target are determined as a fusion target, and a fusion target sequence composed of a plurality of fusion targets is constructed; the second point cloud target is added to the fusion target sequence as a fusion target; a main sensor is determined from the camera and the laser radar; each fusion target in the fusion target sequence is determined as a final detected road target, and a target category corresponding to each road target is determined according to the main sensor.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sensor fusion, and in particular to a multi-source heterogeneous sensor fusion method and device of a camera and a laser radar. BACKGROUND

[0002] With the rapid development of intelligent driving cars, only relying on any single sensor cannot meet the detection needs of vehicles in complex traffic scenarios. For example, a camera can collect the shape, texture, color and other information of an object, and has superior detection performance for pedestrians and cyclists, but it is greatly affected by light and weather, and has poor detection effect in dark environments. Laser radar can easily obtain three-dimensional position information of point cloud, and detect the position, type and heading information of the target through detection algorithm, but laser radar will miss detection due to occlusion problem or near neighbor problem, and has poor detection and classification effect for small targets such as pedestrians and cyclists.

[0003] Since the existing single sensor detection method has low detection accuracy for road targets, the environmental perception ability of intelligent driving cars is poor, and the sensor fusion strategy based on laser radar and camera has been widely used in the field of intelligent driving car perception. However, the existing technology mainly adopts the original fusion strategy, which only determines that the fusion is successful when the target recognized by the laser radar matches the target recognized by the camera, and fuses the detection results of the laser radar and the camera, which has low adaptability. Moreover, the original fusion strategy sets a trusted sensor, and due to the difference of single trusted sensor, there are mainly laser radar-based fusion strategy and camera-based fusion strategy, but limited by the use of single sensor, it cannot meet the detection needs of intelligent driving cars in all-weather, all-time working conditions. SUMMARY

[0004] The present application discloses a multi-source heterogeneous sensor fusion method and device of a camera and a laser radar, which is used to solve the technical problem that the existing sensor fusion strategy cannot meet the perception needs of intelligent driving cars in all-weather, all-time working conditions for complex traffic scenarios.

[0005] In one aspect, the embodiment of the present application provides a multi-source heterogeneous sensor fusion method of a camera and a laser radar, which comprises: determining an image target detection sequence corresponding to the camera and a point cloud three-dimensional target detection sequence corresponding to the laser radar according to road image information collected by the camera and road point cloud information collected by the laser radar; wherein the image target detection sequence comprises a plurality of image targets, and the point cloud three-dimensional target detection sequence comprises a plurality of point cloud targets; calculating an optimal matching result between each of the image targets and each of the point cloud targets, and determining a first point cloud target matched successfully with the image target and a second point cloud target not matched successfully with the image target according to the optimal matching result; determining the first point cloud target and the image target matched with the first point cloud target as a fusion target, and constructing a fusion target sequence composed of a plurality of fusion targets; for each of the second point cloud targets, judging whether a confidence corresponding to the second point cloud target is greater than a preset confidence threshold, and if the confidence is greater than the preset confidence threshold, adding the second point cloud target as a fusion target into the fusion target sequence; for each of the first point cloud targets, determining a main sensor from the camera and the laser radar according to a target category and a confidence corresponding to each of the first point cloud targets and the image target matched with the first point cloud target; determining each fusion target in the fusion target sequence as a final detected road target, and determining a target category corresponding to each of the road targets according to the main sensor.

[0006] In one implementation of the present application, the main sensor is determined from the camera and the laser radar according to a target category and a confidence corresponding to each of the first point cloud targets and the image target matched with the first point cloud target, specifically comprising: determining a first target category and a first confidence corresponding to each of the image targets, and a second target category and a second confidence corresponding to each of the first point cloud targets, and determining whether the first target category and the second target category are consistent; in the case that the first target category and the second target category are consistent, determining that the camera and the laser radar are both main sensors; in the case that the first target category and the second target category are inconsistent, comparing the first confidence and the second confidence to determine that the sensor with higher confidence between the first confidence and the second confidence is the main sensor.

[0007] In an implementation form of the present application, the method further comprises: determining each fusion target in the fusion target sequence as a final detected road target, and determining a target category corresponding to each road target according to the main sensor, specifically comprising: the fusion target comprises a first point cloud target and an image target matched therewith, and a second point cloud target with a confidence greater than a preset confidence threshold; for each first point cloud target and the image target matched therewith, determining a target category corresponding to the main sensor as the target category corresponding to each road target; for the second point cloud target in the fusion target sequence, determining a target category corresponding to the second point cloud target as the target category corresponding to each road target.

[0008] In an implementation form of the present application, before calculating the optimal matching result between each image target and each point cloud target, the method further comprises: projecting the point cloud three-dimensional target detection sequence into an image plane based on a pre-determined joint calibration matrix to obtain a point cloud two-dimensional target detection sequence; wherein the joint calibration matrix is obtained by jointly calibrating the camera and the lidar based on a preset calibration tool.

[0009] In an implementation form of the present application, calculating the optimal matching result between each image target and each point cloud target specifically comprises: determining an image detection frame corresponding to each image target and a point cloud detection frame corresponding to each point cloud target in the point cloud two-dimensional target detection sequence; for each image target in the image target detection sequence, calculating an intersection over union value between each image target and each point cloud target according to the image detection frame and the point cloud detection frame to obtain a correlation matrix between the image target and the point cloud target; wherein the correlation matrix comprises the intersection over union value between each image target and each point cloud target; and calculating the optimal matching result between each image target and each point cloud target based on the correlation matrix.

[0010] In an implementation form of the present application, after determining the image target detection sequence corresponding to the camera and the point cloud three-dimensional target detection sequence corresponding to the lidar, the method further comprises: setting a corresponding time stamp for the image target detection sequence and the point cloud three-dimensional target detection sequence through an in-vehicle industrial computer arranged on the intelligent vehicle; and taking the time stamp of the point cloud three-dimensional target detection sequence as a reference, determining an image target with a minimum time stamp difference from the point cloud three-dimensional target detection sequence from the image target detection sequence to time-synchronize the corresponding road image information and road point cloud information.

[0011] In an implementation form of the present application, the fusion target comprises a target category, two-dimensional plane information, spatial position information, heading information and distance information; after determining the target category corresponding to each road target, the method further comprises: projecting each fusion target in the fusion target sequence into a world coordinate system corresponding to the current intelligent vehicle, to determine the position, distance and heading information of each fusion target relative to the intelligent vehicle; and determining the action decision of the intelligent vehicle at the current time according to the target category corresponding to each fusion target and the position, distance and heading information relative to the intelligent vehicle.

[0012] In an implementation form of the present application, the image target detection sequence corresponding to the camera and the point cloud three-dimensional target detection sequence corresponding to the laser radar are determined according to the road image information collected by the camera and the road point cloud information collected by the laser radar, and specifically comprise: inputting the road image information and the road point cloud information into corresponding pre-trained road target detection models respectively; determining each image target, target detection information corresponding to each image target, each point cloud target and target detection information corresponding to each point cloud target according to each pre-trained road target detection model; wherein the target detection information of the image target comprises the category, center point pixel coordinates and length-width size of the image target, and the target detection information of the point cloud target comprises the category, center point spatial coordinates and length-width-height size of the point cloud target.

[0013] In an implementation form of the present application, before adding the second point cloud target as a fusion target into the fusion target sequence, the method further comprises: projecting each second point cloud target with a confidence greater than a pre-set confidence threshold into an image plane to obtain appearance information of each second point cloud target.

[0014] On the other hand, the embodiments of the present application also provide a multi-source heterogeneous sensor fusion device of camera and laser radar, which comprises: a processor; and a memory having executable code stored thereon, when the executable code is executed, the processor executes a multi-source heterogeneous sensor fusion method of camera and laser radar as described above.

[0015] The multi-source heterogeneous sensor fusion method and system of camera and laser radar provided by the embodiments of the present application at least have the following beneficial effects:

[0016] The road image information and the road point cloud information can detect the target category and the confidence of the road target. The confidence corresponding to the sensor is different when the weather, road conditions and the like are different. Thus, the corresponding main sensor can be adaptively determined according to the confidence, and the target category is determined according to the target category corresponding to the main sensor, the detection and recognition of the road target in a complex environment scene are realized, the reliability is high, and the detection demand in various all-weather and all-time working conditions can be met. Moreover, when the confidence of the point cloud target that fails to be matched with the image target reaches a certain degree, the point cloud target can also be taken as the finally detected target, and the target missing detection phenomenon is reduced to a certain extent. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are included to provide a further understanding of the present application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and serve to explain the present application. In the drawings:

[0018] Figure 1 A multi-source heterogeneous sensor fusion method flow chart of a camera and a laser radar provided for an embodiment of the present application is provided.

[0019] Figure 2 A flow chart of encoding point cloud information into a pseudo image provided for an embodiment of the present application is provided.

[0020] Figure 3 A schematic diagram of time synchronization of a camera and a laser radar provided for an embodiment of the present application is provided.

[0021] Figure 4 A target fusion strategy flow chart provided for an embodiment of the present application is provided.

[0022] Figure 5 A main sensor selection flow chart provided for an embodiment of the present application is provided.

[0023] Figure 6 An image target detection result schematic diagram provided for an embodiment of the present application is provided.

[0024] Figure 7 A point cloud target detection result schematic diagram provided for an embodiment of the present application is provided.

[0025] Figure 8 A fusion target detection result schematic diagram provided for an embodiment of the present application is provided.

[0026] Figure 9 A multi-source heterogeneous sensor fusion device structure schematic diagram of a camera and a laser radar provided for an embodiment of the present application is provided. DETAILED DESCRIPTION

[0027] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below in combination with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0028] Generally, multiple sensors are provided on an intelligent driving vehicle. In the driving process, environmental data is collected through multiple sensors to realize the perception and recognition of complex traffic scenes, and accordingly, the action strategy of the vehicle in the current scene is determined.

[0029] The prior art generally adopts an original fusion strategy, that is, the laser radar point cloud data is projected onto an image, and a fusion method is used to fuse the detection results of the laser radar and the camera. When the targets in the image target detection sequence and the point cloud target detection sequence can be matched two by two, the target is determined as a fusion target. According to the difference in trusted sensors, the fusion strategy is divided into a laser radar-based fusion strategy and a camera-based fusion strategy. However, the use scene and recognition accuracy of a single sensor are insufficient, and the original fusion strategy cannot meet the detection requirements in various all-weather and all-time working conditions. In addition, the traditional camera and laser radar fusion algorithm generally "combines" the category and confidence information of the image and the three-dimensional position information of the point cloud, but is not a true fusion. In the embodiment of the present application, the image information and the point cloud information can both detect the category and confidence of the target, and therefore, true fusion can be realized in the category dimension of the target.

[0030] The present application discloses a multi-source heterogeneous sensor fusion method and device of a camera and a laser radar, which is used to solve the technical problem that the existing sensor fusion strategy cannot meet the perception requirements of intelligent driving vehicles in all-weather and all-time working conditions for complex traffic scenes.

[0031] The technical solutions of the embodiments of the present application will be described in detail below with reference to the drawings.

[0032] Figure 1 A multi-source heterogeneous sensor fusion method flow chart of a camera and a laser radar is provided for the embodiments of the present application. As shown in Figure 1 The multi-source heterogeneous sensor fusion method of a camera and a laser radar provided by the embodiments of the present application mainly includes the following steps:

[0033] S101, the server determines an image target detection sequence corresponding to the camera and a point cloud three-dimensional target detection sequence corresponding to the laser radar according to the road image information collected by the camera and the road point cloud information collected by the laser radar.

[0034] The vehicle-mounted sensors, such as cameras and laser radars, collect road information around the vehicle during the driving process. The server can perform target detection on the road through a pre-trained road target detection model, so as to identify the position and category of different targets on the surrounding road and provide a reference for the behavior decision of the vehicle. The targets include pedestrians, vehicles and cyclists.

[0035] To realize the fusion of the image targets detected by the camera and the point cloud targets detected by the laser radar, the server needs to pre-train a corresponding road target detection model for the image targets and the point cloud targets before obtaining the road image information and the road point cloud information. Let the image target detection sequence be I = {I0, I1, I2, …, I i}, and the point cloud target detection sequence be L = {L0, L1, L2, …, L j}, where each image target in the above sequence is detected by the camera, and each point cloud target in the above sequence is detected by the laser radar.

[0036] In an embodiment, the embodiment of the present application can use the preset vehicle road test data set KITTI as a training set to retrain the pre-set target detection model for road image information, so as to adapt to the road target detection requirement.

[0037] For example, the embodiment of the present application can obtain a target detection model through YOLOV4 training. YOLOV4 trains the MSCOCO data set, which is a network containing 80 categories. The network feature output dimension is 255, and the network output dimension calculation formula is 3 × (5 + 80). Among them, 3 represents that each grid contains 3 scales of Anchor, 5 represents the center coordinates x, y, width and height w, h of the prediction box, and the confidence value c, and 80 represents 80 categories. However, the embodiment of the present application only performs three classifications, i.e. pedestrians, cyclists and vehicles. Therefore, the network output dimension of the improved road target detection model for road image information should be changed to 3 × (5 + 20) = 75 dimensions.

[0038] After inputting the road image information collected by the camera into the above road target detection model for road image information, the road image targets around the vehicle and the target detection information corresponding to the road targets can be determined. The target detection information of the image target includes the category, center point pixel coordinates and length-width size of the image target.

[0039] In an embodiment, the server can use the pre-trained road target detection model to identify the targets on the road to obtain the 3D information of the targets. The implementation is as follows:

[0040] Firstly, the server encodes the road point cloud information (x, y, z, i) as a pseudo image as an initial input of the road target detection model. Figure 2 The flowchart for encoding the point cloud information into a pseudo image is provided for the embodiments of the present application. Figure 2 As shown in the figure, the point cloud input into the road target detection model is evenly divided into columns in the XY plane, P is the number of non-empty columns, and N is the number of points retained in each column. If the number of points in the original column is greater than N, the sampling is reduced to N points, and if it is less than N, zero padding is performed. First, a dense input vector (DxPxN) is generated from the point cloud, and a linear layer is used for feature extraction to obtain a high-dimensional feature vector CxPxN. In the channel dimension, a max-pooling operation is performed to obtain a vector CxP, and P is restored to the HxW dimension in the XY plane to obtain a CxHxW vector, which is a pseudo image vector.

[0041] Secondly, the server performs multi-scale feature extraction on the pseudo image and splices the features at different scales to obtain a corresponding feature map.

[0042] Specifically, the road target detection model can perform two times of two-fold downsampling on the features to obtain feature maps with smaller and smaller resolutions. This can enhance the features of the target at different scales, which is conducive to improving the robustness of the algorithm in dealing with changes in target size. At the same time, it is also helpful for the detection of targets of different sizes, with high resolution and small receptive field, which is conducive to the detection of small target objects such as pedestrians. Then, the features are upsampled. After upsampling, the features at three scales can be sampled to the same resolution, and these features can be spliced together to obtain a feature map.

[0043] Finally, the road target detection model can identify point cloud targets and target detection information corresponding to each point cloud target according to the feature map. The target detection information of the point cloud target includes the category, center point spatial coordinates, and length-width-height size of the point cloud target.

[0044] It should be noted that both the image target and the point cloud target are displayed in the form of a detection frame. The image target detected by the camera can reflect the appearance information of the target, and the point cloud target detected by the laser radar can reflect the three-dimensional information such as the spatial information and the heading information of the target.

[0045] After the server inputs road image information and road point cloud information into the corresponding road target detection models, it can identify road targets, thereby obtaining corresponding image target detection sequences and point cloud 3D target detection sequences. The image target detection sequence includes several image targets, and the point cloud 3D target detection sequence includes several point cloud targets. In this embodiment, targets are divided into three categories: pedestrians, cyclists, and vehicles. Image targets are those identified by the camera as existing around the vehicle, and point cloud targets are those identified by the lidar as existing around the vehicle.

[0046] In one embodiment, after obtaining the image target detection sequence and the point cloud 3D target detection sequence, to achieve their fusion, the camera and LiDAR must first be time-synchronized. This application embodiment employs a soft time synchronization method.

[0047] Specifically, the onboard industrial control computer installed in the intelligent vehicle sets corresponding timestamps for the image target detection sequence and the point cloud 3D target detection sequence. Since the acquisition frequency of the LiDAR is lower than that of the camera, the server uses the timestamp of the point cloud 3D target detection sequence as a benchmark to determine the image target with the smallest timestamp difference from the point cloud 3D target detection sequence. Then, it fuses the unfused data preceding that image target, thereby achieving time synchronization between road image information and road point cloud information. Time synchronization of the camera and LiDAR is a prerequisite for their fusion; only multi-source heterogeneous sensors operating at the same time have the potential for fusion.

[0048] Figure 3 This is a schematic diagram illustrating time synchronization between a camera and a LiDAR system provided in an embodiment of this application. Figure 3 As shown, after the industrial control computer assigns timestamps, both road image information and road point cloud information are in the form of data queues, and each data in the queue has a corresponding timestamp. Assuming the LiDAR acquisition frequency is 10Hz and the camera image acquisition frequency is 25Hz, the image data with the smallest timestamp difference from the point cloud data queue are C2, C4, C6, and C9. Therefore, during time synchronization, {C1, C2} is synchronized with P1, {C3, C4} is synchronized with P2, {C5, C6} is synchronized with P3, and {C7, C8, C9} is synchronized with P4.

[0049] S102. The server calculates the optimal matching result between each image target and each point cloud target, and based on the optimal matching result, determines the first point cloud target that successfully matches the image target and the second point cloud target that fails to match the image target.

[0050] The fusion of the camera and the lidar is essentially a process of obtaining fusion targets. If the image target and the point cloud target are successfully matched, the target is regarded as a fusion target. The fusion target has both appearance information carried by the image target and three-dimensional information carried by the point cloud target, and the process of obtaining the fusion target is actually a process of matching the image target and the point cloud target.

[0051] The matching of the image target sequence and the point cloud target sequence can be classified as a bipartite graph matching problem. The matching refers to that the points in set X and the points in set Y are paired to form a set of edges, and the two vertices of each edge in the set are different from the two vertices of other edges. Therefore, the matching relationship between the image target sequence and the point cloud target sequence can be obtained by using the Hungarian algorithm, and the matching relationship is equivalent to connecting the image target and the point cloud target. The matching relationship between the image target and the point cloud target is actually the optimal matching result between the two. That is, the matching relationship with the maximum sum of the weights of the connections between the image target and the point cloud target is selected after the connections between the image target and the point cloud target are weighted. At this time, the number of fusion targets obtained by the point cloud target and the image target is the largest, and the target recognition is more accurate.

[0052] However, the point cloud information collected by the lidar is three-dimensional information, and each target in the point cloud target detection sequence obtained by the road detection model is also a three-dimensional target. To perform heterogeneous fusion of the camera and the lidar, the three-dimensional target in the point cloud three-dimensional target detection sequence needs to be projected onto the image plane first. Only when the image target detection sequence and the point cloud target detection sequence are in the same coordinate system, can they be associated and matched.

[0053] In an embodiment, to project the point cloud three-dimensional target detection sequence, the server needs to jointly calibrate the camera and the lidar, so as to obtain a joint calibration matrix for projecting the point cloud three-dimensional target detection sequence to the image plane.

[0054] Specifically, the lidar and the camera are rigidly connected, so that in the same space, each point in the lidar coordinate system corresponds to a unique point in the camera coordinate system. Meanwhile, each point in the camera coordinate system corresponds to a unique pixel point in the pixel coordinate system, so that it can be obtained that each point cloud in the lidar coordinate system corresponds to a unique pixel point in the pixel coordinate system. The spatial constraint relationship between the camera image and the lidar point cloud can be used to correspondingly solve the coordinate conversion relationship between the pixel coordinate system and the lidar coordinate system, and obtain the joint calibration matrix. The joint calibration matrix is determined by formula (7) in the embodiment of the present application.

[0055]

[0056] wherein R represents a spatial coordinate rotation, T represents a spatial coordinate translation, u represents a horizontal pixel coordinate system, v represents a vertical pixel coordinate system, u0 represents a horizontal pixel coordinate system u-axis origin, v0 represents a vertical pixel coordinate system v-axis origin, f represents a camera focal length, X L represents an X-axis coordinate system of the lidar, Y L represents a Y-axis coordinate system of the lidar, and Z L represents a Z-axis coordinate system of the lidar.

[0057] It should be noted that when the coordinate conversion relationship is solved by the joint calibration matrix, an optimal solution can be obtained by using a linear least square method.

[0058] In a possible implementation manner, the embodiments of the present application can calibrate the camera and the lidar by using a calibration tool preset in Autoware. The calibration can be performed in the following manner:

[0059] (1) A large enough open area is selected, and a total of 30 poses are placed at a long distance, a short distance, and different directions from the intelligent driving vehicle, respectively, and the poses of the checkerboard in the camera and the lidar coordinate system are recorded. It should be noted that the placement of the corresponding poses in the embodiments of the present application is not limited.

[0060] (2) The data packets are paused, and the camera and the lidar are calibrated. In the corner points in the camera plane, the checkerboard plane is automatically extracted by the calibration tool. In the lidar coordinate system, the checkerboard plane and the normal vector are manually extracted.

[0061] (3) Step (1) is repeated, and the image corner points of the 30 poses of the checkerboard and the plane and the normal vector of the checkerboard in the point cloud are extracted.

[0062] (4) Target optimization is automatically performed, and the optimized rotation matrix and the translation matrix are obtained, and the joint calibration matrix used for coordinate system conversion is determined.

[0063] After the joint calibration matrix for projecting the point cloud three-dimensional target detection sequence to the pixel coordinate system is obtained, the server projects the point cloud three-dimensional target detection sequence into the image plane based on the joint calibration matrix, to obtain a point cloud two-dimensional target detection sequence.

[0064] In an embodiment, after the server obtains the point cloud two-dimensional target detection sequence, the server determines the image detection frame corresponding to each image target and the point cloud detection frame corresponding to each point cloud target in the point cloud two-dimensional target detection sequence. For the image target detection sequence I={I0,I1,I2,...,I i} and the point cloud two-dimensional target detection sequence L={L0,L1,L2,...,L jFor each target in the image and the point cloud, the IOU value between the image detection box corresponding to the image target and the point cloud detection box corresponding to the point cloud target is calculated. Then, an association matrix between the image target and the point cloud target is constructed according to the IOU values. The elements in the association matrix are the IOU values between the image target and the point cloud target. Then, the server calculates the optimal matching result between the image target and the point cloud target based on the association matrix and the Hungarian algorithm. Specifically, the IOU value between each target in the sequence I and L is calculated to obtain the association matrix as shown in formula (8):

[0065]

[0066] wherein, I ij represents the IOU value between the ith image detection box in the image target detection sequence I and the jth point cloud detection box in the point cloud two-dimensional target detection sequence L.

[0067] The association matrix is input into the Hungarian algorithm as the weight value between the image target and the point cloud target, and the final output result is the optimal matching result between I and L. Assuming that the optimal matching result is the set {(I2, L3), (I3, L5)......(I i , L j )}, when the optimal matching result is obtained, the number of successfully matched image targets and point cloud targets is the largest, and the weight sum between the successfully matched image targets and point cloud targets is the largest. At this time, the detected image target and the point cloud target have the highest overlap.

[0068] At this time, the server can obtain the matching relationship between the image target and the point cloud target. If the point cloud target matches the image target, the point cloud target is regarded as a first point cloud target; if the matching is unsuccessful, the point cloud target is regarded as a second point cloud target.

[0069] In S103, the server determines the first point cloud target and the image target matched therewith as a fusion target, and constructs a fusion target sequence composed of a plurality of fusion targets.

[0070] After the single-frame matching of the image target and the point cloud target, three types of targets are obtained. The first type is the target detected by both the image and the point cloud and successfully matched, the second type is the target detected by the image but not matched with any target in the point cloud target sequence, and the third type is the target detected by the point cloud but not matched with any image target. The server regards the first point cloud target and the image target matched therewith as a fusion target, because when the two are successfully matched, the target detected by the vehicle has both the appearance information of the image target and the three-dimensional information of the point cloud target. A plurality of fusion targets can constitute a fusion target sequence.

[0071] The fusion strategy provided in the embodiments of the present application needs to ensure that the target is detected in both the camera and the lidar, so that the occurrence of false detection can be effectively reduced, and the detection accuracy can be improved.

[0072] In S104, the server determines whether the confidence corresponding to each second point cloud target is greater than a preset confidence threshold, and if the confidence is greater than the preset confidence threshold, adds the second point cloud target to the fusion target sequence as a fusion target.

[0073] The camera and the lidar have high requirements for target detection, and thus there may be individual missed detection. The lidar has high reliability compared with the camera, but the lidar only reflects when encountering a solid object, and the detected target object needs to have a certain volume. Therefore, to reduce the possibility of missed detection, if the confidence of a point cloud target in the point cloud target detection sequence that is not matched with the image target is greater than a set confidence threshold, it is considered that the target exists, and the target is added to the fusion sequence.

[0074] In one embodiment, since the fusion targets in the fusion target sequence all have 2D and 3D information, for the second point cloud target that is not matched successfully and has a confidence greater than a preset confidence threshold, the second point cloud target needs to be projected to the image plane to obtain the appearance information (2D information) of the target before being added to the fusion target sequence. In this way, the second point cloud target added to the fusion sequence has both the three-dimensional information (3D information) of the point cloud and the appearance information (2D information) of the image.

[0075] Figure 4 A target fusion strategy flowchart is provided in the embodiments of the present application. As shown in Figure 4 The point cloud detection target carries 3D information of the point cloud, and the image detection target carries 2D information of the image. If the two are successfully matched, the 2D and 3D information is fused in the matched target, and the target can be directly used as a fusion target. For the second point cloud target that is not matched, the second point cloud target needs to be projected to the image plane first. In this way, the second point cloud target obtains 2D information through coordinate conversion on the basis of the original 3D information, and the second point cloud target that has fused 2D and 3D information can be added to the fusion target sequence as a fusion target. That is, the fusion target sequence not only has matched image targets and point cloud targets, but also has point cloud targets with a confidence greater than a preset confidence threshold. In this way, the detection accuracy is ensured, and the possibility of missed detection is reduced.

[0076] In S105, the server determines a main sensor from the camera and the lidar according to the target category and the confidence corresponding to each first point cloud target and the image target matched with the first point cloud target.

[0077] The conventional sensor fusion method is usually provided with a trusted sensor, and when the image target and the point cloud target are matched, the category information detected by the trusted sensor is considered as the final detected target category. However, the trusted sensor is usually a single sensor, and its detection capability is limited, which cannot meet the detection requirements in multiple working conditions. Therefore, the server no longer uses the pre-set trusted sensor to determine the category of the target for the first point cloud target and the image target matched successfully, but selects the main sensor from the camera and the lidar according to the confidence of the target recognized by different sensors. The target category detected by the main sensor is the category corresponding to the final fusion target.

[0078] Specifically, after the server inputs the road point cloud information and the road image information into the corresponding road target detection model, the target categories and the confidence corresponding to the image target and the point cloud target can be obtained. First, the first target category and the first confidence corresponding to each image target, and the second target category and the second confidence corresponding to each first point cloud target are determined, and it is determined whether the first target category and the second target category are consistent. If they are consistent, it means that the target categories detected by the camera and the lidar are consistent, and both the camera and the lidar can be used as the main sensor. If they are not consistent, the server compares the first confidence and the second confidence, and selects the sensor with higher confidence as the main sensor.

[0079] Through the above-mentioned main sensor selection strategy, a flexible fusion strategy can be realized. When strong light or dark light is encountered, the reliability of the camera decreases, and at this time, the classification of the target by the lidar point cloud can be relied on. When encountering objects such as cyclists and pedestrians with similar volumes and slightly overlapping features, the recognition effect of the lidar on these two categories is slightly worse than that of the camera, and at this time, the detection result of the camera can be relied on. Through this strategy, the main sensor can be selected online for different scenes, which changes the original fusion algorithm that needs to pre-select the main sensor offline, and can effectively improve the classification accuracy and environmental adaptability of the fusion algorithm.

[0080] As shown in the main sensor selection flowchart. Figure 5 The category of the image target detected is A, and the confidence is C. The category of the point cloud target detected is B, and the confidence is D. First, compare the categories A and B. If A=B, then the target is of the A=B category. If the categories are different, the confidence of each sensor needs to be compared. When the confidence C of the target detected by the image is greater than or equal to the confidence D of the target detected by the point cloud, the target is considered to be of the A category, and vice versa.

[0081] S106, the server determines each fusion target in the fusion target sequence as a final detected road target, and determines a target category corresponding to each road target according to the main sensor.

[0082] The camera and the lidar perform target detection on the road, and the ultimate goal is to identify the targets around the vehicle through the sensors. The final identified fusion targets are divided into two categories: one category is the image targets and the first point cloud targets matched with each other, and the other category is the second point cloud targets with a confidence greater than a preset confidence threshold. The server determines the target category corresponding to the main sensor for each first point cloud target and the image target matched with the first point cloud target, and the target category corresponding to each road target. In addition, the server determines the target category corresponding to the second point cloud target in the fusion target sequence, and the target category corresponding to each road target. That is, for the matched fusion targets, the category corresponding to the main sensor is the final detected road target category; for the un-matched fusion targets, the target category corresponding to the fusion targets is the final detected road target category. It should be noted that the target categories corresponding to the main sensor and the second point cloud target both include pedestrians, cyclists, and vehicles.

[0083] In one embodiment, after determining the road target category, the server projects each fusion target in the fusion target sequence into the world coordinate system corresponding to the current intelligent vehicle, so as to determine the target category of each fusion target and the position, distance, and heading information of each fusion target relative to the intelligent vehicle according to the information contained in the fusion target. In this way, the intelligent vehicle can determine the corresponding action decision according to the road target.

[0084] The embodiments of the present application adopt two methods of target fusion strategy and main sensor selection strategy to detect road targets. The original fusion strategy is that when the targets in the image target sequence and the point cloud target sequence can be matched two by two, the targets are fusion targets. This algorithm is referred to as an unimproved fusion algorithm. According to different trusted sensors, the fusion strategy is divided into a lidar-based fusion strategy and a camera-based fusion strategy. In order to improve the adaptability of the algorithm, the present application no longer pre-sets the most trusted sensor, but adopts a main sensor selection strategy. At this time, the algorithm is referred to as an improved fusion algorithm. In addition, the unimproved fusion algorithm considers that the sufficient condition for successful target fusion is that the camera and the lidar can both detect the target. However, in the embodiments of the present application, it is considered that the lidar has high reliability, and therefore the target detected by the lidar is also a fusion target. The appearance information of the target can be obtained by projecting the target to a 2D image plane. At this time, the algorithm is referred to as an improved fusion algorithm Plus. The effects of the four fusion strategies are compared as shown in the following table.

[0085] Table 1, comparison results of four fusion strategies

[0086]

[0087] The improved fusion algorithm with the main sensor selection strategy has little improvement compared with the unimproved fusion algorithm based on camera, but has great improvement compared with the unimproved fusion algorithm based on lidar. This is because in the test scene of KITTI dataset, most of them are well-lit environments, and when detecting and classifying targets, the camera has high reliability, so when the main sensor fusion strategy is added in this paper to adaptively select the reliable sensor, there is little improvement for the fusion algorithm based on camera, but there is great modification for the fusion algorithm based on lidar. But when the reliability of lidar is higher than that of camera, such as at night, the camera is invalid in most cases, and the improved fusion algorithm can adaptively select the detection results of lidar, at this time, the detection effect will be greatly improved. As can be seen, the improved fusion algorithm can effectively adapt to target detection in different environments and climate conditions, and autonomously adjust the reliable sensor to ensure that the target detection effect is greater than or equal to the higher level of image and point cloud detection effect.

[0088] The improved fusion algorithm Plus has about 2% improvement in detection accuracy on Car, about 0.5% improvement on Pedestrian, and overall improvement of 0.9%. The improved fusion algorithm Plus has a more significant improvement in accuracy on Car, indicating that the improved fusion algorithm Plus can reduce the occurrence of missed detection problems to a certain extent.

[0089] Based on the above analysis, it can be considered that the main sensor selection strategy with adaptive sensor selection is beneficial to improve the environmental adaptability and robustness of the fusion algorithm. The target fusion strategy can effectively improve the detection accuracy of the algorithm for vehicles and effectively reduce the occurrence of missed detection problems. If you want to further improve the detection accuracy of the fusion algorithm for pedestrians and cyclists, you need to improve the detection accuracy of the point cloud detection algorithm for these two categories.

[0090] Next, according to the improved fusion algorithm Plus, randomly select one scene from the KITTI dataset to test the fusion effect of camera and lidar. Figure 6 The image target detection result is shown in the figure, Figure 7 The point cloud target detection result is shown in the figure, Figure 8 The fusion target detection result is shown in the figure. As Figure 6 , 7As shown in FIG. 8, due to the occlusion problem and the light problem, three cars are missed in the image target detection result, and the pedestrian is misidentified as a rider in the point cloud target detection result. The fusion target detection result corrects the two detection results, and the target object can be completely detected and classified correctly.

[0091] Figure 9 A camera and laser radar multi-source heterogeneous sensor fusion method and device structure schematic diagram are provided for the embodiments of the present application. As shown in FIG. 1, the device stores computer executable instructions, characterized in that the computer executable instructions are configured as a camera and laser radar multi-source heterogeneous sensor fusion method as described above. Figure 9

[0092] Each of the embodiments in the present application is described in a progressive manner, and the same or similar parts of each embodiment can be referred to each other. Each embodiment mainly describes the difference from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0093] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0094] The above only describes the embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the scope of the claims of the present application.​

Claims

1. A multi-source heterogeneous sensor fusion method of camera and lidar, characterized in that, The method comprises: According to the road image information collected by the camera and the road point cloud information collected by the lidar, determine the image target detection sequence corresponding to the camera and the point cloud three-dimensional target detection sequence corresponding to the lidar; wherein the image target detection sequence includes a plurality of image targets, and the point cloud three-dimensional target detection sequence includes a plurality of point cloud targets; Calculate the optimal matching result between each image target and each point cloud target, and determine the first point cloud target matched with the image target successfully and the second point cloud target not matched with the image target successfully according to the optimal matching result; Determine the first point cloud target and the image target matched therewith as a fusion target, and construct a fusion target sequence composed of a plurality of fusion targets; For each second point cloud target, determine whether the confidence corresponding to the second point cloud target is greater than a preset confidence threshold, and if greater than the preset confidence threshold, add the second point cloud target as a fusion target to the fusion target sequence; For each first point cloud target, determine the main sensor from the camera and the lidar according to the target categories and confidences corresponding to each first point cloud target and the image target matched therewith; Determine each fusion target in the fusion target sequence as a final detected road target, and determine the target category corresponding to each road target according to the main sensor; Determine the main sensor from the camera and the lidar according to the target categories and confidences corresponding to each first point cloud target and the image target matched therewith, specifically comprising: Determine the first target category and the first confidence corresponding to each image target, and the second target category and the second confidence corresponding to each first point cloud target, and determine whether the first target category and the second target category are consistent; In the case that the first target category and the second target category are consistent, determine that the camera and the lidar are both main sensors; In the case that the first target category and the second target category are inconsistent, compare the first confidence and the second confidence to determine that the sensor with higher confidence in the first confidence and the second confidence is the main sensor; Determine each fusion target in the fusion target sequence as a final detected road target, and determine the target category corresponding to each road target according to the main sensor, specifically comprising: The fusion target includes a first point cloud target and an image target matched therewith, and a second point cloud target with a confidence greater than a preset confidence threshold; For each first point cloud target and the image target matched therewith, determine the target category corresponding to the main sensor as the target category corresponding to each road target; For the second point cloud target in the fusion target sequence, determine the target category corresponding to the second point cloud target as the target category corresponding to each road target.

2. The multi-source heterogeneous sensor fusion method of camera and lidar according to claim 1, characterized in that, Before calculating the optimal matching result between each image target and each point cloud target, the method further comprises: Project the point cloud three-dimensional target detection sequence into an image plane based on a predetermined joint calibration matrix to obtain a point cloud two-dimensional target detection sequence; wherein the joint calibration matrix is obtained by jointly calibrating the camera and the laser radar based on a preset calibration tool.

3. The multi-source heterogeneous sensor fusion method of claim 2, wherein, Calculate the optimal matching result between each image target and each point cloud target, specifically including: Determine the image detection frame corresponding to each image target and the point cloud detection frame corresponding to each point cloud target in the point cloud two-dimensional target detection sequence; For each image target in the image target detection sequence, calculate the intersection over union value between each image target and each point cloud target based on the image detection frame and the point cloud detection frame to obtain the association matrix between the image target and the point cloud target; wherein the association matrix includes the intersection over union value between each image target and each point cloud target; Calculate the optimal matching result between each image target and each point cloud target based on the association matrix.

4. The multi-source heterogeneous sensor fusion method of camera and lidar according to claim 1, characterized in that, After determining the image target detection sequence corresponding to the camera and the point cloud three-dimensional target detection sequence corresponding to the laser radar, the method further includes: Set corresponding time stamps for the image target detection sequence and the point cloud three-dimensional target detection sequence through the vehicle-mounted industrial computer arranged on the intelligent vehicle; Determine the image target with the smallest time stamp difference from the point cloud three-dimensional target detection sequence from the image target detection sequence based on the time stamp of the point cloud three-dimensional target detection sequence as a reference, to time synchronize the corresponding road image information and road point cloud information.

5. The multi-source heterogeneous sensor fusion method of camera and lidar according to claim 1, characterized in that, The fusion target includes target category, two-dimensional plane information, spatial position information, heading information, and distance information; After determining the target category corresponding to each road target, the method further includes: Project each fusion target in the fusion target sequence into the world coordinate system corresponding to the current intelligent vehicle to determine the position, distance, and heading information of each fusion target relative to the intelligent vehicle; Determine the action decision of the intelligent vehicle at the current time based on the target category corresponding to each fusion target and the position, distance, and heading information relative to the intelligent vehicle.

6. The multi-source heterogeneous sensor fusion method of camera and lidar according to claim 1, characterized in that, Determine the image target detection sequence corresponding to the camera and the point cloud three-dimensional target detection sequence corresponding to the laser radar based on the road image information collected by the camera and the road point cloud information collected by the laser radar, specifically including: Input the road image information and the road point cloud information into the corresponding pre-trained road target detection model respectively; Determine each image target, the target detection information corresponding to each image target, each point cloud target, and the target detection information corresponding to each point cloud target based on each pre-trained road target detection model; wherein the target detection information of the image target includes the category, center point pixel coordinates, and length-width size of the image target, and the target detection information of the point cloud target includes the category, center point spatial coordinates, and length-width-height size of the point cloud target.

7. The multi-source heterogeneous sensor fusion method of camera and lidar according to claim 1, wherein, Before adding the second point cloud target as a fusion target into the fusion target sequence, the method further comprises: Projecting each second point cloud target with a confidence greater than a preset confidence threshold into an image plane to obtain appearance information of each second point cloud target.

8. A multi-source heterogeneous sensor fusion device of camera and lidar, storing computer executable instructions, characterized in that, The computer executable instructions are configured to: A multi-source heterogeneous sensor fusion method of a camera and a laser radar according to any one of claims 1-7. A multi-source heterogeneous sensor fusion method of a camera and a laser radar according to any one of claims 1-7.

Citation Information

Patent Citations

  • Unmanned ship water surface target detection, identification and positioning method based on monocular camera and lidar information fusion

    CN109444911A

  • Vehicle detection method based on laser and vision fusion

    CN110942449A