A target detection method based on multi-modal data fusion
Patent Information
- Application Number
- CN202410854251.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-06-28
AI Technical Summary
[0003]目前已有的多源数据融合的目标检测方法很多都在特征级层面进行融合,实现复杂,在决策层相关的融合方法相对较少
[0031]1、本发明通过融合来自相机和激光雷达的多模态数据,利用各自的优势互补,提高了目标检测的准确性和精度。
Smart Images

Figure CN118840637B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of autonomous driving and target detection technology, specifically involving a target detection method based on multimodal data fusion. Background Technology
[0002] Currently, autonomous vehicles primarily use sensors such as cameras and LiDAR to detect vehicles and pedestrians on the road. These two types of sensors each have their unique data acquisition characteristics, giving them certain advantages in target detection. Cameras can capture rich texture information of objects, making it possible to detect, recognize, and distinguish objects from the background, but they are easily affected by lighting and weather conditions. LiDAR is unaffected by seasons and lighting conditions, has a long detection range, and can provide accurate 3D position information, but LiDAR point cloud data is sparse and struggles to obtain detailed scene information. Due to the limitations of single sensors, we consider fusing multi-source data to leverage their complementary strengths and improve the robustness and accuracy of target detection in autonomous driving scenarios. With the development of deep learning, Convolutional Neural Networks (CNNs) have achieved significant success in image target detection, providing a powerful support platform for multimodal fusion.
[0003] Current object detection methods based on multi-source data fusion often operate at the feature level, which is complex to implement. There are relatively few fusion methods at the decision level. Furthermore, when incorporating bird's-eye view information, the problem of coordinate transformation between the bird's-eye view and the front view's bounding boxes arises. Research indicates that current methods for coordinate transformation between the two views rely on deep learning, which consumes significant time resources and slows down the inference process of the entire fusion framework. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a target detection method based on multimodal data fusion with fast detection speed and high accuracy.
[0005] The technical solution adopted to solve the above-mentioned technical problems is: a target detection method based on multimodal data fusion, comprising the following steps:
[0006] Step 1. Data Preprocessing
[0007] Collect a natural image dataset containing the target and a corresponding LiDAR point cloud dataset. Divide the natural image dataset into a training set and a validation set according to the proportion. Project the LiDAR point cloud dataset onto a two-dimensional plane to generate a sparse depth image set and a bird's-eye view image set. Divide the sparse depth image set into a training set and a validation set according to the proportion. Divide the bird's-eye view image set into a training set and a validation set according to the proportion.
[0008] Step 2. Construct the object detection network
[0009] The YOLOv8-P2 network was used as the object detection network to detect target objects in images of different modalities.
[0010] Step 3. Train the object detection network YOLOv8-P2
[0011] The training sets of the natural image dataset, the bird's-eye view image dataset, and the sparse depth image dataset are respectively input into the object detection network YOLOv8-P2. The object detection network YOLOv8-P2 is trained and optimized using the loss function L, and the weight file obtained after training each training set is saved.
[0012] The loss function L is composed of the classification loss L BCE And bounding box regression loss L Bbox_Loss It consists of two parts;
[0013] The classification loss L BCE for:
[0014]
[0015] In the formula, N is the total number of samples, and y j p is the category label of the j-th sample, where the positive class label is 1 and the negative class label is 0; j It is the probability that the j-th sample is predicted as positive.
[0016] The bounding box regression loss L Bbox_Loss for:
[0017] L Bbox_Loss =L DFL +L WIoU
[0018] L DFL =-((y) i+1 -y)log(w i )+(yy i )log(w i+1 ))
[0019]
[0020] In the formula, y i and y i+1 These are the two values closest to the true label y, y i ≤y≤y i+1 L IoU = 1 - IoU, where IoU is the intersection-over-union ratio between the detected bounding box and the ground truth bounding box, and (x,y) are the x and y coordinates of the center point of the predicted bounding box. gt ,y gt) represents the x and y coordinates of the center point of the true bounding box, W g and H g These are the width and height of the actual bounding box, respectively. The width and height of the actual bounding box are separated from the computation graph. The cross-union ratio (CUB) loss between the detection boxes and the ground truth bounding boxes separates them from the computational graph; α and δ are hyperparameters. It is a moving average with momentum m. n is the training batch size, and t is the batch value when the rate of improvement of average detection accuracy during training slows down significantly.
[0021] Step 4. Generate target detection results in different modalities
[0022] The validation sets of the natural image dataset, the bird's-eye view image dataset, and the sparse depth image dataset are respectively input into the trained object detection network YOLOv8-P2 to obtain the object detection results of the natural modality, the bird's-eye view modality, and the sparse modality.
[0023] Step 5. Overlay the target detection results of the sparse depth image and the target detection results of the natural image in the same front view direction, and remove redundant target detection boxes using the Soft-NMS model to obtain the initial fusion result;
[0024] Step 6. The target detection results of the bird's-eye view image are transformed into coordinate information using a multi-output regression model based on the K-nearest neighbor algorithm, so that the view direction is the same as that of the target detection results of the sparse depth image and the natural image. Then, the coordinate information is superimposed with the initial fusion result and redundant target detection boxes are removed by the Soft-NMS model to obtain the final target detection result.
[0025] As a preferred technical solution, the method for projecting the LiDAR point cloud dataset onto a two-dimensional plane to generate a sparse depth image set and a bird's-eye view image set is as follows: the LiDAR point cloud data is first transformed into the coordinates of the camera coordinate system points through rigid body transformation, then transformed into the coordinates of the image coordinate system points through perspective projection, and finally transformed into the coordinates of the two-dimensional pixel coordinate system points through affine transformation. Among these methods, the pixel values in the sparse depth image are obtained by calculating the longitudinal distance from the original LiDAR point to the origin of the current LiDAR coordinate system, and the pixel values in the bird's-eye view image are obtained by encoding the density, reflection intensity, and height of the LiDAR point cloud in the RGB three color channels.
[0026] As a preferred technical solution, the parameters in step 3 of training the target detection network YOLOv8-P2 are set as follows: momentum m is 0.937, weight decay is 0.0005, training batch n is 16, and the initial learning rate is set to 0.001.
[0027] As a preferred technical solution, the Soft-NMS model formula in steps 5 and 6 is:
[0028]
[0029] In the formula, S k S′ is the initial confidence score of the k-th bounding box. k Here, M represents the confidence score after weight decay update, and b is the location of the bounding box with the highest confidence score. k is the position of the k-th bounding box, D is the set of bounding boxes with the highest scores, IOU(·) is the intersection-union ratio of the areas of two bounding boxes, and σ is the standard deviation of the Gaussian function.
[0030] The beneficial effects of this invention are as follows:
[0031] 1. This invention improves the accuracy and precision of target detection by fusing multimodal data from cameras and lidar, leveraging their complementary advantages.
[0032] 2. This invention employs the YOLOv8-P2 network structure and an improved Wise-IoU loss function to optimize detection speed, enabling the model to perform inference quickly while maintaining high accuracy. The P2 layer in the YOLOv8-P2 network is specifically optimized for small target recognition, improving the detection capability for small-sized targets.
[0033] 3. This invention combines the texture information of natural images with the precise 3D position information of LiDAR, improving the robustness of the detection system under different lighting and weather conditions. When handling the coordinate transformation from a bird's-eye view to a front view, a traditional machine learning regression model based on K-nearest neighbors is used, simplifying the transformation process and reducing the large amount of computational resources that deep learning models might require. The Gaussian-weighted Soft-NMS algorithm effectively suppresses redundant detection boxes, improving the accuracy of the detection results.
[0034] 4. This invention can effectively improve the performance of target detection in the fields of autonomous driving and target detection, and can be widely used. Attached Figure Description
[0035] Figure 1 This is a flowchart illustrating the present invention.
[0036] Figure 2 The input consists of three modalities of data: natural image (a), sparse depth image (b), and bird's-eye view image (c).
[0037] Figure 3 It shows three modalities of image data and the detection results after fusion. Detailed Implementation
[0038] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments, but the present invention is not limited to the following embodiments.
[0039] This embodiment uses vehicles as the detection target.
[0040] exist Figure 1 The target detection method based on multimodal data fusion in this embodiment includes the following steps:
[0041] Step 1. Data Preprocessing
[0042] Using the publicly available KITTI Object Detection Evaluation 2012 dataset, which includes a natural image dataset and a corresponding LiDAR point cloud dataset, the natural image dataset was divided into a training set and a validation set in a 1:1 ratio. The LiDAR point cloud dataset was projected onto a two-dimensional plane to generate a sparse depth image set and a bird's-eye view image set. The sparse depth image set was divided into a training set and a validation set in a 1:1 ratio, and the bird's-eye view image set was divided into a training set and a validation set in a 1:1 ratio.
[0043] The method for projecting a LiDAR point cloud dataset onto a two-dimensional plane to generate sparse depth image sets and bird's-eye view image sets is as follows: First, the LiDAR point cloud data is transformed into the coordinates of the camera coordinate system points through rigid body transformation. Then, the coordinates of the image coordinate system points are obtained through perspective projection. Finally, the coordinates of the points are transformed into the coordinates of the two-dimensional pixel coordinate system through affine transformation. Specifically, the pixel values in the sparse depth image are obtained by calculating the longitudinal distance from the original LiDAR points to the origin in the new coordinate system. The pixel values in the bird's-eye view image are obtained by encoding the density, reflection intensity, and height of the LiDAR point cloud in the RGB three color channels.
[0044] Step 2. Construct the object detection network
[0045] The YOLOv8-P2 network was used as the object detection network to detect vehicles in images of different modalities.
[0046] Step 3. Train the object detection network YOLOv8-P2
[0047] The training sets of the natural image dataset, the bird's-eye view image dataset, and the sparse depth image dataset are respectively input into the object detection network YOLOv8-P2. The object detection network YOLOv8-P2 is trained and optimized using the loss function L, and the weight file obtained after training each training set is saved.
[0048] The training parameters are set as follows: momentum m is 0.937, weight decay is 0.0005, training batch n is 16, the initial learning rate is 0.001, and the number of training iterations is 400.
[0049] The loss function L is composed of the classification loss L BCE And bounding box regression loss L Bbox_Loss It consists of two parts;
[0050] Classification loss L BCE for:
[0051]
[0052] In the formula, N is the total number of samples, and y j p is the category label of the j-th sample, where the positive class label is 1 and the negative class label is 0; j It is the probability that the j-th sample is predicted as positive.
[0053] Boundary regression loss L Bbox_Loss for:
[0054] L Bbox_Loss =L DFL +L WIoU
[0055] L DFL =-((y) i+1 -y)log(w i )+(yy i )log(w i+1 ))
[0056]
[0057] In the formula, y i and y i+1 These are the two values closest to the true label y, y i ≤y≤y i+1 L IoU = 1 - IoU, where IoU is the intersection-over-union ratio between the detected bounding box and the ground truth bounding box, and (x,y) are the x and y coordinates of the center point of the predicted bounding box. gt ,y gt ) represents the x and y coordinates of the center point of the true bounding box, W g and H g These are the width and height of the actual bounding box, respectively. The width and height of the actual bounding box are separated from the computation graph. The cross-union ratio (CUB) loss between the detection boxes and the ground truth bounding boxes separates them from the computational graph; α and δ are hyperparameters. It is a moving average with momentum m. n is the training batch size, and t is the batch value when the rate of improvement of average detection accuracy during training slows down significantly.
[0058] Step 4. Generate target detection results in different modalities
[0059] The validation sets from the natural image dataset, the bird's-eye view image dataset, and the sparse depth image dataset are respectively input into the trained object detection network YOLOv8-P2, such as... Figure 2 The target detection results of natural images, bird's-eye view images, and sparse depth images are obtained.
[0060] Step 5. Overlay the target detection results of the sparse depth image (both in the front view direction) and the target detection results of the natural image, and remove redundant target detection boxes using the Soft-NMS model to obtain the initial fusion result;
[0061] Step 6. The target detection results from the bird's-eye view image are transformed into coordinate information using a multi-output regression model based on the K-nearest neighbor algorithm. This transformation results in a direction consistent with the view orientation of the target detection results from the sparse depth image and the natural image, i.e., the front view orientation. This is then overlaid with the initial fusion result, and redundant target detection boxes are removed using the Soft-NMS model to obtain the final target detection result, such as... Figure 3 In the multi-output regression model based on the K-nearest neighbors algorithm, the number of leaf nodes is initialized to 50.
[0062] The method for removing redundant target detection boxes using the Soft-NMS model is as follows:
[0063] Step A1. Initialize the detection box position information set B = {b1,…,b N} and the corresponding confidence scores S={s1,…,s N} and the set Intersection over Union (IoU) threshold N t =0.7;
[0064] Step A2. Define an empty set D to store the location information of the highest-scoring detection boxes to be retained;
[0065] Step 3. When the detection box position information set At that time, sort all bounding boxes according to their confidence scores S, assign the bounding box with the highest score to M, assign the union of D and M to D, then remove M from set B, and iterate through all the remaining bounding boxes in set B to calculate the Intersection over Union (IOU) with M, and update the confidence score of each bounding box according to the following formula:
[0066]
[0067] In the formula, S kS′ is the initial confidence score of the k-th bounding box. k Here, M represents the confidence score after weight decay update, and b is the location of the bounding box with the highest confidence score. k is the position of the k-th bounding box, IOU(·) is the intersection-union ratio of the areas of two bounding boxes, and σ is the standard deviation of the Gaussian function, which determines the distribution of the Gaussian weights. A smaller σ will lead to faster score decay, while a larger σ will lead to a more gradual score decay. In this example, σ = 0.35.
[0068] experiment
[0069] To verify the superior detection performance of this invention, the target detection method based on multi-source data fusion was compared with various target detection methods, including Faster-RCNN, YOLOv8, MV3D (Lidar), Complex-YOLO, AVOD, DMF, and CMAN. The input data for these methods were RGB natural images, LiDAR point cloud data, and multimodal data, respectively. The target detection performance of each method in Easy, Moderate, and Hard scenarios is shown in the table below. In the table, AP represents the detection accuracy, and mAP represents the average detection accuracy; a higher mAP value indicates better performance.
[0070] Table 1 compares the detection performance (AP) of various detection methods on the KITTI dataset.
[0071]
[0072] As can be seen, the average detection accuracy of this invention in the three scenarios is 86.19% (No BEV) and 86.89% (Add BEV), respectively, which is superior to other detection methods. Compared with many other detection methods, the method presented in this paper achieves a balance and advantage between detection accuracy and detection speed.
Claims
1. A target detection method based on multimodal data fusion, characterized in that, Includes the following steps: Step 1. Data Preprocessing Collect a natural image dataset containing the target and a corresponding LiDAR point cloud dataset. Divide the natural image dataset into a training set and a validation set according to the proportion. Project the LiDAR point cloud dataset onto a two-dimensional plane to generate a sparse depth image set and a bird's-eye view image set. Divide the sparse depth image set into a training set and a validation set according to the proportion. Divide the bird's-eye view image set into a training set and a validation set according to the proportion. Step 2. Construct the object detection network The YOLOv8-P2 network was used as the object detection network to detect target objects in images of different modalities. Step 3. Train the object detection network YOLOv8-P2 The training sets of the natural image dataset, the bird's-eye view image dataset, and the sparse depth image dataset are respectively input into the object detection network YOLOv8-P2. The object detection network YOLOv8-P2 is trained and optimized using the loss function L, and the weight file obtained after training each training set is saved. The loss function L is composed of the classification loss L BCE And bounding box regression loss L Bbox_Loss It consists of two parts; The classification loss L BCE for: In the formula, N is the total number of samples, and y j p is the category label of the j-th sample, where the positive class label is 1 and the negative class label is 0; j It is the probability that the j-th sample is predicted as positive. The bounding box regression loss L Bbox_Loss for: L Bbox_Loss L DFL +L WIoU L DFL =-((y i+1 -y)log(w i )+(y-y i )log(w i+1 )) In the formula, y i and y i+1 These are the two values closest to the true label y, y i ≤y≤y i+1 L IoU = 1 - IoU, where IoU is the intersection-over-union ratio between the detected bounding box and the ground truth bounding box, and (x,y) are the x and y coordinates of the center point of the predicted bounding box. gt ,y gt ) represents the x and y coordinates of the center point of the true bounding box, W g and H g These are the width and height of the actual bounding box, respectively. The width and height of the actual bounding box are separated from the computation graph. The cross-union ratio (CUB) loss between the detection boxes and the ground truth bounding boxes separates them from the computational graph; α and δ are hyperparameters. It is a moving average with momentum m. n is the training batch size, and t is the batch value when the rate of improvement of average detection accuracy during training slows down significantly. Step 4. Generate target detection results in different modalities The validation sets of the natural image dataset, the bird's-eye view image dataset, and the sparse depth image dataset are respectively input into the trained object detection network YOLOv8-P2 to obtain the object detection results of the natural image, the bird's-eye view image, and the sparse depth image. Step 5. Overlay the target detection results of the sparse depth image and the target detection results of the natural image in the same front view direction, and remove redundant target detection boxes using the Soft-NMS model to obtain the initial fusion result; Step 6. The target detection results of the bird's-eye view image are transformed into coordinate information using a multi-output regression model based on the K-nearest neighbor algorithm, so that the view direction is the same as that of the target detection results of the sparse depth image and the natural image. Then, the coordinate information is superimposed with the initial fusion result and redundant target detection boxes are removed using the Soft-NMS model to obtain the final target detection result.
2. The target detection method based on multimodal data fusion according to claim 1, characterized in that, The method for projecting the LiDAR point cloud dataset onto a two-dimensional plane to generate sparse depth image sets and bird's-eye view image sets is as follows: First, the LiDAR point cloud data is transformed into the coordinates of the camera coordinate system points through rigid body transformation. Then, the coordinates of the image coordinate system points are obtained through perspective projection. Finally, the coordinates of the points are transformed into the coordinates of the two-dimensional pixel coordinate system through affine transformation. Specifically, the pixel values in the sparse depth image are obtained by calculating the longitudinal distance from the original LiDAR point to the origin of the current LiDAR coordinate system. The pixel values in the bird's-eye view image are obtained by encoding the density, reflection intensity, and height of the LiDAR point cloud in the RGB three color channels.
3. The target detection method based on multimodal data fusion according to claim 1, characterized in that, In step 3, the parameters of the YOLOv8-P2 target detection network are set as follows: momentum m is 0.937, weight decay is 0.0005, training batch n is 16, and the initial learning rate is 0.
001.
4. The target detection method based on multimodal data fusion according to claim 1, characterized in that, The Soft-NMS model formulas in steps 5 and 6 are as follows: In the formula, S k S′ is the initial confidence score of the k-th bounding box. k Here, M represents the confidence score after weight decay update, and b is the location of the bounding box with the highest confidence score. k is the position of the k-th bounding box, D is the set of bounding boxes with the highest scores, IOU(·) is the intersection-union ratio of the areas of two bounding boxes, and σ is the standard deviation of the Gaussian function.