A 3D Vehicle Target Detection Method Combining Image Texture Features and Prior Information

By combining image texture features and prior information, combined with CascadeRCNN and improved PointPillars network model, the problem of insufficient accuracy and robustness of three-dimensional vehicle object detection in the prior art is solved, and high-precision and stable detection in complex environments are achieved.

CN119580201BActive Publication Date: 2025-07-08CHANGAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411616290.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-07-08
Estimated Expiration
2044-11-13

AI Technical Summary

Technical Problem

The existing three-dimensional vehicle object detection method relies on deep learning technology, making it difficult to effectively deal with the attitude, angle and occlusion problems of the target object, and the vehicle's prior information is not fully utilized to integrate with the image texture features, resulting in insufficient detection accuracy and robustness.

Method used

A 3D vehicle object detection method that fuses image texture features and prior information is adopted. Through pixel-level point cloud fusion and image fusion, the CascadeRCNN algorithm is used to perform image two-dimensional object detection, and the PointPillars network model is improved for point cloud three-dimensional object detection. Combining the decision-making fusion scheme of feature distance and channel attention, the final three-dimensional vehicle object detection result is generated.

Benefits of technology

It significantly improves the accuracy and stability of three-dimensional vehicle detection, especially in complex environments, and meets the real-time needs of autonomous driving and intelligent transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580201B_ABST
    Figure CN119580201B_ABST
Patent Text Reader

Abstract

The present invention provides a 3D vehicle object detection method that fuses image texture features and prior information, including obtaining a dataset, performing pixel-level point cloud and image fusion based on the point cloud data and image data in the dataset, detecting the image data based on a preset image object detector to generate image two-dimensional object bounding boxes, detecting the fused point cloud data based on a preset 3D vehicle object detection model to generate point cloud three-dimensional object bounding boxes, and fusing the image two-dimensional object bounding boxes and the point cloud three-dimensional object bounding boxes based on a decision-level fusion scheme of feature distance and channel attention to obtain the final 3D vehicle object detection result. By fusing image texture features and prior geometric information, the present invention significantly improves the accuracy and stability of 3D vehicle detection, especially performs well in complex environments. The use of an efficient combination technology of geometric optimization and depth information ensures the real-time performance of the system, meeting the actual application requirements of autonomous driving and intelligent transportation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional vehicle target detection, and particularly to a 3D vehicle target detection method that fuses image texture features and prior information. Background Art

[0002] With the rapid development of autonomous driving and intelligent transportation systems, accurate and real-time three-dimensional vehicle target detection technology has become particularly important. Currently, common target detection methods mainly rely on deep learning technology, detect based on two-dimensional images, and then obtain three-dimensional information through post-processing steps. However, such methods usually handle problems such as the pose, angle, and occlusion of target objects poorly. In addition, existing methods rarely make full use of the prior information of vehicles and the fusion of image texture features for detection, resulting in insufficient detection accuracy and robustness. Therefore, it is very necessary to design a 3D vehicle target detection method that fuses image texture features and prior information. Summary of the Invention

[0003] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a 3D vehicle target detection method that fuses image texture features and prior information.

[0004] To achieve the above purpose, the present invention provides the following solutions:

[0005] The present invention provides a 3D vehicle target detection method that fuses image texture features and prior information, including:

[0006] Obtain a data set, and perform pixel-level point cloud and image fusion based on the point cloud data and image data in the data set;

[0007] Detect the image data based on a preset image target detector to generate image two-dimensional target boxes;

[0008] Detect the fused point cloud data based on a preset three-dimensional vehicle target detection model to generate point cloud three-dimensional target boxes;

[0009] Fuse the image two-dimensional target boxes and the point cloud three-dimensional target boxes based on a decision-level fusion scheme of feature distance and channel attention to obtain the final three-dimensional vehicle target detection result.

[0010] Preferably, the data set is the KITTI data set.

[0011] Preferably, performing pixel-level point cloud and image fusion based on the point cloud data and image data in the data set is specifically:

[0012] Obtain the prior information of image targets from the 2D target ground truth label file of the KITTI dataset. Based on the prior information of image targets, obtain the target categories and target box ranges in the image. After mapping the point cloud to the image, filter out the point cloud within the image target box according to the target box range. Based on the method of assigning category information based on Gaussian distance coding, judge the probability that the point cloud within the frustum belongs to the foreground target points through Gaussian distance feature information, and assign the probability as the category score to the point cloud data, and assign the target category information to these point clouds. On this basis, assign the corresponding pixel point R, G, B color values to each point cloud to generate point cloud data with rich feature information.

[0013] Preferably, detect the image data based on a preset image target detector to generate an image 2D target box, specifically:

[0014] Construct an image target detector based on the CascadeRCNN algorithm;

[0015] Train the image target detector based on a preset training dataset;

[0016] Input the image into the trained image target detector for detection to generate an image 2D target box.

[0017] Preferably, detect the fused point cloud data based on a preset 3D vehicle target detection model to generate a point cloud 3D target box, specifically:

[0018] Construct a 3D vehicle target detection model based on the improved PointPillars network model;

[0019] Train the 3D vehicle target detection model based on the KITTI dataset to obtain the trained 3D vehicle target detection model;

[0020] Input the fused point cloud data into the trained 3D vehicle target detection model to generate a point cloud 3D target box.

[0021] Preferably, construct a 3D vehicle target detection model based on the improved PointPillars network model, specifically:

[0022] Improve the PointPillars network model by adding a point cloud local attention mechanism module to the voxel feature network of the traditional PointPillars network model to capture complex features within a specific region of the input point cloud, and integrating an SE attention mechanism module into the 2D pseudo-image network of the traditional PointPillars network model to enhance the network's ability to obtain global features and information;

[0023] Construct a 3D vehicle target detection model based on the improved PointPillars network model.

[0024] Preferably, a decision-level fusion scheme based on feature distance and channel attention fuses the two-dimensional image object bounding box and the three-dimensional point cloud object bounding box to obtain the final three-dimensional vehicle object detection result. Specifically:

[0025] Encode the two-dimensional image object bounding box and the three-dimensional point cloud object bounding box respectively, fuse the class confidence and the object bounding box position, and generate their respective encoded tensors;

[0026] Based on semantic consistency and geometric consistency, fuse the encoded tensors of the two-dimensional image object bounding box and the three-dimensional point cloud object bounding box to obtain a fused feature tensor;

[0027] Input the fused feature tensor into a convolutional neural network structure encoded based on channel attention to finally obtain a fused confidence score;

[0028] Re-distribute the fused confidence score to the three-dimensional point cloud object bounding box, and use the non-maximum suppression (NMS) algorithm to filter redundant and misdetected object bounding boxes to generate the final three-dimensional vehicle object detection result.

[0029] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:

[0030] The present invention provides a 3D vehicle object detection method that fuses image texture features and prior information. The method includes obtaining a data set, performing pixel-level point cloud and image fusion based on the point cloud data and image data in the data set, detecting the image data based on a preset image object detector to generate a two-dimensional image object bounding box, detecting the fused point cloud data based on a preset three-dimensional vehicle object detection model to generate a three-dimensional point cloud object bounding box, and fusing the two-dimensional image object bounding box and the three-dimensional point cloud object bounding box based on a decision-level fusion scheme of feature distance and channel attention to obtain the final three-dimensional vehicle object detection result. By fusing image texture features and prior geometric information, the present invention significantly improves the accuracy and stability of three-dimensional vehicle detection, especially performing well in complex environments. In addition, the use of an efficient combination technology of geometric optimization and depth information ensures the real-time performance of the system, meeting the actual application requirements of autonomous driving and intelligent transportation. Description of the Drawings

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0032] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention;

[0033] Figure 2 It is a schematic diagram of the KITTI lidar coordinate system and the camera coordinate system;

[0034] Figure 3 It is a result diagram of the fusion of point cloud and image data;

[0035] Figure 4 It is a point cloud diagram of the vehicle target frustum;

[0036] Figure 5 It is a point cloud visualization scene diagram under the forward view (the red color indicates the target point cloud);

[0037] Figure 6 It is a point cloud visualization diagram of RGB information fusion based on the target box level;

[0038] Figure 7 It is a point cloud visualization diagram of RGB information fusion based on the scene;

[0039] Figure 8 It is a schematic diagram of the PointPillars network algorithm framework;

[0040] Figure 9 It is a schematic diagram of the improved PointPillars network algorithm framework;

[0041] Figure 10 It is a geometric consistency relationship diagram under the same target;

[0042] Figure 11 It is a schematic diagram of the convolutional network structure based on channel attention coding;

[0043] Figure 12a It is a visualization result diagram of some scenes of other algorithms;

[0044] Figure 12b It is a visualization result diagram of some scenes of the algorithm of the present invention. Specific implementation manners

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] The purpose of the present invention is to provide a 3D vehicle target detection method that fuses image texture features and prior information, which can realize three-dimensional vehicle target detection based on the fused graph texture features and prior information, and is convenient to use.

[0047] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0048] Figure 1 The flowchart of the method provided by the embodiment of the present invention is as Figure 1 shown. The present invention provides a 3D vehicle target detection method that fuses image texture features and prior information, including:

[0049] Step 100: Obtain a dataset, and perform pixel-level point cloud and image fusion based on the point cloud data and image data in the dataset;

[0050] Step 200: Detect the image data based on a preset image target detector to generate an image two-dimensional target box;

[0051] Step 300: Detect the fused point cloud data based on a preset 3D vehicle target detection model to generate a point cloud three-dimensional target box;

[0052] Step 400: Fuse the image two-dimensional target box and the point cloud three-dimensional target box based on a decision-level fusion scheme of feature distance and channel attention to obtain the final 3D vehicle target detection result.

[0053] The dataset is the KITTI dataset. KITTI uses an on-vehicle lidar to collect point cloud data and image data. Since the collected image data does not have depth information, it is difficult to map the image to the point cloud scene using traditional methods. For the fusion method of point cloud and image data, most current methods map the point cloud to the image, reduce the three-dimensional data to two-dimensional data, and retain the depth information. To achieve coordinate transformation, the KITTI dataset provides the calibration parameters of the camera and lidar for each frame of data. According to the calibration parameters, the transformation matrix from the three-dimensional lidar coordinate system to the two-dimensional image coordinate system can be calculated to achieve point cloud mapping. Since the lidar point cloud sampling range is relatively wide and the field of view of the entire point cloud scene is 360°, considering the field of view range of the camera-sampled image and the point cloud distribution in the real traffic scene, the point cloud outside the field of view angle of the image is discarded during the point cloud projection process, and only the point cloud in the forward viewing angle is retained. Figure 2 Shows the axis directions of the lidar coordinate system and the camera coordinate system in the KITTI dataset, where the X-axis of the lidar coordinate system represents forward, the Y-axis represents left, and the Z-axis represents up; the X-axis of the camera coordinate system represents right, the Y-axis represents down, and the Z-axis represents forward.

[0054] (1) Conversion of lidar point cloud coordinates to camera point cloud coordinates

[0055] To convert the point cloud coordinates from the LiDAR coordinate system to the camera coordinate system, it is necessary to use the calibration extrinsic parameter matrix, i.e., the Tr_velo_to_cam transformation matrix, which contains the rotation matrix R and the translation matrix t. This matrix is a 3×4 matrix, and after removing the reflection intensity from the input LiDAR point cloud coordinates, it is an N×3 matrix containing the number of point clouds N and the original three-dimensional coordinate points in the LiDAR coordinate system. To ensure that the point cloud coordinates in the output camera coordinate system are also an N×3 matrix, it is necessary to expand the dimension of the original point cloud, fill it with a column of 1s and transform it into an N×4 matrix, and multiply it by the transposed extrinsic parameter matrix to obtain the point cloud coordinates in the camera coordinate system. Since KITTI provides a total of 4 cameras, and the Tr_velo_to_cam matrix only converts the LiDAR point cloud coordinates to the point cloud coordinates in the 0th camera coordinate system. To ensure that the converted point cloud coordinates can be used by other cameras, it is necessary to multiply by the corrected rotation matrix R0_rect of the 0th camera. The function of this matrix is to make the images of all cameras coplanar and ensure that the optical centers of the 4 cameras are on the same xoy plane. The overall calculation process is shown in the following formula:

[0056]

[0057] Among them, rect0 represents the point cloud coordinate matrix in the uncorrected 0th camera coordinate system, Lidar represents the point cloud matrix in the radar coordinate system, R represents the rotation matrix with a size of 3×3, t represents the translation matrix with a size of 3×1, rect_R0 represents the point cloud coordinate matrix in the corrected 0th camera coordinate system, and in the calculation process, the [R|t] matrix needs to be transposed with the R0_rect matrix.

[0058] (2) Conversion of camera point cloud coordinates to image point cloud coordinates

[0059] To convert the point cloud coordinates from the camera coordinate system to the image coordinate system, it is necessary to use the camera intrinsic calibration matrix, i.e., the P i matrix, and the P i matrix formula is shown in the following formula:

[0060]

[0061] Among them, f u and f v refer to the focal lengths of the camera, c u and c v refer to the offsets from the optical center to the zero point of the CMOS, i.e., the coordinates of the camera optical center in the image coordinate system, and b x refers to the distance offset in the x-axis direction from the current camera to the 0th camera.

[0062] Using the No. 2 color camera as the conversion standard, the point cloud coordinates in the camera coordinate system obtained by calculation are multiplied by the internal parameter P2 of the No. 2 camera to generate the point cloud coordinates in the image coordinate system. Since P2 is a 3×4 matrix and the point cloud coordinates in the output camera coordinate system are an N×3 matrix, it is necessary to fill the point cloud coordinates with 1 in the column to generate an N×4 matrix and multiply it by the transpose matrix of P2 to obtain the point cloud coordinate result in the image coordinate system. To obtain the homogeneous coordinates in the image coordinate system, it is necessary to normalize in the z-axis depth direction. Divide the point cloud coordinates in each image coordinate system by the depth information to obtain the final coordinate result, as shown in the following formula:

[0063]

[0064] Among them, cam represents the point cloud coordinate matrix in the image coordinate system, and cam / z represents the normalized point cloud matrix. Generate the image depth map according to the point cloud coordinates in the calculated image coordinate system. First, filter out the point clouds that are not in the image view according to the length and width of the image. At the same time, since the z-axis in the camera coordinate system represents forward, it is necessary to filter out the point clouds with a z-axis depth value less than 0, and these point clouds are not in the forward view; to make each point cloud an effective value and since the coordinates of each pixel point in the image are integers, it is necessary to round the point cloud coordinates to ensure that each point cloud can be correctly mapped to the image pixel points. The calibration file provided by KITTI is shown in Table 1, and the image and point cloud data fusion diagram is as Figure 3 shown;

[0065] Table 1 Example format of the calibration file for the KITTI dataset

[0066]

[0067]

[0068] In step 100, based on the point cloud data and image data in the dataset, pixel-level point cloud and image fusion is performed. Specifically:

[0069] Obtain the prior information of the image target from the two-dimensional target ground truth label file of the KITTI dataset. According to the prior information of the image target, obtain the target category and the target box range in the image. After mapping the point cloud to the image, filter out the point clouds located within the image target box according to the target box range. Based on the category information assignment method based on Gaussian distance coding, judge the probability that the point cloud in the frustum belongs to the foreground target point through the Gaussian distance feature information, and assign the probability as the category score to the point cloud data, and assign the target category information to these point clouds. On this basis, assign the corresponding pixel point R, G, B color values to each point cloud to generate point cloud data with rich feature information. The specific process is introduced as follows:

[0070] First, the pixel-level fusion based on target prior information and Gaussian coding is introduced. PointPainting first uses a semantic segmentation network to perform semantic segmentation on the original image to obtain the category and category score of each pixel. After mapping the point cloud to the image, according to the correspondence between the pixel and the mapped point cloud, the pixel category score is attached to the corresponding point cloud to generate 8-dimensional point cloud data of (x, y, z, r, s1, s2, s3, s4). The generated new point cloud data is put into a 3D object detector for 3D object detection. Since this algorithm requires a high coupling between the semantic segmentation model and the 3D object detection model, the applicable range of the algorithm is limited. To solve this problem, the present invention proposes a method for fusing point cloud and image based on image target prior information, which uses the prior result obtained from image object detection to expand the information of the original point cloud and generate point cloud data containing rich information. Since both 2D object detection and 3D object detection belong to the object detection task, a low-coupling combination is achieved between them. And during the training process, the KITTI dataset provides the 2D object ground truth label file, and the training data can be directly generated according to the image target prior information given in this file, without reusing the 2D detection model to perform object detection on the training image to generate results and then perform point cloud fusion, which enhances the practicability of the algorithm and reduces the difficulty of data fusion.

[0071] The image target prior information provides the target category and target box range in the current image. After mapping the point cloud to the image, the point cloud located within the image target box is screened out according to the target box range, and the target category information is assigned to these point clouds, which can provide more detailed information during the subsequent 3D object detection process. Since the point cloud data collected by the lidar is in 3D space, while the image data collected by the camera is in 2D space, when the point cloud is mapped to the image, the depth information is lost, which is equivalent to compressing the 3D plane to a 2D plane. As a result, the point cloud within the screened image target box presents the shape of a 3D frustum after being mapped back to 3D space, as Figure 4 shown. It can be seen from the figure that the point cloud within the frustum can completely present the shape of the current target. However, due to the compression in the depth direction, a large number of background points are also within the 2D target box. Therefore, assigning the same numerical category score to the point cloud within this frustum will generate a large number of background noise point clouds. To solve the above problems, this section designs a method for assigning category information based on Gaussian distance coding, which discriminates the probability that the point cloud within the frustum belongs to the foreground target point through Gaussian distance feature information, and uses the probability as the category score to assign to the point cloud data to enrich the original point cloud information.

[0072] Through observation, it is found that the coordinate centers of the two-dimensional target boxes are all located on the detected targets. The point clouds within a certain range of the center points are likely to be target point clouds. Therefore, in this paper, the Gaussian distance is used as a standard to measure the position distance between the point clouds within the target box and the center point of the two-dimensional target box. The larger the Gaussian distance, the closer the point cloud is to the center of the target box in the two-dimensional image, and the more likely it is to be a target point cloud. The smaller the distance, the farther the point cloud is from the center of the target box in the two-dimensional image, and the more likely it is to be a background point cloud. The Gaussian distance calculation formula is shown as follows:

[0073]

[0074] Among them, (x,y) represents the coordinates of the point cloud mapped into the image target box, is the center coordinate of the two-dimensional target box, and w and h respectively represent the width and height of the two-dimensional target box. The present invention only detects vehicle targets. Therefore, during the generation of category information, non-vehicle targets are regarded as background points, and the calculated Gaussian distance information is assigned to the point cloud as the category information of the current target, generating 6-dimensional point cloud data of (x,y,z,r,s1,s2). Among them, s1 represents the probability that the current point cloud is a vehicle target point cloud, and s2 represents the probability that the current point cloud is a background point cloud. Mapping the extended point cloud data back to the three-dimensional space, a visualization scene graph as shown in Figure 5 is obtained. For the convenience of observation, only the point clouds in the forward view are visualized. The red part in the figure represents the point clouds under the frustum, that is, the point cloud data within the image target detection box, and the black part represents the background point cloud data.

[0075] Secondly, the pixel-level fusion based on the RGB texture features of the image is introduced:

[0076] Image data, as a two-dimensional data containing rich texture and color information, is mainly represented by a two-dimensional matrix composed of several pixel points, and each pixel point contains three color channels R, G, and B. Point cloud data, as a data set containing several point clouds, only provides the three-dimensional coordinates x, y, z and reflection intensity r of the original point clouds. In the two-dimensional target detection task, RGB information provides sufficient image information for the convolutional neural network, ensuring the accuracy of target detection. Based on this, the present invention designs a fusion method based on the RGB texture features of the image on the basis of fusing the prior information of the image target, assigns the corresponding pixel point R, G, B color values to each point cloud, and combines the prior result information of the image target to generate point cloud data with rich feature information, including 9-dimensional data of (x,y,z,r,s1,s2,R,G,B);

[0077] The present invention provides two strategies for fusing RGB texture information of images and point clouds. According to different fusion regions, they are divided into the RGB information fusion strategy based on the target box level (Method 1) and the RGB information fusion strategy based on the scene (Method 2):

[0078] Method 1: Assign RGB information to several target frustum point clouds obtained from the prior information of image targets, aiming to enable the subsequent 3D object detector to fully focus on the point clouds in the target region and achieve the tasks of object classification and bounding box regression, as Figure 6 shown;

[0079] Method 2: Assign RGB information to the point clouds from the perspective of the entire scene, and assign values to all point clouds within the image perspective range, aiming to ensure that the point clouds have sufficient context structure, so that the subsequent 3D object detector can extract the overall scene during feature extraction and prevent the loss of some key point cloud information in the design of Method 1, as Figure 7 shown.

[0080] In step 200, the image data is detected based on a preset image object detector to generate an image two-dimensional target box, specifically:

[0081] Construct an image object detector based on the CascadeRCNN algorithm;

[0082] Train the image object detector based on a preset training dataset;

[0083] Input the image into the trained image object detector for detection to generate an image two-dimensional target box.

[0084] In step 300, the fused point cloud data is detected based on a preset 3D vehicle object detection model to generate a point cloud three-dimensional target box, specifically:

[0085] Construct a 3D vehicle object detection model based on the improved PointPillars network model;

[0086] Train the 3D vehicle object detection model based on the KITTI dataset to obtain the trained 3D vehicle object detection model;

[0087] Input the fused point cloud data into the trained 3D vehicle object detection model to generate a point cloud three-dimensional target box.

[0088] Construct a 3D vehicle object detection model based on the improved PointPillars network model, specifically:

[0089] Improve the PointPillars network model by adding a point cloud local attention mechanism module to the pillar feature network of the traditional PointPillars network model to capture complex features in a specific area of the input point cloud, and integrating an SE attention mechanism module into the two-dimensional pseudo-image network of the traditional PointPillars network model to enhance the network's ability to obtain global features and information;

[0090] Construct a 3D vehicle target detection model based on the improved PointPillars network model;

[0091] A detailed introduction to this process is as follows:

[0092] As Figure 8 shown, the main components of the traditional PointPillars network model include a pillar feature network (Pillar Feature Net), a backbone network (Backbone), and an SSD detection head (Detection Head). First, preprocess and extract features from the point cloud data, then project the data into the bird's-eye view through the PFN module and generate the feature representation of each voxel. PSE is used to encode the position information of the voxels to improve the network's understanding ability of the spatial structure. Then, the target proposal generator generates candidate target detection boxes based on the feature map, and finally the target detection head predicts the presence, category, and location information of the target based on these boxes.

[0093] The improvement of the traditional PointPillars network model in the present invention is as follows: First, a point cloud local attention mechanism network for extracting rich features from a single point cloud is added to the point cloud pillar feature extraction module. This attention mechanism allows for a more comprehensive fusion of the feature information of the surrounding point cloud during the feature extraction stage. Second, a channel attention mechanism module is added to the two-dimensional pseudo-image backbone network. The final features are input into the SSD algorithm for target detection, and the overall network structure is as shown in Figure 9.

[0094] The main function of the point cloud local attention module is to enhance the representation ability of point cloud data by introducing an attention mechanism. It enables the network to learn important information from the features of the point cloud while reducing the acquisition of irrelevant information. The input point cloud data is P, and each point cloud contains its own position coordinates (x, y, z). The point cloud to be tested undergoes a series of convolution operations to convert the input data into query, key, and value tensors with different channel dimensions. The similarity between the query and the key is then calculated, and the values ​​are weighted summed using the similarity to obtain the final output representation. The code divides the input data into local regions and adjusts the values ​​of the query, key, and value tensors accordingly. Then the attention score is calculated to determine the importance of the elements in each local region. After further convolution and reshaping, the final output features are generated. This local attention module captures the local neighbor point information of a specific point cloud in the body column and outputs the features of the key information of the local point cloud sampling points around the point cloud.

[0095] The present invention adopts the idea of ​​SE attention mechanism and adds the channel attention mechanism module to the two-dimensional pseudo image backbone network. Specifically, the BaseBackbone model class is defined in Pytorch, and the channel attention mechanism sub-network is added. First, the input feature map is reduced in dimension, and the input three-dimensional tensor (C, H, W) is reduced in dimension to (1, 1, C) using the global average pooling layer, so that the tensor can be regarded as a channel dimension descriptor. Next, through two linear layers and the ReLU activation function, the nonlinear interaction between channels is learned. Finally, the channel weight is mapped to between 0 and 1 through the Sigmoid activation function to form a channel attention weight.

[0096] The detection module uses a Single Shot MultiBox Detector (SSD) for object detection. This is a widely used one-stage detection algorithm, known for its fast detection speed and high accuracy. In the SSD network, the concept of anchors is introduced to adapt to multi-scale object detection tasks, making it more suitable for capturing large-scale transformations in point cloud data. The SSD architecture consists of six main modules. The first module includes the initial five layers (Conv1 - Conv5) of the VGG16 convolutional network. Subsequently, the FC6 layer of VGG16 is incorporated, and the FC7 fully connected layer is converted into the Conv6 layer. The Conv7 layer forms the second module. To extract object information at different scales, four additional modules are introduced, namely the Conv8, Conv9, Conv10, and Conv11 convolutional layers. The final stage involves object classification, detection, and non-maximum suppression of the regression positions. Non-maximum suppression uses the Intersection over Union (IoU) technique to match the prior bounding boxes with the actual objects, ensuring that only the most relevant bounding boxes are retained and redundant detections are eliminated.

[0097] The decision-level fusion scheme based on feature distance and channel attention fuses the 2D object bounding boxes of the image and the 3D object bounding boxes of the point cloud to obtain the final 3D vehicle object detection result. Specifically:

[0098] Encode the 2D object bounding boxes of the image and the 3D object bounding boxes of the point cloud respectively, fuse the class confidence and the object bounding box positions, and generate their respective encoded tensors;

[0099] Based on semantic consistency and geometric consistency, fuse the encoded tensors of the 2D object bounding boxes of the image and the 3D object bounding boxes of the point cloud to obtain the fused feature tensor;

[0100] Input the fused feature tensor into a convolutional neural network structure encoded based on channel attention to finally obtain the fused confidence score;

[0101] Reassign the fused confidence score to the 3D object bounding boxes of the point cloud, and use the non-maximum suppression (NMS) algorithm to filter redundant and misdetected object bounding boxes to generate the final 3D vehicle object detection result;

[0102] Introduce the above process in detail. First, introduce the association between semantic consistency and geometric consistency as:

[0103] The decision-level fusion scheme independently detects the data collected by each sensor, and fuses the detection results of each part through a fusion module. In this fusion scheme, the data of different sensors are independently processed, and there is no interference between sensors. That is, when one sensor fails, it does not affect the overall detection process, improving the robustness of the detection. This scheme includes three modules in the process of point cloud and image fusion: image target detection, point cloud target detection, and target box fusion module. The present invention generates a fused point cloud by fusing the RGB texture features of the image and the prior information of the image target, uses the improved PointPillars network model to detect the vehicle target in the point cloud for the fused point cloud, and at the same time uses the existing CascadeRCNN algorithm as the image target detector to detect the vehicle target in the image. The decision-level fusion is performed on the two different detection results before the non-maximum suppression algorithm. To ensure the effectiveness of the fusion module and improve the target detection accuracy, based on the CLOCs algorithm, the present invention designs a fusion module based on feature distance and channel attention coding;

[0104] The detection results obtained by different sensors are not correlated with each other in terms of class confidence scores. Starting from the perspectives of geometric consistency and semantic consistency, the present invention encodes the image target detection results and the point cloud target detection results into feature tensors with consistent associations, extracts and pools the features of the feature tensors, and remaps the generated probability scores back into the three-dimensional target detection results;

[0105] (1) Geometric consistency

[0106] Essentially, the three-dimensional target detection result box is a bounding box generated by finding 8 angle data in the three-dimensional scene and drawing connecting lines. Therefore, similar to the point cloud data, the 8 corner point data of the three-dimensional detection target result box can be projected onto a two-dimensional plane, and the intersection over union IoU is calculated with the two-dimensional target box to quantify the geometric consistency of the two different detection result boxes. In addition, aiming at the problem that the CLOCs method only uses the intersection over union IoU as the representation method to describe the geometric relationship between the two-dimensional target box and the three-dimensional target box, and the representation relationship is relatively single, a method based on feature distance coding is designed to expand the geometric consistency. As Figure 10 shown, after projecting the three-dimensional target box onto the two-dimensional plane, if the three-dimensional target box and the image two-dimensional target box represent the same target, then they have a large intersection over union IoU. Correspondingly, the feature distance between the center points of the two target boxes is small, indicating that the two target boxes belong to the same target. In this section, the Euclidean distance is used as the feature distance to measure the center point of the target box, and the calculation method is shown in the following formula:

[0107]

[0108] where, (x 3D ,y3D ) represents the center point coordinates of the 3D object detection box, (x 2D , y 2D ) represents the center point coordinates of the 2D object detection box.

[0109] (2) Semantic consistency

[0110] In the object detection task, a frame of data may contain multiple objects. Therefore, the generated object detection results may include object detection boxes and corresponding class confidence scores under multiple categories. To achieve the association of class confidence scores for object detection results of different sensor data, only the object detection boxes of the same category are selected for candidate detection to achieve semantic consistency association.

[0111] Object box encoding method:

[0112] Before generating the fused feature tensor based on semantic consistency and geometric consistency, it is necessary to encode the image detection result box and the point cloud detection result box respectively, fuse the class confidence and the object box position, and generate their respective encoding tensors.

[0113] The image object detection result box contains the coordinate points of the 2D object box and the class confidence score. Encode all the detection results of the same class in a frame of data. This article only focuses on vehicle objects. Assuming there are m 2D vehicle object detection result boxes, the encoding process is as follows:

[0114]

[0115] Among them, P 2D represents the set of encoding results of m object detection result boxes in a frame of image data, P i 2D is the i-th image object detection result box, s i 2D represents the class confidence score of the i-th result box, (x i1 2D , y i1 2D ) and (x i2 2D , y i2 2D ) represent the coordinate values of the upper left and lower right corners of the 2D object box respectively;

[0116] The point cloud object detection result box contains the center point coordinates of the 3D object box, the size of the 3D object box, and the class confidence score. Assuming there are n vehicle object detection result boxes in a frame of point cloud data, using the same encoding method as the image object box, the encoding process is as follows:

[0117]

[0118] Among them, P 3D represents the set of encoding results of n target detection boxes in a frame of point cloud data, and P i 3D is the i-th point cloud target detection result box, and (x i , y i , x i , h i , w i , l i , θ i ) respectively represent the center point coordinates, height, width, length, and orientation angle of the 3D target detection box, and s i 3D represents the class confidence score of the i-th target detection box;

[0119] Combining the encoding results of the image target box and the point cloud target box, for a frame of data with m 2D detection result boxes and n 3D detection result boxes, a fusion tensor T of size m×n×4 is constructed as shown in the following formula:

[0120]

[0121] Among them, T i,j represents the fusion tensor generated by the i-th 2D target result box and the j-th 3D target result box, IoU i,j represents the intersection over union of the projection of the j-th 3D target result box onto the image and the i-th 2D target box, s i 2D and s i 3D are the class confidence scores of the i-th 2D target result box and the j-th 3D target result box respectively, and d j represents the normalized distance of the j-th 3D target result box to the lidar plane;

[0122] At the same time, based on the encoding results of the image target box and the point cloud target box, the center point feature distance is calculated, and a feature distance tensor of size m×n×1 is constructed as shown in the following formula:

[0123] D i,j ={dlc i,j}

[0124] Among them, dlc i,j represents the Euclidean distance between the center point of the i-th 2D target box and the projection of the center point of the j-th 3D target box.

[0125] CLOCs designs a convolutional neural network based on a fully connected structure. By inputting the encoded tensor of the fused target box into the network, feature dimension elevation and max pooling are performed to generate a new probability score map. However, since the network only extracts features on the same target pair (an element in the fused tensor set) during the feature dimension elevation process, the associations between different target pairs are ignored, resulting in the neglect of the importance of different features. To address the above problems, the present invention designs a convolutional neural network structure based on channel attention encoding, uses a 1×1 two-dimensional convolutional kernel for feature dimension elevation, and simultaneously introduces channel attention to strengthen the correlations between different target pairs. The network structure is as shown in Figure 11 shown;

[0126] There are a large number of empty elements in the generated fused tensor T i,j and the feature distance tensor D i,j because the number of two-dimensional object detection boxes is much smaller than that of three-dimensional object detection boxes, and a fused tensor and a feature tensor can only be constructed when there is an intersection between only two object detection boxes. Therefore, before feature extraction, it is necessary to find the non-empty elements p of the two tensors and perform a convolution operation on the non-empty elements p to reduce the operation memory and improve the operation efficiency. For the fused tensor T i,j , first, a 4-layer convolutional neural network as shown in Table 2 is used to expand the feature vector to 96 dimensions. After each layer of convolution, the ReLU activation function is used to filter out the feature values less than or equal to 0, so as to ensure that the generated fused confidence scores are all greater than 0 values. To strengthen the associations between different target pairs, a channel attention module is introduced. Max pooling and average pooling are respectively performed on the non-empty element dimension to obtain the association scores of different target pairs and multiply them with the dimension-elevated fused tensor to generate a fused tensor rich in semantic information. The max pooling method is used to obtain the probability distribution map based on the fused vector. For the feature distance tensor D i,j , a small 4-layer convolutional neural network as shown in Table 3 and max pooling are used to obtain the feature distance weight scores, which are multiplied with the probability distribution map obtained from the fused tensor branch to obtain the final fused confidence score. The fused confidence scores are redistributed to the three-dimensional detection targets, and the non-maximum suppression NMS algorithm is used to filter redundant and misdetected target boxes to generate the final three-dimensional vehicle target detection result.

[0127] Table 2 Convolutional neural network structure table for the fused tensor branch

[0128]

[0129]

[0130] Table 3 Convolutional neural network structure table for the feature distance tensor branch

[0131]

[0132] The present invention uses the Focal Loss function as the loss function of the decision-level fusion module, aiming to solve the problem of foreground-background imbalance, as shown in the following formula:

[0133]

[0134] where L total represents the total loss value, N pos represents the number of positive samples, N a represents the total number of generated 3D object bounding boxes; c i represents the ground truth value of the i-th object bounding box, c i being 1 represents a vehicle target, c i being 0 represents a background target; p i is the predicted class confidence value of the i-th 3D detection bounding box. γ is the hyperparameter of the easy-hard sample factor, aiming to enhance the model's learning of difficult samples; α is the hyperparameter of the balance factor, aiming to adjust the positive-negative sample ratio. The present invention sets α = 0.25 and γ = 2.

[0135] The experimental result analysis of the present invention is as follows:

[0136] (1) Performance comparison between the 3D object detector with point cloud and image fusion and the 3D object detector with single point cloud

[0137] The present invention uses two average detection accuracy metrics, AP 3D _R11 and AP 3D _R40, to evaluate the performance of the multi-modal fusion algorithm proposed in the present invention on the KITTI 3D Object dataset. When calculating the average detection accuracy in the R11 manner, the method proposed in the present invention achieved 3D vehicle object detection accuracies of 90.35%, 80.39%, and 78.43% at the simple, medium, and difficult difficulty levels respectively for AP 3D ; when calculating the average detection accuracy in the R40 manner, the method proposed in the present invention achieved 3D vehicle object detection accuracies of 92.94%, 82.70%, and 78.21% at the simple, medium, and difficult difficulty levels respectively for AP3D, proving that the object detection algorithm with point cloud and image fusion proposed in the present invention can effectively achieve 3D vehicle object detection. To reduce the influence of the increased time overhead caused by adding image information in the pixel-level fusion method and the decision-level fusion method, it is verified that the 3D object detector with point cloud and image fusion has certain advantages compared with the 3D object detector with single point cloud.

[0138] (2) Comparison between two image RGB information fusion methods

[0139] Table 4 shows the comparison results of the detection accuracy of two different image RGB information fusion methods. Method 1 is the image RGB texture information fusion based on the target box level. According to the prior two-dimensional target box, the target point cloud is screened, and the point cloud frustum is generated. Only the RGB information of the target point cloud frustum in the scene is fused to ensure that the algorithm fully focuses on the target point cloud during the feature extraction process. Method 2 is the image RGB texture information fusion based on the scene. The RGB information of the point cloud scene in the entire forward view of the image is fused to ensure that the point cloud information has a sufficient context structure and prevent the loss of key point cloud information. The experimental results show that the method of image RGB information fusion based on the entire scene (Method 2) can provide more point cloud information and scene structure information for the 3D object detector. During the feature extraction process, the algorithm can extract more complete and rich high-dimensional semantic features of the point cloud. Therefore, the average detection accuracy AP 3D of Method 2 is slightly higher than that of Method 1, and the performance is optimal.

[0140] Table 4 Comparison table of the average detection accuracy of two image RGB information fusion methods

[0141]

[0142] (3) Comparison table of the method of the present invention with other methods

[0143] Compare the average detection accuracy AP of the fusion method proposed in the present invention with other 3D object detection methods 3D , as shown in Table 5, the method of using 11 interpolation points (R11) is used for measurement, and the accuracies of other 3D object detection methods are from their respective literatures.

[0144] Table 5 Comparison table of the average detection accuracy of the method of the present invention and other methods on the KITTI 3D Object validation set

[0145]

[0146] According to Table 5, compared with the 3D object detection methods based only on lidar point cloud such as Part-A2, SECOND-L, and IA-SSD, the method of the present invention has achieved better vehicle object detection effects on simple, medium, and difficult samples. For example, on simple samples, the average detection accuracies AP are 0.88%, 1.46%, and 2.01% respectively 3D improvement; compared with MV 3DCompared with other multi-modal feature fusion methods based on view projection, the method proposed in this paper has achieved leading vehicle target detection performance. The detection results of the multi-modal fusion method proposed in this paper are also better than those of other multi-modal fusion methods such as FPointNet, ContFusion, PointPainting, MANet, and UberATG-MMF. The average detection accuracy AP 3D has increased by 9.47%, 7.14%, 2.65%, 2.76%, and 2.96% respectively.

[0147] (4) Qualitative analysis of target detection performance

[0148] Use tools such as Open3D and OpenCV to draw the original point cloud, 3D target detection results, original image, and image detection results. The visualization results are shown in Figure 12. In the 3D scene, the green 3D target box represents the 3D real vehicle target box, the red 3D target box represents the vehicle target prediction box of the method of the present invention, the blue 3D target box represents the vehicle target prediction box of other methods, and the white is the original point cloud; in the 2D image, the green 3D target box represents the real vehicle target box, the red 3D target box represents the vehicle target prediction box of the method of the present invention, and the blue 3D target box represents the vehicle target prediction box of other methods. In both visualization results, the orange circle represents the missed detection target, and the pink circle represents the false detection target.

[0149] It can be seen from the visualization results that when using other algorithms to detect vehicle targets from lidar point clouds, the algorithms can better locate vehicle targets in different poses and different truncation degrees. However, since this method only performs detection based on point clouds, the single point cloud data collected by lidar sensors contains insufficient information, resulting in missed detection and false detection phenomena in some traffic scenarios, such as Figure 12a shown, which are marked with orange circles and pink circles respectively. Compared with the detection method based only on point clouds, the present invention incorporates an image information branch, adds a pixel-level fusion method in the data processing stage, and adds a decision-level fusion method before non-maximum suppression, resulting in a significant improvement in detection performance. From Figure 12b it can be found that compared with the 3D target detection method of point clouds without fusing image information, the algorithm of the present invention fuses the image information branch to make up for the problem of insufficient semantic information contained in the point cloud data collected by a single lidar sensor to a certain extent. During the 3D vehicle target detection feature extraction process, the algorithm can extract denser and richer semantic features from the multi-modal fusion point cloud, effectively improving the 3D vehicle target detection accuracy and enhancing the detection performance for the above-mentioned missed detection and false detection targets. At the same time, it can effectively detect vehicle targets with different occlusion degrees, different poses, and sparse point clouds.

[0150] The present invention also evaluates a pixel-level fusion method based on image RGB texture features and prior results of image targets;

[0151] By applying the fusion scheme to different existing 3D object detectors for comparative experiments and using the R11 method to measure the average detection accuracy, the results are shown in Table 6. According to the experimental results, the method proposed in this paper has a good improvement effect on the detection performance of different 3D object detectors. Comprehensive comparison shows that the fusion scheme has the best improvement effect on the PointRCNN 3D object detection algorithm based on the original point cloud, improving by 0.94%, 1.03% and 1.12% respectively compared with the original method PointRCNN under the three difficulty samples of simple, medium and hard. Since the pixel-level fusion method in this paper directly expands the dimension and enriches the information at the original point cloud data level, the feature extraction network based on the original point cloud can directly obtain better semantic features to achieve high-precision detection. The above results show that the pixel-level fusion method proposed in this paper can effectively enhance the point cloud information and improve the object detection accuracy.

[0152] Table 6 Comparison of the average detection accuracy of the pixel-level fusion scheme of the present invention in different 3D detectors

[0153]

[0154] The present invention also evaluates a decision-level fusion method based on channel attention and feature distance. The decision-level fusion module proposed in this paper can also be regarded as an independent module, and the object box fusion is realized by analyzing the geometric consistency and semantic consistency of the detection results of different sensors. The present invention analyzes the experimental results from two aspects: ablation experiments and performance comparison on different 3D object detectors. The 2D object detector used in the experiment is Cascade-RCNN.

[0155] (1) Ablation experiments

[0156] In this paper, channel attention coding and feature distance coding are introduced on the basis of the CLOCs method, which improves the problems existing in the CLOCs method. To verify the effectiveness of the introduced module, the ablation comparison experiment shown in Table 7 was carried out. The R11 and R40 methods were used for measurement. When the channel attention module and the feature distance module were not added, the fusion module degenerated into the CLOCs method; when the channel attention module was added, the different object pairs between the fusion tensors generated by the two object detection result boxes interacted in the convolutional network, generating more reliable class confidence scores. Under the measurement of the R11 method, AP 3DOn the simple, medium, and difficult difficulty samples, it was improved by 0.12%, 0.03%, and 0.06% respectively; the feature distance module improved the problem of single quantization geometric consistency method in the CLOCs method, enhanced the geometric consistency relationship of different sensor detection results, and under the R11 method metric, AP3D was improved by 0.21%, 0.07%, and 0.02% on the simple, medium, and difficult difficulty samples respectively. When channel attention and feature distance are introduced into the fusion module at the same time, when using the R11 metric, compared with using the original CLOCs method, AP 3D was improved by 0.71%, 0.44%, and 0.36% respectively. The above results show that the method proposed in this paper can effectively improve the problems existing in the CLOCs method.

[0157] Table 7 Experimental table for verifying the effectiveness of channel attention and feature distance coding

[0158]

[0159] (2) Result comparison of decision-level fusion on different 3D object detectors

[0160] In this experiment, decision-level fusion was added to a variety of 3D object detectors for experimental comparison, and the R11 was used for metric in all cases. It can be seen from the results in Table 8 that the detection accuracy of different 3D object detectors has been improved after adding the decision-level fusion method proposed in this paper. For methods with relatively low baseline accuracy, such as the SECOND and PointPillars algorithms, the decision-level fusion scheme has a greater improvement in the detection accuracy of vehicle targets. When using the R11 metric, for the SECOND algorithm, the average detection accuracy AP of the decision-level fusion scheme proposed in this paper under the difficult difficulty level 3D was improved by 6.08%; for the PointPillars algorithm, the average detection accuracy AP of 3.87%, 2.02%, and 3.19% was obtained on the simple, medium, and difficult difficulty samples respectively 3D improvement. For methods with relatively high baseline accuracy, such as the PointRCNN algorithm and the PV-RCNN algorithm, the decision-level fusion scheme has a certain improvement effect on the detection accuracy. When using the R11 metric, for the PointRCNN algorithm and the PV-RCNN algorithm, the average detection accuracy AP of 1.41% and 1.58% was obtained on the medium difficulty samples by the decision-level fusion scheme proposed in this paper respectively 3D improvement. The above results show that the decision-level fusion scheme proposed in this paper has a good improvement effect on the detection performance of 3D object detection methods based on single point cloud.

[0161] Table 8 Comparison of the average detection accuracy of the decision-level fusion scheme proposed in this paper on different 3D object detectors

[0162]

[0163] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

Claims

1. A 3D vehicle target detection method that fuses image texture features and prior information, characterized in that Including: Obtain a dataset, and perform pixel-level point cloud and image fusion based on the point cloud data and image data in the dataset. Specifically: The dataset is the KITTI dataset. Obtain the prior information of image targets from the two-dimensional target ground truth label file of the KITTI dataset. According to the prior information of image targets, obtain the target categories and target box ranges in the image. After mapping the point cloud to the image, filter out the point cloud within the image target box according to the target box range. Based on the category information assignment method encoded by Gaussian distance, judge the probability that the point cloud within the frustum belongs to the foreground target points through the Gaussian distance feature information, and assign the probability as the category score to the point cloud data, and assign the target category information to these point clouds. On this basis, assign the corresponding pixel point R, G, B color values to each point cloud to generate point cloud data with rich feature information; Detect the image data based on a preset image target detector to generate an image two-dimensional target box; Detect the fused point cloud data based on a preset three-dimensional vehicle target detection model to generate a point cloud three-dimensional target box; Fuse the image two-dimensional target box and the point cloud three-dimensional target box based on a decision-level fusion scheme of feature distance and channel attention to obtain the final three-dimensional vehicle target detection result. Specifically: Encode the image two-dimensional target box and the point cloud three-dimensional target box respectively, fuse the class confidence and the target box position to generate their respective encoded tensors; Based on semantic consistency and geometric consistency, fuse the encoded tensors of the image two-dimensional target box and the point cloud three-dimensional target box to obtain a fused feature tensor; Input the fused feature tensor into a convolutional neural network structure encoded by channel attention to finally obtain a fused confidence score; Reassign the fused confidence score to the point cloud three-dimensional target box, and use the non-maximum suppression (NMS) algorithm to filter redundant and misdetected target boxes to generate the final three-dimensional vehicle target detection result.

2. The method according to claim 1, characterized in that Detect the image data based on a preset image target detector to generate an image two-dimensional target box. Specifically: Construct an image target detector based on the CascadeRCNN algorithm; Train the image target detector based on a preset training dataset; Input the image into the trained image target detector for detection to generate an image two-dimensional target box.

3. The method according to claim 2, wherein Detect the fused point cloud data based on a preset three-dimensional vehicle target detection model to generate a point cloud three-dimensional target box. Specifically: Construct a three-dimensional vehicle target detection model based on the improved PointPillars network model; Train the three-dimensional vehicle target detection model based on the KITTI dataset to obtain the trained three-dimensional vehicle target detection model; Input the fused point cloud data into the trained three-dimensional vehicle target detection model to generate a point cloud three-dimensional target box.

4. The method according to claim 3, wherein Construct a three-dimensional vehicle target detection model based on the improved PointPillars network model. Specifically: Improve the PointPillars network model by adding a point cloud local attention mechanism module to the pillar feature network of the traditional PointPillars network model to capture complex features in a specific area of the input point cloud, and integrating an SE attention mechanism module into the 2D pseudo-image network of the traditional PointPillars network model to enhance the network's ability to obtain global features and information; Construct a 3D vehicle target detection model based on the improved PointPillars network model.

Citation Information

Patent Citations

  • Three-dimensional point cloud target detection method fused with two-dimensional image semantics

    CN116597264A

  • Multi-modal fusion target detection method and system based on clustering optimization

    CN117789160A