3D target detection method and device based on sparse radar and binocular stereo image fusion
Through the end-to-end learning framework of sparse radar and binocular stereoscopic image fusion, the problems of high sensor dependence and cost in the existing 3D object detection technology are solved, and efficient and low-cost 3D object detection effect is achieved.
Patent Information
- Application Number
- CN202210405709.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-04-18
AI Technical Summary
Existing 3D object detection technology relies too much on a single sensor, resulting in high cost, low efficiency, low long-distance point cloud resolution and poor texture information.
The end-to-end learning framework of sparse radar and binocular stereoscopic image fusion is adopted to fuse the LiDAR depth map and stereoscopic image feature information through the attention fusion module, and the stereo area extraction network and depth prediction branches are used to predict the position, size and direction of the 3D bounding box.
It realizes efficient 3D object detection of low-cost sensors, improves detection performance and time efficiency, and reduces detection costs.
Smart Images

Figure CN114743079B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence, computer vision, automatic driving and 3D target detection, and in particular to a 3D target detection method and device by fusing sparse radar and binocular stereo images. Background Art
[0002] Currently, 3D object detection for autonomous driving relies heavily on LiDAR (Light Detection And Ranging), which can provide rich information about the surrounding environment. Despite the accurate information, it is not wise to over-rely on a single sensor due to inherent safety risks (e.g., damage, adverse conditions, and blind spots). At the same time, the low resolution and poor texture information of long-distance point clouds are also a big challenge. The most promising candidates are airborne stereo or monocular cameras, which provide both fine-grained texture and three primary color (Red, Green, Blue, RGB) attributes. However, cameras inherently have the problem of depth ambiguity. In addition, stereo or monocular cameras are orders of magnitude cheaper than LiDAR, with high frame rates and dense depth maps. Obviously, each type of sensor has its shortcomings, and the combination can be seen as a possible remedy to failure modes. Some work even points out that multimodal fusion provides redundancy in difficult situations, not just complementarity. Although exploiting synergy is a compelling research hotspot, integrating the advantages of camera views and LiDAR bird's-eye views is not easy. Deep neural networks exploit the property that natural signals consist of hierarchical structures, where fusion strategies may vary and can be divided into the following two categories: sequential fusion and parallel fusion.
[0003] Sequential fusion based methods: These methods utilize multi-stage features in a sequential manner, where the current feature extraction depends heavily on the previous stage.
[0004] Qi et al. proposed Frustum PointNets for 3D Object Detection from RGB-D Data (Frustum PointNets) for 3D object detection from image-depth data. They first used a standard 2D convolutional neural network (CNN) target detector to extract 2D regions, and then projected the 2D candidate boxes to the point cloud in the 3D cone. Next, a block similar to Deep Learning on Point Sets for 3D Classification and Segmentation (PointNet) was used to segment each point in the cone to obtain points of interest for further regression. Frustum PointNets uses mature 2D detection methods to provide prior knowledge, which reduces the 3D search space to a certain extent and inspires its successors. Although Frustum PointNets is very innovative, the disadvantage of this cascade method is that Frustum PointNets relies heavily on the accuracy of the 2D detector. Considering that the depth estimation error grows quadratically at longer distances, You et al. proposed the AccurateDepth for 3D Object Detection in Autonomous Driving (Pseudo-LiDAR++) algorithm to align distant targets. The main contribution of Pseudo-LiDAR++ is that it proposes a graph-based depth correction (GDC) algorithm that utilizes sparse but accurate LiDAR points (e.g., 4 laser beams) to eliminate the bias of stereo-based depth estimation. Specifically, they project a small number of sparse LiDAR points (i.e., "landmarks") to pixel locations and assign them to the corresponding 3D pseudo-LiDAR points as the "real" LiDAR depth. Note that the depth of the 3D pseudo-LiDAR points is obtained through a stereo depth estimation network (Pyramid stereo matching network, PSMNET). To correct the depth values, Pseudo-LiDAR++ first constructs a local graph through k-nearest neighbors (kNN), and then updates the weights of the graph under the supervision of "landmarks". Finally, the information is propagated across the entire graph at negligible cost. Although Pseudo-LiDAR++ cleverly explores a hybrid approach to correct depth bias, it is not an end-to-end approach.
[0005] Parallel fusion based methods: These methods fuse the modalities in the feature space to obtain a multimodal representation which is then fed into a supervised learner.
[0006] Chen et al. proposed a multi-view Figure 3 Multi-View 3D Object Detection Network for Autonomous Driving (MV3D) takes multi-view representations, namely bird's-eye view and front view, as well as images as input. MV3D first generates a set of accurate 3D candidate boxes through the bird's-eye view representation of the point cloud. Given high-quality 3D proposals, MV3D crops the corresponding regions from multiple views according to the coordinates of the 3D proposals. Then, a deep multi-view fusion network is used to fuse regional features. Although MV3D utilizes the multi-view representation of point clouds, its disadvantage is that MV3D relies on manual features, which hinders its further improvement and is quickly surpassed by its successors. Later, Ku et al. proposed Joint 3D Proposal Generation and Object Detection from View Aggregation (AVOD), which is slightly different from MV3D in that it further extends the fusion strategy to the early stage of region proposal. Specifically, given a set of predefined 3D boxes (called anchor boxes), two corresponding regions of interest are cropped and adjusted from the front view feature map and the bird's eye view (BEV) feature map, respectively, and fused through the element-by-element mean operation. Then AVOD inputs the fused features into the fully connected layer to detect the object. AVOD believes that this subtle operation can generate high-recall proposals and benefit positioning accuracy, especially for small objects. Although the fusion strategy proposed by AVOD further improves the quality of proposals, this regional fusion only occurs at the top of the feature pyramid. However, intermediate features are also important for detection. Note that both MV3D and AVOD are instance-level fusion strategies, and then pixel-level fusion is proposed for deep collaboration.
[0007] Most existing technologies use 32 or 64 laser beam LiDAR and RGB images to fuse for 3D target detection, which makes the cost of 3D target detection very high. Although Pseudo-LiDAR++ explores the method of using 4 laser beams LiDAR to correct the depth deviation of stereo images, it is not an end-to-end method, has low time efficiency, and the generation of stereo image depth map uses 64 laser beams of LiDAR information supervision. Summary of the invention
[0008] The present invention provides a sparse radar and binocular stereo fusion network 3D target detection method and device, which fuses the passive stereo camera with the active 4-laser beam LiDAR sensor information to reach the current advanced level and performs high-speed detection in an end-to-end manner, as described below:
[0009] In a first aspect, a 3D target detection method using sparse radar and binocular stereo image fusion is provided, the method comprising:
[0010] After feature encoding of the stereo image and the sparse LiDAR depth map respectively, the feature information of the two paths is fused based on the attention fusion module, wherein the fusion is from the LiDAR depth map to the stereo image;
[0011] Based on the output of the corresponding left and right regions of interest by the stereo region extraction network, the left and right feature maps are input together into the stereo regression network branch and the depth prediction branch to predict the position, size and orientation of the 3D bounding box.
[0012] The stereo regression network branch is used to regress the 2D stereo box, size, viewpoint angle and 2D center; the depth prediction branch is used to predict the univariate depth of the center of the 3D bounding box.
[0013] Furthermore, the attention fusion module fuses the left sparse LiDAR feature map with the corresponding left RGB feature map, and fuses the right sparse LiDAR feature map with the corresponding right RGB feature map.
[0014] The fusion process is as follows:
[0015]
[0016]
[0017]
[0018] Among them, F i Represents the characteristics of fusion, is the feature output of the last block in each stage of the encoder, Refers to the last output feature of the encoder.
[0019] Furthermore, the method further comprises:
[0020] Add the sparse LiDAR features to the image features and set the weight w for each feature level i , by calculating the correlation between the sparse LiDAR and its corresponding stereo image feature map, we get the correlation score w i , defined as:
[0021]
[0022] in, is the i-th pair of stereo image feature map and sparse LiDAR feature map in the feature extractor, w i is the weight of the sparse LiDAR feature map at level i, and cos is the cosine similarity function;
[0023] F i+1 Upsample by 2 to F' f ∈R H×W×C , and apply 1×1 convolution operation to Projected into F' r ∈R H×W×C ,Will Projected into F' s ∈R H×W×C , described as:
[0024] F f = upsample(F i+1 )
[0025]
[0026]
[0027] Among them, upsample is the upsampling operation performed by nearest neighbor interpolation, and f 1×1 represents a 1×1 convolutional layer;
[0028] Upsampled feature maps and corresponding F' r The feature maps are merged by element-wise addition, and a 3×3 convolution is added to each merged feature map to combine the merged features with the applied weight w i The sparse LiDAR features F' s Add them together and the output features are calculated as follows:
[0029] F5=f 3×3 (F' r +w5·F' s )
[0030] Among them, the fusion result F i is the higher-level feature for the next fusion stage, and this process is repeated until the final feature map is generated.
[0031] In a second aspect, a 3D target detection device that fuses sparse radar and binocular stereo images is provided, characterized in that the device comprises: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute any one of the method steps described in the first aspect.
[0032] In a third aspect, a computer-readable storage medium is provided, characterized in that the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes any one of the method steps described in the first aspect.
[0033] The beneficial effects of the technical solution provided by the present invention are:
[0034] 1. This paper proposes a new multi-modal fusion end-to-end learning framework for 3D object detection, which effectively integrates the complementarity of sparse LiDAR and stereo images;
[0035] 2. We propose a deep attention feature fusion module that explores the interdependence of channel features in sparse LiDAR and stereo images while fusing important multimodal spatial features;
[0036] 3. Compared with low-cost sensor methods without depth map supervision, this method achieves state-of-the-art performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a network framework diagram of a sparse radar and binocular stereo fusion network 3D target detection method;
[0038] Figure 2 Schematic diagram of feature fusion module based on attention mechanism;
[0039] Figure 3 It is a structural schematic diagram of a sparse radar and binocular stereo fusion network 3D target detection device. DETAILED DESCRIPTION
[0040] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are described in further detail below.
[0041] 3D object detection is one of the important tasks of scene understanding and has a wide range of applications in the fields of autonomous driving, virtual reality, etc. The embodiments of the present invention observe that sensors such as LiDAR, monocular cameras, and binocular cameras all have their own advantages and disadvantages, and none of them can perform well in all practical scenarios. Therefore, some works study how to fuse multiple modalities to provide more accurate 3D object detection and further improve the performance of object detectors. However, these methods usually use 32 or 64 beams of LiDAR information as input, making the cost of 3D detection very high. Therefore, the embodiments of the present invention consider fusing passive stereo cameras with active 4-beam LiDAR sensor information, which is a practical and popular choice. Compared with 64-beam LiDAR sensors, LiDAR sensors with only 4 beams are two orders of magnitude cheaper and therefore easily affordable. Although the information of 4-beam LiDAR sensors is very sparse and not suitable for capturing the shape of 3D objects alone, if fused with stereo image information, they can learn better feature representations, resulting in better detection performance.
[0042] The embodiment of the present invention proposes a novel multimodal fusion architecture that takes advantage of the fusion of sparse LiDAR and stereo image features to produce rich feature representations. It is worth noting that the architecture proposed in the embodiment of the present invention is designed based on low-cost sensors. Since the 4-beam LiDAR information is extremely sparse, the fusion with the stereo image information is from the LiDAR stream to the image stream to enhance the image features using the accuracy of the LiDAR feature information. To this end, the embodiment of the present invention first obtains sparse but accurate depth information from the 4-beam LiDAR and makes it dense using a simple and fast depth completion method. After feature encoding the stereo image and the sparse LiDAR depth map respectively, an attention fusion module is proposed to fuse the feature information of the two paths. The next module of the network is the stereo region proposal network (RPN), which can output corresponding left and right region of interest (RoI) proposals. Then, the left and right feature maps are input together into two different branches. One is a stereo regression network branch for regressing accurate 2D stereo boxes, sizes, viewpoint angles, and 2D centers. The other is a depth prediction branch, which is used to predict the univariate depth z of the center of the 3D bounding box.
[0043] The goal of the embodiment of the present invention is to detect and locate the 3D bounding box of the target by using stereo RGB images and 4-beam LiDAR. The detection process includes three stages: First, the sparse LiDAR image and the stereo RGB image (including: two left and right pictures, respectively, the left view and the right view) are respectively extracted using the ResNet-50 encoder. Then, the stereo image features are fused with their corresponding sparse LiDAR features using the attention mechanism. Finally, after the fused feature pair passes through the stereo RPN, the position, size and orientation of the 3D bounding box are predicted.
[0044] 1. Depth Completion and Feature Extraction
[0045] In order to enrich the representation of the ordinary stereo (RGBs) 3D object detection network, the embodiment of the present invention decides to add geometric information from the LiDAR point cloud. However, instead of using the 3D point cloud from the LiDAR directly, two sparse LiDAR depth maps corresponding to the stereo images are formed by reprojecting the 4-beam LiDAR scan information to the left and right image coordinates using calibration parameters. LiDAR can provide accurate 3D information for 3D object detection. However, it can be observed that the ordinary 64-beam LiDAR information is sparse, and the 4-beam LiDAR information is even more sparse. Therefore, here, the embodiment of the present invention performs depth completion similar to the method of Ku et al. on the generated sparse LiDAR depth map to generate a dense depth map. First, a simple sequence of morphological operations and Gaussian blur operations are used to fill the holes in the sparse depth image with depth values from nearby valid points. Then, the filled depth image is normalized using the maximum depth value in the dataset so that the depth value is between 0 and 1, and finally, it is input to the encoder to extract features. Stereo images and sparse LiDAR each have a feature encoder, and their feature encoder architecture is the same, the encoder weights are shared by the left and right input views. The feature encoder consists of a series of ResNet blocks followed by convolutions with a stride of 2, which reduces the feature resolution to 1 / 16 of the input image.
[0046] 2. Feature Fusion Based on Attention Mechanism
[0047] The embodiment of the present invention adopts a deep fusion method to fuse sparse LiDAR and RGB features layer by layer. Specifically, in this module, the embodiment of the present invention fuses the left sparse LiDAR feature map with the corresponding left RGB feature map, and fuses the right sparse LiDAR feature map with the corresponding right RGB feature map. The fusion method of layer-by-layer fusion of left and right sparse LiDAR features and RGB features is the same.
[0048] For a network with L layers, early fusion combines features from multiple modalities at the input stage:
[0049]
[0050] Among them, [D l ,l=1,2,…,L] is the feature transformation function, ⊕ is a connection operation (e.g., addition, concatenation), The input information is stereo RGB images and sparse LiDAR data respectively. In contrast, late fusion uses separate sub-networks to independently learn feature transformations and combines their outputs in the prediction stage:
[0051]
[0052] Among them, D rgb , D sl They are the feature transformation functions of the stereo RGB image and the sparse LiDAR data respectively. In order to make the intermediate layer features of different modalities interact more, the embodiment of the present invention designs the following deep fusion process:
[0053]
[0054]
[0055]
[0056] Among them, F i Represents the characteristics of fusion, is the feature output of the last block in each stage of the encoder, Refers to the last output feature of the encoder. Higher-resolution features are produced by upsampling feature maps from higher levels where spatial information is coarser but semantic information is more effective. These features are then enhanced using features from the image path and the LiDAR path through a concatenation operation. Each connection merges feature maps of the same spatial size. The feature maps of the image path and the LiDAR path have lower-level semantics, but because it is subsampled less times, its activation positioning is more accurate. Therefore, the resulting fused features have higher-level semantic information and higher resolution, which is beneficial for 3D object detection. Since the depth information of the input is closely related to the output of the decoder, the features from the sparse LiDAR depth map should provide a greater contribution in the decoder.
[0057] Therefore, the embodiment of the present invention adds the features of the sparse LiDAR depth map to the stereo features in the decoder instead of splicing. This is because summation is beneficial to the features on both sides of the same domain, and can encourage the decoder to learn features that are more relevant to depth so as to be consistent with the features of the sparse LiDAR depth. However, the 4-beam LiDAR information is too sparse to provide enough information for 3D detection alone. Therefore, fusion is from the LiDAR stream to the image stream to enhance the image features. As shown in the above formula, the features between different modalities are on an equal footing when fused, rather than weighted, which may result in the different importance of different modalities not being correctly reflected.
[0058] To solve this problem, the embodiment of the present invention adopts an attention mechanism to add sparse LiDAR features to image features and set a weight w for each feature level. i By calculating the correlation between the sparse LiDAR and its corresponding stereo image feature map, the correlation score w can be obtained. i , which is defined as:
[0059]
[0060] in, is the i-th pair of stereo image feature map and sparse LiDAR feature map in the feature extractor, w i is the weight of the sparse LiDAR feature map at level i, cos is the cosine similarity function, T is the transpose, and R represents the real number domain. Technically speaking, the embodiment of the present invention first converts F i+1 Upsample by 2 to F' f ∈R H×W×C (For simplicity, nearest neighbor upsampling is used), where H, W, C refer to the features F' f Then, a 1×1 convolution operation is applied to convert F i r Projected into F' r ∈R H×W×C ,Will Projected into F' s ∈R H×W×C The process can be described as:
[0061] F' f = upsample(F i+1 ) (7)
[0062]
[0063]
[0064] Among them, upsample is the upsampling operation performed by nearest neighbor interpolation, and f1×1 Refers to a 1×1 convolutional layer. At each stage, the transformed feature F' r , F' s The channels are unified into 256 dimensions.
[0065] In addition, the upsampled feature maps are r The feature maps (after 1×1 convolutional layers to reduce the channel dimension) are combined by element-wise addition. A 3×3 convolution is appended to each combined feature map to reduce the aliasing effect of upsampling. Finally, the combined features are combined with the weight w i The sparse LiDAR features F' s The output features are calculated as follows:
[0066] F5=f 3×3 (F' r +w5·F' s ) (10)
[0067] Among them, f 3×3 represents a 3×3 convolutional layer. The fusion result F i is the higher-level feature for the next fusion stage. This process is repeated until the final feature map is generated. To start the iteration, just generate the initial fused feature map F5, which can be expressed as:
[0068] F i =f 3×3 (F' r +w i ·F' s ) (11)
[0069] Among them, F' r , F' s are the 5th feature level of the stereo image and sparse LiDAR used in the decoder stage, respectively.
[0070] 3. 3D Object Detection
[0071] The embodiment of the present invention uses a stereo RPN module to extract a pair of regions of interest (RoI) for each target in the left and right images, with the aim of avoiding complex matching of all pixels between the left and right images and eliminating the adverse effects of the background on target detection. The stereo RPN creates a joint RoI for each object of the same size and position on the left and right images, so that the joint RoI ensures the starting point of each pair of RoIs. After the stereo RPN, the embodiment of the present invention has corresponding left and right proposal pairs. RoI Align is applied to the left and right feature maps at the appropriate pyramid level respectively. Then, the left and right RoI features are connected and input into the depth prediction branch and the stereo regression branch respectively. The embodiment of the present invention predicts the 3D depth of the target center in the depth prediction branch. z maxand z min The depth between is divided into 24 levels for estimating the center depth of the target. This branch calculates the disparity of each instance to locate its position, and then forms a cost volume of size d×h×w×f by connecting the left and right feature maps at each disparity level. In order to learn from the cost volume and downsample the feature representation from the cost volume, two consecutive 3D convolutional layers are used, each followed by a 3D maximum pooling layer. Since the disparity is inversely proportional to the depth and both represent the position of the target, the disparity is converted into a depth representation after formulating the cost volume. Through network regularization, the downsampled features of the 3D CNN are finally merged into the probability of the center depth of the 3D box. By * By weighted summing according to their normalized probabilities, we can finally get the depth of the center z of a 3D box, as shown below:
[0072]
[0073] Wherein, N represents the number of depth levels, and P(i) refers to the normalized probability. In addition to the depth prediction branch, the embodiment of the present invention also first uses two consecutive fully connected layers in the stereo regression branch to extract semantic features, and then uses four sub-branches to predict the 2D box, dimension, viewpoint angle and 2D center respectively.
[0074] Finally, the state of the 3D bounding box can be represented by the predicted position, orientation, and size of the 3D bounding box, where the position of the 3D bounding box can be represented by its center position (x, y, z).
[0075] The multi-task loss function used by the network proposed in the embodiment of the present invention can be expressed as:
[0076]
[0077] in,(·) s ,(·) r and(·) d They represent stereo RPN, stereo regression and depth prediction respectively. The subscripts box, dim, α and ctr represent the loss functions of 2D stereo box, size, viewpoint and 2D center respectively.
[0078] All the above modules are integrated through the multi-task loss function, and the training data of each module is constrained by the loss function.
[0079] 4. Comparison of 3D Object Detection Results
[0080] As shown in Table 1, the embodiment of the present invention reports the 3D box (AP 3D ) and Bird's Eye View (AP bev). Depending on the input signal, M represents monocular image, S represents stereo image, and L# represents sparse 4-beam LiDAR. PL(AVOD) is the result reported by DSGN without LiDAR supervision. The embodiment of the present invention uses the original KITTI evaluation metric here. The main results are shown in Table 1, where the embodiment of the present invention compares the present method with the previous state-of-the-art methods from monocular to binocular. Compared with the previous monocular-based methods, the present method achieves significant improvements at all levels of all IoU thresholds. Compared with the binocular-based methods, the present method achieves the highest performance at 0.5IoU and 0.7IoU.
[0081] Table 1 Comparison of 3D object detection results evaluated on the KITTI object validation set
[0082]
[0083]
[0084] Specifically, the AP of our method at medium and difficult levels of 0.7 IoU is bev They outperform the previous state-of-the-art IDA-3D method by 1.94% and 1.67% respectively. 3D A similar improvement trend can be seen in , which shows that our method can achieve consistent improvements compared with other methods. 3D (IoU = 0.7), the results of this method in the medium and difficult levels are 2.32% and 1.41% higher than IDA-3D respectively. 3D The performance on the CNN (IoU = 0.7) is only slightly better than IDA-3D, but in the difficult level, the method has an AP 3D A significant improvement of 6.26% is obtained on the IoU=0.5. This may be because the present method focuses on improving the accuracy of the predicted depth of the target and obtains a more accurate depth by introducing sparse LiDAR.
[0085] Table 2 AP of Pseudo-LiDAR++ and this method in the car category on the KITTI validation set bev and AP 3D (%)Compare
[0086]
[0087] The present invention uses 4-beam LiDAR as input instead of 64-beam LiDAR as input or intermediate supervision. It is unfair to compare the present method with the methods in the literature. Therefore, the present method is compared with the Pseudo-LiDAR++ method which also uses stereo images and sparse LiDAR as input. Since Pseudo-LiDAR++ does not report experimental results without 64-beam LiDAR supervision, the present method gives the re-implemented results in Table 2. The experimental results in Table 2 show that the present method outperforms the PL++ (AVOD) method in some indicators. Specifically, in the simple level, when IoU = 0.7, AP 3D Achieved an 11.3% improvement. bev For the 3D point cloud, the present method obtains an improvement of more than 7.82%. This may be because the present method projects the 3D point cloud onto the front view image, while the convolutional network pays more attention to nearby objects. In addition, the comparison of the running time of the present method and the PL++(AVOD) method is also reported in Table 2. The present method has a high speed of 0.116 seconds per frame during inference, which far exceeds the PL++(AVOD) method. The improvement in efficiency is mainly attributed to the network design of the present method. Compared with PSMNet, the network designed in the embodiment of the present invention is an end-to-end network with lightweight modules.
[0088] V. Ablation Experiment Results and Analysis
[0089] Table 3 Ablation experiments on the KITTI validation set
[0090]
[0091]
[0092] Here, we analyze the effectiveness of sparse LiDAR, depth completion, and attention fusion components in our approach.
[0093] When only sparse LiDAR is used, the method directly adds the sparse LiDAR feature map to the corresponding stereo image feature map at the appropriate level in the decoder. When depth completion is not used, the method treats the sparse LiDAR depth map as the input to the depth feature extractor. When attention fusion is not used, the weight of the sparse LiDAR feature map and its corresponding stereo image feature map is 1.
[0094] When only sparse LiDAR is used, the evaluation indicator AP 3D and AP bev The values for the threshold of 0.7 are significantly improved, which shows that sparse LiDAR is crucial for high-quality 3D detection. At the medium-level threshold IoU = 0.7, there is no depth completion component, which makes AP3D The percentage of AP dropped from 38.83% to 37.31%. In addition, when attention fusion is removed, AP bev The performance of is reduced by 1.87% in the simple level 0.7IoU. By combining these three key components, great improvements can be observed in all indicators, and the results almost surpass all previous low-cost based methods.
[0095] The embodiment of the present invention weights each loss to balance the entire multi-task loss behind. Two weighted shared ResNet-50 structures are used as feature encoders for stereo images and sparse LiDAR respectively. For data enhancement, the left and right images in the training set are flipped and swapped, and the image information is mirrored. For sparse LiDAR, the embodiment of the present invention first projects it onto the image plane using calibration parameters, and then applies the same flipping strategy as the previous stereo image. The model of the present invention is implemented under PyTorch 1.1.0, CUDA 10.0. By default, the embodiment of the present invention uses a GPU training network with a batch size of 4 on 4 NVIDIA Tesla V100 GPUs for 65,000 iterations, and the total training time is about 26 hours. The embodiment of the present invention uses a stochastic gradient descent (SGD) optimizer with an initial learning rate of 0.02. The momentum of the SGD optimizer is set to 0.9 and the weight decay is set to 0.0005.
[0096] A 3D target detection device based on sparse radar and binocular stereo image fusion, see Figure 3 The device comprises: a processor 1 and a memory 2,
[0097] After encoding the features of the stereo image and the sparse LiDAR depth map respectively, the feature information of the two paths is fused based on the attention fusion module. The fusion is from the LiDAR depth map to the stereo image.
[0098] Based on the output of the corresponding left and right regions of interest by the stereo region extraction network, the left and right feature maps are input together into the stereo regression network branch and the depth prediction branch to predict the position, size and orientation of the 3D bounding box.
[0099] Among them, the stereo regression network branch is used to regress the 2D stereo box, size, viewpoint angle and 2D center; the depth prediction branch is used to predict the univariate depth of the 3D bounding box center.
[0100] Furthermore, the attention fusion module fuses the left sparse LiDAR feature map with the corresponding left RGB feature map, and the right sparse LiDAR feature map with the corresponding right RGB feature map.
[0101] The fusion process is:
[0102]
[0103]
[0104]
[0105] Among them, F i Represents the characteristics of fusion, is the feature output of the last block in each stage of the encoder, Refers to the last output feature of the encoder.
[0106] Furthermore, it also includes:
[0107] Add the sparse LiDAR features to the image features and set the weight w for each feature level i , by calculating the correlation between the sparse LiDAR and its corresponding stereo image feature map, we get the correlation score w i , defined as:
[0108]
[0109] in, is the i-th pair of stereo image feature map and sparse LiDAR feature map in the feature extractor, w i is the weight of the sparse LiDAR feature map at level i, and cos is the cosine similarity function;
[0110] F i+1 Upsample by 2 to F' f ∈R H×W×C , and apply 1×1 convolution operation to Projected into F' r ∈R H×W×C ,Will Projected into F' s ∈R H×W×C , described as:
[0111] F' f = upsample(F i+1 )
[0112]
[0113]
[0114] Among them, upsample is the upsampling operation performed by nearest neighbor interpolation, and f 1×1 represents a 1×1 convolutional layer;
[0115] Upsampled feature maps and corresponding F' r The feature maps are merged by element-wise addition, and a 3×3 convolution is added to each merged feature map to combine the merged features with the applied weight w i The sparse LiDAR features F' s Add them together and the output features are calculated as follows:
[0116] F5=f 3×3 (F' r +w5·F' s )
[0117] Among them, the fusion result F i is the higher-level feature for the next fusion stage, and this process is repeated until the final feature map is generated.
[0118] It should be pointed out here that the device description in the above embodiment corresponds to the method description in the embodiment, and the embodiment of the present invention will not be described in detail here.
[0119] The execution subjects of the above-mentioned processor 1 and memory 2 can be devices with computing functions such as computers, single-chip microcomputers, and microcontrollers. In specific implementation, the embodiments of the present invention do not limit the execution subjects and are selected according to the needs of actual applications.
[0120] The data signal is transmitted between the memory 2 and the processor 1 via the bus 3, which will not be described in detail in the embodiment of the present invention.
[0121] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, the storage medium includes a stored program, and when the program is running, the device where the storage medium is located is controlled to execute the method steps in the above embodiment.
[0122] The computer-readable storage medium includes but is not limited to a flash memory, a hard disk, a solid-state drive, and the like.
[0123] It should be pointed out here that the description of the readable storage medium in the above embodiment corresponds to the description of the method in the embodiment, and the embodiment of the present invention will not be described in detail here.
[0124] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present invention are generated.
[0125] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions may be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be accessed by the computer or a data storage device such as a server or a data center that includes one or more available media. The available medium may be a magnetic medium or a semiconductor medium, etc.
[0126] Unless otherwise specified, the models of the components in the embodiments of the present invention are not limited, and any device that can perform the above functions may be used.
[0127] Those skilled in the art will appreciate that the accompanying drawing is only a schematic diagram of a preferred embodiment, and the serial numbers of the embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0128] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A 3D target detection method based on sparse radar and binocular stereo image fusion, characterized in that: The method comprises: After feature encoding of the stereo image and the sparse LiDAR depth map respectively, the feature information of the two paths is fused based on the attention fusion module, wherein the fusion is from the LiDAR depth map to the stereo image; Based on the output of the corresponding left and right regions of interest from the stereo region extraction network, the left and right feature maps are input together into the stereo regression network branch and the depth prediction branch to predict the position, size and orientation of the 3D bounding box; The method comprises: Add the sparse LiDAR features to the image features and set the weight w for each feature level i , by calculating the correlation between the sparse LiDAR and its corresponding stereo image feature map, we get the correlation score w i , defined as: in, is the i-th pair of stereo image feature map and sparse LiDAR feature map in the feature extractor, w i is the weight of the sparse LiDAR feature map at level i, and cos is the cosine similarity function; F i+1 Upsample by 2 to F f '∈R H×W×C , apply 1×1 convolution operation to F i r Projected into F′ r ∈R H×W×C , F i s Projected into F′ s ∈R H×W×C , described as: F′ f =upsample(F i+1 ) F′ r =f 1×1 (F i r ) F′ s =f 1×1 (F i s ) Among them, upsample is the upsampling operation performed by nearest neighbor interpolation, and f 1×1 represents a 1×1 convolutional layer; The upsampled feature maps and the corresponding F r 'The feature maps are merged by element-by-element addition, and a 3×3 convolution is added to each merged feature map to combine the merged features with the applied weight w i The sparse LiDAR features F s 'Add, the calculation method of output features is as follows: F5=f 3×3 (F′ r +w5·F′ s ) Fusion result F i is the higher-level feature for the next fusion stage, and this process is repeated until the final feature map is generated.
2. The 3D target detection method of sparse radar and binocular stereo image fusion according to claim 1, characterized in that: The stereo regression network branch is used to regress the 2D stereo box, size, viewpoint angle and 2D center; the depth prediction branch is used to predict the univariate depth of the center of the 3D bounding box.
3. The 3D target detection method of sparse radar and binocular stereo image fusion according to claim 1, characterized in that: The attention fusion module fuses the left sparse LiDAR feature map with the corresponding left RGB feature map, and fuses the right sparse LiDAR feature map with the corresponding right RGB feature map.
4. The 3D target detection method of sparse radar and binocular stereo image fusion according to claim 1, characterized in that: The fusion process is: Among them, F i Represents the characteristics of fusion, is the feature output of the last block in each stage of the encoder, Refers to the last output feature of the encoder.
5. A 3D target detection device based on sparse radar and binocular stereo image fusion, characterized in that: The device comprises: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method according to any one of claims 1 to 4.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is enabled to perform the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Three-dimensional target detection method and device based on multi-sensor information fusion
CN110929692A
Camera and laser radar fused end-to-end target detection method
CN111027401A
Cited By
Attention-based refinement for depth completion
US12737900B2
Attention-based refinement for depth completion
US20250054168A1