Multi-sensor fusion sensing method for automatic driving scene

By extracting multi-view image features, radar features and target probability map information, and using a secondary fusion method to fuse multi-sensor data features, the problem of insufficient accuracy of the existing 3D object detection algorithm is solved, and more efficient detection accuracy and robustness are achieved.

CN119963960AActive Publication Date: 2025-05-09JIANGSU ANZIDA TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510037692.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-09
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The existing 3D object detection algorithms have shortcomings in prediction accuracy, especially the lack of depth information of the 2D image-based methods, resulting in low accuracy; while LiDAR is susceptible to severe weather and lighting, and is priced at a high price, which is not suitable for industrial mass production.

Method used

A multi-sensor fusion perception method for autonomous driving scenarios is proposed. By extracting multi-view image features, radar features and target probability map information, a secondary fusion method is adopted. First, the image features and target probability map information are fused to generate image bev features, and then the radar and image bev features are further fused, and the multi-sensor data features are further extracted using bev encoder.

Benefits of technology

By fully integrating the characteristics of multiple sensors, the detection accuracy is improved, the robustness of the algorithm is improved, and it is suitable for object detection in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963960A_ABST
    Figure CN119963960A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, in particular to a multi-sensor fusion sensing method for an automatic driving scene, which comprises the following steps of: adopting a multi-view image and radar point cloud data as input, respectively extracting features through an image and a radar processing network, carrying out secondary feature fusion at a feature level, and finally realizing the detection of a 3D (three-dimensional) target. Radar semantic features and radar target probability graph information are obtained through radar point cloud data, generation of image bev features is guided by fusing the radar target probability graph information, and the target expression ability of the image features is improved. A secondary feature fusion strategy is adopted, radar semantic features and image bev features are fused, and bev encoder is adopted to further extract multi-sensor data features, so that the multi-sensor features can be fused more fully, and the modeling capability of a network to a target is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a multi-sensor fusion perception method for autonomous driving scenarios. Background Art

[0002] At present, 3D object detection, as one of the hot research directions, has attracted extensive attention from academia and industry in the past. According to the different sensors used, 3D object detection algorithms can be divided into camera-based and fusion-based methods. Traditional camera-based 3D object detection algorithms benefit from the development of 2D detectors. On this basis, they expand the network structure and add prediction tasks to achieve 3D object detection. For example, methods such as FCOS3D and DETR3D achieve 3D object prediction by adding predictions of the 3D object's orientation angle, height, and center point spatial position information.

[0003] However, predicting targets in 3D space from 2D images is an ill-posed problem. Due to the lack of target depth information, the prediction accuracy of this type of method is low. The multi-sensor fusion algorithm is mainly LiDAR-Camera based. LiDAR uses laser pulses to scan the environment instead of radio waves or ultrasound. The smaller wavelength means that the LiDAR system can detect smaller objects with high resolution and high accuracy. By fusing the high-precision spatial information of LiDAR and the rich semantic information of the image, the 3D target detection performance is improved.

[0004] However, LiDAR is easily affected by bad weather and lighting, and is expensive, which is not conducive to industrial mass production. Radar has the characteristics of all-weather operation, long detection distance, and relatively low price. Therefore, Radar-Camera fusion detection methods have also been developed, such as Centerfusion, Craft, etc. This type of method is a decision-level fusion method. First, the target information is predicted by the image and radar network respectively, and then the post-processing method is used to associate and match the target and output the final prediction result. Although this type of method has also achieved certain performance, it has a large amount of calculation, the post-processing rules are relatively complex, and it fails to fully utilize the characteristics of multi-sensor fusion. Summary of the invention

[0005] In response to the problems existing in the background technology, the present invention proposes a multi-sensor fusion perception method for autonomous driving scenarios, which extracts multi-view image features, radar features and target probability map information respectively, and adopts a secondary fusion method. First, the generation of image bev features is guided by fusing image features and target probability map information, and then radar and image bev features are further fused. This is a feature-level fusion strategy that can fully fuse the features of multiple sensors and improve detection accuracy.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A multi-sensor fusion perception method for autonomous driving scenarios uses multi-view images and radar point cloud data as input, extracts features through image and radar processing networks respectively, performs secondary feature fusion at the feature level, and finally realizes 3D target detection.

[0008] Furthermore, the processing of the image data specifically includes the following steps:

[0009] (1.1) Input multi-visual image data;

[0010] (1.2) Preprocess the image;

[0011] (1.3) Extract 2D image features;

[0012] (1.4) Perform feature fusion on the extracted features;

[0013] (1.5) Setting a discrete depth information prediction branch, through which the discrete depth information of each feature map pixel is predicted;

[0014] (1.6) 2D to 3D feature conversion: Based on the predicted depth information and camera internal and external parameters, the coordinates of the 2D feature points corresponding to the 3D space are obtained according to the coordinate conversion formula, that is, the feature representation of the 2D feature in the 3D space is obtained;

[0015] (1.7) BEV-pooling module: According to the ego-vehicle coordinate system, a 128*128 bev grid is pre-divided, and the size of each grid is [0.5m, 0.5m]. According to the divided grids, the points in the 3D space are assigned to the corresponding grids, and their feature representation in the bev space is obtained through sum-pooling;

[0016] (1.8) First fusion: The object probability map in the bev space obtained based on radar data is fused with the image bev features, and the target probability predicted by radar is used to guide the further extraction of the image bev features;

[0017] (1.9) Second fusion: The further extracted image bev features and radar semantic features are concatenated along the channel dimension, and the concatenated features are further fused through the bev encoder;

[0018] (1.10) Prediction head: Predicts the length, width, height, center point position, orientation angle, speed and category information of the target through discrete prediction branches.

[0019] Preferably, the image preprocessing in step (1.2) is performed by using a random flipping, scaling, and rotation preprocessing method.

[0020] Preferably, the feature fusion in step (1.4) is performed through a FPN (feature pyramid network) network.

[0021] Furthermore, the conversion formula of step (1.6) is shown as (1):

[0022]

[0023] where d i is the predicted depth value, K is the camera internal parameter, R is the rotation matrix from the camera coordinate system to the world coordinate system, t is the translation matrix, z i When the depth value of a 2D pixel is d i The corresponding 3D coordinate height value.

[0024] Furthermore, the processing of the radar point cloud data specifically includes the following steps:

[0025] (2.1) Voxelization: First, the irregular point cloud data is voxelized to obtain radar point cloud features;

[0026] (2.2) Radar feature extraction network, to obtain the BEV feature representation of radar data;

[0027] (2.3) Radar bev feature extraction, further extract radar features;

[0028] (2.4) Extraction of radar target overview and semantic features: Through the decoupled radar prediction branch, the radar target probability map and semantic feature representation are obtained, which are then secondary fused with the image features.

[0029] Preferably, the specific steps of step (2.2) are to perform max pooling on the point cloud data along the feature dimension through a fully connected layer network to obtain voxel features, and assign the features to corresponding grids according to the divided bev grids to obtain the bev feature representation of the radar data.

[0030] Preferably, the step (2.3) further extracts radar features through ResNet-18 and FPN structures.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] The present invention designs a new radar-camera fusion network, which can improve the target detection accuracy in complex scenes, make full use of the complementary characteristics of multiple sensors, and enhance the robustness of the algorithm.

[0033] The present invention obtains radar semantic features and radar target probability map information through radar point cloud data, and guides the generation of image bev features by fusing radar target probability map information, thereby improving the image feature's ability to express the target. A secondary feature fusion strategy is adopted, by fusing radar semantic features and image bev features, and using bev encoder to further extract multi-sensor data features, which can more fully fuse multi-sensor features and improve the network's ability to model targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:

[0035] Figure 1 The present invention is a flow chart of the detection method.

[0036] Figure 2 This is a diagram of the radar point cloud data processing network structure of the present invention.

[0037] Figure 3 This is a visualization diagram of some detection results of the embodiments of the present invention. DETAILED DESCRIPTION

[0038] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.

[0039] like Figure 1 and Figure 2 As shown, the multi-sensor fusion perception method for autonomous driving scenarios of the present invention uses multi-view images and radar point cloud data as input, extracts features through image and radar processing networks respectively, fuses them at the feature level, and finally realizes the detection of 3D targets. The algorithm detection process is divided into two main lines, one is the processing of image data, and the other is the processing of radar point cloud data. The two main lines are described in detail below.

[0040] 1. Image data processing

[0041] Step 0: Input multi-view image data.

[0042] Step 1: Image preprocessing: This method uses random flipping, scaling, and rotation for preprocessing.

[0043] Step 2: 2D image backbone network. In this method, ResNet-50 is selected as the backbone network to extract the features of 2D images.

[0044] Step 3: Feature fusion network. In this method, the extracted features are fused through the feature pyramid network (FPN).

[0045] Step 4: Discrete depth information prediction branch. This branch predicts the discrete depth information of each pixel in the feature map.

[0046] Step 5: 2D-to-3D feature conversion module. Based on the predicted depth information and camera internal and external parameters, the coordinates of the 2D feature points corresponding to the 3D space are obtained according to the coordinate conversion formula, that is, the feature representation of the 2D feature in the 3D space is obtained. The conversion formula is shown in (1). i is the predicted depth value. In this method, the predicted depth value range is [1m, 60m], with an interval of 1m. K is the camera internal parameter, R is the rotation matrix from the camera coordinate system to the world coordinate system, t is the translation matrix, and z i When the depth value of a 2D pixel is d i The corresponding 3D coordinate height value.

[0047]

[0048] Step 6: BEV-pooling module. According to the ego-vehicle coordinate system, a 128*128 bev grid is pre-divided, and the size of each grid is [0.5m, 0.5m]. According to the divided grids, the points in the 3D space are assigned to the corresponding grids, and their feature representation in the bev space is obtained through sum-pooling.

[0049] Step 7: First fusion. The probability map of the object in the bev space is extracted based on the radar data, and then fused with the image bev features. The target probability predicted by radar is used to guide the further extraction of the image bev features.

[0050] Step 8: Second fusion. The further extracted image bev features and radar semantic features are concatenated along the channel dimension, and the concatenated features are further fused through the bev encoder.

[0051] Step 9: Prediction head: Through discrete prediction branches, the length, width, height, center point position, orientation angle, speed and category information of the target are predicted.

[0052] 2. Radar data processing

[0053] Step 1: Voxelization. First, the irregular point cloud data is voxelized to obtain radar point cloud features.

[0054] Step 2: Radar feature extraction network. The point cloud data is maxpooled along the feature dimension through the fully connected layer network to obtain voxel features, and the features are assigned to the corresponding grids according to the divided bev grids to obtain the bev feature representation of the radar data.

[0055] Step 3: Radar bev feature extraction: Radar features are further extracted through ResNet-18 and FPN structures.

[0056] Step 4: Extraction of radar object probability map and semantic features. Through the decoupled radar prediction branch, the radar target probability map and semantic feature representation are obtained, and then they are fused with the image features for a second time.

[0057] Example: 3D target detection experiment of the multi-sensor fusion perception method for autonomous driving scenarios (hereinafter referred to as RCFusionNet) of the present invention

[0058] 1) Comparison of test results

[0059] This experiment uses the nuScenes dataset, which contains 1,000 scenes, each lasting 20 seconds. The dataset contains a total of 1.4M images, 1.4M frames of radar data, and 1.4M annotated boxes. These annotated boxes are 23 categories of 3D detection boxes annotated at 2Hz. The annotation content includes category, translation matrix in global coordinate system (unit meter), rotation matrix represented by quaternion, 3D size (unit meter) and other information.

[0060] Table 1 shows the performance comparison between this method and other detection methods. Our method RCFusionNet achieved 37.3% on mAP and 46.6% on NDS. In particular, compared with BEVDet using the Swin-T backbone network, our method improved by 6.1% on mAP and 7.4% on NDS, even though we used a relatively weak 2D backbone network. In addition, compared with the BEVDepth method, our method improved by 2.2% on mAP. Furthermore, when we used ResNet101 as the backbone network, by comparing the performance changes brought by different backbone networks, the model achieved 47.5% on NDS and 38.1% on mAP. This result shows that using a stronger 2D image feature backbone network can effectively improve the performance of the model.

[0061] Table 1-1 Comparison of detection performance of this method and other methods

[0062]

[0063] like Figure 3 The following are some of the detection results of this method. In the comparison of visualization results from the BEV perspective, we show the detection results of baseline, RCFusionNet-C and RCFusionNet respectively. Among them, RCFusionNet-C indicates the version of this method that only uses the image sensor. The yellow solid line box indicates the prediction result, and the blue solid line box indicates the real annotation box. Compared with the baseline and RCFusionNet-C without radar data fusion, this method performs better in the detection of long-distance targets. According to the enlarged image of the local detection result (indicated by the orange box), it actually contains 4 real targets. The baseline model predicts 6 targets, including 2 false detections; RCFusionNet-C predicts 5 targets, including 1 false detection. Through the comparison, it can be intuitively seen that the improvement of the image network has brought about a significant improvement in performance. RCFusionNet successfully detected 4 targets, and the size and orientation of the predicted targets were closer to the actual situation, further verifying that the detection performance was significantly improved after fusing radar data.

[0064] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A multi-sensor fusion perception method for autonomous driving scenarios, characterized by: Multi-view images and radar point cloud data are used as input, features are extracted through image and radar processing networks respectively, and secondary feature fusion is performed at the feature level to finally achieve 3D target detection.

2. The multi-sensor fusion perception method for autonomous driving scenarios according to claim 1, characterized in that: The processing of the image data specifically comprises the following steps: (1.1) Input multi-visual image data; (1.2) Preprocess the image; (1.3) Extract 2D image features; (1.4) Perform feature fusion on the extracted features; (1.5) Setting a discrete depth information prediction branch, through which the discrete depth information of each feature map pixel is predicted; (1.6) 2D to 3D feature conversion: Based on the predicted depth information and camera internal and external parameters, the coordinates of the 2D feature points corresponding to the 3D space are obtained according to the coordinate conversion formula, that is, the feature representation of the 2D feature in the 3D space is obtained; (1.7) BEV-pooling module: According to the ego-vehicle coordinate system, a 128*128 bev grid is pre-divided, and the size of each grid is [0.5m, 0.5m]. According to the divided grids, the points in the 3D space are assigned to the corresponding grids, and their feature representation in the bev space is obtained through sum-pooling; (1.8) First fusion: The target probability map in the bev space obtained based on radar data is fused with the image bev features, and the target probability predicted by radar is used to guide the further extraction of image bev features; (1.9) Second fusion: The further extracted image bev features and radar semantic features are concatenated along the channel dimension, and the concatenated features are further fused through the bev encoder; (1.10) Prediction head: Predicts the length, width, height, center point position, orientation angle, speed and category information of the target through discrete prediction branches.

3. The multi-sensor fusion perception method for autonomous driving scenarios according to claim 2, characterized in that: The image preprocessing in step (1.2) is performed by using a random flipping, scaling, and rotation preprocessing method.

4. The multi-sensor fusion perception method for autonomous driving scenarios according to claim 2, characterized in that: The feature fusion in step (1.4) is performed through the FPN network.

5. The multi-sensor fusion perception method for autonomous driving scenarios according to claim 2, characterized in that: The conversion formula of step (1.6) is shown in (1): where d i is the predicted depth value, K is the camera internal parameter, R is the rotation matrix from the camera coordinate system to the world coordinate system, t is the translation matrix, z i When the depth value of the 2D pixel is d i The corresponding 3D coordinate height value.

6. The multi-sensor fusion perception method for autonomous driving scenarios according to claim 2, characterized in that: The processing of the radar point cloud data specifically includes the following steps: (2.1) Voxelization: First, the irregular point cloud data is voxelized to obtain radar point cloud features; (2.2) Radar feature extraction network, to obtain the BEV feature representation of radar data; (2.3) Radar bev feature extraction, further extract radar features; (2.4) Extraction of radar target overview and semantic features: Through the decoupled radar prediction branch, the radar target probability map and semantic feature representation are obtained, which are then secondary fused with the image features.

7. The multi-sensor fusion perception method for autonomous driving scenarios according to claim 5, characterized in that: The specific steps of step (2.2) are to perform max pooling on the point cloud data along the feature dimension through a fully connected layer network to obtain voxel features, and assign the features to corresponding grids according to the divided bev grids to obtain the bev feature representation of the radar data.

8. The multi-sensor fusion perception method for autonomous driving scenarios according to claim 5, characterized in that: The step (2.3) further extracts radar features through ResNet-18 and FPN structures.

Citation Information

Patent Citations

  • Ground-removed multi-modal fusion 3d target detection method

    CN116994239A

  • 3D target detection system and method based on 4D millimeter wave radar and camera fusion

    CN117452396A

  • Unified BEV representation-based multi-modal fusion target detection method

    CN117727026A

  • Automatic driving scene detection method and device, equipment and medium

    CN118134835A

  • Method and apparatus for target detection, electronic device, and machine-readable storage medium

    WO2024208081A1