Cross-modal fusion method for three-dimensional target detection

Through the bilateral content-aware fusion (BCAF) method, lidar and image features are extracted and aggregated, and feature modal fusion is performed through soft computing, which solves the problem of unfavorable information enhancement when fusion of lidar and image data in the prior art, and achieves more efficient three-dimensional object detection performance.

CN119942516APending Publication Date: 2025-05-06SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311816946.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing lidar camera fusion method lacks an effective bilateral strategy when fusing lidar and image features, resulting in adverse information enhancement and affecting 3D object detection performance. The unilateral fusion strategy cannot fully utilize the complementarity between different data modes.

Method used

A bilateral content-aware fusion (BCAF) method is proposed. By extracting lidar features and image features and performing aggregation operations, the feature modal fusion is performed through soft operations to achieve favorable clue selection of lidar and image data.

Benefits of technology

The performance of cross-modal three-dimensional object detectors is effectively improved, and through the bilateral content perception mechanism and soft fusion module, the complementarity between different data modes is fully utilized, and the accuracy and effectiveness of feature fusion is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942516A_ABST
    Figure CN119942516A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-modal fusion method for three-dimensional target detection, and the method comprises the steps: respectively extracting laser radar features and image features according to an RGB image and an original point cloud; aggregation operation is carried out according to the laser radar features and the image features, and aggregation features are obtained; and carrying out feature modal fusion on the aggregated features through soft operation to realize cross-modal information fusion. According to the cross-modal fusion method for three-dimensional target detection, the features of the image and the laser radar can be effectively enhanced, guidance is provided for feature fusion, favorable clues for selecting the features of the laser radar camera are selected, complementary information of different data modals is fully utilized, and the performance of a cross-modal three-dimensional object detector is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a cross-modal fusion method for three-dimensional target detection. Background Art

[0002] Intelligent agents such as self-driving cars need to understand their surroundings to complete reliable and efficient mobility tasks. As a fundamental component of 3D environment perception, 3D object detection aims to identify and localize objects in the real world. Typically, the 3D perception system of self-driving vehicles is equipped with multiple sensors. According to the application of different sensors, current 3D object detection methods can be divided into camera-only methods, lidar-only methods, and lidar-camera fusion methods. Fusion of lidar and camera modalities can utilize geometric information in point clouds and semantic information in images to generate enhanced representations, which is beneficial for 3D object detection tasks. Despite the impressive improvements, these fusion methods still have two problems. First, they simply concatenate the initial point cloud features and image features, which may inevitably enhance unfavorable information and further affect the performance of 3D object detection tasks. If this information is fused without constraints, false positives may be generated. Second, only considering unilateral fusion strategies cannot fully exploit the complementarity between different data modalities. Distant objects with sparse geometric structures can be enhanced by favorable semantic information, while occluded objects with weak semantic cues can be enhanced by point scan geometric information. Therefore, an effective bilateral strategy is needed when fusing lidar and image features. Summary of the invention

[0003] Aiming at the shortcomings of the existing technology, a cross-modal fusion strategy called bilateral content-aware fusion (BCAF) is proposed to overcome the shortcomings of the lidar camera data point-level fusion method.

[0004] The technical solution adopted by the present invention to achieve the above-mentioned purpose is: a cross-modal fusion method for three-dimensional target detection, comprising the following steps:

[0005] Step 1: Extract lidar features and image features based on RGB images and original point clouds respectively;

[0006] Step 2: Perform aggregation operations based on the lidar features and image features to obtain aggregated features respectively;

[0007] Step 3: Use soft computing to fuse the aggregated features into feature modalities to achieve information fusion of the lidar camera.

[0008] The aggregation operation is performed according to the laser radar features and the image features to obtain the aggregated features, respectively, including:

[0009] LCA based on image features and lidar features As input, then, in The compression layer and encoder layer are then connected to generate the lidar perception core

[0010]

[0011] C LE and C LC The convolutional layers representing the encoder operation and compression operation have parameters respectively. and For image features Any target location l i , i∈4096, in the lidar perception core There is a corresponding position l i ′ ; The entire perceptual convolution is:

[0012]

[0013] in, is the lidar feature.

[0014] The aggregation operation is performed according to the laser radar features and the image features to obtain the aggregated features, respectively, including:

[0015] ICA is based on image features and lidar features As input, then, in The encoder layer and the compression layer are then connected to generate the image-aware kernel

[0016]

[0017] C IE and C IC The convolutional layers representing the encoder operation and compression operation have parameters respectively. and For LiDAR features Any target location l i , i∈4096, in the image perception kernel There is a corresponding position l i ′ ; The entire perceptual convolution is:

[0018]

[0019] in is the perceptual image feature.

[0020] The aggregation operation is performed according to the laser radar features and the image features to obtain the aggregated features, respectively, including:

[0021] Use two parallel fully connected layers and two Sigmoid functions to generate and Aggregate representation of

[0022]

[0023]

[0024] Among them, FCN represents the aggregation operation, the subscript I represents the aggregation of perceived image features, the subscript L represents the aggregation of lidar features; Sigmoid represents the activation function.

[0025] The step of fusing the aggregated features through soft operations on the feature modalities includes the following steps:

[0026] The soft operation converts the aggregate features into and multiply;

[0027] After the soft operation, the final cross-modal features are output through the Conv-BN-ReLU block

[0028]

[0029] Conv represents four 3×3 convolutional layers.

[0030] The Conv-BN-ReLU block is composed of four 3×3 convolutions, BatchNorm, and ReLU connected in sequence.

[0031] A cross-modal fusion system for three-dimensional object detection, comprising:

[0032] A feature extraction module is used to extract lidar features and image features based on the RGB image and the original point cloud, respectively;

[0033] An aggregation operation module, used for performing aggregation operations according to the laser radar features and the image features to obtain aggregate features respectively;

[0034] The modal fusion module is used to fuse the aggregated features into feature modalities through soft operations to achieve information fusion of the lidar camera.

[0035] A cross-modal fusion device for three-dimensional target detection comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement the cross-modal fusion method for three-dimensional target detection when executing the computer program.

[0036] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a cross-modal fusion method for three-dimensional target detection is implemented.

[0037] The present invention has the following beneficial effects and advantages:

[0038] The present invention proposes a new cross-modal fusion method for 3D object detection. Instead of using a unilateral structure that only considers the complementary information of one data with another data, the bilateral structure incorporates the mutual influence between different data modes. First, the multimodal content-aware mechanism aims to utilize different data modes to enhance each other and provide guidance for the final feature fusion. Then, the bilateral soft fusion module fuses the content-aware features and the initial features in a soft manner to select favorable cues for the lidar and image data. Experiments on the KITTI benchmark show that the present invention can effectively improve the performance of cross-modal 3D object detectors. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a structural diagram of the bilateral content-aware fusion strategy;

[0040] Figure 2 This is the structure diagram of the lidar content perception module and the image content perception module. DETAILED DESCRIPTION

[0041] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0042] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation method of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the invention, so the present invention is not limited by the specific implementation disclosed below.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0044] A new cross-modal fusion method for 3D object detection includes the following process:

[0045] Given a single RGB image and a raw point cloud, the region proposal network extracts lidar features and image features, respectively. For lidar features and image features, a bilateral content-aware mechanism is applied to fuse different feature modalities. After fusing the final upsampled features, the initial ROI and point cloud confidence are generated. Finally, the candidate box refinement stage can filter redundant candidate boxes and output a refined 3D bounding box.

[0046] The development process and implementation functions of the present invention are specifically as follows:

[0047] Step 1: Process and expand the RGB image and point cloud data to form the dataset as input.

[0048] Step 2: Establish an image backbone, take a single RGB image as input, and use a lightweight convolutional neural network (CNN) to obtain image features.

[0049] Step 3: Build the LiDAR backbone and use the popular PointNet++ consisting of an encoder part and a decoder part to extract features of 16,384 sampling points from the original point cloud.

[0050] Step 4: A bilateral content-aware soft fusion module is established between the lidar features and the image features, which includes two core components, namely, two content-aware branches for lidar camera functions and a bilateral fusion module to fuse different feature modalities.

[0051] Among them, the bilateral content-aware fusion strategy is a key part of the cross-modal 3D object detector, which is applied between lidar features and image features. It includes two core components, namely two content-aware branches of lidar camera functions and a bilateral fusion module.

[0052] The content-aware (CA) method uses a feature modality as a content-aware kernel to perform convolution operations with another feature modality. It is divided into lidar content-aware (LCA) and image content-aware (ICA). Before feature-aware enhancement, lidar collaborative image features are selected from the initial image features according to the 2D projection coordinates of the point cloud features. The LCA of the first downsampling layer is used as an example for specific explanation. Figure 2 As shown, LCA is based on and The features of are taken as input. Then, The compression layer and encoder layer are then connected to generate the lidar perception core

[0053]

[0054] C LE and C LC The convolutional layers representing the encoder operation and compression operation are respectively, and their parameters are and For image features Any target location l i (i∈4096), in the lidar perception core There is a corresponding position l i ′ (i∈4096). According to the position correspondence, we can use the position l i ′ The sub-nucleus at position l i Where The sub-features of are convolved in 1D. The entire perceptual convolution can be defined as:

[0055]

[0056] in is the lidar feature.

[0057] Similarly, image perception kernels and sensing lidar features can be expressed as:

[0058]

[0059] C IE and C IC The convolutional layers representing the encoder operation and compression operation are respectively, and their parameters are and For image features Any target location l i (i∈4096), in the image perception kernel There is a corresponding position l i ′ (i∈4096). According to the position correspondence, we can use the position l i ′ l i ′ The sub-nucleus and position l i Where The sub-features of are convolved in 1D. The entire perceptual convolution can be defined as:

[0060]

[0061] in is the perceptual image feature.

[0062] The bilateral content-aware fusion strategy bilateral soft fusion (BSF) module fuses different feature modes in a soft way. First, two parallel fully connected layers and two Sigmoid functions are used to generate and Aggregate representation of It can be defined as:

[0063]

[0064]

[0065] The soft operation converts the aggregate features into and After the soft operation, the Conv-BN-ReLU block is cascaded as input and outputs the final cross-modal features

[0066]

[0067] The candidate box refinement part first uses the non-maximum suppression (NMS) method to filter the redundant candidate boxes generated by the region candidate box network. For each remaining candidate box, the method in PointRCNN is used to generate a feature descriptor ROI. Before the final output, three SA layers are set to extract the enhanced features of the ROI. Finally, two branches with two cascaded 1×1 convolutional layers are deployed for fine classification and regression respectively.

Claims

1. A cross-modal fusion method for three-dimensional object detection, characterized in that: The following steps are involved: Step 1: Extract lidar features and image features based on RGB images and original point clouds respectively; Step 2: Perform aggregation operations based on the lidar features and image features to obtain aggregated features respectively; Step 3: Use soft computing to fuse the aggregated features into feature modalities to achieve information fusion of the lidar camera.

2. The cross-modal fusion method for three-dimensional object detection according to claim 1, characterized in that: The aggregation operation is performed according to the laser radar features and the image features to obtain the aggregated features, respectively, including: LCA based on image features and lidar features As input, then, in The compression layer and encoder layer are then connected to generate the lidar perception core C LE and C LC The convolutional layers representing the encoder operation and compression operation have parameters respectively. and For image features Any target location l i , i∈4096, in the lidar perception core There is a corresponding position l′ i ; The entire perceptual convolution is: in, is the lidar feature.

3. The cross-modal fusion method for three-dimensional object detection according to claim 1, characterized in that: The aggregation operation is performed according to the laser radar features and the image features to obtain the aggregated features, respectively, including: ICA is based on image features and lidar features As input, then, in The encoder layer and the compression layer are then connected to generate the image-aware kernel C IE and C IC The convolutional layers representing the encoder operation and compression operation have parameters respectively. and For LiDAR features Any target location l i , i∈4096, in the image perception kernel There is a corresponding position l′ i ; The entire perceptual convolution is: in is the perceptual image feature.

4. The cross-modal fusion method for three-dimensional object detection according to claim 1, characterized in that: The aggregation operation is performed according to the laser radar features and the image features to obtain the aggregated features, respectively, including: Use two parallel fully connected layers and two Sigmoid functions to generate and Aggregate representation of Among them, FCN represents the aggregation operation, the subscript I represents the aggregation of perceived image features, the subscript L represents the aggregation of lidar features; Sigmoid represents the activation function.

5. The cross-modal fusion method for three-dimensional object detection according to claim 1, characterized in that: The step of fusing the aggregated features through soft operations on the feature modalities includes the following steps: The soft operation converts the aggregate features into and multiply; After the soft operation, the final cross-modal features are output through the Conv-BN-ReLU block Conv represents four 3×3 convolutional layers.

6. The cross-modal fusion method for three-dimensional object detection according to claim 1, characterized in that: The Conv-BN-ReLU block is composed of four 3×3 convolutions, BatchNorm, and ReLU connected in sequence.

7. A cross-modal fusion system for three-dimensional object detection, characterized in that: include: A feature extraction module is used to extract lidar features and image features based on the RGB image and the original point cloud, respectively; An aggregation operation module, used for performing aggregation operations according to the laser radar features and the image features to obtain aggregate features respectively; The modal fusion module is used to fuse the aggregated features into feature modalities through soft operations to achieve information fusion of the lidar camera.

8. A cross-modal fusion device for three-dimensional object detection, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement a cross-modal fusion method for three-dimensional target detection as described in any one of claims 1-6 when executing the computer program.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, a cross-modal fusion method for three-dimensional target detection as described in any one of claims 1-6 is implemented.