Multi-modal 3D target detection method based on multi-branch feature extraction

The multi-branch feature extraction method addresses slow inference speeds and suboptimal network designs in 3D target detection by using Split-Inception and Split-Neck modules to enhance point cloud features, achieving real-time performance and improved computational efficiency in autonomous driving.

CN120318789APending Publication Date: 2025-07-15BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510401788.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the existing 3D object detection methods for autonomous driving, the model inference speed is slow, the computing overhead is high, and the point cloud feature extraction network lacks modern design, resulting in the inability to meet real-time requirements and performance improvements.

Method used

The multi-branch feature extraction method is used to extract image and point cloud features respectively through camera branches and lidar branches, and feature processing and fusion is performed using the Split-Inception convolution module and the Split-Neck neck module, including view conversion, multi-layer perceptron and activation function operations to achieve feature richness and multi-scale fusion.

Benefits of technology

It improves the inference speed of the model, reduces the computational overhead, enhances the receptive field of point cloud features, meets the real-time requirements and improves detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318789A_ABST
    Figure CN120318789A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal 3D target detection method based on multi-branch feature extraction, which is characterized in that point cloud data processing is improved on the basis of an existing BEV feature-based multi-modal 3D target detection framework, a Splitt-Inception convolution module is introduced, and corresponding feature extraction is performed on different features divided along a channel; therefore, the point cloud features are enriched and the receptive field is enlarged; and meanwhile, a lightweight Splitt-Neck neck module is used for carrying out extra downsampling of different scales on each channel, and feature multi-scale fusion is realized on the premise of not greatly increasing the number of the channels in multiples, so that the calculation overhead is effectively saved, and the method can provide wider applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving 3D object detection, and specifically relates to a 3D object detection method based on the fusion of lidar point cloud and camera image features in the BEV space. Background Art

[0002] Currently, fusing lidar and camera data in autonomous driving 3D object detection is a generally recognized excellent solution in this field. Most of the existing such methods first generate BEV features of images and point clouds through an independent dual-branch framework, and then perform modality fusion. However, this method has two main problems: First, the model inference speed of the method is relatively slow, generally not exceeding 5FPS / s, far lower than the real-time requirement of 20FPS / s. This is because most of the lidar branches in the model use voxel-based networks (such as VoxelNet), and the computational cost of the 3D sparse convolution operators used by them is significantly greater than that of ordinary 2D convolution operators, so it will lead to long model training and inference times. There are also some models that use a Transformer-based network architecture for feature modeling, and the computational complexity of its multi-head self-attention is O(N 2 ), and the computational cost will increase rapidly with the amount of data. Second, the point cloud feature extraction networks of these models mentioned above lack modern designs, and most of them are still based on the original architectures of SECOND or PointPillars, and the model performance has not been fully exploited. Summary of the Invention

[0003] In view of this, aiming at the technical problems existing in this field, the present invention provides a multi-modal 3D object detection method based on multi-branch feature extraction, which specifically includes the following steps:

[0004] Step 1: Extract image features through the camera branch, predict the pixel depth, and map it to the BEV space by a view transformer to form the image BEV feature F C ;

[0005] Step 2: Process the point cloud data obtained by the lidar branch into voxels, and then process the voxels into a pseudo-image F(H, W, C), where H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map; and use the point cloud feature extraction network to expand the feature channels through staged downsampling, so as to obtain the refined point cloud feature F L ; wherein, each stage of the point cloud feature extraction network includes a plurality of Split-Inception convolution modules;

[0006] Step 3: The refined point cloud feature F LFeed it into a Split-Neck neck module for processing; the point cloud feature F of the neck module L Based on channel grouping, perform downsampling operations at different multiples to obtain corresponding feature maps, and then use transposed convolution and upsampling to obtain a feature map in a unified form. Concatenate and process the feature maps to obtain the point cloud feature F'. L , and finally, after being processed by a multi-layer perceptron and an activation function, output the point cloud feature F". L ;

[0007] Step 4: Concatenate the image BEV feature F obtained in Step 1 C and the point cloud feature F" obtained in Step 3 L along the channel dimension and fuse them to obtain the fused feature F LC ;

[0008] Step 5: Send the fused feature F LC into the detection head to obtain the final object detection result.

[0009] Furthermore, the point cloud feature extraction network used in Step 2 specifically includes 3 stages, and each stage sequentially includes 3, 4, and 6 Split-Inception convolution modules; among them, a downsampling operation is performed once between adjacent two stages, and each time the feature map is downsampled by a factor of 2, and finally a 4x feature map that is downsampled by a factor of 4 is obtained, and the refined point cloud feature F L (H / 4, W / 4, 4C) is output by the point cloud feature extraction network;

[0010] The working process of the Split-Inception convolution module specifically includes:

[0011] First, for the given input feature X(H, W, C), divide it into 4 groups along the channel dimension:

[0012] X s , X l , X p , X i = Split(X)

[0013] = X :,:g , X :g:2g , X :2g:3g , X :3g:

[0014] Among them, g is the number of channels of the convolutional branch, and there is:

[0015] g = r s C

[0016] Among them, r s is the channel division rate, and the default setting is 1 / 8;

[0017] After that, the divided features are sent to parallel feature extraction branches:

[0018]

[0019] X′ p = MaxPool(X p )

[0020] X′ i = X i

[0021] In the formula, represents a 2D convolution with a convolution kernel size of k×k, an input channel of C i , and an output channel of C o ; k s represents the size of the small convolution kernel, taken as 3; k l represents the size of the large convolution kernel, taken as 5; MaxPool represents max pooling of 3×3, without changing the number of channels;

[0022] After that, the outputs of each branch are concatenated along the channel dimension and batch normalization is performed:

[0023] X′ = Concat(X′ s , X′ l , X′ p , X′ i )

[0024] Y = BatchNorm(X′)

[0025] Finally, the features are sent to a multi-layer perceptron (MLP) for inter-channel feature interaction; the MLP contains two fully connected layers and an activation function, and the activation function is located between the two fully connected layers;

[0026] The original input feature X is added to the output of the MLP, and then the final output feature Y′ is obtained through an activation function:

[0027]

[0028] Y′ = σ[MLP(Y) + X]

[0029] In the formula, r is the expansion rate of the MLP, and σ is the activation function ReLU.

[0030] Furthermore, the specific process of using the Split-Neck neck module to output the point cloud feature F″ L in step three includes:

[0031] First, the refined point cloud feature F L(H / 4, W / 4, 4C) is divided into 3 groups along the channel dimension, and 8x downsampling, 16x downsampling, and identity mapping (i.e., 4x downsampling) are performed respectively. The number of channels in each group is as follows:

[0032] C8 = C 16 = 4r' s C

[0033] C4 = 4(1 - 2r' s )C

[0034] where C n represents the number of channels in the n-fold downsampling branch; r' s represents the channel division rate of the neck module, with a default value of 1 / 4, satisfying C4 = 2C and C8 = C 16 = C; then the divided feature representations are F4(H / 4, W / 4, 2C), F8(H / 4, W / 4, C), F 16 (H / 4, W / 4, C); where F n refers to the feature map of the n-fold downsampling branch;

[0035] After that, downsampling operations are performed on F8 and F 16 respectively to obtain the corresponding 8x feature map F'8(H / 8, W / 8, C) and 16x feature map F' 16 (H / 16, W / 16, C):

[0036] F'8 = Down(F8)

[0037] F' 16 = Down(Down(F 16 ))

[0038] In the formula, Down(·) is a convolutional layer composed of a downsampling layer and 3 Split-Inception modules:

[0039] Down(·) = SI(SI(SI(dn(·))))

[0040] In the formula, SI(·) represents a Split-Inception module; dn(·) is the downsampling layer, implemented by a residual block; the residual block consists of two 3x3 convolutions and a shortcut connection; among them, the stride of the first 3x3 convolution is set to 2 to perform downsampling; the stride of the second 3x3 convolution is 1;

[0041] After that, the 8x feature map and the 16x feature map are each upsampled to a 4x feature map through transposed convolution to obtain F''8(H / 4, W / 4, C) and F'' 16(H / 4, W / 4, C), and splice the 4x feature maps of the three branches along the channel dimension to obtain the point cloud feature F' containing multi-scale information L (H / 4, W / 4, 4C):

[0042] F' L = Concat(F4, F″8, F″ 16 )

[0043] Finally, input the point cloud feature F' L into a multi-layer perceptron, and add the output to F L and then pass through an activation function to obtain the final point cloud feature F″ L :

[0044] F″ L = σ[MLP(F' L ) + F L .

[0045] The above multi-modal 3D object detection method based on multi-branch feature extraction provided by the present invention makes improvements in the processing of point cloud data on the basis of the existing multi-modal 3D object detection framework based on BEV features. By introducing the Split-Inception convolution module and performing corresponding feature extraction on different features divided along the channel, the point cloud features are enriched and the receptive field is increased; at the same time, the lightweight Split-Neck neck module is used to perform additional downsampling of different scales on each channel respectively, and feature multi-scale fusion is realized without doubling the number of channels in large quantities, thus effectively saving the computational cost and enabling the method to provide wider applicability. Brief Description of the Drawings

[0046] Figure 1 is the flow framework diagram of the method provided by the present invention;

[0047] Figure 2 is the schematic diagram of the working process of the Split-Inception convolution module;

[0048] Figure 3 is the schematic diagram of the working process of the Split-Neck neck module. Detailed Embodiments

[0049] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] The multi-modal 3D object detection method based on multi-branch feature extraction provided by the present invention, as Figure 1 shown, specifically includes the following steps:

[0051] Step 1: Extract image features through the camera branch, predict the pixel depth, and map it to the BEV space by the view converter to form the image BEV feature F C ;

[0052] Step 2: Process the point cloud data obtained by the lidar branch into voxels, and then process the voxels into a pseudo-image F(H, W, C), where H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map; and use the point cloud feature extraction network to expand the feature channels through staged downsampling, so as to obtain the refined point cloud feature F L ; Among them, each stage of the point cloud feature extraction network includes multiple Split-Inception convolution modules;

[0053] Step 3: Feed the refined point cloud feature F L into a Split-Neck neck module for processing; this neck module groups the point cloud feature F L based on channels, and performs downsampling operations with different multiples to obtain corresponding feature maps, and then uses transposed convolution and upsampling to obtain a unified form of feature maps, and splices and processes the feature maps to obtain the point cloud feature F' L Finally, after being processed by a multi-layer perceptron and an activation function, the point cloud feature F″ L is output;

[0054] Step 4: Splice the image BEV feature F C obtained in Step 1 and the point cloud feature F″ L obtained in Step 3 along the channel dimension, and fuse them to obtain the fused feature F LC ;

[0055] Step 5: Feed the fused feature F LC into the detection head to obtain the final object detection result.

[0056] In a preferred embodiment of the present invention, the point cloud feature extraction network used in Step 2 specifically includes 3 stages, and each stage includes 3, 4, and 6 Split-Inception convolution modules in sequence; among them, a downsampling operation is performed once between adjacent two stages, and each time the feature map is downsampled by a factor of 2, and finally a 4x feature map that is downsampled by a factor of 4 is obtained, and the refined point cloud feature F L (H / 4, W / 4, 4C) is output by the point cloud feature extraction network;

[0057] The working process of the Split-Inception convolution module is asFigure 2 As shown, specifically including:

[0058] First, for the given input feature X(H, W, C), it is divided into 4 groups along the channel dimension:

[0059] X s , X l , X p , X i = Split(X)

[0060] = X :,:g , X :g:2g , X :2g:3g , X :3g:

[0061] Among them, g is the number of channels of the convolutional branch, and there is:

[0062] g = r s C

[0063] Among them, r s is the channel division rate, and the default setting is 1 / 8;

[0064] After that, the divided features are sent into parallel feature extraction branches:

[0065]

[0066] X' p = MaxPool(X p )

[0067] X' i = X i

[0068] In the formula, represents a 2D convolution with a convolution kernel size of k×k, an input channel of C i , and an output channel of C o ; k s represents the size of the small convolution kernel, which is taken as 3; k l represents the size of the large convolution kernel, which is taken as 5; MaxPool represents a 3×3 max pooling, without changing the number of channels;

[0069] After that, the outputs of each branch are concatenated along the channel dimension and batch normalization is performed:

[0070] X' = Concat(X' s , X' l , X' p , X' i )

[0071] Y = BatchNorm(X')

[0072] The purpose of batch normalization is to accelerate the network convergence speed, reduce the risk of gradient vanishing or explosion, and reduce overfitting.

[0073] Finally, the features are fed into a multi-layer perceptron (MLP) for inter-channel feature interaction; the MLP contains two fully connected layers and an activation function, and the activation function is located between the two fully connected layers;

[0074] The original input feature X is added to the output of the MLP, and then the final output feature Y′ is obtained through an activation function:

[0075]

[0076] Y′ = σ[MLP(Y) + X]

[0077] In the formula, r is the expansion rate of the MLP, which is used to adjust the number of channels in the hidden layer; σ is the activation function ReLU, whose role is to enhance the non-linearity of the neural network, improve the expression ability and computational efficiency of the model, and effectively alleviate the problem of gradient vanishing.

[0078] In a preferred embodiment of the present invention, in step three, the Split-Neck neck module is used to output the point cloud feature F″ L The specific process is as Figure 3 shown, including:

[0079] First, the refined point cloud feature F L (H / 4, W / 4, 4C) is divided into 3 groups along the channel dimension, and 8-fold downsampling, 16-fold downsampling, and identity mapping (i.e., 4-fold downsampling) are respectively performed. The number of channels included in each group is:

[0080] C8 = C 16 = 4r′ s C

[0081] C4 = 4(1 - 2r′ s )C

[0082] where C n represents the number of channels of the n-fold downsampling branch; r′ s represents the channel division rate of the neck module, and the default value is 1 / 4, satisfying C4 = 2C and C8 = C 16 = C; then the divided features are represented as F4(H / 4, W / 4, 2C), F8(H / 4, W / 4, C), F 16 (H / 4, W / 4, C); where F n refers to the feature map of the n-fold downsampling branch;

[0083] After that, for F8 and F 16Perform downsampling operations respectively to obtain the corresponding 8x feature map F′8(H / 8, W / 8, C) and 16x feature map F′ 16 (H / 16, W / 16, C):

[0084] F′8 = Down(F8)

[0085] F′ 16 = Down(Down(F 16 ))

[0086] Wherein, Down(·) is a convolutional layer composed of a downsampling layer and 3 Split-Inception modules:

[0087] Down(·) = SI(SI(SI(dn(·))))

[0088] Wherein, SI(·) represents a Split-Inception module; dn(·) is a downsampling layer implemented by a residual block; the residual block is composed of two 3x3 convolutions and a shortcut connection; among them, the stride of the first 3x3 convolution is set to 2 to perform downsampling; the stride of the second 3x3 convolution is 1;

[0089] After that, the 8x feature map and the 16x feature map are each upsampled to a 4x feature map through transposed convolution to obtain F″8(H / 4, W / 4, C) and F″ 16 (H / 4, W / 4, C), and the 4x feature maps of the 3 branches are concatenated along the channel dimension to obtain the point cloud feature E′ L (G / 4, W / 4, 4C):

[0090] F′ L = Concat(F4, F″8, F″ 16 )

[0091] Finally, the point cloud feature F′ L is input into a multi-layer perceptron, and the output is added to F L and then passed through an activation function to obtain the final point cloud feature F″ L :

[0092] F″ L = σ[MLP(F′ L ) + F L .

[0093] It should be understood that the magnitudes of the sequence numbers of the steps in the embodiments of the present invention do not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0094] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-modal 3D object detection method based on multi-branch feature extraction, characterized in that: Specifically, it includes the following steps: Step 1: Extract image features through the camera branch, predict the pixel depth, and map it to the BEV space by the view transformer to form the image BEV feature F C ; Step 2: Process the point cloud data obtained by the lidar branch into voxels, and then process the voxels into a pseudo-image F(H, W, C), where H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map; and use a point cloud feature extraction network to expand the feature channels through staged downsampling to obtain refined point cloud features F L ; among them, each stage of the point cloud feature extraction network includes multiple Split-Inception convolution modules; Step 3: Feed the refined point cloud feature F L into a Split-Neck neck module for processing; the neck module processes the point cloud feature F L based on channel grouping, and performs downsampling operations at different multiples to obtain corresponding feature maps. Then, transposed convolution and upsampling are used to obtain a feature map in a unified form, and the feature maps are concatenated to obtain the point cloud feature F′ L , and finally, after being processed by a multi-layer perceptron and an activation function, the point cloud feature F″L is output; Step 4: Concatenate the image BEV feature F C obtained in Step 1 and the point cloud feature F″ L obtained in Step 3 along the channel dimension, and fuse them to obtain the fused feature F LC ; Step Five: Send the fused feature F LC to the detection head to obtain the final object detection result.

2. The method according to claim 1, wherein: The point cloud feature extraction network used in the second step specifically includes 3 stages, and each stage sequentially includes 3, 4, and 6 Split-Inception convolution modules; among them, a downsampling operation is performed once between adjacent two stages, and each time the feature map is downsampled by a factor of 2, and finally a 4x feature map that is downsampled by 4 times is obtained, and the refined point cloud feature F is output by the point cloud feature extraction network L (H / 4,W / 4,4C); The working process of the Split-Inception convolution module specifically includes: First, for the given input feature X(H, W, C), it is divided into 4 groups along the channel dimension: X s ,X l ,X p ,X i = Split(X) = X :,:g , X :g:2g , X :2g:3g , X :3g: where g is the number of channels of the convolution branch, and there is: where r s is the channel division rate, and the default setting is 1 / 8; After that, the divided features are sent to the parallel feature extraction branches: X′ p = MaxPool(X p ) X′ i = X i In the formula, represents a 2D convolution with a convolution kernel size of k×k, an input channel of C i , and an output channel of C o ; k s represents the size of the small convolution kernel, which is taken as 3; k l represents the size of the large convolution kernel, which is taken as 5; MaxPool represents a 3×3 max pooling that does not change the number of channels; After that, the outputs of each branch are concatenated along the channel dimension and batch normalization is performed: X′ = Concat(X′ s , X′ l , X′ p , X′ i ) Y = BatchNorm(X′) Finally, the features are sent to a multi-layer perceptron MLP for inter-channel feature interaction; the multi-layer perceptron contains two fully-connected layers and an activation function, and the activation function is located between the two fully-connected layers; The original input feature X is added to the output of the MLP, and then the final output feature Y′ is obtained through an activation function: Y′ = σ[MLP(Y) + X] In the formula, r is the expansion rate of the MLP, and σ is the activation function ReLU.

3. The method according to claim 2, wherein: In step three, the Split-Neck neck module is used to output the point cloud feature F″ L The specific process includes: First, divide the refined point cloud feature F L (H / 4, W / 4, 4C) into 3 groups along the channel dimension, and perform 8x downsampling, 16x downsampling, and identity mapping (i.e., 4x downsampling) respectively. The number of channels included in each group is as follows: Among them, C n represents the number of channels of the n-fold downsampling branch; r' s represents the channel division rate of the neck module, defaulting to 1 / 4, satisfying C4 = 2C and C8 = C 16 = C; then the divided feature representations are F4(H / 4, W / 4, 2C), F8(H / 4, W / 4, C), F 16 (H / 4, W / 4, C); where F n refers to the feature map input of the n-fold downsampling branch; After that, perform downsampling operations on F8 and F 16 respectively to obtain the corresponding 8x feature map F′8(H / 8, W / 8, C) and 16x feature map F′ 16 (H / 16, W / 16, C): F′8 = Down(F8) F′ 16 = Down(Down(F 16 )) In the formula, Down(·) is a convolution layer composed of a downsampling layer and 3 Split-Inception modules: Down(·) = SI(SI(SI(dn(·)))) In the formula, SI(·) represents a Split-Inception module; dn(·) is the downsampling layer, which is implemented by a residual block; the residual block consists of two 3x3 convolutions and a shortcut connection; among them, the stride of the first 3x3 convolution is set to 2 to perform downsampling; the stride of the second 3x3 convolution is 1; After that, the 8x feature map and the 16x feature map are each upsampled to the 4x feature map through transposed convolution, obtaining F″8(H / 4, W / 4, C) and F″ 16 (H / 4, W / 4, C), and the 4x feature maps of the three branches are concatenated along the channel dimension to obtain the point cloud feature F′ L (H / 4, W / 4, 4C): F′ L = Concat(F4, F″8, F″ 16 ) Finally, the point cloud feature F′ L is input into a multi-layer perceptron, and the output is added to F L and then passed through an activation function to obtain the final point cloud feature F″ L : F″ L = σ[MLP(F′ L ) + F L 。