A multimodal 3D detection method based on sparse proxy attention

By introducing sparse agent attention module and agent fusion method in three-dimensional object detection, the problems of sparse feature extraction and cross-modal data fusion are solved, and efficient and accurate three-dimensional object detection is achieved.

CN118351410BActive Publication Date: 2025-06-06BEIHANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410475463.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2025-06-06
Estimated Expiration
2044-04-19

AI Technical Summary

Technical Problem

The existing three-dimensional object detection methods are difficult to take into account both speed and accuracy in sparse feature extraction and cross-modal data fusion.

Method used

A multimodal three-dimensional detection method based on sparse proxy attention is proposed, and the proxy concept is used to achieve efficient homomodal feature extraction and end-to-end cross-modal data fusion. Through the sparse proxy attention module and proxy fusion method, the computation and memory pressure are reduced and the receptive field is increased.

Benefits of technology

Efficient sparse feature extraction and cross-modal data fusion are realized, which improves the performance and applicability of model detection and reduces inference delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118351410B_ABST
    Figure CN118351410B_ABST
Patent Text Reader

Abstract

The present invention proposes a multimodal three-dimensional detection method based on sparse proxy attention, which uses the concept of proxy to achieve efficient feature extraction under the same modality and end-to-end cross-modal data fusion. The present invention first proposes a multimodal universal sparse proxy attention module, which simplifies the complexity of attention calculation by using spatial proxy related priors, and achieves efficient parallel acceleration and related optimization of operators. Secondly, the present invention implements an efficient proxy fusion method under the same modality and cross-modality, and realizes a multimodal fusion model based on an agent. Compared with direct fusion based on voxels and images, the agent-based fusion reduces the pressure of calculation and memory, and increases the receptive field during fusion, thereby improving the performance of model detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of video processing and three-dimensional object detection, and in particular to a multimodal three-dimensional detection method based on sparse proxy attention. Background Art

[0002] Multimodal 3D object detection has recently become one of the hot research directions in computer vision and is widely used in fields such as autonomous driving and robotics. The purpose of multimodal 3D detection is to extract corresponding 3D features from multimodal inputs of point clouds and videos, and to detect and mark corresponding objects of interest, such as pedestrians or vehicles.

[0003] At present, the multimodal 3D object detection task mainly faces two difficulties: sparse feature extraction and cross-modal data fusion. First, unlike neat 2D data, 3D point clouds are sparsely and irregularly distributed in continuous space due to the characteristics of 3D sensors, which makes it challenging to directly apply the technology used for traditional conventional data. In order to solve this problem, many methods have proposed different solutions, but none of them can meet the requirements of high receptive field, no need for padding and good local feature extraction at the same time. Traditional sparse convolution benefits from special operator design, which does not require padding of voxels. At the same time, convolution also has good local feature extraction ability, but relatively speaking, the receptive field of convolution is very limited, which affects its feature extraction effect. SST is based on the grouping structure of sliding windows, with a medium-sized receptive field and good local grouping. However, although SST uses certain optimization techniques, it still requires certain padding and ensures parallelism. DSVT uses a grouping method based on coordinate axes, which has a medium-sized receptive field under the premise of zero padding. However, this special grouping method also leads to certain limitations in its local feature extraction ability.

[0004] Secondly, in terms of cross-modal data fusion, current methods are mainly divided into point-based, proposal-based and bird's-eye view (BEV) methods. Point-based and proposal-based methods use the semantic features of two-dimensional images to enrich lidar points and object proposals, respectively. BEV-based methods are currently a hot research topic. They unify the representations of cameras and lidars into a shared BEV space and fuse them through subsequent two-dimensional convolutions. At present, BEV-based methods are mainly divided into point cloud-image separation methods and end-to-end methods. Although the point cloud-image separation fusion scheme has superior performance, it usually requires a specific mode of encoder to process different sensor data in a sequential manner, which leads to increased inference latency and affects their applicability in the real world. End-to-end methods, such as UniTR, are limited by the receptive field of local grouping and cannot be integrated in larger areas.

[0005] In summary, existing 3D detection methods find it difficult to achieve 3D target detection that balances speed and accuracy through sparse feature extraction and cross-modal data fusion. Summary of the invention

[0006] In response to the above problems, the present invention proposes a multimodal three-dimensional detection method based on sparse proxy attention, which uses the concept of proxy to achieve efficient feature extraction under the same modality and end-to-end cross-modal data fusion. The present invention first proposes a multimodal universal sparse proxy attention module, which simplifies the complexity of attention calculation by using spatial proxy related priors, and achieves efficient parallel acceleration and related optimization of operators. Secondly, the present invention implements an efficient proxy fusion method under the same modality and cross-modality, and realizes a multimodal fusion model based on an agent. Compared with direct fusion based on voxels and images, the agent-based fusion reduces the pressure of calculation and memory, and increases the receptive field during fusion, thereby improving the performance of model detection. In order to achieve the above purposes, the present invention adopts the following technical solutions:

[0007] A multimodal three-dimensional detection method based on sparse proxy attention includes the following steps:

[0008] Step (1) divide the annotated 3D detection data set into a training set and a test set, and preprocess the training set and the test set; each data in the training set and the test set contains six images and a point cloud;

[0009] Step (2) using a sparse proxy attention network to extract multimodal three-dimensional features corresponding to the input data;

[0010] Step (3) inputs the multimodal 3D features into the multi-stage 3D target detection head SECOND, predicts the target category and position information of the features extracted by the sparse proxy attention network, and then decodes the position information to obtain the detection results.

[0011] Furthermore, in step (2), the sparse proxy attention network includes eight sparse proxy attention modules. The sparse proxy attention modules include a sparse proxy interaction module, a batch normalization module, and a feedforward network (FFN) module, wherein the parameters of each sparse proxy attention module are consistent, the number of parameter heads of the sparse attention interaction module is 8, and the number of channels is 256. The number of channels of the feedforward network module is 256, and the expansion factor is 4.

[0012] Furthermore, in step (2), the sparse proxy interaction module includes three modules: proxy feature extraction, same-modal and cross-modal proxy feature interaction, and proxy feature feedback. Among them, proxy feature extraction and proxy feature feedback use proxy map-reduce attention operators, whose parameter heads are 8 and channels are 256. Same-modal proxy feature interaction uses convolution for calculation, with a convolution kernel size of 13 and a channel number of 256. Cross-modal proxy feature interaction uses self-attention for calculation, with a parameter head number of 8 and a channel number of 256.

[0013] Furthermore, in step (2), the input data needs to be tokenized. For image data, the token size is 8x8, and the final token dimension is 192. For point cloud data, the token size is (0.3m, 0.3m, 8.0m), and the final dimension is also 192.

[0014] Furthermore, in step (3), the detection head SECOND includes two parts, a BEV feature extraction module and a 3D object detection module. The BEV feature extraction module consists of eight 3x3 convolutions, wherein the number of channels is 256, 256, 256, 256, 512, 512, 512, 512, respectively, the stride is 1, 1, 1, 1, 2, 1, 1, 2, and the padding is 1. The 3D object detection module includes two branches, a classification branch and a regression branch. The classification branch includes two convolutional layers for predicting the confidence of the detection box, and the dimension of its output tensor is the number of target categories. The regression branch includes two convolutional layers for predicting the relevant parameters of the bounding box.

[0015] The present invention also proposes a method for constructing a multimodal three-dimensional detection model based on sparse proxy attention, comprising the following steps:

[0016] Step (1) divide the annotated 3D detection data set into a training set and a test set, and preprocess the training set and the test set; each data in the training set and the test set contains six images and a point cloud;

[0017] Step (2) constructing a neural network model based on sparse proxy attention network;

[0018] Step (3) During the training process, the data in the training set is input into the neural network model to obtain the model loss;

[0019] Step (4) Train the entire network through the adaptive learning rate adjustment algorithm and the automatic derivation mechanism in the Pytorch framework to obtain the trained model parameters and save the network model;

[0020] Step (5) Call the network model to perform inference calculations on the actual test set data to obtain the corresponding confidence prediction results, center point offset, and bounding box parameters, and then obtain the final tracking that should be retained through parameter decoding and NMS to calculate the model accuracy;

[0021] Step (6) deploys the model on A100 and tests the model speed, using TensorRT as the deployment framework on A100.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] (1) This paper first proposes a general attention module that is mainly oriented towards sparse features, which is called sparse proxy attention. This attention is different from traditional grouped attention or global attention. It uses the concept of proxy to extract regular proxy abstract representations from sparse and uncertain inputs, and realizes large-scale information exchange to provide a large receptive field. Sparse proxy attention is mainly divided into three parts: proxy feature extraction, proxy information interaction and proxy information feedback. Both proxy feature extraction and proxy information feedback use the same module, namely sparse cross attention. Through the carefully optimized sparse cross attention operator, the model realizes feature extraction of a large receptive field and improves subsequent detection results.

[0024] (2) The present invention addresses the cross-modal fusion problem and achieves two goals with one stone by using agents: it reduces the pressure of computing and memory, and increases the receptive field during fusion. Since the overall agent association of the model is strongly spatially correlated, if the feature fusion of the agent is not used, the receptive field of the model itself will be greatly reduced, which will greatly affect the effect of model feature extraction. At the same time, the feature fusion of the agent is not limited to a single modality. In response to the multimodal fusion requirements in three-dimensional detection, the model of the present invention divides the fusion into same-modal agent feature fusion and cross-modal agent feature fusion, thereby achieving efficient feature fusion under different modalities and improving the cross-modal detection effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is an overall flow chart of a multimodal three-dimensional detection method based on sparse proxy attention of the present invention;

[0026] Figure 2 is a schematic diagram of a sparse agent attention network;

[0027] Figure 3 is a schematic diagram of the proxy map-reduce attention operator;

[0028] Figure 4 It is a schematic diagram of agent homomodal and cross-modal interactions;

[0029] Figure 5 is the test result obtained by the method of the present invention;

[0030] Figure 6 It is the comparison result of the method of the present invention and other commonly used methods. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0032] The present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0033] like Figure 1 As shown, a multimodal three-dimensional detection method based on sparse proxy attention of the present invention comprises the following steps:

[0034] Step (1) divide the annotated 3D detection data set into a training set and a test set, and preprocess the training set and the test set; each data in the training set and the test set contains six images and a point cloud;

[0035] Step (2) using a sparse proxy attention network model to extract multimodal three-dimensional features corresponding to the input data;

[0036] In step (2), the sparse proxy attention network includes eight sparse proxy attention modules. The sparse proxy attention modules include a sparse proxy interaction module, a batch normalization module, and a feedforward network (FFN) module, wherein the parameters of each sparse proxy attention module are consistent, the number of parameter heads of the sparse attention interaction module is 8, and the number of channels is 256. The number of channels of the feedforward network module is 256, and the expansion factor is 4.

[0037] The sparse agent interaction module in step (2) is as follows: Figure 2 As shown in Figure 2, the sparse agent interaction module includes three modules: agent feature extraction, agent feature interaction, and agent feature feedback. Among them, agent feature extraction and agent feature feedback use the map-reduce attention operator, as shown in Figure 2. Figure 3 As shown in the figure, the number of parameter heads is 8 and the number of channels is 256. The convolution used for interaction of the same modality proxy features is calculated, with a convolution kernel size of 13 and a number of channels of 256.

[0038] The proxy map-reduce calculation method in step (2) is as follows: Let a sparse attention weight matrix calculated between Q queries and K key / values ​​contain N positions that need to be paid attention to. The coordinates of these positions can be represented by a set A containing N binary pairs (src, dst). It describes the association between the query and the key. The calculation formula is:

[0039] ;

[0040] in, and Indicates the token number. Indicates The next query, Indicates The next keys. Represents the feature dimension of query and key. Represents matrix multiplication. Indicates exponential calculation with base e. Indicates The next query and The unnormalized similarity weight of each key.

[0041] Then the similarity weights are normalized on the key dimension to get the attention weights:

[0042] ;

[0043] in, , , Indicates the token number. Indicates The next query and The unnormalized similarity weight of the key, Indicates The next query and The unnormalized similarity weight of each key. Indicates The next query and The normalized similarity weight of each key.

[0044] Finally, the following formula is used to calculate the output features:

[0045] ;

[0046] Indicates The next query and The normalized similarity weight of each key. Shidi The next value. No. The next output features.

[0047] The agent feature interaction module in step (2) includes a cross-modal agent interaction module and a same-modal agent interaction module. Figure 4 As shown in Figure 2. The cross-modal proxy feature interaction is calculated using self-attention, with 8 parameter heads and 256 channels. There are two types of feature alignment in cross-modal interaction: For feature alignment from voxels to images, a feature map from voxel coordinates (x, y, z) to unique two-dimensional image coordinates ( ) where c is the camera number. This model directly uses the center point coordinates of the voxel to represent the voxel coordinates, builds a transformation matrix based on the internal and external parameters of each camera, and projects it onto the imaging plane of each camera. Most points will only fall within the imaging range of one camera. In particular, if the center of the voxel exists in multiple images, an image with a smaller number is selected as the final position to avoid repeated associations; if the center of the voxel does not exist in any image, it is considered to belong to the first image. For feature alignment from image to voxel, a two-dimensional image coordinate ( ) to voxel coordinates( ) is a mapping representation of the depth information of the visible light camera. Since the depth information of the visible light camera has been degraded, this constitutes an ill-posed problem. This model first generates the center position of the dense voxel consistent with the resolution of the current layer voxel, and then uses the scheme of projecting the voxel to the image to obtain the projection point of each voxel. Finally, each pixel will select the original position corresponding to the projection of the nearest voxel on the image as the final projected 3D position.

[0048] In step (2), the input data needs to be tokenized in advance. Start, where represents the original image, It means that there are 6 original images with an image size of 704x256 and three channels. These original images are divided into non-overlapping tokens using a module similar to ViT (Visual Transformer). Each token, called a "tag", represents a combination of the original pixel RGB values, with a size of 8x8, and a final dimension of 192. Then, these patch features are converted to dimension C through linear embedding to obtain image tags , where M is the number of markers. For a LiDAR point cloud , the present invention uses a dynamic voxel feature encoding marker. This marker uses a voxel grid (0.3m, 0.3m, 8.0m) for detection to create lidar voxels. The resulting lidar voxels are represented as , these taggers transform the multimodal input into , which includes N point cloud labels and M image labels. This combined representation is then further processed using a sparse proxy attention network.

[0049] Step (3) Input the multimodal 3D features into the multi-stage 3D target detection head SECOND, predict the target category and position information of the features extracted by the sparse proxy attention network, and then decode the position information to obtain the detection results;

[0050] In step (3), the detection head SECOND includes two parts, a BEV feature extraction module and a 3D object detection module. The BEV feature extraction module consists of eight 3x3 convolutions, where the number of channels is 256, 256, 256, 256, 512, 512, 512, 512, respectively, the stride is 1, 1, 1, 1, 2, 1, 1, 2, and the padding is 1. The 3D object detection module includes two branches, a classification branch and a regression branch. The classification branch includes two convolutional layers for predicting the confidence of the detection frame, and the dimension of its output tensor is the number of target categories. The regression branch includes two convolutional layers for predicting the relevant parameters of the bounding box. The result is the final detection result.

[0051] Example

[0052] The experimental environment is configured as follows: GPU (model gtx3090) is used as the computing platform, GPU parallel computing framework is adopted, Pytorch is selected as the convolutional network framework for training, and model rate verification is performed on gtx3090. The specific steps of the present invention include:

[0053] Step (1) divide the annotated 3D detection data set into a training set and a test set, and preprocess the training set and the test set; each data in the training set and the test set contains six images and a point cloud;

[0054] Step (2) Follow Figure 1 The network architecture diagram in Figure 2 Build a neural network model based on the module architecture diagram;

[0055] Step (3) During the training process, the data in the training set is input into the neural network structure to obtain the model loss;

[0056] Step (4) Train the entire network through the adaptive learning rate adjustment algorithm and the automatic derivation mechanism in the Pytorch framework to obtain the trained model parameters and save the network model;

[0057] Step (5) Call the network model to perform inference calculations on the actual test set data to obtain the corresponding confidence prediction results, center point offset, and bounding box parameters, and then obtain the final tracking that should be retained through parameter decoding and NMS to calculate the model accuracy;

[0058] (6) Deploy the model on A100 and test the model speed, using TensorRT as the deployment framework on A100.

[0059] In combination with the above steps, the present invention includes the following calculation method:

[0060] (1) The calculation method of loss is:

[0061] During the detection process, a 5-dimensional vector is used [ t , r , b , l , p ] To represent the bounding box of the object. All are vectors, representing the offset vector of the midpoint of the four boundaries compared to the center point of the detection box; P is the confidence prediction result;

[0062] Based on this, the loss function of the detection module contains the following parts:

[0063] (1) Classification loss :

[0064] ;

[0065] Among them, i represents the i-th pixel on the image, and is a hyperparameter used to control the weight ratio in the two cases, N is the number of foreground targets, is the predicted classification of the i-th pixel, is the true classification of the i-th pixel, is the classification loss.

[0066] (2) Position regression loss:

[0067] The position regression loss mainly includes the offset loss L of the center point of the bounding box 0 , the offset loss of the bounding box shape parameters L b , the offset loss of the bounding box deflection angle L a , their definitions are:

[0068] ;

[0069] ;

[0070] ;

[0071] ;

[0072] in, and are the true value and predicted value of the center point of the kth bounding box respectively; and are the true value and predicted value of the shape parameter of the kth bounding box respectively; and are the true value and predicted value of the deflection angle of the kth bounding box, N is the number of foreground objects, represents the loss function.

[0073] K is the label of the bounding box. Assuming there are N bounding boxes in total, the label of K starts from 1 and ends at N.

[0074] Compared with the prior art, the present invention uses the concept of proxy to achieve efficient feature extraction in the same modality and end-to-end cross-modal data fusion. Figure 5 As shown, compared with other methods Figure 6 shown.

[0075] It should be emphasized that the above are only preferred embodiments of the present invention and do not limit the present invention in any form. Any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A multimodal three-dimensional detection method based on sparse proxy attention, characterized in that: The method comprises the following steps: Step (1) divide the annotated 3D detection data set into a training set and a test set, and preprocess the training set and the test set; each data in the training set and the test set contains six images and a point cloud; Step (2) using a sparse proxy attention network to extract multimodal three-dimensional features corresponding to the input data; Step (3) Input the multimodal 3D features into the multi-stage 3D target detection head SECOND, predict the target category and position information of the features extracted by the sparse proxy attention network, and then decode the position information to obtain the detection results; In the step (2), the sparse proxy attention network includes eight sparse proxy attention modules, and the sparse proxy attention modules include a sparse proxy interaction module, a batch normalization module, and a feedforward network module, wherein the parameters of each sparse proxy attention module are consistent, the number of parameter heads of the sparse attention interaction module is 8, the number of channels is 256, the number of channels of the feedforward network module is 256, and the expansion factor is 4; In the step (2), the sparse proxy interaction module includes a proxy feature extraction module, a homomodal and cross-modal proxy feature interaction module and a proxy feature feedback module; wherein the proxy feature extraction module and the proxy feature feedback module use a proxy map-reduce attention operator, the number of parameter heads of which is 8 and the number of channels is 256; the homomodal and cross-modal proxy feature interaction module includes a homomodal proxy feature interaction module and a cross-modal proxy feature interaction module; the homomodal proxy feature interaction module uses convolution for calculation, the convolution kernel size is 13 and the number of channels is 256; the cross-modal proxy feature interaction module uses self-attention for calculation, the number of parameter heads is 8 and the number of channels is 256.

2. According to claim 1, a multimodal three-dimensional detection method based on sparse proxy attention is characterized in that: In step (2), the input data needs to be tokenized. For image data, the tokenization size is 8x8, and the final token dimension is 192. For point cloud data, the token size is (0.3m, 0.3m, 8.0m), and the final dimension is 192.

3. According to claim 1, a multimodal three-dimensional detection method based on sparse proxy attention is characterized in that: In the step (3), the detection head SECOND includes two parts, a BEV feature extraction module and a three-dimensional object detection module. The BEV feature extraction module is composed of eight 3x3 convolutions, wherein the number of channels is 256, 256, 256, 256, 512, 512, 512, 512, respectively, and the stride is 1, 1, 1, 1, 2, 1, 1, 2, and the padding is 1. The three-dimensional object detection module includes two branches, a classification branch and a regression branch. The classification branch includes two convolutional layers for predicting the confidence of the detection box, and the dimension of its output tensor is the number of target categories. The regression branch contains two convolutional layers to predict the relevant parameters of the bounding box.

4. A method for constructing a multimodal three-dimensional detection model based on sparse proxy attention, characterized in that: The following steps are involved: Step (1) dividing the annotated 3D detection data set into a training set and a test set, and preprocessing the training set and the test set; Each data set in the training set and the test set contains six images and a point cloud; Step (2) constructing a neural network model based on sparse proxy attention network; Step (3) During the training process, the data in the training set is input into the neural network model to obtain the model loss; Step (4) Train the entire network through the adaptive learning rate adjustment algorithm and the automatic derivation mechanism in the Pytorch framework to obtain the trained model parameters and save the network model; Step (5) Call the network model to perform inference calculations on the actual test set data to obtain the corresponding confidence prediction results, center point offset, and bounding box parameters, and then obtain the final tracking that should be retained through parameter decoding and NMS to calculate the model accuracy; Step (6) Deploy the model on A100 and test the model speed, using TensorRT as the deployment framework on A100; The sparse proxy attention network contains eight sparse proxy attention modules, which include a sparse proxy interaction module, a batch normalization module, and a feedforward network module. The parameters of each sparse proxy attention module are consistent. The number of parameter heads of the sparse attention interaction module is 8, the number of channels is 256, the number of channels of the feedforward network module is 256, and the expansion factor is 4. The sparse proxy interaction module includes a proxy feature extraction module, a homomodal and cross-modal proxy feature interaction module and a proxy feature feedback module; among them, the proxy feature extraction module and the proxy feature feedback module use the proxy map reduction map-reduce attention operator, the number of parameter heads is 8, and the number of channels is 256. The homomodal and cross-modal proxy feature interaction module includes a homomodal proxy feature interaction module and a cross-modal proxy feature interaction module. The homomodal proxy feature interaction module uses convolution for calculation, the convolution kernel size is 13, and the number of channels is 256. The cross-modal proxy feature interaction module uses self-attention for calculation, the number of parameter heads is 8, and the number of channels is 256.

5. The method according to claim 4, characterized in that The calculation method of loss loss is: During the detection process, a 5-dimensional vector is used To represent the bounding box of the object, They are all vectors, representing the offset vectors of the midpoints of the four boundaries compared to the center point of the detection box; is the confidence prediction result; based on this, the loss function contains the following parts: (1) Classification loss : Among them, i represents the i-th pixel on the image, and is a hyperparameter used to control the weight ratio in the two cases, N is the number of foreground targets, is the predicted classification of the i-th pixel, is the true classification of the i-th pixel, is the classification loss; (2) Position regression loss: The position regression loss mainly includes the offset loss L0 of the center point of the bounding box and the offset loss L of the bounding box shape parameter b , the offset loss of the bounding box deflection angle L a , their definitions are: in, and are the true value and predicted value of the center point of the kth bounding box respectively; and are the true value and predicted value of the shape parameter of the kth bounding box respectively; and are the true value and predicted value of the deflection angle of the i-th bounding box, N is the number of foreground objects, Represents the loss function; K is the label of the bounding box. Assuming there are N bounding boxes, the label of K starts from 1 and ends at N.

Citation Information

Patent Citations

  • End-to-end multi-mode multi-task automatic driving perception method and device based on long and short time sequence hybrid coding

    CN117710917A