3D Object Detection Method Based on Sphere Spatial Features and Multimodal Cross-Fusion Network
Through the combination of the spherical feature spatial network and the image semantic feature network, the multimodal data of point clouds and images is fused with the cross attention mechanism, solving the problems of large amount of computing and slow inference speed in the prior art, and achieving more accurate and fast 3D object detection.
Patent Information
- Application Number
- CN202210536412.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-05-12
AI Technical Summary
The existing 3D object detection methods are difficult to effectively fuse multimodal data of point clouds and images, especially when processing relationships between and internal modes, the calculation is large and the inference speed is slow.
The spherical feature spatial network and image semantic feature network are used to reduce the calculation amount through key point selection and sparse convolution to extract spatial features; at the same time, the multi-scale semantic features of the image are extracted using the UNet network, and spatial features and semantic features are fused through the cross attention mechanism.
It realizes more accurate and faster 3D object detection, and through effective feature fusion and computational optimization, the robustness and inference efficiency of the model are improved.
Smart Images

Figure CN114898356B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and more specifically, to a 3D object detection method based on spherical space features and a multi-modal cross-fusion network. Background Art
[0002] The statements in this section only relate to the background art related to the present invention and do not necessarily constitute the prior art.
[0003] 3D object detection technology plays an important role in the field of computer vision and is used to obtain the position and category information of objects in three-dimensional space, mainly including methods based on point clouds and multi-modal data. Among them, multi-modal data fusion has a very wide range of applications. For the same task, applying multi-modal data can fuse the complementary information of different data and make more robust predictions.
[0004] In recent years, the multi-modal fusion technology for point clouds has been booming and plays an important role in research tasks such as 3D shape classification, 3D object detection and tracking, and 3D point cloud segmentation. Point cloud data is mainly obtained by LiDAR devices and has depth and spatial geometric information, but it is sparse, disordered, and irregular. Images usually have rich semantic information and relatively less computational complexity, but they lack the most important depth information for 3D object detection. Therefore, how to use the fusion of point cloud and image features to achieve more accurate object detection is a major challenge. Existing methods either fuse between different modal data or within the same modal data. However, these methods cannot handle the relationships between different modalities and within modalities simultaneously.
[0005] Designing an effective model architecture to obtain more accurate and faster detection targets is a research hotspot in 3D object detection tasks. Convolutional neural networks have been proven to be very effective for two-dimensional image signal processing. However, for three-dimensional point cloud signals, the additional dimension z significantly increases the computational complexity. On the other hand, different from ordinary images, most voxels in three-dimensional point clouds are empty, which makes the point cloud data in three-dimensional voxels usually sparse signals. In addition, how to utilize the complementary information between multiple modal data is also a key issue for accurate 3D object detection. The cross-attention mechanism can achieve the fusion between cross-modal data and within-modal data. Compared with separate cross-modal data fusion or within-modal data fusion, cross-fusion is more conducive to information complementarity. Related methods have been applied in a small number of 3D point cloud object detections and achieved good performance. However, there are still great challenges in applying the attention mechanism to process the multi-modal fusion of point clouds. In addition, the original point cloud data volume is huge and there is relatively more irrelevant data, which will lead to too slow inference speed and difficult model training. Summary of the Invention
[0006] To alleviate the above problems, in this invention, we use a key point selection algorithm to select key points from the original point cloud and propose a Sphere Spatial Features Network to achieve accurate and fast spatial feature extraction tasks. This method consists of two parts, namely point cloud key point selection and spatial sparse convolution. Different from the traditional direct processing of the original point cloud data, the Sphere Spatial Features Network aims to reduce the computational amount by selecting key points and applying sparse convolution, and efficiently obtain the spatial feature group of the point cloud. Then we perform a convolution operation on the spatial feature group to obtain the final spatial geometric features. At the same time, we use the UNet network to extract features from the RGB image and propose an Image-semantic Features Network to achieve efficient semantic feature extraction tasks. This method obtains image semantic features of different scales through the UNet network and expands them through the MLP to obtain semantic features of the same size as the spatial features. Then, the cross-attention mechanism is used to realize the fusion of spatial features and semantic features. Finally, we train the entire network in an end-to-end manner and achieve better prediction performance.
[0007] The technical solution of the present invention provides a 3D object detection method based on spherical spatial features and a multi-modal cross-fusion network, and this method includes the following steps:
[0008] 1. Input the point cloud data, use the voxel centroid downsampling algorithm to select key points from the center of the original point cloud data, and then search for points within a fixed radius adjacent to the key points;
[0009] 1.1) Collect relevant data sets in the fields of image and point cloud 3D object detection, including the KITTI data set, ScanNet V2 data set, Waymo data set, SUN RGB-D data set, and Lyft L5 data set.
[0010] 1.2) In this invention, the KITTI data set with 80,256 object labels is used as the training data set to train the model; the test data set in the KITTI data set is used to detect the generalization performance of the model.
[0011] 1.3) First, use the voxel centroid downsampling algorithm to select key points from the original point cloud, as shown in the following formula:
[0012]
[0013] where m is the number of points in each voxel, and (x i , y i , z i ) are the coordinates of the points.
[0014] 1.4) Subsequently, we use the KD-tree nearest neighbor search algorithm to fix the search radius centered on the key points, and obtain several spherical spatial voxels V spatial , which contain several points P v,i , as shown in the following formula:
[0015] V spatial = KD(K centroid , R), (2)
[0016] where R represents the search radius.
[0017] 2. Use sparse convolution to extract features from the points in each selected spherical space, obtain a spatial group, and then perform a convolution operation on it to obtain spatial features;
[0018] 2.1) First, use sparse convolution (SPConv) to extract features from all the points in each obtained spherical voxel space, and obtain a spatial group. Among them, sparse convolution can be expressed as SC(m, n, f, s), where m represents the number of input feature channels, n represents the number of output feature channels, f represents the size of the filter, s represents the stride, and there is the following relational expression:
[0019]
[0020] 2.2) Subsequently, use the convolution operation to obtain spatial features, expressed as:
[0021] S f = Conv(SC(P v,i ))), (4)
[0022] where P v,i represents the points in each spherical spatial voxel.
[0023] 3. Use a three-layer Unet network to perform multi-scale feature extraction on the RGB image, obtain a semantic group, and expand it to the same dimension as the spatial features through MLP to obtain semantic features;
[0024] 3.1) First, use a 3-layer UNet network to perform multi-scale feature extraction on the RGB image. For the left downsampling module, each layer consists of two 3×3 convolutional layers and applies the ReLU activation function and a 2×2 max pooling layer; for the right upsampling module, each layer consists of an upsampled convolutional layer, a Concat feature concatenation, and two 3×3 convolutional layers and applies the ReLU activation function. Among them, bilinear interpolation is used for skip connections, as shown in the following formula:
[0025]
[0026] Among them, f(0, 0), f(0, 1), f(1, 0), and f(1, 1) represent the coordinates of four known points.
[0027] 3.2) Subsequently, use the MLP to expand the dimension of the above-obtained semantic group to obtain features with the same dimension as the spatial features, as shown in the following formula:
[0028] F s = MLP(f s ), (6)
[0029] Among them, f s represents the semantic group feature.
[0030] 4 Use the cross-attention mechanism to fuse the spatial features and semantic features. The fused features are then input into the detection head for detection, and this model is trained using the hybrid loss function.
[0031] 4.1) First, convert the spatial features into query Q s , convert the semantic features of the image into key K i and value V i , as shown in the following formula:
[0032]
[0033] Among them, S represents the spatial feature vector, I represents the semantic feature vector, W Q , W K , W V represent the query weight matrix, the key weight matrix, and the value weight matrix, respectively.
[0034] 4.2) Then use the attention mechanism for calculation, as shown in the following formula:
[0035]
[0036] Among them, K i T represents the transpose of K i .
[0037] Then input the above-obtained attention into a 1D convolutional layer and a max pooling layer to obtain the finally fused features, and finally input them into the detection head for classification and regression.
[0038] 4.3) The loss generated by classification uses the Focal Loss cross-entropy loss function.
[0039] FL(p t ) = -(1 - p t ) γ log(p t ), (9)
[0040] where 1 - p t is a variable balance factor, γ is a regulation factor and greater than zero.
[0041] 4.4) The loss generated by position regression adopts CIoU Loss.
[0042]
[0043] where α is a weight function, v is used to measure the consistency of the aspect ratio, ρ represents the distance between the center points of the predicted bounding box and the ground truth bounding box, and c represents the diagonal length of the smallest enclosing box of the two boxes.
[0044] 4.5) The bounding box optimization adopts the Smooth L1 loss function.
[0045]
[0046] where f(x i ) represents the predicted value, and y i represents the ground truth value.
[0047] Advantages of the present invention: The present invention effectively reduces the computational consumption of directly processing point cloud data through key point selection and sparse convolution strategy, and obtains rich spatial geometric features. At the same time, the multi-scale information of the RGB image is obtained through the UNet network, which is beneficial to the extraction of more accurate semantic features of the image. Then, the spatial features from the original point cloud and the semantic features of the RGB image are fused through the cross-attention mechanism to achieve the alignment and complementation of information between cross-modalities and within modalities, and strong feature representations are generated through the 1D convolutional layer and the max pooling layer to accurately locate the 3D bounding box and be used for fast and efficient detection of 3D scene targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Network flow framework diagram
[0049] Figure 2 Cross-attention network module DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. In addition, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] The flowchart framework of the present invention is as Figure 1 shown. The 3D object detection method of the present invention based on spherical spatial features and multi-modal cross-fusion network is specifically described as follows:
[0052] 1. Input the point cloud data, select key points from the center of the original point cloud data using the voxel centroid downsampling algorithm, and then search for points within a fixed radius of the neighboring points centered on the key points;
[0053] 1.1) Collect relevant datasets in the field of image and point cloud 3D object detection, including the KITTI dataset, ScanNet V2 dataset, Waymo dataset, SUN RGB-D dataset, and Lyft L5 dataset.
[0054] 1.2) In this invention, the KITTI dataset with 80,256 object labels is used as the training dataset to train the model; the test dataset in the KITTI dataset is used to detect the generalization performance of the model.
[0055] 1.3) First, select key points from the original point cloud using the voxel centroid downsampling algorithm, as shown in the following formula
[0056] where m is the number of points in each voxel, and (x i , y i , z i ) are the coordinates of the points.
[0057] 1.4) Subsequently, we use the KD-tree nearest neighbor search algorithm to fix the search radius centered on the key points to obtain several spherical spatial voxels V spatial , which contain several points P v,i , as shown in the following formula:
[0058] V spatial = KD(K centroid , R), (2)
[0059] where R represents the search radius.
[0060] 2. Use sparse convolution to extract features from the points in each selected spherical space to obtain a spatial group, and then perform a convolution operation on it to obtain spatial features;
[0061] 2.1) First, use sparse convolution (SPConv) to extract features from all the points in each obtained spherical voxel space to obtain a spatial group. Among them, sparse convolution can be expressed as SC(m, n, f, s), where m represents the number of input feature channels, n represents the number of output feature channels, f represents the size of the filter, s represents the stride, and there is the following relationship:
[0062]
[0063] 2.2) Subsequently, use the convolution operation to obtain spatial features, expressed as:
[0064] S f = Conv(SC(P v,i )), (4)
[0065] where P v,i represents the points in each spherical spatial voxel.
[0066] 3. Use a three - layer Unet network to perform multi - scale feature extraction on the RGB image to obtain semantic groups, and expand them to the same dimension as the spatial features through MLP to obtain semantic features;
[0067] 3.1) First, use a 3 - layer UNet network to perform multi - scale feature extraction on the RGB image. For the left - hand downsampling module, each layer consists of two 3×3 convolutional layers with the ReLU activation function applied and a 2×2 max - pooling layer; for the right - hand upsampling module, each layer consists of an upsampling convolutional layer, a Concat feature concatenation, and two 3×3 convolutional layers with the ReLU activation function applied. Among them, bilinear interpolation is used for skip connections, as shown in the following formula:
[0068]
[0069] where f(0, 0), f(0, 1), f(1, 0), f(1, 1) represent the coordinates of four known points.
[0070] 3.2) Subsequently, use MLP to expand the dimension of the above - obtained semantic groups to obtain features of the same dimension as the spatial features, as shown in the following formula:
[0071] F s = MLP(f s ), (6)
[0072] where f s represents the semantic group features.
[0073] 4 Use the cross - attention mechanism to fuse the spatial features and semantic features, as Figure 2 shown, and the fused features are then input into the detection head for detection, and this model is trained using a hybrid loss function.
[0074] 4.1) First, convert the spatial features into query Q s , convert the semantic features of the image into key K i and value V i , as shown in the following formula:
[0075]
[0076] where S represents the spatial feature vector, I represents the semantic feature vector, W Q , W K , WV respectively represent the query weight matrix, the key weight matrix, and the value weight matrix.
[0077] 4.2) Then, the attention mechanism is adopted for calculation, as shown in the following formula:
[0078]
[0079] where K i T represents the transpose of K i .
[0080] Then, the attention obtained above is input into a 1D convolutional layer and a max-pooling layer to obtain the finally fused features, and finally, it is input into the detection head for classification and regression.
[0081] 4.3) The loss generated by classification adopts the Focal Loss cross-entropy loss function.
[0082] FL(p t ) = -(1 - p t ) γ log(p t ), (9)
[0083] where 1 - p t is the variable balance factor, γ is the adjustment factor and is greater than zero.
[0084] 4.4) The loss generated by location regression adopts the CIoU Loss.
[0085]
[0086] where α is the weight function, v is used to measure the consistency of the aspect ratio, ρ represents the distance between the center points of the predicted box and the ground truth box, and c represents the diagonal length of the smallest bounding box of the two boxes.
[0087] 4.5) The bounding box optimization adopts the Smooth L1 loss function.
[0088]
[0089] where f(x i ) represents the predicted value, and y i represents the ground truth value.
[0090] The above is the preferred implementation of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A 3D object detection method based on spherical space features and a multi-modal cross-fusion network, characterized in that, The method includes the following steps: 1) Select key points from the original point cloud data center using the voxel centroid downsampling algorithm, and then search for points within a fixed radius around the key points; 1.1) Collect relevant datasets in the fields of image and point cloud 3D object detection, including the KITTI dataset, ScanNet V2 dataset, Waymo dataset, SUN RGB-D dataset, and Lyft L5 dataset; 1.2) Use the KITTI dataset with 80,256 object labels as the training dataset to train the model; use the test dataset in the KITTI dataset to detect the generalization performance of the model; 1.3) First, select key points from the original point cloud using the voxel centroid downsampling algorithm, as shown in the following formula: where m is the number of points in each voxel, and (x i , y i , z i ) are the coordinates of the points; 1.4) Subsequently, using the KD-tree nearest neighbor search algorithm, with the key points as the center and a fixed search radius, several spherical spatial voxels V are obtained spatial , which contain several points P v,i , as shown in the following formula: V spatial = KD(K centroid ,R), (2) where R represents the search radius; 2) Use sparse convolution to extract features from the points in each selected sphere, obtain a spatial group, and then use convolution operations to obtain spatial features; 3) Use a three-layer Unet network to perform multi-scale feature extraction on the RGB image, obtain a semantic group, and expand it to the same dimension as the spatial features through an MLP to obtain semantic features; 4) Use the cross-attention mechanism to fuse the spatial features and semantic features, input the fused features into the detection head for detection, and use a hybrid loss function to train this model.
2. The 3D object detection method based on spherical space features and multimodal cross-fusion network according to claim 1, characterized in that The specific steps of step 2) are as follows: 2.1) First, use sparse convolution (SPConv) to extract features from all points within each obtained spherical voxel space to obtain a spatial group, where sparse convolution can be expressed as SC(m,n,f,s), m represents the number of input feature channels, n represents the number of output feature channels, f represents the size of the filter, s represents the stride, and there is the following relationship: 2.2) Subsequently, use convolution operations to obtain spatial features, expressed as: S f = Conv(SC(P v,i ), (4) Among which P v,i represents the points in each spherical spatial voxel.
3. The 3D object detection method based on spherical space features and multi-modal cross-fusion network according to claim 1, characterized in that The specific steps of step 3) are as follows: 3.1) First, use a 3-layer UNet network to perform multi-scale feature extraction on the RGB image. For the left downsampling module, each layer consists of two 3×3 convolutional layers and applies the ReLU activation function and a 2×2 max pooling layer; for the right upsampling module, each layer consists of an upsampling convolutional layer, a Concat feature concatenation, and two 3×3 convolutional layers and applies the ReLU activation function, where bilinear interpolation is used for skip connections, as shown in the following formula: where f(0,0), f(0,1), f(1,0), f(1,1) represent the coordinates of four known points; 3.2) Subsequently, use an MLP to expand the dimension of the obtained semantic group to obtain features with the same dimension as the spatial features, as shown in the following formula: F s = MLP(f s ), (6) where f s represents the semantic group feature.
4. The 3D object detection method based on the spherical space feature and the multi-modal cross-fusion network according to claim 1, wherein The specific steps of step 4) are as follows: 4.1) First, convert the spatial features into a query Q s , convert the semantic features of the image into a key K i and a value V i , as shown in the following formula: where S represents the spatial feature vector, I represents the semantic feature vector, and W Q , W K , W V represent the query weight matrix, the key weight matrix, and the value weight matrix, respectively; 4.2) Then use the attention mechanism for calculation, as shown in the following formula: where K i T represents the transpose of K i , and then the obtained attention is input into a 1D convolutional layer and a max pooling layer to obtain the finally fused features, and finally it is input into the detection head for classification and regression; 4.3) The loss generated by classification uses the Focal Loss cross-entropy loss function, FL(p t ) = -(1 - p t ) γ log(p t ), (9) where 1 - p t is a variable balance factor, γ is a regulation factor and greater than zero; 4.4) The loss generated by position regression uses the CIoU Loss loss, where α is the weight function, v is used to measure the consistency of the aspect ratio, ρ represents the distance between the center points of the predicted box and the ground truth box, and c represents the diagonal length of the minimum bounding box of the two boxes; 4.5) The bounding box optimization uses the Smooth L1 loss function, where f(x i ) represents the predicted value, and y i represents the true value.
Citation Information
Patent Citations
Three-dimensional target detection method based on point cloud and image data fusion
CN114092780A
Method, system, and device for automatic calibration of differences in cross-modal target detection
WO2021000664A1