A high-efficiency three-dimensional object detection method based on voxel-point transformer
By combining the voxel-point transformer with sparse convolution and cross-attention mechanism, the problems of high computational resource consumption and accuracy loss in large point cloud detection are solved, efficient three-dimensional object detection is achieved, and detection accuracy and speed are improved.
Patent Information
- Application Number
- CN202310468800.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-27
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2043-04-27
AI Technical Summary
Existing transformer-based 3D object detection methods consume large computational resources and suffer severe accuracy loss when processing large point clouds. In particular, sampling in large outdoor point cloud scenes is time-consuming and quantization errors are introduced during the voxelization process.
A voxel-point transformer combined with sparse convolution and cross-attention mechanism is used to convert point clouds into sparse voxels and sample them within the reference point field, fusing voxel and point features to compensate for quantization errors and improve detection accuracy and efficiency.
It has achieved significant improvements in detection accuracy while ensuring operational efficiency, with the speed increased to 4.4 times, and achieved advanced detection results on the KITTI and Waymo datasets.
Smart Images

Figure CN116630955B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of three-dimensional vision, and particularly relates to a high-efficiency three-dimensional object detection method based on a voxel-point transformer. BACKGROUND
[0002] Three-dimensional object detection based on point clouds has become increasingly popular due to its wide applications, such as autonomous driving and virtual reality.
[0003] A method for object feature part detection based on three-dimensional laser radar point clouds is disclosed in Chinese patent document CN115546267A, which comprises: acquiring three-dimensional laser point cloud data of an object; mapping the three-dimensional laser point cloud to a binary image according to the distribution of laser points in the three-dimensional laser point cloud in horizontal and vertical angles; identifying key feature points of a feature part in the object from the binary image; and reversely mapping the key feature points back to the three-dimensional laser radar point cloud to obtain three-dimensional laser radar point cloud coordinates of the feature part.
[0004] Chinese patent document CN113870160A discloses a point cloud data processing method based on a transformer neural network, which comprises: constructing a three-dimensional object symmetry detection model, acquiring symmetric points of input point cloud data by detecting object symmetry surfaces / axes, converting the projection plane of the point cloud data into a rotation and translation operation of a symmetric structure to obtain a plurality of groups of data-enhanced point cloud graph data; extracting global feature information and local feature information of the plurality of groups of data-enhanced point cloud graph data through a transformer network model to obtain down-sampled point cloud data; and constructing a task-driven task network model in combination with different target task requirements, inputting the down-sampled point cloud data into the task network model to obtain a target task result.
[0005] Due to the invariance of the attention mechanism in the transformer to input arrangement, it has attracted widespread attention to use it to process unordered point clouds. However, due to the quadratic complexity of self-attention, a large amount of computing and memory resources are required when processing large point clouds. In order to overcome this problem, some point-based methods perform attention operations on the down-sampled point set, while some voxel-based methods use attention on local non-empty voxels. However, the former needs to use farthest point sampling to sample the point cloud, which is very time-consuming in large outdoor point cloud scenarios, and the latter inevitably introduces quantization errors in the voxelization process, thereby losing accurate position information. SUMMARY
[0006] The application provides a high-efficiency three-dimensional object detection method based on a voxel-point transformer, which utilizes the advantages of voxel and point cloud representation of the transformer to enable the model to achieve advanced detection accuracy while ensuring running efficiency.
[0007] An efficient three-dimensional object detection method based on voxel-point transformer, comprising:
[0008] (1) Given a laser radar point cloud, the laser radar point cloud is gridded, and then converted into discrete voxels;
[0009] (2) The discretized voxels are input into a three-dimensional backbone network to further capture the high-dimensional semantic features of each voxel;
[0010] (3) The voxels and their high-dimensional semantic features are input into the query initialization network to generate three-dimensional reference points and content queries;
[0011] (4) The laser radar point cloud, three-dimensional reference points, content queries, voxels and their high-dimensional semantic features are simultaneously input into the point-voxel transformer, and the point cloud and voxels within the domain radius of the reference point are sampled as point labels and voxel labels respectively;
[0012] (5) The point labels, voxel labels, three-dimensional reference points and content queries are taken as inputs, and the point-voxel transformer adaptively fuses the features of the point labels and voxel labels into the content queries according to the feature similarity between the point labels and voxel labels and the content queries;
[0013] (6) The content queries with fused features are input into the detection head network to further predict the object category and bounding box corresponding to each content query.
[0014] The present application combines the advantages of voxel and point representation, and overcomes their respective shortcomings. We convert large-scale point cloud into a small number of voxels through sparse convolution, and then sample from non-empty voxels to reduce the large running time caused by sampling. Then, inside the PVT-SSD, voxel features and point cloud features are adaptively fused to make up for the accuracy loss caused by quantization error. In this way, both the long-range context provided by voxels and the accurate position provided by points are preserved.
[0015] The specific process of step (1) is:
[0016] For a given laser radar point cloud, the coordinates of the point cloud in the gridded space are calculated, and they are assigned to the voxels they belong to; if multiple point clouds are contained in a voxel, a point is randomly sampled to represent the voxel.
[0017] In step (2), the three-dimensional backbone network is composed of a plurality of sparse convolutions and sub-manifold convolutions, wherein the sparse convolution can downsample the voxels, greatly reducing the number of voxels, while capturing voxel features containing high-dimensional semantic information.
[0018] The specific process of step (3) is:
[0019] (3-1) input the voxels and high-dimensional semantic features generated in step (2) into the query initialization network; wherein the query initialization network comprises branch 1 and branch 2;
[0020] (3-2) in branch 1, first, voxels located at the same horizontal position and different heights are merged using maximum pooling, then the pooled voxels are sampled, and finally each sampled voxel predicts a central position offset, and the coordinates are added to the offset to generate a three-dimensional reference point;
[0021] (3-3) in branch 2, first, voxels located at the same horizontal position and different heights are merged using concatenation, so that the three-dimensional voxels are converted into a two-dimensional feature map, then a plurality of two-dimensional convolutions are used for feature extraction, finally the three-dimensional reference point generated in branch 1 is projected onto the two-dimensional feature map, and bilinear interpolation is performed to obtain the content query.
[0022] The specific process of step (4) is as follows:
[0023] (4-1) input the three-dimensional reference point of step (3), the content query, the voxels and high-dimensional semantic features generated in step (2), and the laser radar point cloud of step (1) as input;
[0024] (4-2) randomly sample the voxels and high-dimensional semantic features generated in step (2) within the domain radius R1 of the three-dimensional reference point, and use the randomly sampled voxels and voxel features as voxel labels;
[0025] (4-3) randomly sample the laser radar point cloud of step (1) within the domain radius R2 of the three-dimensional reference point, after obtaining the randomly sampled points, obtain the features of each point by performing linear interpolation on the voxels of step (2), and use the sampled points and their corresponding interpolated point features as point labels.
[0026] When randomly sampling the laser radar point cloud of step (1) within the domain radius R2 of the three-dimensional reference point, a fast neighbor query method is used. First, project the point cloud to a distance map according to the following formula:
[0027]
[0028] Where θ is the tilt angle and φ is the azimuth angle.
[0029] Then, neighbor query and sampling are performed on the regular two-dimensional distance map, thereby avoiding the sampling process in irregular three-dimensional space.
[0030] In step (5), cross attention mechanism is used to fuse the features of the labels into the content query according to the feature similarity between the labels and the content query, and the formula is:
[0031]
[0032] Y = FFN(X) + X,
[0033] where Attention is multi-head cross attention, FFN is forward network, P s is the coordinate of the label, F s is the feature of the label, P query is the three-dimensional reference point, F suery is the content query.
[0034] In step (6), the detection head network comprises a category prediction network and a bounding box prediction network, which are respectively composed of a plurality of fully connected layers.
[0035] Compared with the prior art, the present application has the following beneficial effects:
[0036] 1. The present application introduces a three-dimensional object detector based on a voxel-point transformer, which can achieve better accuracy and up to 4.4 times faster running speed compared with previous transformer-based methods.
[0037] 2. Extensive experiments are conducted to verify the effectiveness of the proposed model, and advanced detection accuracy is achieved on KITTI, Waymo and nuScenes datasets. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 is a flowchart of the present application, a high-efficiency three-dimensional object detection method based on voxel-point transformer;
[0039] Figure 2 is a flowchart of the present application, a fast neighbor query method;
[0040] Figure 3 is a schematic diagram of attention weight visualization of the present application;
[0041] Figure 4 is a visualization result diagram of the present application on the task of point cloud three-dimensional object detection. DETAILED DESCRIPTION
[0042] The present application will be further described in detail below in combination with the drawings and examples, it should be pointed out that the following examples are intended to facilitate the understanding of the present application, and do not limit it in any way.
[0043] As Figure 1 shown, a high-efficiency three-dimensional object detection method based on voxel-point transformer, comprising the following steps:
[0044] S01, for a given lidar point cloud, the coordinates of the point cloud in the grid space are calculated, and they are assigned to the voxels to which they belong. Since each voxel may contain multiple point clouds, we randomly sample one of them to represent the voxel.
[0045] S02, the voxel is input into a three-dimensional backbone network, which is composed of several sparse convolutions and submanifold convolutions. Since sparse convolution can downsample voxels, the number of voxels is greatly reduced, while capturing the voxel features containing high-dimensional semantic information.
[0046] S03, the generated voxels and voxel features are input into the query initialization module, and three-dimensional reference points and content queries are generated by branch 1 and branch 2 respectively, as shown in the initialization query of Figure 1
[0047] Specifically: in branch 1, voxels at the same horizontal position and different heights are merged using max pooling, and then the pooled voxels are sampled. Finally, each sampled voxel predicts the center position offset, and the coordinates are added to the offset to generate three-dimensional reference points; in branch 2, voxels at the same horizontal position and different heights are merged using concatenation, so that three-dimensional voxels are converted into two-dimensional feature maps, and then several two-dimensional convolutions are used for feature extraction. Finally, the three-dimensional reference points generated by branch 1 are projected onto the two-dimensional feature map, and bilinear interpolation is performed to obtain the content query.
[0048] S04, the three-dimensional reference points and content queries of S03, the voxels and their high-dimensional semantic features of S02, and the lidar point cloud of S01 are input. Then, voxels and their high-dimensional semantic features are randomly sampled within the radius R1 of the three-dimensional reference point, and the randomly sampled voxels and voxel features are used as voxel labels; the lidar point cloud is randomly sampled within the radius R2 of the three-dimensional reference point, and after obtaining the randomly sampled points, the features of each point are obtained by performing linear interpolation on the voxels. In order to speed up the sampling of the lidar point cloud, the method of Figure 2 is used to speed up the neighbor query process. First, project the point cloud to the distance map according to the following formula:
[0049]
[0050] where θ is the tilt angle and φ is the azimuth angle. Then, neighbor queries and sampling can be performed on the regular two-dimensional distance map, thereby avoiding the sampling process in irregular three-dimensional space.
[0051] S05, the above sampled point labels and voxel labels, as well as three-dimensional reference points and content queries are input, and the cross-attention mechanism is used to fuse the features of the labels into the content query according to the feature similarity of the labels and the content query. The formula of this step is:
[0052]
[0053] Y = FFN(X) + X,
[0054] where Attention is multi-head cross attention, FFN is forward network, P s is the coordinate of the marker, F s is the feature of the marker, P query is the three-dimensional reference point, F query is the content query.
[0055] S06, finally, the content query after the fusion of the marker features is input to the existing detection head network, and the corresponding object class and bounding box are generated for each content query.
[0056] Table 1 is a comparison with the previous method on the Waymo dataset Scalability in perception for autonomous driving: Waymo open dataset, and it can be seen that the method of the present application is much better than the previous method.
[0057] Table 1
[0058]
[0059]
[0060] Table 2 is a comparison of the present application with other transformer-based methods on the KITTI dataset Are we ready for autonomous driving? the KITTI vision benchmark suite, and it can be seen that the method of the present application can also be better than the previous method.
[0061] Table 2
[0062] Method Simple Medium Difficult M3DETR 90.28 81.73 76.96 CT3D 87.83 81.77 77.16 PDV 90.43 81.86 77.36 VoTr-TSD 89.90 82.09 79.14 PVT-SSD (Ours) 90.65 82.29 76.85
[0063] Figure 3 The visualization results of the attention weights of the present application are shown, and it can be seen that the present application can simultaneously focus on features at close range and at long range.
[0064] Figure 4 The visualization results of the present application on the point cloud 3D detection task are shown, and it can be seen that the prediction box generated by the present application can better detect objects.
[0065] The above-described embodiments have described the technical solutions and beneficial effects of the present application in detail, and it should be understood that the above-described is only a specific embodiment of the present application and is not used to limit the present application, and any modification, supplement and equivalent replacement made within the principle range of the present application should be included in the protection range of the present application.
Claims
1. An efficient 3D object detection method based on voxel-to-point transformer, characterized in that: include: (1) Given a LiDAR point cloud, grid the LiDAR point cloud and convert it into discrete voxels; (2) The discretized voxels are input into the 3D backbone network to further capture the high-dimensional semantic features of each voxel; (3) Input the voxels and their high-dimensional semantic features into the query initialization network to generate three-dimensional reference points and content queries; the specific process is as follows: (3-1) Inputting the voxels and high-dimensional semantic features generated in step (2) into the query initialization network; wherein the query initialization network includes branch 1 and branch 2; (3-2) In branch 1, voxels at the same horizontal position but different heights are first merged using maximum pooling, and then the pooled voxels are sampled. Finally, the center position offset of each sampled voxel is predicted, and the coordinates and the offset are added to generate a 3D reference point. (3-3) In branch 2, voxels at the same horizontal position but different heights are first merged using splicing to convert the 3D voxels into a 2D feature map. Several 2D convolutions are then used for feature extraction. Finally, the 3D reference points generated in branch 1 are projected onto the 2D feature map and bilinear interpolation is performed to obtain the content query. (4) The lidar point cloud, 3D reference points, content query, voxels and their high-dimensional semantic features are simultaneously input into the point-to-voxel converter, and the point cloud and voxels are sampled within the range radius of the reference point as point labels and voxel labels, respectively. The specific process is as follows: (4-1) taking the three-dimensional reference points and content query of step (3), the voxels and their high-dimensional semantic features of step (2), and the lidar point cloud of step (1) as input; (4-2) Randomly sampling voxels and their high-dimensional semantic features generated by step (2) within the area radius R1 of the three-dimensional reference point, and using the randomly sampled voxels and voxel features as voxel labels; (4-3) Randomly sampling the lidar point cloud of step (1) within the range radius R2 of the three-dimensional reference point, obtaining the randomly sampled points, and then performing linear interpolation on the voxels of step (2) to obtain the features of each point, and using the sampled points and their corresponding interpolated point features as point markers; (5) Taking point labels, voxel labels, 3D reference points and content queries as input, the point-to-voxel transformer adaptively integrates the features of point labels and voxel labels into the content query based on the feature similarity between the point labels and voxel labels and the content query; (6) The content query after feature fusion is input into the detection head network to further predict the object category and bounding box corresponding to each content query.
2. The efficient three-dimensional object detection method based on voxel-to-point converter according to claim 1, characterized in that: The specific process of step (1) is: For a given lidar point cloud, the coordinates of the point cloud in the gridded space are calculated and assigned to the voxel to which it belongs; if a voxel contains multiple point clouds, one of the points is randomly sampled to represent the voxel.
3. The efficient three-dimensional object detection method based on voxel-to-point converter according to claim 1, characterized in that: In step (2), the three-dimensional backbone network is composed of a number of sparse convolutions and submanifold convolutions, wherein the sparse convolution downsamples the voxels, thereby significantly reducing the number of voxels while capturing voxel features containing high-dimensional semantic information.
4. The efficient three-dimensional object detection method based on voxel-to-point converter according to claim 1, characterized in that: In step (4-3), when randomly sampling the lidar point cloud of step (1) within the range radius R2 of the three-dimensional reference point, a fast neighbor query method is used. First, the point cloud is projected onto the distance map according to the following formula: Where θ is the tilt angle and φ is the azimuth angle; Afterwards, neighbor query and sampling are performed on a regular two-dimensional distance graph, thus avoiding the sampling process in an irregular three-dimensional space.
5. The efficient three-dimensional object detection method based on voxel-to-point converter according to claim 1, characterized in that: In step (5), the cross-attention mechanism is used to integrate the features of the tag into the content query according to the feature similarity between the tag and the content query. The formula is: Y=FFN(X)+X, Among them, Attention is multi-head cross attention, FFN is forward network, P s is the coordinate of the marker, F s is the marked feature, P query is the three-dimensional reference point, F query Query for content.
6. The efficient three-dimensional object detection method based on voxel-to-point converter according to claim 1, characterized in that: In step (6), the detection head network includes a category prediction network and a bounding box prediction network, each of which is composed of several fully connected layers.
Citation Information
Patent Citations
Point cloud data processing method based on converter neural network
CN113870160A
Method for detecting feature part of object based on three-dimensional laser radar point cloud
CN115546267A