Point Cloud 3D Detection Method and Model Based on Adaptive and Multi-Level Feature Dimensionality Reduction

Through adaptive and multi-level feature dimensionality reduction methods, the problem of losing 3D geometric information in the point cloud 3D detection algorithm is solved, the detection accuracy and recall rate are improved, the environmental perception ability of autonomous driving cars is enhanced, and efficient point cloud 3D object detection is achieved.

CN115760983BActive Publication Date: 2025-07-04TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211452093.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-07-04
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

The existing point cloud 3D detection algorithm loses a large amount of 3D geometric information during feature dimensionality reduction, resulting in low detection accuracy and recall, and cannot effectively support the environmental perception of autonomous vehicles.

Method used

Adaptive and multi-level feature dimensionality reduction methods are adopted, and adaptive dynamic feature dimensionality reduction and multi-scale BEV feature extraction of sparse voxel features are combined with multi-level feature fusion to preserve the 3D geometric information of the object and perform multi-scale detection.

Benefits of technology

It improves the accuracy and recall of point cloud 3D object detection, enhances the perception of the environment of autonomous vehicles, has strong real-time performance, and is better than existing algorithms. It currently ranks second on the large-scale autonomous driving dataset nuScenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115760983B_ABST
    Figure CN115760983B_ABST
Patent Text Reader

Abstract

The present invention discloses a point cloud 3D detection method and model based on self-adaptation and multi-level feature dimensionality reduction. The detection method includes: acquiring a real-time point cloud data frame and performing point cloud voxelization to obtain an initial voxel feature and its voxel coordinates; sparsifying the initial voxel feature to obtain a sparse voxel feature; performing self-adaptive dynamic feature dimensionality reduction on the sparse voxel feature to obtain an initial BEV feature, and performing multi-scale BEV feature extraction on the initial BEV feature to obtain a multi-scale BEV feature containing semantic features; performing multi-scale voxel feature extraction on the sparse voxel feature to obtain a multi-scale voxel feature containing geometric features, reducing the dimensionality of the voxel feature at each scale, and then fusing it with the BEV feature at the corresponding scale; using the BEV feature fused in the last layer for 3D detection to estimate the position of the target object in the BEV space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving object perception, and particularly to a point cloud 3D detection method and model based on adaptive and multi-level feature dimensionality reduction. Background Art

[0002] An autonomous driving vehicle is a complex unmanned system that relies on in-vehicle sensors to perceive the environment and make decision control. To achieve the decision and control of autonomous driving, it is necessary to first use sensors (usually including lidar and cameras) to perceive the surrounding environment, process the sensor data to obtain the 3D semantic information of the objects in the surrounding environment, and then make decisions and controls based on this information.

[0003] 3D object detection based on lidar is a key technology to solve the problem of environmental perception in autonomous driving. It uses the real-time acquired point cloud data frames and decodes through a neural network to obtain the 3D semantic information of the objects. To achieve high efficiency, currently, grid-based point cloud 3D object detection algorithms are commonly used in the field of autonomous driving, which can be subdivided into pillar-based and voxel-based. Both of these methods need to reduce the dimensionality of the 3D sparse-structured point cloud through dimensionality reduction operations to obtain a 2D form of BEV (bird's eye view) feature map, and then perform 3D object detection on the 2D BEV feature map. However, the current algorithms use fixed convolution kernels or pooling operations for 3D to 2D feature dimensionality reduction, without considering that these feature dimensionality reduction operations will lose a lot of 3D geometric information. To improve the precision and recall rate of the point cloud 3D object detection algorithm, it is urgent to improve the existing feature dimensionality reduction operations and the structure of the feature extraction backbone network to retain more 3D geometric information. Summary of the Invention

[0004] To solve the problem that the BEV features in the existing point cloud 3D detection technology seriously lose the object geometric information, the present invention proposes a point cloud 3D detection method and model based on adaptive and multi-level feature dimensionality reduction, so that the BEV features can adaptively retain the 3D geometric information and multi-scale geometric information of the objects during the feature extraction process.

[0005] To solve the above problems, one aspect of the present invention proposes the following technical solutions:

[0006] A 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction, comprising: acquiring a real-time point cloud data frame and performing point cloud voxelization to obtain initial voxel features and their voxel coordinates; sparsifying the initial voxel features to obtain sparse voxel features; performing adaptive dynamic feature dimensionality reduction on the sparse voxel features to obtain initial BEV features, and performing multi-scale BEV feature extraction on the initial BEV features to obtain multi-scale BEV features containing semantic features; performing multi-scale voxel feature extraction on the sparse voxel features to obtain multi-scale voxel features containing geometric features, reducing the dimensionality of the voxel features at each scale, and then fusing them with the BEV features at the corresponding scale; using the fused BEV features of the last layer for 3D detection to estimate the position of the target object in the BEV space.

[0007] Further, the step of performing adaptive dynamic feature dimensionality reduction on the sparse voxel features includes: estimating the feature space distribution along the height dimension of the sparse voxel features, and this estimated feature space distribution is used to characterize the importance weights of object features along the height dimension; re-weighting the object features along the height dimension according to the estimated feature space distribution to obtain the initial BEV features.

[0008] Further, the step of estimating the feature space distribution along the height dimension of the sparse voxel features includes: estimating the voxel feature importance weights through a 3D submanifold sparse convolution with a convolution kernel size of 3×3×3, and its number of channels is 1; normalizing the voxel feature importance along the height dimension to obtain the estimated feature space distribution.

[0009] Further, the step of re-weighting the object features along the height dimension according to the estimated feature space distribution includes: dynamically weighting the sparse voxel features along the height dimension using the estimated feature space distribution to obtain the initial BEV features.

[0010] Further, the step of performing multi-scale voxel feature extraction on the sparse voxel features includes: constructing a voxel feature extraction branch for extracting geometric features, using the sparse voxel features as the initial input of this voxel feature extraction branch, and performing multi-stage voxel feature extraction to correspondingly obtain the multi-scale voxel features; wherein, the output result of each stage of voxel feature extraction is used as the input of the next stage of voxel feature extraction.

[0011] Further, the step of reducing the dimensionality of the voxel features at each scale includes: reducing the dimensionality of the output result of each stage of voxel feature extraction to obtain the reduced-dimensional voxel features of the corresponding stage.

[0012] Further, the steps of performing multi-scale BEV feature extraction on the initial BEV features include: constructing a BEV feature extraction branch for extracting semantic features, using the initial BEV features as the initial input of this BEV feature extraction branch, and performing multi-stage BEV feature extraction; fusing the output result of each stage of BEV feature extraction with the reduced-dimensional voxel features of the corresponding stage as the input for the next stage of BEV feature extraction.

[0013] Further, the steps of using the BEV features fused in the last layer for 3D detection and estimating the position of the target object in the BEV space include: sending the BEV features fused in the last layer through a multi-scale Neck network into a Center Head detection head to estimate the position of the object in the BEV space, and simultaneously regressing the center offset residuals of the 3D bounding box of the object, the three-dimensional size of the bounding box, the height of the center point of the bounding box, the heading angle of the object, and the two-axis velocity of the object.

[0014] Further, the steps of sparsifying the initial voxel features include: processing the initial voxel features through a 3D sparse convolution with a convolution kernel size of 5×5×1 to increase the receptive field and obtain the sparse voxel features.

[0015] On the other hand, the present invention proposes a point cloud 3D detection model based on adaptive and multi-level feature dimensionality reduction, including: a dynamic voxelization module for performing point cloud voxelization on a real-time acquired point cloud data frame to obtain initial voxel features and their voxel coordinates; a first 3D sparse convolution connected to the dynamic voxelization module for sparsifying the initial voxel features to obtain sparse voxel features; an adaptive dynamic feature dimensionality reduction module connected to the first 3D sparse convolution for performing adaptive dynamic feature dimensionality reduction on the sparse voxel features to obtain initial BEV features; a voxel feature extraction branch connected to the first 3D sparse convolution for performing multi-scale voxel feature extraction on the sparse voxel features to obtain multi-scale voxel features containing geometric features; a multi-level feature dimensionality reduction module connected to the voxel feature extraction branch for reducing the dimensionality of the voxel features at each scale; a BEV feature extraction branch connected to the adaptive dynamic feature dimensionality reduction module and the multi-level feature dimensionality reduction module for performing multi-scale BEV feature extraction on the initial BEV features to obtain multi-scale BEV features containing semantic features; a 3D detection module connected to the BEV feature extraction branch for using the BEV features fused in the last layer of the BEV feature extraction branch to perform 3D detection and estimate the position of the target object in the BEV space.

[0016] Compared with existing algorithms such as CenterPoint, PV-RCNN, Focals, Voxel-RCNN, SST, PillarNet, and PointPillar, the 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction proposed by the present invention can dynamically focus on regions of an object that are easy to detect, retain richer 3D geometric information, and efficiently utilize multi-scale geometric information through a multi-level feature dimensionality reduction strategy, reducing the loss of 3D information in the feature extraction process and the feature dimensionality reduction process, greatly improving the accuracy and recall rate of the 3D object detection algorithm, and enhancing the perception ability of an autonomous driving vehicle for the surrounding environment. Moreover, the method of the present invention has strong real-time performance, and its speed and accuracy are superior to many algorithms. Currently, it ranks second in the pure point cloud leaderboard on the large-scale autonomous driving scenario dataset nuScenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flowchart of the 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction proposed in an embodiment of the present invention.

[0018] Figure 2 is an operation flowchart of adaptive dynamic feature dimensionality reduction in the 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction proposed in an embodiment of the present invention.

[0019] Figure 3 is an operation flowchart of multi-level feature dimensionality reduction in the 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction proposed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0021] An embodiment of the present invention proposes a 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction. Please refer to Figure 1 , and the method includes the following algorithm steps A1 to A6:

[0022] A1. Real-time obtain a point cloud data frame. The point cloud data frame is obtained, for example, by a lidar.

[0023] A2. Based on the dynamic point cloud voxelization of a multi-layer perceptron, obtain an initial voxel feature and its voxel coordinates. Among them, the voxel coordinates are used as voxel feature indexes when processing voxel features in subsequent steps.

[0024] A3. Enlarge the receptive field of the initial voxel feature through a 3D sparse convolution with a convolution kernel size of 5×5×1 to obtain a sparse voxel feature, denoted as F1 v . For the sparse voxel feature F1 v, it will be processed separately through two branches. One processing branch is step A4, and the other processing branch is step A6.

[0025] A4. For the sparse voxel feature F1 v perform multi-scale voxel feature extraction to obtain multi-scale voxel features including geometric features. Specifically, construct a voxel feature extraction branch to extract geometric features. In some embodiments, the voxel feature extraction branch extracts multi-scale voxel features through a multi-stage extraction method, and each stage extracts voxel features of one scale. For example Figure 1 as shown in, a voxel feature extraction branch including three-stage extraction is constructed to extract voxel features of three scales. Among them, the first-stage voxel feature extraction uses the sparse voxel feature F1 v as the input, and the extracted voxel feature is denoted as And the voxel feature extracted in the first stage is used as the input for the second-stage voxel feature extraction, and the voxel feature extracted in the second stage is denoted as And so on, the voxel feature extracted in the second stage is used as the input for the third-stage voxel feature extraction, and the voxel feature extracted in the third stage is denoted as In this example, voxel features on three scales are obtained The three have different voxel ranges. The voxel range of the voxel feature is denoted as X2, Y2, Z2; the voxel range of the voxel feature is denoted as X3, Y3, Z3; the voxel range of the voxel feature is denoted as X4, Y4, Z4. Among them, (X2, Y2, Z2), (X3, Y3, Z3), (X4, Y4, Z4) represent the maximum values of the voxel coordinates in 3 dimensions.

[0026] In some embodiments, the voxel feature extraction of each stage within the voxel feature extraction branch can be implemented using the same feature extraction unit. For example, each stage uses a 3D submanifold sparse convolution with a kernel size of 3×3×3 and a 3D sparse convolution with a kernel size of 3×3×3 for downsampling.

[0027] A5. Use a multi-level feature dimensionality reduction strategy to reduce the dimensionality of the voxel features at each scale, so that the 3D voxel features are reduced to 2D BEV features at the corresponding scale. Please refer to Figure 3 , and the voxel features at each scale are subjected to dimensionality reduction operations using 3D sparse convolutions with corresponding kernel sizes. Still taking the previous example, for the voxel features on the three scales extracted by the voxel feature extraction branch 3D sparse convolutions with kernel sizes of 1×1×Z2, 1×1×Z3, and 1×1×Z4 are respectively used for dimensionality reduction. Among them, the values of Z2, Z3, and Z4 are usually determined according to the lidar used to obtain the point cloud data frame. For example, for the lidar used in the nuScenes dataset, the values of Z2, Z3, and Z4 here are 21, 11, and 5 respectively.

[0028] A6. For the sparse voxel feature F1 v Perform adaptive dynamic feature dimensionality reduction to obtain the initial BEV (BEV refers to the bird's-eye view, i.e., Bird’s Eye View) feature, denoted as F1 bev . Then, within this processing branch, construct a BEV feature extraction branch to extract semantic features. In some embodiments, the BEV feature extraction branch extracts multi-scale BEV features through a multi-stage extraction method, and each stage represents a different scale. It should be understood that the number of stages of the BEV feature extraction branch should be equal to the number of stages of the voxel feature extraction branch. Please refer to Figure 1 , still taking the aforementioned three-stage as an example, the input of the BEV feature extraction in the first stage of the BEV feature extraction branch is the initial BEV feature F1 bev , and the BEV feature extracted in the first stage is fused with the voxel feature After dimensionality reduction to obtain the feature, which is called the first-stage fused BEV feature, denoted as Then use As the input of the BEV feature extraction in the second stage. Similarly, the BEV feature extracted in the second stage is fused with the voxel feature After dimensionality reduction to obtain the feature, which is called the second-stage fused BEV feature, denoted as And so on, the final output BEV feature after the BEV feature extraction in the third stage In this example, the BEV feature Is used as the last-layer fused feature for subsequent 3D point cloud detection. It should be understood that the three-stage is only exemplary and should not be construed as a limitation of the present invention. In other embodiments, the voxel feature extraction branch and the BEV feature extraction branch may not exclude the use of two-stage, four-stage, five-stage or more stages.

[0029] A7. Send the last-layer BEV feature Into the Center Head detection head through the multi-scale Neck network to estimate the position of the object in the BEV space, and at the same time regress the center offset residual of the object's 3D bounding box, the three-dimensional size of the bounding box, the height of the bounding box center point, the object's heading angle, and the object's two-axis speed.

[0030] In some embodiments, step A6 performs the operation on the sparse voxel feature F1 vThe steps for adaptive dynamic feature dimensionality reduction include: for the sparse voxel feature F1 v Estimate the feature space distribution along the height dimension, and this estimated feature space distribution is used to characterize the importance weight of the object features along the height dimension; then re-weight the object features along the height dimension according to the estimated feature space distribution to obtain the initial BEV feature F1 bev . Specifically, please refer to Figure 2 For the sparse voxel feature F1 v First, estimate the voxel feature importance weight F through a 3D submanifold sparse convolution with a convolution kernel size of 3×3×3 i v , and its number of channels is 1. The importance weight is used to characterize the importance of the feature. The larger the weight, the more important the feature; then perform SoftMax normalization on F i v along the height dimension (e.g., the Z-axis) to obtain the estimated feature space distribution Use the feature space distribution W to perform dynamic weighting on the sparse voxel feature F1 v along the Z-axis to obtain the initial BEV feature where, w i,j,k represents the weight of each feature, and i, j, k respectively represent the three-dimensional coordinates of the voxel feature in the voxel space. i is the X-axis, j is the Y-axis, and k is the Z-axis; Z represents along the height dimension (Z-axis); represents the BEV feature at the (i, j) position in the BEV space, and Z i,j represents all Z coordinates with X and Y axis positions being i and j respectively, represents the voxel feature with coordinates i, j, k in the voxel space.

[0031] Another embodiment of the present invention provides a model adapted to the foregoing point cloud 3D detection method based on adaptive and multi-level feature dimensionality reduction, that is, a point cloud 3D detection model based on adaptive and multi-level feature dimensionality reduction. The model includes: a dynamic voxelization module for performing point cloud voxelization on a real-time acquired point cloud data frame to obtain initial voxel features and their voxel coordinates; a first 3D sparse convolution connected to the dynamic voxelization module for sparsifying the initial voxel features to obtain sparse voxel features; an adaptive dynamic feature dimensionality reduction module connected to the first 3D sparse convolution for performing adaptive dynamic feature dimensionality reduction on the sparse voxel features to obtain initial BEV features; a voxel feature extraction branch connected to the first 3D sparse convolution for performing multi-scale voxel feature extraction on the sparse voxel features to obtain multi-scale voxel features containing geometric features; a multi-level feature dimensionality reduction module connected to the voxel feature extraction branch for reducing the dimensionality of the voxel features at each scale; a BEV feature extraction branch connected to the adaptive dynamic feature dimensionality reduction module and the multi-level feature dimensionality reduction module for performing multi-scale BEV feature extraction on the initial BEV features to obtain multi-scale BEV features containing semantic features; and a 3D detection module connected to the BEV feature extraction branch for performing 3D detection using the finally fused BEV features of the BEV feature extraction branch to estimate the position of the target object in the BEV space.

[0032] The detection method and model of the embodiments of the present invention can be applied to an autonomous driving vehicle equipped with a lidar to perceive the environment, improve the performance of the existing point cloud 3D target detection algorithm without adding extra time consumption, and improve the environmental perception ability of the autonomous driving vehicle and the accuracy and recall rate of target detection; it can also be applied to a robot equipped with a lidar to perceive the environment. It should be understood that its application scenario does not limit the protection scope and can also be applied to other suitable scenarios.

[0033] The method of the embodiments of the present invention can be applied to the existing grid-based point cloud 3D target detection algorithm, achieving real-time performance of more than 10 FPS on a 3090 GPU and reaching SOTA indicators on the nuScenes large-scale autonomous driving dataset. Currently (as of October 1, 2022), it ranks second in the total index NDS on the pure point cloud list, as shown in the following table:

[0034]

[0035]

[0036] Among them, * indicates that test data augmentation is used.

[0037] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several equivalent substitutions or obvious variations can be made, and if the performance or use is the same, they should all be regarded as falling within the protection scope of the present invention.

Claims

1. A 3D point cloud detection method based on self - adaptation and multi - level feature dimensionality reduction, characterized in that, Including: Obtain a real-time point cloud data frame and perform point cloud voxelization to obtain initial voxel features and their voxel coordinates; Sparsify the initial voxel features to obtain sparse voxel features; Perform adaptive dynamic feature dimensionality reduction on the sparse voxel features to obtain initial BEV features, and perform multi-scale BEV feature extraction on the initial BEV features to obtain multi-scale BEV features including semantic features; Perform multi-scale voxel feature extraction on the sparse voxel features to obtain multi-scale voxel features including geometric features, reduce the dimensionality of the voxel features at each scale, and then fuse them with the BEV features at the corresponding scale; Use the last layer of fused BEV features for 3D detection to estimate the position of the target object in the BEV space; The step of performing adaptive dynamic feature dimensionality reduction on the sparse voxel features includes: estimating the feature space distribution along the height dimension of the sparse voxel features, and this estimated feature space distribution is used to characterize the importance weight of the object features along the height dimension; re-weight the object features along the height dimension according to the estimated feature space distribution to obtain the initial BEV features; The step of estimating the feature space distribution along the height dimension of the sparse voxel features includes: estimating the voxel feature importance weight through a 3D sub-manifold sparse convolution with a convolution kernel size of 3×3×3, and its number of channels is 1; normalize the voxel feature importance along the height dimension to obtain the estimated feature space distribution; The step of re-weighting the object features along the height dimension according to the estimated feature space distribution includes: dynamically weighting the sparse voxel features along the height dimension using the estimated feature space distribution to obtain the initial BEV features.

2. The 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction according to claim 1, wherein The step of performing multi-scale voxel feature extraction on the sparse voxel features includes: Construct a voxel feature extraction branch for extracting geometric features, use the sparse voxel features as the initial input of this voxel feature extraction branch, and perform multi-stage voxel feature extraction to correspondingly obtain the multi-scale voxel features; Wherein, the output result of each stage of voxel feature extraction is used as the input of the next stage of voxel feature extraction.

3. The 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction according to claim 2, characterized in that, The step of reducing the dimensionality of the voxel features at each scale includes: Reduce the dimensionality of the output result of each stage of voxel feature extraction to obtain the dimensionality-reduced voxel features at the corresponding stage.

4. The 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction according to claim 3, characterized in that The step of performing multi-scale BEV feature extraction on the initial BEV features includes: Construct a BEV feature extraction branch for extracting semantic features, use the initial BEV features as the initial input of this BEV feature extraction branch, and perform multi-stage BEV feature extraction; the output result of each stage of BEV feature extraction is fused with the dimensionality-reduced voxel features at the corresponding stage and used as the input of the next stage of BEV feature extraction.

5. The 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction according to claim 1, characterized in that, The step of using the last layer of fused BEV features for 3D detection to estimate the position of the target object in the BEV space includes: The fused BEV features of the last layer are fed into the Center Head detection head through a multi-scale Neck network to estimate the position of the object in the BEV space. Meanwhile, the center offset residuals of the 3D bounding box of the object, the three-dimensional size of the bounding box, the height of the center point of the bounding box, the heading angle of the object, and the two-axis velocity of the object are regressed.

6. The 3D point cloud detection method based on adaptive and multi-level feature dimensionality reduction according to claim 1, wherein, The steps for sparsifying the initial voxel features include: Processing the initial voxel features through a 3D sparse convolution with a convolution kernel size of 5×5×1 to increase the receptive field and obtain the sparse voxel features.

Citation Information

Patent Citations

  • Three-dimensional target detection method based on point cloud

    CN112288709A

  • Point cloud target detection method fusing original point cloud and voxel division

    CN113378854A