Point cloud three-dimensional target detection method based on decoupling feature learning
Through the combination of decoupling feature learning and hollow convolution, the target classification and detection box regression features are extracted respectively, which solves the problem of indistinguishable feature requirements in the existing technology, and improves the accuracy of three-dimensional target detection and the detection ability of large-scale targets.
Patent Information
- Application Number
- CN202510349742.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing three-dimensional object detection method fails to fully distinguish feature requirements in the two subtasks of target classification and three-dimensional detection box regression, resulting in errors in the detection results and insufficient receptive fields, which affects the detection accuracy.
The decoupled feature learning method is adopted to extract the target classification and detection box regression features through two decoupled holes ResNet-18 backbone networks, and introduce hollow convolutional expansion receptive fields into the network, combining adaptive weight adjustment and auxiliary detection heads to improve the targetedness of feature learning.
It effectively improves the accuracy of target detection and the detection ability of large-scale targets, and improves the comprehensive detection effect of the model.
Smart Images

Figure CN120279542A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional object detection, and particularly relates to a three-dimensional object detection method for point clouds based on decoupled feature learning. Background Art
[0002] Three-dimensional object detection is an important research direction in the field of computer vision and is widely applied in fields such as autonomous driving, robot navigation, and intelligent monitoring. Since performing large-scale three-dimensional convolutions consumes a large amount of time and computing resources, in order to meet the real-time requirements, the current mainstream three-dimensional object detection methods based on point clouds usually encode point clouds into voxels or point cloud columns and further convert them into two-dimensional BEV (bird eyes view) pseudo-images, and use more efficient two-dimensional convolutions to further extract features from the pseudo-images. In the literature (Li J, Luo C, Yang X. PillarNeXt: Rethinking network designs for 3D object detection in LiDAR point clouds [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2023: 17567-17576.), a three-dimensional object detection method based on point clouds is designed. This method uses PointNet as a point cloud encoder to convert point cloud columns into two-dimensional BEV pseudo-images, uses ResNet-18 as a backbone network to extract features, uses ASPP as a neck feature fusion network, and finally uses a CenterPoint detection head to achieve the object detection task. Chinese Patent "CN116665003A A Three-Dimensional Object Detection Method and Device for Point Clouds Based on Feature Interaction and Fusion" provides a three-dimensional object detection method for point clouds. This patent discloses a three-dimensional object detection method and device for point clouds based on feature interaction and fusion, including the following steps: First, obtain point cloud features with sparsity and integrity in the point cloud information, then obtain global features of BEV feature interaction according to the point cloud features with sparsity and integrity, and then perform multi-scale feature fusion on the global features of BEV feature interaction to achieve three-dimensional object detection for point clouds.
[0003] The 3D object detection task can be divided into two subtasks: object classification and 3D bounding box regression. Existing 3D object detection methods use the features extracted by a shared feature learning network to complete these two subtasks, without considering the differences in the features required for these two subtasks, resulting in certain errors in the detection results. In addition, existing methods rarely make special designs in the backbone network to expand the receptive field, and appropriately expanding the receptive field can effectively improve the model's ability to extract features of large-scale objects, which is of certain significance for improving the detection accuracy. To sum up, guided by the two subtasks of 3D object detection and taking into account the increase in the network receptive field, a new point cloud 3D object detection method based on decoupled feature learning is designed to solve the above problems, which is of certain significance for the field of point cloud 3D object detection. Summary of the Invention
[0004] The object of the present invention is to propose a decoupled feature extraction network, which simultaneously extracts the targeted features required for object classification and detection box regression, and introduces dilated convolution therein to expand the receptive field, thereby improving the object detection accuracy.
[0005] The technical solution of the present invention is as follows: A point cloud 3D object detection method based on decoupled feature learning includes the following steps:
[0006] Step (1) Obtain and augment the 3D point cloud data of the scene to be detected;
[0007] Step (2) Set the horizontal space detection range and the vertical space detection range, and organize the augmented 3D point cloud data within the horizontal space detection range and the vertical space detection range into the form of regular point cloud columns, with each point cloud column having the same size;
[0008] Step (3) The feature extraction network extracts the point cloud depth features from the point cloud columns, then uses the pooling function to perform max pooling on the point cloud depth features within the same point cloud column, and splices the max pooling features with the extracted point cloud depth features;
[0009] Step (4) Repeat step (3) multiple times, use the pooling function to perform max pooling on the point cloud depth features within the same point cloud column, retain the most representative feature values in the point cloud columns, and generate a sparse BEV feature map through coordinate mapping;
[0010] Step (5) Input the sparse BEV feature map into the decoupled feature extraction network to extract classification features and regression features respectively;
[0011] Step (6) Simply add the classification features and the regression features through the feature interaction module to obtain the fused features, and input them into the final CenterPoint detection head to complete the detection.
[0012] Furthermore, the decoupled feature extraction network consists of submanifold sparse convolutions, including two decoupled dilated ResNet-18 backbone networks, an ASPP feature fusion network, an auxiliary classification detection head, and an auxiliary regression detection head; an ASPP feature fusion network and an auxiliary classification detection head are arranged in sequence after one dilated ResNet-18 backbone network, and an ASPP feature fusion network and an auxiliary regression detection head are arranged in sequence after the other dilated ResNet-18 backbone network.
[0013] Furthermore, the dilated ResNet-18 backbone networks are respectively a classification backbone network and a regression backbone network;
[0014] The sparse BEV feature maps are respectively subjected to feature extraction by the classification backbone network and the regression backbone network to obtain two features with 256 channels, and then these two features are respectively input into the ASPP feature fusion networks after their corresponding backbone networks for multi-scale feature extraction and fusion, obtaining two multi-scale fusion features with 256 channels and a size of 168×168, that is, the final classification feature and regression feature.
[0015] Furthermore, the dilated ResNet-18 backbone network includes 4 convolutional modules; each convolutional module contains two convolutional layers, and each convolutional layer consists of two 3*3 convolutions; the dilation rates of the two convolutions in the convolutional layer are set to [1, 2], and the padding parameters are correspondingly modified to [1, 2].
[0016] Furthermore, the auxiliary classification detection head H auxC and the auxiliary regression detection head H auxR are modified from the CenterPoint detection head; the auxiliary classification detection head is obtained by only retaining the classification-related functions of the original CenterPoint detection head, and the auxiliary regression detection head is the same as the original CenterPoint detection head. The two are respectively used to constrain the two backbone networks to learn the corresponding classification features and regression features, and their corresponding loss functions are respectively and where f C and f R are respectively the classification feature and regression feature extracted by their corresponding backbone networks, and are respectively the ground truth of classification and the ground truth of regression.
[0017] Furthermore, the ASPP feature fusion network consists of a 1×1 convolution, three 3×3 dilated convolution modules with different dilation rates, a pooling pyramid, and an ASPP pooling layer; the classification feature or regression feature is respectively input into the ASPP feature fusion network, and then the number of channels is adjusted through a 1×1 convolution to obtain the final multi-scale feature;
[0018] The pooling pyramid includes three parallel 3×3 convolutions with dilation rates set to 6, 12, and 18 respectively; the ASPP pooling layer consists of an average pooling layer, a 1×1 convolution, and an upsampling layer.
[0019] Furthermore, the loss function of the CenterPoint detection head is
[0020] The overall loss function is L total = l U + λ C l C + λ R l R ; Design an adaptive weight adjustment method based on the cosine annealing strategy to adjust λ C and λ R :
[0021]
[0022] where it c and it T represent the current iteration number and the total iteration number during training respectively; and are the maximum and minimum values of the weights; during the training stage, the auxiliary classification detection head H auxR and the auxiliary regression detection head H auxC as well as the final CenterPoint detection head H U will be applied simultaneously; while in the inference stage, only the CenterPoint detection head H U will be applied.
[0023] Furthermore, the horizontal space detection range is [-50.4m, 50.4m], and the vertical space detection range is [-2m, 4m].
[0024] Furthermore, the size of the point cloud pillar is [0.075m, 0.075m, 8m].
[0025] Furthermore, the feature extraction network mainly consists of a linear transformation, batch normalization, and ReLU rectified linear unit.
[0026] The beneficial effects of the present invention: The present invention proposes a three-dimensional object detection method for point clouds based on decoupled feature learning: using two decoupled backbone networks to respectively learn the targeted features of the classification and regression subtasks, effectively improving the accuracy of object detection.
[0027] To enable the decoupled network to fully learn the corresponding classification and regression features, an auxiliary classification detection head and a regression detection head are inserted at the ends of the two decoupled feature extraction networks, and corresponding loss functions are proposed to constrain the backbone network to learn targeted features.
[0028] To train the network effectively and stably, an adaptive weight adjustment method based on the cosine annealing strategy is proposed. In the initial stage of training, the influence of the auxiliary detection head is enlarged to learn classification and regression features targeted. As the training progresses, the influence of the final object detection head is gradually enhanced to improve the final detection effect.
[0029] To enhance the network's detection ability for large-scale target objects and thus improve the accuracy of model object detection, this study introduces dilated convolution into the design of the decoupled backbone network to expand the receptive field. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flowchart of the 3D object detection method for point clouds based on decoupled feature learning;
[0031] Figure 2 It is a design diagram of the dilated ResNet-18 backbone network;
[0032] Figure 3 It is a visualization diagram of the detection results; DETAILED DESCRIPTION OF THE INVENTION
[0033] Figure 1 It is the main flowchart of the technical solution of the present invention. As Figure 1 shown, the 3D object detection method for point clouds based on decoupled feature learning proposed by the present invention includes the following steps:
[0034] (1) Obtain the 3D point cloud data of the scene to be detected, and perform data augmentation on the point cloud by operations such as rotation, flipping, translation, and scaling;
[0035] Rotation: Randomly rotate the point cloud around the Z axis (vertical axis) to simulate different orientations of the object. The rotation angle range is from -45 degrees to 45 degrees. For each augmentation, a value is uniformly sampled from this range, and the scene is rotated.
[0036] Flipping: Flip the point cloud horizontally (X axis) or vertically (Y axis). The horizontal or vertical flipping probability is 50% for each, and each direction is judged independently, and flipping may occur simultaneously.
[0037] Translation: Randomly translate the point cloud along the X / Y / Z axes to simulate the change of the object's position. The translation amount for each axis is uniformly sampled within the range of [-0.5, 0.5] meters.
[0038] Scaling: Isotropically scale the point cloud, with each axis scaled by the same ratio, which is uniformly sampled between 0.9 and 1.1 each time to simulate the change in object size.
[0039] (2) Set the horizontal and vertical space detection ranges to [-50.4m, 50.4m] and [-2m, 4m] respectively, and organize the enhanced point cloud data into a regular point cloud column form, with each point cloud column sized [0.075m, 0.075m, 8m].
[0040] (3) Expand the point cloud feature dimension to 9 through calculation, including: the absolute coordinates x, y, z of the point cloud; the reflection intensity r of the point cloud; the offsets Δx, Δy, Δz of the point cloud relative to the center of the point cloud column; the center coordinates x c 、y c ; Use a feature extraction network composed of linear transformation, batch normalization, and ReLU rectified linear unit to extract the depth features of the point cloud, and then use a pooling function to perform max pooling on the features within the same point cloud column, and splice the pooled features with the extracted depth features of the point cloud.
[0041] (4) Repeat step (3) multiple times. Finally, use a pooling function to perform max pooling on the depth features of the point cloud belonging to the same point cloud column, retain the most representative eigenvalue in the point cloud column, and generate a sparse BEV feature map with 64 channels and a size of 1344×1344 through coordinate mapping.
[0042] (5) Input the sparse BEV feature map into a decoupled feature extraction network composed of submanifold sparse convolution to extract classification features and regression features respectively.
[0043] The decoupled feature extraction network includes two decoupled dilated ResNet-18 backbone networks, two ASPP feature fusion networks after the decoupled backbone networks, and an auxiliary classification detection head and an auxiliary regression detection head inserted respectively thereafter. The sparse BEV feature map undergoes feature extraction through the classification backbone network and the regression backbone network respectively to obtain two features with 256 channels, and then these two features are respectively input into the ASPP feature fusion networks after their corresponding backbone networks for multi-scale feature extraction and fusion, obtaining two multi-scale fusion features with 256 channels and a size of 168×168, which are the final classification features and regression features.
[0044] Regarding the problem of the dilation rate design of the dilated ResNet-18 backbone network, such as Figure 2As shown in the figure. ResNet-18 contains 4 convolutional modules, each of which contains two convolutional layers, and each convolutional layer consists of two 3*3 convolutions. We set the dilation rates of the two convolutions in the convolutional layer to [1, 2], and at the same time modify the padding parameter to [1, 2] correspondingly to ensure that the resolution of the feature map remains unchanged.
[0045] Auxiliary classification detection head H auxC It is transformed from the CenterPoint detection head and retains its classification detection function to constrain the corresponding backbone network to learn classification features; Auxiliary regression detection head H auxR Adopts the original CenterPoint detection head to constrain the corresponding backbone network to learn regression features; Their corresponding loss functions are respectively and where f C and f R are the classification features and regression features extracted by their corresponding backbone networks respectively, and are the ground truth of classification and the ground truth of regression respectively.
[0046] The ASPP feature fusion network consists of a 1×1 convolution, a pooling pyramid, and an ASPP pooling layer. The features are input into these three modules respectively, and then the outputs are concatenated. Finally, the number of channels is adjusted through a 1×1 convolution to obtain the final multi-scale features. The pooling pyramid contains three parallel 3×3 convolutions with dilation rates set to 6, 12, and 18 respectively. The ASPP pooling layer consists of an average pooling layer, a 1×1 convolution, and an upsampling layer.
[0047] (6) Simply add the classification features and regression features extracted by the decoupled feature extraction network through the feature interaction module to obtain the fused features.
[0048] (7) Input the fused features into the final CenterPoint detection head to complete the detection.
[0049] The loss function of the CenterPoint detection head is
[0050] The overall loss function is L total = l U + λ C l C + λ R l R . To ensure the stability of training, an adaptive weight adjustment method based on the cosine annealing strategy is designed:
[0051]
[0052] where it c and itT respectively represent the current iteration number and the total number of iterations during training. and are the maximum and minimum values of the weights, which are set to 1 and 0 respectively in this study. Using this adaptive weight adjustment method, the network can gradually enhance the influence of decoupled feature learning as the training progresses to obtain more targeted classification features and regression features. During the training phase, the auxiliary detection head H auxR and H auxC as well as the final detection head H U will be applied simultaneously; while in the inference phase, only the detection head H U will be applied.
[0053] The present invention selects the lidar point cloud part in the NuScenes dataset as the training and evaluation dataset. This dataset collects point clouds for 10 different target categories, uses a 32-beam lidar, and generates approximately 30,000 points per frame at a frequency of 20 Hz. The official evaluation metrics of the NuScenes dataset include average precision (AP), average translation error (ATE), average scale error (ASE), average orientation error (AOE), average velocity error (AVE), average attribute error (AAE), and the nuScenes detection score (NDS), where the latter is calculated comprehensively based on the above six metrics. In the present invention, we use the mean average precision (mAP), average translation error (ATE), and nuScenes detection score (NDS) to measure the classification ability, regression ability, and overall detection ability of the present invention respectively. The present invention uses an NVIDIA RTX 4090 GPU for training and inference.
[0054] The present invention first conducted ablation experiments on decoupled feature learning and the dilated convolutional backbone network to prove the effectiveness of these two parts in improving the detection effect.
[0055] As shown in Table 1, compared with the baseline network model, the mean average translation error (mATE) of the decoupled feature extraction network model is smaller, which indicates that the present invention produces less error when regressing the attributes of 3D bounding boxes; the mean average precision (mAP) of the decoupled network model is also higher than that of the baseline network model, which indicates that the present invention effectively improves the accuracy of the classification task; in addition, the improvement of the nuScenes detection score (NDS) also confirms that the decoupled feature learning network is effective in improving the comprehensive detection effect of the 3D object detection task. Therefore, the decoupled feature learning network proposed by the present invention is effective for the network to learn targeted classification and regression features, and this decoupled feature learning method can effectively improve the comprehensive object detection ability of the model. In addition, the addition of dilated convolution can also confirm its effectiveness in improving the detection effect for the improvement of the nuScenes detection score (NDS).
[0056] Table 1 Results of Ablation Experiments
[0057]
[0058] The present invention also verifies the effectiveness of dilated convolution for detecting large-scale objects by statistically analyzing the influence of adding dilated convolution on the detection results of different categories of objects. As shown in Table 2, the model with dilated convolution has a significant improvement in the detection ability of large-scale objects such as buses, trailers, and trucks, thus proving the effectiveness of adding dilated convolution in expanding the receptive field of the network.
[0059] Table 2 Influence of Dilated Convolution on the Detection Results of Different Categories of Objects
[0060]
[0061] In addition, the present invention also compares various 3D object detection methods at home and abroad. As shown in Table 3, the present invention reaches 69.1% and 62.9% in terms of nuScenes detection score (NDS) and mean average precision (mAP) respectively, which are 1.1% and 0.8% higher than those of the baseline model PillarNeXt-B, indicating the effectiveness of the decoupled expansion feature learning process.
[0062] Table 3 Comparison Results with Various 3D Object Detection Methods
[0063]
[0064]
[0065] The present invention also visualizes the detection results. As Figure 3 shown, the green box represents the ground truth of the detection box, and the red represents the detection result of the present invention. It can be clearly seen from the figure that the present invention has excellent performance in object detection.
Claims
1. A 3D object detection method for point clouds based on decoupled feature learning, characterized in that, It includes the following steps: Step (1) Obtain and enhance the 3D point cloud data of the scene to be detected; Step (2) Set the horizontal space detection range and the vertical space detection range, and organize the enhanced 3D point cloud data within the horizontal and vertical space detection ranges into the form of regular point cloud columns, with each point cloud column having the same size; Step (3) The feature extraction network extracts the point cloud depth features from the point cloud columns, then uses a pooling function to perform max pooling on the point cloud depth features within the same point cloud column, and splices the max pooling features with the extracted point cloud depth features; Step (4) Repeat Step (3) multiple times, use the pooling function to perform max pooling on the point cloud depth features within the same point cloud column, retain the most representative eigenvalue in the point cloud column, and generate a sparse BEV feature map through coordinate mapping; Step (5) Input the sparse BEV feature map into the decoupled feature extraction network to extract classification features and regression features respectively; Step (6) Simply add the classification features and regression features through the feature interaction module to obtain the fused features, and input them into the final CenterPoint detection head to complete the detection.
2. The 3D object detection method for point clouds based on decoupled feature learning according to claim 1, characterized in that The decoupled feature extraction network consists of submanifold sparse convolutions, including two decoupled dilated ResNet-18 backbone networks, an ASPP feature fusion network, an auxiliary classification detection head, and an auxiliary regression detection head; an ASPP feature fusion network and an auxiliary classification detection head are arranged in sequence after one dilated ResNet-18 backbone network, and an ASPP feature fusion network and an auxiliary regression detection head are arranged in sequence after the other dilated ResNet-18 backbone network.
3. The 3D object detection method for point clouds based on decoupled feature learning according to claim 2, characterized in that The dilated ResNet-18 backbone networks are respectively the classification backbone network and the regression backbone network; The sparse BEV feature map undergoes feature extraction by the classification backbone network and the regression backbone network respectively to obtain two features with 256 channels, and then these two features are respectively input into the ASPP feature fusion networks after their corresponding backbone networks for multi-scale feature extraction and fusion to obtain two multi-scale fusion features with 256 channels and a size of 168×168, which are the final classification features and regression features.
4. The 3D object detection method for point cloud based on decoupled feature learning according to claim 2, characterized in that The dilated ResNet-18 backbone network includes 4 convolutional modules; each convolutional module contains two convolutional layers, and each convolutional layer consists of two 3*3 convolutions; the dilation rates of the two convolutions in the convolutional layer are set to [1, 2], and the padding parameters are modified accordingly to [1, 2].
5. The 3D object detection method for point cloud based on decoupled feature learning according to claim 2, wherein The auxiliary classification detection head H auxC and the auxiliary regression detection head H auxR are transformed from the CenterPoint detection head; the auxiliary classification detection head is obtained by retaining only the classification-related functions of the original CenterPoint detection head, and the auxiliary regression detection head is the same as the original CenterPoint detection head. The two are respectively used to constrain the two backbone networks to learn the corresponding classification features and regression features, and their corresponding loss functions are and where f c and f R are respectively the classification features and regression features extracted by their corresponding backbone networks, and are respectively the ground truth of classification and the ground truth of regression.
6. The 3D object detection method for point clouds based on decoupled feature learning according to claim 2, wherein, The ASPP feature fusion network consists of a 1×1 convolution, three 3×3 dilated convolution modules with different dilation rates, a pooling pyramid, and an ASPP pooling layer; the classification features or regression features are respectively input into the ASPP feature fusion network, and then the channel number is adjusted through a 1×1 convolution to obtain the final multi-scale features; The pooling pyramid contains three parallel 3×3 convolutions with dilation rates set to 6, 12, and 18 respectively; the ASPP pooling layer consists of an average pooling layer, a 1×1 convolution, and an upsampling layer.
7. The 3D object detection method for point clouds based on decoupled feature learning according to claim 2, characterized in that The loss function of the CenterPoint detection head is The overall loss function is L total = l U + λ C l C + λ R l R ; Design an adaptive weight adjustment method based on the cosine annealing strategy to adjust λ C and λ R : Among them, it c and it T represent the current iteration number and the total iteration number during training, respectively; and are the maximum and minimum values of the weights; during the training stage, the auxiliary classification detection head H auxR and the auxiliary regression detection head H auxC as well as the final CenterPoint detection head H U will be applied simultaneously; while in the inference stage, only the CenterPoint detection head H U will be applied.
8. The 3D object detection method for point clouds based on decoupled feature learning according to claim 1, wherein The horizontal space detection range is [-50.4m, 50.4m], and the vertical space detection range is [-2m, 4m].
9. The 3D object detection method for point clouds based on decoupled feature learning according to claim 1, wherein, The size of the point cloud column is [0.075m, 0.075m, 8m].
10. The 3D object detection method for point clouds based on decoupled feature learning according to claim 1, wherein The feature extraction network mainly consists of linear transformation, batch normalization, and ReLU rectified linear unit.
Citation Information
Patent Citations
Point cloud three-dimensional target detection method and device based on feature interaction and fusion
CN116665003A