3D Object Detection Method Based on Multi-Scale Graph Neural Network and Point Cloud Reduction Network
By using multi-scale map neural networks and point cloud reduction networks in 3D object detection, removing irrelevant backgrounds and using key point map neural networks to process point cloud data, the problem of handling complex relationships and related dependencies in the prior art is solved, and faster and more accurate 3D object detection is achieved.
Patent Information
- Application Number
- CN202210477589.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-26
AI Technical Summary
Existing 3D object detection methods are difficult to effectively deal with complex relationships in point cloud data and related dependencies between objects, resulting in slow detection speed and difficult model training.
A 3D object detection method based on multi-scale graph neural network and point cloud reduction network is proposed. The PointNet++ network is used to remove specific irrelevant backgrounds, reduce the point cloud data scale, and use the key point graph neural network module to process the spatial relationship of point clouds.
Faster and accurate 3D object detection is achieved. By removing irrelevant backgrounds and extracting features using graph structures, the calculation consumption and inference time are reduced, and the training efficiency and detection performance of the model are improved.
Smart Images

Figure CN115100641B_ABST
Abstract
Description
Technical Field:
[0001] The present invention relates to the field of computer vision. Specifically, it relates to a 3D object detection method based on a multi-scale graph neural network and a point cloud reduction network. Background Art:
[0002] The statements in this section only relate to the background art related to the present invention and do not necessarily constitute the prior art.
[0003] 3D object detection technology plays an important role in the field of computer vision and is used to obtain the position and category information of objects in three-dimensional space. There are mainly methods based on point clouds, binocular vision, monocular vision, and multi-modal data, etc. Among them, point cloud data has rich geometric information, is closer to the most primitive representation of objects, and is more stable than other modal data. Therefore, 3D object detection technology based on lidar point cloud data is widely used in various fields.
[0004] In recent years, deep learning on point clouds has been booming and plays an important role in research tasks such as 3D shape classification, 3D object detection and tracking, and 3D point cloud segmentation. Point cloud data is mainly obtained by LiDAR devices and has the characteristics of sparsity, disorder, irregularity, and a large quantity. Therefore, how to efficiently process point cloud data is a major challenge. Existing methods either directly process points or indirectly process point clouds by projecting them onto a two-dimensional coordinate system. However, these methods cannot effectively handle the complex relationships between points and the correlation dependencies between objects.
[0005] Designing an effective model architecture to obtain more accurate and faster detection targets is a research hotspot in 3D object detection tasks. In addition, how to utilize the complex relationships between points and the correlation dependencies between objects is also a key issue for accurate 3D object detection. Graph neural networks (GNNs) can transform data in irregular forms into regular representations in Euclidean space. Compared with the most basic network structure of neural networks, the fully connected layer (MLP), where the feature matrix is multiplied by the weight matrix, graph neural networks have an additional adjacency matrix. The calculation form is very simple, multiplying three matrices and then adding a non-linear transformation. Graph neural networks can be divided into recursive graph neural networks (RecGNNs), convolutional graph neural networks (ConvGNNs), graph autoencoders (GAEs), and spatio-temporal graph neural networks (STGNNs). Based on the characteristics of graph neural networks for processing irregular data and complex relationships, related methods have been applied in a small number of 3D point cloud object detections and achieved good performance. However, there are still great challenges in applying graph neural networks to process the spatial relationships of point clouds. In addition, the original point cloud data is large in quantity and has relatively more irrelevant data, which will lead to too slow inference speed and difficult model training. Summary of the Invention:
[0006] To alleviate the above problems, in this invention, we use the backbone network structure of PointNet++ to remove specific irrelevant backgrounds and propose a key-point graph neural network (called KeyPoint-GNN) to achieve accurate and fast 3D object detection tasks. This method consists of two parts, namely the point cloud data reduction module (PCR) and the key-point graph neural network module (KPGNN). Different from traditional methods that directly process the original point cloud data, the PCR method aims to reduce the computational load by removing point cloud data of specific irrelevant backgrounds. More specifically, we set walls and the ground in large outdoor scenes as specific irrelevant backgrounds and use the backbone network of PointNet++ to remove them. Then we sample key points from the reduced point cloud data and use the graph structure to represent the spatial relationship between points. At the same time, we extract features from each graph structure and use these features as vertices to construct a graph again. Finally, we splice feature information of different scales to obtain the detection target quickly and efficiently. Finally, we train the entire network in an end-to-end manner and achieve better prediction performance.
[0007] The technical solution of this invention provides a 3D object detection method based on a multi-scale graph neural network and a point cloud reduction network. This method includes the following steps:
[0008] 1. Input point cloud data and use the feature encoding and feature encoding module in the PointNet++ network to preprocess the point cloud data, removing specific irrelevant backgrounds to reduce the computational input;
[0009] 1.1) Collect relevant datasets in the field of point cloud 3D object detection, including the KITTI dataset, the ModelNet40 dataset, the Waymo dataset, the NuScenes dataset, and the Lyft L5 dataset.
[0010] 1.2) In this invention, use the KITTI dataset training dataset with 80,256 object labels to train the model; use the test dataset in the KITTI dataset to detect the generalization performance of the model.
[0011] 1.3) Set the ground and walls in the point cloud as specific irrelevant backgrounds
[0012] 1.4) First, we use two sets of Set Abstraction for feature extraction. The input of each extraction layer is (N, (d + C)), where N is the number of input points, d is the dimension of the coordinates, and C is the feature dimension. The output is (N′, (d + C′)), where N′ is the number of output points, the dimension of d remains unchanged, and C′ is the new feature dimension.
[0013] 1.5) Subsequently, we use Feature Propagation to map the features of points to the entire point cloud data through upsampling, and use linear interpolation to splice the features with those in Set Abstraction. The relevant formulas are shown as follows:
[0014]
[0015] Here, ω i represents the weight, and f represents the output feature. Finally, the reduced point cloud set P = {p1, p2, p3, … p n} is obtained.
[0016] 2. Select key points in the point cloud using the key point sampling algorithm, and construct a spatial graph with the spatial neighboring points within a fixed radius centered at the key points;
[0017] 2.1) First, use Sectorized Proposal-Centric (SPC) key point sampling to uniformly sample from the surrounding area of the proposal. The sampled key points P′ are represented by the following formula:
[0018]
[0019] where r (s) represents the maximum hyperparameter of the proposal expansion radius, dx j , dy j , dz j represent the size of the proposal, and C, D represent the center and size of the 3D proposal.
[0020] 2.2) Subsequently, use the KD-tree algorithm to search for the neighbor points within a fixed radius of each key point and construct it into a graph G, as shown in the following formula:
[0021] G = KD(P′, R),(3)
[0022] where R represents the specified search radius.
[0023] 3. Use a convolutional neural network to extract features for each graph, construct a planar graph with each feature as a vertex, and start from a certain node to perform feature fusion at different scales with a fixed search radius, as Figure 2 shown;
[0024] 3.1) First, extract features for each spatial graph G, and further construct a graph G′ for the extracted features, as shown in the following formula:
[0025]
[0026] where n represents the number of points in each spatial graph.
[0027] 3.2) Subsequently, the constructed graph G′ is used to perform feature fusion at different scales with a fixed search radius using the nearest neighbor search algorithm, as shown in the following equation:
[0028] f i = A(NNS(G′, R)), (5)
[0029] where the A function represents the feature fusion function, and f i represents the features aggregated at different scales.
[0030] 4. Then, use MLP to convert the features to the same scale and train this model using a hybrid loss function.
[0031] 4.1) Use MLP to transform the features of different scales to the same dimension and concatenate them, as shown in the following equation:
[0032] f = Concat(MLP(f i )), (6)
[0033] where Concat represents the feature concatenation operation.
[0034] 4.2) Input the obtained features into the detection head for 3D proposal optimization and confidence prediction. For confidence prediction, the 3D IoU between the 3D RoI and its corresponding ground truth is used as the training target, as shown in the following equation:
[0035] y k = min(1, max(0, 2IoU k - 0.5)), (7)
[0036] where IoU k is the IoU between the k-th RoI and the true box.
[0037] Then, use cross-entropy loss to train the confidence, as shown in the following equation:
[0038]
[0039] where is the predicted value of the network.
[0040] 4.3) Smooth L1 loss function is used for bounding box optimization.
[0041]
[0042] where f(x i ) represents the predicted value, and y i represents the true value.
[0043] Advantages of the present invention: By removing specific irrelevant backgrounds to reduce the scale of the point cloud, the present invention aims to save computing consumption and time. Through the key point graph neural network module, the scale of processing point cloud data is further reduced, and the spatial relationship of the point cloud is fully utilized through the graph structure to more quickly complete the extraction of the local features of the entire point cloud. Then, by aggregating the features of different scales in the second-stage graph structure, a global feature representation is obtained, and then a strong feature representation is generated to accurately locate the 3D bounding box and be used for quickly and efficiently detecting 3D scene targets. Brief Description of the Drawings:
[0044] Figure 1 Network Flow Framework Diagram
[0045] Figure 2 Multi-scale Graph Neural Network Module Detailed Embodiments:
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. In addition, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art in this research direction without creative efforts shall fall within the protection scope of the present invention.
[0047] The flowchart framework of the present invention is as Figure 1 shown. The 3D object detection method of the present invention based on a multi-scale graph neural network and a point cloud reduction network is specifically described as follows:
[0048] 1. Input point cloud data, and use the feature encoding and feature encoding modules in the PointNet++ network to preprocess the point cloud data to remove specific irrelevant backgrounds to reduce the computing input;
[0049] 1.1) Collect relevant datasets in the field of point cloud 3D object detection, including the KITTI dataset, the ModelNet40 dataset, the Waymo dataset, the NuScenes dataset, and the Lyft L5 dataset.
[0050] 1.2) In this invention, the KITTI dataset with 80,256 object labels is used as the training dataset to train the model; the test dataset in the KITTI dataset is used to detect the generalization performance of the model.
[0051] 1.3) Set the ground and walls in the point cloud as specific irrelevant backgrounds
[0052] 1.4) First, we use two sets of Set Abstraction for feature extraction. The input of each extraction layer is (N, (d + C)), where N is the number of input points, d is the dimension of the coordinates, and C is the feature dimension. The output is (N′, (d + C′)), where N′ is the number of output points, the dimension of d remains unchanged, and C′ is the new feature dimension.
[0053] 1.5) Subsequently, we use Feature Propagation to map the features of points to the entire point cloud data through upsampling and use linear interpolation to splice the features in Set Abstraction. The relevant formula is as follows:
[0054]
[0055] Here, ω i represents the weight, and f represents the output feature. Finally, the reduced point cloud set P = {p1, p2, p3, … p n} is obtained.
[0056] 2. Use the key point sampling algorithm to select key points in the point cloud and construct a spatial graph with the spatial neighboring points within a fixed radius centered at the key points, as Figure 2 shown;
[0057] 2.1) First, use Sectorized Proposal-Centric (SPC) key point sampling to uniformly sample from the surrounding area of the proposal. The key points P′ after sampling are represented by the following formula:
[0058]
[0059] where r (s) represents the maximum hyperparameter of the proposal expansion radius, dx j , dy j , dz j represent the size of the proposal, and C, D represent the center and size of the 3D proposal.
[0060] 2.2) Subsequently, use the KD-tree algorithm to search for the neighbor points within a fixed radius of each key point and construct it into a graph G, as shown in the following formula:
[0061] G = KD(P′, R),(3)
[0062] where R represents the specified search radius.
[0063] 3. Use a convolutional neural network to extract features for each graph, construct a planar graph with each feature as a vertex, and start from a certain node to perform feature fusion at different scales with a fixed search radius, as Figure 2as shown
[0064] 3.1) First, feature extraction is performed on each spatial graph G, and the extracted features are further used to construct graph G′, as shown in the following formula:
[0065]
[0066] where n represents the number of points in each spatial graph.
[0067] 3.2) Subsequently, the constructed graph G′ is used to perform feature fusion at different scales using the nearest neighbor search algorithm with a fixed search radius, as shown in the following formula:
[0068] f i = A(NNS(G′, R)), (5)
[0069] where the A function represents the feature fusion function, and f i represents the features aggregated at different scales.
[0070] 4. Then, use MLP to convert the features to the same scale and train this model using the hybrid loss function.
[0071] 4.1) Use MLP to transform the features at different scales to the same dimension and concatenate them, as shown in the following formula:
[0072] f = Concat(MLP(f i )), (6)
[0073] where Concat represents the feature concatenation operation.
[0074] 4.2) Input the obtained features into the detection head for 3D proposal optimization and confidence prediction. For confidence prediction, use the 3D IoU between the 3D RoI and its corresponding ground truth as the training objective, as shown in the following formula:
[0075] y k = min(1, max(0, 2IoU k - 0.5)), (7)
[0076] where IoU k is the IoU between the kth RoI and the true box.
[0077] Then, use cross - entropy loss to train the confidence, as shown in the following formula:
[0078]
[0079] where is the predicted value of the network.
[0080] 4.3) The boundary box optimization adopts the Smooth L1 loss function.
[0081]
[0082] where f(x i ) represents the predicted value, and y i represents the true value.
[0083] The above is the preferred implementation of the present application and is not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A 3D object detection method based on a multi-scale graph neural network and a point cloud reduction network, characterized in that The method includes the following steps: 1) Input point cloud data, and use the feature encoding and feature encoding modules in the PointNet++ network to preprocess the point cloud data, removing specific irrelevant backgrounds to reduce the computational input; 1.1) Collect relevant datasets in the field of point cloud 3D object detection, including the KITTI dataset, ModelNet40 dataset, Waymo dataset, NuScenes dataset, and Lyft L5 dataset; 1.2) Use the KITTI dataset with 80,256 object labels as the training dataset to train the model; use the test dataset in the KITTI dataset to detect the generalization performance of the model; 1.3) Set the ground and walls in the point cloud as specific irrelevant backgrounds; 1.4) First, use two sets of Set Abstraction for feature extraction. The input of each extraction layer is (N, (d + C)), where N is the number of input points, d is the dimension of the coordinates, and C is the feature dimension. The output is (N′, (d + C′)), where N′ is the number of output points, the dimension of d remains unchanged, and C′ is the new feature dimension; 1.5) Subsequently, use Feature Propagation to map the features of points to the entire point cloud data through upsampling, and use the method of linear interpolation to splice with the features in Set Abstraction. The relevant formula is shown as follows: Here, ω i represents the weight, f represents the output feature, and finally the reduced point cloud set P = {p1, p2, p3, … p n}; 2) Use the key point sampling algorithm to select key points in the point cloud, and construct a spatial graph with the spatial neighboring points within a fixed radius centered on the key points; 3) Use a convolutional neural network to extract features for each graph, and construct a planar graph with each feature as a vertex. Starting from a certain node, perform feature fusion at different scales with a fixed search radius; 3.1) First, extract features for each spatial graph G, and further construct graph G′ for the extracted features, as shown in the following formula: where n represents the number of points in each spatial graph, |p u - p k | represents the distance between point u and point k, and f u represents the feature of point u; 3.2) Subsequently, use the nearest neighbor search algorithm to perform feature fusion at different scales on the constructed graph G′ with a fixed search radius, as shown in the following formula: f i = A(NNS(G′, R)), (5) Among them, the A function represents the feature fusion function, and f i represents the features after aggregation at different scales; 4) Use an MLP to convert the features to the same scale, and train this model using a hybrid loss function.
2. The 3D object detection method based on the multi-scale graph neural network and the point cloud reduction network according to claim 1, characterized in that The above step 2) is specifically as follows: 2.1) First, use Sectorized Proposal-Centric (SPC) key point sampling to uniformly sample from the surrounding area of the proposal. The key point P′ after sampling is represented by the following formula: where r (s) denotes the maximum hyperparameter of the proposal expansion radius, dx j , dy j , dz j denote the size of the proposal, and C, D denote the center and size of the 3D proposal; 2.2) Subsequently, use the KD-tree algorithm to search for the neighbor points within a fixed radius of each key point, and construct it as graph G, as shown in the following formula: G = KD(P', R),(3) where R represents the specified search radius.
3. The 3D object detection method based on the multi-scale graph neural network and the point cloud reduction network according to claim 1, characterized in that The above step 4) is specifically as follows: 4.1) Use an MLP to transform features of different scales to the same dimension and splice them, as shown in the following formula: f = Concat(MLP(f i ),(6) where Concat represents the feature splicing operation; 4.2) Input the obtained features into the detection head for 3D proposal optimization and confidence prediction. For confidence prediction, use the 3D IoU between the 3D RoI and its corresponding ground truth as the training objective, as shown in the following equation: y k = min(1, max(0, 2IoU k - 0.5)), (7) where IoU k is the IoU between the k-th RoI and the ground truth box, and then the confidence is trained using the cross-entropy loss as shown in the following equation: Among them is the predicted value of the network; 4.3) Use the Smooth L1 loss function for bounding box optimization: where f(x i ) represents the predicted value, and y i represents the true value.
Citation Information
Patent Citations
Point cloud 3D target detection method based on key point multi-scale feature fusion
CN113706480A
Three-dimensional dynamic target detection method and device based on voxel point cloud fusion
CN113989797A