3D target detection method and system based on point cloud and voxel dynamic fusion
Through the dynamic fusion method of point cloud and voxel, deep learning and attention mechanisms are used to extract and fuse point cloud and voxel features, solving the problems of high computing complexity and insufficient feature retention in the existing technology, and achieving higher precision 3D object detection.
Patent Information
- Application Number
- CN202510415911.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
AI Technical Summary
When processing point cloud data, the existing 3D object detection methods have problems such as high computational complexity, noise sensitivity and inability to fully retain details and high-frequency characteristics, which are difficult to meet the accuracy requirements of autonomous driving and robot vision.
Using a method based on dynamic fusion of point cloud and voxel, features are extracted through Pointnet++ and Voxelnet networks, feature fusion is performed by combining channel attention, cross attention and self-attention modules, and detection boxes are updated through multiple iterations, and VoxelVFE is used for voxelization, combined with deep learning optimization process.
It improves detection accuracy and generalization performance, can accurately capture local and global features, and generate more accurate target detection results, especially in categories such as pedestrians, bicycles and cars.
Smart Images

Figure CN120299020A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object detection, in particular to a 3D object detection method and system based on dynamic fusion of point cloud and voxel. Background Art
[0002] Three-dimensional lidar object detection is a key technology in the fields of autonomous driving, robot vision, and environmental perception. Its core lies in identifying and locating target objects from the point cloud data obtained by a lidar (LiDAR). Compared with traditional two-dimensional image data, point cloud data has richer spatial information, but it also brings challenges in processing. Point cloud data not only has the problems of high dimension and large data volume, but also due to the uneven sampling and the influence of noise, the complexity of processing and analyzing point cloud data is relatively high. Therefore, how to effectively use point cloud data for accurate object detection has always been a difficult problem that researchers urgently need to solve.
[0003] Currently, 3D object detection methods based on point cloud data can be roughly divided into two categories: direct processing methods based on point cloud and voxelization methods. The direct processing methods usually directly extract features and identify targets from the original point cloud through a deep learning model. The advantage of this type of method is that it can more accurately process the local features of the point cloud, but its computational complexity is relatively high, and it is more sensitive to the density and noise of the point cloud data. In contrast, the voxelization method maps the point cloud data into a regular three-dimensional grid (voxel), thereby converting the point cloud into more regular volume data, which is convenient for subsequent processing and analysis and reduces the computational complexity. However, after voxelization, the details and high-frequency features (such as object boundaries, local structures, etc.) in the original point cloud may not be fully retained, and quantization errors will inevitably be introduced. It can be seen that the existing 3D object detection methods all have limitations and cannot meet the requirements for object detection accuracy in autonomous driving, robot vision, and environmental perception, etc. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide a 3D object detection method and system based on dynamic fusion of point cloud and voxel with high accuracy and strong generalization ability.
[0005] Technical Solution: The 3D object detection method based on dynamic fusion of point cloud and voxel according to the present invention includes the following steps:
[0006] (1) Obtain the three-dimensional point cloud information of the target to be measured, and perform voxelization processing on the three-dimensional point cloud information to obtain three-dimensional voxel information;
[0007] (2) Respectively use the Pointnet++ network and the Voxelnet network to extract features from the point cloud and voxel information to obtain point cloud features and voxel features;
[0008] (3) Screen the point cloud features according to the foreground point probability to obtain the candidate point cloud features;
[0009] (4) Obtain the detection box according to the candidate point cloud features;
[0010] (5) Obtain the candidate point voxel features within the detection box;
[0011] (6) Fuse the candidate point cloud features and the candidate point voxel features through a fusion module to obtain candidate point features; the fusion module includes a channel attention module, a cross-attention module, and a self-attention module. The candidate point voxel features are input into the cross-attention module after passing through the channel attention module, the candidate point cloud features are directly input into the cross-attention module, the output of the cross-attention module is connected to the self-attention module, and the output of the self-attention module is the candidate point features;
[0012] (7) Obtain an updated detection box according to the candidate point features, and repeat steps (5) to (7) to obtain the final detection box and the classification information of the target to be measured within the detection box.
[0013] Further, in step (4), the detection box is obtained through a rough detection module, and the rough detection module includes three sequentially connected fully connected layers.
[0014] Further, in step (7), the updated detection box is obtained through a detection module, and the detection module includes four sequentially connected fully connected layers.
[0015] Further, in step (7), during the process of repeating steps (5) to (7), the loss function is where W i = 2×i / (n 2 + n), is the prediction loss of the detection box obtained in the i-th repetition, W i is the corresponding loss weight, and n is the total number of repetitions.
[0016] Further, in step (6), the cross-attention module uses D as the position encoding, and the self-attention module uses the candidate point coordinates as the position encoding, where D is the Euclidean distance between the voxel coordinates within the detection box and the candidate point coordinates.
[0017] Further, the method for calculating the Euclidean distance between the voxel coordinates and the candidate point coordinates is: convert the voxel coordinates to point cloud coordinates and calculate the Euclidean distance between the point cloud coordinates and the candidate point coordinates.
[0018] Further, in step (1), the DynamicVoxelVFE method is used for voxelization processing.
[0019] The 3D object detection system based on dynamic fusion of point cloud and voxel according to the present invention includes:
[0020] A three-dimensional information acquisition unit, configured to acquire three-dimensional point cloud information of a target to be measured, and perform voxelization processing on the three-dimensional point cloud information to obtain three-dimensional voxel information;
[0021] A three-dimensional feature extraction unit, configured to use a Pointnet++ network and a Voxelnet network respectively to extract features from the point cloud and voxel information to obtain point cloud features and voxel features;
[0022] A candidate point feature extraction unit, configured to screen the point cloud features according to the foreground point probability to obtain candidate point cloud features;
[0023] A detection box extraction unit, configured to obtain a detection box according to the candidate point cloud features;
[0024] A feature screening unit, configured to obtain candidate point voxel features within the detection box;
[0025] A feature fusion unit, configured to fuse the candidate point cloud features and the candidate point voxel features through a fusion module to obtain candidate point features; the fusion module includes a channel attention module, a cross attention module, and a self-attention module. The candidate point voxel features are input into the cross attention module after passing through the channel attention module, the candidate point cloud features are directly input into the cross attention module, the output of the cross attention module is connected to the self-attention module, and the output of the self-attention module is the candidate point feature;
[0026] A detection box update unit, configured to obtain an updated detection box according to the candidate point features, and re-execute the feature screening unit, the feature fusion unit, and the detection box update unit to obtain the final detection box and the classification information of the target to be measured within the detection box.
[0027] The electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the 3D object detection method based on dynamic fusion of point cloud and voxel is implemented.
[0028] The computer-readable storage medium according to the present invention stores a computer program, and when the computer program is executed by a processor, the 3D object detection method based on dynamic fusion of point cloud and voxel is implemented.
[0029] Advantages: Compared with the prior art, the advantages of the present invention are as follows: (1) The present invention utilizes point cloud features to introduce accurate position information, rich geometric information, and local features; utilizes voxel features to introduce rich semantic features and global features; these two types of features are fused through a fusion module based on the attention mechanism, enabling each point to accurately capture and learn the corresponding context information from the voxel features when generating new candidate point features, forming comprehensively enhanced candidate point features. (2) The present invention utilizes a dynamic iterative strategy, combines effective semantic information and global features of voxels, shields redundant voxel features that are unimportant for the target, and the multi-layer iterative method can continuously update the detection box to obtain more accurate prediction results. (3) The present method of the present invention adopts an optimization process of deep learning, can learn the essential features of the 3D environment during the feature fusion process, thereby improving the detection accuracy and generalization performance, and performs excellently in the categories of pedestrians, bicycles, cars, and trucks on the test data. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flowchart of the 3D object detection method of the present invention.
[0031] Figure 2 It is a flowchart of step S3 in the embodiment of the present invention.
[0032] Figure 3 It is a schematic diagram of the fusion module based on the attention mechanism in the embodiment of the present invention.
[0033] Figure 4 It is a detection effect diagram of the highway scene in the KITTI dataset in the embodiment of the present invention.
[0034] Figure 5 It is a detection effect diagram of the street scene in the KITTI dataset in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0035] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0036] The 3D object detection method based on dynamic fusion of point cloud and voxel of the present invention extracts unique features of each modality from the point cloud and voxel, dynamically fuses the two types of features through a fusion module based on the attention mechanism to form enhanced candidate point features, and continuously enhances the candidate point features through multiple iterations, and then performs decoding processing to achieve accurate 3D object detection. As Figure 1 shown, the method of the present invention includes the following steps.
[0037] Step S1, obtain the three-dimensional point cloud information, and voxelize the point cloud information to obtain three-dimensional voxel information, including the voxel center point coordinates and the initial voxel features. Specifically, use the DynamicVoxelVFE method for voxelization:
[0038] [C v ,F′ v =VFE(C p ),
[0039] where C p ∈R N×3 represents the input three-dimensional point cloud information, VFE represents the voxelization method, C v ∈R M×3 represents the coordinates of the voxel points, and F' v ∈R M×Ch represents the initial features of the voxel points. Ch represents the dimension of the features, and N and M represent the number of point cloud points and the number of voxel points respectively.
[0040] Step S2, perform feature extraction on the three-dimensional point cloud information and the three-dimensional voxel information respectively to obtain the three-dimensional point cloud features and voxel features. Specifically, use the Pointnet++ network and the Voxelnet network to perform feature extraction on the point cloud and voxel information respectively:
[0041] F p =PN(C p ),[C v ,F v =VN(C v ,F v );
[0042] where PN and VN represent the Pointnet++ network and the Voxelnet network respectively, F p ∈R N×Ch represents the point cloud features, and F v ∈R M×Ch represents the voxel features; among them, the input three-dimensional point cloud information is encoded by the point cloud encoder (Pointnet++), and the process includes point cloud sampling, point cloud grouping, grouped feature extraction, and feature aggregation to obtain the point cloud features; the voxel is encoded by the voxel encoder (Voxelnet), and the process includes submanifold convolution and sparse convolution to obtain the voxel features.
[0043] Step S3, screen the three-dimensional point cloud features to obtain the candidate point coordinates and the corresponding candidate point cloud features. Use the rough detection module for the candidate point cloud features to obtain the preliminary detection box information, screen the voxels to obtain the candidate point voxel features, as Figure 2 shown, specifically including the following steps.
[0044] Step S31: Perform binary classification on the obtained point cloud features, screen out a certain number of points with the highest foreground probabilities as candidate points, and extract the corresponding candidate point cloud features. The aim is to screen out as many foreground points as possible. Specifically, the classification network consists of two fully connected layers (FC), and then the K points with the highest scores are selected as candidate points, and the corresponding candidate point cloud features are extracted. The specific formula is as follows:
[0045] [IND fg =TOPK(CLS(F P ),K)];
[0046]
[0047] Among them, CLS(·) represents the classification network, and TOPK(·) represents extracting the results of the classification network, that is, the indices IND corresponding to the K values with the largest foreground probabilities fg 。 and represent the candidate point cloud coordinates and candidate point cloud features respectively.
[0048] Step S32: Use the rough detection module to obtain preliminary detection box information. The detection boxes correspond to the candidate points one by one. Specifically, the rough detection module consists of three fully connected layers. The first two fully connected layers are used for linear transformation of features, and the last fully connected layer is used for predicting detection box information. The specific formula is as follows,
[0049]
[0050] Among them, anchor coa ={x,y,z,h,w,l,a}∈R K×6 represents the detection box information, including the center point coordinates (x,y,z), the length, width, and height h,w,l of the box, and the rotation angle a of the box. CoaPre(·) represents the rough detection module.
[0051] Step S33: Screen out the voxel coordinates within the detection box and simultaneously screen out the voxel features to obtain the candidate point voxel features. Specifically, first convert the scale of the detection box to the voxel scale. The specific formula is as follows:
[0052]
[0053] Among them, pcr is the range of the point cloud coordinates, vsize is the size of the voxelized grid. After rotating the voxel coordinates by the angle a, the voxels within the box range are extracted,
[0054]
[0055] Among them, represents the voxel features of the candidate points, Indicates the voxel coordinates corresponding to the candidate point.
[0056] In step S34, when screening out the voxel coordinates, calculate the Euclidean distance between the voxel coordinates within the calculation frame and the candidate point coordinates. Specifically, convert the voxel coordinates into point cloud coordinates, and the formula is as follows:
[0057]
[0058] Then calculate the Euclidean distance with the candidate point coordinates:
[0059]
[0060] In step S4, input the candidate point cloud features and candidate point voxel features into the fusion module and the detection module and repeat n times to dynamically fuse the two features and iteratively update the detection frame information.
[0061] The present invention first determines the effective local voxel features to be used through the point cloud features, and uses the attention-based fusion module to fuse the three-dimensional point cloud features and three-dimensional voxel features and then complete the object detection. Effective feature fusion can achieve higher detection accuracy, reduce redundant voxel information, and reduce the computational complexity.
[0062] In step S41, input the candidate point cloud features and candidate point voxel features into the fusion module based on the attention mechanism to fully fuse the two features and obtain the candidate point features. Specifically, the fusion module includes: channel attention, cross attention, and self-attention calculation. Different from the common decoders, as Figure 3 shown, the present invention first calculates the channel attention of the candidate point voxel features to balance the feature information in the local positions, then uses D as the position encoding to fuse the two features through the cross-attention mechanism, and finally calculates the self-attention of the obtained output to update the features, and the candidate point coordinates are used as the position encoding to obtain the candidate point features, and the specific formula is as follows:
[0063]
[0064] Among them, MCA(·) represents multi-head cross-attention, MSA(·) represents multi-head self-attention, CA(·) represents channel attention, and fc(·) represents the fully connected layer for position encoding. Through the attention mechanism-based fusion module of the present invention, first, the local voxel features are balanced through channel attention calculation; then, through cross-attention calculation, each candidate point can accurately capture and learn the corresponding context information from the voxel features; finally, the candidate point features are enhanced through self-attention calculation. Therefore, better enhanced candidate point features can be obtained.
[0065] Step S42: Send the candidate point features into the detection module to obtain updated detection box information and target classification information. The detection module consists of four fully connected layers. The first two fully connected layers are used for linear transformation of the features, the third fully connected layer is used to predict the detection box information, and the last fully connected layer is used to predict the target classification information. Specifically, the following operations are performed on the fused candidate point features:
[0066] [anchor, cls] = Pre(F can );
[0067] where anchor = {x, y, z, h, w, l, a} ∈ R K×6 , cls ∈ R K×6 represents the classification information, and c represents the number of categories.
[0068] Step S43: Repeat steps S33, S34, S41, and S42 n times to dynamically fuse the voxel features and point cloud features, and continuously iterate the candidate point features and the output of the detection module. To ensure the iteration effect, when calculating the loss, set weights that increase but sum to 1 for the results of n iterations. Specifically, the weight in step S44 is W i = 2×i / (n 2 + n), and the output loss of the detection module is processed according to the following formula:
[0069]
[0070] where is the predicted loss of the detection box obtained in the i-th repetition, W i is the corresponding loss weight, and n is the total number of repetitions.
[0071] To verify the method of the present invention, the KITTI dataset is used for training and verification. The KITTI dataset is one of the most commonly used datasets for evaluating autonomous driving algorithms internationally, and it contains real data collected from scenarios such as urban areas, rural areas, and highways. As Figure 4 shown is the detection effect diagram of the method of the present invention in the highway scenario of the KITTI dataset, Figure 5 shown is the detection effect diagram of the method of the present invention in the street scenario of the KITTI dataset. It can be seen from the figure that the 3D object detection method proposed by the present invention can perform effective object detection in various scenarios.
[0072] In addition, this embodiment also uses three models to compare with the method described in the present invention. The results are shown in Table 1, where SECOND is the work of the Yan team in 2018 [Yan Yan, Yuxing Mao, and Bo Li. Second: Sparsely embedded convolutional detection. Sensors, 18(10)(2018)]; PointRCNN is the work of the Shi team in 2019 [Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. Pointrcnn: 3d object proposal generation and detection from point cloud. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(2019)]; 3DSSD is the work of the Yang team in 2020 [Zetong Yang, Yanan Sun, Shu Liu, and Jiaya Jia. 3dssd: Point-based 3d single stage object detector. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(2020)].
[0073] Table 1 3D detection accuracy and top view detection accuracy of four models for car targets on the KITTI dataset
[0074]
[0075] As can be seen from Table 1, the 3D object detection method based on the dynamic fusion of point cloud and voxel proposed in the present invention has improved detection accuracy compared with the point cloud-based methods PointRCNN and 3DSSD and the voxel-based method SECOND.
[0076] The 3D object detection system based on the dynamic fusion of point cloud and voxel described in the present invention includes:
[0077] A three-dimensional information acquisition unit for acquiring three-dimensional point cloud information of a target to be measured and performing voxelization processing on the three-dimensional point cloud information to obtain three-dimensional voxel information;
[0078] A three-dimensional feature extraction unit, which is used to extract point cloud features and voxel features from point cloud and voxel information respectively by using Pointnet++ network and Voxelnet network;
[0079] A candidate point feature extraction unit, which is used to screen the point cloud features according to the foreground point probability to obtain candidate point cloud features;
[0080] A detection box extraction unit, which is used to obtain a detection box according to the candidate point cloud features;
[0081] A feature screening unit, which is used to obtain candidate point voxel features within the detection box;
[0082] A feature fusion unit, which is used to fuse the candidate point cloud features and candidate point voxel features through a fusion module to obtain candidate point features; the fusion module includes a channel attention module, a cross-attention module and a self-attention module. The candidate point voxel features are input into the cross-attention module after passing through the channel attention module, the candidate point cloud features are directly input into the cross-attention module, the output of the cross-attention module is connected to the self-attention module, and the output of the self-attention module is the candidate point feature;
[0083] A detection box update unit, which is used to obtain an updated detection box according to the candidate point features, and re-execute the feature screening unit, the feature fusion unit and the detection box update unit to obtain the final detection box and the classification information of the target to be detected within the detection box.
[0084] The electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the 3D object detection method based on dynamic fusion of point cloud and voxel is implemented.
[0085] The computer-readable storage medium according to the present invention stores a computer program, and when the computer program is executed by a processor, the 3D object detection method based on dynamic fusion of point cloud and voxel is implemented.
[0086] The computer-readable storage medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory or any other medium that can store program code in the form of instructions or data structures and can be accessed by a computer.
[0087] The processor is used to execute the computer program stored in the memory to implement each step in the method described in the above embodiments.
Claims
1. A 3D object detection method based on dynamic fusion of point cloud and voxel, characterized in that, It includes the following steps: (1) Obtain the three-dimensional point cloud information of the target to be measured, and perform voxelization processing on the three-dimensional point cloud information to obtain three-dimensional voxel information; (2) Use the Pointnet++ network and the Voxelnet network respectively to extract features from the point cloud and voxel information to obtain point cloud features and voxel features; (3) Screen the point cloud features according to the foreground point probability to obtain candidate point cloud features; (4) Obtain a detection box according to the candidate point cloud features; (5) Obtain the candidate point voxel features within the detection box; (6) Use a fusion module to fuse the candidate point cloud features and the candidate point voxel features to obtain candidate point features; the fusion module includes a channel attention module, a cross-attention module, and a self-attention module. The candidate point voxel features are input into the cross-attention module after passing through the channel attention module, the candidate point cloud features are directly input into the cross-attention module, the output of the cross-attention module is connected to the self-attention module, and the output of the self-attention module is the candidate point feature; (7) Obtain an updated detection box according to the candidate point features, and repeat steps (5) to (7) to obtain the final detection box and the classification information of the target to be measured within the detection box.
2. The 3D object detection method based on dynamic fusion of point cloud and voxel according to claim 1, characterized in that, In step (4), a detection box is obtained through a rough detection module, and the rough detection module includes three sequentially connected fully connected layers.
3. The 3D object detection method based on dynamic fusion of point cloud and voxel according to claim 1, characterized in that In step (7), an updated detection box is obtained through a detection module, and the detection module includes four sequentially connected fully connected layers.
4. The 3D object detection method based on dynamic fusion of point cloud and voxel according to claim 1, wherein In step (7), during the process of repeating steps (5) to (7), the loss function is where W i = 2×i / (n 2 + n), is the predicted loss of the detection box obtained in the i-th repetition, and W i is the corresponding loss weight, and n is the total number of repetitions.
5. The 3D object detection method based on dynamic fusion of point cloud and voxel according to claim 1, wherein In step (6), the cross-attention module uses D as the position encoding, and the self-attention module uses the candidate point coordinates as the position encoding, where D is the Euclidean distance between the voxel coordinates and the candidate point coordinates within the detection box.
6. The 3D object detection method based on dynamic fusion of point cloud and voxel according to claim 5, wherein The method for calculating the Euclidean distance between the voxel coordinates and the candidate point coordinates is: convert the voxel coordinates to point cloud coordinates, and calculate the Euclidean distance between the point cloud coordinates and the candidate point coordinates.
7. The 3D object detection method based on dynamic fusion of point cloud and voxel according to claim 1, characterized in that In step (1), the DynamicVoxelVFE method is used for voxelization processing.
8. A 3D object detection system based on dynamic fusion of point cloud and voxel, characterized in that, It includes: A three-dimensional information acquisition unit for obtaining the three-dimensional point cloud information of the target to be measured and performing voxelization processing on the three-dimensional point cloud information to obtain three-dimensional voxel information; A three-dimensional feature extraction unit for using the Pointnet++ network and the Voxelnet network respectively to extract features from the point cloud and voxel information to obtain point cloud features and voxel features; A candidate point feature extraction unit for screening the point cloud features according to the foreground point probability to obtain candidate point cloud features; A detection box extraction unit for obtaining a detection box according to the candidate point cloud features; A feature screening unit for obtaining the candidate point voxel features within the detection box; A feature fusion unit for using a fusion module to fuse the candidate point cloud features and the candidate point voxel features to obtain candidate point features; the fusion module includes a channel attention module, a cross-attention module, and a self-attention module. The candidate point voxel features are input into the cross-attention module after passing through the channel attention module, the candidate point cloud features are directly input into the cross-attention module, the output of the cross-attention module is connected to the self-attention module, and the output of the self-attention module is the candidate point feature; The detection box update unit is used to obtain an updated detection box according to the candidate point features, and re-execute the feature screening unit, the feature fusion unit, and the detection box update unit to obtain the final detection box and the classification information of the target to be measured within the detection box.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements the 3D object detection method based on dynamic fusion of point cloud and voxel according to any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the 3D object detection method based on dynamic fusion of point cloud and voxel according to any one of claims 1-7.